Files
amethyst/relayBench
Claude 5f3a790d56 feat: add relayBench — head-to-head relay benchmark (geode vs strfry vs any)
New :relayBench module that boots relay implementations as real external
processes under equivalent setups (persistent storage, sig verification on,
no auth) and compares them on the same corpus:

- Ingest: receipt->queryable-by-REQ latency (publish on one connection,
  hammer-poll REQ{ids} on another), OK-ack latency, and pipelined corpus
  replay throughput over N connections.
- Queries: client-realistic filters derived from the corpus (home feed,
  thread, notifications, profiles, hashtag, by-ids, ...), time-to-EOSE
  percentiles plus aggregate events/sec under concurrency; result-set
  counts are cross-checked between relays and mismatches flagged.
- NIP-77 negentropy sync between every relay pair: 80%/80% slices with 60%
  overlap, reconcile timing/rounds/wire-bytes per server side, delta
  transfer, steady-state identical-set reconcile, convergence verified.
- Storage footprint after full ingest.

Corpora (all cached as NDJSON + manifest with a sha256 id fingerprint so
results are comparable across runs and machines):
- synthetic (default): deterministic to the byte — seeded keys, fixed
  timestamps, seed-derived BIP-340 aux nonces — with a realistic social
  shape (zipf authors, threads, reactions, reposts, zap request/receipt
  pairs, hashtags);
- the checked-in real dump (quartz test fixture, ~31k unique 2024 events);
- external dumps (NDJSON or JSON array, gzip sniffed by magic bytes),
  e.g. the 2.1M contact-list archive, with --max-event-bytes/--max-tags
  raising both the corpus filter and the strfry config together;
- live download from public relays.

Every source runs through the same preparation: dedup, drop unsigned/
ephemeral/kind-5, enforce relay ingest caps, parallel Schnorr verify,
chronological sort.

relayBench/run.sh is the one-command entry point: builds geode + harness,
resolves strfry (STRFRY_BIN, PATH, or source build into .cache), runs the
suite and renders an ANSI report with per-metric bars and winners, plus
report.md and results.json under relayBench/results/<timestamp>/.

The harness client disables Nagle (TCP_NODELAY): with the JDK default, a
REQ following the previous round's CLOSE stalls a full delayed-ACK
interval and every latency floors at ~44 ms against both geode and strfry
(verified: ~0.3 ms with it off).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NeoCvXnTxsKzqurkmjdC46
2026-07-03 18:37:16 +00:00
..

relayBench

Head-to-head benchmark for Nostr relay implementations. Boots each relay as a real external process on loopback under an equivalent setup — persistent storage, signature verification on, no auth, stock limits — replays the same event corpus into each, and renders a side-by-side report.

./relayBench/run.sh                 # geode vs strfry, 10k-event synthetic corpus
./relayBench/run.sh --quick         # 2k-event smoke run
./relayBench/run.sh --real          # replay the checked-in real-event dump (2024, ~30k events)

run.sh builds :geode:installDist and the harness, and resolves strfry from $STRFRY_BIN, the PATH, or by building it from source into relayBench/.cache/ (first run only). SKIP_STRFRY=1 / SKIP_GEODE=1 skip a side.

What is measured

Ingest — receipt ➜ queryable. One connection publishes an event; from the same instant a second connection hammer-polls REQ {"ids":[id]} until the event comes back. This measures exactly "how long after the relay receives an event can a REQ return it", which is not the same thing as the OK ack — the report also shows OK latency and what fraction of events were already queryable when their OK arrived.

Ingest — throughput. The corpus is replayed over N connections (default 4) with a bounded number of unacked EVENTs in flight, wall-clocked from first send to last OK. Accepted/rejected counts come from the OKs.

Queries. Filters modeled on what real clients send, derived from the corpus itself so they hit meaningful data: global feed, profile hydration, home feed (150 follows), hottest thread, notifications for the most-mentioned pubkey, hashtag feed, 100-id batch fetch, recent time window. Each runs warmup + measured rounds on one connection (time-to-first-event / time-to-EOSE percentiles) and then from 8 connections at once (aggregate events/second). All filters stay inside strfry's default limits, and the number of events each relay returns is cross-checked — a ⚠ in the report means the relays disagree about the result set, which invalidates the speed comparison for that row.

NIP-77 negentropy sync — every pair of relays. Both sides get an 80% slice of the corpus (60% overlap); the harness plays the strfry sync role with one side's dataset as its local set and measures: initial reconciliation against each relay as server (time, NEG-MSG rounds, wire bytes), the delta transfer to convergence, and the steady-state reconcile of identical sets. Convergence is verified, so this doubles as an interop test.

Storage. On-disk footprint after full ingest (LMDB vs SQLite vs whatever).

Corpora

The corpus is the controlled variable: every relay sees the same events in the same order, and reports carry a fingerprint (sha256 over event ids) so two runs are comparable only when fingerprints match.

source flag notes
synthetic (default) --events N --seed S Deterministic to the byte: seeded keys, fixed timestamps, seed-derived BIP-340 nonces. Same spec ⇒ identical NDJSON on any machine. Zipf-popularity authors, threads, reactions, reposts, zap pairs, hashtags.
real dump --real The quartz test fixture nostr_vitor_startup_data.json.gz — ~31k unique real events from 2024 with a rich kind mix (notes, chats, DMs, zaps, reports, communities).
contact lists --corpus contact-lists.gz --limit 100000 --max-event-bytes 1048576 --max-tags 20000 2.1M real kind-3 contact lists (heavy events, ~1.3 kB avg, up to 100+ kB). Grab it with pip install gdown && gdown 1yyC93xY9sDsEsa351ZAMhtAXwBUh3LYT. Raising the size/tag caps reconfigures strfry to match, so both relays still accept the full stream.
any dump --corpus FILE NDJSON or a single JSON array, gzipped or plain (sniffed by magic bytes).
fresh download --download [urls] Pages recent events out of public relays (damus/nos.lol/primal by default).

Every source goes through the same preparation: dedup by id, drop unsigned events (NIP-17 rumors), kind-5 deletions and ephemerals (order-dependent or unqueryable — they would make relays disagree for reasons unrelated to performance), drop events over the size/tag caps, verify every Schnorr signature in parallel, sort chronologically. The prepared corpus is cached in relayBench/.corpus-cache/ as NDJSON next to a manifest.json with the fingerprint and kind histogram — that file pair is a shareable, citable benchmark artifact.

Why this corpus matters: there is no de-facto community benchmark corpus today. Existing relay benchmarks (rnostr's and privkeyio's nostr-bench, mattn's scripts) each synthesize their own events with unspecified distributions, so published numbers aren't reproducible corpus-controlled; the only shared dataset (Wellorder's early-1m) is frozen in January 2023. relayBench's synthetic spec ("seed 1, n=10000, v1" ⇒ byte-identical corpus) and manifest/fingerprint convention are designed so other relay authors can run the exact same workload and publish comparable numbers.

Adding another relay

No code needed if the relay can be launched from a command line:

./relayBench/run.sh --relay 'nostr-rs-relay=/usr/bin/nostr-rs-relay --db {dir} --port {port}'

{port} and {dir} are substituted at launch; the process must listen on 127.0.0.1:{port} with persistent storage under {dir}, verification on and no auth. If the relay needs a config file, point the template at a small wrapper script that writes one (see StrfryRelay in relays/RelayUnderTest.kt for the pattern — adding a first-class subclass is ~20 lines).

Output

The terminal report shows each metric as name / bar / value rows with the winner starred, followed by a head-to-head summary. Every run also writes:

  • relayBench/results/<timestamp>/report.md — shareable Markdown
  • relayBench/results/<timestamp>/results.json — raw numbers for tooling

Direct harness invocation

run.sh is a thin wrapper; the harness itself is relayBench/build/install/relaybench/bin/relaybench — see --help for all options (--samples, --publishers, --window, --query-rounds, --query-conns, --no-sync, --out, --keep-data, …).