New :relayBench module that boots relay implementations as real external
processes under equivalent setups (persistent storage, sig verification on,
no auth) and compares them on the same corpus:
- Ingest: receipt->queryable-by-REQ latency (publish on one connection,
hammer-poll REQ{ids} on another), OK-ack latency, and pipelined corpus
replay throughput over N connections.
- Queries: client-realistic filters derived from the corpus (home feed,
thread, notifications, profiles, hashtag, by-ids, ...), time-to-EOSE
percentiles plus aggregate events/sec under concurrency; result-set
counts are cross-checked between relays and mismatches flagged.
- NIP-77 negentropy sync between every relay pair: 80%/80% slices with 60%
overlap, reconcile timing/rounds/wire-bytes per server side, delta
transfer, steady-state identical-set reconcile, convergence verified.
- Storage footprint after full ingest.
Corpora (all cached as NDJSON + manifest with a sha256 id fingerprint so
results are comparable across runs and machines):
- synthetic (default): deterministic to the byte — seeded keys, fixed
timestamps, seed-derived BIP-340 aux nonces — with a realistic social
shape (zipf authors, threads, reactions, reposts, zap request/receipt
pairs, hashtags);
- the checked-in real dump (quartz test fixture, ~31k unique 2024 events);
- external dumps (NDJSON or JSON array, gzip sniffed by magic bytes),
e.g. the 2.1M contact-list archive, with --max-event-bytes/--max-tags
raising both the corpus filter and the strfry config together;
- live download from public relays.
Every source runs through the same preparation: dedup, drop unsigned/
ephemeral/kind-5, enforce relay ingest caps, parallel Schnorr verify,
chronological sort.
relayBench/run.sh is the one-command entry point: builds geode + harness,
resolves strfry (STRFRY_BIN, PATH, or source build into .cache), runs the
suite and renders an ANSI report with per-metric bars and winners, plus
report.md and results.json under relayBench/results/<timestamp>/.
The harness client disables Nagle (TCP_NODELAY): with the JDK default, a
REQ following the previous round's CLOSE stalls a full delayed-ACK
interval and every latency floors at ~44 ms against both geode and strfry
(verified: ~0.3 ms with it off).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NeoCvXnTxsKzqurkmjdC46
relayBench
Head-to-head benchmark for Nostr relay implementations. Boots each relay as a real external process on loopback under an equivalent setup — persistent storage, signature verification on, no auth, stock limits — replays the same event corpus into each, and renders a side-by-side report.
./relayBench/run.sh # geode vs strfry, 10k-event synthetic corpus
./relayBench/run.sh --quick # 2k-event smoke run
./relayBench/run.sh --real # replay the checked-in real-event dump (2024, ~30k events)
run.sh builds :geode:installDist and the harness, and resolves strfry from
$STRFRY_BIN, the PATH, or by building it from source into
relayBench/.cache/ (first run only). SKIP_STRFRY=1 / SKIP_GEODE=1 skip a
side.
What is measured
Ingest — receipt ➜ queryable. One connection publishes an event; from the
same instant a second connection hammer-polls REQ {"ids":[id]} until the
event comes back. This measures exactly "how long after the relay receives an
event can a REQ return it", which is not the same thing as the OK ack — the
report also shows OK latency and what fraction of events were already
queryable when their OK arrived.
Ingest — throughput. The corpus is replayed over N connections (default 4) with a bounded number of unacked EVENTs in flight, wall-clocked from first send to last OK. Accepted/rejected counts come from the OKs.
Queries. Filters modeled on what real clients send, derived from the corpus itself so they hit meaningful data: global feed, profile hydration, home feed (150 follows), hottest thread, notifications for the most-mentioned pubkey, hashtag feed, 100-id batch fetch, recent time window. Each runs warmup + measured rounds on one connection (time-to-first-event / time-to-EOSE percentiles) and then from 8 connections at once (aggregate events/second). All filters stay inside strfry's default limits, and the number of events each relay returns is cross-checked — a ⚠ in the report means the relays disagree about the result set, which invalidates the speed comparison for that row.
NIP-77 negentropy sync — every pair of relays. Both sides get an 80% slice
of the corpus (60% overlap); the harness plays the strfry sync role with one
side's dataset as its local set and measures: initial reconciliation against
each relay as server (time, NEG-MSG rounds, wire bytes), the delta transfer to
convergence, and the steady-state reconcile of identical sets. Convergence is
verified, so this doubles as an interop test.
Storage. On-disk footprint after full ingest (LMDB vs SQLite vs whatever).
Corpora
The corpus is the controlled variable: every relay sees the same events in
the same order, and reports carry a fingerprint (sha256 over event ids) so
two runs are comparable only when fingerprints match.
| source | flag | notes |
|---|---|---|
| synthetic (default) | --events N --seed S |
Deterministic to the byte: seeded keys, fixed timestamps, seed-derived BIP-340 nonces. Same spec ⇒ identical NDJSON on any machine. Zipf-popularity authors, threads, reactions, reposts, zap pairs, hashtags. |
| real dump | --real |
The quartz test fixture nostr_vitor_startup_data.json.gz — ~31k unique real events from 2024 with a rich kind mix (notes, chats, DMs, zaps, reports, communities). |
| contact lists | --corpus contact-lists.gz --limit 100000 --max-event-bytes 1048576 --max-tags 20000 |
2.1M real kind-3 contact lists (heavy events, ~1.3 kB avg, up to 100+ kB). Grab it with pip install gdown && gdown 1yyC93xY9sDsEsa351ZAMhtAXwBUh3LYT. Raising the size/tag caps reconfigures strfry to match, so both relays still accept the full stream. |
| any dump | --corpus FILE |
NDJSON or a single JSON array, gzipped or plain (sniffed by magic bytes). |
| fresh download | --download [urls] |
Pages recent events out of public relays (damus/nos.lol/primal by default). |
Every source goes through the same preparation: dedup by id, drop unsigned
events (NIP-17 rumors), kind-5 deletions and ephemerals (order-dependent or
unqueryable — they would make relays disagree for reasons unrelated to
performance), drop events over the size/tag caps, verify every Schnorr
signature in parallel, sort chronologically. The prepared corpus is cached in
relayBench/.corpus-cache/ as NDJSON next to a manifest.json with the
fingerprint and kind histogram — that file pair is a shareable, citable
benchmark artifact.
Why this corpus matters: there is no de-facto community benchmark corpus today. Existing relay benchmarks (rnostr's and privkeyio's nostr-bench, mattn's scripts) each synthesize their own events with unspecified distributions, so published numbers aren't reproducible corpus-controlled; the only shared dataset (Wellorder's early-1m) is frozen in January 2023. relayBench's synthetic spec ("seed 1, n=10000, v1" ⇒ byte-identical corpus) and manifest/fingerprint convention are designed so other relay authors can run the exact same workload and publish comparable numbers.
Adding another relay
No code needed if the relay can be launched from a command line:
./relayBench/run.sh --relay 'nostr-rs-relay=/usr/bin/nostr-rs-relay --db {dir} --port {port}'
{port} and {dir} are substituted at launch; the process must listen on
127.0.0.1:{port} with persistent storage under {dir}, verification on and
no auth. If the relay needs a config file, point the template at a small
wrapper script that writes one (see StrfryRelay in
relays/RelayUnderTest.kt for the pattern — adding a first-class subclass is
~20 lines).
Output
The terminal report shows each metric as name / bar / value rows with the winner starred, followed by a head-to-head summary. Every run also writes:
relayBench/results/<timestamp>/report.md— shareable MarkdownrelayBench/results/<timestamp>/results.json— raw numbers for tooling
Direct harness invocation
run.sh is a thin wrapper; the harness itself is
relayBench/build/install/relaybench/bin/relaybench — see --help for all
options (--samples, --publishers, --window, --query-rounds,
--query-conns, --no-sync, --out, --keep-data, …).