Files
ngit-grasp/docs/explanation/sync-scaling-constraints.md
T
DanConwayDev 04b93c9a1d fix(sync): separate transient terminal accounting
Production diagnostics at b55d2de showed watchdog-expired REQs were already absent from rust-nostr's active subscription map. An initial 4,096-message queue removed watchdogs from a 36,054-ID relay.ngit.dev burst but only moved the five-per-cycle signature first to ngit.danconwaydev.com and then git.shakespeare.diy. Empty expired responses disproved event volume as the complete explanation: terminal accounting depended both on a congestible EVENT lane and on permit registration after subscribe returned.

Give transient lifecycle messages an independent control lane by registering a second rust-nostr broadcast receiver before subscriptions can begin. It handles EOSE/CLOSED and connection terminal status without waiting for the bounded processor data queue. Pre-generate transient subscription IDs and register generation-scoped ownership before sending REQ, passing the same ID into rust-nostr and rolling it back on subscribe failure. Keep the original 1,000-message queue; correctness no longer depends on sizing it for traffic.

The terminal listener is the sole transient-release path during a connected session, while the processor listener retains ordered EVENT delivery, pagination signals, and live CLOSED restoration. EOSE still enqueues CLOSE before returning the slot; peer CLOSED and connection teardown are definitive terminal boundaries. Relay notifications are broadcast, so the control listener cannot steal messages from processing.

This deliberately leaves concurrency, filters, pagination, retries, watchdog duration, live lifecycle, and configuration unchanged. It assumes rust-nostr preserves broadcast terminal notifications and that a failed exact-target subscribe sends no REQ; the latter path rolls ownership back.

Validation: a deterministic LocalRelay test blocks a capacity-one EVENT lane behind 1,200 events and still releases all transient permits within three seconds; 25 immediate empty queries prove terminals cannot precede permit registration. All 648 library tests pass, and startup_historic_sync_stays_within_relay_req_concurrency_limit passes standalone with zero proxy REQ rejections. Production acceptance will be repeated on this exact tip.
2026-08-07 15:34:07 +00:00

30 KiB
Raw Blame History

Explanation: Sync Scaling Constraints and Budgets

Purpose: Explains the relay-imposed constraints that bound proactive sync, and justifies how we spend the three budgets they create — filter payload, subscriptions, and concurrency — as the watched item set grows. Audience: Contributors changing sync filter construction, subscription management, or negentropy scheduling; operators reasoning about scale limits.


The Problem

Proactive sync (GRASP-02) watches a growing set of items per relay: repository identifiers, repo references, and root event IDs. Every item must appear in filters twice — once in live subscriptions and once in historic sync (negentropy or REQ+EOSE). As the watched set grows, sync pressure on each relay grows along three axes:

  1. Filter payload — how many items fit in one filter / one message.
  2. Subscription count — how many concurrent subscriptions we hold.
  3. Request concurrency — how many sync operations run at once.

These axes are not independent: relays bound them with shared, mostly undiscoverable limits. This document records the limits we verified, the budget model derived from them, and the levers we use — in order — to scale.

Production motivation (2026-08-04, gitnostr.com): the bootstrap relay received 146 filters in one startup action (869 repos + 3632 root events, chunked at 100 items). Historic sync opened one negentropy round per filter with no bound, drawing 34 "too many concurrent NEG requests" rejections from nos.lol and 61 per-filter timeouts in two minutes.


Constraint Inventory

Verified 2026-08-04 against implementation sources and live NIP-11 documents. Re-verify before relying on exact numbers; defaults change.

strfry (most common large public relay implementation)

Limit Default Source
Tag values per filter (count) none — byte-capped src/filters.h:41
Tag value bytes per filter set 65535 src/filters.h:41
Tag fields per filter 3 (maxTagsPerFilter) golpe.yaml
Filters per REQ 200 (maxReqFilterSize); 3 if optional filterValidation enabled golpe.yaml
Subscriptions per connection 200 (maxSubsPerConnection) golpe.yaml
Concurrent negentropy shares maxSubsPerConnection — no separate knob src/apps/relay/RelayNegentropy.cpp
WebSocket message size 131072 (maxWebsocketPayloadSize) golpe.yaml

The key strfry finding: negentropy views and ordinary subscriptions draw from the same per-connection budget. "ERROR: too many concurrent NEG requests" is emitted when NEG views exceed maxSubsPerConnection.

Live NIP-11 documents (operators tighten defaults)

Relay max_limit max_subscriptions max_message_length
nos.lol (strfry 1.1.0) 500 20 131072
relay.primal.net (strfry 1.0.3-1-g60d35a6) 500 20 1000000
nostr.wine (operator software 0.3.3) 1000 50 524288
relay.damus.io (strfry 1.1.0-1-g691a533f11eb) 500 200 1000000
relay.ditto.pub (Ditto Relay 0.1.0) 1000 20 4000000
relay.nostr.band unavailable — HTTPS timed out twice unavailable unavailable

Live documents fetched 2026-08-06 with Accept: application/nostr+json. The five reachable relays advertise their result cap as limitation.max_limit.

Discoverability gap (NIP-11)

NIP-11 limitation has no field for tag values per filter or filter count per subscription. It does define max_limit (clamp applied to a filter's explicit limit), and default_limit (maximum returned events when limit is omitted — the field pagination actually needs), in addition to max_subscriptions and max_message_length, but implementations and operators advertise these unevenly: default_limit in particular is rarely present (neither nos.lol nor relay.ditto.pub advertises it, checked live 2026-08-06). max_limit also cannot express whether the allowance is per filter or aggregate across a multi-filter REQ. Consequently neither filter sizing nor the pagination model can be negotiated reliably; both need conservative defaults, observation, and reactive fallback.

Our own embedded relay (nostr-sdk LocalRelay, 0.45.0)

  • max_reqs = 500, enforced for REQ only (src/nostr/builder.rs).
  • Negentropy: no concurrency limit at all (upstream TODO), 60000-byte frame limit per NEG message.
  • 20 filters per REQ by default; no limit on tag values per filter.
  • Query result limits (verified 2026-08-06 against the published nostr-sdk-0.45.0 crate source, src/local_relay/local/inner.rs): enforced per filter, not per REQ. A filter without a limit is given default_filter_limit (500); the effective limit is then clamped to min(limit, max_filter_limit, max_query_results) (defaults: no max_filter_limit, max_query_results = 500). Each filter is queried independently and the merged, deduplicated results are sent without aggregate truncation — a source comment suggests the merged set is also capped, but the implementation does not do this.
  • Behaviour change from 0.45.0-alpha.8 and earlier: a filter without an explicit limit previously returned every match; stable 0.45.0 returns at most the newest 500 per filter, so large backlogs arrive via pagination instead of one unbounded response.

khatru (used by the ngit-relay reference implementation) and nostr-rs-relay similarly enforce no filter-size limits by default.

Per-query result limits and the pagination model

NIP-01 defines limit per filter, for the initial query only, and lets relays return fewer events than requested. It neither guarantees that each filter in a multi-filter REQ receives an independent result allowance nor forbids an aggregate cap across the whole REQ. Which model a relay implements is an empirical question, and it decides whether grouped REQ+EOSE pagination (per-filter until cursors inside one grouped subscription) is safe.

Audit result: per-filter everywhere; no aggregate caps

Source audit, 2026-08-06, of nine implementations at release tags — nostr-sdk LocalRelay 0.45.0, strfry 1.1.1, nostr-rs-relay 0.10.0, khatru v0.19.1, relayer v2.2.14, nostream v3.0.0, rnostr v0.4.9, chorus v2.0.2, haven v1.2.2. The full per-implementation table with file:line citations is preserved in this file's history (commit e889ea5).

  • Every implementation applies result limits per filter, and none caps the merged results of a multi-filter REQ, so grouped pagination's per-filter cursor model is sound. until never changes the cap.
  • Defaults for a filter sent without limit: 500 (nostr-sdk, strfry, nostream), 375/250 (haven LMDB/Badger), 300 (rnostr), 1000 or unbounded (nostr-rs-relay by backend), unbounded at the framework layer (khatru, relayer, chorus — stores decide). "Unbounded" means the semantic query limit; timeouts, rate limits, and finite databases still shorten responses, as NIP-01 permits.
  • strfry and rnostr technically accept arbitrarily low operator caps, making any fixed PAGINATION_THRESHOLD formally unsafe at that configuration boundary, but no live deployment anywhere near the threshold was found; live strfry relays advertise max_limit 500, nostr.wine 1000.
  • NIP-11 advertisement of the cap is uneven: strfry, nostream, and rnostr publish max_limit; nostr-rs-relay, khatru-family, and chorus do not.

Implementation limit matrix

These are implementation defaults, not claims about every deployment. ★ means that implementation emits the value in the corresponding standard NIP-11 limitation field; operators can still override or omit advertised values. — means no native limit was found at that layer, not that a reverse proxy, host, storage backend, or embedding application cannot impose one.

Implementation Results / filter Filters / subscription Subscriptions / connection Connections / IP Evidence
nostr-sdk LocalRelay 0.45.0 500 20 500 — (128 global) published crate local_relay/builder.rs:19-29,41-46,338-359
strfry 1.1.1 500 ★ 200 200 ★ — strfry.conf:95-117, RelayWebsocket.cpp:89-97
nostr-rs-relay 0.10.0 SQLite: unbounded; PostgreSQL: 1000 — — — sqlite.rs:1149-1155, postgres.rs:891-900
khatru v0.19.1 store-defined — — — per-filter dispatch in handlers.go:289-324
relayer v2.2.14 store-defined / framework unbounded — — — handlers.go:182-255
nostream v3.0.0 500 default; requested maximum 5000 ★ 10 ★ 10 ★ — base.ts:87, default-settings.yaml:215-222, root-request-handler.ts:87-104
rnostr v0.4.9 300 ★ 10 ★ 20 ★ — setting.rs:123-156,340-352
chorus v2.0.2 unbounded — 128 ★ 5 config.rs:30-47,52-93, nip11.rs:143-153
haven v1.2.2 LMDB: 375; Badger: 250 — — — backend construction in init.go:61-78; eventstore v0.17.5 lmdb/query.go:26-43
Ditto Relay 0.1.0 (cf34437, no release tag) 100 default; requested maximum 1000 ★ 100 ★ 20 ★ — relay.ts:188-206,1256-1307, live relay.ditto.pub NIP-11

The result column distinguishes a filter's implicit default from the largest explicit request where they differ. This matters for pagination: nostream and Ditto normally return 500 and 100 respectively when limit is omitted even though they advertise the larger accepted max_limit.

Admission and rate limits (condensed)

Native rate limiting varies wildly and is invisible to clients. Our own embedded relay enforces per-connection per-minute quotas (120 queries, 6,000 WebSocket messages, 60 event writes); nostream ships per-IP connection-attempt and kind-specific event quotas with EWMA decay; khatru and haven offer discrete leaky counters that drain over minutes; nostr-rs-relay, relayer, and rnostr have token-bucket limiters that are disabled by default; chorus budgets raw bytes per connection (16 MiB burst, 1 MiB/s refill), caps five simultaneous connections per IP, and bans immediate reconnects; strfry and Ditto have no native limiter at all, deferring to deployment infrastructure. The full survey with citations is preserved in this file's history (commit 9723ff4).

NIP-11 describes hard relay limitations, not rate-limit algorithms. The standard fields relevant here are max_limit and max_subscriptions; it has no standard fields for simultaneous connections per IP, connection-attempt rate, message/event/query rate, burst size, window, decay model, or retry-after time. Even an advertised max_limit does not say whether it applies independently to each filter or to the merged REQ, which is why the source audit above remains necessary. Relay-specific extensions can add fields, but clients cannot assume common names or semantics.

The client encodes this model in per-connection RelayPaginationSession state (src/sync/mod.rs). After EOSE it learns the largest raw page seen from that relay and computes max(90, floor(0.9 × estimated_cap)), where estimated_cap also includes an advertised NIP-11 default_limit while that hint remains trusted. A filter meeting the adaptive threshold is fetched again with until set to its oldest raw created_at. Consequences:

  • Every raw delivery matching a tracked filter counts before deduplication or write-policy processing. Purgatory-routed, rejected, and repeated events therefore consume both the relay's allowance and our page count, and the until cursor is derived from that same raw stream.
  • Ditto's 100-event omitted-limit default is now above the adaptive floor and is learned from its first page even though it advertises only max_limit: 1000. max_limit never raises the threshold because it describes explicit limits, not the omitted-limit filters sent here.
  • A relay capping a filter below 90 can still silently truncate history. No such deployment was found in the audit. No audited implementation enforces an aggregate cap across filters in one REQ.
  • Larger learned pages raise the threshold and avoid redundant requests. The 0.9 slack can still produce one final verification-shaped page when a result count falls near the learned cap; this is the deliberate cost of tolerating relay-side page shrinkage.
  • Implemented design (accepted 2026-08-06): keep omitting limit — an explicit limit would cap the relays that serve unbounded pages — count raw deliveries, and adapt the threshold per relay:
    1. Count raw delivered events (implemented). Every delivered event that matches a tracked filter is counted before deduplication and write policy, and the cursor uses the same stream. Purgatory-routed, rejected, and repeated events can no longer consume relay allowance invisibly.
    2. Adaptive per-relay threshold (implemented): estimated_cap = max(largest observed page, advertised default_limit if present); threshold = max(90, floor(0.9 × estimated_cap)). Observed pages are ground truth (always ≤ the true cap, so never unsafe, and converging upward to eliminate redundant pages); the 0.9 slack absorbs relay-side shrinkage such as expired-event skipping; the floor of 90 stays below Ditto's 100, the smallest default found. Learned state is per connection session and NIP-11 is refetched on reconnect, so an operator lowering their cap cannot strand a stale threshold.
    3. NIP-11 fields (implemented): default_limit ("maximum returned events if you send a filter without a limit") is the standard field for exactly this and is used as a hint when advertised — though rarely: neither nos.lol nor relay.ditto.pub advertises it (checked live 2026-08-06). Being self-reported, a wrong-high value is unsafe, so the first page that the hint would declare exhausted triggers one verification page; if it yields new events the hint is discarded in favour of learned-only. max_limit must never raise the threshold while requests omit limit: it bounds accepted explicit requests, not the omitted-limit page size (Ditto: 1000 advertised vs 100 served; nostream: 5000 vs 500).

Working floors

Derived from the tightest commonly observed values; all sizing below assumes:

  • Subscription budget B = 20 per connection (nos.lol, relay.primal.net, Ditto Relay default), shared between live REQs, NEG rounds, and fallback REQs. Caveat found by the 2026-08-06 limit matrix: nostream defaults to 10 subscriptions per connection and 10 filters per REQ (the former is standard NIP-11; nostream emits the latter as a relay-specific field), below this floor — the fixed 4 NEG + 5 REQ + 2 margin pattern alone would overdraw a default nostream before any live subscriptions. Honouring advertised max_subscriptions is therefore required ledger work, not just an optimisation.
  • Message budget M = 128 KB (nos.lol); we target ≤ 96 KB of filter payload per message, a 1.3× margin for the envelope.
  • Per-filter value budget 32 KB (half of strfry's 65535-byte set cap; a full-chunk NEG-OPEN is ~33 KB, ~1.8× under the 60 KB negentropy frame limit our own embedded relay enforces), chosen so three full chunks fit one 96 KB REQ message — see lever 2.
  • A serialized 64-char hex ID costs ~67 bytes ("…",), so: ~489 hex IDs per filter, ~1460 hex IDs per message. Variable-length values (#d identifiers, repo references) must be budgeted by bytes, not count.

Our Approach: A Per-Connection Budget Ledger

Each relay connection owns one implemented budget ledger of B subscription slots. Three consumers share it, in priority order:

  1. Live subscriptions (persistent, limit: 0) — the product; sized first.
  2. Reserved margin (2 slots) — control-plane safety capacity kept beyond the live set (which includes Layer-1) for ad-hoc operations and recovery.
  3. Historic sync and dependency recovery (transient) — negentropy rounds, REQ+EOSE pages/fallbacks/retries, and exact-ID purgatory polls draw from the remainder. NEG retains its four-round class cap and transient REQ its five-request class cap, but neither can exceed the shared residual.

NIP-11 max_subscriptions sets B for each new connection session; when it is absent B falls back to 20. Advertised values below that floor are honoured (notably nostream's default 10). Two slots remain reserved. Live filter groups are packed first and admitted atomically: if the complete live set cannot fit, the existing subscriptions are consolidated into the byte- and filter-count bounded REQ groups first. If the consolidated live set still cannot fit, partial coverage is not opened, historic work is deferred, and a warning surfaces the condition. Reconnect and consolidation build Layer 1, 2, and 3 coverage together and reserve the whole grouped set before opening it. A failure while opening group N sends CLOSE for every earlier group and returns their slots, so callers never inherit hidden partial coverage. Replacement also remembers the exact previous grouping and restores it if the new complete set fails at runtime. Multi-connection sharding remains the later lever.

A transient slot is released only after EOSE has caused CLOSE to be enqueued, after relay CLOSED, or after connection teardown. Because NIP-01 provides no CLOSE acknowledgement, the 120-second recovery path sends CLOSE for only the timed-out subscription and releases only that subscription's generation-scoped slot after the SDK accepts the message; valid production startup pages exceeded 30 seconds, while two minutes remains a bounded escape from a stuck subscription. If CLOSE cannot be enqueued, the slot remains held until ordinary connection teardown so local accounting cannot run ahead of the relay. Exact-ID purgatory polling uses the same transient class bound and shared ledger as historic pagination. Transient subscription IDs and their permits are registered locally before the REQ is sent; this ordering is required because an empty or cached response can deliver EOSE/CLOSED before the SDK subscribe call returns. Subscribe failure rolls that pre-registration back. Unexpected CLOSED for a persistent live subscription is reported to the manager, which recomputes and transactionally reopens complete live coverage. Each reconnect closes the retired ledger and creates a new generation; queued or late borrowers therefore fail before sending on the new SDK session and cannot inflate or bypass its capacity.

The per-relay event processor retains its 1,000-message bounded data queue. Permit release does not depend on that queue draining: a separate listener on rust-nostr's broadcast relay notifications consumes only EOSE/CLOSED terminals and closes/releases transient ownership. The processor-facing listener still delivers ordered EVENT and lifecycle work to the sync actor, but sustained page traffic cannot hide terminal accounting behind EVENT backpressure. The data queue remains finite as a memory-safety boundary for non-conforming peers.

The levers, in the order we reach for them:

Lever 1: Maximise items per filter (byte-budgeted chunking)

Replace the fixed 100-items-per-chunk rule with byte budgets: a filter chunk is full when it reaches 32 KB of serialized tag values (~489 hex IDs), and a message is full at ~96 KB. The 100-item chunk was a guess made when we believed relays capped item counts; the verified constraints are byte caps (strfry 65535 per filter set, message size per NIP-11), so counting items wastes ~4.9× capacity for hex IDs while being unsafe for unbounded-length #d identifiers.

Chunk and REQ budgets are maximised together because they bound different costs: for a total serialized payload T, persistent subscription count scales with how full each REQ is packed (T / 96 KB), while negentropy round count scales with chunk size (T / 32 KB — one round per filter). Bigger chunks do not inflate subscription counts as long as full chunks still pack three to a REQ, so 32 KB chunks in 96 KB REQs minimise both at once — and three full chunks per REQ matches strfry's strict filterValidation limit of three filters per REQ. What eventually bounds filter size is none of the byte caps but per-query result limits (e.g. damus "blocked: too many query results" against filters that match too much at once); accounting for those belongs to the budget-ledger work.

Because the limits are not discoverable (NIP-11 gap), the budget is static and conservative rather than probed; the existing transient-failure cooldown and REQ+EOSE fallback absorb the rare relay with tighter limits.

What this lever cannot do: collapse the three tag-variant filters. NIP-01 ANDs distinct tag conditions within one filter, so a/A/q (and e/E/q) coverage requires three filters per chunk regardless of size. strfry's maxTagsPerFilter = 3 counts tag fields per filter; our filters use one tag field each, so this is not a binding constraint.

Lever 2: Pack filters per REQ — coupled to lever 1 by message size

Live subscriptions send all their filters in one REQ message, so the message budget M caps items per subscription (~1460 hex IDs at the 96 KB payload budget) no matter how items are split into filters. Packing more filters into fewer REQs is what actually shrinks the persistent subscription count, so the rule is a byte budget per REQ message, with filter count as a secondary bound (strfry accepts 200 filters per REQ, but its optional strict filterValidation mode accepts only 3 — matched by three full 32 KB chunks per 96 KB REQ).

Lever 3: Bound and schedule concurrency (coordination with live sync)

Negentropy reconciles one filter per round, and each in-flight round consumes a subscription slot from the same budget as live subscriptions (strfry). So concurrency is not a free scaling axis; it is the residual of the ledger:

  • Per-connection NEG concurrency is at most min(4, B − L − margin − other transients); when no residual remains, historic work waits rather than overdrawing.
  • Rounds queue behind a per-connection semaphore; each completion releases the next. No timed batches or sleeps — throughput degrades smoothly instead of bursting into rejections.
  • Transient REQ+EOSE subscriptions — historic sync groups, fallback filters, exact-ID fetches, retries, and pagination pages — retain a five-request class cap inside the shared ledger: a slot is acquired when the auto-close REQ is sent and released when its EOSE or CLOSED arrives (with a 120 s watchdog that sends CLOSE for only the unresponsive subscription before releasing its slot). Each held permit retains one of a fixed set of request classes (historic page, pagination page, hint verification, negentropy hydration, retry, or semantic fallback); watchdog logs and metrics expose that class without using relay URLs or subscription IDs as metric labels. Live subscriptions are ledgered first, so NEG, transient REQ, and purgatory exact-ID polling share only the remaining capacity.
  • Permit acquisition checks relay health first: while a rate-limit or transient-failure cooldown is active, queued rounds take the REQ+EOSE fallback path (which is itself budget-accounted) instead of firing into a relay that just complained.
  • The reactive machinery (escalating cooldown, NOTICE-based rate-limit pause, per-batch fallback) remains the backstop for relays whose limits are below our floors — prevention first, reaction second.

Lever 4: Multiple connections per relay (last resort)

strfry-family limits are per connection, so a second connection doubles both the subscription budget and the NEG budget at that relay. This is the escalation path when a relay's watched set can no longer fit: needed_live_slots + margin + 1 > B even after levers 1–2.

Costs and risks, which is why it is last:

  • Per-IP connection caps exist but are not advertised anywhere; exceeding them looks like abuse and risks bans. The tightest native cap found by the 2026-08-06 limit matrix is chorus at five simultaneous connections per IP (with a reconnect ban of at least one second), so the ≤ 4 bound now has source evidence rather than being pure caution. Bound connections per relay (≤ 4) and scale in with hysteresis.
  • Each connection re-authenticates (NIP-42) and carries its own health state, file descriptor, and TLS/session overhead.
  • Filter-to-connection assignment must be deterministic (stable sharding of the watched set) so reconnects and consolidation do not reshuffle subscriptions across the pool.

Where the pressure actually lands

Budget pressure is worst where the watched set is largest — today that is our own bootstrap relay (869 repos / 3632 roots ≈ 1 MB of serialized tag values, i.e. ~11 messages minimum even optimally packed). Public relays typically carry small per-relay target sets but tight budgets (B = 20). Two consequences:

  • For infrastructure we control (bootstrap, self-relay), raise and advertise server-side limits rather than spending client-side levers.
  • For public relays, levers 1–3 keep us comfortably inside B = 20 at current scale; lever 4 exists for the point where a single public relay's target set outgrows ~(B − margin) × 1460 hex-ID-equivalents (~26 k items).

Serving-Side Obligations

We are also a relay, and peer GRASP instances run this same sync against us. The rust-nostr 0.45 embedded relay now enforces 10 active negentropy sessions, 20 filters per REQ, and the other bounds recorded in the relay-limits reference. ngit-grasp explicitly selects those defaults and advertises the standard, discoverable subset. The remaining serving-side gap is a bound on tag-value/filter payload size, for which NIP-11 has no standard field. At scale we must:

  1. Retain the existing NEG, REQ, subscription-memory, message, event, and rate bounds, and design a filter-payload bound if production evidence requires it.
  2. Keep NIP-11 limitation aligned with every enforced standard field so well-behaved peers can budget against us.

Trade-offs

Gained: deterministic behaviour against unadvertised limits; startup bursts bounded by design rather than absorbed by cooldowns; a single model (the ledger) that live sync, historic sync, and fallback all account against; a defined escalation path to multi-connection scale.

Given up: peak theoretical throughput on permissive relays (a damus-class relay with 200 subscription slots is used as if it had 20 when limitation is absent — we only relax budgets when NIP-11 advertises headroom); some implementation complexity (byte-budgeted chunking, permit-gated scheduling, eventual sharding).


Alternatives Considered

Adaptive probing (start big, shrink on rejection)

Pros: discovers each relay's true limits; no static guesswork. Cons: rejection signals are non-standard free-text NOTICEs; every startup pays a rejection burst per relay; failure attribution is ambiguous (payload size vs. subscription count vs. rate limit), so the probe can learn the wrong lesson. Why not: we tried the reactive-only posture implicitly and it produced the 2026-08-04 incident; static floors with reactive backstop are deterministic and testable.

NIP-11-driven budgets

Pros: honest relays advertise max_subscriptions and max_message_length; budgets could be exact. Why partial: max_subscriptions and default_limit are consumed when present, but message-size negotiation is not yet implemented, filter count has no current standard NIP-11 field, and many relays omit limitation entirely — so floors remain necessary. Proposing a NIP-11 extension for filter-size limits is worthwhile upstream work.

Timed batching with pause-on-rate-limit

Pros: simple to picture. Cons: reactive by construction (eats one rejection burst per relay per startup), needs heuristic NOTICE parsing as its primary control loop, and fixed pauses waste time on fast relays while still bursting slow ones. Why not: the semaphore ledger achieves the same containment continuously, with the heuristics demoted to backstop.


Rollout Mapping

Lever Status
3 — bounded NEG concurrency Stabilisation cycle 3 (in flight)
3 — bounded transient REQ+EOSE concurrency Landed with cycle 3 (same PR)
1 + 2 — byte-budgeted chunking and REQ packing Landed with cycle 3 (same PR)
Ledger unification (live + historic + fallback + purgatory polling against one NIP-11-aware subscription budget) Implemented 2026-08-06; message-budget negotiation remains follow-up
4 — multi-connection sharding Deferred until a relay's target set approaches the single-connection ceiling
Serving-side limits + NIP-11 advertisement Implemented 2026-08-07 for rust-nostr's enforceable limits and the standard discoverable subset; filter-payload bounding remains follow-up