mirror of
https://relay.ngit.dev/npub15qydau2hjma6ngxkl2cyar74wzyjshvl65za5k5rl69264ar2exs5cyejr/ngit-grasp.git
synced 2026-10-06 07:28:23 +00:00
Historic REQ+EOSE sync subscribed every byte-budgeted filter group of a batch in one loop and left them all awaiting EOSE concurrently, as did the REQ+EOSE fallback, negentropy ID fetches and missing-event retries. On strfry-family relays these REQs share maxSubsPerConnection with negentropy views and live subscriptions; nos.lol (budget 20) answered each gitnostr.com startup with a burst of 'ERROR: too many concurrent REQs' NOTICEs (6 in one second at the 2026-08-05 06:41 startup; 7,739 over the prior three days). A per-connection semaphore (5 permits, shared across clones) now gates every auto-close subscription inside subscribe_filters: the permit is registered against the subscription id on success and released when the connection's own event loop sees that subscription's EOSE or CLOSED frame - before forwarding the notification, so release never depends on downstream channel consumers. This keeps a plain blocking acquire deadlock-free even though the SyncManager actor both creates subscriptions and processes EOSE: bursts pipeline at five in flight, matching the precedent of the actor already stalling inline for negentropy batches. Permits are additionally freed on unsubscribe, disconnect, event-loop termination, and by a 30-second watchdog so a relay that never answers cannot starve later subscriptions. The negentropy semaphore stays separate (different lifetimes); live subscriptions are not gated. Budget: 4 NEG + 5 REQ + 2 margin leaves at least nine slots of the tightest observed budget (20) for live subscriptions. Scheduling is deliberately per-REQ rather than per-core-filter: a paginating filter chain holds no slot between pages, so queued groups interleave breadth-first. Relay-visible concurrency is identical either way and pagination chains have no durable identity across disconnects. Reproduction: a new proxy fixture mimics the strfry limit (rejecting REQs beyond 5 with the production NOTICE, exempting limit:0 live subscriptions, delaying EOSE so REQs provably overlap), and a scenario test syncs 2500 root events into persistent storage, restarts the relay - startup recomputes filters from the full index, the shape that bursts in production - and asserts zero rejections with overlap retained. Unfixed: 3 REQs rejected (opened 91, peak pinned at the limit). Fixed: zero rejections, peak <= 5, sync completes. The TestRelay fixture gains same-port restart support with persistent LMDB storage for this. Deliberately excluded: the unified budget ledger (NIP-11-aware B/M, per-query result caps), raising the 300-ID exact-ID chunks, and any configuration surface for the bound. Validated with the full test suite (nix develop -c cargo test).