mirror of
https://relay.ngit.dev/npub15qydau2hjma6ngxkl2cyar74wzyjshvl65za5k5rl69264ar2exs5cyejr/ngit-grasp.git
synced 2026-10-05 15:08:24 +00:00
docs(sync): record relay constraints and scaling budget design
Document the verified relay-imposed limits that bound proactive sync (strfry per-filter byte caps and shared subscription budget, live NIP-11 limitation documents, the discoverability gap, and our embedded relay's missing negentropy limits) and derive the budget model for scaling: a per-connection subscription ledger shared by live, historic and fallback work, byte-budgeted filter chunking, bounded negentropy concurrency, and multi-connection sharding as a last resort. Motivated by the 2026-08-04 gitnostr.com startup incident where 146 unbounded concurrent negentropy rounds drew rate-limit rejections; the doc provides the constraint justification for the planned bounded-NEG fix and the follow-up chunk-size change. Documentation only; no behaviour changes.
This commit is contained in:
@@ -108,6 +108,20 @@ Explanation documentation helps you **understand concepts** and design decisions
|
||||
|
||||
---
|
||||
|
||||
### [Sync Scaling Constraints and Budgets](sync-scaling-constraints.md)
|
||||
**Relay-imposed limits and how sync spends them at scale**
|
||||
|
||||
**Topics:**
|
||||
- Verified relay limits (strfry, live NIP-11, our embedded relay)
|
||||
- Per-connection budget ledger (live vs historic vs fallback)
|
||||
- Byte-budgeted filter chunking and REQ packing
|
||||
- Bounded negentropy concurrency
|
||||
- Multi-connection escalation and serving-side obligations
|
||||
|
||||
**Read when:** You're changing filter construction, subscription management, or sync concurrency, and need the constraint justification
|
||||
|
||||
---
|
||||
|
||||
### [GRASP-02 Purgatory Git Data Fetching](grasp-02-proactive-sync-purgatory-git-data.md)
|
||||
**Proactive git data fetching from remote servers**
|
||||
|
||||
|
||||
@@ -736,6 +736,12 @@ fn compute_actions(
|
||||
|
||||
## Filter Building (Three-Layer Strategy)
|
||||
|
||||
Chunk sizes, filters-per-REQ packing, and sync concurrency are bounded by
|
||||
relay-imposed limits (subscription budgets, message sizes, filter byte caps).
|
||||
The verified constraints and the budget model that justifies these numbers
|
||||
live in
|
||||
[Sync Scaling Constraints and Budgets](sync-scaling-constraints.md).
|
||||
|
||||
### Layer 1: Announcements
|
||||
|
||||
- **Kinds**: 30617 (Repository Announcements), 30618 (Maintainer Lists)
|
||||
|
||||
@@ -0,0 +1,285 @@
|
||||
# Explanation: Sync Scaling Constraints and Budgets
|
||||
|
||||
**Purpose:** Explains the relay-imposed constraints that bound proactive sync,
|
||||
and justifies how we spend the three budgets they create — filter payload,
|
||||
subscriptions, and concurrency — as the watched item set grows.
|
||||
**Audience:** Contributors changing sync filter construction, subscription
|
||||
management, or negentropy scheduling; operators reasoning about scale limits.
|
||||
|
||||
---
|
||||
|
||||
## The Problem
|
||||
|
||||
Proactive sync ([GRASP-02](grasp-02-proactive-sync.md)) watches a growing set
|
||||
of items per relay: repository identifiers, repo references, and root event
|
||||
IDs. Every item must appear in filters twice — once in live subscriptions and
|
||||
once in historic sync (negentropy or REQ+EOSE). As the watched set grows, sync
|
||||
pressure on each relay grows along three axes:
|
||||
|
||||
1. **Filter payload** — how many items fit in one filter / one message.
|
||||
2. **Subscription count** — how many concurrent subscriptions we hold.
|
||||
3. **Request concurrency** — how many sync operations run at once.
|
||||
|
||||
These axes are not independent: relays bound them with shared, mostly
|
||||
undiscoverable limits. This document records the limits we verified, the
|
||||
budget model derived from them, and the levers we use — in order — to scale.
|
||||
|
||||
Production motivation (2026-08-04, gitnostr.com): the bootstrap relay received
|
||||
146 filters in one startup action (869 repos + 3632 root events, chunked at
|
||||
100 items). Historic sync opened one negentropy round per filter with no
|
||||
bound, drawing 34 "too many concurrent NEG requests" rejections from nos.lol
|
||||
and 61 per-filter timeouts in two minutes.
|
||||
|
||||
---
|
||||
|
||||
## Constraint Inventory
|
||||
|
||||
Verified 2026-08-04 against implementation sources and live NIP-11 documents.
|
||||
Re-verify before relying on exact numbers; defaults change.
|
||||
|
||||
### strfry (most common large public relay implementation)
|
||||
|
||||
| Limit | Default | Source |
|
||||
| --- | --- | --- |
|
||||
| Tag values per filter (count) | none — byte-capped | `src/filters.h:41` |
|
||||
| Tag value bytes per filter set | 65535 | `src/filters.h:41` |
|
||||
| Tag fields per filter | 3 (`maxTagsPerFilter`) | `golpe.yaml` |
|
||||
| Filters per REQ | 200 (`maxReqFilterSize`); 3 if optional `filterValidation` enabled | `golpe.yaml` |
|
||||
| Subscriptions per connection | 200 (`maxSubsPerConnection`) | `golpe.yaml` |
|
||||
| Concurrent negentropy | **shares `maxSubsPerConnection`** — no separate knob | `src/apps/relay/RelayNegentropy.cpp` |
|
||||
| WebSocket message size | 131072 (`maxWebsocketPayloadSize`) | `golpe.yaml` |
|
||||
|
||||
The key strfry finding: **negentropy views and ordinary subscriptions draw
|
||||
from the same per-connection budget.** "ERROR: too many concurrent NEG
|
||||
requests" is emitted when NEG views exceed `maxSubsPerConnection`.
|
||||
|
||||
### Live NIP-11 documents (operators tighten defaults)
|
||||
|
||||
| Relay | `max_subscriptions` | `max_message_length` |
|
||||
| --- | --- | --- |
|
||||
| nos.lol (strfry) | **20** | 131072 |
|
||||
| relay.primal.net | **20** | 1000000 |
|
||||
| nostr.wine | 50 | 524288 |
|
||||
| relay.damus.io | 200 | 1000000 |
|
||||
| relay.nostr.band, relay.ngit.dev | not advertised | not advertised |
|
||||
|
||||
### Discoverability gap (NIP-11)
|
||||
|
||||
NIP-11 `limitation` has **no field for tag values per filter and no field for
|
||||
filters per REQ**. Only `max_subscriptions` and `max_message_length` are
|
||||
advertised, and many relays omit `limitation` entirely. Consequence: filter
|
||||
sizing cannot be negotiated per relay — it must be statically conservative,
|
||||
with reactive fallback as the backstop.
|
||||
|
||||
### Our own embedded relay (nostr-sdk `LocalRelay`, 0.45.0-alpha.8)
|
||||
|
||||
- `max_reqs` = 500, enforced for REQ only (`src/nostr/builder.rs`).
|
||||
- Negentropy: **no concurrency limit at all** (upstream `TODO`), 60000-byte
|
||||
frame limit per NEG message.
|
||||
- No limits on filters per REQ or tag values per filter.
|
||||
|
||||
khatru (used by the ngit-relay reference implementation) and nostr-rs-relay
|
||||
similarly enforce no filter-size limits by default.
|
||||
|
||||
### Working floors
|
||||
|
||||
Derived from the tightest commonly observed values; all sizing below assumes:
|
||||
|
||||
- **Subscription budget B = 20** per connection (nos.lol, relay.primal.net),
|
||||
shared between live REQs, NEG rounds, and fallback REQs.
|
||||
- **Message budget M = 128 KB** (nos.lol); we target ≤ 48 KB of filter payload
|
||||
per message to leave margin for envelope and future growth.
|
||||
- **Per-filter value budget 32 KB** (half of strfry's 65535-byte set cap).
|
||||
- A serialized 64-char hex ID costs ~67 bytes (`"…",`), so:
|
||||
~480 hex IDs per filter, ~700 hex IDs per message. Variable-length values
|
||||
(`#d` identifiers, repo references) must be budgeted by bytes, not count.
|
||||
|
||||
---
|
||||
|
||||
## Our Approach: A Per-Connection Budget Ledger
|
||||
|
||||
Each relay connection owns one budget of B subscription slots. Three
|
||||
consumers share it, in priority order:
|
||||
|
||||
1. **Live subscriptions** (persistent, `limit: 0`) — the product; sized first.
|
||||
2. **Reserved margin** (2 slots) — the Layer-1 announcement subscription plus
|
||||
one spare for ad-hoc operations.
|
||||
3. **Historic sync** (transient) — negentropy rounds and REQ+EOSE fallback
|
||||
subscriptions get the remainder: `N = clamp(B − L − margin, 1, 4)`.
|
||||
|
||||
Historic work is transient, so even `N = 1` makes progress; live coverage is
|
||||
what must never be sacrificed. When even live subscriptions cannot fit
|
||||
(lever 4 below), the budget multiplies across connections rather than being
|
||||
overdrawn.
|
||||
|
||||
The levers, in the order we reach for them:
|
||||
|
||||
### Lever 1: Maximise items per filter (byte-budgeted chunking)
|
||||
|
||||
Replace the fixed 100-items-per-chunk rule with byte budgets: a filter chunk
|
||||
is full when it reaches ~32 KB of serialized tag values (~480 hex IDs), and a
|
||||
message is full at ~48 KB. The 100-item chunk was a guess made when we
|
||||
believed relays capped item counts; the verified constraints are byte caps
|
||||
(strfry 65535 per filter set, message size per NIP-11), so counting items
|
||||
wastes ~5× capacity for hex IDs while being *unsafe* for unbounded-length
|
||||
`#d` identifiers.
|
||||
|
||||
Because the limits are not discoverable (NIP-11 gap), the budget is static
|
||||
and conservative rather than probed; the existing transient-failure cooldown
|
||||
and REQ+EOSE fallback absorb the rare relay with tighter limits.
|
||||
|
||||
What this lever cannot do: collapse the three tag-variant filters. NIP-01
|
||||
ANDs distinct tag conditions within one filter, so `a`/`A`/`q` (and
|
||||
`e`/`E`/`q`) coverage requires three filters per chunk regardless of size.
|
||||
strfry's `maxTagsPerFilter = 3` counts tag *fields* per filter; our filters
|
||||
use one tag field each, so this is not a binding constraint.
|
||||
|
||||
### Lever 2: Pack filters per REQ — coupled to lever 1 by message size
|
||||
|
||||
Live subscriptions send all their filters in one REQ message, so the message
|
||||
budget M caps **items per subscription** (~700 hex IDs at M = 128 KB) no
|
||||
matter how items are split into filters. Bigger chunks (lever 1) therefore
|
||||
mean *fewer filters per REQ*, not more items per subscription: 10 × 500-ID
|
||||
filters in one REQ would be ~330 KB — over nos.lol's cap. The current
|
||||
`MAX_FILTERS_PER_REQ = 10` with 100-item chunks (~67 KB messages) sits near
|
||||
the limit already; the correct rule is a byte budget per REQ message, with
|
||||
filter count as a secondary bound (strfry accepts 200 filters per REQ, but
|
||||
its optional strict `filterValidation` mode accepts only 3).
|
||||
|
||||
### Lever 3: Bound and schedule concurrency (coordination with live sync)
|
||||
|
||||
Negentropy reconciles one filter per round, and each in-flight round consumes
|
||||
a subscription slot **from the same budget as live subscriptions** (strfry).
|
||||
So concurrency is not a free scaling axis; it is the residual of the ledger:
|
||||
|
||||
- Per-connection NEG concurrency `N = clamp(B − L − margin, 1, 4)` — with the
|
||||
B = 20 floor and typical live loads, effectively ≤ 4.
|
||||
- Rounds queue behind a per-connection semaphore; each completion releases
|
||||
the next. No timed batches or sleeps — throughput degrades smoothly instead
|
||||
of bursting into rejections.
|
||||
- Permit acquisition checks relay health first: while a rate-limit or
|
||||
transient-failure cooldown is active, queued rounds take the REQ+EOSE
|
||||
fallback path (which is itself budget-accounted) instead of firing into a
|
||||
relay that just complained.
|
||||
- The reactive machinery (escalating cooldown, NOTICE-based rate-limit pause,
|
||||
per-batch fallback) remains the backstop for relays whose limits are below
|
||||
our floors — prevention first, reaction second.
|
||||
|
||||
### Lever 4: Multiple connections per relay (last resort)
|
||||
|
||||
strfry-family limits are **per connection**, so a second connection doubles
|
||||
both the subscription budget and the NEG budget at that relay. This is the
|
||||
escalation path when a relay's watched set can no longer fit:
|
||||
`needed_live_slots + margin + 1 > B` even after levers 1–2.
|
||||
|
||||
Costs and risks, which is why it is last:
|
||||
|
||||
- Per-IP connection caps exist but are not advertised anywhere; exceeding
|
||||
them looks like abuse and risks bans. Bound connections per relay (≤ 4) and
|
||||
scale in with hysteresis.
|
||||
- Each connection re-authenticates (NIP-42) and carries its own health state,
|
||||
file descriptor, and TLS/session overhead.
|
||||
- Filter-to-connection assignment must be deterministic (stable sharding of
|
||||
the watched set) so reconnects and consolidation do not reshuffle
|
||||
subscriptions across the pool.
|
||||
|
||||
### Where the pressure actually lands
|
||||
|
||||
Budget pressure is worst where the watched set is largest — today that is our
|
||||
own bootstrap relay (869 repos / 3632 roots ≈ 1 MB of serialized tag values,
|
||||
i.e. ~21 messages minimum even optimally packed). Public relays typically
|
||||
carry small per-relay target sets but tight budgets (B = 20). Two
|
||||
consequences:
|
||||
|
||||
- For infrastructure we control (bootstrap, self-relay), raise and advertise
|
||||
server-side limits rather than spending client-side levers.
|
||||
- For public relays, levers 1–3 keep us comfortably inside B = 20 at current
|
||||
scale; lever 4 exists for the point where a single public relay's target
|
||||
set outgrows ~`(B − margin) × 700` hex-ID-equivalents (~12–13 k items).
|
||||
|
||||
---
|
||||
|
||||
## Serving-Side Obligations
|
||||
|
||||
We are also a relay, and peer GRASP instances run this same sync against us.
|
||||
The embedded relay currently enforces **no negentropy concurrency limit**
|
||||
(upstream nostr-sdk `TODO`) and no filter-size limits — the mirror image of
|
||||
the client-side incident that motivated this document. At scale we must:
|
||||
|
||||
1. Enforce server-side bounds (NEG concurrency, filters per REQ, filter
|
||||
payload) so one peer cannot exhaust us.
|
||||
2. Advertise our limits in NIP-11 `limitation` (`max_subscriptions`,
|
||||
`max_message_length`) so well-behaved peers can budget against us —
|
||||
partially compensating for the discoverability gap we suffer as a client.
|
||||
|
||||
---
|
||||
|
||||
## Trade-offs
|
||||
|
||||
**Gained:** deterministic behaviour against unadvertised limits; startup
|
||||
bursts bounded by design rather than absorbed by cooldowns; a single model
|
||||
(the ledger) that live sync, historic sync, and fallback all account against;
|
||||
a defined escalation path to multi-connection scale.
|
||||
|
||||
**Given up:** peak theoretical throughput on permissive relays (a damus-class
|
||||
relay with 200 subscription slots is used as if it had 20 when `limitation`
|
||||
is absent — we only relax budgets when NIP-11 advertises headroom); some
|
||||
implementation complexity (byte-budgeted chunking, permit-gated scheduling,
|
||||
eventual sharding).
|
||||
|
||||
---
|
||||
|
||||
## Alternatives Considered
|
||||
|
||||
### Adaptive probing (start big, shrink on rejection)
|
||||
|
||||
**Pros:** discovers each relay's true limits; no static guesswork.
|
||||
**Cons:** rejection signals are non-standard free-text NOTICEs; every
|
||||
startup pays a rejection burst per relay; failure attribution is ambiguous
|
||||
(payload size vs. subscription count vs. rate limit), so the probe can learn
|
||||
the wrong lesson.
|
||||
**Why not:** we tried the reactive-only posture implicitly and it produced
|
||||
the 2026-08-04 incident; static floors with reactive backstop are
|
||||
deterministic and testable.
|
||||
|
||||
### NIP-11-driven budgets
|
||||
|
||||
**Pros:** honest relays advertise `max_subscriptions` and
|
||||
`max_message_length`; budgets could be exact.
|
||||
**Why partial:** the two advertised fields *are* consumed when present
|
||||
(relaxing B and M above the floors), but per-filter and filters-per-REQ
|
||||
limits simply have no NIP-11 field, and many relays omit `limitation`
|
||||
entirely — so floors remain necessary. Proposing a NIP-11 extension for
|
||||
filter-size limits is worthwhile upstream work.
|
||||
|
||||
### Timed batching with pause-on-rate-limit
|
||||
|
||||
**Pros:** simple to picture.
|
||||
**Cons:** reactive by construction (eats one rejection burst per relay per
|
||||
startup), needs heuristic NOTICE parsing as its *primary* control loop, and
|
||||
fixed pauses waste time on fast relays while still bursting slow ones.
|
||||
**Why not:** the semaphore ledger achieves the same containment continuously,
|
||||
with the heuristics demoted to backstop.
|
||||
|
||||
---
|
||||
|
||||
## Rollout Mapping
|
||||
|
||||
| Lever | Status |
|
||||
| --- | --- |
|
||||
| 3 — bounded NEG concurrency | Stabilisation cycle 3 (in flight) |
|
||||
| 1 + 2 — byte-budgeted chunking and REQ packing | Proposed as cycle 4 |
|
||||
| Ledger unification (live + historic + fallback against one budget, NIP-11-aware B/M) | Design accepted here; implement after cycles 3–4 |
|
||||
| 4 — multi-connection sharding | Deferred until a relay's target set approaches the single-connection ceiling |
|
||||
| Serving-side limits + NIP-11 advertisement | Follow-up work item |
|
||||
|
||||
---
|
||||
|
||||
## Related Documentation
|
||||
|
||||
- [GRASP-02 Proactive Sync](grasp-02-proactive-sync.md) — the sync
|
||||
architecture these budgets apply to (filter layers, live vs historic,
|
||||
negentropy fallback).
|
||||
- [Defensive Measures & Rate Limiting](defensive-measures.md) — the
|
||||
serving-side counterpart.
|
||||
- [Monitoring Overview](monitoring.md) — metrics for observing sync health.
|
||||
Reference in New Issue
Block a user