Files
ngit-grasp/tests
DanConwayDev b3ca015968 fix(sync): retry events lost by incomplete historic-sync batches
Production logs after deploying a8964bb to gitnostr.com showed ~41
incomplete negentropy retries and 20 batches completing with partial
results within six minutes, some batches missing hundreds of events.
Negentropy reconciliation identifies event IDs missing locally, but a
relay's exact-ID response can return only a subset (or nothing on the
retry). Batches without repository/root-event metadata - the generic
Layer 1 announcements batch - cannot build a semantic REQ+EOSE
fallback, so handle_eose finalized them "with partial results" and
dropped the missing IDs entirely. Nothing retried them until the next
daily sync up to 25 hours later, leaving repository announcements and
their dependencies absent indefinitely.

Missing IDs from a batch that finalizes incomplete are now registered
in a per-relay recovery index (sync::missing_events), and the existing
sync maintenance timer refetches them over the relay's live connection
with bounded exponential backoff (30s doubling to 15min, one in-flight
attempt per relay, 300 IDs per fetch). Network I/O runs outside the
sync actor lock. Startup remains non-blocking: the batch still
finalizes as failed, the relay transitions to
ConnectedHistoricSyncFailures, and traffic is served while recovery
runs in the background.

Semantics:
- progress clears only the IDs actually recovered and resets backoff;
- duplicate incomplete responses merge into the pending set without
  duplicating work;
- IDs satisfied by live sync or user submission are cleared on the
  next tick without consuming attempt budget;
- attempts against a disconnected relay are deferred, not counted, so
  an unavailable relay neither expires its work nor loops tightly;
- 12 consecutive zero-progress attempts expire the pending IDs with an
  explicit warning; the relay stays observably degraded until the
  daily sync re-discovers the gap;
- full recovery promotes the relay back to Connected unless an
  unrelated batch failure was observed for it;
- nothing persists across restarts: historic sync re-runs from scratch
  and re-detects any still-missing events, so incomplete work is never
  falsely reported as complete.

Also fixes the retry-subscription-failure path, which confirmed an
incomplete batch without marking it failed (falsely reporting
Connected), and bounds the previously unbounded missing_ids log arrays
to a five-ID sample.

Regression coverage: a new censoring WebSocket proxy fixture sits
between a syncing relay and a real ngit-grasp bootstrap relay,
forwarding NIP-77 frames unchanged while withholding chosen EVENT
frames. The integration test reproduces the full production sequence
(subset response, zero-progress retry, no semantic fallback,
ConnectedHistoricSyncFailures) and proves the withheld event is
recovered and the relay promoted to Connected once the event becomes
available - without a restart and while live sync continues unstarved.
Unit tests cover registration dedupe, partial clears, backoff growth
and cap, explicit expiry, deferral, and health-restoration poisoning.
Full cargo test suite passes.

tokio-tungstenite was added as a dev-dependency for the proxy fixture;
it was already present transitively in Cargo.lock, so no Nix hash
updates are required (crates.io dependency under cargoLock).

Deliberately out of scope: durable persistence of pending recovery
work, retrying missing IDs against other relays, outbound-target
policy changes, and broader logging cleanup.

Confirms the closed issue
nostr:nevent1qqs94up6nnkzjlz4fcy5tesh8yxvr63xqjhg79etmc573fuunjt0qeqpz3mhxue69uhhyetvv9ujumn8d96zuer9wc5tdht6
2026-08-01 20:25:24 +00:00
..