4.5 KiB
Failed historic sync batches not retried after negentropy timeout
ID: b4d0
Problem
When syncing from a relay that doesn't support negentropy (NIP-77), historic sync batches fail and are never retried. The fallback to REQ+EOSE works for the current batch, but failed batches are abandoned, causing permanent data loss for historic events.
Root Cause:
- Archive attempts negentropy sync for historic events
- ngit-relay doesn't support NIP-77, so negentropy times out
- Batch marked as failed → status:
ConnectedHistoricSyncFailures - REQ+EOSE fallback works for current batch, but failed batches are NEVER retried
Impact:
- Archive syncing from ngit-relay lost ~164 state events from before archive start date
- Status shows
ConnectedHistoricSyncFailuresbut no retry mechanism - Workaround: Disable negentropy with
NGIT_SYNC_DISABLE_NEGENTROPY=true
Code References:
src/sync/mod.rs:3152-3512- Historic sync with fallback logicsrc/sync/relay_connection.rs:508-584- Negentropy failure detectionsrc/sync/mod.rs:1109-1113- Batch failure handling (no retry mechanism)
Analysis Document: /persistent/dcdev/clones/ngit-grasp/worktrees/820a-relay-ngit-dev-migration/work/NEGENTROPY-FALLBACK-ANALYSIS.md
Plan
- Implement REQ+EOSE fallback when negentropy retry fails (simpler than 4-phase plan)
Original 4-phase plan replaced with simpler solution: The original plan was overkill. Instead of building new retry infrastructure, we utilize the existing REQ+EOSE mechanism. When negentropy retry makes no progress (relay returns zero events), we now fall back to REQ+EOSE for the missing events instead of marking the batch as failed.
Progress
2026-01-21 [Session Initial]
- Started: Issue created to track negentropy sync bug
- Completed: Applied workaround configuration to archive service (disable negentropy)
- Next: Implement proper retry mechanism for failed batches
2026-01-21 [Session 11:22]
- Completed: Updated NixOS configuration with
syncDisableNegentropy = true - Completed: Deployed configuration to vps1 (commit 37f30bf)
- Completed: Service restarted successfully with REQ+EOSE sync method
- Decision: No session files to delete - sync state is in-memory only
- Verified: Historic sync completed with
sync_method=ReqEose,had_failures=false - Verified: State events (kind 30618) are being received from relay.ngit.dev
- Next: Monitor sync completeness and verify all historic events are retrieved
2026-01-21 [Session 18:45]
- Analysis: Identified root cause - when negentropy retry returns zero events, batch is marked as failed instead of falling back to REQ+EOSE
- Decision: Replaced 4-phase plan with simpler solution - utilize existing REQ+EOSE mechanism
- Completed: Implemented REQ+EOSE fallback in
src/sync/mod.rshandle_eose() function - Completed: cargo check and clippy pass
- Note: No existing fallback tests found (rust-nostr doesn't allow disabling NIP-77)
- Review: Fixed issue where fallback used same ID-based filters that already failed
- Commit: f306256 (amended)
2026-01-21 [Session 19:40]
- Completed: Rebased onto master (
5913c5a) - Completed: Merged to master
- Status: DONE
Notes
- Discovered during relay.ngit.dev migration analysis
- Workaround applied:
NGIT_SYNC_DISABLE_NEGENTROPY=truein archive service - Fix implemented: REQ+EOSE fallback when negentropy retry fails
- Related to NIP-77 support detection and graceful degradation
Solution Details
The Bug: In handle_eose(), when negentropy retry made no progress (relay returned zero events), the batch was marked as failed (batch.failed = true) instead of falling back to REQ+EOSE. The code even had a TODO comment acknowledging this.
The Fix: When negentropy retry fails, instead of marking the batch as failed:
- Create REQ+EOSE subscriptions using semantic filters (kind/author/tags) from
batch.items - This is different from the failed ID-based queries - semantic filters may succeed where ID queries fail
- Change batch
sync_methodtoReqEose - Clear negentropy-specific tracking fields
- Wait for EOSE on the new subscriptions
- Only mark as failed if ALL fallback subscriptions fail
Key insight: The original implementation incorrectly used ID-based filters (Filter::new().ids(...)) for the fallback, which would just repeat the same query that already failed. The correct approach uses build_layer2_and_layer3_filters() to reconstruct semantic filters from the batch's repos and root_events.
This reuses the existing REQ+EOSE mechanism rather than building new retry infrastructure.