Files
ngit-grasp/b4d0-failed-historic-sync-batches-not-retried.md
T
2026-01-21 13:46:36 +00:00

4.5 KiB

Failed historic sync batches not retried after negentropy timeout

ID: b4d0

Problem

When syncing from a relay that doesn't support negentropy (NIP-77), historic sync batches fail and are never retried. The fallback to REQ+EOSE works for the current batch, but failed batches are abandoned, causing permanent data loss for historic events.

Root Cause:

  • Archive attempts negentropy sync for historic events
  • ngit-relay doesn't support NIP-77, so negentropy times out
  • Batch marked as failed → status: ConnectedHistoricSyncFailures
  • REQ+EOSE fallback works for current batch, but failed batches are NEVER retried

Impact:

  • Archive syncing from ngit-relay lost ~164 state events from before archive start date
  • Status shows ConnectedHistoricSyncFailures but no retry mechanism
  • Workaround: Disable negentropy with NGIT_SYNC_DISABLE_NEGENTROPY=true

Code References:

  • src/sync/mod.rs:3152-3512 - Historic sync with fallback logic
  • src/sync/relay_connection.rs:508-584 - Negentropy failure detection
  • src/sync/mod.rs:1109-1113 - Batch failure handling (no retry mechanism)

Analysis Document: /persistent/dcdev/clones/ngit-grasp/worktrees/820a-relay-ngit-dev-migration/work/NEGENTROPY-FALLBACK-ANALYSIS.md

Plan

  • Implement REQ+EOSE fallback when negentropy retry fails (simpler than 4-phase plan)

Original 4-phase plan replaced with simpler solution: The original plan was overkill. Instead of building new retry infrastructure, we utilize the existing REQ+EOSE mechanism. When negentropy retry makes no progress (relay returns zero events), we now fall back to REQ+EOSE for the missing events instead of marking the batch as failed.

Progress

2026-01-21 [Session Initial]

  • Started: Issue created to track negentropy sync bug
  • Completed: Applied workaround configuration to archive service (disable negentropy)
  • Next: Implement proper retry mechanism for failed batches

2026-01-21 [Session 11:22]

  • Completed: Updated NixOS configuration with syncDisableNegentropy = true
  • Completed: Deployed configuration to vps1 (commit 37f30bf)
  • Completed: Service restarted successfully with REQ+EOSE sync method
  • Decision: No session files to delete - sync state is in-memory only
  • Verified: Historic sync completed with sync_method=ReqEose, had_failures=false
  • Verified: State events (kind 30618) are being received from relay.ngit.dev
  • Next: Monitor sync completeness and verify all historic events are retrieved

2026-01-21 [Session 18:45]

  • Analysis: Identified root cause - when negentropy retry returns zero events, batch is marked as failed instead of falling back to REQ+EOSE
  • Decision: Replaced 4-phase plan with simpler solution - utilize existing REQ+EOSE mechanism
  • Completed: Implemented REQ+EOSE fallback in src/sync/mod.rs handle_eose() function
  • Completed: cargo check and clippy pass
  • Note: No existing fallback tests found (rust-nostr doesn't allow disabling NIP-77)
  • Review: Fixed issue where fallback used same ID-based filters that already failed
  • Commit: f306256 (amended)

2026-01-21 [Session 19:40]

  • Completed: Rebased onto master (5913c5a)
  • Completed: Merged to master
  • Status: DONE

Notes

  • Discovered during relay.ngit.dev migration analysis
  • Workaround applied: NGIT_SYNC_DISABLE_NEGENTROPY=true in archive service
  • Fix implemented: REQ+EOSE fallback when negentropy retry fails
  • Related to NIP-77 support detection and graceful degradation

Solution Details

The Bug: In handle_eose(), when negentropy retry made no progress (relay returned zero events), the batch was marked as failed (batch.failed = true) instead of falling back to REQ+EOSE. The code even had a TODO comment acknowledging this.

The Fix: When negentropy retry fails, instead of marking the batch as failed:

  1. Create REQ+EOSE subscriptions using semantic filters (kind/author/tags) from batch.items
  2. This is different from the failed ID-based queries - semantic filters may succeed where ID queries fail
  3. Change batch sync_method to ReqEose
  4. Clear negentropy-specific tracking fields
  5. Wait for EOSE on the new subscriptions
  6. Only mark as failed if ALL fallback subscriptions fail

Key insight: The original implementation incorrectly used ID-based filters (Filter::new().ids(...)) for the fallback, which would just repeat the same query that already failed. The correct approach uses build_layer2_and_layer3_filters() to reconstruct semantic filters from the batch's repos and root_events.

This reuses the existing REQ+EOSE mechanism rather than building new retry infrastructure.