diff --git a/b4d0-failed-historic-sync-batches-not-retried.md b/b4d0-failed-historic-sync-batches-not-retried.md new file mode 100644 index 0000000..0380621 --- /dev/null +++ b/b4d0-failed-historic-sync-batches-not-retried.md @@ -0,0 +1,46 @@ +# Failed historic sync batches not retried after negentropy timeout + +**ID:** b4d0 + +## Problem + +When syncing from a relay that doesn't support negentropy (NIP-77), historic sync batches fail and are never retried. The fallback to REQ+EOSE works for the current batch, but failed batches are abandoned, causing permanent data loss for historic events. + +**Root Cause:** +- Archive attempts negentropy sync for historic events +- ngit-relay doesn't support NIP-77, so negentropy times out +- Batch marked as failed → status: `ConnectedHistoricSyncFailures` +- REQ+EOSE fallback works for current batch, but failed batches are NEVER retried + +**Impact:** +- Archive syncing from ngit-relay lost ~164 state events from before archive start date +- Status shows `ConnectedHistoricSyncFailures` but no retry mechanism +- Workaround: Disable negentropy with `NGIT_SYNC_DISABLE_NEGENTROPY=true` + +**Code References:** +- `src/sync/mod.rs:3152-3512` - Historic sync with fallback logic +- `src/sync/relay_connection.rs:508-584` - Negentropy failure detection +- `src/sync/mod.rs:1109-1113` - Batch failure handling (no retry mechanism) + +**Analysis Document:** `/persistent/dcdev/clones/ngit-grasp/worktrees/820a-relay-ngit-dev-migration/work/NEGENTROPY-FALLBACK-ANALYSIS.md` + +## Plan + +- [ ] Phase 1: Add retry mechanism for failed historic sync batches +- [ ] Phase 2: Implement exponential backoff for retries +- [ ] Phase 3: Force REQ+EOSE fallback after first negentropy timeout +- [ ] Phase 4: Add monitoring for sync completeness (announcement vs state event ratio) + +## Progress + +### 2026-01-21 [Session Initial] +- Started: Issue created to track negentropy sync bug +- Completed: Applied workaround configuration to archive service (disable negentropy) +- Next: Implement proper retry mechanism for failed batches + +## Notes + +- Discovered during relay.ngit.dev migration analysis +- Workaround applied: `NGIT_SYNC_DISABLE_NEGENTROPY=true` in archive service +- Long-term fix requires retry mechanism in sync engine +- Related to NIP-77 support detection and graceful degradation