Files
ngit-grasp/b4d0-failed-historic-sync-batches-not-retried.md
T
2026-01-21 13:46:36 +00:00

86 lines
4.5 KiB
Markdown

# Failed historic sync batches not retried after negentropy timeout
**ID:** b4d0
## Problem
When syncing from a relay that doesn't support negentropy (NIP-77), historic sync batches fail and are never retried. The fallback to REQ+EOSE works for the current batch, but failed batches are abandoned, causing permanent data loss for historic events.
**Root Cause:**
- Archive attempts negentropy sync for historic events
- ngit-relay doesn't support NIP-77, so negentropy times out
- Batch marked as failed → status: `ConnectedHistoricSyncFailures`
- REQ+EOSE fallback works for current batch, but failed batches are NEVER retried
**Impact:**
- Archive syncing from ngit-relay lost ~164 state events from before archive start date
- Status shows `ConnectedHistoricSyncFailures` but no retry mechanism
- Workaround: Disable negentropy with `NGIT_SYNC_DISABLE_NEGENTROPY=true`
**Code References:**
- `src/sync/mod.rs:3152-3512` - Historic sync with fallback logic
- `src/sync/relay_connection.rs:508-584` - Negentropy failure detection
- `src/sync/mod.rs:1109-1113` - Batch failure handling (no retry mechanism)
**Analysis Document:** `/persistent/dcdev/clones/ngit-grasp/worktrees/820a-relay-ngit-dev-migration/work/NEGENTROPY-FALLBACK-ANALYSIS.md`
## Plan
- [x] Implement REQ+EOSE fallback when negentropy retry fails (simpler than 4-phase plan)
**Original 4-phase plan replaced with simpler solution:**
The original plan was overkill. Instead of building new retry infrastructure, we utilize the existing REQ+EOSE mechanism. When negentropy retry makes no progress (relay returns zero events), we now fall back to REQ+EOSE for the missing events instead of marking the batch as failed.
## Progress
### 2026-01-21 [Session Initial]
- Started: Issue created to track negentropy sync bug
- Completed: Applied workaround configuration to archive service (disable negentropy)
- Next: Implement proper retry mechanism for failed batches
### 2026-01-21 [Session 11:22]
- Completed: Updated NixOS configuration with `syncDisableNegentropy = true`
- Completed: Deployed configuration to vps1 (commit 37f30bf)
- Completed: Service restarted successfully with REQ+EOSE sync method
- Decision: No session files to delete - sync state is in-memory only
- Verified: Historic sync completed with `sync_method=ReqEose`, `had_failures=false`
- Verified: State events (kind 30618) are being received from relay.ngit.dev
- Next: Monitor sync completeness and verify all historic events are retrieved
### 2026-01-21 [Session 18:45]
- Analysis: Identified root cause - when negentropy retry returns zero events, batch is marked as failed instead of falling back to REQ+EOSE
- Decision: Replaced 4-phase plan with simpler solution - utilize existing REQ+EOSE mechanism
- Completed: Implemented REQ+EOSE fallback in `src/sync/mod.rs` handle_eose() function
- Completed: cargo check and clippy pass
- Note: No existing fallback tests found (rust-nostr doesn't allow disabling NIP-77)
- Review: Fixed issue where fallback used same ID-based filters that already failed
- Commit: f306256 (amended)
### 2026-01-21 [Session 19:40]
- Completed: Rebased onto master (5913c5a)
- Completed: Merged to master
- Status: DONE
## Notes
- Discovered during relay.ngit.dev migration analysis
- Workaround applied: `NGIT_SYNC_DISABLE_NEGENTROPY=true` in archive service
- Fix implemented: REQ+EOSE fallback when negentropy retry fails
- Related to NIP-77 support detection and graceful degradation
## Solution Details
**The Bug:** In `handle_eose()`, when negentropy retry made no progress (relay returned zero events), the batch was marked as failed (`batch.failed = true`) instead of falling back to REQ+EOSE. The code even had a TODO comment acknowledging this.
**The Fix:** When negentropy retry fails, instead of marking the batch as failed:
1. Create REQ+EOSE subscriptions using **semantic filters** (kind/author/tags) from `batch.items`
2. This is different from the failed ID-based queries - semantic filters may succeed where ID queries fail
3. Change batch `sync_method` to `ReqEose`
4. Clear negentropy-specific tracking fields
5. Wait for EOSE on the new subscriptions
6. Only mark as failed if ALL fallback subscriptions fail
**Key insight:** The original implementation incorrectly used ID-based filters (`Filter::new().ids(...)`) for the fallback, which would just repeat the same query that already failed. The correct approach uses `build_layer2_and_layer3_filters()` to reconstruct semantic filters from the batch's repos and root_events.
This reuses the existing REQ+EOSE mechanism rather than building new retry infrastructure.