mirror of
https://relay.ngit.dev/npub15qydau2hjma6ngxkl2cyar74wzyjshvl65za5k5rl69264ar2exs5cyejr/ngit-grasp.git
synced 2026-10-05 15:08:24 +00:00
86 lines
4.5 KiB
Markdown
86 lines
4.5 KiB
Markdown
# Failed historic sync batches not retried after negentropy timeout
|
|
|
|
**ID:** b4d0
|
|
|
|
## Problem
|
|
|
|
When syncing from a relay that doesn't support negentropy (NIP-77), historic sync batches fail and are never retried. The fallback to REQ+EOSE works for the current batch, but failed batches are abandoned, causing permanent data loss for historic events.
|
|
|
|
**Root Cause:**
|
|
- Archive attempts negentropy sync for historic events
|
|
- ngit-relay doesn't support NIP-77, so negentropy times out
|
|
- Batch marked as failed → status: `ConnectedHistoricSyncFailures`
|
|
- REQ+EOSE fallback works for current batch, but failed batches are NEVER retried
|
|
|
|
**Impact:**
|
|
- Archive syncing from ngit-relay lost ~164 state events from before archive start date
|
|
- Status shows `ConnectedHistoricSyncFailures` but no retry mechanism
|
|
- Workaround: Disable negentropy with `NGIT_SYNC_DISABLE_NEGENTROPY=true`
|
|
|
|
**Code References:**
|
|
- `src/sync/mod.rs:3152-3512` - Historic sync with fallback logic
|
|
- `src/sync/relay_connection.rs:508-584` - Negentropy failure detection
|
|
- `src/sync/mod.rs:1109-1113` - Batch failure handling (no retry mechanism)
|
|
|
|
**Analysis Document:** `/persistent/dcdev/clones/ngit-grasp/worktrees/820a-relay-ngit-dev-migration/work/NEGENTROPY-FALLBACK-ANALYSIS.md`
|
|
|
|
## Plan
|
|
|
|
- [x] Implement REQ+EOSE fallback when negentropy retry fails (simpler than 4-phase plan)
|
|
|
|
**Original 4-phase plan replaced with simpler solution:**
|
|
The original plan was overkill. Instead of building new retry infrastructure, we utilize the existing REQ+EOSE mechanism. When negentropy retry makes no progress (relay returns zero events), we now fall back to REQ+EOSE for the missing events instead of marking the batch as failed.
|
|
|
|
## Progress
|
|
|
|
### 2026-01-21 [Session Initial]
|
|
- Started: Issue created to track negentropy sync bug
|
|
- Completed: Applied workaround configuration to archive service (disable negentropy)
|
|
- Next: Implement proper retry mechanism for failed batches
|
|
|
|
### 2026-01-21 [Session 11:22]
|
|
- Completed: Updated NixOS configuration with `syncDisableNegentropy = true`
|
|
- Completed: Deployed configuration to vps1 (commit 37f30bf)
|
|
- Completed: Service restarted successfully with REQ+EOSE sync method
|
|
- Decision: No session files to delete - sync state is in-memory only
|
|
- Verified: Historic sync completed with `sync_method=ReqEose`, `had_failures=false`
|
|
- Verified: State events (kind 30618) are being received from relay.ngit.dev
|
|
- Next: Monitor sync completeness and verify all historic events are retrieved
|
|
|
|
### 2026-01-21 [Session 18:45]
|
|
- Analysis: Identified root cause - when negentropy retry returns zero events, batch is marked as failed instead of falling back to REQ+EOSE
|
|
- Decision: Replaced 4-phase plan with simpler solution - utilize existing REQ+EOSE mechanism
|
|
- Completed: Implemented REQ+EOSE fallback in `src/sync/mod.rs` handle_eose() function
|
|
- Completed: cargo check and clippy pass
|
|
- Note: No existing fallback tests found (rust-nostr doesn't allow disabling NIP-77)
|
|
- Review: Fixed issue where fallback used same ID-based filters that already failed
|
|
- Commit: f306256 (amended)
|
|
|
|
### 2026-01-21 [Session 19:40]
|
|
- Completed: Rebased onto master (5913c5a)
|
|
- Completed: Merged to master
|
|
- Status: DONE
|
|
|
|
## Notes
|
|
|
|
- Discovered during relay.ngit.dev migration analysis
|
|
- Workaround applied: `NGIT_SYNC_DISABLE_NEGENTROPY=true` in archive service
|
|
- Fix implemented: REQ+EOSE fallback when negentropy retry fails
|
|
- Related to NIP-77 support detection and graceful degradation
|
|
|
|
## Solution Details
|
|
|
|
**The Bug:** In `handle_eose()`, when negentropy retry made no progress (relay returned zero events), the batch was marked as failed (`batch.failed = true`) instead of falling back to REQ+EOSE. The code even had a TODO comment acknowledging this.
|
|
|
|
**The Fix:** When negentropy retry fails, instead of marking the batch as failed:
|
|
1. Create REQ+EOSE subscriptions using **semantic filters** (kind/author/tags) from `batch.items`
|
|
2. This is different from the failed ID-based queries - semantic filters may succeed where ID queries fail
|
|
3. Change batch `sync_method` to `ReqEose`
|
|
4. Clear negentropy-specific tracking fields
|
|
5. Wait for EOSE on the new subscriptions
|
|
6. Only mark as failed if ALL fallback subscriptions fail
|
|
|
|
**Key insight:** The original implementation incorrectly used ID-based filters (`Filter::new().ids(...)`) for the fallback, which would just repeat the same query that already failed. The correct approach uses `build_layer2_and_layer3_filters()` to reconstruct semantic filters from the batch's repos and root_events.
|
|
|
|
This reuses the existing REQ+EOSE mechanism rather than building new retry infrastructure.
|