mirror of
https://relay.ngit.dev/npub15qydau2hjma6ngxkl2cyar74wzyjshvl65za5k5rl69264ar2exs5cyejr/ngit-grasp.git
synced 2026-10-05 15:08:24 +00:00
6.7 KiB
6.7 KiB
Purgatory retry mechanism fails during initial sync when clone URLs are unreachable
ID: 1ce6
Problem
During initial sync (archive/migration scenarios), ngit-grasp's purgatory mechanism can permanently reject events that should be retried, causing significant data loss.
How Purgatory Works
- Events arrive at the relay and go to purgatory (30-minute expiry) while git data is fetched
- If git fetch succeeds within 30 minutes, event is processed
- If git fetch fails, event expires and its ID is added to
expired_events - Critical: Once in
expired_events, the event will NOT be retried until the 7-day cleanup
The Problem
During initial sync, archive fetches events from source relay. When clone URLs are unreachable or throttled:
- Git fetch fails for the event
- After 30 minutes, event expires
- Event ID added to
expired_events(7-day rejection) - Event will NOT be retried even if clone URLs become available
Real-World Impact (relay.ngit.dev Migration)
- Archive synced 1,858 repos from relay.ngit.dev
- 1,581 repos (85.1%) stuck in purgatory - git data never fetched
- Only 277 repos (14.9%) have git data
- State event coverage: 47.0% instead of expected >99%
- Migration blocked until purgatory state manually cleared
Root Cause Analysis
- External relay unavailability: Archive tried to fetch from external relays (git.shakespeare.diy, etc.) that were unreachable or throttling
- No fallback logic: Archive doesn't try alternative clone URLs from the same event
- Permanent rejection: Once event expires, it's rejected for 7 days (no retry mechanism)
- Silent failure: No prominent logging or metrics for purgatory accumulation
Why This Matters
- Initial sync is fragile: If external relays are down during initial sync, repos are stuck for 7 days
- No retry mechanism: Once event expires, it won't be retried even if clone URLs become available
- No fallback logic: Archive doesn't try alternative clone URLs from the same event
- Silent failure: No prominent logging or metrics for purgatory accumulation
- Manual intervention required: Only way to retry is to delete purgatory state file and restart
- Not scalable: Manual intervention doesn't scale for production deployments
Plan
Phase 1: Retry with Exponential Backoff
- Don't permanently reject events after first failure
- Implement retry with increasing delays (1h, 4h, 12h, 24h, 7d)
- Only permanently reject after multiple failures over 7 days
- Track retry count and last attempt time per event
Phase 2: Clone URL Fallback Logic
- Parse all clone URLs from event (not just first one)
- Try secondary URLs if primary fails
- Log which URL succeeded for debugging
- Consider URL priority/preference ordering
Phase 3: Purgatory Metrics and Alerts
- Expose
ngit_purgatory_countmetric - Expose
ngit_expired_events_countmetric - Expose
ngit_purgatory_retry_countmetric - Log warnings when purgatory accumulates (>10%, >50%, >80% of total events)
- Add structured logging for purgatory state changes
Phase 4: Configuration Options
purgatory_retry_enabled(default: true)purgatory_retry_max_attempts(default: 5)purgatory_retry_backoff_multiplier(default: 2)purgatory_expiry_duration(default: 30min, allow longer for initial sync)- Document all options in configuration reference
Phase 5: Initial Sync Mode
- Detect initial sync (empty database or explicit flag)
- Use longer purgatory expiry (4 hours instead of 30 minutes)
- More aggressive retry logic during initial sync
- Better logging and progress reporting
- Consider
NGIT_INITIAL_SYNC_MODE=trueenvironment variable
Acceptance Criteria
- Events are retried multiple times before permanent rejection
- All clone URLs are attempted before giving up
- Purgatory metrics are exposed (count, expired, retries)
- Initial sync has special handling (longer expiry, more retries)
- Manual intervention not required for transient failures
- Configuration options documented and tested
- Integration tests cover retry scenarios
Technical Notes
Current Purgatory Implementation
- Location:
src/purgatory/mod.rs - Expiry tracking:
expired_events: DashMap<EventId, Instant>(line 79) - Expiry duration: 30 minutes (hardcoded)
- Cleanup: 7-day expiry for
expired_eventsentries
Proposed Changes
-
New fields in purgatory:
struct RetryState { attempts: u32, last_attempt: Instant, next_retry: Instant, failed_urls: Vec<String>, } retry_state: DashMap<EventId, RetryState> -
Retry backoff schedule:
- Attempt 1: Immediate
- Attempt 2: 1 hour
- Attempt 3: 4 hours
- Attempt 4: 12 hours
- Attempt 5: 24 hours
- After 5 failures: Permanent rejection (7-day expiry)
-
Clone URL fallback:
- Parse
clonetags from event - Try URLs in order until one succeeds
- Track which URLs failed for debugging
- Parse
Related Context
- Migration issue: 820a-relay-ngit-dev-migration
- Purgatory persistence: a6b6-purgatory-surive-reboot (completed)
- Archive sync results: 1,858 repos, 1,581 in purgatory (85.1%)
- Workaround: Delete
purgatory-state.jsonand restart relay - Investigation: work/state-validation/ARCHIVE-PURGATORY-ANALYSIS.md
Workaround (Until Fixed)
For operators experiencing this issue:
# 1. Stop the relay
systemctl stop ngit-grasp
# 2. Delete purgatory state (forces re-sync of all events)
rm /path/to/data/purgatory-state.json
# 3. Optionally delete rejected events cache
rm /path/to/data/rejected-events-cache.json
# 4. Restart relay
systemctl start ngit-grasp
# 5. Monitor purgatory accumulation
journalctl -u ngit-grasp -f | grep -i purgatory
Note: This workaround requires manual intervention and may cause duplicate processing. The proper fix is implementing the retry mechanism described above.
Progress
2026-01-20 [Session 12:00]
- Created: Issue documenting purgatory retry problem discovered during relay.ngit.dev migration
- Context: Archive sync showed 85.1% of repos stuck in purgatory due to unreachable clone URLs
- Root cause: No retry mechanism after initial git fetch failure
- Impact: Migration blocked, manual intervention required
- Next: Prioritize based on migration timeline
Notes
- Priority: High (blocks migration scenarios)
- Discovered during: 820a relay.ngit.dev migration
- Affects: Archive mode, initial sync, any scenario with unreliable external relays
- Related to but distinct from: a6b6 (persistence) - that issue ensures purgatory survives restarts, this issue addresses retry logic