Files
ngit-grasp/1ce6-purgatory-retry-mechanism.md

6.7 KiB

Purgatory retry mechanism fails during initial sync when clone URLs are unreachable

ID: 1ce6

Problem

During initial sync (archive/migration scenarios), ngit-grasp's purgatory mechanism can permanently reject events that should be retried, causing significant data loss.

How Purgatory Works

  1. Events arrive at the relay and go to purgatory (30-minute expiry) while git data is fetched
  2. If git fetch succeeds within 30 minutes, event is processed
  3. If git fetch fails, event expires and its ID is added to expired_events
  4. Critical: Once in expired_events, the event will NOT be retried until the 7-day cleanup

The Problem

During initial sync, archive fetches events from source relay. When clone URLs are unreachable or throttled:

  1. Git fetch fails for the event
  2. After 30 minutes, event expires
  3. Event ID added to expired_events (7-day rejection)
  4. Event will NOT be retried even if clone URLs become available

Real-World Impact (relay.ngit.dev Migration)

  • Archive synced 1,858 repos from relay.ngit.dev
  • 1,581 repos (85.1%) stuck in purgatory - git data never fetched
  • Only 277 repos (14.9%) have git data
  • State event coverage: 47.0% instead of expected >99%
  • Migration blocked until purgatory state manually cleared

Root Cause Analysis

  1. External relay unavailability: Archive tried to fetch from external relays (git.shakespeare.diy, etc.) that were unreachable or throttling
  2. No fallback logic: Archive doesn't try alternative clone URLs from the same event
  3. Permanent rejection: Once event expires, it's rejected for 7 days (no retry mechanism)
  4. Silent failure: No prominent logging or metrics for purgatory accumulation

Why This Matters

  1. Initial sync is fragile: If external relays are down during initial sync, repos are stuck for 7 days
  2. No retry mechanism: Once event expires, it won't be retried even if clone URLs become available
  3. No fallback logic: Archive doesn't try alternative clone URLs from the same event
  4. Silent failure: No prominent logging or metrics for purgatory accumulation
  5. Manual intervention required: Only way to retry is to delete purgatory state file and restart
  6. Not scalable: Manual intervention doesn't scale for production deployments

Plan

Phase 1: Retry with Exponential Backoff

  • Don't permanently reject events after first failure
  • Implement retry with increasing delays (1h, 4h, 12h, 24h, 7d)
  • Only permanently reject after multiple failures over 7 days
  • Track retry count and last attempt time per event

Phase 2: Clone URL Fallback Logic

  • Parse all clone URLs from event (not just first one)
  • Try secondary URLs if primary fails
  • Log which URL succeeded for debugging
  • Consider URL priority/preference ordering

Phase 3: Purgatory Metrics and Alerts

  • Expose ngit_purgatory_count metric
  • Expose ngit_expired_events_count metric
  • Expose ngit_purgatory_retry_count metric
  • Log warnings when purgatory accumulates (>10%, >50%, >80% of total events)
  • Add structured logging for purgatory state changes

Phase 4: Configuration Options

  • purgatory_retry_enabled (default: true)
  • purgatory_retry_max_attempts (default: 5)
  • purgatory_retry_backoff_multiplier (default: 2)
  • purgatory_expiry_duration (default: 30min, allow longer for initial sync)
  • Document all options in configuration reference

Phase 5: Initial Sync Mode

  • Detect initial sync (empty database or explicit flag)
  • Use longer purgatory expiry (4 hours instead of 30 minutes)
  • More aggressive retry logic during initial sync
  • Better logging and progress reporting
  • Consider NGIT_INITIAL_SYNC_MODE=true environment variable

Acceptance Criteria

  • Events are retried multiple times before permanent rejection
  • All clone URLs are attempted before giving up
  • Purgatory metrics are exposed (count, expired, retries)
  • Initial sync has special handling (longer expiry, more retries)
  • Manual intervention not required for transient failures
  • Configuration options documented and tested
  • Integration tests cover retry scenarios

Technical Notes

Current Purgatory Implementation

  • Location: src/purgatory/mod.rs
  • Expiry tracking: expired_events: DashMap<EventId, Instant> (line 79)
  • Expiry duration: 30 minutes (hardcoded)
  • Cleanup: 7-day expiry for expired_events entries

Proposed Changes

  1. New fields in purgatory:

    struct RetryState {
        attempts: u32,
        last_attempt: Instant,
        next_retry: Instant,
        failed_urls: Vec<String>,
    }
    retry_state: DashMap<EventId, RetryState>
    
  2. Retry backoff schedule:

    • Attempt 1: Immediate
    • Attempt 2: 1 hour
    • Attempt 3: 4 hours
    • Attempt 4: 12 hours
    • Attempt 5: 24 hours
    • After 5 failures: Permanent rejection (7-day expiry)
  3. Clone URL fallback:

    • Parse clone tags from event
    • Try URLs in order until one succeeds
    • Track which URLs failed for debugging
  • Migration issue: 820a-relay-ngit-dev-migration
  • Purgatory persistence: a6b6-purgatory-surive-reboot (completed)
  • Archive sync results: 1,858 repos, 1,581 in purgatory (85.1%)
  • Workaround: Delete purgatory-state.json and restart relay
  • Investigation: work/state-validation/ARCHIVE-PURGATORY-ANALYSIS.md

Workaround (Until Fixed)

For operators experiencing this issue:

# 1. Stop the relay
systemctl stop ngit-grasp

# 2. Delete purgatory state (forces re-sync of all events)
rm /path/to/data/purgatory-state.json

# 3. Optionally delete rejected events cache
rm /path/to/data/rejected-events-cache.json

# 4. Restart relay
systemctl start ngit-grasp

# 5. Monitor purgatory accumulation
journalctl -u ngit-grasp -f | grep -i purgatory

Note: This workaround requires manual intervention and may cause duplicate processing. The proper fix is implementing the retry mechanism described above.

Progress

2026-01-20 [Session 12:00]

  • Created: Issue documenting purgatory retry problem discovered during relay.ngit.dev migration
  • Context: Archive sync showed 85.1% of repos stuck in purgatory due to unreachable clone URLs
  • Root cause: No retry mechanism after initial git fetch failure
  • Impact: Migration blocked, manual intervention required
  • Next: Prioritize based on migration timeline

Notes

  • Priority: High (blocks migration scenarios)
  • Discovered during: 820a relay.ngit.dev migration
  • Affects: Archive mode, initial sync, any scenario with unreliable external relays
  • Related to but distinct from: a6b6 (persistence) - that issue ensures purgatory survives restarts, this issue addresses retry logic