diff --git a/1ce6-purgatory-retry-mechanism.md b/1ce6-purgatory-retry-mechanism.md new file mode 100644 index 0000000..41761a1 --- /dev/null +++ b/1ce6-purgatory-retry-mechanism.md @@ -0,0 +1,174 @@ +# Purgatory retry mechanism fails during initial sync when clone URLs are unreachable + +**ID:** 1ce6 + +## Problem + +During initial sync (archive/migration scenarios), ngit-grasp's purgatory mechanism can permanently reject events that should be retried, causing significant data loss. + +### How Purgatory Works + +1. Events arrive at the relay and go to purgatory (30-minute expiry) while git data is fetched +2. If git fetch succeeds within 30 minutes, event is processed +3. If git fetch fails, event expires and its ID is added to `expired_events` +4. **Critical:** Once in `expired_events`, the event will NOT be retried until the 7-day cleanup + +### The Problem + +During initial sync, archive fetches events from source relay. When clone URLs are unreachable or throttled: + +1. Git fetch fails for the event +2. After 30 minutes, event expires +3. Event ID added to `expired_events` (7-day rejection) +4. **Event will NOT be retried even if clone URLs become available** + +### Real-World Impact (relay.ngit.dev Migration) + +- Archive synced 1,858 repos from relay.ngit.dev +- **1,581 repos (85.1%) stuck in purgatory** - git data never fetched +- Only 277 repos (14.9%) have git data +- State event coverage: 47.0% instead of expected >99% +- Migration blocked until purgatory state manually cleared + +### Root Cause Analysis + +1. **External relay unavailability:** Archive tried to fetch from external relays (git.shakespeare.diy, etc.) that were unreachable or throttling +2. **No fallback logic:** Archive doesn't try alternative clone URLs from the same event +3. **Permanent rejection:** Once event expires, it's rejected for 7 days (no retry mechanism) +4. **Silent failure:** No prominent logging or metrics for purgatory accumulation + +## Why This Matters + +1. **Initial sync is fragile:** If external relays are down during initial sync, repos are stuck for 7 days +2. **No retry mechanism:** Once event expires, it won't be retried even if clone URLs become available +3. **No fallback logic:** Archive doesn't try alternative clone URLs from the same event +4. **Silent failure:** No prominent logging or metrics for purgatory accumulation +5. **Manual intervention required:** Only way to retry is to delete purgatory state file and restart +6. **Not scalable:** Manual intervention doesn't scale for production deployments + +## Plan + +### Phase 1: Retry with Exponential Backoff +- [ ] Don't permanently reject events after first failure +- [ ] Implement retry with increasing delays (1h, 4h, 12h, 24h, 7d) +- [ ] Only permanently reject after multiple failures over 7 days +- [ ] Track retry count and last attempt time per event + +### Phase 2: Clone URL Fallback Logic +- [ ] Parse all clone URLs from event (not just first one) +- [ ] Try secondary URLs if primary fails +- [ ] Log which URL succeeded for debugging +- [ ] Consider URL priority/preference ordering + +### Phase 3: Purgatory Metrics and Alerts +- [ ] Expose `ngit_purgatory_count` metric +- [ ] Expose `ngit_expired_events_count` metric +- [ ] Expose `ngit_purgatory_retry_count` metric +- [ ] Log warnings when purgatory accumulates (>10%, >50%, >80% of total events) +- [ ] Add structured logging for purgatory state changes + +### Phase 4: Configuration Options +- [ ] `purgatory_retry_enabled` (default: true) +- [ ] `purgatory_retry_max_attempts` (default: 5) +- [ ] `purgatory_retry_backoff_multiplier` (default: 2) +- [ ] `purgatory_expiry_duration` (default: 30min, allow longer for initial sync) +- [ ] Document all options in configuration reference + +### Phase 5: Initial Sync Mode +- [ ] Detect initial sync (empty database or explicit flag) +- [ ] Use longer purgatory expiry (4 hours instead of 30 minutes) +- [ ] More aggressive retry logic during initial sync +- [ ] Better logging and progress reporting +- [ ] Consider `NGIT_INITIAL_SYNC_MODE=true` environment variable + +## Acceptance Criteria + +- [ ] Events are retried multiple times before permanent rejection +- [ ] All clone URLs are attempted before giving up +- [ ] Purgatory metrics are exposed (count, expired, retries) +- [ ] Initial sync has special handling (longer expiry, more retries) +- [ ] Manual intervention not required for transient failures +- [ ] Configuration options documented and tested +- [ ] Integration tests cover retry scenarios + +## Technical Notes + +### Current Purgatory Implementation + +- **Location:** `src/purgatory/mod.rs` +- **Expiry tracking:** `expired_events: DashMap` (line 79) +- **Expiry duration:** 30 minutes (hardcoded) +- **Cleanup:** 7-day expiry for `expired_events` entries + +### Proposed Changes + +1. **New fields in purgatory:** + ```rust + struct RetryState { + attempts: u32, + last_attempt: Instant, + next_retry: Instant, + failed_urls: Vec, + } + retry_state: DashMap + ``` + +2. **Retry backoff schedule:** + - Attempt 1: Immediate + - Attempt 2: 1 hour + - Attempt 3: 4 hours + - Attempt 4: 12 hours + - Attempt 5: 24 hours + - After 5 failures: Permanent rejection (7-day expiry) + +3. **Clone URL fallback:** + - Parse `clone` tags from event + - Try URLs in order until one succeeds + - Track which URLs failed for debugging + +## Related Context + +- **Migration issue:** 820a-relay-ngit-dev-migration +- **Purgatory persistence:** a6b6-purgatory-surive-reboot (completed) +- **Archive sync results:** 1,858 repos, 1,581 in purgatory (85.1%) +- **Workaround:** Delete `purgatory-state.json` and restart relay +- **Investigation:** work/state-validation/ARCHIVE-PURGATORY-ANALYSIS.md + +## Workaround (Until Fixed) + +For operators experiencing this issue: + +```bash +# 1. Stop the relay +systemctl stop ngit-grasp + +# 2. Delete purgatory state (forces re-sync of all events) +rm /path/to/data/purgatory-state.json + +# 3. Optionally delete rejected events cache +rm /path/to/data/rejected-events-cache.json + +# 4. Restart relay +systemctl start ngit-grasp + +# 5. Monitor purgatory accumulation +journalctl -u ngit-grasp -f | grep -i purgatory +``` + +**Note:** This workaround requires manual intervention and may cause duplicate processing. The proper fix is implementing the retry mechanism described above. + +## Progress + +### 2026-01-20 [Session 12:00] +- Created: Issue documenting purgatory retry problem discovered during relay.ngit.dev migration +- Context: Archive sync showed 85.1% of repos stuck in purgatory due to unreachable clone URLs +- Root cause: No retry mechanism after initial git fetch failure +- Impact: Migration blocked, manual intervention required +- Next: Prioritize based on migration timeline + +## Notes + +- Priority: High (blocks migration scenarios) +- Discovered during: 820a relay.ngit.dev migration +- Affects: Archive mode, initial sync, any scenario with unreliable external relays +- Related to but distinct from: a6b6 (persistence) - that issue ensures purgatory survives restarts, this issue addresses retry logic