issue: create 1ce6 - purgatory retry mechanism fails during initial sync

This commit is contained in:
DanConwayDev
2026-01-20 13:05:55 +00:00
parent 0e525f3c60
commit 25e2edab40
+174
View File
@@ -0,0 +1,174 @@
# Purgatory retry mechanism fails during initial sync when clone URLs are unreachable
**ID:** 1ce6
## Problem
During initial sync (archive/migration scenarios), ngit-grasp's purgatory mechanism can permanently reject events that should be retried, causing significant data loss.
### How Purgatory Works
1. Events arrive at the relay and go to purgatory (30-minute expiry) while git data is fetched
2. If git fetch succeeds within 30 minutes, event is processed
3. If git fetch fails, event expires and its ID is added to `expired_events`
4. **Critical:** Once in `expired_events`, the event will NOT be retried until the 7-day cleanup
### The Problem
During initial sync, archive fetches events from source relay. When clone URLs are unreachable or throttled:
1. Git fetch fails for the event
2. After 30 minutes, event expires
3. Event ID added to `expired_events` (7-day rejection)
4. **Event will NOT be retried even if clone URLs become available**
### Real-World Impact (relay.ngit.dev Migration)
- Archive synced 1,858 repos from relay.ngit.dev
- **1,581 repos (85.1%) stuck in purgatory** - git data never fetched
- Only 277 repos (14.9%) have git data
- State event coverage: 47.0% instead of expected >99%
- Migration blocked until purgatory state manually cleared
### Root Cause Analysis
1. **External relay unavailability:** Archive tried to fetch from external relays (git.shakespeare.diy, etc.) that were unreachable or throttling
2. **No fallback logic:** Archive doesn't try alternative clone URLs from the same event
3. **Permanent rejection:** Once event expires, it's rejected for 7 days (no retry mechanism)
4. **Silent failure:** No prominent logging or metrics for purgatory accumulation
## Why This Matters
1. **Initial sync is fragile:** If external relays are down during initial sync, repos are stuck for 7 days
2. **No retry mechanism:** Once event expires, it won't be retried even if clone URLs become available
3. **No fallback logic:** Archive doesn't try alternative clone URLs from the same event
4. **Silent failure:** No prominent logging or metrics for purgatory accumulation
5. **Manual intervention required:** Only way to retry is to delete purgatory state file and restart
6. **Not scalable:** Manual intervention doesn't scale for production deployments
## Plan
### Phase 1: Retry with Exponential Backoff
- [ ] Don't permanently reject events after first failure
- [ ] Implement retry with increasing delays (1h, 4h, 12h, 24h, 7d)
- [ ] Only permanently reject after multiple failures over 7 days
- [ ] Track retry count and last attempt time per event
### Phase 2: Clone URL Fallback Logic
- [ ] Parse all clone URLs from event (not just first one)
- [ ] Try secondary URLs if primary fails
- [ ] Log which URL succeeded for debugging
- [ ] Consider URL priority/preference ordering
### Phase 3: Purgatory Metrics and Alerts
- [ ] Expose `ngit_purgatory_count` metric
- [ ] Expose `ngit_expired_events_count` metric
- [ ] Expose `ngit_purgatory_retry_count` metric
- [ ] Log warnings when purgatory accumulates (>10%, >50%, >80% of total events)
- [ ] Add structured logging for purgatory state changes
### Phase 4: Configuration Options
- [ ] `purgatory_retry_enabled` (default: true)
- [ ] `purgatory_retry_max_attempts` (default: 5)
- [ ] `purgatory_retry_backoff_multiplier` (default: 2)
- [ ] `purgatory_expiry_duration` (default: 30min, allow longer for initial sync)
- [ ] Document all options in configuration reference
### Phase 5: Initial Sync Mode
- [ ] Detect initial sync (empty database or explicit flag)
- [ ] Use longer purgatory expiry (4 hours instead of 30 minutes)
- [ ] More aggressive retry logic during initial sync
- [ ] Better logging and progress reporting
- [ ] Consider `NGIT_INITIAL_SYNC_MODE=true` environment variable
## Acceptance Criteria
- [ ] Events are retried multiple times before permanent rejection
- [ ] All clone URLs are attempted before giving up
- [ ] Purgatory metrics are exposed (count, expired, retries)
- [ ] Initial sync has special handling (longer expiry, more retries)
- [ ] Manual intervention not required for transient failures
- [ ] Configuration options documented and tested
- [ ] Integration tests cover retry scenarios
## Technical Notes
### Current Purgatory Implementation
- **Location:** `src/purgatory/mod.rs`
- **Expiry tracking:** `expired_events: DashMap<EventId, Instant>` (line 79)
- **Expiry duration:** 30 minutes (hardcoded)
- **Cleanup:** 7-day expiry for `expired_events` entries
### Proposed Changes
1. **New fields in purgatory:**
```rust
struct RetryState {
attempts: u32,
last_attempt: Instant,
next_retry: Instant,
failed_urls: Vec<String>,
}
retry_state: DashMap<EventId, RetryState>
```
2. **Retry backoff schedule:**
- Attempt 1: Immediate
- Attempt 2: 1 hour
- Attempt 3: 4 hours
- Attempt 4: 12 hours
- Attempt 5: 24 hours
- After 5 failures: Permanent rejection (7-day expiry)
3. **Clone URL fallback:**
- Parse `clone` tags from event
- Try URLs in order until one succeeds
- Track which URLs failed for debugging
## Related Context
- **Migration issue:** 820a-relay-ngit-dev-migration
- **Purgatory persistence:** a6b6-purgatory-surive-reboot (completed)
- **Archive sync results:** 1,858 repos, 1,581 in purgatory (85.1%)
- **Workaround:** Delete `purgatory-state.json` and restart relay
- **Investigation:** work/state-validation/ARCHIVE-PURGATORY-ANALYSIS.md
## Workaround (Until Fixed)
For operators experiencing this issue:
```bash
# 1. Stop the relay
systemctl stop ngit-grasp
# 2. Delete purgatory state (forces re-sync of all events)
rm /path/to/data/purgatory-state.json
# 3. Optionally delete rejected events cache
rm /path/to/data/rejected-events-cache.json
# 4. Restart relay
systemctl start ngit-grasp
# 5. Monitor purgatory accumulation
journalctl -u ngit-grasp -f | grep -i purgatory
```
**Note:** This workaround requires manual intervention and may cause duplicate processing. The proper fix is implementing the retry mechanism described above.
## Progress
### 2026-01-20 [Session 12:00]
- Created: Issue documenting purgatory retry problem discovered during relay.ngit.dev migration
- Context: Archive sync showed 85.1% of repos stuck in purgatory due to unreachable clone URLs
- Root cause: No retry mechanism after initial git fetch failure
- Impact: Migration blocked, manual intervention required
- Next: Prioritize based on migration timeline
## Notes
- Priority: High (blocks migration scenarios)
- Discovered during: 820a relay.ngit.dev migration
- Affects: Archive mode, initial sync, any scenario with unreliable external relays
- Related to but distinct from: a6b6 (persistence) - that issue ensures purgatory survives restarts, this issue addresses retry logic