mirror of
https://relay.ngit.dev/npub15qydau2hjma6ngxkl2cyar74wzyjshvl65za5k5rl69264ar2exs5cyejr/ngit-grasp.git
synced 2026-10-05 15:08:24 +00:00
issue: create 1ce6 - purgatory retry mechanism fails during initial sync
This commit is contained in:
@@ -0,0 +1,174 @@
|
||||
# Purgatory retry mechanism fails during initial sync when clone URLs are unreachable
|
||||
|
||||
**ID:** 1ce6
|
||||
|
||||
## Problem
|
||||
|
||||
During initial sync (archive/migration scenarios), ngit-grasp's purgatory mechanism can permanently reject events that should be retried, causing significant data loss.
|
||||
|
||||
### How Purgatory Works
|
||||
|
||||
1. Events arrive at the relay and go to purgatory (30-minute expiry) while git data is fetched
|
||||
2. If git fetch succeeds within 30 minutes, event is processed
|
||||
3. If git fetch fails, event expires and its ID is added to `expired_events`
|
||||
4. **Critical:** Once in `expired_events`, the event will NOT be retried until the 7-day cleanup
|
||||
|
||||
### The Problem
|
||||
|
||||
During initial sync, archive fetches events from source relay. When clone URLs are unreachable or throttled:
|
||||
|
||||
1. Git fetch fails for the event
|
||||
2. After 30 minutes, event expires
|
||||
3. Event ID added to `expired_events` (7-day rejection)
|
||||
4. **Event will NOT be retried even if clone URLs become available**
|
||||
|
||||
### Real-World Impact (relay.ngit.dev Migration)
|
||||
|
||||
- Archive synced 1,858 repos from relay.ngit.dev
|
||||
- **1,581 repos (85.1%) stuck in purgatory** - git data never fetched
|
||||
- Only 277 repos (14.9%) have git data
|
||||
- State event coverage: 47.0% instead of expected >99%
|
||||
- Migration blocked until purgatory state manually cleared
|
||||
|
||||
### Root Cause Analysis
|
||||
|
||||
1. **External relay unavailability:** Archive tried to fetch from external relays (git.shakespeare.diy, etc.) that were unreachable or throttling
|
||||
2. **No fallback logic:** Archive doesn't try alternative clone URLs from the same event
|
||||
3. **Permanent rejection:** Once event expires, it's rejected for 7 days (no retry mechanism)
|
||||
4. **Silent failure:** No prominent logging or metrics for purgatory accumulation
|
||||
|
||||
## Why This Matters
|
||||
|
||||
1. **Initial sync is fragile:** If external relays are down during initial sync, repos are stuck for 7 days
|
||||
2. **No retry mechanism:** Once event expires, it won't be retried even if clone URLs become available
|
||||
3. **No fallback logic:** Archive doesn't try alternative clone URLs from the same event
|
||||
4. **Silent failure:** No prominent logging or metrics for purgatory accumulation
|
||||
5. **Manual intervention required:** Only way to retry is to delete purgatory state file and restart
|
||||
6. **Not scalable:** Manual intervention doesn't scale for production deployments
|
||||
|
||||
## Plan
|
||||
|
||||
### Phase 1: Retry with Exponential Backoff
|
||||
- [ ] Don't permanently reject events after first failure
|
||||
- [ ] Implement retry with increasing delays (1h, 4h, 12h, 24h, 7d)
|
||||
- [ ] Only permanently reject after multiple failures over 7 days
|
||||
- [ ] Track retry count and last attempt time per event
|
||||
|
||||
### Phase 2: Clone URL Fallback Logic
|
||||
- [ ] Parse all clone URLs from event (not just first one)
|
||||
- [ ] Try secondary URLs if primary fails
|
||||
- [ ] Log which URL succeeded for debugging
|
||||
- [ ] Consider URL priority/preference ordering
|
||||
|
||||
### Phase 3: Purgatory Metrics and Alerts
|
||||
- [ ] Expose `ngit_purgatory_count` metric
|
||||
- [ ] Expose `ngit_expired_events_count` metric
|
||||
- [ ] Expose `ngit_purgatory_retry_count` metric
|
||||
- [ ] Log warnings when purgatory accumulates (>10%, >50%, >80% of total events)
|
||||
- [ ] Add structured logging for purgatory state changes
|
||||
|
||||
### Phase 4: Configuration Options
|
||||
- [ ] `purgatory_retry_enabled` (default: true)
|
||||
- [ ] `purgatory_retry_max_attempts` (default: 5)
|
||||
- [ ] `purgatory_retry_backoff_multiplier` (default: 2)
|
||||
- [ ] `purgatory_expiry_duration` (default: 30min, allow longer for initial sync)
|
||||
- [ ] Document all options in configuration reference
|
||||
|
||||
### Phase 5: Initial Sync Mode
|
||||
- [ ] Detect initial sync (empty database or explicit flag)
|
||||
- [ ] Use longer purgatory expiry (4 hours instead of 30 minutes)
|
||||
- [ ] More aggressive retry logic during initial sync
|
||||
- [ ] Better logging and progress reporting
|
||||
- [ ] Consider `NGIT_INITIAL_SYNC_MODE=true` environment variable
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
- [ ] Events are retried multiple times before permanent rejection
|
||||
- [ ] All clone URLs are attempted before giving up
|
||||
- [ ] Purgatory metrics are exposed (count, expired, retries)
|
||||
- [ ] Initial sync has special handling (longer expiry, more retries)
|
||||
- [ ] Manual intervention not required for transient failures
|
||||
- [ ] Configuration options documented and tested
|
||||
- [ ] Integration tests cover retry scenarios
|
||||
|
||||
## Technical Notes
|
||||
|
||||
### Current Purgatory Implementation
|
||||
|
||||
- **Location:** `src/purgatory/mod.rs`
|
||||
- **Expiry tracking:** `expired_events: DashMap<EventId, Instant>` (line 79)
|
||||
- **Expiry duration:** 30 minutes (hardcoded)
|
||||
- **Cleanup:** 7-day expiry for `expired_events` entries
|
||||
|
||||
### Proposed Changes
|
||||
|
||||
1. **New fields in purgatory:**
|
||||
```rust
|
||||
struct RetryState {
|
||||
attempts: u32,
|
||||
last_attempt: Instant,
|
||||
next_retry: Instant,
|
||||
failed_urls: Vec<String>,
|
||||
}
|
||||
retry_state: DashMap<EventId, RetryState>
|
||||
```
|
||||
|
||||
2. **Retry backoff schedule:**
|
||||
- Attempt 1: Immediate
|
||||
- Attempt 2: 1 hour
|
||||
- Attempt 3: 4 hours
|
||||
- Attempt 4: 12 hours
|
||||
- Attempt 5: 24 hours
|
||||
- After 5 failures: Permanent rejection (7-day expiry)
|
||||
|
||||
3. **Clone URL fallback:**
|
||||
- Parse `clone` tags from event
|
||||
- Try URLs in order until one succeeds
|
||||
- Track which URLs failed for debugging
|
||||
|
||||
## Related Context
|
||||
|
||||
- **Migration issue:** 820a-relay-ngit-dev-migration
|
||||
- **Purgatory persistence:** a6b6-purgatory-surive-reboot (completed)
|
||||
- **Archive sync results:** 1,858 repos, 1,581 in purgatory (85.1%)
|
||||
- **Workaround:** Delete `purgatory-state.json` and restart relay
|
||||
- **Investigation:** work/state-validation/ARCHIVE-PURGATORY-ANALYSIS.md
|
||||
|
||||
## Workaround (Until Fixed)
|
||||
|
||||
For operators experiencing this issue:
|
||||
|
||||
```bash
|
||||
# 1. Stop the relay
|
||||
systemctl stop ngit-grasp
|
||||
|
||||
# 2. Delete purgatory state (forces re-sync of all events)
|
||||
rm /path/to/data/purgatory-state.json
|
||||
|
||||
# 3. Optionally delete rejected events cache
|
||||
rm /path/to/data/rejected-events-cache.json
|
||||
|
||||
# 4. Restart relay
|
||||
systemctl start ngit-grasp
|
||||
|
||||
# 5. Monitor purgatory accumulation
|
||||
journalctl -u ngit-grasp -f | grep -i purgatory
|
||||
```
|
||||
|
||||
**Note:** This workaround requires manual intervention and may cause duplicate processing. The proper fix is implementing the retry mechanism described above.
|
||||
|
||||
## Progress
|
||||
|
||||
### 2026-01-20 [Session 12:00]
|
||||
- Created: Issue documenting purgatory retry problem discovered during relay.ngit.dev migration
|
||||
- Context: Archive sync showed 85.1% of repos stuck in purgatory due to unreachable clone URLs
|
||||
- Root cause: No retry mechanism after initial git fetch failure
|
||||
- Impact: Migration blocked, manual intervention required
|
||||
- Next: Prioritize based on migration timeline
|
||||
|
||||
## Notes
|
||||
|
||||
- Priority: High (blocks migration scenarios)
|
||||
- Discovered during: 820a relay.ngit.dev migration
|
||||
- Affects: Archive mode, initial sync, any scenario with unreliable external relays
|
||||
- Related to but distinct from: a6b6 (persistence) - that issue ensures purgatory survives restarts, this issue addresses retry logic
|
||||
Reference in New Issue
Block a user