mirror of
https://relay.ngit.dev/npub15qydau2hjma6ngxkl2cyar74wzyjshvl65za5k5rl69264ar2exs5cyejr/ngit-grasp.git
synced 2026-10-05 15:08:24 +00:00
issue: create 1ce6 - purgatory retry mechanism fails during initial sync
This commit is contained in:
@@ -0,0 +1,174 @@
|
|||||||
|
# Purgatory retry mechanism fails during initial sync when clone URLs are unreachable
|
||||||
|
|
||||||
|
**ID:** 1ce6
|
||||||
|
|
||||||
|
## Problem
|
||||||
|
|
||||||
|
During initial sync (archive/migration scenarios), ngit-grasp's purgatory mechanism can permanently reject events that should be retried, causing significant data loss.
|
||||||
|
|
||||||
|
### How Purgatory Works
|
||||||
|
|
||||||
|
1. Events arrive at the relay and go to purgatory (30-minute expiry) while git data is fetched
|
||||||
|
2. If git fetch succeeds within 30 minutes, event is processed
|
||||||
|
3. If git fetch fails, event expires and its ID is added to `expired_events`
|
||||||
|
4. **Critical:** Once in `expired_events`, the event will NOT be retried until the 7-day cleanup
|
||||||
|
|
||||||
|
### The Problem
|
||||||
|
|
||||||
|
During initial sync, archive fetches events from source relay. When clone URLs are unreachable or throttled:
|
||||||
|
|
||||||
|
1. Git fetch fails for the event
|
||||||
|
2. After 30 minutes, event expires
|
||||||
|
3. Event ID added to `expired_events` (7-day rejection)
|
||||||
|
4. **Event will NOT be retried even if clone URLs become available**
|
||||||
|
|
||||||
|
### Real-World Impact (relay.ngit.dev Migration)
|
||||||
|
|
||||||
|
- Archive synced 1,858 repos from relay.ngit.dev
|
||||||
|
- **1,581 repos (85.1%) stuck in purgatory** - git data never fetched
|
||||||
|
- Only 277 repos (14.9%) have git data
|
||||||
|
- State event coverage: 47.0% instead of expected >99%
|
||||||
|
- Migration blocked until purgatory state manually cleared
|
||||||
|
|
||||||
|
### Root Cause Analysis
|
||||||
|
|
||||||
|
1. **External relay unavailability:** Archive tried to fetch from external relays (git.shakespeare.diy, etc.) that were unreachable or throttling
|
||||||
|
2. **No fallback logic:** Archive doesn't try alternative clone URLs from the same event
|
||||||
|
3. **Permanent rejection:** Once event expires, it's rejected for 7 days (no retry mechanism)
|
||||||
|
4. **Silent failure:** No prominent logging or metrics for purgatory accumulation
|
||||||
|
|
||||||
|
## Why This Matters
|
||||||
|
|
||||||
|
1. **Initial sync is fragile:** If external relays are down during initial sync, repos are stuck for 7 days
|
||||||
|
2. **No retry mechanism:** Once event expires, it won't be retried even if clone URLs become available
|
||||||
|
3. **No fallback logic:** Archive doesn't try alternative clone URLs from the same event
|
||||||
|
4. **Silent failure:** No prominent logging or metrics for purgatory accumulation
|
||||||
|
5. **Manual intervention required:** Only way to retry is to delete purgatory state file and restart
|
||||||
|
6. **Not scalable:** Manual intervention doesn't scale for production deployments
|
||||||
|
|
||||||
|
## Plan
|
||||||
|
|
||||||
|
### Phase 1: Retry with Exponential Backoff
|
||||||
|
- [ ] Don't permanently reject events after first failure
|
||||||
|
- [ ] Implement retry with increasing delays (1h, 4h, 12h, 24h, 7d)
|
||||||
|
- [ ] Only permanently reject after multiple failures over 7 days
|
||||||
|
- [ ] Track retry count and last attempt time per event
|
||||||
|
|
||||||
|
### Phase 2: Clone URL Fallback Logic
|
||||||
|
- [ ] Parse all clone URLs from event (not just first one)
|
||||||
|
- [ ] Try secondary URLs if primary fails
|
||||||
|
- [ ] Log which URL succeeded for debugging
|
||||||
|
- [ ] Consider URL priority/preference ordering
|
||||||
|
|
||||||
|
### Phase 3: Purgatory Metrics and Alerts
|
||||||
|
- [ ] Expose `ngit_purgatory_count` metric
|
||||||
|
- [ ] Expose `ngit_expired_events_count` metric
|
||||||
|
- [ ] Expose `ngit_purgatory_retry_count` metric
|
||||||
|
- [ ] Log warnings when purgatory accumulates (>10%, >50%, >80% of total events)
|
||||||
|
- [ ] Add structured logging for purgatory state changes
|
||||||
|
|
||||||
|
### Phase 4: Configuration Options
|
||||||
|
- [ ] `purgatory_retry_enabled` (default: true)
|
||||||
|
- [ ] `purgatory_retry_max_attempts` (default: 5)
|
||||||
|
- [ ] `purgatory_retry_backoff_multiplier` (default: 2)
|
||||||
|
- [ ] `purgatory_expiry_duration` (default: 30min, allow longer for initial sync)
|
||||||
|
- [ ] Document all options in configuration reference
|
||||||
|
|
||||||
|
### Phase 5: Initial Sync Mode
|
||||||
|
- [ ] Detect initial sync (empty database or explicit flag)
|
||||||
|
- [ ] Use longer purgatory expiry (4 hours instead of 30 minutes)
|
||||||
|
- [ ] More aggressive retry logic during initial sync
|
||||||
|
- [ ] Better logging and progress reporting
|
||||||
|
- [ ] Consider `NGIT_INITIAL_SYNC_MODE=true` environment variable
|
||||||
|
|
||||||
|
## Acceptance Criteria
|
||||||
|
|
||||||
|
- [ ] Events are retried multiple times before permanent rejection
|
||||||
|
- [ ] All clone URLs are attempted before giving up
|
||||||
|
- [ ] Purgatory metrics are exposed (count, expired, retries)
|
||||||
|
- [ ] Initial sync has special handling (longer expiry, more retries)
|
||||||
|
- [ ] Manual intervention not required for transient failures
|
||||||
|
- [ ] Configuration options documented and tested
|
||||||
|
- [ ] Integration tests cover retry scenarios
|
||||||
|
|
||||||
|
## Technical Notes
|
||||||
|
|
||||||
|
### Current Purgatory Implementation
|
||||||
|
|
||||||
|
- **Location:** `src/purgatory/mod.rs`
|
||||||
|
- **Expiry tracking:** `expired_events: DashMap<EventId, Instant>` (line 79)
|
||||||
|
- **Expiry duration:** 30 minutes (hardcoded)
|
||||||
|
- **Cleanup:** 7-day expiry for `expired_events` entries
|
||||||
|
|
||||||
|
### Proposed Changes
|
||||||
|
|
||||||
|
1. **New fields in purgatory:**
|
||||||
|
```rust
|
||||||
|
struct RetryState {
|
||||||
|
attempts: u32,
|
||||||
|
last_attempt: Instant,
|
||||||
|
next_retry: Instant,
|
||||||
|
failed_urls: Vec<String>,
|
||||||
|
}
|
||||||
|
retry_state: DashMap<EventId, RetryState>
|
||||||
|
```
|
||||||
|
|
||||||
|
2. **Retry backoff schedule:**
|
||||||
|
- Attempt 1: Immediate
|
||||||
|
- Attempt 2: 1 hour
|
||||||
|
- Attempt 3: 4 hours
|
||||||
|
- Attempt 4: 12 hours
|
||||||
|
- Attempt 5: 24 hours
|
||||||
|
- After 5 failures: Permanent rejection (7-day expiry)
|
||||||
|
|
||||||
|
3. **Clone URL fallback:**
|
||||||
|
- Parse `clone` tags from event
|
||||||
|
- Try URLs in order until one succeeds
|
||||||
|
- Track which URLs failed for debugging
|
||||||
|
|
||||||
|
## Related Context
|
||||||
|
|
||||||
|
- **Migration issue:** 820a-relay-ngit-dev-migration
|
||||||
|
- **Purgatory persistence:** a6b6-purgatory-surive-reboot (completed)
|
||||||
|
- **Archive sync results:** 1,858 repos, 1,581 in purgatory (85.1%)
|
||||||
|
- **Workaround:** Delete `purgatory-state.json` and restart relay
|
||||||
|
- **Investigation:** work/state-validation/ARCHIVE-PURGATORY-ANALYSIS.md
|
||||||
|
|
||||||
|
## Workaround (Until Fixed)
|
||||||
|
|
||||||
|
For operators experiencing this issue:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# 1. Stop the relay
|
||||||
|
systemctl stop ngit-grasp
|
||||||
|
|
||||||
|
# 2. Delete purgatory state (forces re-sync of all events)
|
||||||
|
rm /path/to/data/purgatory-state.json
|
||||||
|
|
||||||
|
# 3. Optionally delete rejected events cache
|
||||||
|
rm /path/to/data/rejected-events-cache.json
|
||||||
|
|
||||||
|
# 4. Restart relay
|
||||||
|
systemctl start ngit-grasp
|
||||||
|
|
||||||
|
# 5. Monitor purgatory accumulation
|
||||||
|
journalctl -u ngit-grasp -f | grep -i purgatory
|
||||||
|
```
|
||||||
|
|
||||||
|
**Note:** This workaround requires manual intervention and may cause duplicate processing. The proper fix is implementing the retry mechanism described above.
|
||||||
|
|
||||||
|
## Progress
|
||||||
|
|
||||||
|
### 2026-01-20 [Session 12:00]
|
||||||
|
- Created: Issue documenting purgatory retry problem discovered during relay.ngit.dev migration
|
||||||
|
- Context: Archive sync showed 85.1% of repos stuck in purgatory due to unreachable clone URLs
|
||||||
|
- Root cause: No retry mechanism after initial git fetch failure
|
||||||
|
- Impact: Migration blocked, manual intervention required
|
||||||
|
- Next: Prioritize based on migration timeline
|
||||||
|
|
||||||
|
## Notes
|
||||||
|
|
||||||
|
- Priority: High (blocks migration scenarios)
|
||||||
|
- Discovered during: 820a relay.ngit.dev migration
|
||||||
|
- Affects: Archive mode, initial sync, any scenario with unreliable external relays
|
||||||
|
- Related to but distinct from: a6b6 (persistence) - that issue ensures purgatory survives restarts, this issue addresses retry logic
|
||||||
Reference in New Issue
Block a user