mirror of
https://relay.ngit.dev/npub15qydau2hjma6ngxkl2cyar74wzyjshvl65za5k5rl69264ar2exs5cyejr/ngit-grasp.git
synced 2026-10-05 15:08:24 +00:00
3.9 KiB
3.9 KiB
Retry Individual OIDs on "not our ref" Errors
ID: 1e2a
Problem
When git fetch requests multiple OIDs and encounters a "not our ref" error, git stops at the first missing OID and doesn't attempt to fetch the remaining OIDs. This means if we request 5 OIDs and the first one is missing, we never try the other 4 - even though they may exist on the remote.
Example scenario:
- Request OIDs: [A, B, C, D, E]
- Remote has: [B, C, D, E] but not A
- Git fails on A and stops
- Result: We never fetch B, C, D, E even though they're available
This is particularly problematic because:
- We may have state events (kind 30617) that are critical for repository state
- We may have PR refs that are less critical but still valuable
- We're wasting opportunities to fetch available data
Current behavior:
- Single OID request: Returns clear error "remote missing only oid requested: {oid}"
- Multi OID request: Logs WARNING and returns "remote missing oid {oid} (BUG: {n} other oids not attempted)"
Plan
-
Phase 1: Implement retry logic for "not our ref" errors
- When fetch fails with "not our ref", extract the missing OID
- Remove missing OID from request list
- Retry fetch with remaining OIDs
- Continue until all OIDs attempted or all are missing
- Track which OIDs were missing vs successfully fetched
-
Phase 2: Integrate with throttling system
- Ensure retries respect domain throttling limits
- Don't count retries as separate requests for throttling purposes
- Consider backoff if multiple OIDs are missing
-
Phase 3: Prioritize state events over PR refs
- Separate OIDs by event kind (30617 state events vs others)
- Fetch state event OIDs first
- Fetch PR ref OIDs second
- This ensures critical repository state is prioritized
-
Phase 4: Add metrics and logging
- Track how often we hit this scenario
- Log which OIDs were missing vs fetched
- Monitor impact on sync performance
Progress
2026-01-27 [Session 17:30]
-
Completed: Basic retry logic implemented in migration-v2 branch
- Retry loop that excludes failed OIDs and retries with remaining OIDs (src/purgatory/sync/context.rs:363-471)
- Debug-level logging for retry operations
- Proper error reporting with URLs for all failure cases
-
Remaining work:
- Throttling integration: Respect domain throttling limits between retries
- Metrics: Track retry attempts, successes, and confirmed missing OIDs
- OID prioritization: Implement tiered fetching strategy:
- First: Latest state events (kind 30617)
- Second: Earlier state events
- Third: PR refs
- After successful fetch, retry to get remaining OIDs that aren't confirmed missing
- Note: Rate limiting retries is better handled through throttling integration rather than separate backoff logic
-
Not needed:
- Partial success handling (git servers don't support partial fetches - they fail atomically)
2026-01-27 [Session]
- Implemented Phase 1: retry logic for "not our ref" errors
- When fetch fails with "not our ref", extracts missing OID and retries with remaining OIDs
- Loop continues until fetch succeeds or no OIDs remain
- Changed logging from WARN to DEBUG for retry operations
- Logs final summary showing fetched vs missing OIDs
- All 386 lib tests pass
2026-01-23 [Session 16:00]
- Created issue to track the bug
- Committed short-term fix that improves error messages and warns about the bug
- Commit:
9fd4350"fix: improve 'not our ref' error messages and warn about multi-OID fetch bug"
Notes
- Related code:
src/purgatory/sync/context.rs:411-448(fetch_oids error handling) - Git behavior confirmed by testing: git stops at first missing OID
- Current warning helps identify when this bug is affecting sync
- Proper fix requires retry loop with OID filtering
- Must consider performance impact of multiple fetch attempts
- Should track metrics to understand how often this occurs in production