docs(sync): explain invitation dependency recovery

The rejected-event documentation still described removing cold IDs and waiting for a future broad synchronization. That is the failure mode which allowed reciprocal maintainership invitations to remain at 0/N after their full hot-cache events expired.

Document retained exact IDs, recursive maintainer relay discovery, parallel background requests, authorization ordering, retry timing, and success-only removal. Clarify the cache configuration trade-offs and distinguish the legacy invalidation counters from the new non-destructive recovery path.
This commit is contained in:
DanConwayDev
2026-07-26 01:31:30 +01:00
parent 9e9cc926eb
commit 745c21c5cc
8 changed files with 122 additions and 74 deletions
+26 -12
View File
@@ -506,7 +506,7 @@ async fn test_nip01_websocket_connection() {
## Proactive Sync (GRASP-02)
The ngit-grasp relay implements **Proactive Sync of Nostr Eevents**, which synchronizes repository data from external relays listed in 30617 repository announcements. This enables the relay to maintain complete repository graphs even when events are published to other listed relays.
The ngit-grasp relay implements **Proactive Sync of Nostr Events**, which synchronizes repository data from external relays listed in 30617 repository announcements. This enables the relay to maintain complete repository graphs even when events are published to other listed relays.
**Key Features:**
@@ -516,7 +516,7 @@ The ngit-grasp relay implements **Proactive Sync of Nostr Eevents**, which synch
- **Health tracking** with exponential backoff for failing relays
- **Daily sync** with random 23-25h timer to detect state drift
- **Filter consolidation** when count exceeds 70 to prevent subscription explosion
- **Rejected events index** - prevents wasteful re-fetching during negentropy sync
- **Rejected events index** - prevents wasteful broad re-fetching while retaining exact IDs for dependency recovery
**Architecture:**
@@ -542,15 +542,15 @@ The ngit-grasp relay implements **Proactive Sync of Nostr Eevents**, which synch
│ ┌──────────────────────────────────────────────────────┐ │
│ │ Rejected Events Index (Two-Tier) │ │
│ │ ┌────────────────┐ ┌──────────────────────┐ │ │
│ │ │ Hot Cache │───▶│ Cold Index │ │ │
│ │ │ Hot Cache │ │ Cold Index │ │ │
│ │ │ (2 min) │ │ (7 days) │ │ │
│ │ │ Full events │ │ Metadata only │ │ │
│ │ │ Full events │ │ IDs + metadata │ │ │
│ │ └────────────────┘ └──────────────────────┘ │ │
│ └──────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────┘
```
**Source Code:** [`src/sync/`](src/sync/)
**Source Code:** [`src/sync/`](../../src/sync/)
For full design details, see [grasp-02-proactive-sync-v4.md](grasp-02-proactive-sync-v4.md).
@@ -559,7 +559,7 @@ For full design details, see [grasp-02-proactive-sync-v4.md](grasp-02-proactive-
The rejected events index solves two critical problems during sync:
1. **Negentropy sync efficiency**: Prevents repeatedly downloading events that will be rejected again
2. **Race condition resolution**: Enables immediate re-processing when event dependencies are satisfied
2. **Race condition resolution**: Enables immediate hot-cache re-processing or an exact-ID fetch after the full event expires
**Two-Tier Architecture:**
@@ -576,20 +576,34 @@ Event Rejected (e.g., maintainer before owner announcement)
├──▶ Store full event in Hot Cache (2 min expiry)
└──▶ Store metadata in Cold Index (7 day expiry)
Dependency Arrives (e.g., owner announcement accepted)
Reciprocal Invitation Arrives (owner announcement enters purgatory)
│
├──▶ Invalidate from Cold Index
├──▶ Retrieve from Hot Cache (if still available)
└──▶ Re-process immediately (<1 second vs 24 hours)
├──▶ Keep dependency-resolvable IDs in Cold Index
├──▶ Re-process from Hot Cache when the full event remains available
└──▶ Otherwise fetch only the retained IDs from the maintainer relay chain
Exact-ID Recovery
│
├──▶ Connect to relay hints before starting recovery
├──▶ Query all connected maintainer-chain relays in parallel
├──▶ Process announcements before dependent state events
└──▶ Remove IDs only after policy processing succeeds
Negentropy Sync
│
└──▶ Exclude Cold Index IDs from "missing events" calculation
```
Exact-ID recovery runs as bounded background work, so a slow relay does not hold
the main sync-manager lock. Failed or empty requests retain their IDs and are
retried no more than once every 30 seconds. Each relay request has a five-second
timeout.
**Tracked Events:**
- Repository announcements (kind 30617) rejected for not listing this service or maintainer validation failure
- State events (kind 30618) rejected for missing announcements or unauthorized authors
- Repository announcements (kind 30617) rejected for not listing this service or a dependency-resolvable maintainer validation failure
- State events (kind 30618) rejected for missing announcements or dependency-resolvable authorization failures
- Permanently invalid events classified as `Other` remain excluded and are not exact-refetched
**Source Code:** [`src/sync/rejected_index.rs`](../../src/sync/rejected_index.rs)
+16 -6
View File
@@ -796,14 +796,24 @@ Without the rejected events index, negentropy would repeatedly download events t
**Re-Processing on Dependency Arrival:**
When a dependency is satisfied (e.g., owner announcement accepted):
1. Related entries are **invalidated** (removed) from cold index
2. If event still in hot cache → immediate re-processing
3. If event expired from hot cache → will be re-fetched on next sync (now that dependency exists)
When a dependency is satisfied, an unexpired hot-cache event can be
re-processed immediately. During reciprocal maintainership invitation
bootstrap, the owner announcement may still be in purgatory because promotion
needs the inviter's Git data. In that flow:
This prevents permanently excluding events that could become valid after dependencies arrive.
1. Dependency-resolvable entries remain in the cold index until processing succeeds.
2. If the full event is still in the hot cache, it is re-processed immediately.
3. If the full event expired, the sync manager requests only the retained event IDs from connected relays across the recursive maintainer chain.
4. Relay requests run in parallel as bounded background work; announcements are processed before state events that may depend on them.
5. Successful or duplicate events are removed from both tiers. Failed and empty requests retain their IDs for a throttled retry.
See [work/rejected-events-index-summary.md](../../work/rejected-events-index-summary.md) for complete implementation details.
This prevents both failure modes: broad synchronization does not repeatedly
download known-invalid events, while a dependency-resolvable event cannot become
permanently suppressed merely because its full hot-cache copy expired.
See [Architecture: Rejected Events Index](architecture.md#rejected-events-index)
and [`src/sync/rejected_index.rs`](../../src/sync/rejected_index.rs) for the
design and implementation.
---
+15 -6
View File
@@ -440,14 +440,23 @@ State events rejected during authorization are tracked in the rejected events in
- **Reason: MaintainerNotYetValid** - Author not in maintainer set (may become valid later)
- **Reason: Other** - Other validation failures
When a repository announcement is accepted, rejected state events for that repository are:
1. **Invalidated** from cold index (removed from negentropy exclusion)
2. **Retrieved** from hot cache (if still available within 2 minutes)
3. **Re-processed** immediately with new maintainer set
When a repository announcement is accepted, rejected state events still present
in the hot cache are re-processed immediately. Reciprocal invitation
announcements can enter purgatory before their source state and Git data are
available. During that bootstrap flow, rejected state events are:
This enables rapid recovery from race conditions where state events arrive before maintainer announcements.
1. **Retained** in the cold index until policy processing succeeds
2. **Retrieved** from the hot cache and re-processed immediately when the full event is still available
3. **Fetched by exact event ID** from the recursive maintainer relay chain when the hot-cache copy has expired
4. **Removed from both tiers** only after the event is accepted, enters purgatory, or is already stored
See [work/rejected-events-index-summary.md](../../work/rejected-events-index-summary.md) for complete details on rejection tracking and re-processing.
Repository announcements are processed before dependent state events. Network
requests run outside the main sync-manager lock, and unsuccessful requests retain
their IDs for a throttled retry. This enables rapid recovery from arrival-order
races without reopening a broad synchronization query.
See [GRASP-02: Integration with Rejected Events Index](grasp-02-proactive-sync.md#integration-with-rejected-events-index)
for complete details on rejection tracking and re-processing.
---
+21 -23
View File
@@ -235,16 +235,22 @@ All metrics are parameterized by `event_type` label with values "announcement" o
|--------|------|--------|-------------|
| `ngit_rejected_hot_cache_current` | Gauge | event_type | Current number of entries in hot cache |
| `ngit_rejected_cold_index_current` | Gauge | event_type | Current number of entries in cold index |
| `ngit_rejected_hot_cache_hits` | Counter | event_type | Events successfully retrieved from hot cache for re-processing |
| `ngit_rejected_hot_cache_misses` | Counter | event_type | Events expired from hot cache before dependency arrived |
| `ngit_rejected_hot_cache_hits` | Counter | event_type | Events retrieved by the explicit invalidation API |
| `ngit_rejected_hot_cache_misses` | Counter | event_type | Explicit invalidations whose full event had already expired |
| `ngit_rejected_hot_cache_expired` | Counter | event_type | Entries cleaned up from hot cache (2 min expiry) |
| `ngit_rejected_cold_index_expired` | Counter | event_type | Entries cleaned up from cold index (7 day expiry) |
| `ngit_rejected_invalidated` | Counter | event_type | Entries invalidated when dependency was satisfied |
| `ngit_rejected_invalidated` | Counter | event_type | Entries removed by the explicit invalidation API |
The invitation dependency-recovery path is deliberately non-destructive and does
not increment the hit, miss, or invalidation counters above. It retains cold IDs
until processing succeeds. Operators can observe exact-ID attempts in structured
logs containing `Fetched purgatory dependencies by exact event ID`; the current
and expiry gauges still cover entries used by both paths.
### Example Grafana Queries
```promql
# Hot cache efficiency - how often we successfully re-process from cache
# Explicit invalidation API hot-cache efficiency
rate(ngit_rejected_hot_cache_hits_total[5m])
/ (rate(ngit_rejected_hot_cache_hits_total[5m]) + rate(ngit_rejected_hot_cache_misses_total[5m]))
@@ -254,10 +260,10 @@ ngit_rejected_hot_cache_current{event_type="state"}
ngit_rejected_cold_index_current{event_type="announcement"}
ngit_rejected_cold_index_current{event_type="state"}
# Race condition resolution rate - invalidations indicate successful dependency arrival
# Explicit invalidation activity
rate(ngit_rejected_invalidated_total[5m])
# Cache hit ratio over time (higher is better, means dependencies arriving quickly)
# Explicit invalidation cache hit ratio over time
sum(rate(ngit_rejected_hot_cache_hits_total[5m]))
/ sum(rate(ngit_rejected_hot_cache_hits_total[5m]) + rate(ngit_rejected_hot_cache_misses_total[5m]))
```
@@ -265,19 +271,6 @@ sum(rate(ngit_rejected_hot_cache_hits_total[5m]))
### Example Alerts
```yaml
# Alert if hot cache hit rate is too low (suggests timing issues)
- alert: RejectedEventsCacheMissRate
expr: |
sum(rate(ngit_rejected_hot_cache_misses_total[5m]))
/ sum(rate(ngit_rejected_hot_cache_hits_total[5m]) + rate(ngit_rejected_hot_cache_misses_total[5m]))
> 0.8
for: 15m
labels:
severity: warning
annotations:
summary: "High rejected events cache miss rate ({{ $value | humanizePercentage }})"
description: "Most rejected events are expiring before dependencies arrive"
# Alert if cold index growing too large
- alert: RejectedEventsColdIndexSize
expr: ngit_rejected_cold_index_current > 10000
@@ -300,6 +293,7 @@ sum(rate(ngit_rejected_hot_cache_hits_total[5m]))
**Cold Index (7 days):**
- Stores metadata only (event ID, pubkey, identifier, reason)
- Prevents re-downloading during negentropy sync
- Supplies exact IDs when dependency-resolvable full events have left the hot cache
- Cleaned up daily
- Memory: ~1 MB typical
@@ -307,12 +301,16 @@ sum(rate(ngit_rejected_hot_cache_hits_total[5m]))
**Race Condition Resolution:**
When a maintainer announcement arrives before the owner announcement:
1. Maintainer event rejected → hot cache + cold index
2. Owner announcement accepted → invalidate from cold index
3. If still in hot cache → immediate re-processing (<1 second)
4. If expired from hot cache → will be re-fetched on next sync
2. Reciprocal owner announcement enters purgatory → retain the cold ID until recovery succeeds
3. If still in hot cache → immediate policy re-processing
4. If expired from hot cache → exact-ID requests run in parallel across the recursive maintainer relay chain
5. Empty or failed requests keep the ID for a throttled retry
6. Successful processing removes the event from both tiers
**Negentropy Sync Efficiency:**
During sync, cold index IDs are excluded from "missing events" calculation, preventing wasteful re-download of events that will be rejected again.
See [work/rejected-events-index-summary.md](../../work/rejected-events-index-summary.md) for complete implementation details.
See [GRASP-02: Integration with Rejected Events Index](grasp-02-proactive-sync.md#integration-with-rejected-events-index)
for the recovery flow.
+11 -1
View File
@@ -844,7 +844,17 @@ purgatory.extend_announcement_expiry(&owner_pk, &identifier, Duration::from_secs
### 5. Sync Registration (`src/sync/`)
A background timer (`run_purgatory_announcement_sync`, every 5 seconds) ensures purgatory announcements are registered in `RepoSyncIndex` with `SyncLevel::StateOnly`. When an announcement is promoted, the `SelfSubscriber` upgrades it to `SyncLevel::Full`.
A background timer (`run_purgatory_announcement_sync`, every 5 seconds) ensures purgatory announcements are registered in `RepoSyncIndex` with `SyncLevel::StateOnly`. It connects the announcement's relay hints before recovering rejected dependencies, walks accepted announcements to collect the recursive maintainer relay chain, and re-processes any dependency events still in the hot cache.
If a dependency's full event has expired, the timer starts a background exact-ID
request against all connected relays in that chain. Requests run in parallel and
do not hold the sync-manager lock. Announcements are applied before state events;
IDs remain in the cold index through empty or failed requests and are removed only
after successful policy processing. Production retries are limited to once every
30 seconds, with a five-second timeout per relay request.
When an announcement is promoted, the `SelfSubscriber` upgrades it to
`SyncLevel::Full`.
---
+9 -5
View File
@@ -454,10 +454,11 @@ NGIT_REJECTED_HOT_CACHE_DURATION_SECS=300
**Notes:**
- Hot cache stores full event objects for immediate re-processing when dependencies arrive
- Events expire from hot cache after this duration and move to cold index
- Shorter durations reduce memory usage but may miss dependency arrivals
- Longer durations increase memory but improve race condition resolution
- Rejected events are inserted into the hot cache and cold index at the same time
- The hot cache stores full event objects for immediate re-processing when dependencies arrive
- After the full event expires, dependency-resolvable announcement and state IDs remain in the cold index and can be fetched directly from maintainer-chain relays during reciprocal invitation bootstrap
- Shorter durations reduce memory usage but may add an exact-ID relay round trip and its associated latency
- Longer durations increase memory usage and make immediate, network-free recovery more likely
- Memory impact: ~200 KB typical, ~20 MB worst case
---
@@ -486,8 +487,11 @@ NGIT_REJECTED_COLD_INDEX_EXPIRY_SECS=1209600
- Cold index stores only metadata (event ID, pubkey, identifier, rejection reason)
- Prevents re-downloading rejected events during negentropy sync
- Retains the exact IDs required to recover dependency-resolvable events after their hot-cache copies expire
- Successful recovery removes an ID; failed requests leave it indexed for a throttled retry
- Entries automatically cleaned up daily
- Longer durations prevent more wasteful re-fetching but use slightly more memory
- The retention window must be long enough to cover expected invitation and maintainer-chain bootstrap delays
- Longer durations preserve more recovery opportunities and prevent more wasteful broad re-fetching, but use slightly more memory
- Memory impact: ~1 MB typical
---
+17 -15
View File
@@ -9,7 +9,7 @@
//!
//! 2. **Cold Index (Tier 2)**: Stores metadata only for 7 days
//! - Prevents repeated downloads of rejected events
//! - Enables invalidation when dependencies change
//! - Retains exact IDs for dependency recovery after full events expire
//! - Typical memory: ~1 MB
//!
//! # Problem Solved
@@ -18,16 +18,17 @@
//!
//! ```text
//! 00:00 - Maintainer announcement rejected → Event discarded
//! 00:02 - Owner announcement accepted (lists maintainer) → Want to re-process
//! 00:02 - ❌ Maintainer announcement GONE → Must wait 24h for next sync
//! 00:03 - Owner announcement accepted (lists maintainer) → Want to re-process
//! 00:03 - ❌ Maintainer announcement GONE → Completed historic sync will not fetch it
//! ```
//!
//! With the two-tier system:
//!
//! ```text
//! 00:00 - Maintainer announcement rejected → Stored in hot cache + cold index
//! 00:02 - Owner announcement accepted → Invalidate + get from hot cache
//! 00:02 - ✅ Re-process immediately → Accepted in <1 second
//! 00:03 - Full event expired, but its cold-index ID remains available
//! 00:03 - Exact-ID request recovers the event from the maintainer relay chain
//! 00:03 - ✅ Remove from both tiers only after policy processing succeeds
//! ```
//!
//! # Architecture
@@ -40,18 +41,19 @@
//! │ - Auto-expires after 2 minutes │
//! │ - Memory: ~200 KB typical, ~20 MB worst case │
//! └─────────────────────────────────────────────────────────────┘
//! │
//! │ After 2 minutes
//! ▼
//! +
//! ┌─────────────────────────────────────────────────────────────┐
//! │ Tier 2: Cold Index (7 days) │
//! │ - Stores METADATA only (event_id, pubkey, identifier) │
//! │ - Prevents repeated downloads │
//! │ - Enables invalidation │
//! │ - Enables targeted recovery after the full event expires │
//! │ - Memory: ~1 MB typical │
//! └─────────────────────────────────────────────────────────────┘
//! ```
//!
//! Rejected events enter both tiers at the same time. Expiry removes only the
//! full hot-cache copy; the cold metadata remains until success or cold expiry.
//!
//! # Usage
//!
//! ```rust,ignore
@@ -72,17 +74,17 @@
//! RejectionReason::DoesNotListService,
//! );
//!
//! // Later, when owner announcement accepted...
//! let (removed, hot_events) = index.invalidate_and_get(
//! // Later, when the owner announcement arrives...
//! let (event_ids, hot_events) = index.dependency_candidates(
//! &maintainer_pubkey,
//! "my-repo",
//! Some(EventType::Announcement),
//! );
//!
//! // Re-process events from hot cache immediately
//! for event in hot_events {
//! process_event(&event).await;
//! }
//! // Re-process `hot_events` immediately. For candidate IDs without a full
//! // hot-cache event, request that exact event from the recursive maintainer
//! // relay chain. Keep each entry indexed until policy processing succeeds,
//! // then call `index.remove(&event_id)`.
//! ```
use nostr_sdk::prelude::{Event, EventId, PublicKey};
+7 -6
View File
@@ -5,15 +5,16 @@
//!
//! ## Test design
//!
//! Announcements now require git data before they are released from purgatory and
//! served to other relays. The hot-cache re-processing path we want to exercise is:
//! Announcements require git data before they are released from purgatory and
//! served to other relays. The tests exercise both dependency-recovery paths:
//!
//! relay_b syncs maintainer announcement from relay_a
//! → write policy rejects it (no owner announcement in DB yet)
//! → event stored in hot cache
//! owner git push to relay_b promotes owner announcement from purgatory
//! → our new code calls rejected_events_index.invalidate_and_get()
//! → maintainer announcement re-processed and accepted
//! → event stored in hot cache and cold index
//! reciprocal owner announcement reaches relay_b
//! → hot copy is re-processed immediately when available
//! → expired hot copy is fetched by its retained cold-index event ID
//! → maintainer announcement supplies the source clone and git data syncs
//!
//! To guarantee the maintainer announcements arrive at relay_b *before* the owner
//! git push, relay_b is started with relay_a as its bootstrap relay. That way