From 745c21c5ccea0ab8a533c6d8d1318196deb09993 Mon Sep 17 00:00:00 2001 From: DanConwayDev Date: Sun, 26 Jul 2026 01:31:30 +0100 Subject: [PATCH] docs(sync): explain invitation dependency recovery The rejected-event documentation still described removing cold IDs and waiting for a future broad synchronization. That is the failure mode which allowed reciprocal maintainership invitations to remain at 0/N after their full hot-cache events expired. Document retained exact IDs, recursive maintainer relay discovery, parallel background requests, authorization ordering, retry timing, and success-only removal. Clarify the cache configuration trade-offs and distinguish the legacy invalidation counters from the new non-destructive recovery path. --- docs/explanation/architecture.md | 38 ++++++++++++------ docs/explanation/grasp-02-proactive-sync.md | 22 ++++++++--- docs/explanation/inline-authorization.md | 21 +++++++--- docs/explanation/monitoring.md | 44 ++++++++++----------- docs/explanation/purgatory-design.md | 12 +++++- docs/reference/configuration.md | 14 ++++--- src/sync/rejected_index.rs | 32 ++++++++------- tests/sync/maintainer_reprocessing.rs | 13 +++--- 8 files changed, 122 insertions(+), 74 deletions(-) diff --git a/docs/explanation/architecture.md b/docs/explanation/architecture.md index 38edc8b..386381e 100644 --- a/docs/explanation/architecture.md +++ b/docs/explanation/architecture.md @@ -506,7 +506,7 @@ async fn test_nip01_websocket_connection() { ## Proactive Sync (GRASP-02) -The ngit-grasp relay implements **Proactive Sync of Nostr Eevents**, which synchronizes repository data from external relays listed in 30617 repository announcements. This enables the relay to maintain complete repository graphs even when events are published to other listed relays. +The ngit-grasp relay implements **Proactive Sync of Nostr Events**, which synchronizes repository data from external relays listed in 30617 repository announcements. This enables the relay to maintain complete repository graphs even when events are published to other listed relays. **Key Features:** @@ -516,7 +516,7 @@ The ngit-grasp relay implements **Proactive Sync of Nostr Eevents**, which synch - **Health tracking** with exponential backoff for failing relays - **Daily sync** with random 23-25h timer to detect state drift - **Filter consolidation** when count exceeds 70 to prevent subscription explosion -- **Rejected events index** - prevents wasteful re-fetching during negentropy sync +- **Rejected events index** - prevents wasteful broad re-fetching while retaining exact IDs for dependency recovery **Architecture:** @@ -542,15 +542,15 @@ The ngit-grasp relay implements **Proactive Sync of Nostr Eevents**, which synch │ ┌──────────────────────────────────────────────────────┐ │ │ │ Rejected Events Index (Two-Tier) │ │ │ │ ┌────────────────┐ ┌──────────────────────┐ │ │ -│ │ │ Hot Cache │───▶│ Cold Index │ │ │ +│ │ │ Hot Cache │ │ Cold Index │ │ │ │ │ │ (2 min) │ │ (7 days) │ │ │ -│ │ │ Full events │ │ Metadata only │ │ │ +│ │ │ Full events │ │ IDs + metadata │ │ │ │ │ └────────────────┘ └──────────────────────┘ │ │ │ └──────────────────────────────────────────────────────┘ │ └─────────────────────────────────────────────────────────────┘ ``` -**Source Code:** [`src/sync/`](src/sync/) +**Source Code:** [`src/sync/`](../../src/sync/) For full design details, see [grasp-02-proactive-sync-v4.md](grasp-02-proactive-sync-v4.md). @@ -559,7 +559,7 @@ For full design details, see [grasp-02-proactive-sync-v4.md](grasp-02-proactive- The rejected events index solves two critical problems during sync: 1. **Negentropy sync efficiency**: Prevents repeatedly downloading events that will be rejected again -2. **Race condition resolution**: Enables immediate re-processing when event dependencies are satisfied +2. **Race condition resolution**: Enables immediate hot-cache re-processing or an exact-ID fetch after the full event expires **Two-Tier Architecture:** @@ -576,20 +576,34 @@ Event Rejected (e.g., maintainer before owner announcement) ├──▶ Store full event in Hot Cache (2 min expiry) └──▶ Store metadata in Cold Index (7 day expiry) -Dependency Arrives (e.g., owner announcement accepted) +Reciprocal Invitation Arrives (owner announcement enters purgatory) │ - ├──▶ Invalidate from Cold Index - ├──▶ Retrieve from Hot Cache (if still available) - └──▶ Re-process immediately (<1 second vs 24 hours) + ├──▶ Keep dependency-resolvable IDs in Cold Index + ├──▶ Re-process from Hot Cache when the full event remains available + └──▶ Otherwise fetch only the retained IDs from the maintainer relay chain + +Exact-ID Recovery + │ + ├──▶ Connect to relay hints before starting recovery + ├──▶ Query all connected maintainer-chain relays in parallel + ├──▶ Process announcements before dependent state events + └──▶ Remove IDs only after policy processing succeeds Negentropy Sync │ └──▶ Exclude Cold Index IDs from "missing events" calculation ``` +Exact-ID recovery runs as bounded background work, so a slow relay does not hold +the main sync-manager lock. Failed or empty requests retain their IDs and are +retried no more than once every 30 seconds. Each relay request has a five-second +timeout. + **Tracked Events:** -- Repository announcements (kind 30617) rejected for not listing this service or maintainer validation failure -- State events (kind 30618) rejected for missing announcements or unauthorized authors + +- Repository announcements (kind 30617) rejected for not listing this service or a dependency-resolvable maintainer validation failure +- State events (kind 30618) rejected for missing announcements or dependency-resolvable authorization failures +- Permanently invalid events classified as `Other` remain excluded and are not exact-refetched **Source Code:** [`src/sync/rejected_index.rs`](../../src/sync/rejected_index.rs) diff --git a/docs/explanation/grasp-02-proactive-sync.md b/docs/explanation/grasp-02-proactive-sync.md index 6696e27..ef74375 100644 --- a/docs/explanation/grasp-02-proactive-sync.md +++ b/docs/explanation/grasp-02-proactive-sync.md @@ -796,14 +796,24 @@ Without the rejected events index, negentropy would repeatedly download events t **Re-Processing on Dependency Arrival:** -When a dependency is satisfied (e.g., owner announcement accepted): -1. Related entries are **invalidated** (removed) from cold index -2. If event still in hot cache → immediate re-processing -3. If event expired from hot cache → will be re-fetched on next sync (now that dependency exists) +When a dependency is satisfied, an unexpired hot-cache event can be +re-processed immediately. During reciprocal maintainership invitation +bootstrap, the owner announcement may still be in purgatory because promotion +needs the inviter's Git data. In that flow: -This prevents permanently excluding events that could become valid after dependencies arrive. +1. Dependency-resolvable entries remain in the cold index until processing succeeds. +2. If the full event is still in the hot cache, it is re-processed immediately. +3. If the full event expired, the sync manager requests only the retained event IDs from connected relays across the recursive maintainer chain. +4. Relay requests run in parallel as bounded background work; announcements are processed before state events that may depend on them. +5. Successful or duplicate events are removed from both tiers. Failed and empty requests retain their IDs for a throttled retry. -See [work/rejected-events-index-summary.md](../../work/rejected-events-index-summary.md) for complete implementation details. +This prevents both failure modes: broad synchronization does not repeatedly +download known-invalid events, while a dependency-resolvable event cannot become +permanently suppressed merely because its full hot-cache copy expired. + +See [Architecture: Rejected Events Index](architecture.md#rejected-events-index) +and [`src/sync/rejected_index.rs`](../../src/sync/rejected_index.rs) for the +design and implementation. --- diff --git a/docs/explanation/inline-authorization.md b/docs/explanation/inline-authorization.md index 80bd98f..a15168f 100644 --- a/docs/explanation/inline-authorization.md +++ b/docs/explanation/inline-authorization.md @@ -440,14 +440,23 @@ State events rejected during authorization are tracked in the rejected events in - **Reason: MaintainerNotYetValid** - Author not in maintainer set (may become valid later) - **Reason: Other** - Other validation failures -When a repository announcement is accepted, rejected state events for that repository are: -1. **Invalidated** from cold index (removed from negentropy exclusion) -2. **Retrieved** from hot cache (if still available within 2 minutes) -3. **Re-processed** immediately with new maintainer set +When a repository announcement is accepted, rejected state events still present +in the hot cache are re-processed immediately. Reciprocal invitation +announcements can enter purgatory before their source state and Git data are +available. During that bootstrap flow, rejected state events are: -This enables rapid recovery from race conditions where state events arrive before maintainer announcements. +1. **Retained** in the cold index until policy processing succeeds +2. **Retrieved** from the hot cache and re-processed immediately when the full event is still available +3. **Fetched by exact event ID** from the recursive maintainer relay chain when the hot-cache copy has expired +4. **Removed from both tiers** only after the event is accepted, enters purgatory, or is already stored -See [work/rejected-events-index-summary.md](../../work/rejected-events-index-summary.md) for complete details on rejection tracking and re-processing. +Repository announcements are processed before dependent state events. Network +requests run outside the main sync-manager lock, and unsuccessful requests retain +their IDs for a throttled retry. This enables rapid recovery from arrival-order +races without reopening a broad synchronization query. + +See [GRASP-02: Integration with Rejected Events Index](grasp-02-proactive-sync.md#integration-with-rejected-events-index) +for complete details on rejection tracking and re-processing. --- diff --git a/docs/explanation/monitoring.md b/docs/explanation/monitoring.md index 82dd9ca..70954f4 100644 --- a/docs/explanation/monitoring.md +++ b/docs/explanation/monitoring.md @@ -235,16 +235,22 @@ All metrics are parameterized by `event_type` label with values "announcement" o |--------|------|--------|-------------| | `ngit_rejected_hot_cache_current` | Gauge | event_type | Current number of entries in hot cache | | `ngit_rejected_cold_index_current` | Gauge | event_type | Current number of entries in cold index | -| `ngit_rejected_hot_cache_hits` | Counter | event_type | Events successfully retrieved from hot cache for re-processing | -| `ngit_rejected_hot_cache_misses` | Counter | event_type | Events expired from hot cache before dependency arrived | +| `ngit_rejected_hot_cache_hits` | Counter | event_type | Events retrieved by the explicit invalidation API | +| `ngit_rejected_hot_cache_misses` | Counter | event_type | Explicit invalidations whose full event had already expired | | `ngit_rejected_hot_cache_expired` | Counter | event_type | Entries cleaned up from hot cache (2 min expiry) | | `ngit_rejected_cold_index_expired` | Counter | event_type | Entries cleaned up from cold index (7 day expiry) | -| `ngit_rejected_invalidated` | Counter | event_type | Entries invalidated when dependency was satisfied | +| `ngit_rejected_invalidated` | Counter | event_type | Entries removed by the explicit invalidation API | + +The invitation dependency-recovery path is deliberately non-destructive and does +not increment the hit, miss, or invalidation counters above. It retains cold IDs +until processing succeeds. Operators can observe exact-ID attempts in structured +logs containing `Fetched purgatory dependencies by exact event ID`; the current +and expiry gauges still cover entries used by both paths. ### Example Grafana Queries ```promql -# Hot cache efficiency - how often we successfully re-process from cache +# Explicit invalidation API hot-cache efficiency rate(ngit_rejected_hot_cache_hits_total[5m]) / (rate(ngit_rejected_hot_cache_hits_total[5m]) + rate(ngit_rejected_hot_cache_misses_total[5m])) @@ -254,10 +260,10 @@ ngit_rejected_hot_cache_current{event_type="state"} ngit_rejected_cold_index_current{event_type="announcement"} ngit_rejected_cold_index_current{event_type="state"} -# Race condition resolution rate - invalidations indicate successful dependency arrival +# Explicit invalidation activity rate(ngit_rejected_invalidated_total[5m]) -# Cache hit ratio over time (higher is better, means dependencies arriving quickly) +# Explicit invalidation cache hit ratio over time sum(rate(ngit_rejected_hot_cache_hits_total[5m])) / sum(rate(ngit_rejected_hot_cache_hits_total[5m]) + rate(ngit_rejected_hot_cache_misses_total[5m])) ``` @@ -265,19 +271,6 @@ sum(rate(ngit_rejected_hot_cache_hits_total[5m])) ### Example Alerts ```yaml -# Alert if hot cache hit rate is too low (suggests timing issues) -- alert: RejectedEventsCacheMissRate - expr: | - sum(rate(ngit_rejected_hot_cache_misses_total[5m])) - / sum(rate(ngit_rejected_hot_cache_hits_total[5m]) + rate(ngit_rejected_hot_cache_misses_total[5m])) - > 0.8 - for: 15m - labels: - severity: warning - annotations: - summary: "High rejected events cache miss rate ({{ $value | humanizePercentage }})" - description: "Most rejected events are expiring before dependencies arrive" - # Alert if cold index growing too large - alert: RejectedEventsColdIndexSize expr: ngit_rejected_cold_index_current > 10000 @@ -300,6 +293,7 @@ sum(rate(ngit_rejected_hot_cache_hits_total[5m])) **Cold Index (7 days):** - Stores metadata only (event ID, pubkey, identifier, reason) - Prevents re-downloading during negentropy sync +- Supplies exact IDs when dependency-resolvable full events have left the hot cache - Cleaned up daily - Memory: ~1 MB typical @@ -307,12 +301,16 @@ sum(rate(ngit_rejected_hot_cache_hits_total[5m])) **Race Condition Resolution:** When a maintainer announcement arrives before the owner announcement: + 1. Maintainer event rejected → hot cache + cold index -2. Owner announcement accepted → invalidate from cold index -3. If still in hot cache → immediate re-processing (<1 second) -4. If expired from hot cache → will be re-fetched on next sync +2. Reciprocal owner announcement enters purgatory → retain the cold ID until recovery succeeds +3. If still in hot cache → immediate policy re-processing +4. If expired from hot cache → exact-ID requests run in parallel across the recursive maintainer relay chain +5. Empty or failed requests keep the ID for a throttled retry +6. Successful processing removes the event from both tiers **Negentropy Sync Efficiency:** During sync, cold index IDs are excluded from "missing events" calculation, preventing wasteful re-download of events that will be rejected again. -See [work/rejected-events-index-summary.md](../../work/rejected-events-index-summary.md) for complete implementation details. +See [GRASP-02: Integration with Rejected Events Index](grasp-02-proactive-sync.md#integration-with-rejected-events-index) +for the recovery flow. diff --git a/docs/explanation/purgatory-design.md b/docs/explanation/purgatory-design.md index 8e7d75c..ca64a98 100644 --- a/docs/explanation/purgatory-design.md +++ b/docs/explanation/purgatory-design.md @@ -844,7 +844,17 @@ purgatory.extend_announcement_expiry(&owner_pk, &identifier, Duration::from_secs ### 5. Sync Registration (`src/sync/`) -A background timer (`run_purgatory_announcement_sync`, every 5 seconds) ensures purgatory announcements are registered in `RepoSyncIndex` with `SyncLevel::StateOnly`. When an announcement is promoted, the `SelfSubscriber` upgrades it to `SyncLevel::Full`. +A background timer (`run_purgatory_announcement_sync`, every 5 seconds) ensures purgatory announcements are registered in `RepoSyncIndex` with `SyncLevel::StateOnly`. It connects the announcement's relay hints before recovering rejected dependencies, walks accepted announcements to collect the recursive maintainer relay chain, and re-processes any dependency events still in the hot cache. + +If a dependency's full event has expired, the timer starts a background exact-ID +request against all connected relays in that chain. Requests run in parallel and +do not hold the sync-manager lock. Announcements are applied before state events; +IDs remain in the cold index through empty or failed requests and are removed only +after successful policy processing. Production retries are limited to once every +30 seconds, with a five-second timeout per relay request. + +When an announcement is promoted, the `SelfSubscriber` upgrades it to +`SyncLevel::Full`. --- diff --git a/docs/reference/configuration.md b/docs/reference/configuration.md index 777d02f..4266d5d 100644 --- a/docs/reference/configuration.md +++ b/docs/reference/configuration.md @@ -454,10 +454,11 @@ NGIT_REJECTED_HOT_CACHE_DURATION_SECS=300 **Notes:** -- Hot cache stores full event objects for immediate re-processing when dependencies arrive -- Events expire from hot cache after this duration and move to cold index -- Shorter durations reduce memory usage but may miss dependency arrivals -- Longer durations increase memory but improve race condition resolution +- Rejected events are inserted into the hot cache and cold index at the same time +- The hot cache stores full event objects for immediate re-processing when dependencies arrive +- After the full event expires, dependency-resolvable announcement and state IDs remain in the cold index and can be fetched directly from maintainer-chain relays during reciprocal invitation bootstrap +- Shorter durations reduce memory usage but may add an exact-ID relay round trip and its associated latency +- Longer durations increase memory usage and make immediate, network-free recovery more likely - Memory impact: ~200 KB typical, ~20 MB worst case --- @@ -486,8 +487,11 @@ NGIT_REJECTED_COLD_INDEX_EXPIRY_SECS=1209600 - Cold index stores only metadata (event ID, pubkey, identifier, rejection reason) - Prevents re-downloading rejected events during negentropy sync +- Retains the exact IDs required to recover dependency-resolvable events after their hot-cache copies expire +- Successful recovery removes an ID; failed requests leave it indexed for a throttled retry - Entries automatically cleaned up daily -- Longer durations prevent more wasteful re-fetching but use slightly more memory +- The retention window must be long enough to cover expected invitation and maintainer-chain bootstrap delays +- Longer durations preserve more recovery opportunities and prevent more wasteful broad re-fetching, but use slightly more memory - Memory impact: ~1 MB typical --- diff --git a/src/sync/rejected_index.rs b/src/sync/rejected_index.rs index 9a11b19..092212b 100644 --- a/src/sync/rejected_index.rs +++ b/src/sync/rejected_index.rs @@ -9,7 +9,7 @@ //! //! 2. **Cold Index (Tier 2)**: Stores metadata only for 7 days //! - Prevents repeated downloads of rejected events -//! - Enables invalidation when dependencies change +//! - Retains exact IDs for dependency recovery after full events expire //! - Typical memory: ~1 MB //! //! # Problem Solved @@ -18,16 +18,17 @@ //! //! ```text //! 00:00 - Maintainer announcement rejected → Event discarded -//! 00:02 - Owner announcement accepted (lists maintainer) → Want to re-process -//! 00:02 - ❌ Maintainer announcement GONE → Must wait 24h for next sync +//! 00:03 - Owner announcement accepted (lists maintainer) → Want to re-process +//! 00:03 - ❌ Maintainer announcement GONE → Completed historic sync will not fetch it //! ``` //! //! With the two-tier system: //! //! ```text //! 00:00 - Maintainer announcement rejected → Stored in hot cache + cold index -//! 00:02 - Owner announcement accepted → Invalidate + get from hot cache -//! 00:02 - ✅ Re-process immediately → Accepted in <1 second +//! 00:03 - Full event expired, but its cold-index ID remains available +//! 00:03 - Exact-ID request recovers the event from the maintainer relay chain +//! 00:03 - ✅ Remove from both tiers only after policy processing succeeds //! ``` //! //! # Architecture @@ -40,18 +41,19 @@ //! │ - Auto-expires after 2 minutes │ //! │ - Memory: ~200 KB typical, ~20 MB worst case │ //! └─────────────────────────────────────────────────────────────┘ -//! │ -//! │ After 2 minutes -//! ▼ +//! + //! ┌─────────────────────────────────────────────────────────────┐ //! │ Tier 2: Cold Index (7 days) │ //! │ - Stores METADATA only (event_id, pubkey, identifier) │ //! │ - Prevents repeated downloads │ -//! │ - Enables invalidation │ +//! │ - Enables targeted recovery after the full event expires │ //! │ - Memory: ~1 MB typical │ //! └─────────────────────────────────────────────────────────────┘ //! ``` //! +//! Rejected events enter both tiers at the same time. Expiry removes only the +//! full hot-cache copy; the cold metadata remains until success or cold expiry. +//! //! # Usage //! //! ```rust,ignore @@ -72,17 +74,17 @@ //! RejectionReason::DoesNotListService, //! ); //! -//! // Later, when owner announcement accepted... -//! let (removed, hot_events) = index.invalidate_and_get( +//! // Later, when the owner announcement arrives... +//! let (event_ids, hot_events) = index.dependency_candidates( //! &maintainer_pubkey, //! "my-repo", //! Some(EventType::Announcement), //! ); //! -//! // Re-process events from hot cache immediately -//! for event in hot_events { -//! process_event(&event).await; -//! } +//! // Re-process `hot_events` immediately. For candidate IDs without a full +//! // hot-cache event, request that exact event from the recursive maintainer +//! // relay chain. Keep each entry indexed until policy processing succeeds, +//! // then call `index.remove(&event_id)`. //! ``` use nostr_sdk::prelude::{Event, EventId, PublicKey}; diff --git a/tests/sync/maintainer_reprocessing.rs b/tests/sync/maintainer_reprocessing.rs index a0f3ac8..4cd72bc 100644 --- a/tests/sync/maintainer_reprocessing.rs +++ b/tests/sync/maintainer_reprocessing.rs @@ -5,15 +5,16 @@ //! //! ## Test design //! -//! Announcements now require git data before they are released from purgatory and -//! served to other relays. The hot-cache re-processing path we want to exercise is: +//! Announcements require git data before they are released from purgatory and +//! served to other relays. The tests exercise both dependency-recovery paths: //! //! relay_b syncs maintainer announcement from relay_a //! → write policy rejects it (no owner announcement in DB yet) -//! → event stored in hot cache -//! owner git push to relay_b promotes owner announcement from purgatory -//! → our new code calls rejected_events_index.invalidate_and_get() -//! → maintainer announcement re-processed and accepted +//! → event stored in hot cache and cold index +//! reciprocal owner announcement reaches relay_b +//! → hot copy is re-processed immediately when available +//! → expired hot copy is fetched by its retained cold-index event ID +//! → maintainer announcement supplies the source clone and git data syncs //! //! To guarantee the maintainer announcements arrive at relay_b *before* the owner //! git push, relay_b is started with relay_a as its bootstrap relay. That way