mirror of
https://relay.ngit.dev/npub15qydau2hjma6ngxkl2cyar74wzyjshvl65za5k5rl69264ar2exs5cyejr/ngit-grasp.git
synced 2026-10-05 15:08:24 +00:00
docs(sync): explain invitation dependency recovery
The rejected-event documentation still described removing cold IDs and waiting for a future broad synchronization. That is the failure mode which allowed reciprocal maintainership invitations to remain at 0/N after their full hot-cache events expired. Document retained exact IDs, recursive maintainer relay discovery, parallel background requests, authorization ordering, retry timing, and success-only removal. Clarify the cache configuration trade-offs and distinguish the legacy invalidation counters from the new non-destructive recovery path.
This commit is contained in:
@@ -506,7 +506,7 @@ async fn test_nip01_websocket_connection() {
|
||||
|
||||
## Proactive Sync (GRASP-02)
|
||||
|
||||
The ngit-grasp relay implements **Proactive Sync of Nostr Eevents**, which synchronizes repository data from external relays listed in 30617 repository announcements. This enables the relay to maintain complete repository graphs even when events are published to other listed relays.
|
||||
The ngit-grasp relay implements **Proactive Sync of Nostr Events**, which synchronizes repository data from external relays listed in 30617 repository announcements. This enables the relay to maintain complete repository graphs even when events are published to other listed relays.
|
||||
|
||||
**Key Features:**
|
||||
|
||||
@@ -516,7 +516,7 @@ The ngit-grasp relay implements **Proactive Sync of Nostr Eevents**, which synch
|
||||
- **Health tracking** with exponential backoff for failing relays
|
||||
- **Daily sync** with random 23-25h timer to detect state drift
|
||||
- **Filter consolidation** when count exceeds 70 to prevent subscription explosion
|
||||
- **Rejected events index** - prevents wasteful re-fetching during negentropy sync
|
||||
- **Rejected events index** - prevents wasteful broad re-fetching while retaining exact IDs for dependency recovery
|
||||
|
||||
**Architecture:**
|
||||
|
||||
@@ -542,15 +542,15 @@ The ngit-grasp relay implements **Proactive Sync of Nostr Eevents**, which synch
|
||||
│ ┌──────────────────────────────────────────────────────┐ │
|
||||
│ │ Rejected Events Index (Two-Tier) │ │
|
||||
│ │ ┌────────────────┐ ┌──────────────────────┐ │ │
|
||||
│ │ │ Hot Cache │───▶│ Cold Index │ │ │
|
||||
│ │ │ Hot Cache │ │ Cold Index │ │ │
|
||||
│ │ │ (2 min) │ │ (7 days) │ │ │
|
||||
│ │ │ Full events │ │ Metadata only │ │ │
|
||||
│ │ │ Full events │ │ IDs + metadata │ │ │
|
||||
│ │ └────────────────┘ └──────────────────────┘ │ │
|
||||
│ └──────────────────────────────────────────────────────┘ │
|
||||
└─────────────────────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
**Source Code:** [`src/sync/`](src/sync/)
|
||||
**Source Code:** [`src/sync/`](../../src/sync/)
|
||||
|
||||
For full design details, see [grasp-02-proactive-sync-v4.md](grasp-02-proactive-sync-v4.md).
|
||||
|
||||
@@ -559,7 +559,7 @@ For full design details, see [grasp-02-proactive-sync-v4.md](grasp-02-proactive-
|
||||
The rejected events index solves two critical problems during sync:
|
||||
|
||||
1. **Negentropy sync efficiency**: Prevents repeatedly downloading events that will be rejected again
|
||||
2. **Race condition resolution**: Enables immediate re-processing when event dependencies are satisfied
|
||||
2. **Race condition resolution**: Enables immediate hot-cache re-processing or an exact-ID fetch after the full event expires
|
||||
|
||||
**Two-Tier Architecture:**
|
||||
|
||||
@@ -576,20 +576,34 @@ Event Rejected (e.g., maintainer before owner announcement)
|
||||
├──▶ Store full event in Hot Cache (2 min expiry)
|
||||
└──▶ Store metadata in Cold Index (7 day expiry)
|
||||
|
||||
Dependency Arrives (e.g., owner announcement accepted)
|
||||
Reciprocal Invitation Arrives (owner announcement enters purgatory)
|
||||
│
|
||||
├──▶ Invalidate from Cold Index
|
||||
├──▶ Retrieve from Hot Cache (if still available)
|
||||
└──▶ Re-process immediately (<1 second vs 24 hours)
|
||||
├──▶ Keep dependency-resolvable IDs in Cold Index
|
||||
├──▶ Re-process from Hot Cache when the full event remains available
|
||||
└──▶ Otherwise fetch only the retained IDs from the maintainer relay chain
|
||||
|
||||
Exact-ID Recovery
|
||||
│
|
||||
├──▶ Connect to relay hints before starting recovery
|
||||
├──▶ Query all connected maintainer-chain relays in parallel
|
||||
├──▶ Process announcements before dependent state events
|
||||
└──▶ Remove IDs only after policy processing succeeds
|
||||
|
||||
Negentropy Sync
|
||||
│
|
||||
└──▶ Exclude Cold Index IDs from "missing events" calculation
|
||||
```
|
||||
|
||||
Exact-ID recovery runs as bounded background work, so a slow relay does not hold
|
||||
the main sync-manager lock. Failed or empty requests retain their IDs and are
|
||||
retried no more than once every 30 seconds. Each relay request has a five-second
|
||||
timeout.
|
||||
|
||||
**Tracked Events:**
|
||||
- Repository announcements (kind 30617) rejected for not listing this service or maintainer validation failure
|
||||
- State events (kind 30618) rejected for missing announcements or unauthorized authors
|
||||
|
||||
- Repository announcements (kind 30617) rejected for not listing this service or a dependency-resolvable maintainer validation failure
|
||||
- State events (kind 30618) rejected for missing announcements or dependency-resolvable authorization failures
|
||||
- Permanently invalid events classified as `Other` remain excluded and are not exact-refetched
|
||||
|
||||
**Source Code:** [`src/sync/rejected_index.rs`](../../src/sync/rejected_index.rs)
|
||||
|
||||
|
||||
@@ -796,14 +796,24 @@ Without the rejected events index, negentropy would repeatedly download events t
|
||||
|
||||
**Re-Processing on Dependency Arrival:**
|
||||
|
||||
When a dependency is satisfied (e.g., owner announcement accepted):
|
||||
1. Related entries are **invalidated** (removed) from cold index
|
||||
2. If event still in hot cache → immediate re-processing
|
||||
3. If event expired from hot cache → will be re-fetched on next sync (now that dependency exists)
|
||||
When a dependency is satisfied, an unexpired hot-cache event can be
|
||||
re-processed immediately. During reciprocal maintainership invitation
|
||||
bootstrap, the owner announcement may still be in purgatory because promotion
|
||||
needs the inviter's Git data. In that flow:
|
||||
|
||||
This prevents permanently excluding events that could become valid after dependencies arrive.
|
||||
1. Dependency-resolvable entries remain in the cold index until processing succeeds.
|
||||
2. If the full event is still in the hot cache, it is re-processed immediately.
|
||||
3. If the full event expired, the sync manager requests only the retained event IDs from connected relays across the recursive maintainer chain.
|
||||
4. Relay requests run in parallel as bounded background work; announcements are processed before state events that may depend on them.
|
||||
5. Successful or duplicate events are removed from both tiers. Failed and empty requests retain their IDs for a throttled retry.
|
||||
|
||||
See [work/rejected-events-index-summary.md](../../work/rejected-events-index-summary.md) for complete implementation details.
|
||||
This prevents both failure modes: broad synchronization does not repeatedly
|
||||
download known-invalid events, while a dependency-resolvable event cannot become
|
||||
permanently suppressed merely because its full hot-cache copy expired.
|
||||
|
||||
See [Architecture: Rejected Events Index](architecture.md#rejected-events-index)
|
||||
and [`src/sync/rejected_index.rs`](../../src/sync/rejected_index.rs) for the
|
||||
design and implementation.
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -440,14 +440,23 @@ State events rejected during authorization are tracked in the rejected events in
|
||||
- **Reason: MaintainerNotYetValid** - Author not in maintainer set (may become valid later)
|
||||
- **Reason: Other** - Other validation failures
|
||||
|
||||
When a repository announcement is accepted, rejected state events for that repository are:
|
||||
1. **Invalidated** from cold index (removed from negentropy exclusion)
|
||||
2. **Retrieved** from hot cache (if still available within 2 minutes)
|
||||
3. **Re-processed** immediately with new maintainer set
|
||||
When a repository announcement is accepted, rejected state events still present
|
||||
in the hot cache are re-processed immediately. Reciprocal invitation
|
||||
announcements can enter purgatory before their source state and Git data are
|
||||
available. During that bootstrap flow, rejected state events are:
|
||||
|
||||
This enables rapid recovery from race conditions where state events arrive before maintainer announcements.
|
||||
1. **Retained** in the cold index until policy processing succeeds
|
||||
2. **Retrieved** from the hot cache and re-processed immediately when the full event is still available
|
||||
3. **Fetched by exact event ID** from the recursive maintainer relay chain when the hot-cache copy has expired
|
||||
4. **Removed from both tiers** only after the event is accepted, enters purgatory, or is already stored
|
||||
|
||||
See [work/rejected-events-index-summary.md](../../work/rejected-events-index-summary.md) for complete details on rejection tracking and re-processing.
|
||||
Repository announcements are processed before dependent state events. Network
|
||||
requests run outside the main sync-manager lock, and unsuccessful requests retain
|
||||
their IDs for a throttled retry. This enables rapid recovery from arrival-order
|
||||
races without reopening a broad synchronization query.
|
||||
|
||||
See [GRASP-02: Integration with Rejected Events Index](grasp-02-proactive-sync.md#integration-with-rejected-events-index)
|
||||
for complete details on rejection tracking and re-processing.
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -235,16 +235,22 @@ All metrics are parameterized by `event_type` label with values "announcement" o
|
||||
|--------|------|--------|-------------|
|
||||
| `ngit_rejected_hot_cache_current` | Gauge | event_type | Current number of entries in hot cache |
|
||||
| `ngit_rejected_cold_index_current` | Gauge | event_type | Current number of entries in cold index |
|
||||
| `ngit_rejected_hot_cache_hits` | Counter | event_type | Events successfully retrieved from hot cache for re-processing |
|
||||
| `ngit_rejected_hot_cache_misses` | Counter | event_type | Events expired from hot cache before dependency arrived |
|
||||
| `ngit_rejected_hot_cache_hits` | Counter | event_type | Events retrieved by the explicit invalidation API |
|
||||
| `ngit_rejected_hot_cache_misses` | Counter | event_type | Explicit invalidations whose full event had already expired |
|
||||
| `ngit_rejected_hot_cache_expired` | Counter | event_type | Entries cleaned up from hot cache (2 min expiry) |
|
||||
| `ngit_rejected_cold_index_expired` | Counter | event_type | Entries cleaned up from cold index (7 day expiry) |
|
||||
| `ngit_rejected_invalidated` | Counter | event_type | Entries invalidated when dependency was satisfied |
|
||||
| `ngit_rejected_invalidated` | Counter | event_type | Entries removed by the explicit invalidation API |
|
||||
|
||||
The invitation dependency-recovery path is deliberately non-destructive and does
|
||||
not increment the hit, miss, or invalidation counters above. It retains cold IDs
|
||||
until processing succeeds. Operators can observe exact-ID attempts in structured
|
||||
logs containing `Fetched purgatory dependencies by exact event ID`; the current
|
||||
and expiry gauges still cover entries used by both paths.
|
||||
|
||||
### Example Grafana Queries
|
||||
|
||||
```promql
|
||||
# Hot cache efficiency - how often we successfully re-process from cache
|
||||
# Explicit invalidation API hot-cache efficiency
|
||||
rate(ngit_rejected_hot_cache_hits_total[5m])
|
||||
/ (rate(ngit_rejected_hot_cache_hits_total[5m]) + rate(ngit_rejected_hot_cache_misses_total[5m]))
|
||||
|
||||
@@ -254,10 +260,10 @@ ngit_rejected_hot_cache_current{event_type="state"}
|
||||
ngit_rejected_cold_index_current{event_type="announcement"}
|
||||
ngit_rejected_cold_index_current{event_type="state"}
|
||||
|
||||
# Race condition resolution rate - invalidations indicate successful dependency arrival
|
||||
# Explicit invalidation activity
|
||||
rate(ngit_rejected_invalidated_total[5m])
|
||||
|
||||
# Cache hit ratio over time (higher is better, means dependencies arriving quickly)
|
||||
# Explicit invalidation cache hit ratio over time
|
||||
sum(rate(ngit_rejected_hot_cache_hits_total[5m]))
|
||||
/ sum(rate(ngit_rejected_hot_cache_hits_total[5m]) + rate(ngit_rejected_hot_cache_misses_total[5m]))
|
||||
```
|
||||
@@ -265,19 +271,6 @@ sum(rate(ngit_rejected_hot_cache_hits_total[5m]))
|
||||
### Example Alerts
|
||||
|
||||
```yaml
|
||||
# Alert if hot cache hit rate is too low (suggests timing issues)
|
||||
- alert: RejectedEventsCacheMissRate
|
||||
expr: |
|
||||
sum(rate(ngit_rejected_hot_cache_misses_total[5m]))
|
||||
/ sum(rate(ngit_rejected_hot_cache_hits_total[5m]) + rate(ngit_rejected_hot_cache_misses_total[5m]))
|
||||
> 0.8
|
||||
for: 15m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "High rejected events cache miss rate ({{ $value | humanizePercentage }})"
|
||||
description: "Most rejected events are expiring before dependencies arrive"
|
||||
|
||||
# Alert if cold index growing too large
|
||||
- alert: RejectedEventsColdIndexSize
|
||||
expr: ngit_rejected_cold_index_current > 10000
|
||||
@@ -300,6 +293,7 @@ sum(rate(ngit_rejected_hot_cache_hits_total[5m]))
|
||||
**Cold Index (7 days):**
|
||||
- Stores metadata only (event ID, pubkey, identifier, reason)
|
||||
- Prevents re-downloading during negentropy sync
|
||||
- Supplies exact IDs when dependency-resolvable full events have left the hot cache
|
||||
- Cleaned up daily
|
||||
- Memory: ~1 MB typical
|
||||
|
||||
@@ -307,12 +301,16 @@ sum(rate(ngit_rejected_hot_cache_hits_total[5m]))
|
||||
|
||||
**Race Condition Resolution:**
|
||||
When a maintainer announcement arrives before the owner announcement:
|
||||
|
||||
1. Maintainer event rejected → hot cache + cold index
|
||||
2. Owner announcement accepted → invalidate from cold index
|
||||
3. If still in hot cache → immediate re-processing (<1 second)
|
||||
4. If expired from hot cache → will be re-fetched on next sync
|
||||
2. Reciprocal owner announcement enters purgatory → retain the cold ID until recovery succeeds
|
||||
3. If still in hot cache → immediate policy re-processing
|
||||
4. If expired from hot cache → exact-ID requests run in parallel across the recursive maintainer relay chain
|
||||
5. Empty or failed requests keep the ID for a throttled retry
|
||||
6. Successful processing removes the event from both tiers
|
||||
|
||||
**Negentropy Sync Efficiency:**
|
||||
During sync, cold index IDs are excluded from "missing events" calculation, preventing wasteful re-download of events that will be rejected again.
|
||||
|
||||
See [work/rejected-events-index-summary.md](../../work/rejected-events-index-summary.md) for complete implementation details.
|
||||
See [GRASP-02: Integration with Rejected Events Index](grasp-02-proactive-sync.md#integration-with-rejected-events-index)
|
||||
for the recovery flow.
|
||||
|
||||
@@ -844,7 +844,17 @@ purgatory.extend_announcement_expiry(&owner_pk, &identifier, Duration::from_secs
|
||||
|
||||
### 5. Sync Registration (`src/sync/`)
|
||||
|
||||
A background timer (`run_purgatory_announcement_sync`, every 5 seconds) ensures purgatory announcements are registered in `RepoSyncIndex` with `SyncLevel::StateOnly`. When an announcement is promoted, the `SelfSubscriber` upgrades it to `SyncLevel::Full`.
|
||||
A background timer (`run_purgatory_announcement_sync`, every 5 seconds) ensures purgatory announcements are registered in `RepoSyncIndex` with `SyncLevel::StateOnly`. It connects the announcement's relay hints before recovering rejected dependencies, walks accepted announcements to collect the recursive maintainer relay chain, and re-processes any dependency events still in the hot cache.
|
||||
|
||||
If a dependency's full event has expired, the timer starts a background exact-ID
|
||||
request against all connected relays in that chain. Requests run in parallel and
|
||||
do not hold the sync-manager lock. Announcements are applied before state events;
|
||||
IDs remain in the cold index through empty or failed requests and are removed only
|
||||
after successful policy processing. Production retries are limited to once every
|
||||
30 seconds, with a five-second timeout per relay request.
|
||||
|
||||
When an announcement is promoted, the `SelfSubscriber` upgrades it to
|
||||
`SyncLevel::Full`.
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -454,10 +454,11 @@ NGIT_REJECTED_HOT_CACHE_DURATION_SECS=300
|
||||
|
||||
**Notes:**
|
||||
|
||||
- Hot cache stores full event objects for immediate re-processing when dependencies arrive
|
||||
- Events expire from hot cache after this duration and move to cold index
|
||||
- Shorter durations reduce memory usage but may miss dependency arrivals
|
||||
- Longer durations increase memory but improve race condition resolution
|
||||
- Rejected events are inserted into the hot cache and cold index at the same time
|
||||
- The hot cache stores full event objects for immediate re-processing when dependencies arrive
|
||||
- After the full event expires, dependency-resolvable announcement and state IDs remain in the cold index and can be fetched directly from maintainer-chain relays during reciprocal invitation bootstrap
|
||||
- Shorter durations reduce memory usage but may add an exact-ID relay round trip and its associated latency
|
||||
- Longer durations increase memory usage and make immediate, network-free recovery more likely
|
||||
- Memory impact: ~200 KB typical, ~20 MB worst case
|
||||
|
||||
---
|
||||
@@ -486,8 +487,11 @@ NGIT_REJECTED_COLD_INDEX_EXPIRY_SECS=1209600
|
||||
|
||||
- Cold index stores only metadata (event ID, pubkey, identifier, rejection reason)
|
||||
- Prevents re-downloading rejected events during negentropy sync
|
||||
- Retains the exact IDs required to recover dependency-resolvable events after their hot-cache copies expire
|
||||
- Successful recovery removes an ID; failed requests leave it indexed for a throttled retry
|
||||
- Entries automatically cleaned up daily
|
||||
- Longer durations prevent more wasteful re-fetching but use slightly more memory
|
||||
- The retention window must be long enough to cover expected invitation and maintainer-chain bootstrap delays
|
||||
- Longer durations preserve more recovery opportunities and prevent more wasteful broad re-fetching, but use slightly more memory
|
||||
- Memory impact: ~1 MB typical
|
||||
|
||||
---
|
||||
|
||||
+17
-15
@@ -9,7 +9,7 @@
|
||||
//!
|
||||
//! 2. **Cold Index (Tier 2)**: Stores metadata only for 7 days
|
||||
//! - Prevents repeated downloads of rejected events
|
||||
//! - Enables invalidation when dependencies change
|
||||
//! - Retains exact IDs for dependency recovery after full events expire
|
||||
//! - Typical memory: ~1 MB
|
||||
//!
|
||||
//! # Problem Solved
|
||||
@@ -18,16 +18,17 @@
|
||||
//!
|
||||
//! ```text
|
||||
//! 00:00 - Maintainer announcement rejected → Event discarded
|
||||
//! 00:02 - Owner announcement accepted (lists maintainer) → Want to re-process
|
||||
//! 00:02 - ❌ Maintainer announcement GONE → Must wait 24h for next sync
|
||||
//! 00:03 - Owner announcement accepted (lists maintainer) → Want to re-process
|
||||
//! 00:03 - ❌ Maintainer announcement GONE → Completed historic sync will not fetch it
|
||||
//! ```
|
||||
//!
|
||||
//! With the two-tier system:
|
||||
//!
|
||||
//! ```text
|
||||
//! 00:00 - Maintainer announcement rejected → Stored in hot cache + cold index
|
||||
//! 00:02 - Owner announcement accepted → Invalidate + get from hot cache
|
||||
//! 00:02 - ✅ Re-process immediately → Accepted in <1 second
|
||||
//! 00:03 - Full event expired, but its cold-index ID remains available
|
||||
//! 00:03 - Exact-ID request recovers the event from the maintainer relay chain
|
||||
//! 00:03 - ✅ Remove from both tiers only after policy processing succeeds
|
||||
//! ```
|
||||
//!
|
||||
//! # Architecture
|
||||
@@ -40,18 +41,19 @@
|
||||
//! │ - Auto-expires after 2 minutes │
|
||||
//! │ - Memory: ~200 KB typical, ~20 MB worst case │
|
||||
//! └─────────────────────────────────────────────────────────────┘
|
||||
//! │
|
||||
//! │ After 2 minutes
|
||||
//! ▼
|
||||
//! +
|
||||
//! ┌─────────────────────────────────────────────────────────────┐
|
||||
//! │ Tier 2: Cold Index (7 days) │
|
||||
//! │ - Stores METADATA only (event_id, pubkey, identifier) │
|
||||
//! │ - Prevents repeated downloads │
|
||||
//! │ - Enables invalidation │
|
||||
//! │ - Enables targeted recovery after the full event expires │
|
||||
//! │ - Memory: ~1 MB typical │
|
||||
//! └─────────────────────────────────────────────────────────────┘
|
||||
//! ```
|
||||
//!
|
||||
//! Rejected events enter both tiers at the same time. Expiry removes only the
|
||||
//! full hot-cache copy; the cold metadata remains until success or cold expiry.
|
||||
//!
|
||||
//! # Usage
|
||||
//!
|
||||
//! ```rust,ignore
|
||||
@@ -72,17 +74,17 @@
|
||||
//! RejectionReason::DoesNotListService,
|
||||
//! );
|
||||
//!
|
||||
//! // Later, when owner announcement accepted...
|
||||
//! let (removed, hot_events) = index.invalidate_and_get(
|
||||
//! // Later, when the owner announcement arrives...
|
||||
//! let (event_ids, hot_events) = index.dependency_candidates(
|
||||
//! &maintainer_pubkey,
|
||||
//! "my-repo",
|
||||
//! Some(EventType::Announcement),
|
||||
//! );
|
||||
//!
|
||||
//! // Re-process events from hot cache immediately
|
||||
//! for event in hot_events {
|
||||
//! process_event(&event).await;
|
||||
//! }
|
||||
//! // Re-process `hot_events` immediately. For candidate IDs without a full
|
||||
//! // hot-cache event, request that exact event from the recursive maintainer
|
||||
//! // relay chain. Keep each entry indexed until policy processing succeeds,
|
||||
//! // then call `index.remove(&event_id)`.
|
||||
//! ```
|
||||
|
||||
use nostr_sdk::prelude::{Event, EventId, PublicKey};
|
||||
|
||||
@@ -5,15 +5,16 @@
|
||||
//!
|
||||
//! ## Test design
|
||||
//!
|
||||
//! Announcements now require git data before they are released from purgatory and
|
||||
//! served to other relays. The hot-cache re-processing path we want to exercise is:
|
||||
//! Announcements require git data before they are released from purgatory and
|
||||
//! served to other relays. The tests exercise both dependency-recovery paths:
|
||||
//!
|
||||
//! relay_b syncs maintainer announcement from relay_a
|
||||
//! → write policy rejects it (no owner announcement in DB yet)
|
||||
//! → event stored in hot cache
|
||||
//! owner git push to relay_b promotes owner announcement from purgatory
|
||||
//! → our new code calls rejected_events_index.invalidate_and_get()
|
||||
//! → maintainer announcement re-processed and accepted
|
||||
//! → event stored in hot cache and cold index
|
||||
//! reciprocal owner announcement reaches relay_b
|
||||
//! → hot copy is re-processed immediately when available
|
||||
//! → expired hot copy is fetched by its retained cold-index event ID
|
||||
//! → maintainer announcement supplies the source clone and git data syncs
|
||||
//!
|
||||
//! To guarantee the maintainer announcements arrive at relay_b *before* the owner
|
||||
//! git push, relay_b is started with relay_a as its bootstrap relay. That way
|
||||
|
||||
Reference in New Issue
Block a user