Files
ngit-grasp/tests/sync/mod.rs
T
DanConwayDev e942199a55 fix(sync): bound concurrent transient REQ subscriptions per relay connection
Historic REQ+EOSE sync subscribed every byte-budgeted filter group of a
batch in one loop and left them all awaiting EOSE concurrently, as did
the REQ+EOSE fallback, negentropy ID fetches and missing-event retries.
On strfry-family relays these REQs share maxSubsPerConnection with
negentropy views and live subscriptions; nos.lol (budget 20) answered
each gitnostr.com startup with a burst of 'ERROR: too many concurrent
REQs' NOTICEs (6 in one second at the 2026-08-05 06:41 startup; 7,739
over the prior three days).

A per-connection semaphore (5 permits, shared across clones) now gates
every auto-close subscription inside subscribe_filters: the permit is
registered against the subscription id on success and released when the
connection's own event loop sees that subscription's EOSE or CLOSED
frame - before forwarding the notification, so release never depends on
downstream channel consumers. This keeps a plain blocking acquire
deadlock-free even though the SyncManager actor both creates
subscriptions and processes EOSE: bursts pipeline at five in flight,
matching the precedent of the actor already stalling inline for
negentropy batches. Permits are additionally freed on unsubscribe,
disconnect, event-loop termination, and by a 30-second watchdog so a
relay that never answers cannot starve later subscriptions. The
negentropy semaphore stays separate (different lifetimes); live
subscriptions are not gated. Budget: 4 NEG + 5 REQ + 2 margin leaves at
least nine slots of the tightest observed budget (20) for live
subscriptions.

Scheduling is deliberately per-REQ rather than per-core-filter: a
paginating filter chain holds no slot between pages, so queued groups
interleave breadth-first. Relay-visible concurrency is identical either
way and pagination chains have no durable identity across disconnects.

Reproduction: a new proxy fixture mimics the strfry limit (rejecting
REQs beyond 5 with the production NOTICE, exempting limit:0 live
subscriptions, delaying EOSE so REQs provably overlap), and a scenario
test syncs 2500 root events into persistent storage, restarts the relay
- startup recomputes filters from the full index, the shape that bursts
in production - and asserts zero rejections with overlap retained.
Unfixed: 3 REQs rejected (opened 91, peak pinned at the limit). Fixed:
zero rejections, peak <= 5, sync completes. The TestRelay fixture gains
same-port restart support with persistent LMDB storage for this.

Deliberately excluded: the unified budget ledger (NIP-11-aware B/M,
per-query result caps), raising the 300-ID exact-ID chunks, and any
configuration surface for the bound.

Validated with the full test suite (nix develop -c cargo test).
2026-08-05 07:50:45 +00:00

140 lines
4.9 KiB
Rust

//! Proactive Sync Integration Tests
//!
//! This module organizes tests for ngit-grasp's proactive sync functionality.
//! Tests are grouped by sync scenario:
//!
//! - Historic sync (relay syncs from pre-configured bootstrap relay)
//! - Relay discovery (relay discovers other relays from announcement events)
//! - Live sync (events sync in real-time after connection established)
//! - Tag variations (testing different Layer 2/3 tag types: a/A/q, e/E/q)
//! - Catchup sync (events from disconnected period sync on reconnect)
//! - Metrics (Prometheus metrics for sync operations)
//!
//! # Test Files
//!
//! - `historic_sync.rs` - Bootstrap and replay tests (uses `run_sync_test()` helper)
//! - `discovery.rs` - Relay discovery from announcements (manual setup required)
//! - `live_sync.rs` - Real-time sync after connection (manual setup required)
//! - `tag_variations.rs` - Layer 2/3 tag type coverage (manual setup required)
//! - `catchup.rs` - Catchup after disconnect (stub, `#[ignore]`)
//! - `metrics.rs` - Prometheus metrics integration tests
//!
//! # Test Patterns
//!
//! This module uses two main testing approaches, each suited to different scenarios:
//!
//! ## Pattern 1: Helper-Based Tests (Historic Sync)
//!
//! **Use `run_sync_test()` for:**
//! - Verifying historic event sync (events published before relay starts)
//! - Bootstrap and initialization tests
//! - Simple count-based event verification
//! - Single-relay scenarios
//!
//! **Example from `historic_sync.rs`:**
//! ```rust
//! use common::sync_helpers::{run_sync_test, build_layer2_issue_event};
//!
//! #[tokio::test]
//! async fn test_bootstrap_syncs_existing_layer2_events() {
//! let repo_event = /* create repo announcement */;
//! let issue1 = build_layer2_issue_event(&repo_event, "Issue 1");
//! let issue2 = build_layer2_issue_event(&repo_event, "Issue 2");
//!
//! run_sync_test(
//! &[&repo_event], // Bootstrap events
//! &[&issue1, &issue2], // Events to verify
//! 2, // Expected count
//! ).await;
//! }
//! ```
//!
//! **Helper Architecture:**
//! - Publishes all events to bootstrap relay before target relay starts
//! - Automatically starts target relay with bootstrap relay configured
//! - Verifies event counts after sync completes
//! - Handles all relay lifecycle management
//!
//! ## Pattern 2: Manual Setup Tests (Live, Discovery, Tag Variations)
//!
//! **Use manual setup for:**
//! - Live sync (events published *during* relay operation)
//! - Multi-relay coordination (discovery chains)
//! - Detailed event inspection (tag format verification)
//! - Precise timing control
//!
//! **Example from `live_sync.rs`:**
//! ```rust
//! #[tokio::test]
//! async fn test_live_sync_layer2_events() {
//! let bootstrap = TestRelay::start().await;
//! let target = TestRelay::start_with_bootstrap(bootstrap.url()).await;
//!
//! // Publish AFTER relay is running (live sync)
//! let event = build_layer2_issue_event(&repo, "Live Issue");
//! client.publish_event(event).await;
//!
//! // Verify with timing control
//! wait_for_event_on_relay(&target, &event.id, timeout).await;
//! }
//! ```
//!
//! **Example from `discovery.rs`:**
//! ```rust
//! #[tokio::test]
//! async fn test_discovers_layer3_via_layer2() {
//! // Multi-relay orchestration
//! let relay_a = TestRelay::start().await;
//! let relay_b = TestRelay::start_with_sync(None).await;
//!
//! // relay_b receives announcement listing relay_a, discovers and syncs from it
//! }
//! ```
//!
//! **Example from `tag_variations.rs`:**
//! ```rust
//! #[tokio::test]
//! async fn test_layer2_sync_with_uppercase_a_tag() {
//! // Detailed tag format verification
//! let event = build_event_with_uppercase_A();
//!
//! // Custom assertions about tag normalization
//! assert!(synced_event.tags.contains_uppercase_a());
//! }
//! ```
//!
//! ## Why Two Patterns?
//!
//! The `run_sync_test()` helper embodies a specific pattern:
//! ```
//! Setup → Publish Batch → Start Relay → Verify Counts
//! ```
//!
//! This pattern is **incompatible** with tests needing:
//! - Event publication *during* relay operation (live sync)
//! - Multiple relay coordination (discovery)
//! - Detailed event inspection beyond counts (tag variations)
//! - Precise timing control
//!
//! For these scenarios, manual setup provides necessary flexibility.
//!
//! # Shared Imports
//!
//! All sync tests use helpers from `common::sync_helpers`:
//! - `TestClient` - Client with retry logic
//! - `run_sync_test()` - Helper for historic sync tests
//! - Event builders for Layer 2/3 events
//! - `wait_for_event_on_relay()` - Non-panicking assertion helper
//! - `fetch_metrics()` - Prometheus metrics fetching
// Test modules
pub mod historic_recovery;
pub mod historic_sync;
pub mod catchup;
pub mod discovery;
pub mod live_sync;
pub mod maintainer_reprocessing;
pub mod metrics;
pub mod neg_concurrency;
pub mod req_concurrency;
pub mod tag_variations;