fix(relay): admit repository-scale live coverage

Production gitnostr.com repeatedly received CLOSED responses from relay.ngit.dev after reconnect: its 34-filter live set for 924 full repositories, 89 state-only repositories, and 3,748 roots crossed rust-nostr 0.45's newly selected 1 MiB cumulative subscription-state limit. Fourteen representative REQs were accepted and the remainder were refused, silently leaving persistent live coverage incomplete; cooldown recovery only recreated the same impossible set.

Raise the explicitly selected per-connection retained subscription-state allowance to 5 MiB. This remains a finite boundary, matches the largest individual WebSocket message already admitted, and leaves roughly four times the observed working-set headroom without promising the theoretical 500 x 96 KiB maximum. Document the exact serving policy and its lack of NIP-11 negotiation.

A scenario opens 17 persistent sub-96 KiB REQs carrying 34 filters. It failed unchanged 2.1.1 at live-14 with the production CLOSED reason and passes with the new bound. Correctness assumes this allowance is enforced per connection by rust-nostr; client-side adaptation to unknown third-party cumulative byte limits and multi-connection sharding remain excluded because NIP-11 exposes no such capability.

Validation: focused scenario passed; 658 library tests passed; git diff --check passed; nix build .#ngit-grasp passed.
This commit is contained in:
DanConwayDev
2026-08-08 10:05:49 +00:00
parent e9c2edf540
commit c5e52c5a1e
5 changed files with 97 additions and 5 deletions
+8
View File
@@ -7,6 +7,14 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
## [Unreleased]
### Fixed
- Raise retained subscription state per connection from 1 MiB to 5 MiB. A
production 34-filter repository-sync live set reached roughly 1.2 MiB, so
rust-nostr's newly introduced default repeatedly closed part of persistent
coverage. The replacement remains a finite per-connection bound and matches
the relay's maximum admitted WebSocket message size.
## [2.1.1] - 2026-08-08
ngit-grasp 2.1.1 is a patch release improving relay-to-relay sync stability
+5 -3
View File
@@ -33,9 +33,11 @@ The relay advertises the standard `max_subscriptions`, `max_limit`,
`default_limit`, `max_message_length`, and `max_subid_length` NIP-11 fields.
rust-nostr 0.45 also enforces fixed ngit-grasp-selected defaults of 1,200 queries,
30 authentication events, and 6,000 WebSocket messages per minute; 20 filters
per REQ; 1 MiB subscription state; 10 active negentropy sessions and 50,000
negentropy items per connection; a 5 MiB WebSocket message; and a 10-second
handshake deadline. NIP-11 has no standard fields for most of those controls.
per REQ; 5 MiB subscription state (raised from rust-nostr's 1 MiB default so
repository-scale persistent live filters remain admitted); 10 active
negentropy sessions and 50,000 negentropy items per connection; a 5 MiB
WebSocket message; and a 10-second handshake deadline. NIP-11 has no standard
fields for most of those controls.
The query allowance is a temporary 10× override of rust-nostr's newly added
120/minute default because NIP-77 currently charges each SDK-managed `NEG-MSG`
continuation separately; it must be reviewed when upstream revises that
+9 -1
View File
@@ -20,7 +20,7 @@ production admission policy.
| Handshake deadline | 10 seconds | Fixed | No standard field |
| Subscription-ID length | 250 bytes | Fixed | `max_subid_length` |
| Filters per REQ | 20 | Fixed | No standard field |
| Subscription state per connection | 1 MiB | Fixed | No standard field |
| Subscription state per connection | 5 MiB | Fixed | No standard field |
| Active negentropy sessions per connection | 10 | Fixed | No standard field |
| Negentropy items per connection | 50,000 | Fixed | No standard field |
| Negentropy frame | 60,000 bytes | Fixed upstream | No standard field |
@@ -35,6 +35,14 @@ Production history contains a valid NIP-34 patch event of about 149 KiB, so
64 KiB is incompatible with ngit-grasp's purpose. The raised limit remains
bounded and below the 5 MiB WebSocket message ceiling.
The 5 MiB subscription-state allowance raises rust-nostr's 1 MiB default.
Repository sync keeps multiple byte-budgeted filters live: production's
34-filter coverage reached roughly 1.2 MiB and the smaller bound silently left
part of that coverage closed. The new allowance remains finite per connection,
matches the maximum admitted WebSocket message, and provides about four times
the observed working-set headroom. NIP-11 has no field for this cumulative
byte limit, so clients cannot negotiate it.
The 1,200-query allowance temporarily overrides rust-nostr 0.45's newly
introduced 120/minute default. A finite per-connection bound remains one useful
DoS layer, but the upstream default is unusually restrictive for sync clients
+11 -1
View File
@@ -52,6 +52,16 @@ const CLIENT_MESSAGES_PER_MINUTE: u32 = 6_000;
/// this 10× value after upstream separates or otherwise revises that accounting.
const CLIENT_QUERIES_PER_MINUTE: u32 = 1_200;
/// Cumulative serialized REQ state retained for one client connection.
///
/// Repository sync legitimately keeps multiple byte-budgeted live filters
/// open at once. Production's 34-filter coverage reached about 1.2 MiB and
/// was partially rejected by rust-nostr's 1 MiB default. Five MiB remains a
/// finite per-connection allocation boundary while matching the largest
/// individual WebSocket message we already admit and leaving roughly 4x room
/// above the observed working set.
const MAX_SUBSCRIPTION_STATE_BYTES: usize = 5 * 1024 * 1024;
/// NIP-34 Write Policy — admission and routing for GRASP-01 events
///
/// Acts as the top-level admission gate and router. Each incoming event is:
@@ -1038,7 +1048,7 @@ pub async fn create_relay(
.websocket_handshake_timeout(Duration::from_secs(10))
.max_subid_length(250)
.max_filters_per_req(20)
.max_subscription_bytes(1024 * 1024)
.max_subscription_bytes(MAX_SUBSCRIPTION_STATE_BYTES)
.max_negentropy_subscriptions(10)
.max_negentropy_items(50_000)
.max_filter_limit(config.relay_filter_limit)
+64
View File
@@ -0,0 +1,64 @@
//! Connection-wide active-subscription state regression scenarios.
//!
//! Repository sync can legitimately keep many byte-budgeted filters live on
//! one relay connection. The cumulative allowance must accommodate that
//! coverage even though each individual REQ remains conservatively bounded.
mod common;
use std::time::Duration;
use common::TestRelay;
use futures_util::{SinkExt, StreamExt};
use tokio_tungstenite::tungstenite::Message;
#[tokio::test]
async fn repository_scale_live_filter_set_remains_admitted() {
let relay = TestRelay::start().await;
let (mut stream, _) = tokio_tungstenite::connect_async(relay.url())
.await
.expect("connect to relay");
// Two 550-value filters are about 74 KiB. Seventeen persistent REQs model
// the 34-filter live set that production rejected after crossing the old
// 1 MiB cumulative default, while every individual message stays below
// the sync client's 96 KiB REQ budget.
let ids = (0..550)
.map(|index| format!("{index:064x}"))
.collect::<Vec<_>>();
let filter = serde_json::json!({"#e": ids, "limit": 0});
for index in 0..17 {
let request = serde_json::json!(["REQ", format!("live-{index}"), filter, filter]);
assert!(request.to_string().len() < 96 * 1024);
stream
.send(Message::Text(request.to_string().into()))
.await
.expect("send persistent live REQ");
}
let admitted = tokio::time::timeout(Duration::from_secs(10), async {
let mut eose = 0;
while let Some(message) = stream.next().await {
let text = message.expect("valid relay response").into_text().unwrap();
assert!(
!text.contains("active subscriptions exceed max size"),
"repository-scale live coverage was rejected: {text}"
);
if text.starts_with("[\"EOSE\",\"live-") {
eose += 1;
if eose == 17 {
return;
}
}
}
panic!("relay closed before acknowledging every live REQ");
})
.await;
assert!(
admitted.is_ok(),
"relay did not acknowledge live coverage in time"
);
relay.stop().await;
}