Merge #86376f5d: fix(relay): preserve bursty client sessions

nostr:nevent1qgsx2lyl2e4zvfadwcvkd9fkrcwczj7mf858hy85mwqclwgut8wpg2spz3mhxue69uhhyetvv9ujumn8d96zuer9wcq3yamnwvaz7tm8d96xummnw3ezucm0d5q3kamnwvaz7tmwva5hgtnyv9hxxmmwwashjer9wchxxmmdqqsgvdm0t5fnvz2q5zf2p6myctssejzp8t6h5cq9fz84qrkv8u784sc3dgkww

PR-Author: DanConwayDev's Agent
nostr:npub1v47f74n2ycn66asev62nv8sas99akj0g0wg0fkup37u3ckwuzs4q7cwtp0

CoverNote:

## Production problem

rust-nostr 0.45 introduced a connection-wide 300-frame-per-minute bucket and closes the WebSocket when it is exhausted. gitnostr.com recorded 5,020 `too many client messages` disconnects between 2026-08-06 10:00 and 2026-08-07 07:55 UTC. This predates and is independent of the subscription-budget ledger.

Correlating relay timestamps with Caddy's real-client access log showed two populations:

- a rapid source sent more than 300 frames and roughly 1.2 MiB in about 1.5 seconds, which should remain bounded;
- legitimate browser sessions from `gitworkshop.dev`, `gittr.space`, `armada.buzz`, and a localhost development origin also crossed the ceiling.

## Approach

Raise only the fixed connection-wide allowance to 6,000 messages/minute (100/second). All activity for a relay is multiplexed through one gitworkshop WebSocket: independent loaders, pagination, reconnects, publishes, AUTH, REQs, and CLOSEs. Batching occurs within individual loaders rather than across the whole connection.

An initial 1,200/minute production candidate was deliberately rejected after it disconnected a legitimate localhost browser session sustaining roughly 44–45 frames/second. The final allowance provides about 2.2x sustained headroom over that observed workload, plus the token bucket's initial burst capacity, while retaining a finite catch-all for malformed or non-operation frames.

The tighter independent quotas remain unchanged: 60 EVENT writes, 120 queries, and 30 authentication events per minute. No new configuration option is added because this non-standard limit cannot be discovered by clients and ngit sync cannot adapt to it.

## Correctness evidence

- A 1,201-frame burst keeps the WebSocket usable and completes a subsequent REQ/EOSE exchange.
- A rapid 6,001-frame burst is still disconnected.
- `nix develop -c cargo test --test relay_message_rate`: 46 passed.
- `nix develop -c cargo test --lib`: 643 passed on the initial candidate; the final change only raises the tested constant and the final targeted suite passed 46/46.
- `nix build .#ngit-grasp`: passed on the initial candidate; deployment built the exact final package.

## Production validation

The exact final tip `a6b5e3a9f66f8bb6e4cc3f23ef62c9be1a9b37f2` was deployed as `/nix/store/bp02cmfviwmf03px52xbjqs8rk4a01ih-ngit-grasp-2.0.0` at 2026-08-07 09:43:58 UTC.

Through 09:57:33 UTC it admitted 874 WebSocket connections with:

- zero `too many client messages` closures;
- zero concurrent-REQ rejections, connection-permit rejections, or subscription-budget failures;
- zero restarts and zero panics;
- 18 expected transient-REQ watchdog CLOSE recoveries;
- approximately 995.9 MB peak memory.

This exact-tip production evidence, together with the failed lower-limit experiment and regression scenarios, supports merge.

## Excluded scope

Per-IP admission fairness and changing rust-nostr's upstream default are separate concerns. A 6,000/minute upstream default is worth proposing as a pragmatic compatibility floor, but this PR does not alter upstream policy.
This commit is contained in:
DanConwayDev
2026-08-07 10:59:28 +01:00
6 changed files with 102 additions and 7 deletions
+3 -1
View File
@@ -18,7 +18,9 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
release, including the upstream NEG-OPEN handling fix and new local-relay
resource hardening. The embedded relay now imposes 500 active REQs per
connection; per-minute connection quotas of 60 event writes, 120 queries,
30 authentication events, and 300 text messages; 20 filters per REQ; 500
30 authentication events, and 6,000 WebSocket messages (raised from
rust-nostr's 300/minute default after it disconnected legitimate bursty
clients in production); 20 filters per REQ; 500
results per filter; 250-byte
subscription IDs; 1 MiB retained subscription state; 10 active negentropy
sessions and 50,000 negentropy items per connection; 5 MiB WebSocket
+2 -2
View File
@@ -32,8 +32,8 @@ These limits prevent individual connections from overwhelming the relay.
The relay advertises the standard `max_subscriptions`, `max_limit`,
`default_limit`, `max_message_length`, and `max_subid_length` NIP-11 fields.
rust-nostr 0.45 also enforces fixed ngit-grasp-selected defaults of 120 queries,
30 authentication events, and 300 text messages per minute; 20 filters per
REQ; 1 MiB subscription state; 10 active negentropy sessions and 50,000
30 authentication events, and 6,000 WebSocket messages per minute; 20 filters
per REQ; 1 MiB subscription state; 10 active negentropy sessions and 50,000
negentropy items per connection; a 5 MiB WebSocket message; and a 10-second
handshake deadline. NIP-11 has no standard fields for most of those controls.
+2 -2
View File
@@ -170,8 +170,8 @@ though they advertise the larger accepted `max_limit`.
#### Admission and rate limits (condensed)
Native rate limiting varies wildly and is invisible to clients. Our own
embedded relay enforces per-connection per-minute quotas (120 queries, 300
text messages, 60 event writes); nostream ships per-IP connection-attempt
embedded relay enforces per-connection per-minute quotas (120 queries, 6,000
WebSocket messages, 60 event writes); nostream ships per-IP connection-attempt
and kind-specific event quotas with EWMA decay; khatru and haven offer
discrete leaky counters that drain over minutes; nostr-rs-relay, relayer,
and rnostr have token-bucket limiters that are disabled by default; chorus
+1 -1
View File
@@ -15,7 +15,7 @@ production admission policy.
| Event writes per minute | 60 | Fixed | No standard field |
| Queries per minute | 120 | Fixed | No standard field |
| Authentication events per minute | 30 | Fixed | No standard field |
| Text messages per minute | 300 | Fixed | No standard field |
| WebSocket messages per minute | 6,000 | Fixed | No standard field |
| WebSocket message size | 5 MiB | Fixed | `max_message_length` |
| Handshake deadline | 10 seconds | Fixed | No standard field |
| Subscription-ID length | 250 bytes | Fixed | `max_subid_length` |
+10 -1
View File
@@ -34,6 +34,15 @@ use crate::nostr::SharedDatabase;
use crate::purgatory::promotion_hooks::NostrPurgatoryPromotionHooks;
use crate::sync::rejected_index::RejectedEventsIndex;
/// Connection-wide frame allowance layered above the operation-specific
/// write, query, and authentication quotas.
///
/// rust-nostr's 300/minute default closed production browser clients during
/// ordinary bursty subscription churn. One hundred frames per second preserves a
/// bounded catch-all for malformed/non-operation traffic while leaving the
/// tighter operation quotas in charge of valid protocol work.
const CLIENT_MESSAGES_PER_MINUTE: u32 = 6_000;
/// NIP-34 Write Policy — admission and routing for GRASP-01 events
///
/// Acts as the top-level admission gate and router. Each incoming event is:
@@ -1014,7 +1023,7 @@ pub async fn create_relay(
})
.queries_per_minute(120)
.auth_events_per_minute(30)
.messages_per_minute(300)
.messages_per_minute(CLIENT_MESSAGES_PER_MINUTE)
.max_websocket_message_size(5 * 1024 * 1024)
.max_event_size(config.relay_max_event_size_bytes)
.websocket_handshake_timeout(Duration::from_secs(10))
+84
View File
@@ -0,0 +1,84 @@
//! Connection-wide relay message-rate regression scenarios.
//!
//! rust-nostr 0.45 added a catch-all 300-frame-per-minute bucket on top of
//! its operation-specific quotas. Production browser clients crossed that
//! threshold during legitimate subscription churn and were disconnected.
mod common;
use std::time::Duration;
use common::TestRelay;
use futures_util::{SinkExt, StreamExt};
use tokio_tungstenite::tungstenite::Message;
const CLIENT_MESSAGES_PER_MINUTE: usize = 6_000;
async fn send_close_frames(
stream: &mut tokio_tungstenite::WebSocketStream<
tokio_tungstenite::MaybeTlsStream<tokio::net::TcpStream>,
>,
count: usize,
) {
for index in 0..count {
stream
.send(Message::Text(
format!(r#"["CLOSE","burst-{index}"]"#).into(),
))
.await
.expect("connection should accept frame");
}
}
#[tokio::test]
async fn legitimate_1201_frame_burst_keeps_connection_usable() {
let relay = TestRelay::start().await;
let (mut stream, _) = tokio_tungstenite::connect_async(relay.url())
.await
.expect("connect to relay");
send_close_frames(&mut stream, 1_201).await;
stream
.send(Message::Text(r#"["REQ","proof",{"kinds":[1]}]"#.into()))
.await
.expect("send proof query");
let response = tokio::time::timeout(Duration::from_secs(5), async {
while let Some(message) = stream.next().await {
let text = message.expect("valid relay response").into_text().unwrap();
if text.contains(r#"["EOSE","proof"]"#) {
return;
}
}
panic!("relay closed before answering proof query");
})
.await;
assert!(response.is_ok(), "relay did not answer proof query in time");
relay.stop().await;
}
#[tokio::test]
async fn catch_all_limit_still_closes_excessive_frame_burst() {
let relay = TestRelay::start().await;
let (mut stream, _) = tokio_tungstenite::connect_async(relay.url())
.await
.expect("connect to relay");
send_close_frames(&mut stream, CLIENT_MESSAGES_PER_MINUTE + 1).await;
let closed = tokio::time::timeout(Duration::from_secs(5), async {
while let Some(message) = stream.next().await {
if message.is_err() {
return;
}
}
})
.await;
assert!(
closed.is_ok(),
"relay did not close excessive sender in time"
);
relay.stop().await;
}