mirror of
https://github.com/jmcorgan/fips.git
synced 2026-08-11 09:07:44 +00:00
Optional peer discovery and NAT hole-punching path gated behind a new
`nostr-discovery` cargo feature. Nodes publish signed overlay endpoint
adverts to public Nostr relays, consume peer adverts to populate
fallback dial addresses, and use STUN-assisted UDP hole punching with
NIP-59 gift-wrap offer/answer signaling to establish direct UDP paths
between NATed peers. Once a punched socket is up, it is handed into
the existing FIPS UDP transport and the standard Noise/FMP session
stack takes over unchanged.
The cargo feature is in the default feature set
(`default = ["nostr-discovery"]`) so stock builds include it; a
build that explicitly disables default features (or selects a
feature set without `nostr-discovery`) does not link the nostr /
nostr-sdk crates and does not emit a no-op poll in the tick loop.
Runtime behavior is independently gated by
`node.discovery.nostr.enabled`, which defaults to false; if the
config enables Nostr on a non-feature build, startup logs a
warning and continues without it.
== Cargo feature and dependencies
- New cargo feature `nostr-discovery = ["dep:nostr", "dep:nostr-sdk"]`.
Not in the default feature set.
- New optional Linux-only dependencies: `nostr 0.44` (features: std,
nip59) and `nostr-sdk 0.44`. Gift-wrap unwrap is hand-rolled in
`src/discovery/nostr/signal.rs` rather than relying on the SDK's
rumor-author check, which FIPS sidesteps by trusting `seal.pubkey`
exclusively.
== Wire format
Overlay advert event: `kind 37195`, parameterized replaceable
(NIP-01 application-defined replaceable range 30000-39999), with
`d = "fips-overlay-v1"`. The digits visually spell FIPS (7=F, 1=I,
9=P, 5=S); a relay survey confirmed the kind is unused.
Advert content carries the version tag, endpoint list
(`udp|tcp|tor` + addr), optional signal-relay and stun-server
metadata, and `issuedAt` / `expiresAt` timestamps. Endpoint
`addr: "nat"` is the sentinel that triggers traversal on the peer
side. NIP-40 `expiration` tag bounds staleness on permanent
shutdown. Lifecycle relies on parameterized-replaceable
supersession; the daemon does not emit NIP-09 kind-5 deletes —
strict relays (Damus, Primal) race delete-against-replace and can
silently drop the replacement.
Gift-wrapped signal event: `kind 21059`. Punch packets carry magic
values `PUNCH_MAGIC` / `PUNCH_ACK_MAGIC`, a sequence number, and a
16-byte session hash.
== Discovery surface
- `src/discovery.rs` (always compiled)
- `EstablishedTraversal`: bound UDP socket + selected remote +
peer npub + optional transport name/config tuning overrides.
- `BootstrapHandoffResult`: returned on successful handoff —
allocated transport id, local/remote addrs, peer NodeAddr,
session id.
- `src/discovery/nostr/` (`#![cfg(feature = "nostr-discovery")]`)
- `types.rs`: wire and control types described above. `ADVERT_KIND`
constant. `BootstrapError` enumerates failure modes (disabled,
missing advert, missing NAT endpoint, no usable relays, invalid
advert, invalid npub, signal timeout, punch timeout, replay,
STUN failure, protocol, nostr, io, serde, event-parse).
- `runtime.rs`: `NostrDiscovery` coordinator. Owns the shared
nostr-sdk `Client`, subscribes to advert + signal event kinds,
maintains a bounded advert cache and a bounded seen-sessions
replay set, drains `BootstrapEvent::{Established, Failed}` for
the node to consume, exposes `update_local_advert`,
`request_connect`, `advert_endpoints_for_peer`,
`cached_open_discovery_candidates`, and `shutdown`.
- `signal.rs`: NIP-59 gift-wrap encode/decode. Outbound wraps are
built against per-attempt ephemeral keys; inbound events are
unwrapped against the node identity.
- `stun.rs`: RFC 5389/8489 Binding Request client with
XOR-MAPPED-ADDRESS parsing for both IPv4 and IPv6; used only to
observe the initiator's own reflexive address against its
locally configured STUN list (peer-advertised STUN is
informational, never an egress target).
- `traversal.rs`: per-attempt candidate-pair punch planner.
Allocates a fresh `0.0.0.0:0` UDP socket per attempt, enumerates
LAN-private and ULA interface addresses alongside the STUN
reflexive address, schedules probe/ack exchanges at the
configured interval for the configured duration, and picks the
first candidate pair that authenticates end-to-end.
Strategy ordering is Reflexive↔Reflexive first, then LAN, then
Mixed. The STUN-observed pair is the only candidate that's reliable
across arbitrary network topologies; trying it first prevents the
planner from latching onto a misleading host-candidate path before
the reflexive path gets a chance. There is no catch-all
Local↔Local strategy: a previous design that paired every local
host candidate from one side with every local host candidate from
the other could declare success on a one-way reachable asymmetric
L3 path (corporate VPN, Tailscale subnet route, overlapping private
address space), only for the FMP handshake to stall because the
return path didn't match. The legitimate `Lan` strategy still pairs
candidates that share a subnet.
== Configuration surface
`node.discovery.nostr.*` (`NostrDiscoveryConfig`), all `serde(default)`
with `deny_unknown_fields`:
- `enabled` (default false), `advertise` (default true)
- `advert_relays`, `dm_relays`, `stun_servers`: defaults are
`wss://relay.damus.io`, `wss://nos.lol`, `wss://offchain.pub`
for both relay lists, and Google / Cloudflare / Twilio for STUN.
Operators are expected to override for production. Other
verified-working public relays for reference:
`nostr.bitcoiner.social`, `nostr-pub.wellorder.net`,
`nostr.oxtr.dev`, `nostr.mom`.
- `app` (default `"fips-overlay-v1"`), `signal_ttl_secs` (120)
- `policy`: `NostrDiscoveryPolicy::{Disabled, ConfiguredOnly (default),
Open}` — controls whether advert-derived endpoints are consumed
only for peers carrying `via_nostr = true`, or also for
non-configured peers within a budget cap.
- `share_local_candidates` (default false) — when false, the offer's
`local_addresses` list is empty and peers see only the reflexive
address. Enable per-node only for genuinely same-LAN deployments;
off-by-default eliminates the misleading-path failure mode for
the common case where peers are not on the same broadcast domain.
- `open_discovery_max_pending` (64) — caps queued open-discovery
retries; bounded by available outbound slots.
- `max_concurrent_incoming_offers` (16) — semaphore against offer
spam; excess offers are debug-logged and dropped.
- `advert_cache_max_entries` (2048) and `seen_sessions_max_entries`
(2048) — bound memory under ambient relay volume; overflow
evictions are debug-logged.
- `attempt_timeout_secs` (10), `replay_window_secs` (300)
- `punch_start_delay_ms` (2000), `punch_interval_ms` (200),
`punch_duration_ms` (10000)
- `advert_ttl_secs` (3600), `advert_refresh_secs` (1800)
Per-peer and per-transport flags:
- `PeerConfig.via_nostr: bool` — when true (and Nostr is enabled),
advert-derived addresses are appended as fallback dial candidates
after static addresses for that peer.
- `PeerConfig.addresses` is now `serde(default)` and may be empty
when `via_nostr: true`; validation requires at least one of the
two to be present per peer, and the error message names the
peer's npub.
- `UdpConfig.advertise_on_nostr: Option<bool>` and
`UdpConfig.public: Option<bool>` — UDP transports can be
advertised either as direct `host:port` (public = true) or as the
`addr: "nat"` sentinel that triggers rendezvous on the peer side.
- `TcpConfig.advertise_on_nostr` and `TorConfig.advertise_on_nostr`
— TCP and Tor onion endpoints can be advertised as directly
reachable.
- A reserved peer address `transport: udp, addr: "nat"` parses without
special-casing in YAML and routes through the bootstrap runtime.
Cross-field validation (`Config::validate`, called from `Node::new`
and `Node::with_identity`):
- Any transport with `advertise_on_nostr = true` requires
`node.discovery.nostr.enabled = true`.
- Any peer with `via_nostr = true` requires
`node.discovery.nostr.enabled = true`.
- A non-public UDP advert (`advertise_on_nostr = true`,
`public = false` — i.e. `udp:nat`) additionally requires at least
one `dm_relay` and at least one `stun_server`.
Surfaced as `ConfigError::Validation`.
== Node integration
`src/node/lifecycle.rs` is the main integration point.
- At node start (after transports are up, before TUN), if Nostr is
enabled and the feature is compiled in, `NostrDiscovery::start` is
invoked, the initial local overlay advert is built from the live
transport set and published, and the runtime handle is stored.
- The rx tick loop calls `poll_nostr_discovery` (feature-gated both
at method definition and call site), which refreshes the local
advert, drains bootstrap events, adopts established traversals,
schedules retries for failed traversals, and — under `policy:
open` — enqueues outbound retries for non-configured peers
visible in the advert cache, bounded by
`open_discovery_max_pending` and the remaining outbound slots.
- Outbound peer dialing is refactored to `try_peer_addresses`, which
first exhausts the static address list in priority order and only
then appends advert-derived fallback addresses; both lists run
through the same `attempt_peer_address_list` code path. The
`udp:nat` sentinel address triggers `NostrDiscovery::request_connect`
for the peer instead of a direct dial and returns `Ok(())`.
- `build_overlay_advert` walks operational transports, consults
per-instance `UdpConfig` / `TcpConfig` / `TorConfig` (matching by
optional transport instance name), and emits an `OverlayAdvert`
including `signalRelays` and `stunServers` when any UDP endpoint
is advertised as NAT.
- `adopt_established_traversal` is the bootstrap handoff API:
allocates a new `TransportId`, constructs a `UdpTransport` with
the user-supplied (or default) `UdpConfig`, calls the new
`adopt_socket_async` to reuse the punched socket verbatim,
registers the transport in the normal transport map, records it
in `bootstrap_transports`, and calls `initiate_connection` so the
normal handshake path runs. On failure, the transport is stopped
and removed cleanly and the set membership is rolled back.
- On clean shutdown, `NostrDiscovery::shutdown` is awaited so
background tasks stop before transports are torn down. (The
advert is not explicitly retracted; NIP-40 expiration plus the
next refresh from any live publisher supersedes it.)
New `Node` fields:
- `nostr_discovery: Option<Arc<NostrDiscovery>>` (feature-gated).
- `bootstrap_transports: HashSet<TransportId>` — per-peer UDP
transports adopted from NAT traversal, cleaned up via
`cleanup_bootstrap_transport_if_unused` whenever the link,
connection, peer, or pending-connect referencing them is removed.
Retry and error surface:
- `RetryState.expires_at_ms: Option<u64>` — optional absolute expiry
for a retry entry. `pump_retries` drops expired entries with an
info log. Used for open-discovery retries, which expire at two
times the advert TTL.
- New `NodeError::BootstrapHandoff(String)` returned from
`adopt_established_traversal` when the underlying transport
adoption fails or local address discovery fails.
- New `ConfigError::Validation(String)`.
- A small refactor extracts `Node::now_ms()` and reuses it across
lifecycle, rx-loop tick, and timeout bookkeeping.
== UDP transport
`src/transport/udp/`:
- `UdpRawSocket::adopt(std::net::UdpSocket, recv_buf, send_buf)`:
adopts an externally bound socket, makes it non-blocking, applies
the configured buffer sizes (warning if the kernel clamps), and
reports the resulting local address. Preserves the NAT mapping —
no rebind.
- `UdpTransport::adopt_socket_async(std::net::UdpSocket)`: the
`start_async` analogue for an already-bound socket, wiring the
async socket and recv task exactly as the fresh-bind path would.
- `Drop` impl for `UdpTransport`: if a transport is dropped while
still holding a recv task or socket (for example on error
teardown), aborts the task, clears the socket, and emits a debug
log so the cleanup is visible in tracing rather than silent.
== Logging and observability
Default `EnvFilter` demotes third-party relay-pool DEBUG output to
TRACE-only: `nostr_relay_pool`, `nostr_sdk`, and `nostr` are pinned
at INFO when our level is anything below TRACE, and at TRACE when
our level is TRACE — so the raw frames are still reachable when
explicitly asked for. RUST_LOG continues to override completely.
Concise one-line DEBUG events are emitted at the meaningful points
in the discovery / hole-punch sequence:
- `advert: published` (event id, relay count, endpoints, ttl)
- `advert: peer cached` (notify-loop ingress for non-self)
- `advert: resolved` (cache hit / relay fetch outcome)
- `traversal: initiator starting`
- `traversal: initiator STUN observed` (reflexive, local count)
- `traversal: offer sent` (session id, relay count, event id)
- `traversal: answer received` (accepted, reflexive, local)
- `traversal: initiator punch succeeded` (remote addr)
- `traversal: offer received` (responder side)
- `traversal: responder STUN observed`
- `traversal: answer sent`
- `traversal: responder punch succeeded`
Npubs are shortened to `npub1<4>..<4>` and event/session ids to
their first 8 hex characters.
Other operator-facing logs:
- `UdpTransport` adoption and drop paths log at info / debug.
- `adopt_established_traversal` logs at debug on entry and info on
successful return, tagged with peer npub, session id, transport
id, and both socket endpoints, so the bootstrap handoff is
traceable end-to-end alongside the `UdpTransport::drop` log.
- `cleanup_bootstrap_transport_if_unused` logs at debug when the
reference-count check drops an adopted transport.
- `connect_peer` tags its entry `debug!` with `peer_npub` so
downstream STUN, punch, and handshake logs for the same peer
correlate for operators.
- Advert-cache and seen-sessions overflow evictions log at debug so
mis-sized caps are visible under ambient relay volume.
- Gift-wrap unwrap failures on `SIGNAL_KIND` events log at trace
(hot path: fires for every unrelated signal event on the same
relay).
- Traversal-offer handler failures log at debug. Expected conditions
such as punch timeout on symmetric NAT are covered there; real
problems are reported upstream via `BootstrapEvent::Failed`.
- Inbound-offer rate-limit messages name the governing config field
(`max_concurrent_incoming_offers`) and state that the offer was
rate-limited rather than failing.
== Tests
- 18 new unit tests in `src/discovery/nostr/tests.rs` covering advert
encoding, signal envelope round-trip, STUN parsing, punch-packet
codec, and replay-window enforcement. Run under the
`nostr-discovery` feature.
- Config-validation tests in `src/config/mod.rs` covering the three
cross-field invariants and YAML parsing of the full
`node.discovery.nostr` block plus `peers[].via_nostr`, empty
`addresses` with `via_nostr: true`, and a `udp: nat` address.
- `src/node/tests/bootstrap.rs` integration tests that drive a
synthetic traversal (bound UDP socket pair + synthetic peer
identity) through `adopt_established_traversal` and assert the
Noise handshake completes over the adopted socket.
- Punch-planner tests assert reflexive-before-LAN ordering and that
same-LAN scenarios still include the LAN target in the plan.
- `testing/nat/` Docker NAT lab harness:
- Local `strfry` relay, local STUN responder, and one or two
router containers performing `iptables` NAT.
- Node LAN interfaces are provisioned with explicit `veth` pairs
injected into the node and router namespaces so every packet
traverses the router namespace (plain Docker bridges are not
used for the LAN).
- `cone` scenario: both peers behind full-cone-emulation NAT
(SNAT with source-port preservation, inbound DNAT back to the
single LAN host regardless of remote source); asserts UDP
traversal succeeds and link remote addresses are on the router
WAN subnet.
- `symmetric` scenario: `MASQUERADE --random-fully`; asserts UDP
traversal fails and TCP fallback converges over router-
published WAN addresses.
- `lan` scenario: both peers share a LAN subnet; asserts LAN
addresses are preferred over reflexive ones.
- Cleanup tears down all profile-gated services
(`--profile cone --profile symmetric --profile lan`) so no
orphan containers survive a run.
- `testing/scripts/build.sh` builds the Docker test image with
`--features "tui nostr-discovery"` by default so NAT-harness
binaries include bootstrap support.
== CI
- Linux release build and nextest unit-test job both use
`--features "gateway nostr-discovery"` so the feature-gated code
and its unit tests compile and run in CI.
- Three new integration matrix entries (`nat-cone`, `nat-symmetric`,
`nat-lan`) invoke `testing/nat/scripts/nat-test.sh`, collect
`docker compose logs` on failure, and always stop containers.
== Packaging and operations
- `packaging/common/fips.yaml` ships a fully commented
`node.discovery.nostr.*` block, plus documented
`advertise_on_nostr` / `public` examples under the UDP transport,
an `advertise_on_nostr` example under TCP, and a `via_nostr: true`
example under the static peer section with both a direct
`host:port` UDP address and a `udp: nat` fallback.
- `.github/workflows/package-openwrt.yml`: NIP-94 release event
publishes target the new default relay set.
== Documentation
- `README.md`: overlay discovery + NAT traversal moved from
"Near-term priorities" into "What works today".
- `docs/design/fips-intro.md`: rewrites the paragraphs that
previously described Nostr discovery and NAT traversal as future
work; describes the shipped mechanism and the feature gate.
- `docs/design/fips-transport-layer.md`: drops the "(future
direction)" qualifier from the Nostr Relay Discovery section,
expands with the `udp:nat` advertisement and bootstrap handoff
description, and updates the Current State callout.
- `docs/design/fips-mesh-layer.md`: notes that mid-session NAT
rebinding (roaming) and initial NAT traversal (Nostr path) are
distinct mechanisms.
- `docs/design/fips-configuration.md`: documents the full
`node.discovery.nostr.*` surface, including the three resource
caps and `share_local_candidates`.
- `docs/design/fips-nostr-discovery.md`: design and configuration
reference for the shipped mechanism, including the empty-
`addresses`-with-`via_nostr` shorthand.
- `docs/proposals/nostr-udp-hole-punch-protocol.md`: adds an
Implemented status callout, clarifies that the punch socket is
per-peer and per-attempt rather than shared with the application
listener, aligns field names with the shipped JSON
(`sessionId`, `issuedAt` / `expiresAt`, `reflexiveAddress`,
`localAddresses`, `stunServer`), sets the `d`-tag to
`fips-overlay-v1`, names the kind as 37195, and notes that
advertised STUN entries are informational.
- `docs/proposals/README.md`: adds a Status column and marks the
hole-punching proposal Implemented.
- `CHANGELOG.md`: Unreleased > Added entry covering the discovery
path, STUN/punch path, configuration surface, and Docker NAT lab.
Co-authored-by: Johnathan Corgan <johnathan@corganlabs.com>
274 lines
10 KiB
Rust
274 lines
10 KiB
Rust
//! Timeout management for stale handshake connections, idle sessions,
|
|
//! and handshake message resend scheduling.
|
|
|
|
use crate::node::Node;
|
|
use crate::peer::HandshakeState;
|
|
use crate::transport::LinkId;
|
|
use tracing::{debug, info};
|
|
|
|
impl Node {
|
|
/// Check for timed-out handshake connections and clean them up.
|
|
///
|
|
/// Called periodically by the RX event loop. Removes connections that have
|
|
/// been idle longer than the configured handshake timeout or are in Failed state.
|
|
pub(in crate::node) fn check_timeouts(&mut self) {
|
|
if self.connections.is_empty() {
|
|
return;
|
|
}
|
|
|
|
let now_ms = Self::now_ms();
|
|
let timeout_ms = self.config.node.rate_limit.handshake_timeout_secs * 1000;
|
|
|
|
let stale: Vec<LinkId> = self
|
|
.connections
|
|
.iter()
|
|
.filter(|(_, conn)| conn.is_timed_out(now_ms, timeout_ms) || conn.is_failed())
|
|
.map(|(link_id, _)| *link_id)
|
|
.collect();
|
|
|
|
for link_id in stale {
|
|
// Log and schedule retry before cleanup (need connection state)
|
|
if let Some(conn) = self.connections.get(&link_id) {
|
|
let direction = conn.direction();
|
|
let idle_ms = conn.idle_time(now_ms);
|
|
if conn.is_failed() {
|
|
debug!(
|
|
link_id = %link_id,
|
|
direction = %direction,
|
|
"Failed handshake connection cleaned up"
|
|
);
|
|
} else {
|
|
debug!(
|
|
link_id = %link_id,
|
|
direction = %direction,
|
|
idle_secs = idle_ms / 1000,
|
|
"Stale handshake connection timed out"
|
|
);
|
|
}
|
|
|
|
// Schedule retry for failed outbound auto-connect peers
|
|
if conn.is_outbound()
|
|
&& let Some(identity) = conn.expected_identity()
|
|
{
|
|
self.schedule_retry(*identity.node_addr(), now_ms);
|
|
}
|
|
}
|
|
self.cleanup_stale_connection(link_id, now_ms);
|
|
}
|
|
}
|
|
|
|
/// Remove a handshake connection and all associated state.
|
|
///
|
|
/// Frees the session index, removes pending_outbound entry, and cleans up
|
|
/// the link and address mapping. Does not log — callers provide context-appropriate
|
|
/// log messages.
|
|
fn cleanup_stale_connection(&mut self, link_id: LinkId, _now_ms: u64) {
|
|
let conn = match self.connections.remove(&link_id) {
|
|
Some(c) => c,
|
|
None => return,
|
|
};
|
|
let transport_id = conn.transport_id();
|
|
|
|
// Free session index and pending_outbound if allocated
|
|
if let Some(idx) = conn.our_index() {
|
|
if let Some(tid) = conn.transport_id() {
|
|
self.pending_outbound.remove(&(tid, idx.as_u32()));
|
|
}
|
|
let _ = self.index_allocator.free(idx);
|
|
}
|
|
|
|
// Remove link and addr_to_link
|
|
self.remove_link(&link_id);
|
|
if let Some(transport_id) = transport_id {
|
|
self.cleanup_bootstrap_transport_if_unused(transport_id);
|
|
}
|
|
}
|
|
|
|
/// Resend handshake messages for pending connections.
|
|
///
|
|
/// For outbound connections in SentMsg1 state, resends the stored msg1
|
|
/// with exponential backoff. Called periodically from the RX event loop.
|
|
pub(in crate::node) async fn resend_pending_handshakes(&mut self, now_ms: u64) {
|
|
if self.connections.is_empty() {
|
|
return;
|
|
}
|
|
|
|
let max_resends = self.config.node.rate_limit.handshake_max_resends;
|
|
let interval_ms = self.config.node.rate_limit.handshake_resend_interval_ms;
|
|
let backoff = self.config.node.rate_limit.handshake_resend_backoff;
|
|
|
|
// Collect resend candidates: outbound, in SentMsg1, with stored msg1,
|
|
// under max resends, and past the scheduled time.
|
|
let candidates: Vec<(LinkId, Vec<u8>)> = self
|
|
.connections
|
|
.iter()
|
|
.filter(|(_, conn)| {
|
|
conn.is_outbound()
|
|
&& conn.handshake_state() == HandshakeState::SentMsg1
|
|
&& conn.resend_count() < max_resends
|
|
&& conn.next_resend_at_ms() > 0
|
|
&& now_ms >= conn.next_resend_at_ms()
|
|
})
|
|
.filter_map(|(link_id, conn)| {
|
|
conn.handshake_msg1().map(|msg1| (*link_id, msg1.to_vec()))
|
|
})
|
|
.collect();
|
|
|
|
for (link_id, msg1_bytes) in candidates {
|
|
// Get transport and address info from the connection
|
|
let (transport_id, remote_addr) = match self.connections.get(&link_id) {
|
|
Some(conn) => match (conn.transport_id(), conn.source_addr()) {
|
|
(Some(tid), Some(addr)) => (tid, addr.clone()),
|
|
_ => continue,
|
|
},
|
|
None => continue,
|
|
};
|
|
|
|
// Send the stored msg1
|
|
let sent = if let Some(transport) = self.transports.get(&transport_id) {
|
|
match transport.send(&remote_addr, &msg1_bytes).await {
|
|
Ok(_) => true,
|
|
Err(e) => {
|
|
debug!(
|
|
link_id = %link_id,
|
|
error = %e,
|
|
"Handshake msg1 resend failed"
|
|
);
|
|
false
|
|
}
|
|
}
|
|
} else {
|
|
false
|
|
};
|
|
|
|
if sent && let Some(conn) = self.connections.get_mut(&link_id) {
|
|
let count = conn.resend_count() + 1;
|
|
let next = now_ms + (interval_ms as f64 * backoff.powi(count as i32)) as u64;
|
|
conn.record_resend(next);
|
|
debug!(
|
|
link_id = %link_id,
|
|
resend = count,
|
|
"Resent handshake msg1"
|
|
);
|
|
}
|
|
}
|
|
}
|
|
|
|
/// Resend session-layer handshake messages and timeout stale handshakes.
|
|
///
|
|
/// For sessions in Initiating or AwaitingMsg3 state:
|
|
/// - If the handshake has exceeded the timeout window, remove the session.
|
|
/// - If a resend is due and under max resends, resend the stored payload
|
|
/// wrapped in a fresh SessionDatagram (so routing can adapt).
|
|
pub(in crate::node) async fn resend_pending_session_handshakes(&mut self, now_ms: u64) {
|
|
if self.sessions.is_empty() {
|
|
return;
|
|
}
|
|
|
|
let timeout_ms = self.config.node.rate_limit.handshake_timeout_secs * 1000;
|
|
let max_resends = self.config.node.rate_limit.handshake_max_resends;
|
|
let interval_ms = self.config.node.rate_limit.handshake_resend_interval_ms;
|
|
let backoff = self.config.node.rate_limit.handshake_resend_backoff;
|
|
let ttl = self.config.node.session.default_ttl;
|
|
|
|
// First pass: find timed-out sessions to remove
|
|
let timed_out: Vec<crate::NodeAddr> = self
|
|
.sessions
|
|
.iter()
|
|
.filter(|(_, entry)| {
|
|
!entry.is_established() && now_ms.saturating_sub(entry.last_activity()) > timeout_ms
|
|
})
|
|
.map(|(addr, _)| *addr)
|
|
.collect();
|
|
|
|
for addr in &timed_out {
|
|
let name = self.peer_display_name(addr);
|
|
info!(dest = %name, "Session handshake timed out, removing");
|
|
self.sessions.remove(addr);
|
|
self.pending_tun_packets.remove(addr);
|
|
}
|
|
|
|
// Second pass: collect resend candidates
|
|
let my_addr = *self.node_addr();
|
|
let candidates: Vec<(crate::NodeAddr, Vec<u8>)> = self
|
|
.sessions
|
|
.iter()
|
|
.filter(|(_, entry)| {
|
|
!entry.is_established()
|
|
&& entry.handshake_payload().is_some()
|
|
&& entry.resend_count() < max_resends
|
|
&& entry.next_resend_at_ms() > 0
|
|
&& now_ms >= entry.next_resend_at_ms()
|
|
})
|
|
.map(|(addr, entry)| (*addr, entry.handshake_payload().unwrap().to_vec()))
|
|
.collect();
|
|
|
|
for (dest_addr, payload) in candidates {
|
|
use crate::protocol::SessionDatagram;
|
|
|
|
let mut datagram = SessionDatagram::new(my_addr, dest_addr, payload).with_ttl(ttl);
|
|
let sent = match self.send_session_datagram(&mut datagram).await {
|
|
Ok(_) => true,
|
|
Err(e) => {
|
|
debug!(
|
|
dest = %self.peer_display_name(&dest_addr),
|
|
error = %e,
|
|
"Session handshake resend failed"
|
|
);
|
|
false
|
|
}
|
|
};
|
|
|
|
if sent && let Some(entry) = self.sessions.get_mut(&dest_addr) {
|
|
let count = entry.resend_count() + 1;
|
|
let next = now_ms + (interval_ms as f64 * backoff.powi(count as i32)) as u64;
|
|
entry.record_resend(next);
|
|
debug!(
|
|
dest = %self.peer_display_name(&dest_addr),
|
|
resend = count,
|
|
"Resent session handshake"
|
|
);
|
|
}
|
|
}
|
|
}
|
|
|
|
/// Remove established sessions that have been idle too long.
|
|
///
|
|
/// Only targets sessions in the Established state. Initiating/AwaitingMsg3
|
|
/// sessions are handled by the handshake timeout.
|
|
pub(in crate::node) fn purge_idle_sessions(&mut self, now_ms: u64) {
|
|
let timeout_ms = self.config.node.session.idle_timeout_secs * 1000;
|
|
if timeout_ms == 0 {
|
|
return; // disabled
|
|
}
|
|
|
|
let idle: Vec<_> = self
|
|
.sessions
|
|
.iter()
|
|
.filter(|(_, entry)| {
|
|
entry.is_established() && now_ms.saturating_sub(entry.last_activity()) > timeout_ms
|
|
})
|
|
.map(|(addr, _)| *addr)
|
|
.collect();
|
|
|
|
for addr in idle {
|
|
// Compute display name before removing the session
|
|
let name = self.peer_display_name(&addr);
|
|
|
|
// Log MMP teardown metrics before removing the session
|
|
if let Some(entry) = self.sessions.get(&addr)
|
|
&& let Some(mmp) = entry.mmp()
|
|
{
|
|
Self::log_session_mmp_teardown(&name, mmp);
|
|
}
|
|
self.sessions.remove(&addr);
|
|
self.pending_tun_packets.remove(&addr);
|
|
debug!(
|
|
dest = %name,
|
|
idle_secs = timeout_ms / 1000,
|
|
"Idle session removed (no application data)"
|
|
);
|
|
}
|
|
}
|
|
}
|