mirror of
https://github.com/jmcorgan/fips.git
synced 2026-08-09 08:14:42 +00:00
83c4e800a500d978b1c6fe8bb411b4b6f0bbd4bb
33
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
7fe1d75637 |
peer: delete the pending-connection type, leaving the control machine whole
The per-peer control machine has absorbed every field the pending connection carried. What remained was a struct holding two Noise handles beside a duplicate copy of bookkeeping nobody read. Replace it with a small carrier for the two handles and delete the type. Presence of that carrier, not the state of the handles inside it, is what marks a machine as mid-handshake. The distinction is essential rather than stylistic: a failed handshake drops its initiation handle and is deliberately retained so the stale sweep can reclaim it, and a completed one has its session taken before disposal. Deriving presence from the handles would make both invisible to the sweep, the connection count, and the peering budget at once, leaking the slot permanently. A test drives an empty carrier past every presence predicate and then detaches it, so a future edit cannot quietly couple the two. The remote startup epoch now comes from the surviving carrier, which the handshake operations already wrote at the same two points with the same value. The paired writes onto the pending connection's own bookkeeping had no readers left and are gone. The handshake-phase surface leaves the public API: it was public by accident rather than design, and the machine behind it is crate internal. Callers outside the crate that need a view of pending handshakes go through the operator queries, which are unchanged. ConnectionState::inbound_with_transport loses its last non-test caller with the inbound seed and is marked test-only. |
||
|
|
e21e09d7e6 |
peer: source the handshake identity and index family from the control machine
The pending connection and the per-peer control machine have carried duplicate copies of the handshake-phase fields since the machine gained its own connection state. Read them from the machine and drop the connection's projections. The peer identity is the sharp one. The connection learned it from msg1 and the machine's carrier did not, so the two views genuinely disagreed for inbound connections until the Noise operations moved onto the machine and began recording each result on both. That is now in place, so every reader can take the machine's copy unchanged: the establish snapshot, the promotion sweep for competing connections, the dial and path in-progress checks, the peering observations and per-peer in-flight budget, and the stale-connection sweep's retry address. Also repointed: our_index, their_index, transport_id, started_at, the stored handshake message bytes, and the idle-timeout check. Two duplicate index writes on the connection are dropped, both immediately preceded by the machine-side write of the same value. Reads that previously came off a connection detached just before its machine was disposed now capture the machine's value first. Test seeding follows the establish paths: the seed builders write our_index to the carrier, which promotion now reads. A new test pins the inbound identity learn on the carrier and the retry address a failed inbound connection reports to the sweep, so a silent regression to a blank identity cannot pass. ConnectionState::duration loses its last non-test caller with the connection accessor and is marked test-only. |
||
|
|
e7537929ba |
node: move link, direction, and peer address onto the control machine
These three had no counterpart on the control machine, so readers still reached them through the pending connection. Add machine-side accessors and repoint every reader, then drop the connection's. The peer address needed its writes lifted, not just its reads repointed: the machine's copy was never written. It is now written at each of the points the connection's copy was — the inbound seed, the dial, message-2 completion, and the two paths that seed a machine from a pre-built connection — so promotion and the resend path read a value with the same provenance at the same time as before. Link and direction need no lift. Both machine constructors already seed them from the same arguments the connection is built with, so an outbound machine carries outbound state and an inbound one inbound, and the machine's link always equals the connection's. The handshake operations' direction guards read the machine's copy for the same reason. The stale-connection sweep's teardown log and its resend path both now take the transport and address from the machine, and each keeps its own check that a pending connection is still attached rather than relying on the caller to have established it. Add a test pinning link, direction, and address on the two shapes that seed a carrier independently — the dial and an accepted message 1. The cross-connection winner reads both values from a carrier one of those two already seeded. |
||
|
|
6a80790742 |
node: source connection index/transport/stats from the peer machine
The connection leg's their_index, transport_id and link_stats duplicated the peer machine's own connection-state copy. Route every production reader through the machine copy: promotion, the stale-connection reaper, the outbound msg1 resend, and the in-use / connecting-path checks. Seed the inbound machine's transport_id from the msg1 packet (the same source the leg used), matching the outbound dial seed, so the hard-required promote read never sees a missing transport. link_stats is a never-mutated zero seed, so its leg accessors are removed outright; the leg their_index/transport_id getters/setters remain only for the test-only connection builder. Also delete the write-only next_resend_at_ms field from the connection state (the resend deadline is carried by the machine-armed retransmit timer, not this field). Byte-identical. |
||
|
|
60b8acf716 |
node: source handshake resend buffers from the peer machine
The connection leg's stored msg1/msg2 handshake-resend buffers duplicated the peer machine's own connection-state copy. Write the machine copy at the outbound msg1-prep and the two inbound authorize sites, then source the retransmit resend-bytes reader, send_stored_msg1, and the pending tier of find_stored_msg2 from the machine. The post-promote active-peer msg2 copy is written from the same wire bytes, so the pending tier now matching after promotion is value-identical. Byte-identical resend behavior. |
||
|
|
b3f2018fce |
node: source connection start/activity timestamps from the peer machine
The connection leg's started_at and last_activity duplicated the peer machine's own connection-state copy. Make the machine copy the sole live telemetry and timeout carrier: re-stamp it with the leg's provenance when the leg is attached at handshake start (the msg1-prep clock, not the earlier dial-time constructor value) and mirror the completion touch, then delete the leg's last_activity/touch accessors and source the connection-row projection, show_connections, and the timeout readers from the machine. Byte-identical for all normal paths. |
||
|
|
56bbc81a40 |
node: derive handshake state from the peer machine, delete the leg field
The connection leg's HandshakeState field duplicated the peer machine's handshake phase. Delete it and derive the displayed handshake state string from the machine's PeerState, looked up by link id (mirroring the resend count). Move the failure signal onto the machine: a send_failed flag that preserves retransmit eligibility (the machine stays in its handshake phase), alongside the existing Failed state. The leg's crypto self-gates now guard on Noise handle presence instead of the deleted phase field, and mark_failed only drops the handle. Telemetry strings, wire bytes, index allocation, and the stale-connection reaping are byte-identical for all normal paths. |
||
|
|
cf1c957336 |
node: embed the handshake leg in the peer machine, drop the connections map
The pending-handshake PeerConnection map and the per-peer control machine map were parallel LinkId-keyed structures whose keysets must stay coherent by hand. With every leg now born with a machine, the leg becomes storage inside its machine (leg: Option<PeerConnection>, pure storage the machine never reads or drives) and Node loses the connections field; every access routes through the machine. The non-mechanical lowerings, each argued at the site: the rekey-vs-establish gate in handle_msg2 tests leg-absence (an established peer's machine stays keyed by its link, so machine-presence would misclassify every rekey msg2 as a fresh establish); the connecting-predicates, peering observation, and handshake-slot budget iterate machines-with-legs so connect-window machines (leg not yet born) are excluded exactly as before and never double-counted against their pending-connect slot; the cross-connection extract takes the leg before disposing the machine; the stale reaper takes the leg and leaves the machine untouched when none is present, matching the old early return. The map-coherence debug check keeps its machine-has-carrier direction with the embedded leg as a carrier; the leg-to-machine direction is now true by construction and its gate const is gone. connection_count() counts machines with legs; the connections() iterator, the test seams, and the control-socket connection rows are re-implemented over the embedded legs with unchanged output. |
||
|
|
765819f52b |
node: reap the handshake timeout through the per-peer machine timer
Move the outbound handshake-timeout reap off the unconditional check_timeouts scan onto the machine-armed HandshakeTimeout timer. A new drive_handshake_timeouts (run before the retransmit drive, so a timed-out leg is reaped rather than resent on the same tick) reaps the outbound legs that carry a HandshakeTimeout timer and have idle-timed-out this tick. The timer's presence selects the leg (only outbound legs arm one; IK inbound arms none); the reap threshold is the shell is_timed_out(now, config) predicate, not the timer's stored deadline. The machine arms the timer from a hardcoded constant at dial, which is not authoritative for an operator-tuned handshake_timeout_secs, so reading the threshold from config each tick keeps the reap neutral for any timeout value and on the last_activity clock exactly as before. check_timeouts keeps reaping everything else: failed connections (all of them, promptly) and the idle-timeout of legs without a machine timer (inbound legs, and any machine-less connection). Total coverage is unchanged. check_timeouts runs before the timer drive, so a failed outbound leg is reaped there first, its timers dropped, and the timeout drive never double-fires on it. Split drive_peer_timers into the timeout drive and the (byte-identical) retransmit drive. The dormant machine timeout handler is not dispatched, so the session index is freed once, by cleanup_stale_connection. A regression test drives a timed-out outbound leg to reap; the existing check_timeouts tests continue to anchor the residual sweep. Test count 1631 to 1632. |
||
|
|
5021197f5c |
node: drive the handshake msg1 resend from the per-peer machine timer
Move the outbound msg1 resend off the unconditional tick function onto the machine-armed retransmit timer. The per-peer machine already arms a HandshakeRetransmit deadline at dial; a new drive_peer_timers fires the due ones (kind-filtered to the retransmit timer) and homes the resend counter on the machine, where the operator-visible count now reads from. The resend decision reads the same operator config as before (interval, backoff, max), the wire bytes and transport target still come from the shell connection, and the pure core computes the backoff schedule. As with the deleted resend_pending_handshakes, the count and reschedule advance only on a successful send: a failed send neither advances the count nor marks the connection failed, it just retries on the next tick. The handshake-timeout timer stays on the legacy check_timeouts path, and the rekey/liveness timers keep their own shell drivers, so drive_peer_timers deliberately fires only the retransmit kind. The show_connections resend count is relocated to read from the machine (the counter's new home) via connection_resend_count; machine-less and inbound connections report 0, matching what the shell connection reported before. Delete resend_pending_handshakes and resend_candidates (the latter also from the LifecycleView trait, its only user). The resend unit test is re-expressed against the machine + timer path, covering the due check and the record-on-success semantics. Known limitation: the first resend interval is armed from the machine's hardcoded 1000ms constant, which equals the config default. Under an operator override of handshake_resend_interval_ms the first resend diverges from the pre-change dial+interval; subsequent resends and the cap remain config-driven. Neutralizing the first-resend override needs the interval threaded into the sync core (the shell arm would be clobbered by the machine's own timer arming at dial), deferred to when the connection and machine entities merge. Test count unchanged at 1631. |
||
|
|
31f5a8c1b7 |
peer: add a per-peer timer store populated by the machine timer actions
The control machine already emits SetTimer/CancelTimer actions when it arms the handshake retransmit and timeout deadlines, but their executor arms were a single no-op stub and the machine's Timeout event was never dispatched -- the real work still runs on the legacy tick (check_timeouts, resend_pending_handshakes). Add the storage the time-as-input driver will read: a peer_timers map keyed by LinkId then TimerKind, holding each armed timer's absolute deadline. The SetTimer arm now inserts (overwrite = reschedule) and CancelTimer removes; the outbound msg2 promote cancels the two dial-armed handshake timers, since the machine survives promotion and the entries would otherwise linger in the store. This store is a shadow: it is written and cleared but no driver reads it yet, so behavior is unchanged. The legacy tick stays authoritative until the driver that feeds Timeout and deletes the overlapping tick paths lands. Route every machine removal through a new remove_peer_machine(link) choke-point that drops the timer store alongside the machine, replacing the twelve direct peer_machines.remove sites so no armed timer outlives its machine. TimerKind gains Hash (to key the store) and Ord (for deterministic driver collection); it is internal and never serialized. Test count unchanged at 1631. |
||
|
|
3b99a416ad |
peer: persist the outbound control machine at dial
Create and persist the per-peer control machine when an outbound handshake is dialed, keyed by its link, instead of building a transient at msg2. The msg2 completion path now looks up that persisted machine to drive the promote, falling back to a transient only if none is present (e.g. a direct-seeded test). The machine parks in the Discovered state until promotion and is inert to the liveness reap and rekey cadence while unpromoted, since it is absent from the peers map. It is removed on every path that ends the outbound leg without promoting -- the stale-connection reaper, the msg2 authorization-failure arm, and the cross-connection resolution block -- mirroring the connection's own lifetime so no dangling machine survives. Its session index is deliberately left unset on the machine (the shell owns the index on its connection), so a later inbound restart does not emit a spurious decrypt-session unregister. A regression test covers that invariant. |
||
|
|
bf81f422ea |
node/peering: drive the mandatory-floor and retry mechanism through the reconciler
Route the auto-connect floor and the connection-retry mechanism through the sans-IO reconciler instead of the imperative Node methods. A thin peering driver (note_handshake_timeout / note_link_dead) centralizes the reflex path with the gate guard and the already-connected check; the per-tick retry-dial and the startup floor build reconciler inputs and execute the emitted Connect intents. The three imperative retry methods are removed and their call sites rerouted. Wire the drain and startup gates. Entering a bounded drain now clears the retry schedule and suppresses the peer-loss reconnect reflex (the drain window runs under the Suspended gate), so the drain no longer fights a reconnect for the peers it just closed. The startup floor runs under an explicit Reconciling gate at the existing peer-connect seam, preserving the current dial position. Behavior-neutral except the intended drain change. Overlay discovery and opportunistic transport-neighbor growth remain imperative for now; they observe the relocated retry schedule and are cut over next. Unit tests migrated to the driver API with assertions unchanged (1607). |
||
|
|
4ad5940114 |
proto/fsp: extract FSP session protocol into sans-IO layout, retire src/protocol
Migrate the FSP end-to-end session subsystem into src/proto/fsp/ following the established sans-IO shape, and retire the src/protocol grab-bag now that FSP was its last occupant. Relocate the FSP session wire (node/session_wire.rs plus the FSP message types from protocol/session.rs) into proto/fsp/wire.rs. Hoist the pure decision logic into proto/fsp/core.rs over plain-data SessionSnapshots returning an ordered FspAction list the shell drives: session-rekey policy, msg3-resend classification, post-decrypt epoch reaction, setup/dual-init tie-break, coords/path-MTU emit-policy, bounded pending-queue, and IPv6 ECN. The crypto-owning SessionEntry stays shell-side in node/session.rs (matching the FMP ActivePeer pattern); proto/fsp is wire + core + limits only, with no proto->noise dependency and no crypto. Move the coords helpers to proto/stp/ (they serialize TreeCoordinate), and split SessionMessageType: the encrypted-inner 0x10-0x1F variants stay in proto/fsp/wire.rs while the 0x20-0x2F routing signals become a new RoutingSignalType in proto/routing/wire.rs. Migrate the session-MMP shell adapter, which continues to drive proto/mmp/. Retire src/protocol: LinkMessageType and SessionDatagram move to a new shared proto/link.rs, ProtocolError becomes proto::Error (relocated verbatim), the deprecated MessageType alias and the unimported PROTOCOL_VERSION are dropped, and src/protocol/ is deleted along with its lib.rs module declaration. Behavior-neutral: wire bytes unchanged, oracle tests pass unedited except mod-path relocation; adds rekey/epoch characterization tests and pure poll/emit-policy core tests. |
||
|
|
4802792e38 | proto/fmp: sans-IO connection-lifecycle state machine | ||
|
|
08b8b3908e |
node: extract immutable state into a shared context and atomic metric registry
Store node counters in an atomic metric registry read through &self, and introduce a shared NodeContext bundle holding the effectively-immutable fields (config, identity, startup epoch, capability limits). Source the immutable config and identity reads across the receive hot path, the handshake/session/mmp/encrypted state machines, and the discovery, tree, bloom, retry, and lifecycle modules through the context accessors rather than direct field reads. The Node fields and the context are rebuilt in lockstep at every mutation site. |
||
|
|
f396d71826 |
node: deterministic tie-breaker for cross-init NAT traversal adoption
When both peers' Nostr-mediated UDP punches complete within the same scheduling window, each side's `BootstrapEvent::Established` event arrives with `is_connecting_to_peer` already true: each side received an inbound msg1 from the peer's pre-punch outbound attempt, which created a connecting-state record. The deduplication skip then fires on both sides, neither installs the fresh traversal socket as canonical, and the peer-adoption budget (45 s) expires. Cross-node wall-clock alignment of the skip log line in observed failures was within ~1 ms — simultaneous dual- fire under contention, the dual-initiation pattern. Apply the deterministic NodeAddr tie-breaker already used at `handlers/handshake.rs:269` for rekey dual-initiation and in `peer::cross_connection_winner` for cross-connection resolution. Smaller NodeAddr wins as adopter: enumerate the in-flight connections whose `expected_identity` points at this peer, tear them down via the canonical `cleanup_stale_connection` helper, and fall through to `adopt_established_traversal`. Larger NodeAddr loses and keeps the existing `continue` semantics; the loser's in-flight outbound is reconciled by `handle_msg1`'s cross- connection logic when the winner's fresh msg1 arrives over the adopted socket. `cleanup_stale_connection` visibility bumped from module-private to `pub(in crate::node)` so it is callable from `lifecycle.rs`. The defensive re-check inside `adopt_established_traversal` itself is left as-is — after the outer cleanup the winner reaches it with `is_connecting_to_peer == false`, so the inner skip won't trip. The `BootstrapEvent::Failed` arm is unchanged: there is no winning outcome on dual failure, and the existing skip + retry-schedule semantics are correct. |
||
|
|
34e00b9f6e |
Add Nostr-mediated overlay discovery and UDP NAT traversal (#53)
Optional peer discovery and NAT hole-punching path gated behind a new
`nostr-discovery` cargo feature. Nodes publish signed overlay endpoint
adverts to public Nostr relays, consume peer adverts to populate
fallback dial addresses, and use STUN-assisted UDP hole punching with
NIP-59 gift-wrap offer/answer signaling to establish direct UDP paths
between NATed peers. Once a punched socket is up, it is handed into
the existing FIPS UDP transport and the standard Noise/FMP session
stack takes over unchanged.
The cargo feature is in the default feature set
(`default = ["nostr-discovery"]`) so stock builds include it; a
build that explicitly disables default features (or selects a
feature set without `nostr-discovery`) does not link the nostr /
nostr-sdk crates and does not emit a no-op poll in the tick loop.
Runtime behavior is independently gated by
`node.discovery.nostr.enabled`, which defaults to false; if the
config enables Nostr on a non-feature build, startup logs a
warning and continues without it.
== Cargo feature and dependencies
- New cargo feature `nostr-discovery = ["dep:nostr", "dep:nostr-sdk"]`.
Not in the default feature set.
- New optional Linux-only dependencies: `nostr 0.44` (features: std,
nip59) and `nostr-sdk 0.44`. Gift-wrap unwrap is hand-rolled in
`src/discovery/nostr/signal.rs` rather than relying on the SDK's
rumor-author check, which FIPS sidesteps by trusting `seal.pubkey`
exclusively.
== Wire format
Overlay advert event: `kind 37195`, parameterized replaceable
(NIP-01 application-defined replaceable range 30000-39999), with
`d = "fips-overlay-v1"`. The digits visually spell FIPS (7=F, 1=I,
9=P, 5=S); a relay survey confirmed the kind is unused.
Advert content carries the version tag, endpoint list
(`udp|tcp|tor` + addr), optional signal-relay and stun-server
metadata, and `issuedAt` / `expiresAt` timestamps. Endpoint
`addr: "nat"` is the sentinel that triggers traversal on the peer
side. NIP-40 `expiration` tag bounds staleness on permanent
shutdown. Lifecycle relies on parameterized-replaceable
supersession; the daemon does not emit NIP-09 kind-5 deletes —
strict relays (Damus, Primal) race delete-against-replace and can
silently drop the replacement.
Gift-wrapped signal event: `kind 21059`. Punch packets carry magic
values `PUNCH_MAGIC` / `PUNCH_ACK_MAGIC`, a sequence number, and a
16-byte session hash.
== Discovery surface
- `src/discovery.rs` (always compiled)
- `EstablishedTraversal`: bound UDP socket + selected remote +
peer npub + optional transport name/config tuning overrides.
- `BootstrapHandoffResult`: returned on successful handoff —
allocated transport id, local/remote addrs, peer NodeAddr,
session id.
- `src/discovery/nostr/` (`#![cfg(feature = "nostr-discovery")]`)
- `types.rs`: wire and control types described above. `ADVERT_KIND`
constant. `BootstrapError` enumerates failure modes (disabled,
missing advert, missing NAT endpoint, no usable relays, invalid
advert, invalid npub, signal timeout, punch timeout, replay,
STUN failure, protocol, nostr, io, serde, event-parse).
- `runtime.rs`: `NostrDiscovery` coordinator. Owns the shared
nostr-sdk `Client`, subscribes to advert + signal event kinds,
maintains a bounded advert cache and a bounded seen-sessions
replay set, drains `BootstrapEvent::{Established, Failed}` for
the node to consume, exposes `update_local_advert`,
`request_connect`, `advert_endpoints_for_peer`,
`cached_open_discovery_candidates`, and `shutdown`.
- `signal.rs`: NIP-59 gift-wrap encode/decode. Outbound wraps are
built against per-attempt ephemeral keys; inbound events are
unwrapped against the node identity.
- `stun.rs`: RFC 5389/8489 Binding Request client with
XOR-MAPPED-ADDRESS parsing for both IPv4 and IPv6; used only to
observe the initiator's own reflexive address against its
locally configured STUN list (peer-advertised STUN is
informational, never an egress target).
- `traversal.rs`: per-attempt candidate-pair punch planner.
Allocates a fresh `0.0.0.0:0` UDP socket per attempt, enumerates
LAN-private and ULA interface addresses alongside the STUN
reflexive address, schedules probe/ack exchanges at the
configured interval for the configured duration, and picks the
first candidate pair that authenticates end-to-end.
Strategy ordering is Reflexive↔Reflexive first, then LAN, then
Mixed. The STUN-observed pair is the only candidate that's reliable
across arbitrary network topologies; trying it first prevents the
planner from latching onto a misleading host-candidate path before
the reflexive path gets a chance. There is no catch-all
Local↔Local strategy: a previous design that paired every local
host candidate from one side with every local host candidate from
the other could declare success on a one-way reachable asymmetric
L3 path (corporate VPN, Tailscale subnet route, overlapping private
address space), only for the FMP handshake to stall because the
return path didn't match. The legitimate `Lan` strategy still pairs
candidates that share a subnet.
== Configuration surface
`node.discovery.nostr.*` (`NostrDiscoveryConfig`), all `serde(default)`
with `deny_unknown_fields`:
- `enabled` (default false), `advertise` (default true)
- `advert_relays`, `dm_relays`, `stun_servers`: defaults are
`wss://relay.damus.io`, `wss://nos.lol`, `wss://offchain.pub`
for both relay lists, and Google / Cloudflare / Twilio for STUN.
Operators are expected to override for production. Other
verified-working public relays for reference:
`nostr.bitcoiner.social`, `nostr-pub.wellorder.net`,
`nostr.oxtr.dev`, `nostr.mom`.
- `app` (default `"fips-overlay-v1"`), `signal_ttl_secs` (120)
- `policy`: `NostrDiscoveryPolicy::{Disabled, ConfiguredOnly (default),
Open}` — controls whether advert-derived endpoints are consumed
only for peers carrying `via_nostr = true`, or also for
non-configured peers within a budget cap.
- `share_local_candidates` (default false) — when false, the offer's
`local_addresses` list is empty and peers see only the reflexive
address. Enable per-node only for genuinely same-LAN deployments;
off-by-default eliminates the misleading-path failure mode for
the common case where peers are not on the same broadcast domain.
- `open_discovery_max_pending` (64) — caps queued open-discovery
retries; bounded by available outbound slots.
- `max_concurrent_incoming_offers` (16) — semaphore against offer
spam; excess offers are debug-logged and dropped.
- `advert_cache_max_entries` (2048) and `seen_sessions_max_entries`
(2048) — bound memory under ambient relay volume; overflow
evictions are debug-logged.
- `attempt_timeout_secs` (10), `replay_window_secs` (300)
- `punch_start_delay_ms` (2000), `punch_interval_ms` (200),
`punch_duration_ms` (10000)
- `advert_ttl_secs` (3600), `advert_refresh_secs` (1800)
Per-peer and per-transport flags:
- `PeerConfig.via_nostr: bool` — when true (and Nostr is enabled),
advert-derived addresses are appended as fallback dial candidates
after static addresses for that peer.
- `PeerConfig.addresses` is now `serde(default)` and may be empty
when `via_nostr: true`; validation requires at least one of the
two to be present per peer, and the error message names the
peer's npub.
- `UdpConfig.advertise_on_nostr: Option<bool>` and
`UdpConfig.public: Option<bool>` — UDP transports can be
advertised either as direct `host:port` (public = true) or as the
`addr: "nat"` sentinel that triggers rendezvous on the peer side.
- `TcpConfig.advertise_on_nostr` and `TorConfig.advertise_on_nostr`
— TCP and Tor onion endpoints can be advertised as directly
reachable.
- A reserved peer address `transport: udp, addr: "nat"` parses without
special-casing in YAML and routes through the bootstrap runtime.
Cross-field validation (`Config::validate`, called from `Node::new`
and `Node::with_identity`):
- Any transport with `advertise_on_nostr = true` requires
`node.discovery.nostr.enabled = true`.
- Any peer with `via_nostr = true` requires
`node.discovery.nostr.enabled = true`.
- A non-public UDP advert (`advertise_on_nostr = true`,
`public = false` — i.e. `udp:nat`) additionally requires at least
one `dm_relay` and at least one `stun_server`.
Surfaced as `ConfigError::Validation`.
== Node integration
`src/node/lifecycle.rs` is the main integration point.
- At node start (after transports are up, before TUN), if Nostr is
enabled and the feature is compiled in, `NostrDiscovery::start` is
invoked, the initial local overlay advert is built from the live
transport set and published, and the runtime handle is stored.
- The rx tick loop calls `poll_nostr_discovery` (feature-gated both
at method definition and call site), which refreshes the local
advert, drains bootstrap events, adopts established traversals,
schedules retries for failed traversals, and — under `policy:
open` — enqueues outbound retries for non-configured peers
visible in the advert cache, bounded by
`open_discovery_max_pending` and the remaining outbound slots.
- Outbound peer dialing is refactored to `try_peer_addresses`, which
first exhausts the static address list in priority order and only
then appends advert-derived fallback addresses; both lists run
through the same `attempt_peer_address_list` code path. The
`udp:nat` sentinel address triggers `NostrDiscovery::request_connect`
for the peer instead of a direct dial and returns `Ok(())`.
- `build_overlay_advert` walks operational transports, consults
per-instance `UdpConfig` / `TcpConfig` / `TorConfig` (matching by
optional transport instance name), and emits an `OverlayAdvert`
including `signalRelays` and `stunServers` when any UDP endpoint
is advertised as NAT.
- `adopt_established_traversal` is the bootstrap handoff API:
allocates a new `TransportId`, constructs a `UdpTransport` with
the user-supplied (or default) `UdpConfig`, calls the new
`adopt_socket_async` to reuse the punched socket verbatim,
registers the transport in the normal transport map, records it
in `bootstrap_transports`, and calls `initiate_connection` so the
normal handshake path runs. On failure, the transport is stopped
and removed cleanly and the set membership is rolled back.
- On clean shutdown, `NostrDiscovery::shutdown` is awaited so
background tasks stop before transports are torn down. (The
advert is not explicitly retracted; NIP-40 expiration plus the
next refresh from any live publisher supersedes it.)
New `Node` fields:
- `nostr_discovery: Option<Arc<NostrDiscovery>>` (feature-gated).
- `bootstrap_transports: HashSet<TransportId>` — per-peer UDP
transports adopted from NAT traversal, cleaned up via
`cleanup_bootstrap_transport_if_unused` whenever the link,
connection, peer, or pending-connect referencing them is removed.
Retry and error surface:
- `RetryState.expires_at_ms: Option<u64>` — optional absolute expiry
for a retry entry. `pump_retries` drops expired entries with an
info log. Used for open-discovery retries, which expire at two
times the advert TTL.
- New `NodeError::BootstrapHandoff(String)` returned from
`adopt_established_traversal` when the underlying transport
adoption fails or local address discovery fails.
- New `ConfigError::Validation(String)`.
- A small refactor extracts `Node::now_ms()` and reuses it across
lifecycle, rx-loop tick, and timeout bookkeeping.
== UDP transport
`src/transport/udp/`:
- `UdpRawSocket::adopt(std::net::UdpSocket, recv_buf, send_buf)`:
adopts an externally bound socket, makes it non-blocking, applies
the configured buffer sizes (warning if the kernel clamps), and
reports the resulting local address. Preserves the NAT mapping —
no rebind.
- `UdpTransport::adopt_socket_async(std::net::UdpSocket)`: the
`start_async` analogue for an already-bound socket, wiring the
async socket and recv task exactly as the fresh-bind path would.
- `Drop` impl for `UdpTransport`: if a transport is dropped while
still holding a recv task or socket (for example on error
teardown), aborts the task, clears the socket, and emits a debug
log so the cleanup is visible in tracing rather than silent.
== Logging and observability
Default `EnvFilter` demotes third-party relay-pool DEBUG output to
TRACE-only: `nostr_relay_pool`, `nostr_sdk`, and `nostr` are pinned
at INFO when our level is anything below TRACE, and at TRACE when
our level is TRACE — so the raw frames are still reachable when
explicitly asked for. RUST_LOG continues to override completely.
Concise one-line DEBUG events are emitted at the meaningful points
in the discovery / hole-punch sequence:
- `advert: published` (event id, relay count, endpoints, ttl)
- `advert: peer cached` (notify-loop ingress for non-self)
- `advert: resolved` (cache hit / relay fetch outcome)
- `traversal: initiator starting`
- `traversal: initiator STUN observed` (reflexive, local count)
- `traversal: offer sent` (session id, relay count, event id)
- `traversal: answer received` (accepted, reflexive, local)
- `traversal: initiator punch succeeded` (remote addr)
- `traversal: offer received` (responder side)
- `traversal: responder STUN observed`
- `traversal: answer sent`
- `traversal: responder punch succeeded`
Npubs are shortened to `npub1<4>..<4>` and event/session ids to
their first 8 hex characters.
Other operator-facing logs:
- `UdpTransport` adoption and drop paths log at info / debug.
- `adopt_established_traversal` logs at debug on entry and info on
successful return, tagged with peer npub, session id, transport
id, and both socket endpoints, so the bootstrap handoff is
traceable end-to-end alongside the `UdpTransport::drop` log.
- `cleanup_bootstrap_transport_if_unused` logs at debug when the
reference-count check drops an adopted transport.
- `connect_peer` tags its entry `debug!` with `peer_npub` so
downstream STUN, punch, and handshake logs for the same peer
correlate for operators.
- Advert-cache and seen-sessions overflow evictions log at debug so
mis-sized caps are visible under ambient relay volume.
- Gift-wrap unwrap failures on `SIGNAL_KIND` events log at trace
(hot path: fires for every unrelated signal event on the same
relay).
- Traversal-offer handler failures log at debug. Expected conditions
such as punch timeout on symmetric NAT are covered there; real
problems are reported upstream via `BootstrapEvent::Failed`.
- Inbound-offer rate-limit messages name the governing config field
(`max_concurrent_incoming_offers`) and state that the offer was
rate-limited rather than failing.
== Tests
- 18 new unit tests in `src/discovery/nostr/tests.rs` covering advert
encoding, signal envelope round-trip, STUN parsing, punch-packet
codec, and replay-window enforcement. Run under the
`nostr-discovery` feature.
- Config-validation tests in `src/config/mod.rs` covering the three
cross-field invariants and YAML parsing of the full
`node.discovery.nostr` block plus `peers[].via_nostr`, empty
`addresses` with `via_nostr: true`, and a `udp: nat` address.
- `src/node/tests/bootstrap.rs` integration tests that drive a
synthetic traversal (bound UDP socket pair + synthetic peer
identity) through `adopt_established_traversal` and assert the
Noise handshake completes over the adopted socket.
- Punch-planner tests assert reflexive-before-LAN ordering and that
same-LAN scenarios still include the LAN target in the plan.
- `testing/nat/` Docker NAT lab harness:
- Local `strfry` relay, local STUN responder, and one or two
router containers performing `iptables` NAT.
- Node LAN interfaces are provisioned with explicit `veth` pairs
injected into the node and router namespaces so every packet
traverses the router namespace (plain Docker bridges are not
used for the LAN).
- `cone` scenario: both peers behind full-cone-emulation NAT
(SNAT with source-port preservation, inbound DNAT back to the
single LAN host regardless of remote source); asserts UDP
traversal succeeds and link remote addresses are on the router
WAN subnet.
- `symmetric` scenario: `MASQUERADE --random-fully`; asserts UDP
traversal fails and TCP fallback converges over router-
published WAN addresses.
- `lan` scenario: both peers share a LAN subnet; asserts LAN
addresses are preferred over reflexive ones.
- Cleanup tears down all profile-gated services
(`--profile cone --profile symmetric --profile lan`) so no
orphan containers survive a run.
- `testing/scripts/build.sh` builds the Docker test image with
`--features "tui nostr-discovery"` by default so NAT-harness
binaries include bootstrap support.
== CI
- Linux release build and nextest unit-test job both use
`--features "gateway nostr-discovery"` so the feature-gated code
and its unit tests compile and run in CI.
- Three new integration matrix entries (`nat-cone`, `nat-symmetric`,
`nat-lan`) invoke `testing/nat/scripts/nat-test.sh`, collect
`docker compose logs` on failure, and always stop containers.
== Packaging and operations
- `packaging/common/fips.yaml` ships a fully commented
`node.discovery.nostr.*` block, plus documented
`advertise_on_nostr` / `public` examples under the UDP transport,
an `advertise_on_nostr` example under TCP, and a `via_nostr: true`
example under the static peer section with both a direct
`host:port` UDP address and a `udp: nat` fallback.
- `.github/workflows/package-openwrt.yml`: NIP-94 release event
publishes target the new default relay set.
== Documentation
- `README.md`: overlay discovery + NAT traversal moved from
"Near-term priorities" into "What works today".
- `docs/design/fips-intro.md`: rewrites the paragraphs that
previously described Nostr discovery and NAT traversal as future
work; describes the shipped mechanism and the feature gate.
- `docs/design/fips-transport-layer.md`: drops the "(future
direction)" qualifier from the Nostr Relay Discovery section,
expands with the `udp:nat` advertisement and bootstrap handoff
description, and updates the Current State callout.
- `docs/design/fips-mesh-layer.md`: notes that mid-session NAT
rebinding (roaming) and initial NAT traversal (Nostr path) are
distinct mechanisms.
- `docs/design/fips-configuration.md`: documents the full
`node.discovery.nostr.*` surface, including the three resource
caps and `share_local_candidates`.
- `docs/design/fips-nostr-discovery.md`: design and configuration
reference for the shipped mechanism, including the empty-
`addresses`-with-`via_nostr` shorthand.
- `docs/proposals/nostr-udp-hole-punch-protocol.md`: adds an
Implemented status callout, clarifies that the punch socket is
per-peer and per-attempt rather than shared with the application
listener, aligns field names with the shipped JSON
(`sessionId`, `issuedAt` / `expiresAt`, `reflexiveAddress`,
`localAddresses`, `stunServer`), sets the `d`-tag to
`fips-overlay-v1`, names the kind as 37195, and notes that
advertised STUN entries are informational.
- `docs/proposals/README.md`: adds a Status column and marks the
hole-punching proposal Implemented.
- `CHANGELOG.md`: Unreleased > Added entry covering the discovery
path, STUN/punch path, configuration surface, and Docker NAT lab.
Co-authored-by: Johnathan Corgan <johnathan@corganlabs.com>
|
||
|
|
6196307f0e |
Merge branch 'maint'
# Conflicts: # src/bin/fips.rs # src/bin/fipstop/app.rs # src/config/mod.rs # src/config/node.rs # src/config/transport.rs # src/mmp/receiver.rs # src/mmp/sender.rs # src/node/handlers/handshake.rs # src/node/handlers/rekey.rs # src/node/lifecycle.rs # src/node/mod.rs # src/transport/ethernet/socket.rs # src/transport/mod.rs # src/upper/tun.rs |
||
|
|
13c0b70dc3 |
Add rustfmt formatting policy and reformat codebase
Add rustfmt.toml with stable defaults and apply cargo fmt to all source files. This establishes a consistent formatting baseline for CI enforcement. |
||
|
|
b8fbecc575 |
Demote 35 info-level log messages to debug for cleaner production output
Reduce info-level noise by moving intermediate steps, periodic telemetry, cross-connection resolution details, and redundant messages to debug. Info output now focuses on operator-relevant state changes: lifecycle events, peer promotions, session establishment, parent switches, and transport start/stop. Key categories demoted: - Handshake cross-connection resolution mechanics (10 messages) - Periodic MMP link/session metric reports (4 messages) - TUN cleanup messages redundant with lifecycle shutdown (4 messages) - Transport "packet channel closed" shutdown messages (4 messages) - Retry scheduling, discovery lookup initiation, other intermediate steps Change default RUST_LOG from debug to info in systemd unit files. |
||
|
|
324535e76d |
Make auto-connect peers retry indefinitely on initial connection failure
Previously, static peers configured with AutoConnect gave up after 6 attempts (1 initial + 5 retries). If the remote peer was offline at startup, the node permanently abandoned the connection. The reconnect path (after MMP link-dead) already retried indefinitely but only activated after a peer had been previously connected. Remove the reconnect parameter from schedule_retry() and always set reconnect=true when creating retry entries, since only auto-connect peers reach this code path. The 300s backoff cap prevents resource waste. The max_retries=0 config still works as an explicit kill switch. |
||
|
|
2293f7d2d5 |
Replace Noise IK with Noise XK at the FSP session layer
The session-layer handshake now uses the 3-message XK pattern instead of the 2-message IK pattern, providing stronger initiator identity hiding. The initiator static key is deferred to msg3 and encrypted under the es+ee DH chain, so eavesdroppers cannot identify the initiator from the handshake. XK pattern: -> e, es (msg1) / <- e, ee + epoch (msg2) / -> s, se + epoch (msg3) Key changes: - Add XK handshake methods alongside existing IK methods in noise module - Add SessionMsg3 wire format and FSP_PHASE_MSG3 (0x03) prefix - Replace Responding state with AwaitingMsg3 in session state machine - Rewrite session handlers: handle_session_setup defers identity to msg3, handle_session_ack processes msg2 and sends msg3, new handle_session_msg3 completes the responder handshake and registers identity - Link-layer (FMP) continues to use Noise IK unchanged - Add comprehensive XK unit tests and update all integration tests |
||
|
|
78a73e1749 |
Auto-reconnect after MMP peer removal, directed outbound configs, sim improvements
Auto-reconnect: - Add per-peer auto_reconnect config (default true) to PeerConfig - schedule_reconnect() feeds removed peers back into retry system with unlimited retries and exponential backoff after MMP dead timeout - RetryState gains reconnect flag to distinguish startup retries (max_retries-limited) from auto-reconnect (unlimited) Retry re-fire fix: - process_pending_retries() now pushes retry_after_ms past the handshake timeout window after successful initiate_peer_connection(), preventing retries from firing every tick with no backoff Chaos sim improvements: - Directed outbound configs: BFS spanning tree + lower-ID-first assignment eliminates dual-connect race conditions in simulation - Save runner log (runner.log) alongside per-node logs for event correlation - Increase churn-20 traffic aggressiveness and node churn (max_down_nodes 3→5, traffic interval min 0s, duration max 90s, concurrent flows 5→10) |
||
|
|
5d1783edd5 |
Session-layer handshake message retry with exponential backoff
Add resend logic for SessionSetup/SessionAck messages routed through the mesh. Stores the encoded payload on SessionEntry for resend in a fresh SessionDatagram (so routing can adapt to topology changes). Uses the same config parameters as link-layer retry. Also fixes a latent bug: Initiating/Responding sessions previously had no timeout — a stuck handshake would live forever. Now cleaned up after handshake_timeout_secs (default 30s). Responder idempotency: duplicate SessionSetup triggers resend of stored SessionAck instead of being silently dropped. Initiator-side duplicate SessionAck already handled safely (entry.take_state() sees Established, puts it back and returns). Handshake payload cleared on Established transition at both initiator (handle_session_ack) and responder (handle_encrypted_session_msg). |
||
|
|
6a10e9228b |
Link-layer handshake message retry with exponential backoff
Add message-level retry for Noise IK handshake within the 30s timeout window. Previously, a lost msg1 or msg2 required the full timeout to expire before cleanup and retry. Under 10% bidirectional loss (~19% per attempt), this made connection establishment unreliable. Initiator resends stored msg1 bytes with exponential backoff (1s, 2s, 4s, 8s, 16s — 5 resends). Responder stores msg2 and resends on duplicate msg1 receipt. Duplicate msg2 at initiator drops silently via existing pending_outbound cleanup. Config: handshake_resend_interval_ms (1000), handshake_resend_backoff (2.0), handshake_max_resends (5) on node.rate_limit. P(all 6 attempts fail) under 19% loss = 0.19^6 ≈ 0.005%. |
||
|
|
9605cbafe3 |
Human-readable peer identifiers in log messages
Replace raw NodeAddr hex strings in log output with human-readable identifiers using a four-tier lookup: configured alias, active peer short npub, session endpoint short npub, or truncated hex fallback. - Add PeerIdentity::short_npub() for compact npub display (npub1xxxx...yyyy) - Add NodeAddr::short_hex() for compact hex fallback (first 4 bytes + ...) - Add peer_aliases map populated at startup from peer config - Add Node::peer_display_name() with four-tier resolution - Add SessionEntry::remote_pubkey() accessor for session-layer lookups - Update all 120 log field occurrences across 11 handler/node files - Pass pre-computed display names to MMP static metric/teardown methods to work around borrow checker constraints in iterator loops |
||
|
|
c5c7e68a6e |
Session idle timeout now based on application data only
MMP reports (SenderReport, ReceiverReport, PathMtuNotification) were resetting the idle timer on both RX and TX paths, preventing sessions from ever timing out when MMP traffic kept flowing. Changed touch() to only be called for DataPacket send/receive and session establishment, so sessions with no application data tear down after the idle timeout even with active MMP measurement traffic. |
||
|
|
04d9fd625d |
FSP wire format revision and session-layer MMP implementation
FSP wire format revision (TASK-2026-0007): Introduce the FIPS Session Protocol (FSP) wire format with a 4-byte common prefix [ver_phase:1][flags:1][payload_len:2 LE] replacing the old 1-byte msg_type dispatch. All session messages share this prefix with phase-based dispatch (Established, Setup, Ack, Unencrypted). - New session_wire.rs: FSP constants, header types, parse/build helpers - SessionMessageType enum: DataPacket (0x10), SenderReport (0x11), ReceiverReport (0x12), PathMtuNotification (0x13) - FspFlags (CP/K/U) and FspInnerFlags (SP) for flag management - SessionSenderReport, SessionReceiverReport, PathMtuNotification message structs with encode/decode - FSP send pipeline: 12-byte header as AAD, 6-byte inner header (timestamp + msg_type + inner_flags), encrypt_with_aad() - FSP receive pipeline: parse header, extract cleartext coords (CP), AEAD decrypt with AAD, strip inner header, msg_type dispatch - Forwarding: transit nodes parse cleartext coords without decryption - Removed DataPacket struct and associated types - SessionEntry: session_start_ms, mark_established(), session_timestamp() - FIPS_OVERHEAD: 144 → 150 bytes (+6 for FSP inner header) - Design docs updated for new wire format Session-layer MMP implementation (TASK-2026-0008): Implement complete session-layer MMP reusing the link-layer algorithm modules (SenderState, ReceiverState, MmpMetrics, SpinBitState) with independent configuration and higher report interval clamps. - SessionMmpConfig: separate config section (node.session_mmp.*) - MmpSessionState: session-specific wrapper with PathMtuState tracking - Session-layer constants (500ms-10s report intervals, 1s cold start) - Parameterized interval methods (new_with_cold_start, update_report_interval_with_bounds) on SenderState/ReceiverState - Bidirectional From conversions between link/session report types - SessionEntry: mmp and is_initiator fields, initialized on Established - send_session_msg() for reports/notifications - Per-message RX recording with spin bit state tracking - Handlers for SenderReport, ReceiverReport, PathMtuNotification - path_mtu threaded from SessionDatagram envelope through to handlers - check_session_mmp_reports() tick handler with collect-then-send pattern - Periodic and teardown operator logging for session metrics - PathMtuState: destination observes incoming MTU on all session messages, source seeded from outbound transport MTU, decrease-immediate / increase-requires-3-consecutive rules Link-layer MMP fix: - Stop feeding spin bit RTT samples into SRTT estimator; inter-frame timing in the mesh is irregular, inflating spin-bit RTT by variable processing delays; timestamp-echo provides accurate RTT 29 files changed, 602 tests pass, 0 clippy warnings. |
||
|
|
5f1c6c2c7c |
Add session idle timeout (90s) and identity cache expiry (60s)
Sessions in the Established state that have no activity for 90 seconds are now automatically removed. This ensures idle sessions are torn down before transit node coord_cache entries expire (300s TTL), so that when traffic resumes a fresh SessionSetup re-warms transit node caches with current coordinates. The identity cache now stores registration timestamps and expires entries after 60 seconds via lazy expiry on lookup. This prevents unbounded growth while allowing natural repopulation through DNS resolution on next use. Timer ordering: identity (60s) < session (90s) < coord_cache (300s). Both timeouts are configurable: node.session.idle_timeout_secs and node.cache.identity_ttl_secs. Setting idle_timeout_secs to 0 disables session idle purging. Changes: - Add idle_timeout_secs (default 90) to SessionConfig - Add identity_ttl_secs (default 60) to CacheConfig - Add timestamp to identity_cache entries, lazy expiry on lookup - Add purge_idle_sessions() called from tick loop - Remove #[cfg(test)] from SessionEntry::last_activity() - 7 new tests covering timeout behavior and edge cases |
||
|
|
b8a1f322c2 |
Module reorganization and clippy cleanup
Move single-consumer modules into node/:
- rate_limit.rs, wire.rs, dns.rs — exclusively used by node subsystem
- Reduces top-level lib.rs from 16 to 13 modules
Split large files into focused subdirectories:
- noise.rs (1475 lines) → noise/{mod, handshake, session, replay, tests}.rs
- tree.rs (1479 lines) → tree/{mod, coordinate, declaration, state, tests}.rs
- bloom.rs (849 lines) → bloom/{mod, filter, state, tests}.rs
- All public APIs re-exported from mod.rs, no external import changes
Remove unused rate_limit defaults:
- HANDSHAKE_TIMEOUT_SECS, MAX_PENDING_INBOUND constants
- Default constructor eliminated in favor of with_params() taking config values
Fix all clippy warnings across codebase:
- Remove .clone() on Copy types, collapse nested ifs, replace match-return-None
with ?, remove/gate unused code, fix loop indexing, remove unnecessary casts
- Box large PeerSlot enum variants to reduce size disparity
- cargo clippy --all-targets now reports zero warnings
|
||
|
|
7463d8799a |
Promote 27 hardcoded constants to configurable parameters
Add 9 config subsection structs (LimitsConfig, RateLimitConfig, RetryConfig, CacheConfig, DiscoveryConfig, TreeConfig, BloomConfig, SessionConfig, BuffersConfig) under node.* with serde defaults. Wire all configurable values through to consuming code: - Resource limits (max_connections, max_peers, max_links, max_pending_inbound) - Rate limiting (handshake_burst, handshake_rate, handshake_timeout_secs) - Retry/backoff (consolidate max_retries, base_interval_secs under node.retry.*, add max_backoff_secs) - Cache sizes/TTL (coord_size, coord_ttl_secs, route_size) - Discovery (ttl, timeout_secs, recent_expiry_secs) - Spanning tree (root_refresh_secs, announce_min_interval_ms, parent_switch_threshold) - Bloom filter (update_debounce_ms) - Session/data plane (default_hop_limit, pending_packets_per_dest, pending_max_destinations) - Internal buffers (packet_channel, tun_channel, dns_channel) - Network internals (base_rtt_ms, tick_interval_secs) - DNS responder TTL (dns.ttl) REPLAY_WINDOW_SIZE kept as compile-time constant (array sizing). Disable flaky test_discovery_100_nodes (run with --ignored). |
||
|
|
cc29c51cac |
Refactor node/handlers.rs and node/tests.rs into subdirectories
Split handlers.rs (986 lines) into handlers/ with 5 subfiles organized by responsibility: rx_loop, encrypted, handshake, dispatch, timeout. Split tests.rs (2350 lines) into tests/ with 4 subfiles: unit tests, handshake integration, spanning tree convergence, and bloom filter tests. Shared test helpers extracted to tests/mod.rs. Visibility adjusted from pub(super) to pub(in crate::node) for handler methods now two levels deep. Unused imports cleaned up in node/mod.rs. All 316 tests pass, zero warnings. |