From dca938c10427b11446ac6fe771aaf5a7c2493b0b Mon Sep 17 00:00:00 2001 From: Arjen <18398758+Origami74@users.noreply.github.com> Date: Fri, 11 Sep 2026 20:37:26 -0300 Subject: [PATCH] feat(peer): switchover between a peer's transports, with no handshake MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A peer reachable over more than one transport keeps one Noise session and moves its traffic between transports on failure or degradation. Implements docs/design/fips-multi-path-switchover.md §4-§10 and closes the two-interface case in #143. Three inner link messages next to Heartbeat: `0x52 PathProbe` and `0x53 PathAck`, carrying a probe id, the sender's path id and a `remote_active` bit, and `0x54 PathClose`, naming the receiver's path id and a reason. A probe is an ordinary encrypted frame sent on a candidate transport; the receiver, having decrypted it against the session found by index, adds the path as `Probing`, marks it `rx_live` and answers on that same path. The prober's receipt of the ack marks the path `Live`, `tx_live`, and takes an RTT sample. No handshake, no key material, no index allocation. Old nodes drop the unknown types at debug, so a path to one stays `Probing` and never becomes eligible. The discovery gate changes shape: a live peer beaconing on a transport we hold no path to it over becomes a path candidate rather than being skipped, and the heartbeat tick probes it. The active path's first probe is small — the handshake proved it and seeded its MTU — while a standby's discovery probes are padded to the link MTU, as is one a minute on every path, so a medium that passes small frames and drops large ones never proves itself. A standby the peer never answers on is given up after eight probes; the active path never is. Detection is per path and takes each medium's own failure signal: a carrier edge, an unreachable-on-send (`ENETUNREACH`/`EHOSTUNREACH`), an interface going away, or two unanswered heartbeats on a path the peer is also silent on. Any of them marks the path `Suspect` and selection leaves it at once, because the standby is warm: heartbeats run at `node.path.active_heartbeat_ms` where either side sends and `standby_heartbeat_ms` elsewhere, both stretched by the path's own round trip so a Tor or Nym path is neither flooded nor declared dead every round trip. A peer holding one live path is not heartbeated here at all — selection has nothing to move to, and the link heartbeat keeps its liveness. Soft signals (the peer's `remote_active` flipping away, silence here while a standby hears the peer) trigger a probe, never `Suspect`: reading them as a verdict forces both sides onto one path and loops under a one-way failure. A node that loses a path tells the peer with a `PathClose` on a surviving one, so the peer moves at once rather than after its own timeout. A transport that returns inside the five-minute grace revives its dead paths as `Probing` with their RTT window and ETX intact. Selection is measured, not configured: each path scores `quality_index(etx, min_rtt)`, and traffic moves when the active path is no longer eligible (mandatory) or when a standby beats it by `switch_margin` for `switch_dwell_secs` (discretionary). Min RTT over a window rather than SRTT, because SRTT inflates under load on the path carrying traffic while an idle standby looks pristine — a ping-pong generator. A `role: backup` transport carries a peer's traffic only while no normal path is eligible, and yields outright when one becomes eligible. `fipsctl path pin` overrides both while its path is eligible. A switch re-seeds the path MTU from the new path, tightens the session MTUs and refreshes the MSS ceiling, so the first frames after a switch are not black-holed. The link record and `addr_to_link` follow the active path. The link cost the tree sees is held at its pre-switch value for the dwell, and until the two receiver reports that span the switch have arrived — the first counts every frame in flight on the old path as lost and spikes the per-report ETX for one interval, the second replaces it — so neither a short flap nor that spike ripples mesh-wide through parent selection or the next-hop order. Operator surface: `role: backup` on any transport, `node.path.*` (`switch_margin`, validated finite and at least 1.0, `switch_dwell_secs`, `min_samples`, `active_heartbeat_ms`, `standby_heartbeat_ms`), and `fipsctl path show|pin|unpin` over the `path_show`, `path_pin` and `path_unpin` control commands. Two chaos scenarios calibrate the defaults and are wired into both runners: `dual-path-flap` (a raw-Ethernet veth as the cable, the Docker bridge over UDP as the wifi) and `dual-udp-flap` (two interface-bound UDP instances). Each flaps one path under iperf and carries detectors that can fail on a switchover that did not carry traffic — a per-node ceiling on "Peer promoted to active" (a second is a re-peering), a one-second ceiling from link-down to the first switch, and a two-second ceiling on any zero-byte iperf interval run — alongside the `path_switches` band. Neither has been run to calibrate; the defaults are chosen, not derived, and the design doc says so. Refs #143 --- .github/workflows/ci.yml | 6 + CHANGELOG.md | 45 +- docs/design/fips-mesh-layer.md | 14 + docs/reference/cli-fipsctl.md | 17 + docs/reference/configuration.md | 59 +- docs/reference/control-socket.md | 3 + docs/reference/wire-formats.md | 49 + src/bin/fipsctl.rs | 43 + src/config/mod.rs | 6 +- src/config/node.rs | 85 ++ src/config/transport.rs | 71 + src/control/commands.rs | 61 + src/control/queries.rs | 3 +- src/node/dataplane/dispatch.rs | 19 + src/node/dataplane/encrypted.rs | 28 +- src/node/dataplane/rx_loop.rs | 19 +- src/node/handlers/handshake.rs | 28 +- src/node/handlers/mmp.rs | 86 +- src/node/handlers/mod.rs | 1 + src/node/handlers/path.rs | 775 ++++++++++ src/node/handlers/session.rs | 2 +- src/node/lifecycle/mod.rs | 17 +- src/node/mod.rs | 129 +- src/node/tests/handshake.rs | 2 + src/node/tests/heartbeat.rs | 4 +- src/node/tests/multi_path.rs | 1450 +++++++++++++++++++ src/node/tests/routing.rs | 25 +- src/node/tests/spanning_tree.rs | 4 +- src/node/tree.rs | 9 +- src/peer/active.rs | 1003 ++++++++++++- src/peer/mod.rs | 5 +- src/proto/link.rs | 167 +++ src/transport/ble/mod.rs | 4 + src/transport/ethernet/mod.rs | 7 + src/transport/loopback.rs | 35 +- src/transport/mod.rs | 54 + src/transport/nym/mod.rs | 4 + src/transport/tcp/mod.rs | 4 + src/transport/tor/mod.rs | 4 + src/transport/udp/mod.rs | 4 + testing/chaos/README.md | 14 + testing/chaos/scenarios/dual-path-flap.yaml | 123 ++ testing/chaos/scenarios/dual-udp-flap.yaml | 121 ++ testing/chaos/sim/assertions.py | 240 +++ testing/chaos/sim/config_gen.py | 56 +- testing/chaos/sim/links.py | 13 +- testing/chaos/sim/netem.py | 20 +- testing/chaos/sim/runner.py | 112 +- testing/chaos/sim/scenario.py | 143 +- testing/chaos/sim/topology.py | 119 +- testing/chaos/sim/veth.py | 31 +- testing/ci-local.sh | 4 +- testing/docker/entrypoint.sh | 14 +- testing/lib/log_analysis.py | 9 + 54 files changed, 5174 insertions(+), 196 deletions(-) create mode 100644 src/node/handlers/path.rs create mode 100644 testing/chaos/scenarios/dual-path-flap.yaml create mode 100644 testing/chaos/scenarios/dual-udp-flap.yaml diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index b42f1f10..6d15b695 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -686,6 +686,12 @@ jobs: - suite: ethernet-churn type: chaos scenario: ethernet-churn + - suite: dual-path-flap + type: chaos + scenario: dual-path-flap + - suite: dual-udp-flap + type: chaos + scenario: dual-udp-flap - suite: tcp-mesh type: chaos scenario: tcp-mesh diff --git a/CHANGELOG.md b/CHANGELOG.md index 036a0197..ed06cb66 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -9,6 +9,48 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ### Added +- Multi-path switchover: a peer reachable over more than one transport keeps + one Noise session and moves its traffic between transports on failure or + degradation, with no handshake. Encrypted frames are demuxed by session + index alone, so a frame from a known peer decrypts whichever transport + delivered it. A peer holds a *path* per transport: the first from the + handshake, further ones proven by a `PathProbe`/`PathAck` exchange under + the existing session. Every path is heartbeated on its own once a peer + has more than one — fast on a path either side sends on, slow on a + standby — and a lost carrier, a route gone on send, an interface gone, or + two unanswered heartbeats on a path the peer is also silent on moves + traffic to the best proven standby inside a second; only a peer with no + path left is dropped. Selection is measured, not configured: per-path + `etx × (1 + min_rtt / 100)` with a margin and a dwell (`node.path.*`), so + a cable coming back under working wifi does not take traffic back until + the wifi degrades. A switch re-seeds the path MTU, tightens session MTUs + and holds the tree-visible link cost for the dwell, and until the two + receiver reports that span the switch have arrived, so neither the + switch nor the one-report loss spike it leaves ripples mesh-wide. A transport that returns inside the five-minute grace revives + its dead paths with their history. Design: + `docs/design/fips-multi-path-switchover.md`; defaults are placeholders, + and `testing/chaos/scenarios/dual-path-flap` and `dual-udp-flap` are the + scenarios that calibrate them, each failing on a re-peering, a switch + slower than a second, or an iperf stall past two seconds. + +- Wire, for multi-path: three inner link-message types in the link-control + block, `0x52 PathProbe`, `0x53 PathAck` (probe id, `remote_active` flag, + the sender's path id) and `0x54 PathClose` (the receiver's path id and a + reason). Ordinary encrypted FMP frames under the session; no header, + handshake or index change. Old nodes drop them at debug after + authenticating the frame. The first probe on a standby and one a minute + after on every path are padded to the link MTU, so a medium that passes + small frames and drops large ones never proves itself; a node that loses + a path tells the peer with a `PathClose` on a surviving path, so the + peer moves at once rather than after its own timeout. + +- Operator surface, for multi-path: `role: backup` on any transport (never + carries a peer's traffic while a normal path is eligible); + `fipsctl path show|pin|unpin` and the `path_show`, `path_pin`, + `path_unpin` control commands; `node.path.*` (`switch_margin`, validated + finite and at least 1.0, `switch_dwell_secs`, `min_samples`, + `active_heartbeat_ms`, `standby_heartbeat_ms`). + - An authentic frame arriving on a transport the peer has no path on no longer re-pins the peer's send side to that transport, and a decrypt failure on such a transport is not counted toward force-removal. Both @@ -42,7 +84,8 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 on a wifi-only router) is bound and healthy rather than permanently `Degraded`, and starts carrying traffic the moment a port comes up. Whether an interface has carrier is reported separately as `interface.carrier` in - `show_transports`, never acted on. + `show_transports`; path selection reads it, presence does not. + This closes the OpenWrt boot race (procd starts `fips` before wifi has created `fips-mesh0` / `fips-ap0`; both transports were skipped for the life of the process while the 802.11s peer link formed anyway, so the node looked diff --git a/docs/design/fips-mesh-layer.md b/docs/design/fips-mesh-layer.md index b437fa22..0a4a2fe7 100644 --- a/docs/design/fips-mesh-layer.md +++ b/docs/design/fips-mesh-layer.md @@ -446,6 +446,16 @@ message type). Any successfully decrypted frame — data, gossip, MMP report, or heartbeat — resets the peer's last-receive timestamp tracked by the MMP receiver. +### Path probes + +A peer reachable over more than one transport holds a path per transport +under the one session; each path is heartbeated on its own with a +PathProbe (0x52) answered by a PathAck (0x53) on the same path, and a +path that is going is announced with a PathClose (0x54) on a surviving +one. Layouts in `wire-formats.md`; the mechanism in +`docs/design/fips-multi-path-switchover.md`. A peer with one path is not +path-probed: the bare Heartbeat above keeps its liveness. + ### Dead Timeout When no traffic (of any kind) is received from a peer for the @@ -483,6 +493,9 @@ established-frame envelope. They group naturally by purpose: - **Liveness and lifecycle**: Heartbeat is a minimal frame sent peer-to-peer to keep the link alive; Disconnect carries an orderly teardown reason code peer-to-peer. +- **Paths**: PathProbe, PathAck and PathClose (0x52–0x54) prove, measure + and withdraw the individual transports a multi-path peer is reachable + over, peer-to-peer under the one session. Handshake messages (phase 0x1 msg1, phase 0x2 msg2) travel before encryption is established and are identified by the FMP common-prefix @@ -559,6 +572,7 @@ an attacker sends invalid packets to elicit responses. | Rate limiting (token bucket) | **Implemented** | | Disconnect with reason codes | **Implemented** | | Heartbeat liveness detection | **Implemented** | +| Multi-path (per-path probes, switchover without a handshake) | **Implemented** (`feat/multi-path-switchover`) | | Reconnection handling | **Implemented** | | Auto-reconnect after link-dead removal | **Implemented** | | Handshake message retry (link + session layer) | **Implemented** | diff --git a/docs/reference/cli-fipsctl.md b/docs/reference/cli-fipsctl.md index 3818669e..15fa5c60 100644 --- a/docs/reference/cli-fipsctl.md +++ b/docs/reference/cli-fipsctl.md @@ -140,6 +140,23 @@ Tell the daemon to drop a peer link. | -------- | ----------- | | `peer` | npub (bech32) or hostname from `/etc/fips/hosts`. | +### `path ` + +A peer reachable over more than one transport holds one Noise session and a +*path* per transport; traffic moves between paths on failure or degradation +without a handshake. These commands show and override that choice. + +| Subcommand | Description | +| ---------- | ----------- | +| `path show ` | Every path to the peer: its transport, address, state (`probing`, `live`, `suspect`, `dead`), whether it is the one this node sends on (`active`) and the one the peer sends on (`remote_active`), role, pin, how long since each direction was last proven, the min/last RTT, sample count, per-path ETX and score. Also the link cost the tree sees and whether a post-switch hold is in force. | +| `path pin ` | Pin this node's traffic to the peer to one transport, named by its configured instance name (e.g. `cable`) or numeric id. The pin holds while the path is eligible (live, and answering probes); while it is not — suspect, dead, or unproven — selection is measured as if unpinned, and the pin re-applies the moment the path is eligible again, with no margin or dwell. `unpin` clears it. | +| `path unpin ` | Clear the pin; selection is measured again from the next tick. | + +`peer` is an npub (bech32) or a hostname from `/etc/fips/hosts`. Selection +between paths is measured, not configured (see `node.path.*` in +[configuration.md](configuration.md)); the pin and a transport's +`role: backup` are the only overrides. + ### `probe ` Diagnose whether a mesh endpoint is reachable, in five stages, and diff --git a/docs/reference/configuration.md b/docs/reference/configuration.md index 6c49376e..4d3c0951 100644 --- a/docs/reference/configuration.md +++ b/docs/reference/configuration.md @@ -446,6 +446,41 @@ Controls tree construction and parent selection. | `node.tree.flap_window_secs` | u64 | `60` | Sliding window for counting parent switches | | `node.tree.flap_dampening_secs` | u64 | `120` | Extended hold-down duration when flap threshold exceeded | +### Path Selection (`node.path.*`) + +A peer reachable over more than one transport keeps one Noise session and +holds a *path* per transport. Further paths are added by a probe under the +existing session (a beacon from a live peer on a transport with no path to +it yet is probed, not dialled), each path is heartbeated on its own, and +this node's traffic moves between them on failure or degradation with no +handshake. Selection is measured, not configured: a path's score is +`etx × (1 + min_rtt_ms / 100)` from its own probes. These knobs bound when a +measured difference is acted on; their defaults are placeholders pending +calibration. + +| Parameter | Type | Default | Description | +|-----------|------|---------|-------------| +| `node.path.switch_margin` | f64 | `1.3` | Discretionary switch margin `K`: the active path's score must exceed the best standby's by this factor. Encodes the fail-back policy: a cable returning under working wifi (1.00 vs about 1.10) is under `K`, so traffic stays until the wifi degrades. | +| `node.path.switch_dwell_secs` | u64 | `2` | The margin must hold this long before a discretionary switch. Also how long the link cost the tree sees is held at its pre-switch value after any switch. | +| `node.path.standby_heartbeat_ms` | u64 | `1000` | Heartbeat interval on a standby path. A standby is only as warm as its last echo; this bounds how stale a proven standby can be when the active path dies. A standby the peer has never acknowledged (an old node, or a medium it cannot hear us on) is probed full-size on a backoff that doubles from this up to `node.heartbeat_interval_secs`, eight times, then given up as dead and forgotten after the five-minute grace; a fresh address starts the count over. | +| `node.path.min_samples` | u32 | `2` | RTT samples a standby needs before it is eligible. | +| `node.path.active_heartbeat_ms` | u64 | `200` | Heartbeat interval on a path that either side sends on. A heartbeat unanswered for two of these on a path the peer has acknowledged before, with nothing heard from the peer on it either, marks it suspect, and selection leaves it at once; a late echo on a path still carrying the peer's frames is counted as loss only, and a late echo still measures the path, so the timeout can stretch on a medium slower than it. Both the interval and the timeout stretch with a path's own measured round trip, so a Tor or Nym path is neither flooded nor declared dead every round trip. The first probe on a standby, and one a minute after on every path, is padded to the link MTU so a medium that passes small frames and drops large ones never proves itself. A peer with a single live path is not path-heartbeated at all — there is nothing to switch to — so on a mesh of single-homed peers this costs nothing. | + +A path losing its transport (interface gone), carrier (cable unplugged), or a +route (`ENETUNREACH` on send) is left immediately when another proven path +exists, and the peer is told on a surviving path so it moves too rather than +waiting for its own timeout; only a peer with no path left is dropped. The +interface and carrier signals exist for Ethernet transports, which are bound +per interface; a UDP instance bound with `transports.udp.interface` gets no +presence or carrier signal, so the loss of its path is detected by the +unreachable-on-send error and the heartbeat echo timeout only (two +`active_heartbeat_ms`, about half a second). A transport that returns +within five minutes of going revives its dead paths with their measured +history rather than starting from nothing. On a connection-oriented +transport (`tcp`, `tor`, `nym`) a path's connection is the one a dial or +a probe opened; a `disconnect` closes every path's connection. See +`fipsctl path` in [cli-fipsctl.md](cli-fipsctl.md). + ### Bloom Filter (`node.bloom.*`) | Parameter | Type | Default | Description | @@ -645,6 +680,12 @@ adding entries and the precedence rules: ## Transports (`transports.*`) +Every transport accepts `role: normal | backup` (default `normal`). A +`backup` transport never carries a peer's traffic while any `normal` path to +that peer is eligible, whatever the measurements say; it is a statement about +the transport's purpose ("drop BLE when something better is stable"), not a +rank. See `node.path.*`. + ### UDP (`transports.udp.*`) | Parameter | Type | Default | Description | @@ -656,8 +697,9 @@ adding entries and the precedence rules: | `transports.udp.advertise_on_nostr` | bool | `false` | Include this UDP transport in Nostr endpoint adverts. Implicitly forced false when `outbound_only: true`. | | `transports.udp.public` | bool | `false` | If advertised: `true` publishes direct `host:port`; `false` publishes `udp:nat` rendezvous | | `transports.udp.external_addr` | string | *(none)* | Explicit advertise-as override. Bare IP (`"203.0.113.45"` — bind port is appended) or full `host:port`. Takes precedence over the bound address and STUN autodiscovery. Useful when the public IP isn't on a local interface (cloud 1:1 NAT, EIP) or to skip STUN for a deterministic value. | -| `transports.udp.interface` | string | *(none)* | Bind the socket to one interface (e.g. `en0`), making this instance one path. Two instances bound to two interfaces give a peer reachable over both two paths. Linux binds both directions (`SO_BINDTODEVICE`); macOS binds egress only (`IP_BOUND_IF`), so inbound on a wildcard `bind_addr` still arrives from any interface there. Unsupported elsewhere (fails to start). The interface must exist when the daemon starts: unlike an Ethernet transport, an interface-bound UDP instance is not retried when its interface appears later, and it gets no presence or carrier signal while running (see `node.path.*` above). | | `transports.udp.outbound_only` | bool | `false` | Pure-client posture. When `true`, the transport binds to `0.0.0.0:0` (kernel-assigned ephemeral port) regardless of `bind_addr`, refuses inbound handshake msg1, and is never advertised on Nostr regardless of `advertise_on_nostr`. | +| `transports.udp.interface` | string | *(none)* | Bind the socket to one interface (e.g. `en0`), making this instance one path. Two instances bound to two interfaces give a peer reachable over both two paths. Linux binds both directions (`SO_BINDTODEVICE`); macOS binds egress only (`IP_BOUND_IF`), so inbound on a wildcard `bind_addr` still arrives from any interface there. Unsupported elsewhere (fails to start). The interface must exist when the daemon starts: unlike an Ethernet transport, an interface-bound UDP instance is not retried when its interface appears later, and it gets no presence or carrier signal while running (see `node.path.*` above). | +| `transports.udp.role` | string | `normal` | `normal` or `backup`; see above. | | `transports.udp.accept_connections` | bool | `true` | Accept inbound handshake msg1 from new peers. Combine with `outbound_only: false` and `accept_connections: false` (plus `auto_connect` on peer entries) for a node that initiates outbound links but rejects fresh inbound handshakes. The handshake handler carves out msg1 from peers already established on this transport so rekey continues to work. | ### Ethernet (`transports.ethernet.*`) @@ -1032,6 +1074,15 @@ Static peer list. Each entry defines a peer to connect to. | `peers[].auto_reconnect` | bool | `true` | Automatically reconnect after MMP link-dead removal (exponential backoff, unlimited retries) | | `peers[].via_nostr` | bool | `false` | Append Nostr advert-derived endpoints after static addresses for this peer | +**Several addresses, one session.** A peer entry may list addresses on +several transports (`udp/main` and `udp/eth0`, or `udp` and `tor`). All +are dialled; the first handshake to complete makes the session, and a +later one to the same live peer over a transport that has no path yet is +kept as a *path* under that session at both ends, not as a second +session. A configured address whose transport was down at dial time is +added as a path once the transport is up and the peer is live. Paths are +listed by `fipsctl path show `. + **Named UDP instances.** Where several UDP transports are configured under named sub-keys, a peer address can name the one it belongs to by writing the transport field as `udp/`, for example @@ -1282,6 +1333,12 @@ node: heartbeat_interval_secs: 10 link_dead_timeout_secs: 30 # drain_timeout_secs: 2 # bounded Draining phase; absent = 2s + path: + switch_margin: 1.3 # K: active score must exceed best standby's by this + switch_dwell_secs: 2 # D: margin must hold this long + min_samples: 2 # N: RTT samples before a standby is eligible + active_heartbeat_ms: 200 # heartbeat on a path either side sends on + standby_heartbeat_ms: 1000 # heartbeat on a standby path limits: max_connections: 256 max_peers: 128 diff --git a/docs/reference/control-socket.md b/docs/reference/control-socket.md index f4a71943..2fe2f7fe 100644 --- a/docs/reference/control-socket.md +++ b/docs/reference/control-socket.md @@ -173,6 +173,9 @@ not reproduced here to avoid duplicating the source. | `probe_start` | `npub` (bech32) | Admits a diagnostic probe job and returns immediately. `data`: `probe_id`, `npub`, `node_addr`, `display_name`, `budget_ms`. | | `probe_poll` | `probe_id` (integer) | Reports a probe's progress. `data`: `state` (`running` / `done`) and `report`. A terminal job is removed on the poll that observes it, so the report is delivered once. | | `probe_cancel` | `probe_id` (integer) | Runs the probe's terminal actions immediately, without the teardown grace tick. | +| `path_show` | `npub` (bech32) | Every path to the peer. `data`: `peer`, `link_cost`, `link_cost_held`, and `paths[]` with `transport_id`, `transport` (instance name or null), `addr`, `state`, `active`, `remote_active`, `role`, `pinned`, `rx_live_ms_ago`, `tx_live_ms_ago`, `acked_once`, `last_rtt_ms`, `min_rtt_ms`, `rtt_samples`, `etx`, `score`. | +| `path_pin` | `npub` (bech32), `transport` (instance name or numeric id) | Pins this node's traffic to the peer to that transport's path. Applies on the next selection tick. Error if the peer has no path there. | +| `path_unpin` | `npub` (bech32) | Clears the pin. | `connect` on a peer the node is **already connected to** neither tears the live link down nor ignores the address: the address is tried as an alternate diff --git a/docs/reference/wire-formats.md b/docs/reference/wire-formats.md index 818d1073..28363985 100644 --- a/docs/reference/wire-formats.md +++ b/docs/reference/wire-formats.md @@ -21,6 +21,9 @@ The FMP link layer defines the following message types, dispatched by the | 0x31 | LookupResponse | Forwarded — reverse-path via `recent_requests` | | 0x50 | Disconnect | Peer-to-peer (orderly link teardown) | | 0x51 | Heartbeat | Peer-to-peer (link liveness) | +| 0x52 | PathProbe | Peer-to-peer (per-path liveness and path discovery) | +| 0x53 | PathAck | Peer-to-peer (echo of a PathProbe, on the same path) | +| 0x54 | PathClose | Peer-to-peer (a path is going; sent on a surviving path) | Handshake messages travel before encryption is established and are identified by the FMP common-prefix `phase` field rather than a `msg_type` byte @@ -155,6 +158,9 @@ the 1-byte message type and message-specific fields. | 0x31 | LookupResponse | Coordinate discovery response | | 0x50 | Disconnect | Orderly link teardown | | 0x51 | Heartbeat | Link liveness probe | +| 0x52 | PathProbe | Per-path liveness probe; proves a transport as a path | +| 0x53 | PathAck | Echo of a PathProbe, sent back on the same path | +| 0x54 | PathClose | Notice that a path is going, sent on a surviving path | ### Noise IK Message 1 (phase 0x1) @@ -405,6 +411,49 @@ Orderly link teardown with reason code. | 0x07 | Timeout | Heartbeat liveness timeout | | 0xFF | Other | Unspecified reason | +### PathProbe (0x52) and PathAck (0x53) + +A node reachable over more than one transport holds a *path* per +transport under one session +(`docs/design/fips-multi-path-switchover.md`). A PathProbe is sent on a +candidate or standby path — and, once a peer has more than one path, on +the active one as its heartbeat; the receiver, having decrypted it under +the session, has proof the sender is reachable there and answers with a +PathAck **on the same path**. Both share one layout. + +| Offset | Field | Size | Encoding | +| ------ | ----- | ---- | -------- | +| 0 | msg_type | 1 | `0x52` (probe) or `0x53` (ack) | +| 1 | probe_id | 4 | u32 LE — the path's probe sequence; the ack echoes the probe's | +| 5 | flags | 1 | bit 0 `remote_active`: "this path is where I send"; other bits zero | +| 6 | path_id | 4 | u32 LE — the sender's own identifier for the path this travels on | +| 10 | padding | 0..MTU | Zero; ignored by the decoder | + +**Fixed part: 10 bytes.** The first probe on a standby, and one a minute +after on every path, are padded to the link MTU so a medium that passes +small frames and drops large ones never proves itself; the ack echoes the +probe's size. Transport ids are local to each node, so each side learns +the other's `path_id` from its probes and acks; a PathClose names a path +by the *receiver's* id. + +### PathClose (0x54) + +Sent on a surviving path when the sender loses a path to the receiver +(interface gone, carrier lost, operator), so the receiver moves at once +rather than after its own timeout. + +| Offset | Field | Size | Encoding | +| ------ | ----- | ---- | -------- | +| 0 | msg_type | 1 | `0x54` | +| 1 | path_id | 4 | u32 LE — the *receiver's* id for the path being closed | +| 5 | reason | 1 | `0` unspecified, `1` interface gone, `2` carrier lost, `3` operator | + +**Total: 6 bytes.** + +A node that predates these types drops them at debug after +authenticating the frame; a probe to such a node is never acked and the +path never becomes eligible. + ### SenderReport (0x01) Sent by the frame sender to provide interval-based transmission statistics. diff --git a/src/bin/fipsctl.rs b/src/bin/fipsctl.rs index 440c52ea..f15ca403 100644 --- a/src/bin/fipsctl.rs +++ b/src/bin/fipsctl.rs @@ -67,6 +67,11 @@ enum Commands { #[arg(short = 'k', long = "key", conflicts_with = "identity")] key: Option, }, + /// Paths to a peer: show them, or pin traffic to one + Path { + #[command(subcommand)] + what: PathCommands, + }, /// Connect to a peer Connect { /// Peer identifier: npub (bech32) or hostname from /etc/fips/hosts @@ -190,6 +195,27 @@ enum ShowCommands { NativeFlows, } +#[derive(Subcommand, Debug)] +enum PathCommands { + /// Show every path to a peer, per direction + Show { + /// Peer npub or hostname + peer: String, + }, + /// Pin the peer's traffic to one transport until unpinned or the path dies + Pin { + /// Peer npub or hostname + peer: String, + /// Transport instance name (as configured) or numeric id + transport: String, + }, + /// Clear the pin + Unpin { + /// Peer npub or hostname + peer: String, + }, +} + #[derive(Subcommand, Debug)] enum AclCommands { /// Loaded peer ACL state @@ -597,6 +623,23 @@ fn main() { let npub = resolve_peer(peer); build_command("disconnect", serde_json::json!({"npub": npub})) } + Commands::Path { what } => match what { + PathCommands::Show { peer } => { + let npub = resolve_peer(peer); + build_command("path_show", serde_json::json!({"npub": npub})) + } + PathCommands::Pin { peer, transport } => { + let npub = resolve_peer(peer); + build_command( + "path_pin", + serde_json::json!({"npub": npub, "transport": transport}), + ) + } + PathCommands::Unpin { peer } => { + let npub = resolve_peer(peer); + build_command("path_unpin", serde_json::json!({"npub": npub})) + } + }, Commands::Probe { target, json, diff --git a/src/config/mod.rs b/src/config/mod.rs index 2c864736..582dd628 100644 --- a/src/config/mod.rs +++ b/src/config/mod.rs @@ -39,13 +39,13 @@ pub use gateway::{ConntrackConfig, GatewayConfig, GatewayDnsConfig, PortForward, pub use node::{ BloomConfig, BuffersConfig, CacheConfig, ControlConfig, LimitsConfig, LookupConfig, MmpConfig, NativeApiConfig, NetmonConfig, NodeConfig, NostrRendezvousConfig, NostrRendezvousPolicy, - RateLimitConfig, RekeyConfig, RendezvousConfig, RetryConfig, SessionConfig, SessionMmpConfig, - TreeConfig, + PathConfig, RateLimitConfig, RekeyConfig, RendezvousConfig, RetryConfig, SessionConfig, + SessionMmpConfig, TreeConfig, }; pub use peer::{ConnectPolicy, PeerAddress, PeerConfig, TransportSpec}; pub use transport::{ BleConfig, DirectoryServiceConfig, EthernetConfig, NymConfig, TcpConfig, TorConfig, - TransportInstances, TransportsConfig, UdpConfig, + TransportInstances, TransportRole, TransportsConfig, UdpConfig, }; /// Default config filename. diff --git a/src/config/node.rs b/src/config/node.rs index 25e6695e..70a4e0f5 100644 --- a/src/config/node.rs +++ b/src/config/node.rs @@ -1192,6 +1192,86 @@ impl BuffersConfig { // ECN Congestion Signaling // ============================================================================ +// ============================================================================ +// Path Selection +// ============================================================================ + +/// Multi-path switchover (`node.path.*`). +/// +/// A peer reachable over more than one transport keeps one Noise session and +/// moves its traffic between paths on failure or degradation. Selection is +/// measured, not configured: a path's score is `etx × (1 + min_rtt_ms / 100)` +/// from its own probes. These knobs bound *when* a measured difference is +/// acted on. Their defaults are placeholders to calibrate against +/// `testing/chaos`, not values to reason about +/// (`docs/design/fips-multi-path-switchover.md`, "Calibration"). +#[derive(Debug, Clone, Serialize, Deserialize)] +pub struct PathConfig { + /// Discretionary switch margin `K` (`node.path.switch_margin`): the + /// active path's score must exceed the best standby's by this factor. + /// This is what encodes the fail-back policy: a cable returning under + /// wifi is 1.00 against 1.10, under the margin, so traffic stays. + #[serde(default = "PathConfig::default_switch_margin")] + pub switch_margin: f64, + + /// Discretionary switch dwell `D` in seconds + /// (`node.path.switch_dwell_secs`): the margin must hold this long. + #[serde(default = "PathConfig::default_switch_dwell_secs")] + pub switch_dwell_secs: u64, + + /// RTT samples `N` a standby needs before it is eligible + /// (`node.path.min_samples`). + #[serde(default = "PathConfig::default_min_samples")] + pub min_samples: u32, + + /// Heartbeat interval on a path that either side sends on, in ms + /// (`node.path.active_heartbeat_ms`). Standby paths use + /// `node.path.standby_heartbeat_ms`. A floor: on a path whose round trip + /// is longer than this (Tor, Nym) the interval and the echo timeout + /// stretch with the measured round trip, so one setting serves a cable + /// and a circuit. + #[serde(default = "PathConfig::default_active_heartbeat_ms")] + pub active_heartbeat_ms: u64, + + /// Heartbeat interval on a standby path, in ms + /// (`node.path.standby_heartbeat_ms`). A standby is only as warm as its + /// last echo: this bounds how stale a "proven" standby can be when the + /// active path dies and selection reaches for it. Stretched by the + /// path's round trip like the active interval. + #[serde(default = "PathConfig::default_standby_heartbeat_ms")] + pub standby_heartbeat_ms: u64, +} + +impl Default for PathConfig { + fn default() -> Self { + Self { + switch_margin: Self::default_switch_margin(), + switch_dwell_secs: Self::default_switch_dwell_secs(), + min_samples: Self::default_min_samples(), + active_heartbeat_ms: Self::default_active_heartbeat_ms(), + standby_heartbeat_ms: Self::default_standby_heartbeat_ms(), + } + } +} + +impl PathConfig { + fn default_switch_margin() -> f64 { + 1.3 + } + fn default_switch_dwell_secs() -> u64 { + 2 + } + fn default_min_samples() -> u32 { + 2 + } + fn default_active_heartbeat_ms() -> u64 { + 200 + } + fn default_standby_heartbeat_ms() -> u64 { + 1000 + } +} + /// Rekey / session rekeying configuration (`node.rekey.*`). /// /// Controls periodic full rekey for both FMP (link layer) and FSP @@ -1359,6 +1439,10 @@ pub struct NodeConfig { #[serde(default)] pub tree: TreeConfig, + /// Multi-path switchover (`node.path.*`). + #[serde(default)] + pub path: PathConfig, + /// Bloom filter (`node.bloom.*`). #[serde(default)] pub bloom: BloomConfig, @@ -1429,6 +1513,7 @@ impl Default for NodeConfig { rendezvous: RendezvousConfig::default(), discovery: None, tree: TreeConfig::default(), + path: PathConfig::default(), bloom: BloomConfig::default(), session: SessionConfig::default(), buffers: BuffersConfig::default(), diff --git a/src/config/transport.rs b/src/config/transport.rs index d9c7cd17..311b597b 100644 --- a/src/config/transport.rs +++ b/src/config/transport.rs @@ -41,6 +41,23 @@ const DEFAULT_UDP_RECV_BUF: usize = 2 * 1024 * 1024; /// Default UDP send buffer size (2 MB). const DEFAULT_UDP_SEND_BUF: usize = 2 * 1024 * 1024; +/// What a transport is *for*, as far as path selection is concerned. +/// +/// Not a rank. Selection between paths is measured, never configured +/// (`docs/design/fips-multi-path-switchover.md` §8); this is the one +/// statement an operator can make about a transport's purpose: a `backup` +/// transport never carries a peer's traffic while any non-backup path to +/// that peer is eligible. "Drop BLE when something better is stable." +#[derive(Debug, Default, Clone, Copy, PartialEq, Eq, Serialize, Deserialize)] +#[serde(rename_all = "lowercase")] +pub enum TransportRole { + /// Carries traffic whenever it is the best measured path. + #[default] + Normal, + /// Carries traffic only while no `normal` path is eligible. + Backup, +} + /// UDP transport instance configuration. #[derive(Debug, Clone, Default, Serialize, Deserialize)] #[serde(deny_unknown_fields)] @@ -54,6 +71,10 @@ pub struct UdpConfig { #[serde(default, skip_serializing_if = "Option::is_none")] pub interface: Option, + /// Path-selection role (`role: normal | backup`). Default: normal. + #[serde(default, skip_serializing_if = "Option::is_none")] + pub role: Option, + /// Bind address (`bind_addr`). Defaults to "0.0.0.0:2121". /// /// When `outbound_only = true`, this field is ignored and the transport @@ -118,6 +139,11 @@ pub struct UdpConfig { } impl UdpConfig { + /// Path-selection role. Default: normal. + pub fn role(&self) -> TransportRole { + self.role.unwrap_or_default() + } + /// Get the bind address, using default if not configured. /// /// When `outbound_only = true`, returns `0.0.0.0:0` so the kernel picks @@ -269,6 +295,10 @@ const MIN_BEACON_INTERVAL_SECS: u64 = 10; #[derive(Debug, Clone, Default, Serialize, Deserialize)] #[serde(deny_unknown_fields)] pub struct EthernetConfig { + /// Path-selection role (`role: normal | backup`). Default: normal. + #[serde(default, skip_serializing_if = "Option::is_none")] + pub role: Option, + /// Network interface name (e.g., "eth0", "enp3s0"). Required. pub interface: String, @@ -333,6 +363,11 @@ pub struct EthernetConfig { } impl EthernetConfig { + /// Path-selection role. Default: normal. + pub fn role(&self) -> TransportRole { + self.role.unwrap_or_default() + } + /// Get the EtherType, using default if not configured. pub fn ethertype(&self) -> u16 { self.ethertype.unwrap_or(DEFAULT_ETHERNET_ETHERTYPE) @@ -407,6 +442,10 @@ const DEFAULT_TCP_MAX_INBOUND: usize = 256; #[derive(Debug, Clone, Default, Serialize, Deserialize)] #[serde(deny_unknown_fields)] pub struct TcpConfig { + /// Path-selection role (`role: normal | backup`). Default: normal. + #[serde(default, skip_serializing_if = "Option::is_none")] + pub role: Option, + /// Listen address (e.g., "0.0.0.0:443"). If not set, outbound-only. #[serde(default, skip_serializing_if = "Option::is_none")] pub bind_addr: Option, @@ -457,6 +496,11 @@ pub struct TcpConfig { } impl TcpConfig { + /// Path-selection role. Default: normal. + pub fn role(&self) -> TransportRole { + self.role.unwrap_or_default() + } + /// Get the default MTU. pub fn mtu(&self) -> u16 { self.mtu.unwrap_or(DEFAULT_TCP_MTU) @@ -555,6 +599,10 @@ const DEFAULT_TOR_ADVERTISED_PORT: u16 = 443; #[derive(Debug, Clone, Default, Serialize, Deserialize)] #[serde(deny_unknown_fields)] pub struct TorConfig { + /// Path-selection role (`role: normal | backup`). Default: normal. + #[serde(default, skip_serializing_if = "Option::is_none")] + pub role: Option, + /// Tor access mode: "socks5", "control_port", or "directory". /// Default: "socks5". #[serde(default, skip_serializing_if = "Option::is_none")] @@ -652,6 +700,11 @@ impl DirectoryServiceConfig { } impl TorConfig { + /// Path-selection role. Default: normal. + pub fn role(&self) -> TransportRole { + self.role.unwrap_or_default() + } + /// Get the access mode. Default: "socks5". pub fn mode(&self) -> &str { self.mode.as_deref().unwrap_or("socks5") @@ -739,6 +792,10 @@ const DEFAULT_BLE_PROBE_COOLDOWN_SECS: u64 = 30; #[derive(Debug, Clone, Default, Serialize, Deserialize)] #[serde(deny_unknown_fields)] pub struct BleConfig { + /// Path-selection role (`role: normal | backup`). Default: normal. + #[serde(default, skip_serializing_if = "Option::is_none")] + pub role: Option, + /// HCI adapter name (e.g., "hci0"). Required. #[serde(default, skip_serializing_if = "Option::is_none")] pub adapter: Option, @@ -788,6 +845,11 @@ pub struct BleConfig { } impl BleConfig { + /// Path-selection role. Default: normal. + pub fn role(&self) -> TransportRole { + self.role.unwrap_or_default() + } + /// Get the adapter name. Default: "hci0". pub fn adapter(&self) -> &str { self.adapter.as_deref().unwrap_or("hci0") @@ -868,6 +930,10 @@ const DEFAULT_NYM_STARTUP_TIMEOUT_SECS: u64 = 120; #[derive(Debug, Clone, Default, Serialize, Deserialize)] #[serde(deny_unknown_fields)] pub struct NymConfig { + /// Path-selection role (`role: normal | backup`). Default: normal. + #[serde(default, skip_serializing_if = "Option::is_none")] + pub role: Option, + /// SOCKS5 proxy address (host:port). Defaults to "127.0.0.1:1080". #[serde(default, skip_serializing_if = "Option::is_none")] pub socks5_addr: Option, @@ -887,6 +953,11 @@ pub struct NymConfig { } impl NymConfig { + /// Path-selection role. Default: normal. + pub fn role(&self) -> TransportRole { + self.role.unwrap_or_default() + } + /// Get the SOCKS5 proxy address. Default: "127.0.0.1:1080". pub fn socks5_addr(&self) -> &str { self.socks5_addr diff --git a/src/control/commands.rs b/src/control/commands.rs index c423887d..7d44c0aa 100644 --- a/src/control/commands.rs +++ b/src/control/commands.rs @@ -16,6 +16,9 @@ pub async fn dispatch(node: &mut Node, command: &str, params: Option<&Value>) -> "probe_start" => probe_start(node, params).await, "probe_poll" => probe_poll(node, params), "probe_cancel" => probe_cancel(node, params).await, + "path_show" => path_show(node, params), + "path_pin" => path_pin(node, params), + "path_unpin" => path_unpin(node, params), _ => Response::error(format!("unknown command: {command}")), } } @@ -124,6 +127,64 @@ fn probe_id(params: Option<&Value>) -> Option { params?.get("probe_id").and_then(|v| v.as_u64()) } +fn npub_param<'a>(params: Option<&'a Value>, what: &str) -> Result<&'a str, Response> { + params + .and_then(|p| p.get("npub")) + .and_then(|v| v.as_str()) + .ok_or_else(|| Response::error(format!("missing 'npub' parameter for {what}"))) +} + +/// Show every path to a peer. +/// +/// Params: `{"npub": "npub1..."}` +fn path_show(node: &mut Node, params: Option<&Value>) -> Response { + let npub = match npub_param(params, "path_show") { + Ok(v) => v, + Err(r) => return r, + }; + match node.api_path_show(npub) { + Ok(data) => Response::ok(data), + Err(msg) => Response::error(msg), + } +} + +/// Pin a peer's traffic to one transport. +/// +/// Params: `{"npub": "npub1...", "transport": "cable"}` +fn path_pin(node: &mut Node, params: Option<&Value>) -> Response { + let npub = match npub_param(params, "path_pin") { + Ok(v) => v, + Err(r) => return r, + }; + let transport = match params + .and_then(|p| p.get("transport")) + .and_then(|v| v.as_str()) + { + Some(v) => v, + None => return Response::error("missing 'transport' parameter"), + }; + debug!(npub = %npub, transport = %transport, "API path pin requested"); + match node.api_path_pin(npub, transport) { + Ok(data) => Response::ok(data), + Err(msg) => Response::error(msg), + } +} + +/// Clear a peer's path pin. +/// +/// Params: `{"npub": "npub1..."}` +fn path_unpin(node: &mut Node, params: Option<&Value>) -> Response { + let npub = match npub_param(params, "path_unpin") { + Ok(v) => v, + Err(r) => return r, + }; + debug!(npub = %npub, "API path unpin requested"); + match node.api_path_unpin(npub) { + Ok(data) => Response::ok(data), + Err(msg) => Response::error(msg), + } +} + #[cfg(test)] mod tests { use super::*; diff --git a/src/control/queries.rs b/src/control/queries.rs index b85dd3bf..98eeb828 100644 --- a/src/control/queries.rs +++ b/src/control/queries.rs @@ -306,7 +306,7 @@ pub fn show_peers(node: &Node) -> Value { if any_peer_has_srtt && !peer.has_srtt() { None } else { - Some(coords.depth() as f64 + peer.link_cost()) + Some(coords.depth() as f64 + peer.link_cost(crate::time::mono_ms())) } }); peer_json["effective_depth"] = match effective_depth { @@ -2746,6 +2746,7 @@ mod tests { // An interface no host has, so presence is deterministically absent // and carrier deterministically false on every machine this runs on. let config = EthernetConfig { + role: None, interface: "fips-absent-x0".to_string(), ethertype: None, mtu: None, diff --git a/src/node/dataplane/dispatch.rs b/src/node/dataplane/dispatch.rs index c1ab8512..b2266719 100644 --- a/src/node/dataplane/dispatch.rs +++ b/src/node/dataplane/dispatch.rs @@ -2,17 +2,24 @@ use crate::NodeAddr; use crate::node::Node; +use crate::transport::{TransportAddr, TransportId}; use tracing::{debug, info, trace}; impl Node { /// Dispatch a decrypted link message to the appropriate handler. /// /// Link messages are protocol messages exchanged between authenticated peers. + /// + /// `arrival` is the transport and address the frame came in on. Most + /// handlers do not care: the frame was authenticated by the session, + /// which is per peer, not per path. The path probe exchange does: it + /// answers on the path the probe used, and that is the whole point. pub(in crate::node) async fn dispatch_link_message( &mut self, from: &NodeAddr, plaintext: &[u8], ce_flag: bool, + arrival: (TransportId, &TransportAddr), ) { if plaintext.is_empty() { return; @@ -58,6 +65,18 @@ impl Node { // Heartbeat — no-op, last_recv_time already updated by record_recv() trace!(peer = %self.peer_display_name(from), "Received heartbeat"); } + 0x52 => { + // PathProbe + self.handle_path_probe(from, payload, arrival).await; + } + 0x53 => { + // PathAck + self.handle_path_ack(from, payload, arrival); + } + 0x54 => { + // PathClose + self.handle_path_close(from, payload).await; + } _ => { debug!(msg_type = msg_type, "Unknown link message type"); } diff --git a/src/node/dataplane/encrypted.rs b/src/node/dataplane/encrypted.rs index 3dfb7319..f425bcce 100644 --- a/src/node/dataplane/encrypted.rs +++ b/src/node/dataplane/encrypted.rs @@ -277,6 +277,7 @@ impl Node { } address_changed = peer.set_current_addr(packet.transport_id, packet.remote_addr.clone()); + peer.note_path_rx(packet.transport_id, now_ms); peer.link_stats_mut() .record_recv(packet.data.len(), packet.timestamp_ms); peer.touch(packet.timestamp_ms); @@ -298,8 +299,13 @@ impl Node { let _ = address_changed; // Dispatch to link message handler - self.dispatch_link_message(&node_addr, link_message, ce_flag) - .await; + self.dispatch_link_message( + &node_addr, + link_message, + ce_flag, + (packet.transport_id, &packet.remote_addr), + ) + .await; } /// Log a decryption failure with replay suppression. @@ -377,6 +383,7 @@ impl Node { if let Some(peer) = self.peers.get_mut(node_addr) { peer.reset_decrypt_failures(); address_changed = peer.set_current_addr(transport_id, remote_addr.clone()); + peer.note_path_rx(transport_id, now_ms); peer.link_stats_mut() .record_recv(packet_len, packet_timestamp_ms); peer.touch(packet_timestamp_ms); @@ -398,8 +405,13 @@ impl Node { let _ = address_changed; } let link_message = &fmp_plaintext[INNER_TIMESTAMP_LEN..]; - self.dispatch_link_message(node_addr, link_message, ce_flag) - .await; + self.dispatch_link_message( + node_addr, + link_message, + ce_flag, + (transport_id, remote_addr), + ) + .await; } /// Process a decrypt-worker bounce (FMP plaintext only — the @@ -546,10 +558,10 @@ impl Node { node_addr: &crate::NodeAddr, transport_id: crate::transport::TransportId, ) { - let on_path = self.peers.get(node_addr).is_some_and(|peer| { - peer.transport_id() - .is_none_or(|bound| bound == transport_id) - }); + let on_path = self + .peers + .get(node_addr) + .is_some_and(|peer| peer.paths().is_empty() || peer.path_on(transport_id).is_some()); if !on_path { trace!( peer = %self.peer_display_name(node_addr), diff --git a/src/node/dataplane/rx_loop.rs b/src/node/dataplane/rx_loop.rs index bb26d78a..60df8704 100644 --- a/src/node/dataplane/rx_loop.rs +++ b/src/node/dataplane/rx_loop.rs @@ -121,6 +121,13 @@ impl Node { let tick_period = Duration::from_secs(self.config().node.tick_interval_secs); let mut tick = tokio::time::interval(tick_period); + // The fast path tick: per-path heartbeats and the carrier edge. Its + // own timer because the maintenance tick is seconds and a dead path + // is meant to be noticed inside one. + let mut path_tick = tokio::time::interval(Duration::from_millis( + self.config().node.path.active_heartbeat_ms.max(50), + )); + path_tick.set_missed_tick_behavior(tokio::time::MissedTickBehavior::Delay); // Set up control socket channel let (control_tx, mut control_rx) = @@ -410,9 +417,14 @@ impl Node { // Not policy-filtered: whether an interface's absence // is normal is a statement about node *health*, not // about whether the routes over it still work. - if !edge.present { + if edge.present { + // The medium came back: whatever went unanswered + // on it before says nothing about now, so the + // next discovery tick may probe it at once. + self.reset_probe_backoff_on_transport(edge.transport_id); + } else { let reaped = - self.reap_peers_on_transport(edge.transport_id).await; + self.withdraw_transport(edge.transport_id).await; if reaped > 0 { info!( transport_id = %edge.transport_id, @@ -490,6 +502,9 @@ impl Node { ).await; let _ = response_tx.send(response); } + _ = path_tick.tick() => { + self.run_path_heartbeats().await; + } deadline = tick.tick() => { // Tick-body instrumentation. The gate is read ONCE per tick // into `instr_on`, which is then passed explicitly to every diff --git a/src/node/handlers/handshake.rs b/src/node/handlers/handshake.rs index f63637c7..072a42b0 100644 --- a/src/node/handlers/handshake.rs +++ b/src/node/handlers/handshake.rs @@ -206,9 +206,13 @@ impl Node { { return true; } - if self.peers.values().any(|p| { - p.transport_id() == Some(transport_id) && p.current_addr() == Some(remote_addr) - }) { + // Any path, not only the one we send on: a rekey msg1 arrives on the + // path the *peer* sends on. + if self + .peers + .values() + .any(|p| p.is_reachable_at(transport_id, remote_addr)) + { return true; } false @@ -281,9 +285,7 @@ impl Node { // yields an identity when it matches. self.peers .values() - .find(|p| { - p.transport_id() == Some(transport_id) && p.current_addr() == Some(remote_addr) - }) + .find(|p| p.is_reachable_at(transport_id, remote_addr)) .map(|p| Msg1Waiver::Expect(*p.node_addr())) .unwrap_or(Msg1Waiver::Unattributed) } @@ -1804,6 +1806,13 @@ impl Node { self.config().node.tree.announce_min_interval_ms, ); + new_peer.set_path_role( + transport_id, + self.transports + .get(&transport_id) + .map(|t| t.role()) + .unwrap_or_default(), + ); self.peers.insert(peer_node_addr, new_peer); self.peers_by_index .insert(our_index.as_u32(), peer_node_addr); @@ -1914,6 +1923,13 @@ impl Node { new_peer.set_last_tree_announce_sent_ms(ts); } + new_peer.set_path_role( + transport_id, + self.transports + .get(&transport_id) + .map(|t| t.role()) + .unwrap_or_default(), + ); self.peers.insert(peer_node_addr, new_peer); self.peers_by_index .insert(our_index.as_u32(), peer_node_addr); diff --git a/src/node/handlers/mmp.rs b/src/node/handlers/mmp.rs index 43c0af8a..234e276d 100644 --- a/src/node/handlers/mmp.rs +++ b/src/node/handlers/mmp.rs @@ -35,6 +35,14 @@ use tracing::{debug, info, trace, warn}; /// bounded this can come down to the tick. const HEARTBEAT_RETRY_INTERVAL: Duration = Duration::from_secs(2); +/// How long a `Dead` path keeps its history before it is forgotten. +/// +/// Presence flaps on a cable (dock sleep, autoneg bounce) are what the +/// binder's churn guard exists for; a path that came back inside this window +/// is re-probed with its RTT intact rather than measured from nothing. See +/// `docs/design/fips-multi-path-switchover.md` §6. +const DEAD_PATH_GRACE_MS: u64 = 5 * 60 * 1000; + /// Decide whether a peer is due a heartbeat, from the two timestamps it keeps. /// /// Two gates rather than one. `sent` is when a heartbeat last *landed*, and it @@ -200,6 +208,8 @@ impl Node { // Get session timestamp before taking mutable borrow on MMP let our_timestamp_ms = peer.session_elapsed_ms(); + // One report closer to releasing a post-switch cost hold. + peer.note_receiver_report(); let Some(mmp) = peer.mmp_mut() else { return; @@ -244,7 +254,7 @@ impl Node { .peers .iter() .filter(|(_, p)| p.has_srtt()) - .map(|(a, p)| (*a, p.link_cost())) + .map(|(a, p)| (*a, p.link_cost(now_ms))) .collect(); // Wall-clock seconds for the escaping declaration timestamp; // monotonic ms for the flap-dampening / hold-down timers. @@ -549,6 +559,11 @@ impl Node { let actions = self.mmp.plan_heartbeats(&snapshots); + // Dead-path history expires here, on the same cadence as liveness. + for peer in self.peers.values_mut() { + peer.prune_dead_paths(now_ms, DEAD_PATH_GRACE_MS); + } + // Wall-clock basis for reconnect scheduling, sourced once (as before). let now_ms = std::time::SystemTime::now() .duration_since(std::time::UNIX_EPOCH) @@ -601,68 +616,25 @@ impl Node { } } - /// Reap every active peer reachable only through `transport_id`. - /// - /// Called on a transport's detach edge. Until this existed, losing an - /// interface withdrew nothing: the peers stayed in the registry, the - /// routes through them stayed selectable, and this node kept advertising - /// reachability it no longer had — so transit traffic was dropped in - /// silence, and other nodes kept routing toward us for those destinations, - /// until the liveness reaper noticed up to `link_dead_timeout_secs` later. - /// Measured on real hardware that was 27 seconds of routing through a link - /// that had already gone, with four alternative peers available the whole - /// time. - /// - /// The detach edge is both earlier and more certain than inactivity, so it - /// is the better trigger. This routes through the same - /// [`Self::route_link_dead`] the liveness reaper uses rather than - /// open-coding a second teardown — every consequence of losing a peer - /// (sessions, path MTU, session indices, the link, the control machine, - /// tree cleanup and re-announce, bloom withdrawal) already hangs off that - /// one path, and a parallel one would drift from it. - /// - /// Deliberately undamped. A flapping interface cannot drive a reap storm - /// through here, because `ChurnGuard` stops publishing presence edges - /// after three short-lived bindings and does not resume until one lasts — - /// so the edges this reacts to are already rate-limited at the source, and - /// a second damper here would only add a way for the two to disagree. - /// - /// Returns how many peers were reaped. - pub(in crate::node) async fn reap_peers_on_transport( + /// Reap one peer whose last path went away. `now_ms` is the wall-clock + /// reconnect basis, as for the liveness reaper. + pub(in crate::node) async fn reap_peer_without_path( &mut self, + node_addr: NodeAddr, transport_id: TransportId, - ) -> usize { - let doomed: Vec = self - .peers - .iter() - .filter(|(_, peer)| peer.transport_id() == Some(transport_id)) - .map(|(node_addr, _)| *node_addr) - .collect(); - - if doomed.is_empty() { - return 0; - } - - let now_ms = std::time::SystemTime::now() - .duration_since(std::time::UNIX_EPOCH) - .map(|d| d.as_millis() as u64) - .unwrap_or(0); - - let reaped = doomed.len(); - for node_addr in doomed { - debug!( - peer = %self.peer_display_name(&node_addr), - %transport_id, - "Removing peer: its interface went away" - ); - self.route_link_dead(node_addr, now_ms).await; - } - reaped + now_ms: u64, + ) { + debug!( + peer = %self.peer_display_name(&node_addr), + %transport_id, + "Removing peer: its interface went away" + ); + self.route_link_dead(node_addr, now_ms).await; } /// Route a link-dead reap through the peer machine + executor. Two callers /// decide: the tick sweep's `plan_heartbeats` batch emits a `ReapPeer` for a - /// peer that has gone quiet, and [`Self::reap_peers_on_transport`] withdraws + /// peer that has gone quiet, and [`Self::withdraw_transport`] reaps /// a transport's peers when its interface goes away. Mirrors /// [`route_rekey_cadence`](Node::route_rekey_cadence): the shell has already /// decided by the time this runs, so the machine only CONSUMES the decision diff --git a/src/node/handlers/mod.rs b/src/node/handlers/mod.rs index 607f160b..cfc26c33 100644 --- a/src/node/handlers/mod.rs +++ b/src/node/handlers/mod.rs @@ -6,6 +6,7 @@ mod mmp; mod native; pub(in crate::node) use native::PendingNative; pub(in crate::node) mod netmon; +pub(in crate::node) mod path; pub(crate) mod probe; // Widened from private by the rekey drain cap: `node::session` calls // `rekey::drain_max_retention_ms` to bound how long a superseded epoch is diff --git a/src/node/handlers/path.rs b/src/node/handlers/path.rs new file mode 100644 index 00000000..ea3fe45a --- /dev/null +++ b/src/node/handlers/path.rs @@ -0,0 +1,775 @@ +//! Path probe / path ack: adding a second path to a peer under the session +//! it already has. +//! +//! `docs/design/fips-multi-path-switchover.md` §4. A probe is an ordinary +//! encrypted frame sent on a candidate transport. The receiver, having +//! decrypted it against the session found by index, has proof the peer is +//! reachable on that `(transport, addr)`: it adds the path as `Probing`, +//! marks it `rx_live`, and answers with an ack **on that same path**. The +//! prober's receipt of the ack proves the reverse direction: the path goes +//! `Live`, `tx_live`, and takes an RTT sample. No handshake, no new key +//! material, no index allocation. +//! +//! This is also where the active path moves: selection on the fast tick +//! (`run_path_selection`), withdrawal when a transport goes +//! (`withdraw_transport`), a peer's `PathClose`, and the switch side +//! effects (`apply_path_switch`). The carrier edge and the per-path +//! heartbeats that feed selection live here too. + +use crate::NodeAddr; +use crate::node::Node; +use crate::peer::{HeartbeatTiming, PathPolicy, PathSwitch, PathWithdrawal}; +use crate::proto::link::{PathClose, PathCloseReason, PathMessage}; +use crate::transport::{TransportAddr, TransportId}; +use tracing::{debug, info, trace}; + +impl Node { + /// The selection knobs, from `node.path.*`. + pub(in crate::node) fn path_policy(&self) -> PathPolicy { + let cfg = &self.config().node.path; + let standby_ms = cfg.standby_heartbeat_ms.max(1); + PathPolicy { + margin: cfg.switch_margin, + dwell_ms: cfg.switch_dwell_secs.saturating_mul(1000), + min_samples: cfg.min_samples, + rtt_window_ms: standby_ms + .saturating_mul(u64::from(cfg.min_samples).max(1)) + .saturating_mul(2) + .max(30_000), + } + } + + /// The role of the transport `transport_id`, or `Normal` if it is not + /// registered. + fn transport_role(&self, transport_id: TransportId) -> crate::config::TransportRole { + self.transports + .get(&transport_id) + .map(|t| t.role()) + .unwrap_or_default() + } + + /// Run selection for every peer. Called from the tick. A switch here is + /// discretionary or pinned, or mandatory after a `Suspect` mark that + /// nothing else acted on; the presence edge runs its own. + pub(in crate::node) fn run_path_selection(&mut self) { + let policy = self.path_policy(); + let now_ms = crate::time::mono_ms(); + let switches: Vec<(NodeAddr, PathSwitch)> = self + .peers + .iter_mut() + .filter_map(|(addr, peer)| peer.select_path(now_ms, &policy).map(|s| (*addr, s))) + .collect(); + for (node_addr, switch) in switches { + info!( + peer = %self.peer_display_name(&node_addr), + from_transport = %switch.from.0, + to_transport = %switch.to.0, + to_addr = %switch.to.1, + reason = ?switch.reason, + "Path switched, session kept" + ); + self.apply_path_switch(&node_addr, switch.to); + } + } + + /// Pin a peer's traffic to its path on `transport_id`. Applies on the + /// next selection run. `false` if the peer or the path is unknown. + pub(crate) fn pin_peer_path( + &mut self, + node_addr: &NodeAddr, + transport_id: TransportId, + ) -> bool { + self.peers + .get_mut(node_addr) + .is_some_and(|p| p.pin_path(transport_id)) + } + + /// Clear a peer's pin. `false` if the peer is unknown. + pub(crate) fn unpin_peer_path(&mut self, node_addr: &NodeAddr) -> bool { + match self.peers.get_mut(node_addr) { + Some(p) => { + p.unpin_paths(); + true + } + None => false, + } + } + + /// The peer named by `npub`, if it is an active peer. + pub(in crate::node) fn resolve_peer_npub(&self, npub: &str) -> Result { + let identity = crate::PeerIdentity::from_npub(npub) + .map_err(|e| format!("invalid npub '{npub}': {e}"))?; + let node_addr = *identity.node_addr(); + if !self.peers.contains_key(&node_addr) { + return Err(format!("peer not found: {npub}")); + } + Ok(node_addr) + } + + /// A transport named by its instance name (`cable`, `main`) or its + /// numeric id. + fn resolve_transport_name(&self, name: &str) -> Result { + if let Some((id, _)) = self.transports.iter().find(|(_, t)| t.name() == Some(name)) { + return Ok(*id); + } + if let Ok(n) = name.parse::() + && self.transports.contains_key(&TransportId::new(n)) + { + return Ok(TransportId::new(n)); + } + Err(format!("transport not found: {name}")) + } + + /// `fipsctl path show `: every path to the peer, per direction. + pub(crate) fn api_path_show(&self, npub: &str) -> Result { + let node_addr = self.resolve_peer_npub(npub)?; + let peer = &self.peers[&node_addr]; + let now_ms = crate::time::mono_ms(); + let active = peer.transport_id(); + let paths: Vec = peer + .paths() + .iter() + .map(|path| { + let ago = |at: Option| at.map(|t| now_ms.saturating_sub(t)); + serde_json::json!({ + "transport_id": path.transport_id().as_u32(), + "transport": self + .transports + .get(&path.transport_id()) + .and_then(|t| t.name().map(str::to_string)), + "addr": path.addr().to_string(), + "state": format!("{:?}", path.state()).to_lowercase(), + "active": Some(path.transport_id()) == active, + "remote_active": path.remote_active(), + "role": format!("{:?}", path.role()).to_lowercase(), + "pinned": path.pinned(), + "rx_live_ms_ago": ago(path.rx_live_at_ms()), + "tx_live_ms_ago": ago(path.tx_live_at_ms()), + "acked_once": path.acked_once(), + "last_rtt_ms": path.last_rtt_ms(), + "min_rtt_ms": path.min_rtt_ms(), + "rtt_samples": path.rtt_samples(), + "etx": path.etx(), + "score": path.score(), + }) + }) + .collect(); + Ok(serde_json::json!({ + "peer": npub, + "link_cost": peer.link_cost(now_ms), + "link_cost_held": peer.link_cost_held(now_ms), + "paths": paths, + })) + } + + /// `fipsctl path pin `. + pub(crate) fn api_path_pin( + &mut self, + npub: &str, + transport: &str, + ) -> Result { + let node_addr = self.resolve_peer_npub(npub)?; + let transport_id = self.resolve_transport_name(transport)?; + if !self.pin_peer_path(&node_addr, transport_id) { + return Err(format!("peer {npub} has no path on transport {transport}")); + } + info!(peer = %self.peer_display_name(&node_addr), %transport_id, "Path pinned by operator"); + Ok(serde_json::json!({ "pinned": transport_id.as_u32() })) + } + + /// `fipsctl path unpin `. + pub(crate) fn api_path_unpin(&mut self, npub: &str) -> Result { + let node_addr = self.resolve_peer_npub(npub)?; + self.unpin_peer_path(&node_addr); + info!(peer = %self.peer_display_name(&node_addr), "Path unpinned by operator"); + Ok(serde_json::json!({ "pinned": serde_json::Value::Null })) + } + + /// A live peer beaconed on `transport_id` at `remote_addr`, a transport + /// we hold no path to it over: add the path as `Probing`. + /// + /// Nothing is sent here. The heartbeat tick is the one issuer of + /// probes: it picks the new path up within one fast interval, probes + /// it full-size, and applies the discovery backoff if the peer never + /// answers (an old node that drops `0x52` at debug), so the path is + /// probed at the capped cadence and never becomes eligible. The backoff + /// is reset when the transport's presence cycles. + pub(in crate::node) fn add_path_candidate( + &mut self, + node_addr: NodeAddr, + transport_id: TransportId, + remote_addr: TransportAddr, + ) { + let role = self.transport_role(transport_id); + let Some(peer) = self.peers.get_mut(&node_addr) else { + return; + }; + if peer.transport_id() == Some(transport_id) { + // The active path: the handshake proved it. + return; + } + let was_new = peer.path_on(transport_id).is_none(); + peer.add_path(transport_id, remote_addr.clone()) + .set_role(role); + if was_new { + debug!( + peer = %self.peer_display_name(&node_addr), + transport_id = %transport_id, + remote_addr = %remote_addr, + "Peer beaconed on a new transport; path added, probing" + ); + } + } + + /// Tests: add the path and probe it at once, as one heartbeat tick + /// would, without the tick's other sends. Uses the test-only + /// `take_probe`, which honours the per-path backoff. + #[cfg(test)] + pub(in crate::node) async fn maybe_probe_path( + &mut self, + node_addr: NodeAddr, + transport_id: TransportId, + remote_addr: TransportAddr, + ) { + self.add_path_candidate(node_addr, transport_id, remote_addr.clone()); + let now_ms = crate::time::mono_ms(); + let timing = self.heartbeat_timing(); + let Some(peer) = self.peers.get_mut(&node_addr) else { + return; + }; + if peer.transport_id() == Some(transport_id) { + return; + } + let Some((probe_id, remote_active, path_id)) = peer.take_probe( + transport_id, + now_ms, + timing.fast_ms, + timing.discovery_cap_ms, + ) else { + return; + }; + let probe = PathMessage { + probe_id, + remote_active, + path_id, + }; + let wire = self.pad_to_link_mtu(probe.encode_probe().to_vec(), transport_id, &remote_addr); + if let Err(e) = self + .send_encrypted_link_message_on_path(&node_addr, &wire, transport_id, remote_addr) + .await + { + debug!(peer = %self.peer_display_name(&node_addr), error = %e, "Path probe send failed"); + } + } + + /// A `PathProbe` arrived from `from` on `arrival`. + /// + /// The frame decrypted under `from`'s session, so `from` is reachable + /// over `arrival`. Record the path and answer on it. The ack says + /// whether `arrival` is the path *we* send on, which it usually is not. + pub(in crate::node) async fn handle_path_probe( + &mut self, + from: &NodeAddr, + payload: &[u8], + arrival: (TransportId, &TransportAddr), + ) { + let probe = match PathMessage::decode(payload) { + Ok(p) => p, + Err(e) => { + debug!(peer = %self.peer_display_name(from), error = %e, "Malformed path probe"); + return; + } + }; + let (transport_id, remote_addr) = arrival; + let now_ms = crate::time::mono_ms(); + let role = self.transport_role(transport_id); + let Some(peer) = self.peers.get_mut(from) else { + return; + }; + let was_new = peer.path_on(transport_id).is_none(); + peer.note_path_probe( + transport_id, + remote_addr.clone(), + probe.remote_active, + probe.path_id, + now_ms, + ); + if was_new { + peer.set_path_role(transport_id, role); + } + let ours_active = peer.transport_id() == Some(transport_id); + let our_path_id = peer + .path_on(transport_id) + .map(|p| p.local_id()) + .unwrap_or(0); + if was_new { + debug!( + peer = %self.peer_display_name(from), + transport_id = %transport_id, + remote_addr = %remote_addr, + "Peer probed a new path; added" + ); + } + + // The ack echoes the probe's size, so a full-size probe proves the + // path for data-sized frames in both directions. + let ack = PathMessage { + probe_id: probe.probe_id, + remote_active: ours_active, + path_id: our_path_id, + }; + let mut wire = ack.encode_ack().to_vec(); + let probe_len = payload.len() + 1; + if wire.len() < probe_len { + wire.resize(probe_len, 0); + } + if let Err(e) = self + .send_encrypted_link_message_on_path(from, &wire, transport_id, remote_addr.clone()) + .await + { + debug!( + peer = %self.peer_display_name(from), + transport_id = %transport_id, + error = %e, + "Path ack send failed" + ); + } + } + + /// A `PathAck` arrived from `from` on `arrival`: our probe on that path + /// reached the peer and the answer reached us, so the path works in + /// both directions. + pub(in crate::node) fn handle_path_ack( + &mut self, + from: &NodeAddr, + payload: &[u8], + arrival: (TransportId, &TransportAddr), + ) { + let ack = match PathMessage::decode(payload) { + Ok(a) => a, + Err(e) => { + debug!(peer = %self.peer_display_name(from), error = %e, "Malformed path ack"); + return; + } + }; + let (transport_id, _) = arrival; + let now_ms = crate::time::mono_ms(); + let window_ms = self.path_policy().rtt_window_ms; + let Some(peer) = self.peers.get_mut(from) else { + return; + }; + let was_live = peer + .path_on(transport_id) + .is_some_and(|p| p.state() == crate::peer::PathState::Live); + match peer.note_path_ack( + transport_id, + ack.probe_id, + ack.remote_active, + ack.path_id, + now_ms, + window_ms, + ) { + Some(rtt_ms) if !was_live => info!( + peer = %self.peer_display_name(from), + transport_id = %transport_id, + rtt_ms, + "Path live: the peer answers on this transport" + ), + Some(rtt_ms) => trace!( + peer = %self.peer_display_name(from), + transport_id = %transport_id, + rtt_ms, + "Path ack" + ), + None => trace!( + peer = %self.peer_display_name(from), + transport_id = %transport_id, + probe_id = ack.probe_id, + "Path ack matched no outstanding probe" + ), + } + } + + /// A transport's presence went away: withdraw the path every peer held + /// over it. Returns how many peers were reaped for want of another path. + /// + /// For each peer: the path goes `Dead` with its history kept. If it was + /// a standby, nothing else happens. If it was the active path and an + /// eligible standby exists, traffic moves there now, under the same + /// session, and the switch side effects run + /// ([`apply_path_switch`](Self::apply_path_switch)). Only a peer with + /// no eligible path left is reaped, through the same routed link-dead + /// teardown the liveness reaper uses. + /// + /// The peer machine sees nothing while any path remains: a switch is + /// not a link event. Deliberately undamped, like the reap it grew from. + pub(in crate::node) async fn withdraw_transport(&mut self, transport_id: TransportId) -> usize { + let now_ms = crate::time::mono_ms(); + let policy = self.path_policy(); + let affected: Vec = self + .peers + .iter() + .filter(|(_, peer)| peer.path_on(transport_id).is_some()) + .map(|(node_addr, _)| *node_addr) + .collect(); + if affected.is_empty() { + return 0; + } + + let wall_ms = Self::now_ms(); + + let mut reaped = 0; + for node_addr in affected { + let outcome = match self.peers.get_mut(&node_addr) { + Some(peer) => peer.withdraw_path(transport_id, now_ms, &policy), + None => continue, + }; + match outcome { + PathWithdrawal::NoPath => {} + PathWithdrawal::Standby => { + debug!( + peer = %self.peer_display_name(&node_addr), + %transport_id, + "Standby path withdrawn: its interface went away" + ); + self.send_path_close(&node_addr, transport_id, PathCloseReason::InterfaceGone) + .await; + } + PathWithdrawal::Switched { from, to } => { + info!( + peer = %self.peer_display_name(&node_addr), + from_transport = %from.0, + to_transport = %to.0, + to_addr = %to.1, + "Active path withdrawn: traffic moved to the standby, session kept" + ); + self.apply_path_switch(&node_addr, to); + self.send_path_close(&node_addr, transport_id, PathCloseReason::InterfaceGone) + .await; + } + PathWithdrawal::NoAlternative => { + self.reap_peer_without_path(node_addr, transport_id, wall_ms) + .await; + reaped += 1; + } + } + } + reaped + } + + /// Everything that follows the peer's active path changing to `to`. + /// + /// A switch is also an MTU change, and three things size traffic from + /// the peer's transport without re-running on their own + /// (`docs/design/fips-multi-path-switchover.md` §5): + /// + /// - the peer's `path_mtu_lookup` seed, which only ever tightens within + /// a link and would leave one cable→BLE excursion clamping every new + /// flow to this peer at the BLE MTU after fail-back; its relinked + /// branch is the hook, so re-seed from the new path; + /// - the per-session source MTU, which tightens on the next send anyway + /// but only loosens after tens of seconds; tighten it now for every + /// session this peer is the next hop of, so the TUN gate answers with + /// PTB instead of losing the first packet per flow at the transport; + /// - the node-wide MSS ceiling. + /// + /// The link record follows the traffic so everything that reports the + /// peer's transport and address by link stays truthful; the control + /// machine is keyed on the link and is untouched. + pub(in crate::node) fn apply_path_switch( + &mut self, + node_addr: &NodeAddr, + to: (TransportId, TransportAddr), + ) { + let (transport_id, addr) = to; + if let Some(link_id) = self.peers.get(node_addr).map(|p| p.link_id()) + && let Some(link) = self.links.get_mut(&link_id) + { + self.addr_to_link.retain(|_, mapped| *mapped != link_id); + link.rebind(transport_id, addr.clone()); + self.addr_to_link + .insert((transport_id, addr.clone()), link_id); + } + + self.seed_path_mtu_for_link_peer(node_addr, transport_id, &addr); + + let link_mtu = self + .transports + .get(&transport_id) + .map(|t| t.link_mtu(&addr)); + if let Some(link_mtu) = link_mtu { + let dests: Vec = self.sessions.keys().copied().collect(); + for dest in dests { + let via_peer = self + .find_next_hop(&dest) + .is_some_and(|hop| hop.node_addr() == node_addr); + if !via_peer { + continue; + } + if let Some(mmp) = self.sessions.get_mut(&dest).and_then(|s| s.mmp_mut()) { + mmp.path_mtu.seed_source_mtu(link_mtu); + } + } + } + + self.refresh_tun_mss_ceiling(); + } + + /// The fast path tick: per-path heartbeats, the carrier edge, and the + /// selection that a `Suspect` mark may call for. + /// + /// Runs every `node.path.active_heartbeat_ms`. Detection is near-instant + /// for direct peers because each medium's own failure signal is used, + /// not a faster timer: carrier here, a failed echo from the heartbeat + /// plan, an unreachable send from the send path. All three mark a path + /// `Suspect`; selection acts on `Suspect` at once because the standby is + /// warm. + pub(in crate::node) async fn run_path_heartbeats(&mut self) { + let carrier_closes = self.poll_carrier_edges(); + + let now_ms = crate::time::mono_ms(); + let timing = self.heartbeat_timing(); + + let mut sends = Vec::new(); + for (node_addr, peer) in self.peers.iter_mut() { + let plan = peer.plan_heartbeats(now_ms, &timing); + for transport_id in plan.suspects { + debug!( + peer = %node_addr, + %transport_id, + "Path suspect: heartbeat echo timed out" + ); + } + for send in plan.sends { + sends.push((*node_addr, send)); + } + } + for (node_addr, send) in sends { + let probe = PathMessage { + probe_id: send.probe_id, + remote_active: send.remote_active, + path_id: send.path_id, + }; + let mut wire = probe.encode_probe().to_vec(); + if send.full_size { + wire = self.pad_to_link_mtu(wire, send.transport_id, &send.addr); + } + if let Err(e) = self + .send_encrypted_link_message_on_path( + &node_addr, + &wire, + send.transport_id, + send.addr, + ) + .await + { + trace!( + peer = %self.peer_display_name(&node_addr), + transport_id = %send.transport_id, + error = %e, + "Path heartbeat send failed" + ); + } + } + + self.run_path_selection(); + + // After selection: a close for the path we were sending on can only + // go out once traffic has moved off it, and `send_path_close` sends + // nothing for the path that is still active. + for (node_addr, transport_id) in carrier_closes { + self.send_path_close(&node_addr, transport_id, PathCloseReason::CarrierLost) + .await; + } + } + + /// The heartbeat intervals from `node.path.*` and `node.heartbeat_interval_secs`. + fn heartbeat_timing(&self) -> HeartbeatTiming { + let cfg = &self.config().node; + let fast_ms = cfg.path.active_heartbeat_ms.max(50); + let slow_ms = cfg.path.standby_heartbeat_ms.max(fast_ms); + HeartbeatTiming { + fast_ms, + slow_ms, + // Two fast intervals: one echo lost is loss, two is a path. + // Stretched per path by its own round trip inside + // `plan_heartbeats`. + timeout_ms: fast_ms.saturating_mul(2), + // A path the peer never acknowledges is probed full-size at + // this cadence for as long as it exists: the link heartbeat + // interval, not the standby one. + discovery_cap_ms: cfg + .heartbeat_interval_secs + .saturating_mul(1000) + .max(slow_ms), + } + } + + /// Read carrier on every interface-bound transport and mark the paths + /// over one that just lost it `Suspect`. Unplugging a cable drops + /// carrier on both NICs, so both ends see this inside one fast tick. + /// Returns the `(peer, transport)` pairs to send a `PathClose` for. + fn poll_carrier_edges(&mut self) -> Vec<(NodeAddr, TransportId)> { + let mut closes = Vec::new(); + let readings: Vec<(TransportId, bool)> = self + .transports + .iter() + .filter_map(|(id, t)| t.interface_presence().map(|p| (*id, p.carrier))) + .collect(); + for (transport_id, carrier) in readings { + let previous = self.carrier_seen.insert(transport_id, carrier); + if previous == Some(true) && !carrier { + let mut marked = Vec::new(); + for (node_addr, peer) in self.peers.iter_mut() { + if peer.mark_path_suspect(transport_id) { + marked.push(*node_addr); + } + } + if !marked.is_empty() { + info!(%transport_id, paths = marked.len(), "Carrier lost: paths suspect"); + closes.extend(marked.into_iter().map(|a| (a, transport_id))); + } + } + } + closes + } + + /// Tell `node_addr` that our path over `transport_id` is closing, on + /// whichever path we now send on. Best effort: the peer would learn + /// from the echo timeout anyway, this just makes it immediate. Nothing + /// is sent if that path is the one we send on (there is no other way to + /// reach the peer) or the peer never told us its id for it, which it + /// does with its first probe or ack on the path. + pub(in crate::node) async fn send_path_close( + &mut self, + node_addr: &NodeAddr, + transport_id: TransportId, + reason: PathCloseReason, + ) { + let Some(peer) = self.peers.get(node_addr) else { + return; + }; + if peer.transport_id() == Some(transport_id) { + return; + } + let Some(remote_id) = peer.path_on(transport_id).and_then(|p| p.remote_id()) else { + return; + }; + let close = PathClose { + path_id: remote_id, + reason, + }; + if let Err(e) = self + .send_encrypted_link_message(node_addr, &close.encode()) + .await + { + trace!( + peer = %self.peer_display_name(node_addr), + %transport_id, + error = %e, + "Path close send failed" + ); + } + } + + /// The peer is closing the path it calls `path_id` (a `PathClose` + /// arrived). Withdraw our side of it as if its transport had gone: Dead + /// with history, traffic moved if it was there. Advisory: the peer's + /// next probe on it revives it. + pub(in crate::node) async fn handle_path_close(&mut self, from: &NodeAddr, payload: &[u8]) { + let close = match PathClose::decode(payload) { + Ok(c) => c, + Err(e) => { + debug!(peer = %self.peer_display_name(from), error = %e, "Malformed path close"); + return; + } + }; + let now_ms = crate::time::mono_ms(); + let policy = self.path_policy(); + let Some(peer) = self.peers.get_mut(from) else { + return; + }; + let Some((transport_id, outcome)) = + peer.withdraw_path_by_local_id(close.path_id, now_ms, &policy) + else { + trace!(peer = %self.peer_display_name(from), path_id = close.path_id, "Path close named no path"); + return; + }; + match outcome { + PathWithdrawal::NoPath => {} + PathWithdrawal::Standby => debug!( + peer = %self.peer_display_name(from), + %transport_id, + reason = ?close.reason, + "Peer closed a standby path" + ), + PathWithdrawal::Switched { from: was, to } => { + info!( + peer = %self.peer_display_name(from), + from_transport = %was.0, + to_transport = %to.0, + reason = ?close.reason, + "Peer closed our active path: traffic moved to the standby, session kept" + ); + self.apply_path_switch(from, to); + } + PathWithdrawal::NoAlternative => { + // The peer says the only path we have to it is going. Leave + // the peer to the echo timeout and the liveness reaper: a + // close is advisory, and the path may outlive the warning. + debug!( + peer = %self.peer_display_name(from), + %transport_id, + reason = ?close.reason, + "Peer closed our only path; waiting for liveness to confirm" + ); + } + } + } + + /// Pad a link message to fill the link MTU on `transport_id` to `addr`, + /// so the frame is data-sized: outer header, inner timestamp and AEAD + /// tag are accounted for. + fn pad_to_link_mtu( + &self, + mut wire: Vec, + transport_id: TransportId, + addr: &TransportAddr, + ) -> Vec { + let Some(transport) = self.transports.get(&transport_id) else { + return wire; + }; + let room = usize::from(transport.link_mtu(addr)) + .saturating_sub(super::session::LINK_FRAME_OVERHEAD); + if wire.len() < room { + wire.resize(room, 0); + } + wire + } + + /// The kernel refused a send to the peer on `transport_id` for want of + /// a route: the path is `Suspect` now, not after an echo timeout. + pub(in crate::node) fn note_path_unreachable( + &mut self, + node_addr: &NodeAddr, + transport_id: TransportId, + ) { + if let Some(peer) = self.peers.get_mut(node_addr) + && peer.mark_path_suspect(transport_id) + { + debug!( + peer = %self.peer_display_name(node_addr), + %transport_id, + "Path suspect: send unreachable" + ); + } + } + + /// A transport's presence came back: clear the probe backoff on every + /// path over it so the next discovery tick may probe at once. + pub(in crate::node) fn reset_probe_backoff_on_transport(&mut self, transport_id: TransportId) { + for peer in self.peers.values_mut() { + peer.reset_probe_backoff_on(transport_id); + } + } +} diff --git a/src/node/handlers/session.rs b/src/node/handlers/session.rs index 15aad913..79fc64a4 100644 --- a/src/node/handlers/session.rs +++ b/src/node/handlers/session.rs @@ -67,7 +67,7 @@ pub(in crate::node) const PATH_MTU_RELEASE_MIN_INTERVAL: std::time::Duration = /// import above, which is `#[cfg(unix)]`. This constant feeds `link_wire_len`, /// whose caller `send_session_datagram` is compiled on every platform, so /// taking the name from that import fails to build on Windows. -const LINK_FRAME_OVERHEAD: usize = +pub(in crate::node) const LINK_FRAME_OVERHEAD: usize = crate::proto::fmp::wire::ESTABLISHED_HEADER_SIZE + 4 + crate::noise::TAG_SIZE; /// Wire size of an encoded `SessionDatagram` of `encoded_len` bytes. diff --git a/src/node/lifecycle/mod.rs b/src/node/lifecycle/mod.rs index 7315ab82..deeb03d8 100644 --- a/src/node/lifecycle/mod.rs +++ b/src/node/lifecycle/mod.rs @@ -779,6 +779,9 @@ impl Node { // dataplane maps are unmutated, so the core's per-peer cap sees a stable // in-flight count — the same guarantee the old collect-then-dial had. let mut transport_neighbors: Vec = Vec::new(); + // Live peers beaconing on a transport we hold no path to them over. + // Added after the loop, which borrows the transport table. + let mut path_candidates: Vec<(NodeAddr, TransportId, TransportAddr)> = Vec::new(); for (transport_id, transport) in &self.transports { if !transport.is_operational() { continue; @@ -840,6 +843,11 @@ impl Node { // again. What is given up is switching away from a link // that is working, which is not a thing worth doing. if self.active_peer_link_is_live(&node_addr) { + // A live peer beaconing on a transport we hold no + // path to it over is a path to add, not a link to + // replace: the heartbeat tick probes it under the + // existing session instead of dialling. + path_candidates.push((node_addr, candidate_transport_id, remote_addr)); continue; } if self.is_connecting_to_peer_on_path( @@ -867,6 +875,10 @@ impl Node { } } + for (node_addr, transport_id, remote_addr) in path_candidates { + self.add_path_candidate(node_addr, transport_id, remote_addr); + } + if transport_neighbors.is_empty() { return; } @@ -3434,10 +3446,7 @@ impl Node { /// Notifies the peer, removes it locally, closes the transport connection /// it was using, and suppresses auto-reconnect. pub(crate) async fn api_disconnect(&mut self, npub: &str) -> Result { - let peer_identity = - PeerIdentity::from_npub(npub).map_err(|e| format!("invalid npub '{npub}': {e}"))?; - let node_addr = *peer_identity.node_addr(); - + let node_addr = self.resolve_peer_npub(npub)?; let Some(peer) = self.peers.get(&node_addr) else { return Err(format!("peer not found: {npub}")); }; diff --git a/src/node/mod.rs b/src/node/mod.rs index 264be9f0..ae2ce3dd 100644 --- a/src/node/mod.rs +++ b/src/node/mod.rs @@ -636,6 +636,9 @@ pub struct Node { /// the peer entry itself. Pruned on insert; see /// `EPOCH_RESTART_MIN_INTERVAL_SECS`. restart_dampener: HashMap, + /// Last carrier reading per interface-bound transport, for the carrier + /// edge the fast path tick detects. Absent until first read. + carrier_seen: HashMap, // === Rate Limiting === /// Rate limiter for msg1 processing (DoS protection). @@ -921,6 +924,7 @@ impl Node { peers_by_index: HashMap::new(), pending_outbound: HashMap::new(), restart_dampener: HashMap::new(), + carrier_seen: HashMap::new(), msg1_rate_limiter, setup_rate_limiter, icmp_rate_limiter: IcmpRateLimiter::new(), @@ -1095,6 +1099,7 @@ impl Node { peers_by_index: HashMap::new(), pending_outbound: HashMap::new(), restart_dampener: HashMap::new(), + carrier_seen: HashMap::new(), msg1_rate_limiter, setup_rate_limiter, icmp_rate_limiter: IcmpRateLimiter::new(), @@ -2279,6 +2284,7 @@ impl Node { // (their effective_depth is `None`); during cold start (no peer has // SRTT) every peer falls back to the default link cost of 1.0. let any_peer_has_srtt = self.peers().any(|p| p.has_srtt()); + let now_ms = crate::time::mono_ms(); let now_ms = Self::now_ms(); let peer_rows: Vec = self @@ -2324,7 +2330,7 @@ impl Node { if any_peer_has_srtt && !peer.has_srtt() { None } else { - Some(coords.depth() as f64 + peer.link_cost()) + Some(coords.depth() as f64 + peer.link_cost(now_ms)) } }); @@ -3741,6 +3747,41 @@ impl Node { node_addr: &NodeAddr, plaintext: &[u8], ce_flag: bool, + ) -> Result<(), NodeError> { + self.send_encrypted_link_message_via(node_addr, plaintext, ce_flag, None) + .await + } + + /// Like `send_encrypted_link_message` but on a chosen path rather than + /// the peer's active one. + /// + /// The path probe exchange uses this to reach a peer over a transport it + /// is not (yet) sending on. Same session, same counter, same key: only + /// the transport and address differ. + pub(super) async fn send_encrypted_link_message_on_path( + &mut self, + node_addr: &NodeAddr, + plaintext: &[u8], + transport_id: TransportId, + remote_addr: TransportAddr, + ) -> Result<(), NodeError> { + self.send_encrypted_link_message_via( + node_addr, + plaintext, + false, + Some((transport_id, remote_addr)), + ) + .await + } + + /// The one send path for encrypted link messages. `via` picks the + /// transport and address; `None` means the peer's active path. + async fn send_encrypted_link_message_via( + &mut self, + node_addr: &NodeAddr, + plaintext: &[u8], + ce_flag: bool, + via: Option<(TransportId, TransportAddr)>, ) -> Result<(), NodeError> { let peer = self .peers @@ -3751,17 +3792,24 @@ impl Node { node_addr: *node_addr, reason: "no their_index".into(), })?; - let transport_id = peer.transport_id().ok_or_else(|| NodeError::SendFailed { - node_addr: *node_addr, - reason: "no transport_id".into(), - })?; - let remote_addr = peer - .current_addr() - .cloned() - .ok_or_else(|| NodeError::SendFailed { - node_addr: *node_addr, - reason: "no current_addr".into(), - })?; + let on_active_path = via.is_none(); + let (transport_id, remote_addr) = match via { + Some(target) => target, + None => { + let transport_id = peer.transport_id().ok_or_else(|| NodeError::SendFailed { + node_addr: *node_addr, + reason: "no transport_id".into(), + })?; + let remote_addr = + peer.current_addr() + .cloned() + .ok_or_else(|| NodeError::SendFailed { + node_addr: *node_addr, + reason: "no current_addr".into(), + })?; + (transport_id, remote_addr) + } + }; // Prepend 4-byte session-relative timestamp (inner header) let timestamp_ms = peer.session_elapsed_ms(); @@ -3779,8 +3827,16 @@ impl Node { // Snapshot the per-peer connect()-ed UDP socket BEFORE the // session borrow so the encrypt-worker dispatch can refcount- // clone the Arc without re-borrowing self.peers later. + // The connected socket is pinned to the active path's 5-tuple, so a + // send on any other path must go through the listen socket. #[cfg(any(target_os = "linux", target_os = "macos"))] - let connected_socket = peer.connected_udp(); + let connected_socket = if on_active_path { + peer.connected_udp() + } else { + None + }; + #[cfg(not(any(target_os = "linux", target_os = "macos")))] + let _ = on_active_path; let session = peer .noise_session_mut() @@ -3918,28 +3974,29 @@ impl Node { } } - let bytes_sent = transport - .send(&remote_addr, &wire_packet) - .await - .map_err(|e| match e { - TransportError::MtuExceeded { packet_size, mtu } => NodeError::MtuExceeded { - node_addr: *node_addr, - packet_size, - mtu, - }, - // Preserve the transport's own classification instead of - // flattening every non-MTU failure into one string. A caller - // that wants to keep its half-built state across an interface - // flap can only do that if the distinction survives to it. - other if other.is_transient() => NodeError::SendUnavailable { - node_addr: *node_addr, - reason: format!("transport send: {}", other), - }, - other => NodeError::SendFailed { - node_addr: *node_addr, - reason: format!("transport send: {}", other), - }, - })?; + let sent = transport.send(&remote_addr, &wire_packet).await; + if sent.as_ref().is_err_and(|e| e.is_unreachable()) { + self.note_path_unreachable(node_addr, transport_id); + } + let bytes_sent = sent.map_err(|e| match e { + TransportError::MtuExceeded { packet_size, mtu } => NodeError::MtuExceeded { + node_addr: *node_addr, + packet_size, + mtu, + }, + // Preserve the transport's own classification instead of + // flattening every non-MTU failure into one string. A caller + // that wants to keep its half-built state across an interface + // flap can only do that if the distinction survives to it. + other if other.is_transient() => NodeError::SendUnavailable { + node_addr: *node_addr, + reason: format!("transport send: {}", other), + }, + other => NodeError::SendFailed { + node_addr: *node_addr, + reason: format!("transport send: {}", other), + }, + })?; // Update send statistics if let Some(peer) = self.peers.get_mut(node_addr) { @@ -4006,7 +4063,7 @@ impl routing::RoutingView for NodeRoutingView<'_> { } fn peer_link_cost<'a>(&'a self, peer: Self::Peer<'a>) -> f64 { - peer.1.link_cost() + peer.1.link_cost(crate::time::mono_ms()) } fn peer_coords<'a>(&'a self, peer: Self::Peer<'a>) -> Option<&'a TreeCoordinate> { diff --git a/src/node/tests/handshake.rs b/src/node/tests/handshake.rs index 68794d11..0993ade9 100644 --- a/src/node/tests/handshake.rs +++ b/src/node/tests/handshake.rs @@ -2212,6 +2212,7 @@ async fn a_transient_msg2_failure_keeps_the_link_for_the_retry() { // `InterfaceUnavailable` — the real error, from the real code path, // rather than a stub that merely returns something transient. let config = EthernetConfig { + role: None, interface: "fips-absent-x0".to_string(), ethertype: None, mtu: None, @@ -2294,6 +2295,7 @@ async fn a_transient_msg2_failure_on_the_restart_path_leaves_the_fresh_leg_pendi // An interface no host has, so every send off this transport reports // `InterfaceUnavailable` — the real error from the real code path. let config = EthernetConfig { + role: None, interface: "fips-absent-x0".to_string(), ethertype: None, mtu: None, diff --git a/src/node/tests/heartbeat.rs b/src/node/tests/heartbeat.rs index a8e1c984..5b58a60f 100644 --- a/src/node/tests/heartbeat.rs +++ b/src/node/tests/heartbeat.rs @@ -368,7 +368,7 @@ async fn a_detached_transport_withdraws_the_peers_that_needed_it() { .transport_id() .expect("an established peer names its transport"); - let reaped = nodes[0].node.reap_peers_on_transport(transport_id).await; + let reaped = nodes[0].node.withdraw_transport(transport_id).await; assert_eq!(reaped, 1); assert!( @@ -399,7 +399,7 @@ async fn a_detached_transport_leaves_other_transports_peers_alone() { // A transport this peer was never reachable through. let unrelated = TransportId::new(peer_transport.as_u32() + 100); - let reaped = nodes[0].node.reap_peers_on_transport(unrelated).await; + let reaped = nodes[0].node.withdraw_transport(unrelated).await; assert_eq!(reaped, 0, "an unrelated transport withdraws nothing"); assert!( diff --git a/src/node/tests/multi_path.rs b/src/node/tests/multi_path.rs index 0b6880b1..c4a27b98 100644 --- a/src/node/tests/multi_path.rs +++ b/src/node/tests/multi_path.rs @@ -253,6 +253,1191 @@ fn a_promoted_peer_holds_one_path_and_rebind_repoints_it() { // UDP `interface:` binding // ============================================================================ +// ============================================================================ +// Path probe / path ack (design §4) +// ============================================================================ + +use super::spanning_tree::{ + LOOPBACK_REGISTRY, TestNode, initiate_handshake, make_test_node, next_loopback_addr, + process_available_packets, +}; +use crate::peer::PathState; +use crate::proto::link::PathMessage; +use crate::transport::TransportHandle; +use crate::transport::loopback::LoopbackTransport; + +/// The second transport both nodes share, standing in for the wifi next +/// to the cable that `TestNode` comes with. +fn wifi() -> TransportId { + TransportId::new(2) +} + +/// Give `node` a second loopback transport, `wifi()`, on a fresh address that +/// delivers into the node's existing packet channel. Returns that address. +fn add_wifi(node: &mut TestNode) -> TransportAddr { + let addr = next_loopback_addr(); + let tx = LOOPBACK_REGISTRY + .lock() + .unwrap() + .get(&node.addr) + .cloned() + .expect("the node's cable address is registered"); + LOOPBACK_REGISTRY.lock().unwrap().insert(addr.clone(), tx); + let transport = LoopbackTransport::new(wifi(), addr.clone(), LOOPBACK_REGISTRY.clone()); + node.node + .transports + .insert(wifi(), TransportHandle::Loopback(transport)); + addr +} + +/// Two nodes peered over the cable, each also reachable over `wifi()`. +/// Returns `(nodes, wifi_addr_of_0, wifi_addr_of_1)`. +async fn dual_homed_pair() -> (Vec, TransportAddr, TransportAddr) { + let mut nodes = vec![make_test_node().await, make_test_node().await]; + let wifi_0 = add_wifi(&mut nodes[0]); + let wifi_1 = add_wifi(&mut nodes[1]); + initiate_handshake(&mut nodes, 0, 1).await; + for _ in 0..10 { + if process_available_packets(&mut nodes).await == 0 { + break; + } + } + assert_eq!(nodes[0].node.peer_count(), 1, "peered over the cable"); + assert_eq!(nodes[1].node.peer_count(), 1, "peered over the cable"); + (nodes, wifi_0, wifi_1) +} + +#[test] +fn path_message_round_trips_on_the_wire() { + let probe = PathMessage { + probe_id: 0xDEAD_BEEF, + remote_active: true, + path_id: 7, + }; + let wire = probe.encode_probe(); + assert_eq!(wire[0], 0x52); + assert_eq!(PathMessage::decode(&wire[1..]).unwrap(), probe); + + let ack = PathMessage { + probe_id: 7, + remote_active: false, + path_id: 0xFFFF_0000, + }; + let wire = ack.encode_ack(); + assert_eq!(wire[0], 0x53); + assert_eq!(PathMessage::decode(&wire[1..]).unwrap(), ack); + + assert!(PathMessage::decode(&wire[1..8]).is_err(), "short payload"); + let mut padded = wire.to_vec(); + padded.extend_from_slice(&[0u8; 1200]); + assert_eq!( + PathMessage::decode(&padded[1..]).unwrap(), + ack, + "padding is ignored" + ); +} + +#[tokio::test] +async fn a_probe_adds_a_path_at_both_ends_and_the_ack_makes_it_live() { + let (mut nodes, wifi_0, _wifi_1) = dual_homed_pair().await; + let addr_0 = *nodes[0].node.node_addr(); + let addr_1 = *nodes[1].node.node_addr(); + let cable = nodes[0].transport_id; + + // Node 1 probes node 0 over the wifi. + nodes[1] + .node + .maybe_probe_path(addr_0, wifi(), wifi_0.clone()) + .await; + { + let peer = nodes[1].node.get_peer(&addr_0).unwrap(); + let path = peer + .path_on(wifi()) + .expect("the prober adds the path first"); + assert_eq!(path.state(), PathState::Probing); + assert!(path.tx_live_at_ms().is_none()); + } + + // Probe reaches node 0. + assert_eq!(process_available_packets(&mut nodes).await, 1); + { + let peer = nodes[0].node.get_peer(&addr_1).unwrap(); + let path = peer.path_on(wifi()).expect("the receiver adds the path"); + assert_eq!( + path.state(), + PathState::Probing, + "hearing is not proof of the reverse" + ); + assert!(path.rx_live_at_ms().is_some()); + assert!(!path.remote_active(), "node 1 still sends on the cable"); + assert_eq!(peer.transport_id(), Some(cable), "nothing switched"); + assert_eq!(peer.paths().len(), 2); + } + + // Ack reaches node 1. + assert_eq!(process_available_packets(&mut nodes).await, 1); + { + let peer = nodes[1].node.get_peer(&addr_0).unwrap(); + let path = peer.path_on(wifi()).unwrap(); + assert_eq!(path.state(), PathState::Live); + assert!(path.tx_live_at_ms().is_some()); + assert!(path.last_rtt_ms().is_some()); + assert!(!path.remote_active(), "node 0 still sends on the cable"); + assert_eq!(peer.transport_id(), Some(cable), "nothing switched"); + } + + // One session, one index: no handshake was started anywhere. + assert_eq!(nodes[0].node.peer_count(), 1); + assert_eq!(nodes[1].node.peer_count(), 1); + assert!( + !nodes[1] + .node + .is_connecting_to_peer_on_path(&addr_0, wifi(), &wifi_0) + ); +} + +#[tokio::test] +async fn a_beacon_from_a_live_peer_on_a_new_transport_probes_instead_of_dialling() { + let (mut nodes, wifi_0, _wifi_1) = dual_homed_pair().await; + let addr_0 = *nodes[0].node.node_addr(); + let pubkey_0 = nodes[0].node.identity().pubkey(); + + // Node 1 hears node 0 beacon on the wifi. + match nodes[1].node.transports.get(&wifi()).unwrap() { + TransportHandle::Loopback(t) => t.inject_discovered(wifi_0.clone(), pubkey_0), + _ => unreachable!(), + } + nodes[1].node.poll_transport_discovery().await; + + assert!( + !nodes[1] + .node + .is_connecting_to_peer_on_path(&addr_0, wifi(), &wifi_0), + "a live peer is probed, not dialled" + ); + assert_eq!( + nodes[1] + .node + .get_peer(&addr_0) + .unwrap() + .path_on(wifi()) + .map(|p| p.state()), + Some(PathState::Probing) + ); + + // Discovery sends nothing; the heartbeat tick probes the new path. + assert_eq!(nodes[0].packet_rx.len(), 0); + nodes[1].node.run_path_heartbeats().await; + + // Probes out (cable heartbeat and wifi probe), acks back. + for _ in 0..4 { + if process_available_packets(&mut nodes).await == 0 { + break; + } + } + assert_eq!( + nodes[1] + .node + .get_peer(&addr_0) + .unwrap() + .path_on(wifi()) + .map(|p| p.state()), + Some(PathState::Live) + ); +} + +#[tokio::test] +async fn a_beaconed_path_is_probed_by_the_next_heartbeat_tick_and_once_only() { + let (mut nodes, wifi_0, _wifi_1) = dual_homed_pair().await; + let addr_0 = *nodes[0].node.node_addr(); + + // The beacon adds the path; nothing is sent until the tick. + nodes[1] + .node + .add_path_candidate(addr_0, wifi(), wifi_0.clone()); + assert_eq!( + nodes[0].packet_rx.len(), + 0, + "discovery sends nothing itself" + ); + + // Two ticks back to back: one probe on the wifi, the second held + // while the first is in flight. + nodes[1].node.run_path_heartbeats().await; + nodes[1].node.run_path_heartbeats().await; + let mut on_wifi = 0; + while let Ok(packet) = nodes[0].packet_rx.try_recv() { + if packet.transport_id == wifi() { + on_wifi += 1; + } + } + assert_eq!(on_wifi, 1, "one probe in flight per path"); + + // The path is still unproven. + assert_eq!( + nodes[1] + .node + .get_peer(&addr_0) + .unwrap() + .path_on(wifi()) + .map(|p| p.state()), + Some(PathState::Probing) + ); +} + +#[tokio::test] +async fn an_ack_for_no_outstanding_probe_changes_nothing() { + let (mut nodes, _wifi_0, wifi_1) = dual_homed_pair().await; + let addr_1 = *nodes[1].node.node_addr(); + + let stale = PathMessage { + probe_id: 99, + remote_active: false, + path_id: 1, + }; + nodes[0] + .node + .handle_path_ack(&addr_1, &stale.encode_ack()[1..], (wifi(), &wifi_1)); + assert!( + nodes[0] + .node + .get_peer(&addr_1) + .unwrap() + .path_on(wifi()) + .is_none(), + "an ack never creates a path; only a probe does" + ); +} + +#[tokio::test] +async fn the_active_path_is_not_probed() { + let (mut nodes, _wifi_0, _wifi_1) = dual_homed_pair().await; + let addr_0 = *nodes[0].node.node_addr(); + let cable = nodes[0].transport_id; + let cable_addr_0 = nodes[0].addr.clone(); + + nodes[1] + .node + .maybe_probe_path(addr_0, cable, cable_addr_0) + .await; + assert_eq!( + nodes[0].packet_rx.len(), + 0, + "the handshake proved the active path" + ); +} + +// ============================================================================ +// Presence loss withdraws a path (design §5–6) +// ============================================================================ + +/// A dual-homed pair with the wifi path `Live` at both ends: each side has +/// probed the other and heard the ack. +async fn pair_with_wifi_live() -> (Vec, TransportAddr, TransportAddr) { + let (mut nodes, wifi_0, wifi_1) = dual_homed_pair().await; + let addr_0 = *nodes[0].node.node_addr(); + let addr_1 = *nodes[1].node.node_addr(); + nodes[1] + .node + .maybe_probe_path(addr_0, wifi(), wifi_0.clone()) + .await; + nodes[0] + .node + .maybe_probe_path(addr_1, wifi(), wifi_1.clone()) + .await; + for _ in 0..4 { + if process_available_packets(&mut nodes).await == 0 { + break; + } + } + for (node, peer) in [(&nodes[0], addr_1), (&nodes[1], addr_0)] { + let path = node.node.get_peer(&peer).unwrap().path_on(wifi()).unwrap(); + assert_eq!(path.state(), PathState::Live, "precondition"); + assert!(path.tx_live_at_ms().is_some(), "precondition"); + } + (nodes, wifi_0, wifi_1) +} + +#[tokio::test] +async fn losing_the_active_transport_moves_traffic_to_the_live_standby() { + let (mut nodes, wifi_0, _wifi_1) = pair_with_wifi_live().await; + let addr_0 = *nodes[0].node.node_addr(); + let cable = nodes[1].transport_id; + + let reaped = nodes[1].node.withdraw_transport(cable).await; + assert_eq!(reaped, 0, "a peer with a live standby is not reaped"); + + let peer = nodes[1].node.get_peer(&addr_0).expect("peer survives"); + assert_eq!(peer.transport_id(), Some(wifi()), "traffic moved"); + assert_eq!(peer.current_addr(), Some(&wifi_0)); + assert_eq!(peer.path_on(cable).unwrap().state(), PathState::Dead); + assert_eq!(peer.paths().len(), 2, "the dead path keeps its history"); + + // The link record follows the traffic. + let link = nodes[1].node.get_link(&peer.link_id()).expect("link kept"); + assert_eq!(link.transport_id(), wifi()); + assert_eq!(link.remote_addr(), &wifi_0); + assert!( + nodes[1] + .node + .addr_to_link + .contains_key(&(wifi(), wifi_0.clone())) + ); + + // The path MTU seed now describes the wifi. + let seeded_by = nodes[1] + .node + .path_mtu_seeded_by + .read() + .unwrap() + .get(&crate::FipsAddress::from_node_addr(&addr_0)) + .copied(); + assert_eq!(seeded_by, Some(wifi())); + + // The same session carries on: a frame sent now goes out on the wifi + // and decrypts at the far end, which never saw a switch. + let before = nodes[0].packet_rx.len(); + nodes[1] + .node + .send_encrypted_link_message(&addr_0, &[0x51]) + .await + .expect("send over the standby"); + assert_eq!(nodes[0].packet_rx.len(), before + 1); + let packet = nodes[0].packet_rx.try_recv().unwrap(); + assert_eq!(packet.transport_id, wifi()); + nodes[0].node.handle_encrypted_frame(packet).await; + let far = nodes[0] + .node + .get_peer(nodes[1].node.node_addr()) + .expect("no re-peering"); + assert_eq!(far.consecutive_decrypt_failures(), 0); + assert_eq!(far.transport_id(), Some(cable), "the far end did not move"); +} + +#[tokio::test] +async fn losing_the_active_transport_with_only_a_probing_standby_reaps() { + let (mut nodes, wifi_0, _wifi_1) = dual_homed_pair().await; + let addr_0 = *nodes[0].node.node_addr(); + let cable = nodes[1].transport_id; + + // Probe sent, ack never processed: the wifi path is unproven. + nodes[1] + .node + .maybe_probe_path(addr_0, wifi(), wifi_0.clone()) + .await; + assert_eq!( + nodes[1] + .node + .get_peer(&addr_0) + .unwrap() + .path_on(wifi()) + .unwrap() + .state(), + PathState::Probing + ); + + let reaped = nodes[1].node.withdraw_transport(cable).await; + assert_eq!(reaped, 1, "an unproven path is not a path to switch to"); + assert!(nodes[1].node.get_peer(&addr_0).is_none()); +} + +#[tokio::test] +async fn losing_a_standby_leaves_traffic_where_it_is() { + let (mut nodes, _wifi_0, _wifi_1) = pair_with_wifi_live().await; + let addr_0 = *nodes[0].node.node_addr(); + let cable = nodes[1].transport_id; + + let reaped = nodes[1].node.withdraw_transport(wifi()).await; + assert_eq!(reaped, 0); + let peer = nodes[1].node.get_peer(&addr_0).unwrap(); + assert_eq!(peer.transport_id(), Some(cable)); + let dead = peer.path_on(wifi()).unwrap(); + assert_eq!(dead.state(), PathState::Dead); + assert!(dead.last_rtt_ms().is_some(), "history kept"); + assert!(!dead.is_eligible()); +} + +#[tokio::test] +async fn a_dead_path_is_reprobed_when_its_transport_returns_and_forgotten_after_the_grace() { + let (mut nodes, wifi_0, _wifi_1) = pair_with_wifi_live().await; + let addr_0 = *nodes[0].node.node_addr(); + + nodes[1].node.withdraw_transport(wifi()).await; + + // Presence returns: the next discovery tick probes it again, and the + // ack brings it back Live with its history. (The withdrawal also sent + // node 0 a PathClose on the cable, which it processes here too.) + nodes[1].node.reset_probe_backoff_on_transport(wifi()); + nodes[1] + .node + .maybe_probe_path(addr_0, wifi(), wifi_0.clone()) + .await; + for _ in 0..4 { + if process_available_packets(&mut nodes).await == 0 { + break; + } + } + assert_eq!( + nodes[1] + .node + .get_peer(&addr_0) + .unwrap() + .path_on(wifi()) + .unwrap() + .state(), + PathState::Live + ); + + // Dead again, and this time the grace expires. + nodes[1].node.withdraw_transport(wifi()).await; + let now = crate::time::mono_ms(); + let peer = nodes[1].node.get_peer_mut(&addr_0).unwrap(); + peer.prune_dead_paths(now, 60_000); + assert!( + peer.path_on(wifi()).is_some(), + "inside the grace, history kept" + ); + peer.prune_dead_paths(now + 60_001, 60_000); + assert!( + peer.path_on(wifi()).is_none(), + "past the grace, a fresh path" + ); + assert_eq!(peer.paths().len(), 1); + assert_eq!(peer.transport_id(), Some(nodes[1].transport_id)); +} + +#[tokio::test] +async fn a_rekey_msg1_on_a_standby_path_is_recognised_as_the_established_peer() { + let (nodes, wifi_0, _wifi_1) = pair_with_wifi_live().await; + assert!( + nodes[1].node.is_established_link_msg1(wifi(), &wifi_0), + "the peer sends its rekey on the path *it* uses, which may be our standby" + ); + assert!( + !nodes[1] + .node + .is_established_link_msg1(wifi(), &TransportAddr::from_string("loopback:none")) + ); +} + +#[tokio::test] +async fn garbage_on_a_standby_path_counts_against_the_peer() { + let (mut nodes, _wifi_0, wifi_1) = pair_with_wifi_live().await; + let addr_1 = *nodes[1].node.node_addr(); + let our_index = nodes[0] + .node + .get_peer(&addr_1) + .unwrap() + .our_index() + .unwrap(); + for counter in 0..3u64 { + nodes[0] + .node + .handle_encrypted_frame(ReceivedPacket::new( + wifi(), + wifi_1.clone(), + garbage_frame(our_index, counter), + )) + .await; + } + assert_eq!( + nodes[0] + .node + .get_peer(&addr_1) + .unwrap() + .consecutive_decrypt_failures(), + 3, + "a path in the set is a transport the peer is on" + ); +} + +// ============================================================================ +// Selection (design §8) +// ============================================================================ + +use crate::config::TransportRole; +use crate::peer::{ActivePeer, HeartbeatTiming, PathPolicy, SwitchReason}; + +const CABLE: u32 = 1; +const WIFI: u32 = 2; + +fn tid(n: u32) -> TransportId { + TransportId::new(n) +} + +/// Feed the path on `t` one acknowledged probe with round trip `rtt_ms`, +/// as the heartbeat exchange would. Returns the clock after the ack. +fn sample(peer: &mut ActivePeer, t: u32, now_ms: u64, rtt_ms: u64) -> u64 { + let (id, _, _) = peer + .take_probe(tid(t), now_ms, 1, 1) + .expect("a probe is due: the last one was acked"); + peer.note_path_ack(tid(t), id, false, 1, now_ms + rtt_ms, u64::MAX) + .expect("the ack matches"); + now_ms + rtt_ms +} + +/// A peer active on the cable with three samples at `cable_rtt`, and a +/// wifi standby with three samples at `wifi_rtt`. +fn dual_path_peer(cable_rtt: u64, wifi_rtt: u64) -> ActivePeer { + let mut peer = ActivePeer::new(make_peer_identity(), LinkId::new(1), 0); + peer.rebind_transport(tid(CABLE), TransportAddr::from_string("10.0.0.1:1")); + peer.add_path(tid(WIFI), TransportAddr::from_string("10.0.0.7:1")); + let mut now = 1_000; + for _ in 0..3 { + now = sample(&mut peer, CABLE, now, cable_rtt); + now = sample(&mut peer, WIFI, now, wifi_rtt); + } + assert_eq!(peer.transport_id(), Some(tid(CABLE))); + peer +} + +fn policy() -> PathPolicy { + PathPolicy { + margin: 1.5, + dwell_ms: 5_000, + min_samples: 3, + rtt_window_ms: u64::MAX, + } +} + +#[test] +fn a_cable_under_a_slightly_better_wifi_is_not_left() { + // 1.00 vs 1.10 in the design's table: ratio under K, stay. + let mut peer = dual_path_peer(10, 1); + assert!(peer.select_path(100_000, &policy()).is_none()); + assert!(peer.select_path(200_000, &policy()).is_none()); + assert_eq!(peer.transport_id(), Some(tid(CABLE))); +} + +#[test] +fn a_degraded_active_path_is_left_after_the_dwell() { + // cable 200 ms → score 3.0; wifi 5 ms → 1.05. Ratio 2.9 > K. + let mut peer = dual_path_peer(200, 5); + assert!( + peer.select_path(100_000, &policy()).is_none(), + "dwell starts" + ); + assert!( + peer.select_path(104_999, &policy()).is_none(), + "dwell running" + ); + let switch = peer + .select_path(105_000, &policy()) + .expect("margin held for the dwell"); + assert_eq!(switch.reason, SwitchReason::Discretionary); + assert_eq!(switch.to.0, tid(WIFI)); + assert_eq!(peer.transport_id(), Some(tid(WIFI))); +} + +#[test] +fn the_dwell_restarts_when_the_margin_stops_holding() { + let mut peer = dual_path_peer(200, 5); + assert!(peer.select_path(100_000, &policy()).is_none()); + // The cable recovers for a moment: enough fast samples to pull the + // window min down. + let mut now = 101_000; + for _ in 0..3 { + now = sample(&mut peer, CABLE, now, 1); + } + assert!(peer.select_path(now, &policy()).is_none()); + // Min RTT is min over the window, so the recovery sticks: the path + // never trips the margin again in this test. + assert!(peer.select_path(now + 10_000, &policy()).is_none()); + assert_eq!(peer.transport_id(), Some(tid(CABLE))); +} + +#[test] +fn a_standby_with_too_few_samples_is_not_selectable() { + let mut peer = ActivePeer::new(make_peer_identity(), LinkId::new(1), 0); + peer.rebind_transport(tid(CABLE), TransportAddr::from_string("10.0.0.1:1")); + peer.add_path(tid(WIFI), TransportAddr::from_string("10.0.0.7:1")); + let mut now = 1_000; + for _ in 0..3 { + now = sample(&mut peer, CABLE, now, 200); + } + now = sample(&mut peer, WIFI, now, 5); + now = sample(&mut peer, WIFI, now, 5); + assert!(peer.select_path(now, &policy()).is_none()); + assert!(peer.select_path(now + 60_000, &policy()).is_none()); +} + +#[test] +fn a_backup_path_is_not_selected_while_a_normal_one_is_selectable() { + let mut peer = dual_path_peer(200, 5); + peer.set_path_role(tid(WIFI), TransportRole::Backup); + assert!(peer.select_path(100_000, &policy()).is_none()); + assert!(peer.select_path(200_000, &policy()).is_none()); + assert_eq!(peer.transport_id(), Some(tid(CABLE))); +} + +#[test] +fn a_backup_active_path_yields_to_a_normal_one_at_once() { + // Traffic landed on the backup (say, the cable was gone); the cable is + // back and selectable, and it is slower. Role wins: leave the backup. + let mut peer = dual_path_peer(20, 1); + peer.set_path_role(tid(WIFI), TransportRole::Backup); + peer.pin_path(tid(WIFI)); + let pinned = peer.select_path(100_000, &policy()).expect("pin honoured"); + assert_eq!(pinned.reason, SwitchReason::Pinned); + peer.unpin_paths(); + let back = peer + .select_path(100_001, &policy()) + .expect("a backup yields without margin or dwell"); + assert_eq!(back.reason, SwitchReason::Discretionary); + assert_eq!(peer.transport_id(), Some(tid(CABLE))); +} + +#[test] +fn a_pinned_path_wins_and_holds() { + let mut peer = dual_path_peer(1, 200); + assert!(peer.pin_path(tid(WIFI))); + let switch = peer.select_path(100_000, &policy()).expect("pinned"); + assert_eq!(switch.reason, SwitchReason::Pinned); + assert_eq!(peer.transport_id(), Some(tid(WIFI))); + // The cable is far better now, and it does not matter. + assert!(peer.select_path(200_000, &policy()).is_none()); + assert!(!peer.pin_path(tid(9)), "no path there"); +} + +#[test] +fn a_tie_keeps_the_current_path() { + let mut peer = dual_path_peer(5, 5); + assert!(peer.select_path(100_000, &policy()).is_none()); + assert!(peer.select_path(200_000, &policy()).is_none()); + assert_eq!(peer.transport_id(), Some(tid(CABLE))); +} + +#[test] +fn withdrawal_prefers_a_selectable_standby_over_a_barely_live_one() { + let mut peer = dual_path_peer(1, 5); + let ble = 3; + peer.add_path(tid(ble), TransportAddr::from_string("ble:1")); + let now = sample(&mut peer, ble, 500_000, 1); // Live, one sample only + let outcome = peer.withdraw_path(tid(CABLE), now, &policy()); + match outcome { + crate::peer::PathWithdrawal::Switched { to, .. } => { + assert_eq!(to.0, tid(WIFI), "three samples beat one, whatever the RTT"); + } + other => panic!("expected a switch, got {other:?}"), + } +} + +#[test] +fn transport_role_and_path_config_parse() { + let cfg: crate::config::UdpConfig = serde_yaml::from_str("role: backup\n").unwrap(); + assert_eq!(cfg.role(), TransportRole::Backup); + let cfg: crate::config::UdpConfig = serde_yaml::from_str("bind_addr: 0.0.0.0:1\n").unwrap(); + assert_eq!(cfg.role(), TransportRole::Normal); + let node: crate::config::PathConfig = + serde_yaml::from_str("switch_margin: 2.0\nmin_samples: 5\n").unwrap(); + assert_eq!(node.switch_margin, 2.0); + assert_eq!(node.min_samples, 5); + assert_eq!(node.switch_dwell_secs, 2); + assert_eq!(node.active_heartbeat_ms, 200); +} + +// ============================================================================ +// Detection (design §7) +// ============================================================================ + +const FAST: u64 = 250; +const SLOW: u64 = 10_000; +const TIMEOUT: u64 = 750; +const TIMING: HeartbeatTiming = HeartbeatTiming { + fast_ms: FAST, + slow_ms: SLOW, + timeout_ms: TIMEOUT, + discovery_cap_ms: SLOW, +}; + +#[test] +fn heartbeats_are_fast_on_the_active_path_and_slow_on_a_standby() { + let mut peer = dual_path_peer(1, 5); + let t0 = 1_000_000; + let plan = peer.plan_heartbeats(t0, &TIMING); + let on: Vec<_> = plan.sends.iter().map(|s| s.transport_id).collect(); + assert!( + on.contains(&tid(CABLE)) && on.contains(&tid(WIFI)), + "both due at once: {on:?}" + ); + assert!(plan.suspects.is_empty()); + let cable = plan + .sends + .iter() + .find(|s| s.transport_id == tid(CABLE)) + .unwrap(); + assert!(cable.remote_active, "the active path says so"); + + // Nothing more while both are in flight. + assert!(peer.plan_heartbeats(t0 + 100, &TIMING).sends.is_empty()); + + // Acks land. Then only the active path is due again inside a second. + for t in [CABLE, WIFI] { + let id = plan + .sends + .iter() + .find(|s| s.transport_id == tid(t)) + .unwrap() + .probe_id; + peer.note_path_ack(tid(t), id, false, 1, t0 + 200, u64::MAX) + .unwrap(); + } + let plan = peer.plan_heartbeats(t0 + FAST + 1, &TIMING); + let on: Vec<_> = plan.sends.iter().map(|s| s.transport_id).collect(); + assert_eq!(on, vec![tid(CABLE)]); + let plan = peer.plan_heartbeats(t0 + SLOW + 1, &TIMING); + assert!(plan.sends.iter().any(|s| s.transport_id == tid(WIFI))); +} + +#[test] +fn a_timed_out_echo_on_an_acknowledged_path_makes_it_suspect_and_selection_leaves_it() { + let mut peer = dual_path_peer(1, 5); + let t0 = 1_000_000; + let plan = peer.plan_heartbeats(t0, &TIMING); + assert!(plan.sends.iter().any(|s| s.transport_id == tid(CABLE))); + // The wifi ack lands; the cable's never does. + let wifi_id = plan + .sends + .iter() + .find(|s| s.transport_id == tid(WIFI)) + .unwrap() + .probe_id; + peer.note_path_ack(tid(WIFI), wifi_id, false, 1, t0 + 5, u64::MAX); + + let before = peer.path_on(tid(CABLE)).unwrap().etx(); + let plan = peer.plan_heartbeats(t0 + TIMEOUT, &TIMING); + assert_eq!(plan.suspects, vec![tid(CABLE)]); + let cable = peer.path_on(tid(CABLE)).unwrap(); + assert_eq!(cable.state(), PathState::Suspect); + assert!(cable.etx() > before, "a lost echo is a lost sample"); + assert!( + plan.sends.iter().any(|s| s.transport_id == tid(CABLE)), + "and it is probed again at once" + ); + + let switch = peer + .select_path(t0 + TIMEOUT, &policy()) + .expect("mandatory: the active path is not tx_live"); + assert_eq!(switch.reason, SwitchReason::Mandatory); + assert_eq!(peer.transport_id(), Some(tid(WIFI))); + + // The cable answers after all: Live again, and now the standby. + let cable_id = plan + .sends + .iter() + .find(|s| s.transport_id == tid(CABLE)) + .unwrap() + .probe_id; + peer.note_path_ack(tid(CABLE), cable_id, false, 1, t0 + TIMEOUT + 1, u64::MAX) + .unwrap(); + assert_eq!(peer.path_on(tid(CABLE)).unwrap().state(), PathState::Live); + assert_eq!( + peer.transport_id(), + Some(tid(WIFI)), + "no ping-pong: K decides fail-back" + ); +} + +#[test] +fn a_path_the_peer_never_acknowledged_backs_off_instead_of_going_suspect() { + let mut peer = ActivePeer::new(make_peer_identity(), LinkId::new(1), 0); + peer.rebind_transport(tid(CABLE), TransportAddr::from_string("10.0.0.1:1")); + // Promotion-style: Live and tx_live from the handshake, never acked. + let t0 = 1_000_000; + let plan = peer.plan_heartbeats(t0, &TIMING); + assert_eq!(plan.sends.len(), 1); + let plan = peer.plan_heartbeats(t0 + TIMEOUT, &TIMING); + assert!(plan.suspects.is_empty(), "an old node is not a dead path"); + let cable = peer.path_on(tid(CABLE)).unwrap(); + assert_eq!(cable.state(), PathState::Live); + assert_eq!(plan.sends.len(), 1, "tried again, with the backoff doubled"); + assert!( + peer.plan_heartbeats(t0 + TIMEOUT + 2 * FAST - 1, &TIMING) + .sends + .is_empty(), + "the unanswered probe pushed the next one out" + ); +} + +#[test] +fn withdrawing_our_only_path_leaves_it_suspect_and_still_probed() { + let mut peer = ActivePeer::new(make_peer_identity(), LinkId::new(1), 0); + peer.rebind_transport(tid(CABLE), TransportAddr::from_string("10.0.0.1:1")); + let t0 = 1_000_000; + // An advisory close from the peer names our only path. + assert_eq!( + peer.withdraw_path(tid(CABLE), t0, &policy()), + crate::peer::PathWithdrawal::NoAlternative + ); + let cable = peer.path_on(tid(CABLE)).unwrap(); + assert_eq!( + cable.state(), + PathState::Suspect, + "not Dead: a Dead path is never probed" + ); + assert_eq!( + peer.transport_id(), + Some(tid(CABLE)), + "traffic stays: nowhere else to go" + ); + // The heartbeat tick keeps probing it, and an ack brings it back. + let plan = peer.plan_heartbeats(t0 + 1, &TIMING); + let id = plan + .sends + .iter() + .find(|s| s.transport_id == tid(CABLE)) + .expect("still probed") + .probe_id; + peer.note_path_ack(tid(CABLE), id, true, 1, t0 + 3, u64::MAX) + .unwrap(); + assert_eq!(peer.path_on(tid(CABLE)).unwrap().state(), PathState::Live); +} + +#[test] +fn presence_return_resets_the_discovery_backoff() { + let mut peer = ActivePeer::new(make_peer_identity(), LinkId::new(1), 0); + peer.rebind_transport(tid(CABLE), TransportAddr::from_string("10.0.0.1:1")); + let t0 = 1_000_000; + // Never answered: each timeout doubles the wait. + assert_eq!(peer.plan_heartbeats(t0, &TIMING).sends.len(), 1); + assert_eq!(peer.plan_heartbeats(t0 + TIMEOUT, &TIMING).sends.len(), 1); + assert_eq!( + peer.plan_heartbeats(t0 + 2 * TIMEOUT, &TIMING).sends.len(), + 1 + ); + let t = t0 + 3 * TIMEOUT; + assert!( + peer.plan_heartbeats(t, &TIMING).sends.is_empty(), + "the third timeout pushed the next probe past now" + ); + // The transport's presence cycles: probed at once. + peer.reset_probe_backoff_on(tid(CABLE)); + assert_eq!(peer.plan_heartbeats(t, &TIMING).sends.len(), 1); +} + +#[test] +fn the_discovery_backoff_is_capped_and_never_goes_suspect() { + let mut peer = ActivePeer::new(make_peer_identity(), LinkId::new(1), 0); + peer.rebind_transport(tid(CABLE), TransportAddr::from_string("10.0.0.1:1")); + peer.add_path(tid(WIFI), TransportAddr::from_string("10.0.0.7:1")); + // A standby is probed at `slow_ms`; the backoff doubles from there + // and the cap is what bounds it. + let timing = HeartbeatTiming { + slow_ms: 1_000, + discovery_cap_ms: 2_000, + ..TIMING + }; + // Drive the wifi standby through many unanswered probes. + let mut t = 1_000_000; + let mut sent_at = Vec::new(); + for _ in 0..40 { + let plan = peer.plan_heartbeats(t, &timing); + assert!( + plan.suspects.is_empty(), + "never acknowledged: not a dead path" + ); + if plan.sends.iter().any(|s| s.transport_id == tid(WIFI)) { + sent_at.push(t); + } + t += TIMEOUT; + } + let gaps: Vec = sent_at.windows(2).map(|w| w[1] - w[0]).collect(); + assert!(gaps.len() >= 4, "kept probing: {sent_at:?}"); + assert!( + gaps.iter().all(|g| *g <= 2_000 + TIMEOUT), + "the backoff is capped at discovery_cap_ms: {gaps:?}" + ); + assert!( + peer.plan_heartbeats(t, &timing) + .sends + .iter() + .filter(|s| s.transport_id == tid(WIFI)) + .all(|s| s.full_size), + "every probe on an unproven path is full-size" + ); +} + +#[test] +fn a_late_ack_after_the_echo_timeout_still_measures_the_path() { + let mut peer = ActivePeer::new(make_peer_identity(), LinkId::new(1), 0); + peer.rebind_transport(tid(CABLE), TransportAddr::from_string("10.0.0.1:1")); + peer.add_path(tid(WIFI), TransportAddr::from_string("10.0.0.7:1")); + let t0 = 1_000_000; + let plan = peer.plan_heartbeats(t0, &TIMING); + let id = plan + .sends + .iter() + .find(|s| s.transport_id == tid(WIFI)) + .expect("the new path is probed") + .probe_id; + + // A circuit slower than the timeout floor: the echo times out first. + let plan = peer.plan_heartbeats(t0 + TIMEOUT, &TIMING); + assert!(plan.suspects.is_empty()); + assert_eq!(peer.path_on(tid(WIFI)).unwrap().state(), PathState::Probing); + + // Then the ack arrives. It is still the answer to our probe. + let rtt = peer + .note_path_ack(tid(WIFI), id, false, 1, t0 + TIMEOUT + 250, u64::MAX) + .expect("a late ack is still an ack"); + assert_eq!(rtt, TIMEOUT + 250); + let wifi = peer.path_on(tid(WIFI)).unwrap(); + assert_eq!(wifi.state(), PathState::Live); + assert!(wifi.acked_once()); + assert_eq!(wifi.last_rtt_ms(), Some(TIMEOUT + 250)); + + // A second late ack for the same probe is a duplicate. + assert!( + peer.note_path_ack(tid(WIFI), id, false, 1, t0 + TIMEOUT + 300, u64::MAX) + .is_none() + ); + + // With a round trip on record the timeout stretches: the next probe + // is not timed out at the floor. + let plan = peer.plan_heartbeats(t0 + 2 * TIMEOUT, &TIMING); + let id = plan + .sends + .iter() + .find(|s| s.transport_id == tid(WIFI)) + .map(|s| s.probe_id); + let t1 = t0 + 2 * TIMEOUT; + let plan = peer.plan_heartbeats(t1 + TIMEOUT, &TIMING); + assert!(plan.suspects.is_empty(), "3 × last RTT > the floor"); + if let Some(id) = id { + assert!( + peer.note_path_ack(tid(WIFI), id, false, 1, t1 + TIMEOUT + 1, u64::MAX) + .is_some(), + "still outstanding, not timed out" + ); + } +} + +#[test] +fn a_remote_active_flip_away_from_a_path_triggers_a_probe_not_a_suspect() { + let mut peer = dual_path_peer(1, 5); + let t0 = 1_000_000; + // The peer said it sends on the wifi, then said it does not. + peer.note_path_probe( + tid(WIFI), + TransportAddr::from_string("10.0.0.7:1"), + true, + 1, + t0, + ); + let plan = peer.plan_heartbeats(t0, &TIMING); + for s in plan.sends { + peer.note_path_ack( + s.transport_id, + s.probe_id, + s.transport_id == tid(WIFI), + 1, + t0 + 1, + u64::MAX, + ); + } + // Wifi is now on the fast cadence (the peer sends there); it was just + // acked, so nothing is due for a while. + assert!(peer.plan_heartbeats(t0 + 10, &TIMING).sends.is_empty()); + peer.note_path_probe( + tid(WIFI), + TransportAddr::from_string("10.0.0.7:1"), + false, + 1, + t0 + 20, + ); + let plan = peer.plan_heartbeats(t0 + 21, &TIMING); + assert!( + plan.sends.iter().any(|s| s.transport_id == tid(WIFI)), + "probe now" + ); + assert!(plan.suspects.is_empty()); + assert_eq!(peer.path_on(tid(WIFI)).unwrap().state(), PathState::Live); +} + +#[test] +fn silence_on_the_active_path_while_a_standby_hears_the_peer_triggers_a_probe() { + let mut peer = dual_path_peer(1, 5); + let t0 = 1_000_000; + let plan = peer.plan_heartbeats(t0, &TIMING); + for s in plan.sends { + peer.note_path_ack(s.transport_id, s.probe_id, false, 1, t0 + 1, u64::MAX); + } + // Two of the peer's (slow) intervals of silence on the cable, while + // the wifi keeps hearing it. + let later = t0 + 2 * SLOW + 1; + peer.note_path_rx(tid(WIFI), later); + let plan = peer.plan_heartbeats(later, &TIMING); + assert!(plan.sends.iter().any(|s| s.transport_id == tid(CABLE))); + assert_eq!( + peer.path_on(tid(CABLE)).unwrap().state(), + PathState::Live, + "a hint, not a verdict" + ); +} + +#[test] +fn a_hard_signal_marks_a_live_path_suspect_only() { + let mut peer = dual_path_peer(1, 5); + assert!(peer.mark_path_suspect(tid(WIFI))); + assert!(!peer.mark_path_suspect(tid(WIFI)), "already suspect"); + assert!(!peer.mark_path_suspect(tid(9)), "no such path"); + assert!(!peer.path_on(tid(WIFI)).unwrap().is_eligible()); +} + +#[test] +fn unreachable_send_errors_are_classified() { + use crate::transport::TransportError; + let e = TransportError::Io(std::io::Error::from(std::io::ErrorKind::NetworkUnreachable)); + assert!(e.is_unreachable()); + let e = TransportError::Io(std::io::Error::from(std::io::ErrorKind::HostUnreachable)); + assert!(e.is_unreachable()); + let e = TransportError::Io(std::io::Error::from(std::io::ErrorKind::ConnectionRefused)); + assert!(!e.is_unreachable()); + assert!(!TransportError::Timeout.is_unreachable()); +} + +#[tokio::test] +async fn the_fast_tick_heartbeats_the_active_path_and_the_ack_measures_it() { + let (mut nodes, _wifi_0, _wifi_1) = pair_with_wifi_live().await; + let addr_0 = *nodes[0].node.node_addr(); + let cable = nodes[1].transport_id; + assert!( + !nodes[1] + .node + .get_peer(&addr_0) + .unwrap() + .path_on(cable) + .unwrap() + .acked_once(), + "the handshake proved the cable; no probe has yet" + ); + + nodes[1].node.run_path_heartbeats().await; + let queued = nodes[0].packet_rx.len(); + assert!(queued >= 1, "a heartbeat probe went out on the cable"); + for _ in 0..4 { + if process_available_packets(&mut nodes).await == 0 { + break; + } + } + let cable_path = nodes[1] + .node + .get_peer(&addr_0) + .unwrap() + .path_on(cable) + .unwrap(); + assert!(cable_path.acked_once()); + assert!(cable_path.min_rtt_ms().is_some()); + assert_eq!(cable_path.state(), PathState::Live); + // Node 0 learned that node 1 sends on the cable. + let far = nodes[0].node.get_peer(nodes[1].node.node_addr()).unwrap(); + assert!(far.path_on(cable).unwrap().remote_active()); +} + +// ============================================================================ +// Tree dampening, fipsctl path, UDP interface (design §10, step 7) +// ============================================================================ + +#[test] +fn a_switch_holds_the_reported_link_cost_for_the_dwell_and_the_reports_that_span_it() { + let mut peer = dual_path_peer(200, 5); + let now = crate::time::mono_ms(); + let before = peer.link_cost(now); + assert!(!peer.link_cost_held(now)); + let p = PathPolicy { + dwell_ms: 60_000, + ..policy() + }; + peer.pin_path(tid(WIFI)); + peer.select_path(now, &p).expect("pinned switch"); + assert!(peer.link_cost_held(now)); + assert_eq!(peer.link_cost(now), before, "held at the pre-switch cost"); + // The dwell alone does not release it: the receiver report that spans + // the switch counts the frames in flight on the old path as lost and + // spikes the per-report ETX for one interval, and the one after that + // replaces it. Two reports, then the dwell, whichever is later. + assert!( + peer.link_cost_held(now + 60_001), + "still held: no report has arrived since the switch" + ); + peer.note_receiver_report(); + assert!(peer.link_cost_held(now + 60_001), "the spanning report"); + peer.note_receiver_report(); + assert!(!peer.link_cost_held(now + 60_001), "replaced: released"); + assert!( + peer.link_cost_held(now + 1), + "reports in, dwell not: still held" + ); + + let p0 = PathPolicy { + dwell_ms: 0, + ..policy() + }; + let mut peer = dual_path_peer(200, 5); + peer.pin_path(tid(WIFI)); + peer.select_path(now, &p0).expect("pinned switch"); + assert!(!peer.link_cost_held(now), "a zero dwell holds nothing"); +} + +#[tokio::test] +async fn fipsctl_path_show_pin_and_unpin_go_through_the_control_api() { + let (nodes, _wifi_0, _wifi_1) = pair_with_wifi_live().await; + let mut nodes = nodes; + let npub_0 = nodes[0].node.identity().npub(); + + let shown = nodes[1].node.api_path_show(&npub_0).expect("known peer"); + let paths = shown["paths"].as_array().unwrap(); + assert_eq!(paths.len(), 2); + assert_eq!( + paths.iter().filter(|p| p["active"] == true).count(), + 1, + "exactly one active path" + ); + assert!(paths.iter().all(|p| p["state"] == "live")); + assert!(paths.iter().all(|p| p["pinned"] == false)); + + let err = nodes[1].node.api_path_show("npub1notapeer").unwrap_err(); + assert!(err.contains("invalid npub"), "{err}"); + + // Pin by numeric id (the loopback transports carry no name). + let pinned = nodes[1] + .node + .api_path_pin(&npub_0, &wifi().as_u32().to_string()) + .expect("pin"); + assert_eq!(pinned["pinned"], wifi().as_u32()); + nodes[1].node.run_path_selection(); + let addr_0 = *nodes[0].node.node_addr(); + assert_eq!( + nodes[1].node.get_peer(&addr_0).unwrap().transport_id(), + Some(wifi()), + "the pin took on the next selection run" + ); + let shown = nodes[1].node.api_path_show(&npub_0).unwrap(); + assert!( + shown["paths"] + .as_array() + .unwrap() + .iter() + .any(|p| p["pinned"] == true && p["active"] == true) + ); + + assert!( + nodes[1] + .node + .api_path_pin(&npub_0, "no-such-transport") + .is_err() + ); + nodes[1].node.api_path_unpin(&npub_0).expect("unpin"); + let shown = nodes[1].node.api_path_show(&npub_0).unwrap(); + assert!( + shown["paths"] + .as_array() + .unwrap() + .iter() + .all(|p| p["pinned"] == false) + ); +} + #[test] fn udp_interface_config_parses() { let cfg: crate::config::UdpConfig = serde_yaml::from_str("interface: en0\n").unwrap(); @@ -274,3 +1459,268 @@ fn binding_udp_to_a_missing_interface_fails_to_start() { .expect("an absent interface cannot be bound"); assert!(err.to_string().contains("fips-absent-x0"), "{err}"); } + +// ============================================================================ +// Path close, full-size probes, revival +// ============================================================================ + +#[test] +fn path_close_round_trips_on_the_wire() { + use crate::proto::link::{PathClose, PathCloseReason}; + let close = PathClose { + path_id: 0x1234_5678, + reason: PathCloseReason::CarrierLost, + }; + let wire = close.encode(); + assert_eq!(wire[0], 0x54); + assert_eq!(PathClose::decode(&wire[1..]).unwrap(), close); + assert_eq!( + PathClose::decode(&[1, 0, 0, 0, 200]).unwrap().reason, + PathCloseReason::Unspecified, + "an unknown reason byte is not an error" + ); +} + +#[test] +fn a_probe_reaching_a_dead_path_revives_it() { + let mut peer = dual_path_peer(1, 5); + peer.withdraw_path(tid(WIFI), 1_000_000, &policy()); + assert_eq!(peer.path_on(tid(WIFI)).unwrap().state(), PathState::Dead); + peer.note_path_probe( + tid(WIFI), + TransportAddr::from_string("10.0.0.7:1"), + false, + 9, + 1_000_500, + ); + let wifi = peer.path_on(tid(WIFI)).unwrap(); + assert_eq!( + wifi.state(), + PathState::Probing, + "heard again, unproven our way" + ); + assert_eq!(wifi.remote_id(), Some(9)); + assert!(wifi.last_rtt_ms().is_some(), "history kept"); +} + +#[tokio::test] +async fn the_first_probe_on_a_path_is_full_size_and_so_is_its_ack() { + let (mut nodes, wifi_0, _wifi_1) = dual_homed_pair().await; + let addr_0 = *nodes[0].node.node_addr(); + nodes[1] + .node + .maybe_probe_path(addr_0, wifi(), wifi_0.clone()) + .await; + let probe = nodes[0].packet_rx.try_recv().expect("probe queued"); + let mtu = usize::from( + nodes[1] + .node + .transports + .get(&wifi()) + .unwrap() + .link_mtu(&wifi_0), + ); + assert_eq!(probe.data.len(), mtu, "the first probe fills the link MTU"); + nodes[0].node.handle_encrypted_frame(probe).await; + let ack = nodes[1].packet_rx.try_recv().expect("ack queued"); + assert_eq!(ack.data.len(), mtu, "and the ack echoes its size"); + nodes[1].node.handle_encrypted_frame(ack).await; + assert_eq!( + nodes[1] + .node + .get_peer(&addr_0) + .unwrap() + .path_on(wifi()) + .unwrap() + .state(), + PathState::Live + ); +} + +#[test] +fn a_proven_path_is_probed_full_size_once_a_minute() { + let mut peer = dual_path_peer(1, 5); + let t0 = 1_000_000; + let plan = peer.plan_heartbeats(t0, &TIMING); + // dual_path_peer acked via take_probe, which sets no full-size stamp, + // so the first heartbeat is full-size; after that, small until a + // minute has passed. + assert!(plan.sends.iter().all(|s| s.full_size)); + for s in plan.sends { + peer.note_path_ack(s.transport_id, s.probe_id, false, 1, t0 + 1, u64::MAX); + } + let plan = peer.plan_heartbeats(t0 + FAST + 1, &TIMING); + assert!(plan.sends.iter().all(|s| !s.full_size)); + for s in plan.sends { + peer.note_path_ack( + s.transport_id, + s.probe_id, + false, + 1, + t0 + FAST + 2, + u64::MAX, + ); + } + let plan = peer.plan_heartbeats(t0 + 61_000, &TIMING); + assert!(plan.sends.iter().any(|s| s.full_size)); +} + +#[tokio::test] +async fn losing_a_transport_tells_the_peer_which_closes_its_side() { + let (mut nodes, _wifi_0, _wifi_1) = pair_with_wifi_live().await; + let addr_0 = *nodes[0].node.node_addr(); + let addr_1 = *nodes[1].node.node_addr(); + assert!( + nodes[0] + .node + .get_peer(&addr_1) + .unwrap() + .path_on(wifi()) + .unwrap() + .remote_id() + .is_some(), + "the probe exchange taught each side the other's path id" + ); + + // Node 1 loses its wifi: it tells node 0 over the cable. + nodes[1].node.withdraw_transport(wifi()).await; + assert_eq!(nodes[0].packet_rx.len(), 1, "one PathClose on the cable"); + let packet = nodes[0].packet_rx.try_recv().unwrap(); + assert_eq!(packet.transport_id, nodes[0].transport_id); + nodes[0].node.handle_encrypted_frame(packet).await; + let wifi_path = nodes[0] + .node + .get_peer(&addr_1) + .unwrap() + .path_on(wifi()) + .unwrap(); + assert_eq!( + wifi_path.state(), + PathState::Dead, + "closed at once, no echo timeout" + ); + assert!(wifi_path.last_rtt_ms().is_some(), "history kept"); + assert_eq!( + nodes[0].node.get_peer(&addr_1).unwrap().transport_id(), + Some(nodes[0].transport_id) + ); + let _ = addr_0; +} + +#[tokio::test] +async fn carrier_loss_on_the_active_path_moves_traffic_and_tells_the_peer() { + let (mut nodes, _wifi_0, _wifi_1) = pair_with_wifi_live().await; + let addr_0 = *nodes[0].node.node_addr(); + let addr_1 = *nodes[1].node.node_addr(); + let cable = nodes[1].transport_id; + let set_carrier = |node: &TestNode, carrier: bool| match node.node.transports.get(&cable) { + Some(TransportHandle::Loopback(t)) => t.set_carrier(Some(carrier)), + _ => unreachable!("the cable is a loopback transport"), + }; + + // One tick with carrier up, so the drop below is an edge. + set_carrier(&nodes[1], true); + nodes[1].node.run_path_heartbeats().await; + for _ in 0..8 { + if process_available_packets(&mut nodes).await == 0 { + break; + } + } + assert_eq!( + nodes[1].node.get_peer(&addr_0).unwrap().transport_id(), + Some(cable), + "precondition: traffic on the cable" + ); + + // The cable loses carrier. Inside one tick: Suspect, selection moves + // traffic to the wifi, and the peer is told on the wifi. + set_carrier(&nodes[1], false); + nodes[1].node.run_path_heartbeats().await; + { + let peer = nodes[1].node.get_peer(&addr_0).unwrap(); + assert_eq!(peer.path_on(cable).unwrap().state(), PathState::Suspect); + assert_eq!(peer.transport_id(), Some(wifi()), "moved to the standby"); + } + + // Node 0 takes the tick's frames: the heartbeats it acks, and the + // PathClose that withdraws its cable path and moves its traffic too. + for _ in 0..8 { + if process_available_packets(&mut nodes).await == 0 { + break; + } + } + let peer = nodes[0].node.get_peer(&addr_1).unwrap(); + assert_eq!( + peer.path_on(cable).unwrap().state(), + PathState::Dead, + "closed by the peer's PathClose, not by an echo timeout" + ); + assert_eq!(peer.transport_id(), Some(wifi()), "and traffic followed"); +} + +#[tokio::test] +async fn a_peer_closing_our_active_path_moves_our_traffic() { + let (mut nodes, _wifi_0, _wifi_1) = pair_with_wifi_live().await; + let addr_0 = *nodes[0].node.node_addr(); + let addr_1 = *nodes[1].node.node_addr(); + let cable = nodes[0].transport_id; + + // One heartbeat exchange on the cable so each side knows the other's + // id for it (promotion proves the path but names nothing). + nodes[0].node.run_path_heartbeats().await; + nodes[1].node.run_path_heartbeats().await; + for _ in 0..4 { + if process_available_packets(&mut nodes).await == 0 { + break; + } + } + + // Node 0 sends on the cable. Node 1 loses the cable and says so on the wifi. + nodes[1].node.withdraw_transport(cable).await; + assert_eq!( + nodes[1].node.get_peer(&addr_0).unwrap().transport_id(), + Some(wifi()) + ); + let packet = nodes[0] + .packet_rx + .try_recv() + .expect("PathClose on the wifi"); + assert_eq!(packet.transport_id, wifi()); + nodes[0].node.handle_encrypted_frame(packet).await; + let peer = nodes[0].node.get_peer(&addr_1).unwrap(); + assert_eq!( + peer.transport_id(), + Some(wifi()), + "moved without waiting for a timeout" + ); + assert_eq!(peer.path_on(cable).unwrap().state(), PathState::Dead); +} + +#[test] +fn an_outage_does_not_keep_charging_the_path_s_etx() { + let mut peer = dual_path_peer(1, 5); + let t0 = 1_000_000; + let plan = peer.plan_heartbeats(t0, &TIMING); + let wifi_id = plan + .sends + .iter() + .find(|s| s.transport_id == tid(WIFI)) + .unwrap() + .probe_id; + peer.note_path_ack(tid(WIFI), wifi_id, false, 1, t0 + 5, u64::MAX); + // The cable goes silent: first timeout is one loss and the Suspect mark. + let plan = peer.plan_heartbeats(t0 + TIMEOUT, &TIMING); + assert_eq!(plan.suspects, vec![tid(CABLE)]); + let after_one = peer.path_on(tid(CABLE)).unwrap().etx(); + // Sixteen more seconds of timeouts while Suspect: no further charge. + let mut now = t0 + TIMEOUT; + for _ in 0..16 { + now += 1_000; + peer.plan_heartbeats(now, &TIMING); + } + assert_eq!( + peer.path_on(tid(CABLE)).unwrap().etx(), + after_one, + "an outage is one event, not a lossy medium" + ); +} diff --git a/src/node/tests/routing.rs b/src/node/tests/routing.rs index fb8fd58b..07bf83ca 100644 --- a/src/node/tests/routing.rs +++ b/src/node/tests/routing.rs @@ -1717,8 +1717,14 @@ fn test_seam_link_cost_etx_orders_bloom_candidates() { seam_set_cost(&mut node, &low, 3.0, 1_000); seam_set_cost(&mut node, &high, 1.0, 1_000); - let cost_low = node.get_peer(&low).unwrap().link_cost(); - let cost_high = node.get_peer(&high).unwrap().link_cost(); + let cost_low = node + .get_peer(&low) + .unwrap() + .link_cost(crate::time::mono_ms()); + let cost_high = node + .get_peer(&high) + .unwrap() + .link_cost(crate::time::mono_ms()); assert!( cost_high < cost_low, "fixture: ETX alone must make high cheaper ({cost_high} vs {cost_low})" @@ -1749,8 +1755,14 @@ fn test_seam_link_cost_srtt_orders_bloom_candidates() { seam_set_cost(&mut node, &low, 1.0, 50_000); // 50 ms -> cost 1.5 seam_set_cost(&mut node, &high, 1.0, 1_000); // 1 ms -> cost 1.01 - let cost_low = node.get_peer(&low).unwrap().link_cost(); - let cost_high = node.get_peer(&high).unwrap().link_cost(); + let cost_low = node + .get_peer(&low) + .unwrap() + .link_cost(crate::time::mono_ms()); + let cost_high = node + .get_peer(&high) + .unwrap() + .link_cost(crate::time::mono_ms()); assert!( cost_high < cost_low, "fixture: SRTT alone must make high cheaper ({cost_high} vs {cost_low})" @@ -1951,7 +1963,10 @@ fn test_seam_routing_view_reads_match_live_peer_state() { let addr = view.peer_addr(*peer); let live = node.peers.get(&addr).unwrap(); assert_eq!(view.peer_may_reach(*peer, &dest), live.may_reach(&dest)); - assert_eq!(view.peer_link_cost(*peer), live.link_cost()); + assert_eq!( + view.peer_link_cost(*peer), + live.link_cost(crate::time::mono_ms()) + ); assert_eq!( view.peer_coords(*peer), node.tree_state().peer_coords(&addr) diff --git a/src/node/tests/spanning_tree.rs b/src/node/tests/spanning_tree.rs index eb851b76..4eebe42c 100644 --- a/src/node/tests/spanning_tree.rs +++ b/src/node/tests/spanning_tree.rs @@ -19,13 +19,13 @@ static LARGE_NETWORK_TEST_LOCK: std::sync::LazyLock> = /// address. Each node gets a unique synthetic address (`loopback:{n}`) from /// `LOOPBACK_ADDR_COUNTER`, so addresses never collide across concurrently /// running tests and stale entries from finished tests are harmless. -static LOOPBACK_REGISTRY: std::sync::LazyLock = +pub(super) static LOOPBACK_REGISTRY: std::sync::LazyLock = std::sync::LazyLock::new(new_registry); static LOOPBACK_ADDR_COUNTER: std::sync::atomic::AtomicU64 = std::sync::atomic::AtomicU64::new(0); /// Allocate the next globally-unique loopback address. -fn next_loopback_addr() -> TransportAddr { +pub(super) fn next_loopback_addr() -> TransportAddr { let n = LOOPBACK_ADDR_COUNTER.fetch_add(1, std::sync::atomic::Ordering::Relaxed); TransportAddr::from_string(&format!("loopback:{}", n)) } diff --git a/src/node/tree.rs b/src/node/tree.rs index f228db8c..b82d2c6f 100644 --- a/src/node/tree.rs +++ b/src/node/tree.rs @@ -329,11 +329,12 @@ impl Node { // Re-evaluate parent selection with current link costs. // Exclude peers without MMP RTT data — they are not yet eligible // as parent candidates (prevents oscillation from optimistic defaults). + let now_ms = crate::time::mono_ms(); let peer_costs: BTreeMap = self .peers .iter() .filter(|(_, peer)| peer.has_srtt()) - .map(|(addr, peer)| (*addr, peer.link_cost())) + .map(|(addr, peer)| (*addr, peer.link_cost(now_ms))) .collect(); // No peers are excluded from parent candidacy on this branch; the // non-full/leaf skip is a next-only shell refinement. @@ -600,11 +601,12 @@ impl Node { self.last_parent_reeval = Some(now); + let now_ms = crate::time::mono_ms(); let peer_costs: BTreeMap = self .peers .iter() .filter(|(_, peer)| peer.has_srtt()) - .map(|(addr, peer)| (*addr, peer.link_cost())) + .map(|(addr, peer)| (*addr, peer.link_cost(now_ms))) .collect(); // No peers are excluded from parent candidacy on this branch; the // non-full/leaf skip is a next-only shell refinement. @@ -752,11 +754,12 @@ impl Node { // before the recovery mutation — exactly as before. self.metrics().tree.parent_losses.inc(); + let now_ms = crate::time::mono_ms(); let peer_costs: BTreeMap = self .peers .iter() .filter(|(_, peer)| peer.has_srtt()) - .map(|(addr, peer)| (*addr, peer.link_cost())) + .map(|(addr, peer)| (*addr, peer.link_cost(now_ms))) .collect(); // Wall-clock seconds stamped onto the new declaration; monotonic ms for diff --git a/src/peer/active.rs b/src/peer/active.rs index c4cb9e91..aa5a93b8 100644 --- a/src/peer/active.rs +++ b/src/peer/active.rs @@ -3,7 +3,7 @@ //! Represents a fully authenticated peer after successful Noise handshake. //! ActivePeer holds tree state, Bloom filter, and routing information. -use crate::config::MmpConfig; +use crate::config::{MmpConfig, TransportRole}; use crate::node::REKEY_JITTER_SECS; use crate::noise::{HandshakeState as NoiseHandshakeState, NoiseError, NoiseSession}; use crate::proto::bloom::BloomFilter; @@ -15,9 +15,25 @@ use crate::utils::index::SessionIndex; use crate::{FipsAddress, NodeAddr, PeerIdentity}; use rand::RngExt; use secp256k1::XOnlyPublicKey; +use std::collections::VecDeque; use std::fmt; use std::time::{Duration, Instant}; +/// How often a full-size (MTU-padded) probe goes out on a proven path. +const FULL_SIZE_PROBE_INTERVAL_MS: u64 = 60_000; + +/// Fold one probe outcome into a path's ETX: the long EWMA (α = 1/32) of +/// the delivery ratio, inverted and clamped like the link ETX. Per report a +/// raw value is a flap generator on a lightly loaded link; the long average +/// is what selection reads. +fn smooth_etx(etx: f64, delivered: bool) -> f64 { + let alpha = crate::proto::mmp::EWMA_LONG_ALPHA; + let ratio = (1.0 / etx).clamp(0.01, 1.0); + let sample = if delivered { 1.0 } else { 0.0 }; + let next = ratio + alpha * (sample - ratio); + (1.0 / next.max(0.01)).clamp(1.0, 100.0) +} + /// Draw a fresh per-session rekey jitter from `[-REKEY_JITTER_SECS, +REKEY_JITTER_SECS]`. fn draw_rekey_jitter() -> i64 { rand::rng().random_range(-REKEY_JITTER_SECS..=REKEY_JITTER_SECS) @@ -54,6 +70,156 @@ impl fmt::Display for ConnectivityState { } } +/// Where a path is in its life. See `docs/design/fips-multi-path-switchover.md` §4, §7. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub enum PathState { + /// Added, probe outstanding or never answered. Never eligible to carry + /// traffic: an unmeasured path is not assumed good. + Probing, + /// The peer has acknowledged a probe on this path: both directions work. + Live, + /// A hard signal (carrier, send error, failed ack) says the path may be + /// gone. Selection acts on this at once because the standby is warm. + Suspect, + /// Confirmed gone. Kept with its history so a returning transport does + /// not start from scratch. + Dead, +} + +/// Probe bookkeeping for one path: what is outstanding, and when the next +/// one may go. +#[derive(Clone, Copy, Debug, Default)] +struct ProbeState { + /// The next `probe_id` to use. Per path, both directions. + next_id: u32, + /// `(probe_id, sent_at_ms)` of the probe awaiting its ack. + outstanding: Option<(u32, u64)>, + /// `(probe_id, sent_at_ms)` of the last probe whose echo timed out. + /// Its ack is still an ack: on a medium whose round trip exceeds the + /// timeout floor (Tor, Nym, satellite) the first echo is always late, + /// and discarding it would leave the path unmeasured, so the timeout + /// never stretches and the path never proves itself. + timed_out: Option<(u32, u64)>, + /// Probes sent since the last ack, for the backoff. + unanswered: u32, + /// Earliest monotonic ms at which another probe may be sent. + next_at_ms: u64, +} + +/// The knobs selection is bounded by. Built from `node.path.*`. +#[derive(Clone, Copy, Debug)] +pub struct PathPolicy { + /// Discretionary switch margin `K`. + pub margin: f64, + /// Discretionary switch dwell `D`, ms. + pub dwell_ms: u64, + /// RTT samples `N` a path needs before it is selectable. + pub min_samples: u32, + /// The min-RTT window, ms. At least `N` standby heartbeat intervals, + /// otherwise a standby never accumulates a min. + pub rtt_window_ms: u64, +} + +impl PathPolicy { + /// Everything selectable at once; for tests. + pub const PERMISSIVE: Self = Self { + margin: 1.5, + dwell_ms: 0, + min_samples: 0, + rtt_window_ms: u64::MAX, + }; +} + +/// A post-switch hold on the link cost the tree sees. +#[derive(Clone, Copy, Debug)] +struct CostHold { + /// The cost as it was at the switch. + cost: f64, + /// Monotonic ms the dwell runs to. + until_ms: u64, + /// Receiver reports still to arrive before the hold may release: the + /// one that spans the switch, and the one that replaces its ETX. + reports_pending: u8, +} + +/// Why selection moved the active path. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub enum SwitchReason { + /// The active path was not `tx_live`: switched at once, no margin. + Mandatory, + /// The active path scored worse than the best standby by the margin, + /// for the dwell. + Discretionary, + /// The operator pinned another path. + Pinned, +} + +/// The intervals [`ActivePeer::plan_heartbeats`] runs on, all in ms. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub struct HeartbeatTiming { + /// Interval on a path either side sends on. + pub fast_ms: u64, + /// Interval on a standby. + pub slow_ms: u64, + /// Echo timeout floor; stretched per path by its round trip. + pub timeout_ms: u64, + /// Ceiling on the discovery backoff of a path the peer has never + /// acknowledged: an old node, or a medium the peer cannot hear us on. + /// Every such probe is full-size, so this bounds a permanent cost. + pub discovery_cap_ms: u64, +} + +/// One heartbeat probe to put on the wire, from +/// [`ActivePeer::plan_heartbeats`]. +#[derive(Clone, Debug, PartialEq, Eq)] +pub struct HeartbeatSend { + pub transport_id: TransportId, + pub addr: TransportAddr, + pub probe_id: u32, + /// Whether this is the path we send on. + pub remote_active: bool, + /// Our identifier for the path. + pub path_id: u32, + /// Pad this probe to the link MTU: the first probe on a path, and one a + /// minute after, so a medium that forwards small frames and drops large + /// ones never proves itself. + pub full_size: bool, +} + +/// What one pass of [`ActivePeer::plan_heartbeats`] decided. +#[derive(Clone, Debug, Default, PartialEq, Eq)] +pub struct HeartbeatPlan { + /// Probes to send now. + pub sends: Vec, + /// Paths whose outstanding probe timed out and went `Suspect`. + pub suspects: Vec, +} + +/// Selection moved the active path from `from` to `to`. +#[derive(Clone, Debug, PartialEq, Eq)] +pub struct PathSwitch { + pub from: (TransportId, TransportAddr), + pub to: (TransportId, TransportAddr), + pub reason: SwitchReason, +} + +/// What withdrawing a path did to the peer's send side. +#[derive(Clone, Debug, PartialEq, Eq)] +pub enum PathWithdrawal { + /// The peer had no path on that transport. + NoPath, + /// A standby went `Dead`; traffic was never on it. + Standby, + /// The active path went `Dead` and an eligible path took over. + Switched { + from: (TransportId, TransportAddr), + to: (TransportId, TransportAddr), + }, + /// The active path went `Dead` and nothing eligible remains: the peer + /// is unreachable and the caller reaps it. + NoAlternative, +} + /// One transport-level path to a peer. /// /// Path identity is the transport instance: one transport holds at most one @@ -62,15 +228,53 @@ impl fmt::Display for ConnectivityState { /// the peer; a path carries only what is bound to the medium it runs over. /// See `docs/design/fips-multi-path-switchover.md` §1–2. /// -/// Today a peer holds at most one path, added at promotion. The probe -/// exchange that adds further paths under the existing session is the next -/// step of that design; nothing here assumes a single path. +/// The first path is added at promotion, proven by the handshake. Further +/// ones are added by the probe exchange under the existing session (§4). #[derive(Debug)] pub struct PeerPath { /// The transport instance this path runs over. transport_id: TransportId, /// The peer's current address on that transport (roams). addr: TransportAddr, + state: PathState, + /// Monotonic ms of the last authentic frame heard on this path. Free: + /// any authentic frame proves the peer can reach us here. + rx_live_at_ms: Option, + /// Monotonic ms of the last ack proving the peer hears us here. Needs + /// the echo: hearing the peer is a hint, not proof our direction works. + tx_live_at_ms: Option, + /// The peer's last word on whether this is the path it sends on. + remote_active: bool, + /// Most recent probe round trip on this path, ms. + last_rtt_ms: Option, + /// When the path went `Dead`, for the history grace period. + dead_since_ms: Option, + probe: ProbeState, + /// `(sampled_at_ms, rtt_ms)` probe round trips inside the min-RTT + /// window. Min, not smoothed: srtt inflates under load (wifi + /// bufferbloat) while an idle standby looks pristine, which is a + /// ping-pong generator; min RTT is a property of the medium. + rtt_window: VecDeque<(u64, u64)>, + /// Per-path expected transmission count, from the probe/heartbeat ack + /// ratio, long-EWMA smoothed. 1.0 until measured. + etx: f64, + /// Whether the peer has ever acknowledged a probe on this path. Gates + /// the failed-echo signal: a peer that never answers probes is an old + /// node, not a dead path. + acked_once: bool, + /// From the transport's config: a `Backup` path never carries traffic + /// while a `Normal` one is eligible. + role: TransportRole, + /// Our identifier for this path, carried in every probe and ack we + /// send on it, so the peer can name it in a `PathClose`. + local_id: u32, + /// The peer's identifier for its side of this path, learned from its + /// probes and acks. + remote_id: Option, + /// When a full-size probe last went out on this path. + last_full_probe_ms: Option, + /// Operator override: wins selection while it is `tx_live`. + pinned: bool, /// Unix UDP fast-path: per-path `connect()`-ed socket (paired with /// the listen socket via `SO_REUSEPORT`). The kernel demux prefers @@ -90,10 +294,25 @@ pub struct PeerPath { } impl PeerPath { - fn new(transport_id: TransportId, addr: TransportAddr) -> Self { + fn new(transport_id: TransportId, addr: TransportAddr, state: PathState) -> Self { Self { transport_id, addr, + state, + rx_live_at_ms: None, + tx_live_at_ms: None, + remote_active: false, + last_rtt_ms: None, + dead_since_ms: None, + probe: ProbeState::default(), + rtt_window: VecDeque::new(), + etx: 1.0, + acked_once: false, + role: TransportRole::Normal, + local_id: rand::rng().random::(), + remote_id: None, + last_full_probe_ms: None, + pinned: false, #[cfg(any(target_os = "linux", target_os = "macos"))] connected_udp: None, #[cfg(any(target_os = "linux", target_os = "macos"))] @@ -111,6 +330,120 @@ impl PeerPath { &self.addr } + /// Where the path is in its life. + pub fn state(&self) -> PathState { + self.state + } + + /// Monotonic ms of the last authentic frame heard on this path. + pub fn rx_live_at_ms(&self) -> Option { + self.rx_live_at_ms + } + + /// Monotonic ms of the last ack proving the peer hears us here. + pub fn tx_live_at_ms(&self) -> Option { + self.tx_live_at_ms + } + + /// Whether the peer last said it sends on this path. + pub fn remote_active(&self) -> bool { + self.remote_active + } + + /// Most recent probe round trip on this path, ms. + pub fn last_rtt_ms(&self) -> Option { + self.last_rtt_ms + } + + /// Whether the path may carry our traffic: `Live`, and the peer has + /// acknowledged hearing us on it. + pub fn is_eligible(&self) -> bool { + self.state == PathState::Live && self.tx_live_at_ms.is_some() + } + + /// The transport's role, from its config. + pub fn role(&self) -> TransportRole { + self.role + } + + /// Record the transport's role. + pub fn set_role(&mut self, role: TransportRole) { + self.role = role; + } + + /// Whether the operator pinned traffic to this path. + pub fn pinned(&self) -> bool { + self.pinned + } + + /// Whether the peer has ever acknowledged a probe here. + pub fn acked_once(&self) -> bool { + self.acked_once + } + + /// Our identifier for this path on the wire. + pub fn local_id(&self) -> u32 { + self.local_id + } + + /// The peer's identifier for its side of this path, once heard. + pub fn remote_id(&self) -> Option { + self.remote_id + } + + /// Minimum probe round trip inside the window, ms. + pub fn min_rtt_ms(&self) -> Option { + self.rtt_window.iter().map(|(_, rtt)| *rtt).min() + } + + /// RTT samples inside the window. + pub fn rtt_samples(&self) -> u32 { + self.rtt_window.len() as u32 + } + + /// Per-path expected transmission count, smoothed. + pub fn etx(&self) -> f64 { + self.etx + } + + /// The path's quality index, [`quality_index`](crate::proto::mmp::quality_index) + /// of its ETX and min RTT, lower is better. `None` until an RTT has been + /// measured. + pub fn score(&self) -> Option { + self.min_rtt_ms() + .map(|rtt| crate::proto::mmp::quality_index(self.etx, rtt as f64)) + } + + /// Whether selection may pick this path: eligible, with enough samples, + /// per `policy`. + pub fn is_selectable(&self, policy: &PathPolicy) -> bool { + self.is_eligible() && self.rtt_samples() >= policy.min_samples + } + + fn record_rtt(&mut self, now_ms: u64, rtt_ms: u64, window_ms: u64) { + self.rtt_window.push_back((now_ms, rtt_ms)); + while let Some((at, _)) = self.rtt_window.front() { + if now_ms.saturating_sub(*at) > window_ms { + self.rtt_window.pop_front(); + } else { + break; + } + } + } + + /// Whether a probe may be sent now, per the backoff. + pub fn probe_due(&self, now_ms: u64) -> bool { + now_ms >= self.probe.next_at_ms + } + + /// Reset the probe backoff so the next tick may probe at once. Called + /// when the transport's presence cycles: the medium has changed, so what + /// went unanswered before says nothing about now. + pub fn reset_probe_backoff(&mut self) { + self.probe.unanswered = 0; + self.probe.next_at_ms = 0; + } + /// Drop the connected socket and its drain. The drain goes first so /// its last fd reference is released cleanly; the kernel fd closes on /// the last `Arc` drop, so in-flight worker jobs holding the old `Arc` @@ -178,6 +511,17 @@ struct PeerSendState { /// Index into `paths` of the path *our* frames go out on. `None` only /// while `paths` is empty. active: Option, + /// When the active path first scored worse than the best standby by + /// the margin; the discretionary dwell counts from here. + discretionary_since_ms: Option, + /// The link cost reported to the tree is held at its pre-switch value + /// after a switch, so a short switch (cable flap, replug) does not + /// ripple mesh-wide (design §10). Released once the dwell has passed + /// *and* the receiver reports that span the switch have been replaced: + /// the first report after a switch counts every frame in flight on the + /// old path as lost and spikes the per-report ETX for one interval, + /// and that spike must not reach parent selection. + cost_hold: Option, /// Link used to reach this peer. link_id: LinkId, @@ -212,6 +556,8 @@ impl PeerSendState { session_start, paths: Vec::new(), active: None, + discretionary_since_ms: None, + cost_hold: None, link_id, link_stats: LinkStats::new(), last_seen, @@ -418,7 +764,11 @@ impl ActivePeer { send.noise_session = Some(noise_session); send.our_index = Some(our_index); send.their_index = Some(their_index); - send.paths.push(PeerPath::new(transport_id, current_addr)); + let mut path = PeerPath::new(transport_id, current_addr, PathState::Live); + let now_ms = crate::time::mono_ms(); + path.rx_live_at_ms = Some(now_ms); + path.tx_live_at_ms = Some(now_ms); + send.paths.push(path); send.active = Some(0); send.link_stats = link_stats; send.mmp = Some(MmpPeerState::new( @@ -714,12 +1064,607 @@ impl ActivePeer { path.transport_id = transport_id; path.addr = addr; } else { - self.send.paths.push(PeerPath::new(transport_id, addr)); + self.send + .paths + .push(PeerPath::new(transport_id, addr, PathState::Live)); self.send.active = Some(0); } true } + // === Path set === + + /// The path on `transport_id`, if the peer has one. + pub fn path_on(&self, transport_id: TransportId) -> Option<&PeerPath> { + self.send + .paths + .iter() + .find(|path| path.transport_id == transport_id) + } + + fn path_on_mut(&mut self, transport_id: TransportId) -> Option<&mut PeerPath> { + self.send + .paths + .iter_mut() + .find(|path| path.transport_id == transport_id) + } + + /// Add a `Probing` path on `transport_id` at `addr`, or return the one + /// already there. The only way a path is created after promotion: both + /// ends of the probe exchange call this, the prober before it sends and + /// the receiver when a probe arrives. + pub fn add_path(&mut self, transport_id: TransportId, addr: TransportAddr) -> &mut PeerPath { + let idx = match self + .send + .paths + .iter() + .position(|path| path.transport_id == transport_id) + { + Some(idx) => idx, + None => { + self.send + .paths + .push(PeerPath::new(transport_id, addr, PathState::Probing)); + self.send.paths.len() - 1 + } + }; + &mut self.send.paths[idx] + } + + /// An authentic frame arrived on `transport_id`: the path there, if any, + /// is `rx_live` as of `now_ms`. + pub fn note_path_rx(&mut self, transport_id: TransportId, now_ms: u64) { + if let Some(path) = self.path_on_mut(transport_id) { + path.rx_live_at_ms = Some(now_ms); + } + } + + /// Take the next probe to send on `transport_id`: its `probe_id`, and + /// whether the path is the one we send on. Records it as outstanding and + /// advances the backoff; `None` if the peer has no path there or the + /// backoff has not expired. `backoff_cap_ms` bounds the retry interval, + /// which doubles from `base_ms` per unanswered probe. + /// + /// Tests only. In production [`plan_heartbeats`](Self::plan_heartbeats) + /// is the one issuer of probes, so no two writers race for + /// `probe.outstanding`. + #[cfg(test)] + pub fn take_probe( + &mut self, + transport_id: TransportId, + now_ms: u64, + base_ms: u64, + backoff_cap_ms: u64, + ) -> Option<(u32, bool, u32)> { + let active = self.transport_id() == Some(transport_id); + let path = self.path_on_mut(transport_id)?; + if !path.probe_due(now_ms) { + return None; + } + let id = path.probe.next_id; + path.probe.next_id = path.probe.next_id.wrapping_add(1); + path.probe.outstanding = Some((id, now_ms)); + let shift = path.probe.unanswered.min(16); + let delay = base_ms.saturating_mul(1u64 << shift).min(backoff_cap_ms); + path.probe.next_at_ms = now_ms.saturating_add(delay.max(1)); + path.probe.unanswered = path.probe.unanswered.saturating_add(1); + let path_id = path.local_id; + Some((id, active, path_id)) + } + + /// A `PathProbe` arrived on `transport_id` from `addr`. Adds the path if + /// it is new, roams its address if not, marks it `rx_live` and records + /// what the peer said about sending here. + pub fn note_path_probe( + &mut self, + transport_id: TransportId, + addr: TransportAddr, + remote_active: bool, + remote_id: u32, + now_ms: u64, + ) { + let path = self.add_path(transport_id, addr.clone()); + if path.addr != addr { + path.addr = addr; + #[cfg(any(target_os = "linux", target_os = "macos"))] + path.clear_connected_udp(); + } + path.remote_id = Some(remote_id); + if path.state == PathState::Dead { + // The peer is probing a path we had given up on: it is back, + // unproven in our direction until its ack. + path.state = PathState::Probing; + path.dead_since_ms = None; + } + path.rx_live_at_ms = Some(now_ms); + if path.remote_active && !remote_active { + // The peer stopped sending here. It may have stopped hearing us + // here too: a hint, so probe now, never `Suspect` (design §7). + path.probe.next_at_ms = 0; + } + path.remote_active = remote_active; + } + + /// A `PathAck` arrived on `transport_id`. If it answers the outstanding + /// probe, or the one whose echo last timed out, the path is `tx_live` + /// and `Live`, the backoff is cleared and the round trip is sampled. + /// Returns the RTT sample in ms, or `None` if the ack matched nothing + /// (a stale or duplicate ack is ignored). + pub fn note_path_ack( + &mut self, + transport_id: TransportId, + probe_id: u32, + remote_active: bool, + remote_id: u32, + now_ms: u64, + rtt_window_ms: u64, + ) -> Option { + let path = self.path_on_mut(transport_id)?; + path.remote_id = Some(remote_id); + let sent_at_ms = match (path.probe.outstanding, path.probe.timed_out) { + (Some((id, at)), _) if id == probe_id => { + path.probe.outstanding = None; + at + } + (_, Some((id, at))) if id == probe_id => { + path.probe.timed_out = None; + at + } + _ => return None, + }; + path.probe.unanswered = 0; + let rtt_ms = now_ms.saturating_sub(sent_at_ms); + path.last_rtt_ms = Some(rtt_ms); + path.record_rtt(now_ms, rtt_ms, rtt_window_ms); + path.acked_once = true; + path.etx = smooth_etx(path.etx, true); + path.rx_live_at_ms = Some(now_ms); + path.tx_live_at_ms = Some(now_ms); + path.remote_active = remote_active; + path.state = PathState::Live; + path.dead_since_ms = None; + Some(rtt_ms) + } + + /// Whether the peer has a path on `transport_id` at `addr`, whichever + /// path it is. The msg1 classifier asks this: a rekey from the peer + /// arrives on the path *the peer* sends on, which need not be ours. + pub fn is_reachable_at(&self, transport_id: TransportId, addr: &TransportAddr) -> bool { + self.path_on(transport_id) + .is_some_and(|path| path.addr == *addr) + } + + /// The transport `transport_id` went away: the path on it is `Dead`. + /// + /// The path keeps its history (RTT, liveness marks) so a returning + /// transport is re-probed rather than re-measured from nothing; the + /// history expires with [`prune_dead_paths`](Self::prune_dead_paths). + /// If the withdrawn path was the active one, the best eligible path takes + /// over: `Live` and `tx_live`, lowest last RTT first. With no eligible + /// path the active index is left where it was, the path is `Suspect` + /// rather than `Dead` so it keeps being probed, and the caller decides; + /// a `Probing` path is never promoted, because nothing has proven it + /// carries anything. + pub fn withdraw_path( + &mut self, + transport_id: TransportId, + now_ms: u64, + policy: &PathPolicy, + ) -> PathWithdrawal { + let Some(idx) = self + .send + .paths + .iter() + .position(|path| path.transport_id == transport_id) + else { + return PathWithdrawal::NoPath; + }; + { + let path = &mut self.send.paths[idx]; + path.state = PathState::Dead; + path.dead_since_ms = Some(now_ms); + path.probe.outstanding = None; + #[cfg(any(target_os = "linux", target_os = "macos"))] + path.clear_connected_udp(); + } + if self.send.active != Some(idx) { + return PathWithdrawal::Standby; + } + let from = { + let path = &self.send.paths[idx]; + (path.transport_id, path.addr.clone()) + }; + match self.best_alternative(idx, policy) { + Some(next) => { + self.hold_link_cost(now_ms, policy.dwell_ms); + self.send.active = Some(next); + let path = &self.send.paths[next]; + PathWithdrawal::Switched { + from, + to: (path.transport_id, path.addr.clone()), + } + } + None => { + // Nothing to move to. The caller either reaps the peer + // (transport gone) or keeps it (advisory close): in the + // latter case the path must stay probed, and only a + // non-`Dead` path is, so it is `Suspect`, not `Dead`. + let path = &mut self.send.paths[idx]; + path.state = PathState::Suspect; + path.dead_since_ms = None; + PathWithdrawal::NoAlternative + } + } + } + + /// The best selectable path other than `exclude`, with its score: + /// `Normal` before `Backup`, then lowest score; a selectable path with + /// no score yet ranks last. The one rule both selection and withdrawal + /// pick by. A `Backup` is offered only while no selectable `Normal` + /// path exists at all, `exclude` included: a `Normal` active path is + /// never left for a `Backup`, however it scores. + fn best_selectable(&self, exclude: usize, policy: &PathPolicy) -> Option<(usize, Option)> { + let any_normal = self + .send + .paths + .iter() + .any(|p| p.is_selectable(policy) && p.role == TransportRole::Normal); + self.send + .paths + .iter() + .enumerate() + .filter(|(i, p)| *i != exclude && p.is_selectable(policy)) + .filter(|(_, p)| p.role == TransportRole::Normal || !any_normal) + .min_by(|(_, a), (_, b)| { + a.score() + .unwrap_or(f64::MAX) + .total_cmp(&b.score().unwrap_or(f64::MAX)) + }) + .map(|(i, p)| (i, p.score())) + } + + /// The best path other than `exclude` to move traffic to, if any. + /// + /// Selectable paths first ([`best_selectable`](Self::best_selectable)); + /// failing that, any eligible path by last RTT: a `Live` path with + /// fewer than `N` samples beats no path. A `Probing` path is never + /// returned. + fn best_alternative(&self, exclude: usize, policy: &PathPolicy) -> Option { + if let Some((i, _)) = self.best_selectable(exclude, policy) { + return Some(i); + } + let candidates = || { + self.send + .paths + .iter() + .enumerate() + .filter(move |(i, p)| *i != exclude && p.is_eligible()) + }; + let any_normal = candidates().any(|(_, p)| p.role == TransportRole::Normal); + candidates() + .filter(|(_, p)| p.role == TransportRole::Normal || !any_normal) + .min_by_key(|(_, p)| p.last_rtt_ms.unwrap_or(u64::MAX)) + .map(|(i, _)| i) + } + + /// Run selection over the path set (design §8). Returns the switch if + /// the active path changed. + /// + /// - **Pinned:** a pinned path wins while it is `tx_live`. + /// - **Mandatory:** active path not `tx_live` (`Suspect`/`Dead`) → + /// best alternative now, no margin, no dwell. + /// - **Discretionary:** `active.score > best.score × K`, sustained for + /// the dwell → switch. Only selectable paths compete here, and only + /// when the active path itself has a score. + /// - Ties keep the current path. + pub fn select_path(&mut self, now_ms: u64, policy: &PathPolicy) -> Option { + let active = self.send.active?; + let pinned = self + .send + .paths + .iter() + .position(|p| p.pinned && p.is_eligible()); + if let Some(pin) = pinned { + if pin == active { + self.send.discretionary_since_ms = None; + return None; + } + self.hold_link_cost(now_ms, policy.dwell_ms); + return Some(self.switch_to(active, pin, SwitchReason::Pinned)); + } + if !self.send.paths[active].is_eligible() { + let next = self.best_alternative(active, policy)?; + self.hold_link_cost(now_ms, policy.dwell_ms); + return Some(self.switch_to(active, next, SwitchReason::Mandatory)); + } + let Some(active_score) = self.send.paths[active].score() else { + self.send.discretionary_since_ms = None; + return None; + }; + let Some((next, Some(best_score))) = self.best_selectable(active, policy) else { + self.send.discretionary_since_ms = None; + return None; + }; + // A `Backup` active path yields to a selectable `Normal` one outright: + // its role says it should not be carrying traffic at all. + let active_is_backup_yielding = self.send.paths[active].role == TransportRole::Backup + && self.send.paths[next].role == TransportRole::Normal; + if !active_is_backup_yielding && active_score <= best_score * policy.margin { + self.send.discretionary_since_ms = None; + return None; + } + let since = *self.send.discretionary_since_ms.get_or_insert(now_ms); + if !active_is_backup_yielding && now_ms.saturating_sub(since) < policy.dwell_ms { + return None; + } + self.hold_link_cost(now_ms, policy.dwell_ms); + Some(self.switch_to(active, next, SwitchReason::Discretionary)) + } + + /// Whether a post-switch cost hold is in force at `now_ms`. + pub fn link_cost_held(&self, now_ms: u64) -> bool { + self.send + .cost_hold + .is_some_and(|hold| now_ms < hold.until_ms || hold.reports_pending > 0) + } + + /// A receiver report arrived: one fewer to wait for before a + /// post-switch cost hold may release. + pub fn note_receiver_report(&mut self) { + if let Some(hold) = self.send.cost_hold.as_mut() { + hold.reports_pending = hold.reports_pending.saturating_sub(1); + } + } + + fn switch_to(&mut self, from: usize, to: usize, reason: SwitchReason) -> PathSwitch { + self.send.discretionary_since_ms = None; + let from_path = &self.send.paths[from]; + let from_key = (from_path.transport_id, from_path.addr.clone()); + self.send.active = Some(to); + let to_path = &self.send.paths[to]; + PathSwitch { + from: from_key, + to: (to_path.transport_id, to_path.addr.clone()), + reason, + } + } + + /// Pin traffic to the path on `transport_id`. Returns `false` if the + /// peer has no path there. Selection honours the pin on its next run. + pub fn pin_path(&mut self, transport_id: TransportId) -> bool { + let Some(idx) = self + .send + .paths + .iter() + .position(|p| p.transport_id == transport_id) + else { + return false; + }; + for (i, p) in self.send.paths.iter_mut().enumerate() { + p.pinned = i == idx; + } + true + } + + /// Clear any pin. + pub fn unpin_paths(&mut self) { + for p in self.send.paths.iter_mut() { + p.pinned = false; + } + } + + /// Record the transport's role on the path over `transport_id`. + pub fn set_path_role(&mut self, transport_id: TransportId, role: TransportRole) { + if let Some(path) = self.path_on_mut(transport_id) { + path.role = role; + } + } + + /// A hard signal (carrier lost, unreachable on send) says the path on + /// `transport_id` may be gone: `Suspect`, so selection leaves it now + /// and a later ack restores it. Only a `Live` path can become suspect. + pub fn mark_path_suspect(&mut self, transport_id: TransportId) -> bool { + match self.path_on_mut(transport_id) { + Some(path) if path.state == PathState::Live => { + path.state = PathState::Suspect; + true + } + _ => false, + } + } + + /// Decide this tick's per-path heartbeats (design §7). + /// + /// Every path that is not `Dead` is heartbeated with a `PathProbe` + /// whose `probe_id` is the path's sequence: fast (`fast_ms`) on a path + /// that either side sends on, slow (`slow_ms`) on a standby, one probe + /// in flight per path. A probe unanswered for `timeout_ms` is a failed + /// echo: on a path the peer has acknowledged before, that is a hard + /// signal and the path goes `Suspect` (a never-acknowledged path is an + /// old node, not a dead path; it keeps the discovery backoff instead, + /// capped at `discovery_cap_ms`). The timed-out probe is remembered so + /// its late ack still samples the round trip: until a path has one, + /// its timeout cannot stretch. + /// Each timeout is one lost sample for the path's ETX, and a verdict + /// only if the peer has also been silent on the path for the timeout: + /// a late echo on a path still carrying the peer's frames is load, not + /// death. Both the interval + /// and the timeout stretch with the path's measured round trip, so a + /// circuit whose round trip exceeds `fast_ms` is neither flooded nor + /// declared dead every round trip: interval is at least the min RTT, + /// timeout at least three times the last RTT. + /// + /// The silence hint: our active path silent for two of the peer's + /// intervals on it while a standby hears the peer triggers a probe now, + /// never `Suspect` (see §7 for the loop that would otherwise follow). + pub fn plan_heartbeats(&mut self, now_ms: u64, timing: &HeartbeatTiming) -> HeartbeatPlan { + let HeartbeatTiming { + fast_ms, + slow_ms, + timeout_ms, + discovery_cap_ms, + } = *timing; + let mut plan = HeartbeatPlan::default(); + let active = self.send.active; + let newest_rx = self.send.paths.iter().filter_map(|p| p.rx_live_at_ms).max(); + for (i, path) in self.send.paths.iter_mut().enumerate() { + if path.state == PathState::Dead { + continue; + } + let ours = active == Some(i); + let interval = if ours || path.remote_active { + fast_ms.max(path.min_rtt_ms().unwrap_or(0)) + } else { + slow_ms + }; + let timeout_ms = timeout_ms.max(path.last_rtt_ms.unwrap_or(0).saturating_mul(3)); + + if let Some((id, sent_at)) = path.probe.outstanding + && now_ms.saturating_sub(sent_at) >= timeout_ms + { + path.probe.outstanding = None; + path.probe.timed_out = Some((id, sent_at)); + if path.acked_once { + // Loss is sampled while the path is Live. Once it is + // Suspect the state already says it is down, and every + // further timeout is the same outage, not a lossier + // medium; counting them would keep a returning cable + // scoring like a bad link for the next minute and stall + // the fail-back. + if path.state == PathState::Live { + path.etx = smooth_etx(path.etx, false); + } + // A late echo on a path we are still hearing the peer on + // is a loss sample, not a verdict: under load the echo + // queues behind data and comes back late while the path + // is plainly carrying traffic. Only a path silent in + // both directions for the timeout goes Suspect. + let heard_recently = path + .rx_live_at_ms + .is_some_and(|rx| now_ms.saturating_sub(rx) < timeout_ms); + if path.state == PathState::Live && !heard_recently { + path.state = PathState::Suspect; + plan.suspects.push(path.transport_id); + } + path.probe.next_at_ms = 0; + } else { + path.probe.unanswered = path.probe.unanswered.saturating_add(1); + } + } + + // Silence hint on our active path. + if ours + && let Some(last_rx) = path.rx_live_at_ms + && let Some(newest) = newest_rx + { + let peer_interval = if path.remote_active { fast_ms } else { slow_ms }; + if newest > last_rx && now_ms.saturating_sub(last_rx) >= 2 * peer_interval { + path.probe.next_at_ms = 0; + } + } + + if path.probe.outstanding.is_some() || now_ms < path.probe.next_at_ms { + continue; + } + let id = path.probe.next_id; + path.probe.next_id = path.probe.next_id.wrapping_add(1); + path.probe.outstanding = Some((id, now_ms)); + let delay = if path.acked_once { + interval + } else { + // Discovery backoff: doubles per unanswered probe, capped at + // `discovery_cap_ms`. Every one of these is full-size, and + // an old node never answers, so the cap is a permanent + // per-path cost. + let shift = path.probe.unanswered.min(16); + interval + .saturating_mul(1u64 << shift) + .min(discovery_cap_ms.max(interval)) + }; + path.probe.next_at_ms = now_ms.saturating_add(delay.max(1)); + let full_size = !path.acked_once + || path + .last_full_probe_ms + .is_none_or(|t| now_ms.saturating_sub(t) >= FULL_SIZE_PROBE_INTERVAL_MS); + if full_size { + path.last_full_probe_ms = Some(now_ms); + } + plan.sends.push(HeartbeatSend { + transport_id: path.transport_id, + addr: path.addr.clone(), + probe_id: id, + remote_active: ours, + path_id: path.local_id, + full_size, + }); + } + plan + } + + /// The peer closed the path we call `local_id` (a `PathClose` names + /// the receiver's id). Same as losing the transport under it: `Dead` + /// with history, traffic moved if it was there. Returns the transport + /// it was on alongside the outcome. + pub fn withdraw_path_by_local_id( + &mut self, + local_id: u32, + now_ms: u64, + policy: &PathPolicy, + ) -> Option<(TransportId, PathWithdrawal)> { + let transport_id = self + .send + .paths + .iter() + .find(|p| p.local_id == local_id)? + .transport_id; + Some(( + transport_id, + self.withdraw_path(transport_id, now_ms, policy), + )) + } + + /// Forget `Dead` paths older than `grace_ms`. The active path is never + /// pruned, whatever its state: the peer is reaped, not trimmed. + pub fn prune_dead_paths(&mut self, now_ms: u64, grace_ms: u64) { + let Some(active) = self.send.active else { + return; + }; + let keep: Vec = self + .send + .paths + .iter() + .enumerate() + .map(|(i, path)| { + i == active + || path.state != PathState::Dead + || path + .dead_since_ms + .is_none_or(|since| now_ms.saturating_sub(since) < grace_ms) + }) + .collect(); + if keep.iter().all(|k| *k) { + return; + } + let mut i = 0; + let mut new_active = active; + self.send.paths.retain(|_| { + let k = keep[i]; + if !k && i < active { + new_active -= 1; + } + i += 1; + k + }); + self.send.active = Some(new_active); + } + + /// Clear the probe backoff on every path over `transport_id`. + pub fn reset_probe_backoff_on(&mut self, transport_id: TransportId) { + if let Some(path) = self.path_on_mut(transport_id) { + path.reset_probe_backoff(); + } + } + // === Handshake Resend === /// Store wire-format msg2 for resend on duplicate msg1. @@ -871,19 +1816,47 @@ impl ActivePeer { /// /// Returns 1.0 (optimistic default) when MMP metrics are not yet /// available, matching depth-only parent selection behavior. - pub fn link_cost(&self) -> f64 { + /// + /// Reads the smoothed (long EWMA) ETX rather than the per-report value: + /// after a path switch the next report spans the gap and produces one + /// ETX spike, which must not reach the tree. While a cost hold is in + /// force after a switch the pre-switch cost is returned instead. + pub fn link_cost(&self, now_ms: u64) -> f64 { + if let Some(hold) = self.send.cost_hold + && self.link_cost_held(now_ms) + { + return hold.cost; + } + self.raw_link_cost() + } + + /// The per-report quality index, as the tree has always read it. + fn raw_link_cost(&self) -> f64 { match self.mmp() { - Some(mmp) => { - let etx = mmp.metrics.etx; - match mmp.metrics.srtt_ms() { - Some(srtt_ms) => crate::proto::mmp::quality_index(etx, srtt_ms), - None => 1.0, - } - } + Some(mmp) => match mmp.metrics.srtt_ms() { + Some(srtt_ms) => crate::proto::mmp::quality_index(mmp.metrics.etx, srtt_ms), + None => 1.0, + }, None => 1.0, } } + /// Hold the reported link cost at its current value: for `hold_ms`, + /// and until the two receiver reports after the switch have arrived + /// (the one that spans it, whose ETX carries the frames lost in flight, + /// and the one that replaces that ETX). + fn hold_link_cost(&mut self, now_ms: u64, hold_ms: u64) { + if hold_ms == 0 { + return; + } + let cost = self.link_cost(now_ms); + self.send.cost_hold = Some(CostHold { + cost, + until_ms: now_ms.saturating_add(hold_ms), + reports_pending: 2, + }); + } + /// Whether this peer has at least one MMP RTT measurement. pub fn has_srtt(&self) -> bool { self.mmp() diff --git a/src/peer/mod.rs b/src/peer/mod.rs index 3f493fe3..08212a85 100644 --- a/src/peer/mod.rs +++ b/src/peer/mod.rs @@ -8,7 +8,10 @@ mod active; pub(crate) mod machine; -pub use active::{ActivePeer, ConnectivityState, PeerPath}; +pub use active::{ + ActivePeer, ConnectivityState, HeartbeatPlan, HeartbeatSend, HeartbeatTiming, PathPolicy, + PathState, PathSwitch, PathWithdrawal, PeerPath, SwitchReason, +}; use crate::NodeAddr; use crate::transport::LinkId; diff --git a/src/proto/link.rs b/src/proto/link.rs index ebc4d26b..9cb390cb 100644 --- a/src/proto/link.rs +++ b/src/proto/link.rs @@ -47,6 +47,16 @@ pub enum LinkMessageType { /// Periodic heartbeat for link liveness detection. /// No payload — the msg_type byte alone is sufficient. Heartbeat = 0x51, + /// Probe a candidate path to a peer under the existing session. + /// Payload is a [`PathMessage`]. + PathProbe = 0x52, + /// Answer to a [`LinkMessageType::PathProbe`], sent back on the path + /// the probe arrived on. Payload is a [`PathMessage`]. + PathAck = 0x53, + /// "I am closing this path": sent on any other path when the sender + /// knows a path is going (interface gone, carrier lost). Payload is a + /// [`PathClose`]. + PathClose = 0x54, } impl LinkMessageType { @@ -62,6 +72,9 @@ impl LinkMessageType { 0x31 => Some(LinkMessageType::LookupResponse), 0x50 => Some(LinkMessageType::Disconnect), 0x51 => Some(LinkMessageType::Heartbeat), + 0x52 => Some(LinkMessageType::PathProbe), + 0x53 => Some(LinkMessageType::PathAck), + 0x54 => Some(LinkMessageType::PathClose), _ => None, } } @@ -84,11 +97,165 @@ impl fmt::Display for LinkMessageType { LinkMessageType::LookupResponse => "LookupResponse", LinkMessageType::Disconnect => "Disconnect", LinkMessageType::Heartbeat => "Heartbeat", + LinkMessageType::PathProbe => "PathProbe", + LinkMessageType::PathAck => "PathAck", + LinkMessageType::PathClose => "PathClose", }; write!(f, "{}", name) } } +// ============================================================================ +// Path Probe / Path Ack +// ============================================================================ + +/// Payload shared by `PathProbe` (0x52) and `PathAck` (0x53). +/// +/// A probe is an ordinary encrypted frame under the current session, sent on +/// a candidate transport. The receiver, having decrypted it against the +/// session found by index, has proof the peer is reachable there: it adds +/// the path and answers with an ack **on that same path**. The prober's +/// receipt of the ack proves the reverse direction. One round trip, no +/// handshake, no new key material, no index allocation. +/// +/// ## Wire Format +/// +/// | Offset | Field | Size | Notes | +/// |--------|---------------|---------|-----------------------------------------| +/// | 0 | msg_type | 1 byte | 0x52 or 0x53 | +/// | 1 | probe_id | 4 bytes | LE; the ack echoes the probe's | +/// | 5 | flags | 1 byte | bit 0: `remote_active` | +/// | 6 | path_id | 4 bytes | LE; the sender's id for its path | +/// | 10 | padding | any | ignored; a full-size probe pads to MTU | +/// +/// Trailing bytes are ignored on decode, so a probe may be padded to the +/// link MTU: a path that forwards small frames and drops large ones then +/// never proves itself. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub struct PathMessage { + /// Per-path sequence chosen by the prober; the ack carries it back. + pub probe_id: u32, + /// "This path is where I currently send." Costs one bit and is a free + /// detection signal: a peer that stops sending here may have stopped + /// hearing us here too. + pub remote_active: bool, + /// The sender's own identifier for the path this travels on. Transport + /// ids are local to each node; each side learns the other's id from its + /// probes and acks, and a later [`PathClose`] names the path by the + /// *receiver's* id, so the receiver needs no lookup table. + pub path_id: u32, +} + +impl PathMessage { + /// Encoded size including the msg_type byte. + pub const WIRE_SIZE: usize = 10; + + const FLAG_REMOTE_ACTIVE: u8 = 0x01; + + /// Encode as a `PathProbe` link message (msg_type included). + pub fn encode_probe(&self) -> [u8; Self::WIRE_SIZE] { + self.encode(LinkMessageType::PathProbe) + } + + /// Encode as a `PathAck` link message (msg_type included). + pub fn encode_ack(&self) -> [u8; Self::WIRE_SIZE] { + self.encode(LinkMessageType::PathAck) + } + + fn encode(&self, kind: LinkMessageType) -> [u8; Self::WIRE_SIZE] { + let mut out = [0u8; Self::WIRE_SIZE]; + out[0] = kind.to_byte(); + out[1..5].copy_from_slice(&self.probe_id.to_le_bytes()); + if self.remote_active { + out[5] |= Self::FLAG_REMOTE_ACTIVE; + } + out[6..10].copy_from_slice(&self.path_id.to_le_bytes()); + out + } + + /// Decode from the link-layer payload (after the msg_type byte). + pub fn decode(payload: &[u8]) -> Result { + let mut reader = crate::proto::codec::Reader::new(payload); + let probe_id = reader.read_u32_le()?; + let flags = reader.read_u8()?; + let path_id = reader.read_u32_le()?; + Ok(Self { + probe_id, + remote_active: flags & Self::FLAG_REMOTE_ACTIVE != 0, + path_id, + }) + } +} + +/// Why a path is being closed. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +#[repr(u8)] +pub enum PathCloseReason { + /// No particular reason given. + Unspecified = 0, + /// The interface under the path went away. + InterfaceGone = 1, + /// The interface lost carrier. + CarrierLost = 2, + /// An operator asked for it. + Operator = 3, +} + +impl PathCloseReason { + fn from_byte(b: u8) -> Self { + match b { + 1 => Self::InterfaceGone, + 2 => Self::CarrierLost, + 3 => Self::Operator, + _ => Self::Unspecified, + } + } +} + +/// `PathClose` (0x54): the sender is closing the path it identifies. +/// +/// Sent on any path that still works, so the peer learns at once rather +/// than after an echo timeout. The peer withdraws its side of that path +/// (keeping its history) and moves its traffic if it was on it. Advisory: +/// a later probe on the path revives it. +/// +/// ## Wire Format +/// +/// | Offset | Field | Size | Notes | +/// |--------|----------|---------|-------------------------------------------| +/// | 0 | msg_type | 1 byte | 0x54 | +/// | 1 | path_id | 4 bytes | LE; the *receiver's* id for the path | +/// | 5 | reason | 1 byte | [`PathCloseReason`] | +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub struct PathClose { + /// The receiver's identifier for the path, as it carried in its own + /// probes and acks on it. + pub path_id: u32, + pub reason: PathCloseReason, +} + +impl PathClose { + /// Encoded size including the msg_type byte. + pub const WIRE_SIZE: usize = 6; + + /// Encode as a link message (msg_type included). + pub fn encode(&self) -> [u8; Self::WIRE_SIZE] { + let mut out = [0u8; Self::WIRE_SIZE]; + out[0] = LinkMessageType::PathClose.to_byte(); + out[1..5].copy_from_slice(&self.path_id.to_le_bytes()); + out[5] = self.reason as u8; + out + } + + /// Decode from the link-layer payload (after the msg_type byte). + pub fn decode(payload: &[u8]) -> Result { + let mut reader = crate::proto::codec::Reader::new(payload); + let path_id = reader.read_u32_le()?; + let reason = PathCloseReason::from_byte(reader.read_u8()?); + Ok(Self { path_id, reason }) + } +} + // ============================================================================ // Session Datagram (Link-Layer Encapsulation) // ============================================================================ diff --git a/src/transport/ble/mod.rs b/src/transport/ble/mod.rs index ed29e456..8ddd0cd6 100644 --- a/src/transport/ble/mod.rs +++ b/src/transport/ble/mod.rs @@ -790,6 +790,10 @@ impl BleTransport { } impl Transport for BleTransport { + fn role(&self) -> crate::config::TransportRole { + self.config.role() + } + fn transport_id(&self) -> TransportId { self.transport_id } diff --git a/src/transport/ethernet/mod.rs b/src/transport/ethernet/mod.rs index ba81dc64..3a94ad52 100644 --- a/src/transport/ethernet/mod.rs +++ b/src/transport/ethernet/mod.rs @@ -553,6 +553,10 @@ impl Drop for EthernetTransport { } impl Transport for EthernetTransport { + fn role(&self) -> crate::config::TransportRole { + self.config.role() + } + fn transport_id(&self) -> TransportId { self.transport_id } @@ -1637,6 +1641,7 @@ mod tests { // A name no host has. `fips` is not a valid netdev prefix anywhere and // the suffix keeps it clear of the test harness's own veth pairs. let config = EthernetConfig { + role: None, interface: "fips-absent-x0".to_string(), ethertype: None, mtu: None, @@ -1821,6 +1826,7 @@ mod tests { let can_open = PacketSocket::open(loopback, 0x2121).is_ok(); let config = EthernetConfig { + role: None, interface: loopback.to_string(), ethertype: None, mtu: None, @@ -2401,6 +2407,7 @@ mod tests { } let config = EthernetConfig { + role: None, interface: iface.to_string(), ethertype: None, mtu: None, diff --git a/src/transport/loopback.rs b/src/transport/loopback.rs index d0f2d335..6cf21fb0 100644 --- a/src/transport/loopback.rs +++ b/src/transport/loopback.rs @@ -50,6 +50,13 @@ pub struct LoopbackTransport { mtu: u16, /// Shared address-to-receiver registry. registry: LoopbackRegistry, + /// Beacons a test has queued for the next `discover()` drain, standing + /// in for the transport-neighbor beacons a real transport hears. + discovered: Mutex>, + /// Carrier a test has set, standing in for an interface-bound + /// transport's `IFF_RUNNING`. `None`: not interface-bound, no presence + /// reported, which is the default. + carrier: Mutex>, } impl LoopbackTransport { @@ -75,9 +82,30 @@ impl LoopbackTransport { my_addr, mtu, registry, + discovered: Mutex::new(Vec::new()), + carrier: Mutex::new(None), } } + /// Pretend this transport is bound to an interface with (or without) + /// carrier; `None` returns it to reporting no presence at all. + pub fn set_carrier(&self, carrier: Option) { + *self.carrier.lock().unwrap() = carrier; + } + + /// The carrier a test set, if any. + pub fn carrier(&self) -> Option { + *self.carrier.lock().unwrap() + } + + /// Queue a beacon for the next `discover()` drain, as if this transport + /// had heard `addr` announce `pubkey_hint`. + pub fn inject_discovered(&self, addr: TransportAddr, pubkey_hint: secp256k1::XOnlyPublicKey) { + let mut peer = DiscoveredPeer::new(self.transport_id, addr); + peer.pubkey_hint = Some(pubkey_hint); + self.discovered.lock().unwrap().push(peer); + } + /// This transport's synthetic loopback address. pub fn my_addr(&self) -> &TransportAddr { &self.my_addr @@ -169,6 +197,11 @@ impl Transport for LoopbackTransport { } fn discover(&self) -> Result, TransportError> { - Ok(Vec::new()) + Ok(std::mem::take(&mut *self.discovered.lock().unwrap())) + } + + /// Beacons a test injects are meant to be acted on. + fn auto_connect(&self) -> bool { + true } } diff --git a/src/transport/mod.rs b/src/transport/mod.rs index 81a25fee..cc11123e 100644 --- a/src/transport/mod.rs +++ b/src/transport/mod.rs @@ -268,6 +268,20 @@ impl TransportError { /// which is a statement about the peer rather than about this node's /// ability to transmit, and the existing retry paths for them already sit /// at a different layer. + /// Whether the kernel refused the send for want of a route: the + /// interface is up but nothing is reachable through it. A hard signal + /// that the path is gone (`ENETUNREACH`, `EHOSTUNREACH`), distinct from + /// `is_transient`: the binder is not going to fix this. + pub fn is_unreachable(&self) -> bool { + match self { + Self::Io(e) => matches!( + e.kind(), + std::io::ErrorKind::NetworkUnreachable | std::io::ErrorKind::HostUnreachable + ), + _ => false, + } + } + pub fn is_transient(&self) -> bool { match self { // The interface is absent or mid-rebind. The binder is polling for @@ -565,6 +579,15 @@ impl Link { link } + /// Point the link at another transport and address. + /// + /// A peer whose active path switched keeps its link (the control machine + /// is keyed on it); the record follows the traffic. + pub fn rebind(&mut self, transport_id: TransportId, remote_addr: TransportAddr) { + self.transport_id = transport_id; + self.remote_addr = remote_addr; + } + /// Get the link ID. pub fn link_id(&self) -> LinkId { self.link_id @@ -748,6 +771,12 @@ pub trait Transport { true } + /// The transport's path-selection role. Default: normal. Concrete + /// transports read from their own config. + fn role(&self) -> crate::config::TransportRole { + crate::config::TransportRole::Normal + } + /// Close a specific connection (connection-oriented transports only). /// /// For connectionless transports (UDP, Ethernet), this is a no-op. @@ -1026,6 +1055,15 @@ impl TransportHandle { failed_attempts: state.attempts(), }) } + #[cfg(test)] + TransportHandle::Loopback(t) => t.carrier().map(|carrier| InterfacePresence { + presence: "present", + carrier, + policy: "optional", + since_secs: 0, + binds: 1, + failed_attempts: 0, + }), _ => None, } } @@ -1107,6 +1145,22 @@ impl TransportHandle { } } + /// The transport's path-selection role. + pub fn role(&self) -> crate::config::TransportRole { + match self { + TransportHandle::Udp(t) => t.role(), + #[cfg(any(target_os = "linux", target_os = "macos"))] + TransportHandle::Ethernet(t) => t.role(), + TransportHandle::Tcp(t) => t.role(), + TransportHandle::Tor(t) => t.role(), + TransportHandle::Nym(t) => t.role(), + #[cfg(ble_available)] + TransportHandle::Ble(t) => t.role(), + #[cfg(test)] + TransportHandle::Loopback(t) => t.role(), + } + } + /// Whether this transport auto-connects to discovered peers. pub fn auto_connect(&self) -> bool { match self { diff --git a/src/transport/nym/mod.rs b/src/transport/nym/mod.rs index f1a5dea0..e9eb9630 100644 --- a/src/transport/nym/mod.rs +++ b/src/transport/nym/mod.rs @@ -611,6 +611,10 @@ impl NymTransport { } impl Transport for NymTransport { + fn role(&self) -> crate::config::TransportRole { + self.config.role() + } + fn transport_id(&self) -> TransportId { self.transport_id } diff --git a/src/transport/tcp/mod.rs b/src/transport/tcp/mod.rs index d4a8f54c..6facfe82 100644 --- a/src/transport/tcp/mod.rs +++ b/src/transport/tcp/mod.rs @@ -786,6 +786,10 @@ impl TcpTransport { } impl Transport for TcpTransport { + fn role(&self) -> crate::config::TransportRole { + self.config.role() + } + fn transport_id(&self) -> TransportId { self.transport_id } diff --git a/src/transport/tor/mod.rs b/src/transport/tor/mod.rs index ed93dcb6..088df86e 100644 --- a/src/transport/tor/mod.rs +++ b/src/transport/tor/mod.rs @@ -1047,6 +1047,10 @@ impl TorTransport { } impl Transport for TorTransport { + fn role(&self) -> crate::config::TransportRole { + self.config.role() + } + fn transport_id(&self) -> TransportId { self.transport_id } diff --git a/src/transport/udp/mod.rs b/src/transport/udp/mod.rs index 9c708c05..8e2c4b31 100644 --- a/src/transport/udp/mod.rs +++ b/src/transport/udp/mod.rs @@ -437,6 +437,10 @@ impl UdpTransport { } impl Transport for UdpTransport { + fn role(&self) -> crate::config::TransportRole { + self.config.role() + } + fn transport_id(&self) -> TransportId { self.transport_id } diff --git a/testing/chaos/README.md b/testing/chaos/README.md index b86d08e2..38c3f1f1 100644 --- a/testing/chaos/README.md +++ b/testing/chaos/README.md @@ -101,11 +101,25 @@ Explicit topologies exercising non-UDP transports. | ethernet-only | 4 | Ethernet | Ring | 30s | yes | -- | AF_PACKET transport with beacon discovery | | ethernet-mesh | 6 | UDP + Ethernet | Mesh | 120s | yes | yes | Mixed UDP/Ethernet, netem mutation + flaps | | tcp-mesh | 6 | UDP + TCP | Mesh | 120s | yes | yes | Mixed UDP/TCP, netem mutation + flaps | +| dual-path-flap | 2 | Ethernet + UDP | Pair | 180s | yes | yes | One session, two paths; cable flaps, traffic moves without re-peering | +| dual-udp-flap | 2 | UDP + UDP | Pair | 180s | yes | yes | Same, all-IP: two interface-bound UDP instances as two paths | - **ethernet-only**: 4-node ring on raw Ethernet (AF_PACKET). Peers discovered via beacons, not static config. Minimal netem (1-5ms delay). - **ethernet-mesh**: Mirrors `tcp-mesh` topology but with Ethernet instead of TCP. UDP edges use static config; Ethernet edges use beacon discovery. +- **dual-path-flap**: two nodes joined twice, by a raw-Ethernet veth and by + UDP over the bridge (`[n01, n02, ethernet+udp]`). The cable flaps + (`link_flaps.only_transport: ethernet`); traffic must move to the UDP path + under the same session and never re-peer. The calibration scenario for + `node.path.*`; see the file header for what to read from a run. +- **dual-udp-flap**: the all-IP twin (`[n01, n02, udp-veth+udp]`): a veth + carrying IP with an interface-bound UDP instance at each end, plus UDP + over the bridge. The veth half is handed to the daemon by the runner + (`fipsctl connect ... udp/`) once the pair has peered, and becomes a + path under the existing session. Exercises `udp.interface` on the listen + and per-peer connected sockets, and a configured address on a new + transport becoming a path. - **tcp-mesh**: 6-node mesh with 4 UDP and 3 TCP edges. Both transports use static peer config. Netem mutation (30% fraction, every 20-40s) and link flaps (1 link max, 10-20s down). diff --git a/testing/chaos/scenarios/dual-path-flap.yaml b/testing/chaos/scenarios/dual-path-flap.yaml new file mode 100644 index 00000000..25eaaa51 --- /dev/null +++ b/testing/chaos/scenarios/dual-path-flap.yaml @@ -0,0 +1,123 @@ +# One peer, two paths: switchover without a second handshake +# +# The calibration scenario for `node.path.*` +# (docs/design/fips-multi-path-switchover.md, "Calibration"). Two nodes joined +# twice: a raw-Ethernet veth (the cable) and the Docker bridge over UDP (the +# wifi). The UDP half is dialled from static config, the Ethernet half is +# found by beacon and added as a second path under the same session by the +# probe exchange. iperf runs across the pair throughout while the cable is +# taken down and brought back. +# +# What "down" means here: `link_flaps` sets 100% loss on the veth, so the +# interface stays IFF_UP and keeps carrier. That is deliberate. Presence and +# carrier are covered by unit tests; the thing only a network can exercise +# is the heartbeat echo timeout — the detector every medium falls back to — +# and the discretionary fail-back when the cable returns. +# +# What to read from a run (sim-results//analysis.txt and node logs): +# - "Path switched" / "traffic moved to the standby": one per node per +# flap edge, so two flaps give at least four; the ceiling below allows +# one extra per flap for a discretionary fail-back and nothing more. +# - the timestamps between "Link DOWN" in runner.log and the first switch +# on each node: the switch latency. The target is under one second. +# - "Peer promoted to active" more than once on a node: a re-peering, +# which is the failure this scenario exists to catch. `max_promotions` +# fails the run on it; `switch_latency` and `max_stall` fail it on a +# slow switch and on a switch that did not carry traffic. +# +# n01 ===eth=== n02 (flapped) +# n01 ---udp--- n02 (never touched) + +scenario: + name: "dual-path-flap" + seed: 42 + duration_secs: 180 + +topology: + algorithm: explicit + num_nodes: 2 + default_transport: udp + params: + adjacency: + - [n01, n02, ethernet+udp] + +# The cable is mild; the UDP half carries 40 ms each way. Score is +# etx × (1 + min_rtt/100): about 1.0 on the cable against about 1.8 over +# UDP, well past the 1.3 margin, so the pair moves onto the cable soon after +# both paths are live (the UDP static dial lands first) and returns to it +# after each flap once the cable is proven again. Every flap therefore hits +# the *active* path, and the fail-back is exercised too. +netem: + enabled: true + default_policy: + delay_ms: [1, 3] + jitter_ms: [0, 1] + loss_pct: [0, 0.2] + dual_udp_policy: + delay_ms: [40, 40] + jitter_ms: [0, 1] + loss_pct: [0, 0.2] + +# The mechanism. Only the cable flaps; the UDP path is the standby that +# has to be warm when it does. `protect_connectivity` is off because the +# harness counts the dual edge once and would refuse to flap the "only" +# link; the UDP half keeps the pair connected throughout. +link_flaps: + enabled: true + only_transport: ethernet + interval_secs: {min: 35, max: 45} + max_down_links: 1 + down_duration_secs: {min: 12, max: 18} + protect_connectivity: false + +traffic: + enabled: true + max_concurrent: 1 + interval_secs: {min: 5, max: 10} + duration_secs: {min: 20, max: 30} + parallel_streams: 2 + +node_churn: + enabled: false + +assertions: + baseline: + min_nodes_reporting: 2 + max_roots: 1 + min_nodes_parented: 1 + + # Traffic crossed the pair while the cable was being taken away. + min_traffic: + min_sessions_ok: 2 + + # Three to four flaps in 180 s at the interval above. Each flap moves both + # nodes off the cable (2) and its return moves them back (2): four per + # flap, plus the initial move onto the cable (2). Floor: the first flap + # moved something. Ceiling: four flaps, both directions, both nodes, the + # initial move, and nothing else. + path_switches: + min_total: 4 + max_total: 18 + + # The detectors that can fail on a switchover that did not carry. A + # promotion past the first is a re-peering: the session was lost to the + # link-dead reaper and rebuilt, which standby probes alone cannot show. + # The switch must land inside a second of the runner taking the link + # down. And no iperf3 session may sit at zero bytes for longer than two + # of its one-second intervals: min_traffic passes on bytes either side + # of a hole, this fails on the hole. + max_promotions: + per_node: 1 + switch_latency: + max_ms: 1000 + max_stall: + max_secs: 2 + + max_errors: + max_total: 0 + +logging: + # The path handler at debug: "Path suspect" and "Sent path probe" are what + # the switch latency is read from. + rust_log: "info,fips::node::handlers::path=debug" + output_dir: "./sim-results" diff --git a/testing/chaos/scenarios/dual-udp-flap.yaml b/testing/chaos/scenarios/dual-udp-flap.yaml new file mode 100644 index 00000000..963312b8 --- /dev/null +++ b/testing/chaos/scenarios/dual-udp-flap.yaml @@ -0,0 +1,121 @@ +# One peer, two UDP paths: switchover between two interface-bound UDP instances +# +# The all-IP twin of dual-path-flap. Two nodes joined twice, both by UDP: a +# dedicated veth carrying IP (`10.222.0.0/24`), on which each node runs a +# UDP instance bound to that interface (`transports.udp..interface`), +# and the Docker bridge on the ordinary wildcard instance. Both are static +# peer addresses on the dial owner (`udp/main` first, `udp/` second), +# so both handshakes run at startup; the first to complete makes the +# session, the second settles as a cross-connection and creates no path +# state — the dialled address is a candidate, and the probe exchange under +# the session proves it as a path at both ends. +# +# What this exercises that dual-path-flap cannot: two UDP instances as two +# paths, `udp.interface` binding on the listen socket *and* on the per-peer +# connected sockets (Linux `SO_BINDTODEVICE`), and a configured address on a +# new transport becoming a path. The flap and fail-back mechanics are the +# same: `link_flaps` sets 100% loss on the veth, the interface stays up, and +# the heartbeat echo timeout is the detector. +# +# What to read from a run: as dual-path-flap. "Path live: the peer answers +# on this transport" in both node logs is the second instance being proven +# as a path; `fipsctl path show` on either node should list both instances +# live before the first flap, and "Peer promoted to active" appears once +# per node (`max_promotions` fails the run otherwise). +# +# n01 ===udp/veth=== n02 (flapped) +# n01 ----udp/main--- n02 (never touched) + +scenario: + name: "dual-udp-flap" + seed: 42 + duration_secs: 180 + +topology: + algorithm: explicit + num_nodes: 2 + default_transport: udp + params: + adjacency: + - [n01, n02, udp-veth+udp] + +# The veth is mild; the bridge half carries 40 ms each way. Score is +# etx × (1 + min_rtt/100): about 1.0 on the veth against about 1.8 over the +# bridge, well past the 1.3 margin, so the pair moves onto the veth soon +# after the runner adds it and returns to it after each flap once it is +# proven again. Every flap therefore hits the *active* path, and the +# fail-back is exercised too. +netem: + enabled: true + default_policy: + delay_ms: [1, 3] + jitter_ms: [0, 1] + loss_pct: [0, 0.2] + dual_udp_policy: + delay_ms: [40, 40] + jitter_ms: [0, 1] + loss_pct: [0, 0.2] + +# The mechanism. Only the veth flaps; the bridge path is the standby that +# has to be warm when it does. `protect_connectivity` is off because the +# harness counts the dual edge once and would refuse to flap the "only" +# link; the bridge half keeps the pair connected throughout. +link_flaps: + enabled: true + only_transport: udp-veth + interval_secs: {min: 35, max: 45} + max_down_links: 1 + down_duration_secs: {min: 12, max: 18} + protect_connectivity: false + +traffic: + enabled: true + max_concurrent: 1 + interval_secs: {min: 5, max: 10} + duration_secs: {min: 20, max: 30} + parallel_streams: 2 + +node_churn: + enabled: false + +assertions: + baseline: + min_nodes_reporting: 2 + max_roots: 1 + min_nodes_parented: 1 + + # Traffic crossed the pair while the cable was being taken away. + min_traffic: + min_sessions_ok: 2 + + # Three to four flaps in 180 s at the interval above. Each flap moves both + # nodes off the veth (2) and its return moves them back (2): four per + # flap, plus the initial move onto the veth (2). Floor: the first flap + # moved something. Ceiling: four flaps, both directions, both nodes, the + # initial move, and nothing else. + path_switches: + min_total: 4 + max_total: 18 + + # The detectors that can fail on a switchover that did not carry. A + # promotion past the first is a re-peering: the session was lost to the + # link-dead reaper and rebuilt, which standby probes alone cannot show. + # The switch must land inside a second of the runner taking the link + # down. And no iperf3 session may sit at zero bytes for longer than two + # of its one-second intervals: min_traffic passes on bytes either side + # of a hole, this fails on the hole. + max_promotions: + per_node: 1 + switch_latency: + max_ms: 1000 + max_stall: + max_secs: 2 + + max_errors: + max_total: 0 + +logging: + # The path handler at debug: "Path suspect" and "Sent path probe" are what + # the switch latency is read from. + rust_log: "info,fips::node::handlers::path=debug,fips::node::lifecycle=debug" + output_dir: "./sim-results" diff --git a/testing/chaos/sim/assertions.py b/testing/chaos/sim/assertions.py index 4652198f..f7e9591b 100644 --- a/testing/chaos/sim/assertions.py +++ b/testing/chaos/sim/assertions.py @@ -25,7 +25,10 @@ from .scenario import ( CongestionSignalsAssertion, MaxErrorsAssertion, MaxParentSwitchesAssertion, + MaxPromotionsAssertion, + MaxStallAssertion, MinParentSwitchesAssertion, + SwitchLatencyAssertion, TreeParentsAssertion, ) from .topology import SimTopology @@ -501,3 +504,240 @@ def evaluate_min_traffic( f"means the data path did not survive what the scenario did to it." ), ) + + +def evaluate_path_switches(cfg, count: int) -> AssertionOutcome: + """Band on path switches (traffic moving between transports under one + session) over the run. + + ``min_total`` catches a harness that flapped a link nothing was + switching over: a dual-path scenario in which no switch happened + tested nothing. ``max_total`` is the stability ceiling: a healthy + dual-path pair should switch only when a link goes and comes back, + never on its own. + """ + if cfg.min_total is not None and count < cfg.min_total: + return AssertionOutcome( + name="path_switches", + passed=False, + detail=( + f"FAIL path_switches: {count} switches < floor {cfg.min_total} " + f"— the flaps did not move traffic between paths. Check that " + f"both paths came up (fipsctl path show) before the first flap." + ), + ) + if cfg.max_total is not None and count > cfg.max_total: + return AssertionOutcome( + name="path_switches", + passed=False, + detail=( + f"FAIL path_switches: {count} switches > ceiling {cfg.max_total} " + f"— traffic is moving between paths more than the flaps " + f"account for. Look for discretionary switches on a healthy " + f"pair: the margin or the dwell is too small." + ), + ) + return AssertionOutcome( + name="path_switches", + passed=True, + detail=( + f"PASS path_switches: {count} switches within " + f"[{cfg.min_total if cfg.min_total is not None else 0}, " + f"{cfg.max_total if cfg.max_total is not None else 'inf'}]" + ), + ) + + +def evaluate_max_promotions( + cfg: MaxPromotionsAssertion, + promotions: list[tuple[str, str]], +) -> AssertionOutcome: + """Per-node ceiling on "Peer promoted to active" lines. + + ``promotions`` is ``AnalysisResult.peers_promoted``: ``(source, line)`` + pairs, one per handshake that completed on that node. The first per + peer is the pair meeting; any beyond the ceiling is a re-peering — the + session was torn down and rebuilt — which a switchover scenario exists + to prove does not happen. Liveness is no backstop here: standby probes + keep a peer alive while its data path is blackholed, so this counts + the one event a blackhole that lasts to the reaper cannot avoid. + """ + per_node: dict[str, int] = {} + for source, _line in promotions: + per_node[source] = per_node.get(source, 0) + 1 + over = {src: n for src, n in per_node.items() if n > cfg.per_node} + if not over: + return AssertionOutcome( + name="max_promotions", + passed=True, + detail=( + f"PASS max_promotions: every node promoted at most " + f"{cfg.per_node} time(s) ({len(promotions)} total)" + ), + ) + breakdown = ", ".join(f"{src}={n}" for src, n in sorted(over.items())) + samples = "\n".join( + f" [{src}] {line.strip()}" + for src, line in promotions + if src in over + ) + return AssertionOutcome( + name="max_promotions", + passed=False, + detail=( + f"FAIL max_promotions: {breakdown} exceed(s) the per-node ceiling " + f"of {cfg.per_node}. A second promotion is a re-peering: the " + f"session was lost and rebuilt, so a switchover did not carry.\n" + f"{samples}" + ), + ) + + +def _line_epoch(line: str) -> float | None: + """Epoch seconds of a node log line's leading RFC 3339 timestamp, if any. + + ``tracing`` writes ``2026-09-13T13:39:01.123456Z`` first on every line. + Anything else (a bare stderr line, a runner line) is not timed. + """ + from datetime import datetime, timezone + + head = line.strip().split(" ", 1)[0] + if not head.endswith("Z"): + return None + try: + return datetime.fromisoformat(head.replace("Z", "+00:00")).timestamp() + except ValueError: + return None + + +def evaluate_switch_latency( + cfg: SwitchLatencyAssertion, + flap_events: list[tuple[float, str, str, str]], + switches: list[tuple[str, str]], +) -> AssertionOutcome: + """Ceiling on the time from each link-down to the first switch on + either endpoint. + + ``flap_events`` is the link manager's record: ``(epoch, "down" | "up", + a, b)``. ``switches`` is ``AnalysisResult.path_switches``: ``(source, + line)``, where ``source`` is the node id and the line carries its own + timestamp. For each down edge, the latency is the earliest switch line + on ``a`` or ``b`` stamped at or after the down; a down with no switch + inside ``max_ms`` fails. The worst flap is what is reported. + """ + downs = [(t, a, b) for t, kind, a, b in flap_events if kind == "down"] + if not downs: + return AssertionOutcome( + name="switch_latency", + passed=False, + detail="FAIL switch_latency: no link was taken down, nothing measured", + ) + timed: list[tuple[str, float]] = [] + for source, line in switches: + t = _line_epoch(line) + if t is not None: + timed.append((source, t)) + limit = cfg.max_ms / 1000.0 + worst: tuple[float, str, str] | None = None + missing: list[str] = [] + for down_at, a, b in downs: + after = [ + t - down_at + for src, t in timed + if src in (a, b) and t >= down_at and t - down_at <= limit + ] + if not after: + from datetime import datetime, timezone + + when = datetime.fromtimestamp(down_at, timezone.utc).strftime("%H:%M:%S") + missing.append(f"{a}--{b} down at {when}Z") + continue + latency = min(after) + if worst is None or latency > worst[0]: + worst = (latency, a, b) + if missing: + return AssertionOutcome( + name="switch_latency", + passed=False, + detail=( + f"FAIL switch_latency: {len(missing)} of {len(downs)} link-down(s) " + f"had no path switch on either endpoint within {cfg.max_ms} ms: " + + "; ".join(missing) + ), + ) + assert worst is not None + return AssertionOutcome( + name="switch_latency", + passed=True, + detail=( + f"PASS switch_latency: worst {worst[0] * 1000:.0f} ms " + f"({worst[1]}--{worst[2]}) over {len(downs)} link-down(s), " + f"ceiling {cfg.max_ms} ms" + ), + ) + + +def _longest_stall_secs(result: dict) -> float: + """Longest run of consecutive zero-byte intervals in one iperf3 result, + in seconds. 0 for a result with no intervals.""" + intervals = result.get("intervals") if isinstance(result, dict) else None + if not isinstance(intervals, list): + return 0.0 + longest = 0.0 + run = 0.0 + for iv in intervals: + summary = iv.get("sum") if isinstance(iv, dict) else None + if not isinstance(summary, dict): + continue + seconds = summary.get("seconds", 1.0) + if not isinstance(seconds, (int, float)) or seconds <= 0: + seconds = 1.0 + if summary.get("bytes", 0) == 0: + run += seconds + longest = max(longest, run) + else: + run = 0.0 + return longest + + +def evaluate_max_stall( + cfg: MaxStallAssertion, + results: list[dict], +) -> AssertionOutcome: + """Ceiling on the longest zero-byte run inside any iperf3 session. + + ``min_traffic`` cannot see a hole: a session that stalls for ten + seconds mid-run still moves bytes before and after. This reads the + per-interval totals iperf3 records and fails on the longest run of + zeros across every session, which is the stall a switchover leaves + when it does not carry. + """ + if not results: + return AssertionOutcome( + name="max_stall", + passed=False, + detail="FAIL max_stall: no iperf3 session ran, nothing measured", + ) + stalls = [(_longest_stall_secs(r), r) for r in results] + worst_secs, worst = max(stalls, key=lambda pair: pair[0]) + if worst_secs <= cfg.max_secs: + return AssertionOutcome( + name="max_stall", + passed=True, + detail=( + f"PASS max_stall: longest zero-byte run {worst_secs:.0f} s " + f"across {len(results)} session(s), ceiling {cfg.max_secs:g} s" + ), + ) + start = worst.get("start", {}) if isinstance(worst, dict) else {} + when = start.get("timestamp", {}).get("time", "?") if isinstance(start, dict) else "?" + return AssertionOutcome( + name="max_stall", + passed=False, + detail=( + f"FAIL max_stall: a session starting {when} moved nothing for " + f"{worst_secs:.0f} s, over the ceiling of {cfg.max_secs:g} s. Bytes " + f"either side of the hole satisfied min_traffic; the hole is a " + f"switchover that did not carry." + ), + ) diff --git a/testing/chaos/sim/config_gen.py b/testing/chaos/sim/config_gen.py index 39463456..640ac8ef 100644 --- a/testing/chaos/sim/config_gen.py +++ b/testing/chaos/sim/config_gen.py @@ -7,7 +7,7 @@ from copy import deepcopy import yaml -from .topology import SimTopology +from .topology import UDP_VETH, SimTopology def _deep_merge(base: dict, override: dict) -> dict: @@ -33,6 +33,7 @@ def _load_template() -> str: _TRANSPORT_PORTS = { "udp": 2121, + "udp-veth": 2122, "tcp": 443, } @@ -49,16 +50,32 @@ def generate_peers_block( if not outbound_peers: return " []" + # With interface-bound UDP instances the UDP transport is named, and a + # bare ``udp`` address would resolve to the lowest instance id, which + # may be the one bound to a veth: qualify the bridge half. + bridge = "udp/main" if topology.udp_veth_links(node_id) else "udp" lines = [] for peer_id in sorted(outbound_peers): peer = topology.nodes[peer_id] transport = topology.transport_for_edge(node_id, peer_id) - port = _TRANSPORT_PORTS.get(transport, 2121) + if topology.is_dual_udp_edge(node_id, peer_id): + # The veth half of a dual edge is found by beacon (Ethernet) or + # added as a path by the runner after the pair has peered + # (udp-veth); the bridge half is dialled from here. + transport = "udp" + if transport == UDP_VETH: + link = next(l for l in topology.udp_veth_links(node_id) if l.peer_id == peer_id) + transport = f"udp/{link.instance}" + addr = link.peer_addr + else: + addr = f"{peer.docker_ip}:{_TRANSPORT_PORTS.get(transport, 2121)}" + if transport == "udp": + transport = bridge lines.append(f' - npub: "{peer.npub}"') lines.append(f' alias: "{peer_id}"') lines.append(f" addresses:") lines.append(f" - transport: {transport}") - lines.append(f' addr: "{peer.docker_ip}:{port}"') + lines.append(f' addr: "{addr}"') lines.append(f" connect_policy: auto_connect") return "\n".join(lines) @@ -110,6 +127,28 @@ def _inject_ethernet_transports(parsed: dict, eth_ifaces: list[str]): } +def _inject_udp_instances(parsed: dict, topology: SimTopology, node_id: str, has_udp: bool): + """Turn the template's single UDP transport into named instances: ``main`` + (the bridge, kept only if the node has bridge-UDP peers) plus one + interface-bound instance per ``udp-veth`` edge, named after its veth. + """ + links = topology.udp_veth_links(node_id) + if not links: + return + transports = parsed.setdefault("transports", {}) + main = transports.pop("udp", None) or {"bind_addr": "0.0.0.0:2121"} + instances = {} + if has_udp: + instances["main"] = main + for link in links: + instances[link.instance] = { + "bind_addr": f"0.0.0.0:{_TRANSPORT_PORTS['udp-veth']}", + "interface": link.iface, + "mtu": main.get("mtu", 1472), + } + transports["udp"] = instances + + def _inject_tcp_transport(parsed: dict): """Inject TCP transport config into a parsed FIPS config.""" transports = parsed.setdefault("transports", {}) @@ -152,9 +191,12 @@ def generate_node_config( eth_ifaces = topology.ethernet_interfaces(node_id) has_tcp = bool(topology.tcp_peers(node_id)) has_udp = _has_transport_peers(topology, node_id, "udp") + has_udp_veth = bool(topology.udp_veth_links(node_id)) # Inject non-UDP transport configs and handle pure-transport nodes - needs_yaml_rewrite = eth_ifaces or has_tcp or not has_udp or fips_overrides + needs_yaml_rewrite = ( + eth_ifaces or has_tcp or has_udp_veth or not has_udp or fips_overrides + ) if needs_yaml_rewrite: parsed = yaml.safe_load(config) @@ -164,7 +206,9 @@ def generate_node_config( _inject_ethernet_transports(parsed, eth_ifaces) if has_tcp: _inject_tcp_transport(parsed) - if not has_udp: + if has_udp_veth: + _inject_udp_instances(parsed, topology, node_id, has_udp) + elif not has_udp: # No UDP edges: remove UDP transport transports = parsed.get("transports", {}) transports.pop("udp", None) @@ -179,6 +223,8 @@ def _has_transport_peers(topology: SimTopology, node_id: str, transport: str) -> edge = (min(node_id, peer_id), max(node_id, peer_id)) if topology.edge_transport.get(edge, "udp") == transport: return True + if transport == "udp" and edge in topology.dual_udp_edges: + return True return False diff --git a/testing/chaos/sim/links.py b/testing/chaos/sim/links.py index e6967957..129f434a 100644 --- a/testing/chaos/sim/links.py +++ b/testing/chaos/sim/links.py @@ -52,6 +52,11 @@ class LinkManager: self.link_states: dict[tuple[str, str], LinkState] = { edge: LinkState(edge=edge) for edge in topology.edges } + # Every edge taken down or restored, as (epoch seconds, "down" | + # "up", a, b). The switch-latency assertion reads the down edges + # against the nodes' own log timestamps, so this is wall-clock time + # (the containers share the host clock). + self.flap_events: list[tuple[float, str, str, str]] = [] @property def down_count(self) -> int: @@ -68,6 +73,10 @@ class LinkManager: up_links = [ e for e, ls in self.link_states.items() if not ls.is_down and e[0] not in down and e[1] not in down + and ( + self.config.only_transport is None + or self.topology.transport_for_edge(*e) == self.config.only_transport + ) ] if not up_links: return @@ -115,6 +124,7 @@ class LinkManager: state.is_down = True state.down_since = now state.restore_at = now + duration + self.flap_events.append((now, "down", a, b)) log.info("Link DOWN: %s -- %s (restore in %.0fs)", a, b, duration) @@ -130,6 +140,7 @@ class LinkManager: self._set_held(b, a, False) down_for = time.time() - state.down_since if state.down_since else 0 + self.flap_events.append((time.time(), "up", a, b)) state.is_down = False state.down_since = None state.restore_at = None @@ -155,7 +166,7 @@ class LinkManager: container = self.topology.container_name(src_node) transport = self.topology.transport_for_edge(src_node, dst_node) - if transport == "ethernet": + if self.topology.is_veth_transport(transport): iface = veth_interface_name(src_node, dst_node) state = self.netem_mgr.veth_states.get(container, {}).get(iface) if state is None: diff --git a/testing/chaos/sim/netem.py b/testing/chaos/sim/netem.py index 8085e2e0..252bdb59 100644 --- a/testing/chaos/sim/netem.py +++ b/testing/chaos/sim/netem.py @@ -210,8 +210,12 @@ class NetemManager: eth_peers = [] for peer_id in sorted(node.peers): transport = self.topology.transport_for_edge(node_id, peer_id) - if transport == "ethernet": + if self.topology.is_veth_transport(transport): eth_peers.append(peer_id) + # A dual edge also has a UDP half over the bridge, which + # gets its own HTB class like any IP peer. + if self.topology.is_dual_udp_edge(node_id, peer_id): + ip_peers[peer_id] = self.topology.nodes[peer_id].docker_ip else: ip_peers[peer_id] = self.topology.nodes[peer_id].docker_ip @@ -232,6 +236,11 @@ class NetemManager: netem_handle = f"{idx + 10}:" policy = self._policy_for_edge(node_id, peer_id) + if ( + self.topology.is_dual_udp_edge(node_id, peer_id) + and self.config.dual_udp_policy is not None + ): + policy = self.config.dual_udp_policy params = self._sample_policy(policy) rate = self._htb_rate(node_id, peer_id) @@ -399,7 +408,9 @@ class NetemManager: for peer_id in sorted(self.topology.nodes[node_id].peers): if peer_id in self.down_nodes: continue - if self.topology.transport_for_edge(node_id, peer_id) != "ethernet": + if not self.topology.is_veth_transport( + self.topology.transport_for_edge(node_id, peer_id) + ): continue peer_container = self.topology.container_name(peer_id) state = self.veth_states.get(peer_container, {}).get( @@ -477,8 +488,9 @@ class NetemManager: self.down_nodes.add(src) continue - if transport == "ethernet": - # Ethernet: simple netem replace on veth + if self.topology.is_veth_transport(transport): + # Ethernet, or a UDP instance bound to a veth: simple netem + # replace on the veth iface = veth_interface_name(src, dst) veth_states = self.veth_states.get(container, {}) state = veth_states.get(iface) diff --git a/testing/chaos/sim/runner.py b/testing/chaos/sim/runner.py index 1d9df9de..4e3bcf93 100644 --- a/testing/chaos/sim/runner.py +++ b/testing/chaos/sim/runner.py @@ -20,6 +20,10 @@ from .assertions import ( evaluate_max_errors, evaluate_max_parent_switches, evaluate_min_parent_switches, + evaluate_path_switches, + evaluate_max_promotions, + evaluate_switch_latency, + evaluate_max_stall, evaluate_min_traffic, evaluate_tree_parents, ) @@ -318,9 +322,9 @@ class SimRunner: # The entrypoint script waits for configured Ethernet interfaces # to appear before starting FIPS, so we just need to create the # veth pairs promptly after containers are running. - if self.topology.has_ethernet(): + if self.topology.has_veth(): self.veth_mgr = VethManager(self.topology) - log.info("Setting up Ethernet veth pairs...") + log.info("Setting up veth pairs...") self.veth_mgr.setup_all() # 7. Initialize managers @@ -385,6 +389,11 @@ class SimRunner: self._sleep(wait) self._take_snapshot("warmup") + # The veth half of every udp-veth+udp edge: UDP has no beacon, so + # the runner hands the daemon the address once the pair has peered + # over the bridge, and it becomes a path under that session. + self._add_udp_veth_paths() + # Populate npub cache after convergence (nodes must be running) if self.peer_churn_mgr: self.peer_churn_mgr.refresh_all_npubs() @@ -395,13 +404,67 @@ class SimRunner: if self.link_swap_mgr: self.link_swap_mgr.setup_initial() + def _add_udp_veth_paths(self, only_node: str | None = None): + """Give each dual udp-veth edge its veth path. + + Sent from the edge's dial owner (the side whose static config holds + the bridge address) as a control-socket ``connect`` naming the + interface-bound instance: to a peer it already holds a session with, + the daemon adds that as a path rather than dialling. Waits for the + bridge session first, so the command cannot become the first dial. + """ + from .control import send_command + + outbound = self.topology.directed_outbound() + for node_id in sorted(self.topology.nodes): + if only_node is not None and node_id != only_node: + continue + for link in self.topology.udp_veth_links(node_id): + if not self.topology.is_dual_udp_edge(node_id, link.peer_id): + continue + if link.peer_id not in outbound.get(node_id, []): + continue + if node_id in self._down_nodes or link.peer_id in self._down_nodes: + continue + container = self.topology.container_name(node_id) + npub = self.topology.nodes[link.peer_id].npub + params = { + "npub": npub, + "address": link.peer_addr, + "transport": f"udp/{link.instance}", + } + added = None + for _ in range(30): + if send_command(container, "path_show", {"npub": npub}) is None: + time.sleep(1) # not peered over the bridge yet + continue + added = send_command(container, "connect", params) + break + if added is None: + log.warning( + "udp-veth path %s -> %s via %s not added", + node_id, link.peer_id, link.instance, + ) + else: + log.info( + "udp-veth path %s -> %s via %s (%s)", + node_id, link.peer_id, link.instance, link.peer_addr, + ) + def _handle_node_restart(self, node_id: str): """Called after a node container is restarted. - For ephemeral identity nodes, waits briefly for the daemon to - start, then queries its new npub and updates the peer churn - manager's cache. + Re-adds the node's udp-veth paths once it has re-peered, and for + ephemeral identity nodes waits briefly for the daemon to start, + then queries its new npub and updates the peer churn manager's + cache. """ + if any( + self.topology.is_dual_udp_edge(node_id, link.peer_id) + for link in self.topology.udp_veth_links(node_id) + ): + time.sleep(2) + self._add_udp_veth_paths(only_node=node_id) if not self.peer_churn_mgr: return if node_id not in self.peer_churn_mgr.ephemeral_nodes: @@ -769,6 +832,45 @@ class SimRunner: else: log.error("%s", outcome.detail) + ps_cfg = self.scenario.assertions.path_switches + if ps_cfg is not None: + outcome = evaluate_path_switches(ps_cfg, len(result.path_switches)) + self.assertion_outcomes.append(outcome) + if outcome.passed: + log.info("%s", outcome.detail) + else: + log.error("%s", outcome.detail) + + mp_cfg = self.scenario.assertions.max_promotions + if mp_cfg is not None: + outcome = evaluate_max_promotions(mp_cfg, result.peers_promoted) + self.assertion_outcomes.append(outcome) + if outcome.passed: + log.info("%s", outcome.detail) + else: + log.error("%s", outcome.detail) + + sl_cfg = self.scenario.assertions.switch_latency + if sl_cfg is not None: + flap_events = self.link_mgr.flap_events if self.link_mgr else [] + outcome = evaluate_switch_latency( + sl_cfg, flap_events, result.path_switches + ) + self.assertion_outcomes.append(outcome) + if outcome.passed: + log.info("%s", outcome.detail) + else: + log.error("%s", outcome.detail) + + stall_cfg = self.scenario.assertions.max_stall + if stall_cfg is not None: + outcome = evaluate_max_stall(stall_cfg, iperf_results) + self.assertion_outcomes.append(outcome) + if outcome.passed: + log.info("%s", outcome.detail) + else: + log.error("%s", outcome.detail) + # Applied to every scenario by default, so this is the one # assertion that is present even when the YAML declares no # assertions block at all. diff --git a/testing/chaos/sim/scenario.py b/testing/chaos/sim/scenario.py index de6e5dee..dfad3ebf 100644 --- a/testing/chaos/sim/scenario.py +++ b/testing/chaos/sim/scenario.py @@ -28,7 +28,10 @@ class Range: raise ValueError(f"{name}: min ({self.min}) must be >= 0") -VALID_TRANSPORTS = ("udp", "ethernet", "tcp") +VALID_TRANSPORTS = ("udp", "ethernet", "tcp", "udp-veth") +# Dual edges: a veth half and a UDP-over-the-bridge half between the same +# two nodes (see topology.py). +DUAL_TRANSPORTS = ("ethernet+udp", "udp-veth+udp") @dataclass @@ -100,6 +103,10 @@ class NetemConfig: default_policy: NetemPolicy = field(default_factory=NetemPolicy) link_policies: list[LinkPolicyOverride] = field(default_factory=list) mutation: NetemMutationConfig = field(default_factory=NetemMutationConfig) + # Policy for the UDP half of ``ethernet+udp`` dual edges. The Ethernet + # half takes the edge's ordinary policy. ``None`` means the UDP half + # gets the default policy too. + dual_udp_policy: NetemPolicy | None = None @dataclass @@ -109,6 +116,9 @@ class LinkFlapsConfig: max_down_links: int = 2 down_duration_secs: Range = field(default_factory=lambda: Range(10, 30)) protect_connectivity: bool = True + # Only flap edges of this transport type (``ethernet``, ``udp``, ...). + # ``None`` flaps any edge. + only_transport: str | None = None @dataclass @@ -226,6 +236,62 @@ class MinParentSwitchesAssertion: min_total: int = 1 +@dataclass +class PathSwitchesAssertion: + """Band on path switches: a peer's traffic moving to another + transport under the same session. ``min_total`` proves the flaps moved + something; ``max_total`` is the stability ceiling on a pair that should + switch only when a link goes and comes back. + """ + + min_total: int | None = None + max_total: int | None = None + + +@dataclass +class MaxPromotionsAssertion: + """Per-node ceiling on "Peer promoted to active". + + A promotion is a handshake completing: the first one per peer is how a + pair meets, every later one is a re-peering — the session was lost and + rebuilt, which is the failure a switchover scenario exists to catch. + Counted per node, because a pair re-peering shows up on both sides and + a mesh-wide total would hide which one started it. + """ + + per_node: int = 1 + + +@dataclass +class SwitchLatencyAssertion: + """Ceiling on the time from a link going down to the first path switch. + + Read from the runner's own record of when each edge was taken down + against the timestamp of the first "Path switched" line on either of + its endpoints after that moment. This is the number the design puts a + target on ("under one second"); without it a scenario could switch + thirty seconds after the flap and still pass ``path_switches``. + ``max_ms`` bounds the worst flap; a flap with no switch at all within + ``max_ms`` fails outright. + """ + + max_ms: int = 1000 + + +@dataclass +class MaxStallAssertion: + """Ceiling on the longest run of iperf3 intervals that moved no bytes. + + ``min_traffic`` passes on any session that moved any data, and a + switchover that blackholes traffic for ten seconds still leaves plenty + of bytes on either side of the hole. iperf3 reports per-interval + totals (one second by default); this reads the longest run of zeros + inside any session, in seconds, and fails if it exceeds ``max_secs``. + """ + + max_secs: float = 2.0 + + @dataclass class MaxParentSwitchesAssertion: """Stability ceiling: fail if parent switches exceed ``max_total``. @@ -351,6 +417,10 @@ class AssertionsConfig: bloom_send_rate: BloomSendRateAssertion | None = None min_parent_switches: MinParentSwitchesAssertion | None = None max_parent_switches: MaxParentSwitchesAssertion | None = None + path_switches: PathSwitchesAssertion | None = None + max_promotions: MaxPromotionsAssertion | None = None + switch_latency: SwitchLatencyAssertion | None = None + max_stall: MaxStallAssertion | None = None max_errors: MaxErrorsAssertion | None = None congestion_signals: CongestionSignalsAssertion | None = None tree_parents: TreeParentsAssertion | None = None @@ -413,12 +483,12 @@ _SECTION_KEYS = { "num_nodes", "algorithm", "params", "ensure_connected", "subnet", "ip_start", "default_transport", "transport_mix", "pin_root", }, - "netem": {"enabled", "default_policy", "link_policies", "mutation"}, + "netem": {"enabled", "default_policy", "link_policies", "mutation", "dual_udp_policy"}, "netem.link_policies[]": {"edges", "policy", "policy_name"}, "netem.mutation": {"interval_secs", "fraction", "policies", "exclude_edges"}, "link_flaps": { "enabled", "interval_secs", "max_down_links", "down_duration_secs", - "protect_connectivity", + "protect_connectivity", "only_transport", }, "traffic": { "enabled", "max_concurrent", "interval_secs", "duration_secs", @@ -435,8 +505,9 @@ _SECTION_KEYS = { "link_swap.edges[]": {"edge", "policy"}, "assertions": { "bloom_send_rate", "min_parent_switches", "max_parent_switches", - "max_errors", "congestion_signals", "tree_parents", "baseline", - "min_traffic", + "path_switches", "max_promotions", "switch_latency", "max_stall", + "max_errors", "congestion_signals", "tree_parents", + "baseline", "min_traffic", }, "logging": {"rust_log", "output_dir"}, } @@ -444,6 +515,10 @@ _ASSERTION_KEYS = { "bloom_send_rate": {"window_secs", "max_per_node"}, "min_parent_switches": {"min_total"}, "max_parent_switches": {"max_total", "node"}, + "path_switches": {"min_total", "max_total"}, + "max_promotions": {"per_node"}, + "switch_latency": {"max_ms"}, + "max_stall": {"max_secs"}, "max_errors": {"max_total"}, "min_traffic": {"min_sessions_ok", "min_bytes_total"}, "congestion_signals": { @@ -556,6 +631,10 @@ def load_scenario(path: str) -> Scenario: s.netem.default_policy = _parse_netem_policy( nc["default_policy"], "netem.default_policy" ) + if "dual_udp_policy" in nc: + s.netem.dual_udp_policy = _parse_netem_policy( + nc["dual_udp_policy"], "netem.dual_udp_policy" + ) if "link_policies" in nc: for lp_data in nc["link_policies"]: _reject_unknown( @@ -597,6 +676,8 @@ def load_scenario(path: str) -> Scenario: lf["down_duration_secs"], "link_flaps.down_duration_secs" ) s.link_flaps.protect_connectivity = lf.get("protect_connectivity", True) + only = lf.get("only_transport") + s.link_flaps.only_transport = str(only) if only is not None else None # Traffic section tf = raw.get("traffic", {}) @@ -696,6 +777,54 @@ def load_scenario(path: str) -> Scenario: s.assertions.min_parent_switches = MinParentSwitchesAssertion( min_total=int(mps.get("min_total", 1)), ) + if "path_switches" in asrt: + ps = asrt["path_switches"] + _reject_unknown(ps, _ASSERTION_KEYS["path_switches"], "assertions.path_switches") + if "min_total" not in ps and "max_total" not in ps: + raise ValueError( + "assertions.path_switches: give min_total, max_total or both " + "(an empty band asserts nothing)" + ) + for key in ("min_total", "max_total"): + if key in ps and (isinstance(ps[key], bool) or not isinstance(ps[key], int) or ps[key] < 0): + raise ValueError( + f"assertions.path_switches: {key} must be a non-negative " + f"integer, got {ps[key]!r}" + ) + s.assertions.path_switches = PathSwitchesAssertion( + min_total=ps.get("min_total"), + max_total=ps.get("max_total"), + ) + if "max_promotions" in asrt: + mp = asrt["max_promotions"] + _reject_unknown(mp, _ASSERTION_KEYS["max_promotions"], "assertions.max_promotions") + per_node = mp.get("per_node", 1) + if isinstance(per_node, bool) or not isinstance(per_node, int) or per_node < 1: + raise ValueError( + "assertions.max_promotions: per_node must be a positive integer, " + f"got {per_node!r} — a pair has to meet once" + ) + s.assertions.max_promotions = MaxPromotionsAssertion(per_node=per_node) + if "switch_latency" in asrt: + sl = asrt["switch_latency"] + _reject_unknown(sl, _ASSERTION_KEYS["switch_latency"], "assertions.switch_latency") + max_ms = sl.get("max_ms", 1000) + if isinstance(max_ms, bool) or not isinstance(max_ms, int) or max_ms < 1: + raise ValueError( + "assertions.switch_latency: max_ms must be a positive integer, " + f"got {max_ms!r}" + ) + s.assertions.switch_latency = SwitchLatencyAssertion(max_ms=max_ms) + if "max_stall" in asrt: + ms = asrt["max_stall"] + _reject_unknown(ms, _ASSERTION_KEYS["max_stall"], "assertions.max_stall") + max_secs = ms.get("max_secs", 2.0) + if isinstance(max_secs, bool) or not isinstance(max_secs, (int, float)) or max_secs <= 0: + raise ValueError( + "assertions.max_stall: max_secs must be a positive number, " + f"got {max_secs!r}" + ) + s.assertions.max_stall = MaxStallAssertion(max_secs=float(max_secs)) if "max_parent_switches" in asrt: xps = asrt["max_parent_switches"] _reject_unknown( @@ -1012,10 +1141,10 @@ def _validate(s: Scenario): node_ids.update(str(p) for p in entry[:2]) if len(entry) == 3: transport = str(entry[2]) - if transport not in VALID_TRANSPORTS: + if transport not in VALID_TRANSPORTS and transport not in DUAL_TRANSPORTS: raise ValueError( f"explicit adjacency[{i}]: transport '{transport}' " - f"not in {VALID_TRANSPORTS}" + f"not in {VALID_TRANSPORTS} or {DUAL_TRANSPORTS}" ) if len(node_ids) != s.topology.num_nodes: raise ValueError( diff --git a/testing/chaos/sim/topology.py b/testing/chaos/sim/topology.py index 2fe277ec..44dd233c 100644 --- a/testing/chaos/sim/topology.py +++ b/testing/chaos/sim/topology.py @@ -11,6 +11,41 @@ from .keys import derive_full from .naming import name_suffix, veth_token from .scenario import TopologyConfig +# An edge carried by UDP over a dedicated veth pair, each end an +# interface-bound UDP instance (``transports.udp..interface``). The +# harness's stand-in for "wifi and cable, both IP": two UDP instances on +# two interfaces, so a peer reachable over both holds two paths. +UDP_VETH = "udp-veth" + +# Port of the interface-bound UDP instances. Not 2121: the bridge instance +# binds the wildcard on that port, and a second wildcard bind on the same +# port would conflict. +UDP_VETH_PORT = 2122 + +# Second octet of the /24s the veth pairs carry. Clear of docker's default +# pool (172.17-31), the sim's claimed 10.30.x ranges and sidecar's 10.40.x. +_UDP_VETH_NET = "10.222" + + +@dataclass(frozen=True) +class UdpVethLink: + """One end of a ``udp-veth`` edge, as a node sees it.""" + + peer_id: str + # The veth interface in this node's container, and the UDP instance + # name bound to it. + iface: str + local_ip: str + peer_ip: str + + @property + def instance(self) -> str: + return self.iface + + @property + def peer_addr(self) -> str: + return f"{self.peer_ip}:{UDP_VETH_PORT}" + @dataclass class SimNode: @@ -29,6 +64,11 @@ class SimTopology: edges: set[tuple[str, str]] = field(default_factory=set) # Per-edge transport type; edges not in this dict default to "udp" edge_transport: dict[tuple[str, str], str] = field(default_factory=dict) + # Edges declared ``ethernet+udp``: an Ethernet veth (found by beacon) + # *and* a UDP static-peer entry over the bridge, so the pair holds two + # paths under one session. ``edge_transport`` says ``ethernet`` for + # these, which is what netem and link flaps act on. + dual_udp_edges: set[tuple[str, str]] = field(default_factory=set) # Suffix scoping globally-visible names to this run and scenario; empty # outside the CI harness, which keeps a bare run's names unchanged. name_suffix: str = "" @@ -42,6 +82,60 @@ class SimTopology: """ return veth_token(self.name_suffix) + def is_dual_udp_edge(self, a: str, b: str) -> bool: + """Whether the edge also carries a UDP static-peer link.""" + return _make_edge(a, b) in self.dual_udp_edges + + def is_veth_transport(self, transport: str) -> bool: + """Whether edges of this transport run over a dedicated veth pair.""" + return transport in ("ethernet", UDP_VETH) + + def veth_edges(self) -> list[tuple[str, str]]: + """Every edge that needs a veth pair: Ethernet and ``udp-veth``.""" + return sorted( + e for e, t in self.edge_transport.items() if self.is_veth_transport(t) + ) + + def has_veth(self) -> bool: + return bool(self.veth_edges()) + + def udp_veth_edges(self) -> list[tuple[str, str]]: + """Edges carried by UDP over a veth, in canonical order. The index + of an edge here is what its /24 is numbered by.""" + return sorted(e for e, t in self.edge_transport.items() if t == UDP_VETH) + + def udp_veth_links(self, node_id: str) -> list[UdpVethLink]: + """This node's ends of its ``udp-veth`` edges, with addressing. + + Edge ``k`` (in ``udp_veth_edges`` order) is ``10.222.k.0/24``: the + lower node id is ``.1``, the higher ``.2``. + """ + links = [] + for k, (a, b) in enumerate(self.udp_veth_edges()): + if node_id not in (a, b): + continue + if k > 255: + raise ValueError("more than 256 udp-veth edges are not addressable") + local, peer = (a, b) if node_id == a else (b, a) + local_ip = f"{_UDP_VETH_NET}.{k}.{1 if node_id == a else 2}" + peer_ip = f"{_UDP_VETH_NET}.{k}.{2 if node_id == a else 1}" + links.append( + UdpVethLink( + peer_id=peer, + iface=veth_interface_name(local, peer), + local_ip=local_ip, + peer_ip=peer_ip, + ) + ) + return links + + def udp_veth_ip(self, node_id: str, peer_id: str) -> str | None: + """The veth IP ``node_id`` has on its ``udp-veth`` edge to ``peer_id``.""" + for link in self.udp_veth_links(node_id): + if link.peer_id == peer_id: + return link.local_ip + return None + def transport_for_edge(self, a: str, b: str) -> str: """Get the transport type for an edge (defaults to 'udp').""" edge = _make_edge(a, b) @@ -161,6 +255,7 @@ class SimTopology: static_edges = { e for e in self.edges if self.edge_transport.get(e, "udp") != "ethernet" + or e in self.dual_udp_edges } outbound: dict[str, list[str]] = {nid: [] for nid in self.nodes} @@ -249,7 +344,7 @@ def generate_topology( adjacency = config.params.get("adjacency") if not adjacency: raise ValueError("explicit topology requires params.adjacency") - edges, edge_transport = _generate_explicit( + edges, edge_transport, dual_udp_edges = _generate_explicit( adjacency, config.default_transport ) # Validate all referenced nodes exist @@ -264,6 +359,7 @@ def generate_topology( # Assign transport types to edges if config.algorithm != "explicit": edge_transport = _assign_edge_transports(edges, config, rng) + dual_udp_edges = set() # Build peer lists from edges for a, b in edges: @@ -276,6 +372,7 @@ def generate_topology( nodes=nodes, edges=edges, edge_transport=edge_transport, + dual_udp_edges=dual_udp_edges, name_suffix=name_suffix(), ) @@ -354,17 +451,24 @@ def _generate_erdos_renyi( def _generate_explicit( adjacency: list, default_transport: str = "udp" -) -> tuple[set[tuple[str, str]], dict[tuple[str, str], str]]: +) -> tuple[set[tuple[str, str]], dict[tuple[str, str], str], set[tuple[str, str]]]: """Build edges from an explicit adjacency list. Each entry is a 2-element list ``[nodeA, nodeB]`` (uses default - transport) or a 3-element list ``[nodeA, nodeB, transport]``. + transport) or a 3-element list ``[nodeA, nodeB, transport]``. The + transport ``ethernet+udp`` declares a dual edge: an Ethernet veth and + a UDP static-peer link between the same two nodes, so the pair holds + two paths under one session. ``udp-veth+udp`` is the all-IP dual edge: + UDP over a dedicated veth (an interface-bound UDP instance at each + end) and UDP over the bridge. - Returns ``(edges, edge_transport)`` where ``edge_transport`` maps - each edge to its transport type. + Returns ``(edges, edge_transport, dual_udp_edges)`` where + ``edge_transport`` maps each edge to its transport type (``ethernet`` + for a dual edge) and ``dual_udp_edges`` is the set of dual edges. """ edges = set() edge_transport: dict[tuple[str, str], str] = {} + dual_udp_edges: set[tuple[str, str]] = set() for i, entry in enumerate(adjacency): if not isinstance(entry, (list, tuple)) or len(entry) not in (2, 3): raise ValueError( @@ -374,8 +478,11 @@ def _generate_explicit( edge = _make_edge(str(entry[0]), str(entry[1])) edges.add(edge) transport = str(entry[2]) if len(entry) == 3 else default_transport + if transport in ("ethernet+udp", f"{UDP_VETH}+udp"): + transport = transport[: -len("+udp")] + dual_udp_edges.add(edge) edge_transport[edge] = transport - return edges, edge_transport + return edges, edge_transport, dual_udp_edges def _assign_edge_transports( diff --git a/testing/chaos/sim/veth.py b/testing/chaos/sim/veth.py index cb1a2228..aea547b1 100644 --- a/testing/chaos/sim/veth.py +++ b/testing/chaos/sim/veth.py @@ -37,7 +37,7 @@ import subprocess import time from .docker_exec import DockerExecError, docker_exec, docker_exec_quiet -from .topology import SimTopology, veth_interface_name +from .topology import UDP_VETH, SimTopology, veth_interface_name log = logging.getLogger(__name__) @@ -111,14 +111,14 @@ class VethManager: 4. Rename to final names and bring up 5. Query MACs and store in SimNode.ethernet_macs """ - eth_edges = self.topology.ethernet_edges() - if not eth_edges: + edges = self.topology.veth_edges() + if not edges: return image = self._get_image() - log.info("Setting up %d Ethernet veth pairs (helper image: %s)...", len(eth_edges), image) + log.info("Setting up %d veth pairs (helper image: %s)...", len(edges), image) - for a, b in eth_edges: + for a, b in edges: self._create_veth_pair(a, b, image) log.info( @@ -135,7 +135,7 @@ class VethManager: stopped, whose pairs are left for their own restart. """ image = self._get_image() - for a, b in self.topology.ethernet_edges(): + for a, b in self.topology.veth_edges(): if a != node_id and b != node_id: continue # Remove existing pair if any (host-side might still exist) @@ -216,8 +216,15 @@ class VethManager: _require_host(["ip", "link", "set", host_a, "netns", str(pid_a)], image) _require_host(["ip", "link", "set", host_b, "netns", str(pid_b)], image) - _raise_link(container_a, host_a, final_a) - _raise_link(container_b, host_b, final_b) + # A udp-veth edge carries IP: address each end before it comes up, + # so the daemon's interface wait never sees the interface without + # its address. + ip_a = ip_b = None + if self.topology.transport_for_edge(node_a, node_b) == UDP_VETH: + ip_a = self.topology.udp_veth_ip(node_a, node_b) + ip_b = self.topology.udp_veth_ip(node_b, node_a) + _raise_link(container_a, host_a, final_a, ip_a) + _raise_link(container_b, host_b, final_b, ip_b) _await_up(container_a, final_a) _await_up(container_b, final_b) @@ -272,11 +279,13 @@ def _require_host(cmd: list[str], image: str): raise VethSetupError(f"host command failed: {' '.join(cmd)}") -def _raise_link(container: str, temp: str, final: str): - """Rename a moved veth end to its final name and set it up.""" +def _raise_link(container: str, temp: str, final: str, ip: str | None = None): + """Rename a moved veth end to its final name, address it if asked, + and set it up.""" + addr = f" && ip addr add {ip}/24 dev {final}" if ip else "" _in_container( container, - f"ip link set {temp} name {final} && ip link set {final} up", + f"ip link set {temp} name {final}{addr} && ip link set {final} up", f"renaming {temp} to {final}", ) diff --git a/testing/ci-local.sh b/testing/ci-local.sh index c9d91bb2..12c2150a 100755 --- a/testing/ci-local.sh +++ b/testing/ci-local.sh @@ -31,7 +31,7 @@ # nat-lan, nostr-publish-consume, stun-faults, # chaos-churn-mixed-10, chaos-ethernet-mesh, # chaos-ethernet-only, chaos-ethernet-churn, chaos-tcp-mesh, -# chaos-congestion-stress, +# chaos-congestion-stress, dual-path-flap, dual-udp-flap, # sidecar, dns-resolver, deb-install, medium-change # # Opt-in (require --with-tor; depend on live Tor network): @@ -166,6 +166,8 @@ CHAOS_SUITES=( "ethernet-mesh ethernet-mesh" "ethernet-only ethernet-only" "ethernet-churn ethernet-churn" + "dual-path-flap dual-path-flap" + "dual-udp-flap dual-udp-flap" "tcp-mesh tcp-mesh" "congestion-stress congestion-stress" ) diff --git a/testing/docker/entrypoint.sh b/testing/docker/entrypoint.sh index 644b308b..fd9788c7 100644 --- a/testing/docker/entrypoint.sh +++ b/testing/docker/entrypoint.sh @@ -34,14 +34,14 @@ enable_ecn() { } wait_for_ethernet() { - # If config references ethernet transports, wait for interfaces to appear. - # Veth pairs are created from the host after the container starts. + # If config binds any transport to an interface (Ethernet, or a UDP + # instance with `interface:`), wait for those interfaces to appear. Veth + # pairs are created from the host after the container starts, and a UDP + # instance bound to an interface that is not there yet fails to start. local eth_ifaces="" - if grep -q 'ethernet:' "$CONFIG" 2>/dev/null; then - eth_ifaces=$(grep '^\s*interface:' "$CONFIG" \ - | sed 's/.*interface:\s*//' \ - | tr -d ' ' || true) - fi + eth_ifaces=$(grep '^\s*interface:' "$CONFIG" 2>/dev/null \ + | sed 's/.*interface:\s*//' \ + | tr -d ' "' || true) if [ -n "$eth_ifaces" ]; then echo "Waiting for Ethernet interfaces: $eth_ifaces" diff --git a/testing/lib/log_analysis.py b/testing/lib/log_analysis.py index 27039cfa..8d72c38d 100644 --- a/testing/lib/log_analysis.py +++ b/testing/lib/log_analysis.py @@ -40,6 +40,7 @@ class AnalysisResult: peers_promoted: list[tuple[str, str]] = field(default_factory=list) peer_removals: list[tuple[str, str]] = field(default_factory=list) parent_switches: list[tuple[str, str]] = field(default_factory=list) + path_switches: list[tuple[str, str]] = field(default_factory=list) mmp_link_metrics: list[tuple[str, str]] = field(default_factory=list) mmp_session_metrics: list[tuple[str, str]] = field(default_factory=list) handshake_timeouts: list[tuple[str, str]] = field(default_factory=list) @@ -70,6 +71,7 @@ class AnalysisResult: f"Peers promoted: {len(self.peers_promoted)}", f"Peer removals: {len(self.peer_removals)}", f"Parent switches: {len(self.parent_switches)}", + f"Path switches: {len(self.path_switches)}", f"Handshake timeouts: {len(self.handshake_timeouts)}", f"MMP link samples: {len(self.mmp_link_metrics)}", f"MMP session samples: {len(self.mmp_session_metrics)}", @@ -160,6 +162,13 @@ def _analyze_lines(result: AnalysisResult, source: str, log_text: str): # Parent switches if "Parent switched" in line: result.parent_switches.append((source, line)) + # Path switches: a peer's traffic moved to another transport under + # the same session. Three emitters, one per trigger (selection, the + # presence edge, a peer's PathClose); all say "session kept". + if "session kept" in line and ( + "Path switched" in line or "traffic moved to the standby" in line + ): + result.path_switches.append((source, line)) # Handshake timeouts if "timed out" in line and ("handshake" in line.lower() or "Handshake" in line): result.handshake_timeouts.append((source, line))