feat(peer): switchover between a peer's transports, with no handshake

A peer reachable over more than one transport keeps one Noise session
and moves its traffic between transports on failure or degradation.
Implements docs/design/fips-multi-path-switchover.md §4-§10 and closes
the two-interface case in #143.

Three inner link messages next to Heartbeat: `0x52 PathProbe` and
`0x53 PathAck`, carrying a probe id, the sender's path id and a
`remote_active` bit, and `0x54 PathClose`, naming the receiver's path
id and a reason. A probe is an ordinary encrypted frame sent on a
candidate transport; the receiver, having decrypted it against the
session found by index, adds the path as `Probing`, marks it `rx_live`
and answers on that same path. The prober's receipt of the ack marks
the path `Live`, `tx_live`, and takes an RTT sample. No handshake, no
key material, no index allocation. Old nodes drop the unknown types at
debug, so a path to one stays `Probing` and never becomes eligible.

The discovery gate changes shape: a live peer beaconing on a transport
we hold no path to it over becomes a path candidate rather than being
skipped, and the heartbeat tick probes it. The active path's first
probe is small — the handshake proved it and seeded its MTU — while a
standby's discovery probes are padded to the link MTU, as is one a
minute on every path, so a medium that passes small frames and drops
large ones never proves itself. A standby the peer never answers on is
given up after eight probes; the active path never is.

Detection is per path and takes each medium's own failure signal: a
carrier edge, an unreachable-on-send (`ENETUNREACH`/`EHOSTUNREACH`), an
interface going away, or two unanswered heartbeats on a path the peer
is also silent on. Any of them marks the path `Suspect` and selection
leaves it at once, because the standby is warm: heartbeats run at
`node.path.active_heartbeat_ms` where either side sends and
`standby_heartbeat_ms` elsewhere, both stretched by the path's own
round trip so a Tor or Nym path is neither flooded nor declared dead
every round trip. A peer holding one live path is not heartbeated here
at all — selection has nothing to move to, and the link heartbeat keeps
its liveness. Soft signals (the peer's `remote_active` flipping away,
silence here while a standby hears the peer) trigger a probe, never
`Suspect`: reading them as a verdict forces both sides onto one path
and loops under a one-way failure.

A node that loses a path tells the peer with a `PathClose` on a
surviving one, so the peer moves at once rather than after its own
timeout. A transport that returns inside the five-minute grace revives
its dead paths as `Probing` with their RTT window and ETX intact.

Selection is measured, not configured: each path scores
`quality_index(etx, min_rtt)`, and traffic moves when the active path
is no longer eligible (mandatory) or when a standby beats it by
`switch_margin` for `switch_dwell_secs` (discretionary). Min RTT over a
window rather than SRTT, because SRTT inflates under load on the path
carrying traffic while an idle standby looks pristine — a ping-pong
generator. A `role: backup` transport carries a peer's traffic only
while no normal path is eligible, and yields outright when one becomes
eligible. `fipsctl path pin` overrides both while its path is eligible.

A switch re-seeds the path MTU from the new path, tightens the
session MTUs and refreshes the MSS ceiling, so the first frames after a
switch are not black-holed. The link record and `addr_to_link` follow
the active path. The link cost the tree sees is held at its pre-switch
value for the dwell, and until the two receiver reports that span the
switch have arrived — the first counts every frame in flight on the old
path as lost and spikes the per-report ETX for one interval, the second
replaces it — so neither a short flap nor that spike ripples mesh-wide
through parent selection or the next-hop order.

Operator surface: `role: backup` on any transport, `node.path.*`
(`switch_margin`, validated finite and at least 1.0, `switch_dwell_secs`,
`min_samples`, `active_heartbeat_ms`, `standby_heartbeat_ms`), and
`fipsctl path show|pin|unpin` over the `path_show`, `path_pin` and
`path_unpin` control commands.

Two chaos scenarios calibrate the defaults and are wired into both
runners: `dual-path-flap` (a raw-Ethernet veth as the cable, the Docker
bridge over UDP as the wifi) and `dual-udp-flap` (two interface-bound
UDP instances). Each flaps one path under iperf and carries detectors
that can fail on a switchover that did not carry traffic — a per-node
ceiling on "Peer promoted to active" (a second is a re-peering), a
one-second ceiling from link-down to the first switch, and a two-second
ceiling on any zero-byte iperf interval run — alongside the
`path_switches` band. Neither has been run to calibrate; the defaults
are chosen, not derived, and the design doc says so.

Refs #143
This commit is contained in:
Arjen
2026-09-23 11:48:33 -03:00
parent ddb9a8336a
commit dca938c104
54 changed files with 5174 additions and 196 deletions
+6
View File
@@ -686,6 +686,12 @@ jobs:
- suite: ethernet-churn
type: chaos
scenario: ethernet-churn
- suite: dual-path-flap
type: chaos
scenario: dual-path-flap
- suite: dual-udp-flap
type: chaos
scenario: dual-udp-flap
- suite: tcp-mesh
type: chaos
scenario: tcp-mesh
+44 -1
View File
@@ -9,6 +9,48 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
### Added
- Multi-path switchover: a peer reachable over more than one transport keeps
one Noise session and moves its traffic between transports on failure or
degradation, with no handshake. Encrypted frames are demuxed by session
index alone, so a frame from a known peer decrypts whichever transport
delivered it. A peer holds a *path* per transport: the first from the
handshake, further ones proven by a `PathProbe`/`PathAck` exchange under
the existing session. Every path is heartbeated on its own once a peer
has more than one — fast on a path either side sends on, slow on a
standby — and a lost carrier, a route gone on send, an interface gone, or
two unanswered heartbeats on a path the peer is also silent on moves
traffic to the best proven standby inside a second; only a peer with no
path left is dropped. Selection is measured, not configured: per-path
`etx × (1 + min_rtt / 100)` with a margin and a dwell (`node.path.*`), so
a cable coming back under working wifi does not take traffic back until
the wifi degrades. A switch re-seeds the path MTU, tightens session MTUs
and holds the tree-visible link cost for the dwell, and until the two
receiver reports that span the switch have arrived, so neither the
switch nor the one-report loss spike it leaves ripples mesh-wide. A transport that returns inside the five-minute grace revives
its dead paths with their history. Design:
`docs/design/fips-multi-path-switchover.md`; defaults are placeholders,
and `testing/chaos/scenarios/dual-path-flap` and `dual-udp-flap` are the
scenarios that calibrate them, each failing on a re-peering, a switch
slower than a second, or an iperf stall past two seconds.
- Wire, for multi-path: three inner link-message types in the link-control
block, `0x52 PathProbe`, `0x53 PathAck` (probe id, `remote_active` flag,
the sender's path id) and `0x54 PathClose` (the receiver's path id and a
reason). Ordinary encrypted FMP frames under the session; no header,
handshake or index change. Old nodes drop them at debug after
authenticating the frame. The first probe on a standby and one a minute
after on every path are padded to the link MTU, so a medium that passes
small frames and drops large ones never proves itself; a node that loses
a path tells the peer with a `PathClose` on a surviving path, so the
peer moves at once rather than after its own timeout.
- Operator surface, for multi-path: `role: backup` on any transport (never
carries a peer's traffic while a normal path is eligible);
`fipsctl path show|pin|unpin` and the `path_show`, `path_pin`,
`path_unpin` control commands; `node.path.*` (`switch_margin`, validated
finite and at least 1.0, `switch_dwell_secs`, `min_samples`,
`active_heartbeat_ms`, `standby_heartbeat_ms`).
- An authentic frame arriving on a transport the peer has no path on no
longer re-pins the peer's send side to that transport, and a decrypt
failure on such a transport is not counted toward force-removal. Both
@@ -42,7 +84,8 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
on a wifi-only router) is bound and healthy rather than permanently
`Degraded`, and starts carrying traffic the moment a port comes up. Whether
an interface has carrier is reported separately as `interface.carrier` in
`show_transports`, never acted on.
`show_transports`; path selection reads it, presence does not.
This closes the OpenWrt boot race (procd starts `fips` before wifi has
created `fips-mesh0` / `fips-ap0`; both transports were skipped for the life
of the process while the 802.11s peer link formed anyway, so the node looked
+14
View File
@@ -446,6 +446,16 @@ message type). Any successfully decrypted frame — data, gossip, MMP report,
or heartbeat — resets the peer's last-receive timestamp tracked by the MMP
receiver.
### Path probes
A peer reachable over more than one transport holds a path per transport
under the one session; each path is heartbeated on its own with a
PathProbe (0x52) answered by a PathAck (0x53) on the same path, and a
path that is going is announced with a PathClose (0x54) on a surviving
one. Layouts in `wire-formats.md`; the mechanism in
`docs/design/fips-multi-path-switchover.md`. A peer with one path is not
path-probed: the bare Heartbeat above keeps its liveness.
### Dead Timeout
When no traffic (of any kind) is received from a peer for the
@@ -483,6 +493,9 @@ established-frame envelope. They group naturally by purpose:
- **Liveness and lifecycle**: Heartbeat is a minimal frame sent
peer-to-peer to keep the link alive; Disconnect carries an orderly
teardown reason code peer-to-peer.
- **Paths**: PathProbe, PathAck and PathClose (0x52–0x54) prove, measure
and withdraw the individual transports a multi-path peer is reachable
over, peer-to-peer under the one session.
Handshake messages (phase 0x1 msg1, phase 0x2 msg2) travel before
encryption is established and are identified by the FMP common-prefix
@@ -559,6 +572,7 @@ an attacker sends invalid packets to elicit responses.
| Rate limiting (token bucket) | **Implemented** |
| Disconnect with reason codes | **Implemented** |
| Heartbeat liveness detection | **Implemented** |
| Multi-path (per-path probes, switchover without a handshake) | **Implemented** (`feat/multi-path-switchover`) |
| Reconnection handling | **Implemented** |
| Auto-reconnect after link-dead removal | **Implemented** |
| Handshake message retry (link + session layer) | **Implemented** |
+17
View File
@@ -140,6 +140,23 @@ Tell the daemon to drop a peer link.
| -------- | ----------- |
| `peer` | npub (bech32) or hostname from `/etc/fips/hosts`. |
### `path <what>`
A peer reachable over more than one transport holds one Noise session and a
*path* per transport; traffic moves between paths on failure or degradation
without a handshake. These commands show and override that choice.
| Subcommand | Description |
| ---------- | ----------- |
| `path show <peer>` | Every path to the peer: its transport, address, state (`probing`, `live`, `suspect`, `dead`), whether it is the one this node sends on (`active`) and the one the peer sends on (`remote_active`), role, pin, how long since each direction was last proven, the min/last RTT, sample count, per-path ETX and score. Also the link cost the tree sees and whether a post-switch hold is in force. |
| `path pin <peer> <transport>` | Pin this node's traffic to the peer to one transport, named by its configured instance name (e.g. `cable`) or numeric id. The pin holds while the path is eligible (live, and answering probes); while it is not — suspect, dead, or unproven — selection is measured as if unpinned, and the pin re-applies the moment the path is eligible again, with no margin or dwell. `unpin` clears it. |
| `path unpin <peer>` | Clear the pin; selection is measured again from the next tick. |
`peer` is an npub (bech32) or a hostname from `/etc/fips/hosts`. Selection
between paths is measured, not configured (see `node.path.*` in
[configuration.md](configuration.md)); the pin and a transport's
`role: backup` are the only overrides.
### `probe <target>`
Diagnose whether a mesh endpoint is reachable, in five stages, and
+58 -1
View File
@@ -446,6 +446,41 @@ Controls tree construction and parent selection.
| `node.tree.flap_window_secs` | u64 | `60` | Sliding window for counting parent switches |
| `node.tree.flap_dampening_secs` | u64 | `120` | Extended hold-down duration when flap threshold exceeded |
### Path Selection (`node.path.*`)
A peer reachable over more than one transport keeps one Noise session and
holds a *path* per transport. Further paths are added by a probe under the
existing session (a beacon from a live peer on a transport with no path to
it yet is probed, not dialled), each path is heartbeated on its own, and
this node's traffic moves between them on failure or degradation with no
handshake. Selection is measured, not configured: a path's score is
`etx × (1 + min_rtt_ms / 100)` from its own probes. These knobs bound when a
measured difference is acted on; their defaults are placeholders pending
calibration.
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `node.path.switch_margin` | f64 | `1.3` | Discretionary switch margin `K`: the active path's score must exceed the best standby's by this factor. Encodes the fail-back policy: a cable returning under working wifi (1.00 vs about 1.10) is under `K`, so traffic stays until the wifi degrades. |
| `node.path.switch_dwell_secs` | u64 | `2` | The margin must hold this long before a discretionary switch. Also how long the link cost the tree sees is held at its pre-switch value after any switch. |
| `node.path.standby_heartbeat_ms` | u64 | `1000` | Heartbeat interval on a standby path. A standby is only as warm as its last echo; this bounds how stale a proven standby can be when the active path dies. A standby the peer has never acknowledged (an old node, or a medium it cannot hear us on) is probed full-size on a backoff that doubles from this up to `node.heartbeat_interval_secs`, eight times, then given up as dead and forgotten after the five-minute grace; a fresh address starts the count over. |
| `node.path.min_samples` | u32 | `2` | RTT samples a standby needs before it is eligible. |
| `node.path.active_heartbeat_ms` | u64 | `200` | Heartbeat interval on a path that either side sends on. A heartbeat unanswered for two of these on a path the peer has acknowledged before, with nothing heard from the peer on it either, marks it suspect, and selection leaves it at once; a late echo on a path still carrying the peer's frames is counted as loss only, and a late echo still measures the path, so the timeout can stretch on a medium slower than it. Both the interval and the timeout stretch with a path's own measured round trip, so a Tor or Nym path is neither flooded nor declared dead every round trip. The first probe on a standby, and one a minute after on every path, is padded to the link MTU so a medium that passes small frames and drops large ones never proves itself. A peer with a single live path is not path-heartbeated at all — there is nothing to switch to — so on a mesh of single-homed peers this costs nothing. |
A path losing its transport (interface gone), carrier (cable unplugged), or a
route (`ENETUNREACH` on send) is left immediately when another proven path
exists, and the peer is told on a surviving path so it moves too rather than
waiting for its own timeout; only a peer with no path left is dropped. The
interface and carrier signals exist for Ethernet transports, which are bound
per interface; a UDP instance bound with `transports.udp.interface` gets no
presence or carrier signal, so the loss of its path is detected by the
unreachable-on-send error and the heartbeat echo timeout only (two
`active_heartbeat_ms`, about half a second). A transport that returns
within five minutes of going revives its dead paths with their measured
history rather than starting from nothing. On a connection-oriented
transport (`tcp`, `tor`, `nym`) a path's connection is the one a dial or
a probe opened; a `disconnect` closes every path's connection. See
`fipsctl path` in [cli-fipsctl.md](cli-fipsctl.md).
### Bloom Filter (`node.bloom.*`)
| Parameter | Type | Default | Description |
@@ -645,6 +680,12 @@ adding entries and the precedence rules:
## Transports (`transports.*`)
Every transport accepts `role: normal | backup` (default `normal`). A
`backup` transport never carries a peer's traffic while any `normal` path to
that peer is eligible, whatever the measurements say; it is a statement about
the transport's purpose ("drop BLE when something better is stable"), not a
rank. See `node.path.*`.
### UDP (`transports.udp.*`)
| Parameter | Type | Default | Description |
@@ -656,8 +697,9 @@ adding entries and the precedence rules:
| `transports.udp.advertise_on_nostr` | bool | `false` | Include this UDP transport in Nostr endpoint adverts. Implicitly forced false when `outbound_only: true`. |
| `transports.udp.public` | bool | `false` | If advertised: `true` publishes direct `host:port`; `false` publishes `udp:nat` rendezvous |
| `transports.udp.external_addr` | string | *(none)* | Explicit advertise-as override. Bare IP (`"203.0.113.45"` — bind port is appended) or full `host:port`. Takes precedence over the bound address and STUN autodiscovery. Useful when the public IP isn't on a local interface (cloud 1:1 NAT, EIP) or to skip STUN for a deterministic value. |
| `transports.udp.interface` | string | *(none)* | Bind the socket to one interface (e.g. `en0`), making this instance one path. Two instances bound to two interfaces give a peer reachable over both two paths. Linux binds both directions (`SO_BINDTODEVICE`); macOS binds egress only (`IP_BOUND_IF`), so inbound on a wildcard `bind_addr` still arrives from any interface there. Unsupported elsewhere (fails to start). The interface must exist when the daemon starts: unlike an Ethernet transport, an interface-bound UDP instance is not retried when its interface appears later, and it gets no presence or carrier signal while running (see `node.path.*` above). |
| `transports.udp.outbound_only` | bool | `false` | Pure-client posture. When `true`, the transport binds to `0.0.0.0:0` (kernel-assigned ephemeral port) regardless of `bind_addr`, refuses inbound handshake msg1, and is never advertised on Nostr regardless of `advertise_on_nostr`. |
| `transports.udp.interface` | string | *(none)* | Bind the socket to one interface (e.g. `en0`), making this instance one path. Two instances bound to two interfaces give a peer reachable over both two paths. Linux binds both directions (`SO_BINDTODEVICE`); macOS binds egress only (`IP_BOUND_IF`), so inbound on a wildcard `bind_addr` still arrives from any interface there. Unsupported elsewhere (fails to start). The interface must exist when the daemon starts: unlike an Ethernet transport, an interface-bound UDP instance is not retried when its interface appears later, and it gets no presence or carrier signal while running (see `node.path.*` above). |
| `transports.udp.role` | string | `normal` | `normal` or `backup`; see above. |
| `transports.udp.accept_connections` | bool | `true` | Accept inbound handshake msg1 from new peers. Combine with `outbound_only: false` and `accept_connections: false` (plus `auto_connect` on peer entries) for a node that initiates outbound links but rejects fresh inbound handshakes. The handshake handler carves out msg1 from peers already established on this transport so rekey continues to work. |
### Ethernet (`transports.ethernet.*`)
@@ -1032,6 +1074,15 @@ Static peer list. Each entry defines a peer to connect to.
| `peers[].auto_reconnect` | bool | `true` | Automatically reconnect after MMP link-dead removal (exponential backoff, unlimited retries) |
| `peers[].via_nostr` | bool | `false` | Append Nostr advert-derived endpoints after static addresses for this peer |
**Several addresses, one session.** A peer entry may list addresses on
several transports (`udp/main` and `udp/eth0`, or `udp` and `tor`). All
are dialled; the first handshake to complete makes the session, and a
later one to the same live peer over a transport that has no path yet is
kept as a *path* under that session at both ends, not as a second
session. A configured address whose transport was down at dial time is
added as a path once the transport is up and the peer is live. Paths are
listed by `fipsctl path show <peer>`.
**Named UDP instances.** Where several UDP transports are configured
under named sub-keys, a peer address can name the one it belongs to by
writing the transport field as `udp/<instance>`, for example
@@ -1282,6 +1333,12 @@ node:
heartbeat_interval_secs: 10
link_dead_timeout_secs: 30
# drain_timeout_secs: 2 # bounded Draining phase; absent = 2s
path:
switch_margin: 1.3 # K: active score must exceed best standby's by this
switch_dwell_secs: 2 # D: margin must hold this long
min_samples: 2 # N: RTT samples before a standby is eligible
active_heartbeat_ms: 200 # heartbeat on a path either side sends on
standby_heartbeat_ms: 1000 # heartbeat on a standby path
limits:
max_connections: 256
max_peers: 128
+3
View File
@@ -173,6 +173,9 @@ not reproduced here to avoid duplicating the source.
| `probe_start` | `npub` (bech32) | Admits a diagnostic probe job and returns immediately. `data`: `probe_id`, `npub`, `node_addr`, `display_name`, `budget_ms`. |
| `probe_poll` | `probe_id` (integer) | Reports a probe's progress. `data`: `state` (`running` / `done`) and `report`. A terminal job is removed on the poll that observes it, so the report is delivered once. |
| `probe_cancel` | `probe_id` (integer) | Runs the probe's terminal actions immediately, without the teardown grace tick. |
| `path_show` | `npub` (bech32) | Every path to the peer. `data`: `peer`, `link_cost`, `link_cost_held`, and `paths[]` with `transport_id`, `transport` (instance name or null), `addr`, `state`, `active`, `remote_active`, `role`, `pinned`, `rx_live_ms_ago`, `tx_live_ms_ago`, `acked_once`, `last_rtt_ms`, `min_rtt_ms`, `rtt_samples`, `etx`, `score`. |
| `path_pin` | `npub` (bech32), `transport` (instance name or numeric id) | Pins this node's traffic to the peer to that transport's path. Applies on the next selection tick. Error if the peer has no path there. |
| `path_unpin` | `npub` (bech32) | Clears the pin. |
`connect` on a peer the node is **already connected to** neither tears the
live link down nor ignores the address: the address is tried as an alternate
+49
View File
@@ -21,6 +21,9 @@ The FMP link layer defines the following message types, dispatched by the
| 0x31 | LookupResponse | Forwarded — reverse-path via `recent_requests` |
| 0x50 | Disconnect | Peer-to-peer (orderly link teardown) |
| 0x51 | Heartbeat | Peer-to-peer (link liveness) |
| 0x52 | PathProbe | Peer-to-peer (per-path liveness and path discovery) |
| 0x53 | PathAck | Peer-to-peer (echo of a PathProbe, on the same path) |
| 0x54 | PathClose | Peer-to-peer (a path is going; sent on a surviving path) |
Handshake messages travel before encryption is established and are identified
by the FMP common-prefix `phase` field rather than a `msg_type` byte
@@ -155,6 +158,9 @@ the 1-byte message type and message-specific fields.
| 0x31 | LookupResponse | Coordinate discovery response |
| 0x50 | Disconnect | Orderly link teardown |
| 0x51 | Heartbeat | Link liveness probe |
| 0x52 | PathProbe | Per-path liveness probe; proves a transport as a path |
| 0x53 | PathAck | Echo of a PathProbe, sent back on the same path |
| 0x54 | PathClose | Notice that a path is going, sent on a surviving path |
### Noise IK Message 1 (phase 0x1)
@@ -405,6 +411,49 @@ Orderly link teardown with reason code.
| 0x07 | Timeout | Heartbeat liveness timeout |
| 0xFF | Other | Unspecified reason |
### PathProbe (0x52) and PathAck (0x53)
A node reachable over more than one transport holds a *path* per
transport under one session
(`docs/design/fips-multi-path-switchover.md`). A PathProbe is sent on a
candidate or standby path — and, once a peer has more than one path, on
the active one as its heartbeat; the receiver, having decrypted it under
the session, has proof the sender is reachable there and answers with a
PathAck **on the same path**. Both share one layout.
| Offset | Field | Size | Encoding |
| ------ | ----- | ---- | -------- |
| 0 | msg_type | 1 | `0x52` (probe) or `0x53` (ack) |
| 1 | probe_id | 4 | u32 LE — the path's probe sequence; the ack echoes the probe's |
| 5 | flags | 1 | bit 0 `remote_active`: "this path is where I send"; other bits zero |
| 6 | path_id | 4 | u32 LE — the sender's own identifier for the path this travels on |
| 10 | padding | 0..MTU | Zero; ignored by the decoder |
**Fixed part: 10 bytes.** The first probe on a standby, and one a minute
after on every path, are padded to the link MTU so a medium that passes
small frames and drops large ones never proves itself; the ack echoes the
probe's size. Transport ids are local to each node, so each side learns
the other's `path_id` from its probes and acks; a PathClose names a path
by the *receiver's* id.
### PathClose (0x54)
Sent on a surviving path when the sender loses a path to the receiver
(interface gone, carrier lost, operator), so the receiver moves at once
rather than after its own timeout.
| Offset | Field | Size | Encoding |
| ------ | ----- | ---- | -------- |
| 0 | msg_type | 1 | `0x54` |
| 1 | path_id | 4 | u32 LE — the *receiver's* id for the path being closed |
| 5 | reason | 1 | `0` unspecified, `1` interface gone, `2` carrier lost, `3` operator |
**Total: 6 bytes.**
A node that predates these types drops them at debug after
authenticating the frame; a probe to such a node is never acked and the
path never becomes eligible.
### SenderReport (0x01)
Sent by the frame sender to provide interval-based transmission statistics.
+43
View File
@@ -67,6 +67,11 @@ enum Commands {
#[arg(short = 'k', long = "key", conflicts_with = "identity")]
key: Option<PathBuf>,
},
/// Paths to a peer: show them, or pin traffic to one
Path {
#[command(subcommand)]
what: PathCommands,
},
/// Connect to a peer
Connect {
/// Peer identifier: npub (bech32) or hostname from /etc/fips/hosts
@@ -190,6 +195,27 @@ enum ShowCommands {
NativeFlows,
}
#[derive(Subcommand, Debug)]
enum PathCommands {
/// Show every path to a peer, per direction
Show {
/// Peer npub or hostname
peer: String,
},
/// Pin the peer's traffic to one transport until unpinned or the path dies
Pin {
/// Peer npub or hostname
peer: String,
/// Transport instance name (as configured) or numeric id
transport: String,
},
/// Clear the pin
Unpin {
/// Peer npub or hostname
peer: String,
},
}
#[derive(Subcommand, Debug)]
enum AclCommands {
/// Loaded peer ACL state
@@ -597,6 +623,23 @@ fn main() {
let npub = resolve_peer(peer);
build_command("disconnect", serde_json::json!({"npub": npub}))
}
Commands::Path { what } => match what {
PathCommands::Show { peer } => {
let npub = resolve_peer(peer);
build_command("path_show", serde_json::json!({"npub": npub}))
}
PathCommands::Pin { peer, transport } => {
let npub = resolve_peer(peer);
build_command(
"path_pin",
serde_json::json!({"npub": npub, "transport": transport}),
)
}
PathCommands::Unpin { peer } => {
let npub = resolve_peer(peer);
build_command("path_unpin", serde_json::json!({"npub": npub}))
}
},
Commands::Probe {
target,
json,
+3 -3
View File
@@ -39,13 +39,13 @@ pub use gateway::{ConntrackConfig, GatewayConfig, GatewayDnsConfig, PortForward,
pub use node::{
BloomConfig, BuffersConfig, CacheConfig, ControlConfig, LimitsConfig, LookupConfig, MmpConfig,
NativeApiConfig, NetmonConfig, NodeConfig, NostrRendezvousConfig, NostrRendezvousPolicy,
RateLimitConfig, RekeyConfig, RendezvousConfig, RetryConfig, SessionConfig, SessionMmpConfig,
TreeConfig,
PathConfig, RateLimitConfig, RekeyConfig, RendezvousConfig, RetryConfig, SessionConfig,
SessionMmpConfig, TreeConfig,
};
pub use peer::{ConnectPolicy, PeerAddress, PeerConfig, TransportSpec};
pub use transport::{
BleConfig, DirectoryServiceConfig, EthernetConfig, NymConfig, TcpConfig, TorConfig,
TransportInstances, TransportsConfig, UdpConfig,
TransportInstances, TransportRole, TransportsConfig, UdpConfig,
};
/// Default config filename.
+85
View File
@@ -1192,6 +1192,86 @@ impl BuffersConfig {
// ECN Congestion Signaling
// ============================================================================
// ============================================================================
// Path Selection
// ============================================================================
/// Multi-path switchover (`node.path.*`).
///
/// A peer reachable over more than one transport keeps one Noise session and
/// moves its traffic between paths on failure or degradation. Selection is
/// measured, not configured: a path's score is `etx × (1 + min_rtt_ms / 100)`
/// from its own probes. These knobs bound *when* a measured difference is
/// acted on. Their defaults are placeholders to calibrate against
/// `testing/chaos`, not values to reason about
/// (`docs/design/fips-multi-path-switchover.md`, "Calibration").
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct PathConfig {
/// Discretionary switch margin `K` (`node.path.switch_margin`): the
/// active path's score must exceed the best standby's by this factor.
/// This is what encodes the fail-back policy: a cable returning under
/// wifi is 1.00 against 1.10, under the margin, so traffic stays.
#[serde(default = "PathConfig::default_switch_margin")]
pub switch_margin: f64,
/// Discretionary switch dwell `D` in seconds
/// (`node.path.switch_dwell_secs`): the margin must hold this long.
#[serde(default = "PathConfig::default_switch_dwell_secs")]
pub switch_dwell_secs: u64,
/// RTT samples `N` a standby needs before it is eligible
/// (`node.path.min_samples`).
#[serde(default = "PathConfig::default_min_samples")]
pub min_samples: u32,
/// Heartbeat interval on a path that either side sends on, in ms
/// (`node.path.active_heartbeat_ms`). Standby paths use
/// `node.path.standby_heartbeat_ms`. A floor: on a path whose round trip
/// is longer than this (Tor, Nym) the interval and the echo timeout
/// stretch with the measured round trip, so one setting serves a cable
/// and a circuit.
#[serde(default = "PathConfig::default_active_heartbeat_ms")]
pub active_heartbeat_ms: u64,
/// Heartbeat interval on a standby path, in ms
/// (`node.path.standby_heartbeat_ms`). A standby is only as warm as its
/// last echo: this bounds how stale a "proven" standby can be when the
/// active path dies and selection reaches for it. Stretched by the
/// path's round trip like the active interval.
#[serde(default = "PathConfig::default_standby_heartbeat_ms")]
pub standby_heartbeat_ms: u64,
}
impl Default for PathConfig {
fn default() -> Self {
Self {
switch_margin: Self::default_switch_margin(),
switch_dwell_secs: Self::default_switch_dwell_secs(),
min_samples: Self::default_min_samples(),
active_heartbeat_ms: Self::default_active_heartbeat_ms(),
standby_heartbeat_ms: Self::default_standby_heartbeat_ms(),
}
}
}
impl PathConfig {
fn default_switch_margin() -> f64 {
1.3
}
fn default_switch_dwell_secs() -> u64 {
2
}
fn default_min_samples() -> u32 {
2
}
fn default_active_heartbeat_ms() -> u64 {
200
}
fn default_standby_heartbeat_ms() -> u64 {
1000
}
}
/// Rekey / session rekeying configuration (`node.rekey.*`).
///
/// Controls periodic full rekey for both FMP (link layer) and FSP
@@ -1359,6 +1439,10 @@ pub struct NodeConfig {
#[serde(default)]
pub tree: TreeConfig,
/// Multi-path switchover (`node.path.*`).
#[serde(default)]
pub path: PathConfig,
/// Bloom filter (`node.bloom.*`).
#[serde(default)]
pub bloom: BloomConfig,
@@ -1429,6 +1513,7 @@ impl Default for NodeConfig {
rendezvous: RendezvousConfig::default(),
discovery: None,
tree: TreeConfig::default(),
path: PathConfig::default(),
bloom: BloomConfig::default(),
session: SessionConfig::default(),
buffers: BuffersConfig::default(),
+71
View File
@@ -41,6 +41,23 @@ const DEFAULT_UDP_RECV_BUF: usize = 2 * 1024 * 1024;
/// Default UDP send buffer size (2 MB).
const DEFAULT_UDP_SEND_BUF: usize = 2 * 1024 * 1024;
/// What a transport is *for*, as far as path selection is concerned.
///
/// Not a rank. Selection between paths is measured, never configured
/// (`docs/design/fips-multi-path-switchover.md` §8); this is the one
/// statement an operator can make about a transport's purpose: a `backup`
/// transport never carries a peer's traffic while any non-backup path to
/// that peer is eligible. "Drop BLE when something better is stable."
#[derive(Debug, Default, Clone, Copy, PartialEq, Eq, Serialize, Deserialize)]
#[serde(rename_all = "lowercase")]
pub enum TransportRole {
/// Carries traffic whenever it is the best measured path.
#[default]
Normal,
/// Carries traffic only while no `normal` path is eligible.
Backup,
}
/// UDP transport instance configuration.
#[derive(Debug, Clone, Default, Serialize, Deserialize)]
#[serde(deny_unknown_fields)]
@@ -54,6 +71,10 @@ pub struct UdpConfig {
#[serde(default, skip_serializing_if = "Option::is_none")]
pub interface: Option<String>,
/// Path-selection role (`role: normal | backup`). Default: normal.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub role: Option<TransportRole>,
/// Bind address (`bind_addr`). Defaults to "0.0.0.0:2121".
///
/// When `outbound_only = true`, this field is ignored and the transport
@@ -118,6 +139,11 @@ pub struct UdpConfig {
}
impl UdpConfig {
/// Path-selection role. Default: normal.
pub fn role(&self) -> TransportRole {
self.role.unwrap_or_default()
}
/// Get the bind address, using default if not configured.
///
/// When `outbound_only = true`, returns `0.0.0.0:0` so the kernel picks
@@ -269,6 +295,10 @@ const MIN_BEACON_INTERVAL_SECS: u64 = 10;
#[derive(Debug, Clone, Default, Serialize, Deserialize)]
#[serde(deny_unknown_fields)]
pub struct EthernetConfig {
/// Path-selection role (`role: normal | backup`). Default: normal.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub role: Option<TransportRole>,
/// Network interface name (e.g., "eth0", "enp3s0"). Required.
pub interface: String,
@@ -333,6 +363,11 @@ pub struct EthernetConfig {
}
impl EthernetConfig {
/// Path-selection role. Default: normal.
pub fn role(&self) -> TransportRole {
self.role.unwrap_or_default()
}
/// Get the EtherType, using default if not configured.
pub fn ethertype(&self) -> u16 {
self.ethertype.unwrap_or(DEFAULT_ETHERNET_ETHERTYPE)
@@ -407,6 +442,10 @@ const DEFAULT_TCP_MAX_INBOUND: usize = 256;
#[derive(Debug, Clone, Default, Serialize, Deserialize)]
#[serde(deny_unknown_fields)]
pub struct TcpConfig {
/// Path-selection role (`role: normal | backup`). Default: normal.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub role: Option<TransportRole>,
/// Listen address (e.g., "0.0.0.0:443"). If not set, outbound-only.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub bind_addr: Option<String>,
@@ -457,6 +496,11 @@ pub struct TcpConfig {
}
impl TcpConfig {
/// Path-selection role. Default: normal.
pub fn role(&self) -> TransportRole {
self.role.unwrap_or_default()
}
/// Get the default MTU.
pub fn mtu(&self) -> u16 {
self.mtu.unwrap_or(DEFAULT_TCP_MTU)
@@ -555,6 +599,10 @@ const DEFAULT_TOR_ADVERTISED_PORT: u16 = 443;
#[derive(Debug, Clone, Default, Serialize, Deserialize)]
#[serde(deny_unknown_fields)]
pub struct TorConfig {
/// Path-selection role (`role: normal | backup`). Default: normal.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub role: Option<TransportRole>,
/// Tor access mode: "socks5", "control_port", or "directory".
/// Default: "socks5".
#[serde(default, skip_serializing_if = "Option::is_none")]
@@ -652,6 +700,11 @@ impl DirectoryServiceConfig {
}
impl TorConfig {
/// Path-selection role. Default: normal.
pub fn role(&self) -> TransportRole {
self.role.unwrap_or_default()
}
/// Get the access mode. Default: "socks5".
pub fn mode(&self) -> &str {
self.mode.as_deref().unwrap_or("socks5")
@@ -739,6 +792,10 @@ const DEFAULT_BLE_PROBE_COOLDOWN_SECS: u64 = 30;
#[derive(Debug, Clone, Default, Serialize, Deserialize)]
#[serde(deny_unknown_fields)]
pub struct BleConfig {
/// Path-selection role (`role: normal | backup`). Default: normal.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub role: Option<TransportRole>,
/// HCI adapter name (e.g., "hci0"). Required.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub adapter: Option<String>,
@@ -788,6 +845,11 @@ pub struct BleConfig {
}
impl BleConfig {
/// Path-selection role. Default: normal.
pub fn role(&self) -> TransportRole {
self.role.unwrap_or_default()
}
/// Get the adapter name. Default: "hci0".
pub fn adapter(&self) -> &str {
self.adapter.as_deref().unwrap_or("hci0")
@@ -868,6 +930,10 @@ const DEFAULT_NYM_STARTUP_TIMEOUT_SECS: u64 = 120;
#[derive(Debug, Clone, Default, Serialize, Deserialize)]
#[serde(deny_unknown_fields)]
pub struct NymConfig {
/// Path-selection role (`role: normal | backup`). Default: normal.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub role: Option<TransportRole>,
/// SOCKS5 proxy address (host:port). Defaults to "127.0.0.1:1080".
#[serde(default, skip_serializing_if = "Option::is_none")]
pub socks5_addr: Option<String>,
@@ -887,6 +953,11 @@ pub struct NymConfig {
}
impl NymConfig {
/// Path-selection role. Default: normal.
pub fn role(&self) -> TransportRole {
self.role.unwrap_or_default()
}
/// Get the SOCKS5 proxy address. Default: "127.0.0.1:1080".
pub fn socks5_addr(&self) -> &str {
self.socks5_addr
+61
View File
@@ -16,6 +16,9 @@ pub async fn dispatch(node: &mut Node, command: &str, params: Option<&Value>) ->
"probe_start" => probe_start(node, params).await,
"probe_poll" => probe_poll(node, params),
"probe_cancel" => probe_cancel(node, params).await,
"path_show" => path_show(node, params),
"path_pin" => path_pin(node, params),
"path_unpin" => path_unpin(node, params),
_ => Response::error(format!("unknown command: {command}")),
}
}
@@ -124,6 +127,64 @@ fn probe_id(params: Option<&Value>) -> Option<u64> {
params?.get("probe_id").and_then(|v| v.as_u64())
}
fn npub_param<'a>(params: Option<&'a Value>, what: &str) -> Result<&'a str, Response> {
params
.and_then(|p| p.get("npub"))
.and_then(|v| v.as_str())
.ok_or_else(|| Response::error(format!("missing 'npub' parameter for {what}")))
}
/// Show every path to a peer.
///
/// Params: `{"npub": "npub1..."}`
fn path_show(node: &mut Node, params: Option<&Value>) -> Response {
let npub = match npub_param(params, "path_show") {
Ok(v) => v,
Err(r) => return r,
};
match node.api_path_show(npub) {
Ok(data) => Response::ok(data),
Err(msg) => Response::error(msg),
}
}
/// Pin a peer's traffic to one transport.
///
/// Params: `{"npub": "npub1...", "transport": "cable"}`
fn path_pin(node: &mut Node, params: Option<&Value>) -> Response {
let npub = match npub_param(params, "path_pin") {
Ok(v) => v,
Err(r) => return r,
};
let transport = match params
.and_then(|p| p.get("transport"))
.and_then(|v| v.as_str())
{
Some(v) => v,
None => return Response::error("missing 'transport' parameter"),
};
debug!(npub = %npub, transport = %transport, "API path pin requested");
match node.api_path_pin(npub, transport) {
Ok(data) => Response::ok(data),
Err(msg) => Response::error(msg),
}
}
/// Clear a peer's path pin.
///
/// Params: `{"npub": "npub1..."}`
fn path_unpin(node: &mut Node, params: Option<&Value>) -> Response {
let npub = match npub_param(params, "path_unpin") {
Ok(v) => v,
Err(r) => return r,
};
debug!(npub = %npub, "API path unpin requested");
match node.api_path_unpin(npub) {
Ok(data) => Response::ok(data),
Err(msg) => Response::error(msg),
}
}
#[cfg(test)]
mod tests {
use super::*;
+2 -1
View File
@@ -306,7 +306,7 @@ pub fn show_peers(node: &Node) -> Value {
if any_peer_has_srtt && !peer.has_srtt() {
None
} else {
Some(coords.depth() as f64 + peer.link_cost())
Some(coords.depth() as f64 + peer.link_cost(crate::time::mono_ms()))
}
});
peer_json["effective_depth"] = match effective_depth {
@@ -2746,6 +2746,7 @@ mod tests {
// An interface no host has, so presence is deterministically absent
// and carrier deterministically false on every machine this runs on.
let config = EthernetConfig {
role: None,
interface: "fips-absent-x0".to_string(),
ethertype: None,
mtu: None,
+19
View File
@@ -2,17 +2,24 @@
use crate::NodeAddr;
use crate::node::Node;
use crate::transport::{TransportAddr, TransportId};
use tracing::{debug, info, trace};
impl Node {
/// Dispatch a decrypted link message to the appropriate handler.
///
/// Link messages are protocol messages exchanged between authenticated peers.
///
/// `arrival` is the transport and address the frame came in on. Most
/// handlers do not care: the frame was authenticated by the session,
/// which is per peer, not per path. The path probe exchange does: it
/// answers on the path the probe used, and that is the whole point.
pub(in crate::node) async fn dispatch_link_message(
&mut self,
from: &NodeAddr,
plaintext: &[u8],
ce_flag: bool,
arrival: (TransportId, &TransportAddr),
) {
if plaintext.is_empty() {
return;
@@ -58,6 +65,18 @@ impl Node {
// Heartbeat — no-op, last_recv_time already updated by record_recv()
trace!(peer = %self.peer_display_name(from), "Received heartbeat");
}
0x52 => {
// PathProbe
self.handle_path_probe(from, payload, arrival).await;
}
0x53 => {
// PathAck
self.handle_path_ack(from, payload, arrival);
}
0x54 => {
// PathClose
self.handle_path_close(from, payload).await;
}
_ => {
debug!(msg_type = msg_type, "Unknown link message type");
}
+18 -6
View File
@@ -277,6 +277,7 @@ impl Node {
}
address_changed =
peer.set_current_addr(packet.transport_id, packet.remote_addr.clone());
peer.note_path_rx(packet.transport_id, now_ms);
peer.link_stats_mut()
.record_recv(packet.data.len(), packet.timestamp_ms);
peer.touch(packet.timestamp_ms);
@@ -298,7 +299,12 @@ impl Node {
let _ = address_changed;
// Dispatch to link message handler
self.dispatch_link_message(&node_addr, link_message, ce_flag)
self.dispatch_link_message(
&node_addr,
link_message,
ce_flag,
(packet.transport_id, &packet.remote_addr),
)
.await;
}
@@ -377,6 +383,7 @@ impl Node {
if let Some(peer) = self.peers.get_mut(node_addr) {
peer.reset_decrypt_failures();
address_changed = peer.set_current_addr(transport_id, remote_addr.clone());
peer.note_path_rx(transport_id, now_ms);
peer.link_stats_mut()
.record_recv(packet_len, packet_timestamp_ms);
peer.touch(packet_timestamp_ms);
@@ -398,7 +405,12 @@ impl Node {
let _ = address_changed;
}
let link_message = &fmp_plaintext[INNER_TIMESTAMP_LEN..];
self.dispatch_link_message(node_addr, link_message, ce_flag)
self.dispatch_link_message(
node_addr,
link_message,
ce_flag,
(transport_id, remote_addr),
)
.await;
}
@@ -546,10 +558,10 @@ impl Node {
node_addr: &crate::NodeAddr,
transport_id: crate::transport::TransportId,
) {
let on_path = self.peers.get(node_addr).is_some_and(|peer| {
peer.transport_id()
.is_none_or(|bound| bound == transport_id)
});
let on_path = self
.peers
.get(node_addr)
.is_some_and(|peer| peer.paths().is_empty() || peer.path_on(transport_id).is_some());
if !on_path {
trace!(
peer = %self.peer_display_name(node_addr),
+17 -2
View File
@@ -121,6 +121,13 @@ impl Node {
let tick_period = Duration::from_secs(self.config().node.tick_interval_secs);
let mut tick = tokio::time::interval(tick_period);
// The fast path tick: per-path heartbeats and the carrier edge. Its
// own timer because the maintenance tick is seconds and a dead path
// is meant to be noticed inside one.
let mut path_tick = tokio::time::interval(Duration::from_millis(
self.config().node.path.active_heartbeat_ms.max(50),
));
path_tick.set_missed_tick_behavior(tokio::time::MissedTickBehavior::Delay);
// Set up control socket channel
let (control_tx, mut control_rx) =
@@ -410,9 +417,14 @@ impl Node {
// Not policy-filtered: whether an interface's absence
// is normal is a statement about node *health*, not
// about whether the routes over it still work.
if !edge.present {
if edge.present {
// The medium came back: whatever went unanswered
// on it before says nothing about now, so the
// next discovery tick may probe it at once.
self.reset_probe_backoff_on_transport(edge.transport_id);
} else {
let reaped =
self.reap_peers_on_transport(edge.transport_id).await;
self.withdraw_transport(edge.transport_id).await;
if reaped > 0 {
info!(
transport_id = %edge.transport_id,
@@ -490,6 +502,9 @@ impl Node {
).await;
let _ = response_tx.send(response);
}
_ = path_tick.tick() => {
self.run_path_heartbeats().await;
}
deadline = tick.tick() => {
// Tick-body instrumentation. The gate is read ONCE per tick
// into `instr_on`, which is then passed explicitly to every
+22 -6
View File
@@ -206,9 +206,13 @@ impl Node {
{
return true;
}
if self.peers.values().any(|p| {
p.transport_id() == Some(transport_id) && p.current_addr() == Some(remote_addr)
}) {
// Any path, not only the one we send on: a rekey msg1 arrives on the
// path the *peer* sends on.
if self
.peers
.values()
.any(|p| p.is_reachable_at(transport_id, remote_addr))
{
return true;
}
false
@@ -281,9 +285,7 @@ impl Node {
// yields an identity when it matches.
self.peers
.values()
.find(|p| {
p.transport_id() == Some(transport_id) && p.current_addr() == Some(remote_addr)
})
.find(|p| p.is_reachable_at(transport_id, remote_addr))
.map(|p| Msg1Waiver::Expect(*p.node_addr()))
.unwrap_or(Msg1Waiver::Unattributed)
}
@@ -1804,6 +1806,13 @@ impl Node {
self.config().node.tree.announce_min_interval_ms,
);
new_peer.set_path_role(
transport_id,
self.transports
.get(&transport_id)
.map(|t| t.role())
.unwrap_or_default(),
);
self.peers.insert(peer_node_addr, new_peer);
self.peers_by_index
.insert(our_index.as_u32(), peer_node_addr);
@@ -1914,6 +1923,13 @@ impl Node {
new_peer.set_last_tree_announce_sent_ms(ts);
}
new_peer.set_path_role(
transport_id,
self.transports
.get(&transport_id)
.map(|t| t.role())
.unwrap_or_default(),
);
self.peers.insert(peer_node_addr, new_peer);
self.peers_by_index
.insert(our_index.as_u32(), peer_node_addr);
+23 -51
View File
@@ -35,6 +35,14 @@ use tracing::{debug, info, trace, warn};
/// bounded this can come down to the tick.
const HEARTBEAT_RETRY_INTERVAL: Duration = Duration::from_secs(2);
/// How long a `Dead` path keeps its history before it is forgotten.
///
/// Presence flaps on a cable (dock sleep, autoneg bounce) are what the
/// binder's churn guard exists for; a path that came back inside this window
/// is re-probed with its RTT intact rather than measured from nothing. See
/// `docs/design/fips-multi-path-switchover.md` §6.
const DEAD_PATH_GRACE_MS: u64 = 5 * 60 * 1000;
/// Decide whether a peer is due a heartbeat, from the two timestamps it keeps.
///
/// Two gates rather than one. `sent` is when a heartbeat last *landed*, and it
@@ -200,6 +208,8 @@ impl Node {
// Get session timestamp before taking mutable borrow on MMP
let our_timestamp_ms = peer.session_elapsed_ms();
// One report closer to releasing a post-switch cost hold.
peer.note_receiver_report();
let Some(mmp) = peer.mmp_mut() else {
return;
@@ -244,7 +254,7 @@ impl Node {
.peers
.iter()
.filter(|(_, p)| p.has_srtt())
.map(|(a, p)| (*a, p.link_cost()))
.map(|(a, p)| (*a, p.link_cost(now_ms)))
.collect();
// Wall-clock seconds for the escaping declaration timestamp;
// monotonic ms for the flap-dampening / hold-down timers.
@@ -549,6 +559,11 @@ impl Node {
let actions = self.mmp.plan_heartbeats(&snapshots);
// Dead-path history expires here, on the same cadence as liveness.
for peer in self.peers.values_mut() {
peer.prune_dead_paths(now_ms, DEAD_PATH_GRACE_MS);
}
// Wall-clock basis for reconnect scheduling, sourced once (as before).
let now_ms = std::time::SystemTime::now()
.duration_since(std::time::UNIX_EPOCH)
@@ -601,55 +616,14 @@ impl Node {
}
}
/// Reap every active peer reachable only through `transport_id`.
///
/// Called on a transport's detach edge. Until this existed, losing an
/// interface withdrew nothing: the peers stayed in the registry, the
/// routes through them stayed selectable, and this node kept advertising
/// reachability it no longer had — so transit traffic was dropped in
/// silence, and other nodes kept routing toward us for those destinations,
/// until the liveness reaper noticed up to `link_dead_timeout_secs` later.
/// Measured on real hardware that was 27 seconds of routing through a link
/// that had already gone, with four alternative peers available the whole
/// time.
///
/// The detach edge is both earlier and more certain than inactivity, so it
/// is the better trigger. This routes through the same
/// [`Self::route_link_dead`] the liveness reaper uses rather than
/// open-coding a second teardown — every consequence of losing a peer
/// (sessions, path MTU, session indices, the link, the control machine,
/// tree cleanup and re-announce, bloom withdrawal) already hangs off that
/// one path, and a parallel one would drift from it.
///
/// Deliberately undamped. A flapping interface cannot drive a reap storm
/// through here, because `ChurnGuard` stops publishing presence edges
/// after three short-lived bindings and does not resume until one lasts —
/// so the edges this reacts to are already rate-limited at the source, and
/// a second damper here would only add a way for the two to disagree.
///
/// Returns how many peers were reaped.
pub(in crate::node) async fn reap_peers_on_transport(
/// Reap one peer whose last path went away. `now_ms` is the wall-clock
/// reconnect basis, as for the liveness reaper.
pub(in crate::node) async fn reap_peer_without_path(
&mut self,
node_addr: NodeAddr,
transport_id: TransportId,
) -> usize {
let doomed: Vec<NodeAddr> = self
.peers
.iter()
.filter(|(_, peer)| peer.transport_id() == Some(transport_id))
.map(|(node_addr, _)| *node_addr)
.collect();
if doomed.is_empty() {
return 0;
}
let now_ms = std::time::SystemTime::now()
.duration_since(std::time::UNIX_EPOCH)
.map(|d| d.as_millis() as u64)
.unwrap_or(0);
let reaped = doomed.len();
for node_addr in doomed {
now_ms: u64,
) {
debug!(
peer = %self.peer_display_name(&node_addr),
%transport_id,
@@ -657,12 +631,10 @@ impl Node {
);
self.route_link_dead(node_addr, now_ms).await;
}
reaped
}
/// Route a link-dead reap through the peer machine + executor. Two callers
/// decide: the tick sweep's `plan_heartbeats` batch emits a `ReapPeer` for a
/// peer that has gone quiet, and [`Self::reap_peers_on_transport`] withdraws
/// peer that has gone quiet, and [`Self::withdraw_transport`] reaps
/// a transport's peers when its interface goes away. Mirrors
/// [`route_rekey_cadence`](Node::route_rekey_cadence): the shell has already
/// decided by the time this runs, so the machine only CONSUMES the decision
+1
View File
@@ -6,6 +6,7 @@ mod mmp;
mod native;
pub(in crate::node) use native::PendingNative;
pub(in crate::node) mod netmon;
pub(in crate::node) mod path;
pub(crate) mod probe;
// Widened from private by the rekey drain cap: `node::session` calls
// `rekey::drain_max_retention_ms` to bound how long a superseded epoch is
+775
View File
@@ -0,0 +1,775 @@
//! Path probe / path ack: adding a second path to a peer under the session
//! it already has.
//!
//! `docs/design/fips-multi-path-switchover.md` §4. A probe is an ordinary
//! encrypted frame sent on a candidate transport. The receiver, having
//! decrypted it against the session found by index, has proof the peer is
//! reachable on that `(transport, addr)`: it adds the path as `Probing`,
//! marks it `rx_live`, and answers with an ack **on that same path**. The
//! prober's receipt of the ack proves the reverse direction: the path goes
//! `Live`, `tx_live`, and takes an RTT sample. No handshake, no new key
//! material, no index allocation.
//!
//! This is also where the active path moves: selection on the fast tick
//! (`run_path_selection`), withdrawal when a transport goes
//! (`withdraw_transport`), a peer's `PathClose`, and the switch side
//! effects (`apply_path_switch`). The carrier edge and the per-path
//! heartbeats that feed selection live here too.
use crate::NodeAddr;
use crate::node::Node;
use crate::peer::{HeartbeatTiming, PathPolicy, PathSwitch, PathWithdrawal};
use crate::proto::link::{PathClose, PathCloseReason, PathMessage};
use crate::transport::{TransportAddr, TransportId};
use tracing::{debug, info, trace};
impl Node {
/// The selection knobs, from `node.path.*`.
pub(in crate::node) fn path_policy(&self) -> PathPolicy {
let cfg = &self.config().node.path;
let standby_ms = cfg.standby_heartbeat_ms.max(1);
PathPolicy {
margin: cfg.switch_margin,
dwell_ms: cfg.switch_dwell_secs.saturating_mul(1000),
min_samples: cfg.min_samples,
rtt_window_ms: standby_ms
.saturating_mul(u64::from(cfg.min_samples).max(1))
.saturating_mul(2)
.max(30_000),
}
}
/// The role of the transport `transport_id`, or `Normal` if it is not
/// registered.
fn transport_role(&self, transport_id: TransportId) -> crate::config::TransportRole {
self.transports
.get(&transport_id)
.map(|t| t.role())
.unwrap_or_default()
}
/// Run selection for every peer. Called from the tick. A switch here is
/// discretionary or pinned, or mandatory after a `Suspect` mark that
/// nothing else acted on; the presence edge runs its own.
pub(in crate::node) fn run_path_selection(&mut self) {
let policy = self.path_policy();
let now_ms = crate::time::mono_ms();
let switches: Vec<(NodeAddr, PathSwitch)> = self
.peers
.iter_mut()
.filter_map(|(addr, peer)| peer.select_path(now_ms, &policy).map(|s| (*addr, s)))
.collect();
for (node_addr, switch) in switches {
info!(
peer = %self.peer_display_name(&node_addr),
from_transport = %switch.from.0,
to_transport = %switch.to.0,
to_addr = %switch.to.1,
reason = ?switch.reason,
"Path switched, session kept"
);
self.apply_path_switch(&node_addr, switch.to);
}
}
/// Pin a peer's traffic to its path on `transport_id`. Applies on the
/// next selection run. `false` if the peer or the path is unknown.
pub(crate) fn pin_peer_path(
&mut self,
node_addr: &NodeAddr,
transport_id: TransportId,
) -> bool {
self.peers
.get_mut(node_addr)
.is_some_and(|p| p.pin_path(transport_id))
}
/// Clear a peer's pin. `false` if the peer is unknown.
pub(crate) fn unpin_peer_path(&mut self, node_addr: &NodeAddr) -> bool {
match self.peers.get_mut(node_addr) {
Some(p) => {
p.unpin_paths();
true
}
None => false,
}
}
/// The peer named by `npub`, if it is an active peer.
pub(in crate::node) fn resolve_peer_npub(&self, npub: &str) -> Result<NodeAddr, String> {
let identity = crate::PeerIdentity::from_npub(npub)
.map_err(|e| format!("invalid npub '{npub}': {e}"))?;
let node_addr = *identity.node_addr();
if !self.peers.contains_key(&node_addr) {
return Err(format!("peer not found: {npub}"));
}
Ok(node_addr)
}
/// A transport named by its instance name (`cable`, `main`) or its
/// numeric id.
fn resolve_transport_name(&self, name: &str) -> Result<TransportId, String> {
if let Some((id, _)) = self.transports.iter().find(|(_, t)| t.name() == Some(name)) {
return Ok(*id);
}
if let Ok(n) = name.parse::<u32>()
&& self.transports.contains_key(&TransportId::new(n))
{
return Ok(TransportId::new(n));
}
Err(format!("transport not found: {name}"))
}
/// `fipsctl path show <peer>`: every path to the peer, per direction.
pub(crate) fn api_path_show(&self, npub: &str) -> Result<serde_json::Value, String> {
let node_addr = self.resolve_peer_npub(npub)?;
let peer = &self.peers[&node_addr];
let now_ms = crate::time::mono_ms();
let active = peer.transport_id();
let paths: Vec<serde_json::Value> = peer
.paths()
.iter()
.map(|path| {
let ago = |at: Option<u64>| at.map(|t| now_ms.saturating_sub(t));
serde_json::json!({
"transport_id": path.transport_id().as_u32(),
"transport": self
.transports
.get(&path.transport_id())
.and_then(|t| t.name().map(str::to_string)),
"addr": path.addr().to_string(),
"state": format!("{:?}", path.state()).to_lowercase(),
"active": Some(path.transport_id()) == active,
"remote_active": path.remote_active(),
"role": format!("{:?}", path.role()).to_lowercase(),
"pinned": path.pinned(),
"rx_live_ms_ago": ago(path.rx_live_at_ms()),
"tx_live_ms_ago": ago(path.tx_live_at_ms()),
"acked_once": path.acked_once(),
"last_rtt_ms": path.last_rtt_ms(),
"min_rtt_ms": path.min_rtt_ms(),
"rtt_samples": path.rtt_samples(),
"etx": path.etx(),
"score": path.score(),
})
})
.collect();
Ok(serde_json::json!({
"peer": npub,
"link_cost": peer.link_cost(now_ms),
"link_cost_held": peer.link_cost_held(now_ms),
"paths": paths,
}))
}
/// `fipsctl path pin <peer> <transport>`.
pub(crate) fn api_path_pin(
&mut self,
npub: &str,
transport: &str,
) -> Result<serde_json::Value, String> {
let node_addr = self.resolve_peer_npub(npub)?;
let transport_id = self.resolve_transport_name(transport)?;
if !self.pin_peer_path(&node_addr, transport_id) {
return Err(format!("peer {npub} has no path on transport {transport}"));
}
info!(peer = %self.peer_display_name(&node_addr), %transport_id, "Path pinned by operator");
Ok(serde_json::json!({ "pinned": transport_id.as_u32() }))
}
/// `fipsctl path unpin <peer>`.
pub(crate) fn api_path_unpin(&mut self, npub: &str) -> Result<serde_json::Value, String> {
let node_addr = self.resolve_peer_npub(npub)?;
self.unpin_peer_path(&node_addr);
info!(peer = %self.peer_display_name(&node_addr), "Path unpinned by operator");
Ok(serde_json::json!({ "pinned": serde_json::Value::Null }))
}
/// A live peer beaconed on `transport_id` at `remote_addr`, a transport
/// we hold no path to it over: add the path as `Probing`.
///
/// Nothing is sent here. The heartbeat tick is the one issuer of
/// probes: it picks the new path up within one fast interval, probes
/// it full-size, and applies the discovery backoff if the peer never
/// answers (an old node that drops `0x52` at debug), so the path is
/// probed at the capped cadence and never becomes eligible. The backoff
/// is reset when the transport's presence cycles.
pub(in crate::node) fn add_path_candidate(
&mut self,
node_addr: NodeAddr,
transport_id: TransportId,
remote_addr: TransportAddr,
) {
let role = self.transport_role(transport_id);
let Some(peer) = self.peers.get_mut(&node_addr) else {
return;
};
if peer.transport_id() == Some(transport_id) {
// The active path: the handshake proved it.
return;
}
let was_new = peer.path_on(transport_id).is_none();
peer.add_path(transport_id, remote_addr.clone())
.set_role(role);
if was_new {
debug!(
peer = %self.peer_display_name(&node_addr),
transport_id = %transport_id,
remote_addr = %remote_addr,
"Peer beaconed on a new transport; path added, probing"
);
}
}
/// Tests: add the path and probe it at once, as one heartbeat tick
/// would, without the tick's other sends. Uses the test-only
/// `take_probe`, which honours the per-path backoff.
#[cfg(test)]
pub(in crate::node) async fn maybe_probe_path(
&mut self,
node_addr: NodeAddr,
transport_id: TransportId,
remote_addr: TransportAddr,
) {
self.add_path_candidate(node_addr, transport_id, remote_addr.clone());
let now_ms = crate::time::mono_ms();
let timing = self.heartbeat_timing();
let Some(peer) = self.peers.get_mut(&node_addr) else {
return;
};
if peer.transport_id() == Some(transport_id) {
return;
}
let Some((probe_id, remote_active, path_id)) = peer.take_probe(
transport_id,
now_ms,
timing.fast_ms,
timing.discovery_cap_ms,
) else {
return;
};
let probe = PathMessage {
probe_id,
remote_active,
path_id,
};
let wire = self.pad_to_link_mtu(probe.encode_probe().to_vec(), transport_id, &remote_addr);
if let Err(e) = self
.send_encrypted_link_message_on_path(&node_addr, &wire, transport_id, remote_addr)
.await
{
debug!(peer = %self.peer_display_name(&node_addr), error = %e, "Path probe send failed");
}
}
/// A `PathProbe` arrived from `from` on `arrival`.
///
/// The frame decrypted under `from`'s session, so `from` is reachable
/// over `arrival`. Record the path and answer on it. The ack says
/// whether `arrival` is the path *we* send on, which it usually is not.
pub(in crate::node) async fn handle_path_probe(
&mut self,
from: &NodeAddr,
payload: &[u8],
arrival: (TransportId, &TransportAddr),
) {
let probe = match PathMessage::decode(payload) {
Ok(p) => p,
Err(e) => {
debug!(peer = %self.peer_display_name(from), error = %e, "Malformed path probe");
return;
}
};
let (transport_id, remote_addr) = arrival;
let now_ms = crate::time::mono_ms();
let role = self.transport_role(transport_id);
let Some(peer) = self.peers.get_mut(from) else {
return;
};
let was_new = peer.path_on(transport_id).is_none();
peer.note_path_probe(
transport_id,
remote_addr.clone(),
probe.remote_active,
probe.path_id,
now_ms,
);
if was_new {
peer.set_path_role(transport_id, role);
}
let ours_active = peer.transport_id() == Some(transport_id);
let our_path_id = peer
.path_on(transport_id)
.map(|p| p.local_id())
.unwrap_or(0);
if was_new {
debug!(
peer = %self.peer_display_name(from),
transport_id = %transport_id,
remote_addr = %remote_addr,
"Peer probed a new path; added"
);
}
// The ack echoes the probe's size, so a full-size probe proves the
// path for data-sized frames in both directions.
let ack = PathMessage {
probe_id: probe.probe_id,
remote_active: ours_active,
path_id: our_path_id,
};
let mut wire = ack.encode_ack().to_vec();
let probe_len = payload.len() + 1;
if wire.len() < probe_len {
wire.resize(probe_len, 0);
}
if let Err(e) = self
.send_encrypted_link_message_on_path(from, &wire, transport_id, remote_addr.clone())
.await
{
debug!(
peer = %self.peer_display_name(from),
transport_id = %transport_id,
error = %e,
"Path ack send failed"
);
}
}
/// A `PathAck` arrived from `from` on `arrival`: our probe on that path
/// reached the peer and the answer reached us, so the path works in
/// both directions.
pub(in crate::node) fn handle_path_ack(
&mut self,
from: &NodeAddr,
payload: &[u8],
arrival: (TransportId, &TransportAddr),
) {
let ack = match PathMessage::decode(payload) {
Ok(a) => a,
Err(e) => {
debug!(peer = %self.peer_display_name(from), error = %e, "Malformed path ack");
return;
}
};
let (transport_id, _) = arrival;
let now_ms = crate::time::mono_ms();
let window_ms = self.path_policy().rtt_window_ms;
let Some(peer) = self.peers.get_mut(from) else {
return;
};
let was_live = peer
.path_on(transport_id)
.is_some_and(|p| p.state() == crate::peer::PathState::Live);
match peer.note_path_ack(
transport_id,
ack.probe_id,
ack.remote_active,
ack.path_id,
now_ms,
window_ms,
) {
Some(rtt_ms) if !was_live => info!(
peer = %self.peer_display_name(from),
transport_id = %transport_id,
rtt_ms,
"Path live: the peer answers on this transport"
),
Some(rtt_ms) => trace!(
peer = %self.peer_display_name(from),
transport_id = %transport_id,
rtt_ms,
"Path ack"
),
None => trace!(
peer = %self.peer_display_name(from),
transport_id = %transport_id,
probe_id = ack.probe_id,
"Path ack matched no outstanding probe"
),
}
}
/// A transport's presence went away: withdraw the path every peer held
/// over it. Returns how many peers were reaped for want of another path.
///
/// For each peer: the path goes `Dead` with its history kept. If it was
/// a standby, nothing else happens. If it was the active path and an
/// eligible standby exists, traffic moves there now, under the same
/// session, and the switch side effects run
/// ([`apply_path_switch`](Self::apply_path_switch)). Only a peer with
/// no eligible path left is reaped, through the same routed link-dead
/// teardown the liveness reaper uses.
///
/// The peer machine sees nothing while any path remains: a switch is
/// not a link event. Deliberately undamped, like the reap it grew from.
pub(in crate::node) async fn withdraw_transport(&mut self, transport_id: TransportId) -> usize {
let now_ms = crate::time::mono_ms();
let policy = self.path_policy();
let affected: Vec<NodeAddr> = self
.peers
.iter()
.filter(|(_, peer)| peer.path_on(transport_id).is_some())
.map(|(node_addr, _)| *node_addr)
.collect();
if affected.is_empty() {
return 0;
}
let wall_ms = Self::now_ms();
let mut reaped = 0;
for node_addr in affected {
let outcome = match self.peers.get_mut(&node_addr) {
Some(peer) => peer.withdraw_path(transport_id, now_ms, &policy),
None => continue,
};
match outcome {
PathWithdrawal::NoPath => {}
PathWithdrawal::Standby => {
debug!(
peer = %self.peer_display_name(&node_addr),
%transport_id,
"Standby path withdrawn: its interface went away"
);
self.send_path_close(&node_addr, transport_id, PathCloseReason::InterfaceGone)
.await;
}
PathWithdrawal::Switched { from, to } => {
info!(
peer = %self.peer_display_name(&node_addr),
from_transport = %from.0,
to_transport = %to.0,
to_addr = %to.1,
"Active path withdrawn: traffic moved to the standby, session kept"
);
self.apply_path_switch(&node_addr, to);
self.send_path_close(&node_addr, transport_id, PathCloseReason::InterfaceGone)
.await;
}
PathWithdrawal::NoAlternative => {
self.reap_peer_without_path(node_addr, transport_id, wall_ms)
.await;
reaped += 1;
}
}
}
reaped
}
/// Everything that follows the peer's active path changing to `to`.
///
/// A switch is also an MTU change, and three things size traffic from
/// the peer's transport without re-running on their own
/// (`docs/design/fips-multi-path-switchover.md` §5):
///
/// - the peer's `path_mtu_lookup` seed, which only ever tightens within
/// a link and would leave one cable→BLE excursion clamping every new
/// flow to this peer at the BLE MTU after fail-back; its relinked
/// branch is the hook, so re-seed from the new path;
/// - the per-session source MTU, which tightens on the next send anyway
/// but only loosens after tens of seconds; tighten it now for every
/// session this peer is the next hop of, so the TUN gate answers with
/// PTB instead of losing the first packet per flow at the transport;
/// - the node-wide MSS ceiling.
///
/// The link record follows the traffic so everything that reports the
/// peer's transport and address by link stays truthful; the control
/// machine is keyed on the link and is untouched.
pub(in crate::node) fn apply_path_switch(
&mut self,
node_addr: &NodeAddr,
to: (TransportId, TransportAddr),
) {
let (transport_id, addr) = to;
if let Some(link_id) = self.peers.get(node_addr).map(|p| p.link_id())
&& let Some(link) = self.links.get_mut(&link_id)
{
self.addr_to_link.retain(|_, mapped| *mapped != link_id);
link.rebind(transport_id, addr.clone());
self.addr_to_link
.insert((transport_id, addr.clone()), link_id);
}
self.seed_path_mtu_for_link_peer(node_addr, transport_id, &addr);
let link_mtu = self
.transports
.get(&transport_id)
.map(|t| t.link_mtu(&addr));
if let Some(link_mtu) = link_mtu {
let dests: Vec<NodeAddr> = self.sessions.keys().copied().collect();
for dest in dests {
let via_peer = self
.find_next_hop(&dest)
.is_some_and(|hop| hop.node_addr() == node_addr);
if !via_peer {
continue;
}
if let Some(mmp) = self.sessions.get_mut(&dest).and_then(|s| s.mmp_mut()) {
mmp.path_mtu.seed_source_mtu(link_mtu);
}
}
}
self.refresh_tun_mss_ceiling();
}
/// The fast path tick: per-path heartbeats, the carrier edge, and the
/// selection that a `Suspect` mark may call for.
///
/// Runs every `node.path.active_heartbeat_ms`. Detection is near-instant
/// for direct peers because each medium's own failure signal is used,
/// not a faster timer: carrier here, a failed echo from the heartbeat
/// plan, an unreachable send from the send path. All three mark a path
/// `Suspect`; selection acts on `Suspect` at once because the standby is
/// warm.
pub(in crate::node) async fn run_path_heartbeats(&mut self) {
let carrier_closes = self.poll_carrier_edges();
let now_ms = crate::time::mono_ms();
let timing = self.heartbeat_timing();
let mut sends = Vec::new();
for (node_addr, peer) in self.peers.iter_mut() {
let plan = peer.plan_heartbeats(now_ms, &timing);
for transport_id in plan.suspects {
debug!(
peer = %node_addr,
%transport_id,
"Path suspect: heartbeat echo timed out"
);
}
for send in plan.sends {
sends.push((*node_addr, send));
}
}
for (node_addr, send) in sends {
let probe = PathMessage {
probe_id: send.probe_id,
remote_active: send.remote_active,
path_id: send.path_id,
};
let mut wire = probe.encode_probe().to_vec();
if send.full_size {
wire = self.pad_to_link_mtu(wire, send.transport_id, &send.addr);
}
if let Err(e) = self
.send_encrypted_link_message_on_path(
&node_addr,
&wire,
send.transport_id,
send.addr,
)
.await
{
trace!(
peer = %self.peer_display_name(&node_addr),
transport_id = %send.transport_id,
error = %e,
"Path heartbeat send failed"
);
}
}
self.run_path_selection();
// After selection: a close for the path we were sending on can only
// go out once traffic has moved off it, and `send_path_close` sends
// nothing for the path that is still active.
for (node_addr, transport_id) in carrier_closes {
self.send_path_close(&node_addr, transport_id, PathCloseReason::CarrierLost)
.await;
}
}
/// The heartbeat intervals from `node.path.*` and `node.heartbeat_interval_secs`.
fn heartbeat_timing(&self) -> HeartbeatTiming {
let cfg = &self.config().node;
let fast_ms = cfg.path.active_heartbeat_ms.max(50);
let slow_ms = cfg.path.standby_heartbeat_ms.max(fast_ms);
HeartbeatTiming {
fast_ms,
slow_ms,
// Two fast intervals: one echo lost is loss, two is a path.
// Stretched per path by its own round trip inside
// `plan_heartbeats`.
timeout_ms: fast_ms.saturating_mul(2),
// A path the peer never acknowledges is probed full-size at
// this cadence for as long as it exists: the link heartbeat
// interval, not the standby one.
discovery_cap_ms: cfg
.heartbeat_interval_secs
.saturating_mul(1000)
.max(slow_ms),
}
}
/// Read carrier on every interface-bound transport and mark the paths
/// over one that just lost it `Suspect`. Unplugging a cable drops
/// carrier on both NICs, so both ends see this inside one fast tick.
/// Returns the `(peer, transport)` pairs to send a `PathClose` for.
fn poll_carrier_edges(&mut self) -> Vec<(NodeAddr, TransportId)> {
let mut closes = Vec::new();
let readings: Vec<(TransportId, bool)> = self
.transports
.iter()
.filter_map(|(id, t)| t.interface_presence().map(|p| (*id, p.carrier)))
.collect();
for (transport_id, carrier) in readings {
let previous = self.carrier_seen.insert(transport_id, carrier);
if previous == Some(true) && !carrier {
let mut marked = Vec::new();
for (node_addr, peer) in self.peers.iter_mut() {
if peer.mark_path_suspect(transport_id) {
marked.push(*node_addr);
}
}
if !marked.is_empty() {
info!(%transport_id, paths = marked.len(), "Carrier lost: paths suspect");
closes.extend(marked.into_iter().map(|a| (a, transport_id)));
}
}
}
closes
}
/// Tell `node_addr` that our path over `transport_id` is closing, on
/// whichever path we now send on. Best effort: the peer would learn
/// from the echo timeout anyway, this just makes it immediate. Nothing
/// is sent if that path is the one we send on (there is no other way to
/// reach the peer) or the peer never told us its id for it, which it
/// does with its first probe or ack on the path.
pub(in crate::node) async fn send_path_close(
&mut self,
node_addr: &NodeAddr,
transport_id: TransportId,
reason: PathCloseReason,
) {
let Some(peer) = self.peers.get(node_addr) else {
return;
};
if peer.transport_id() == Some(transport_id) {
return;
}
let Some(remote_id) = peer.path_on(transport_id).and_then(|p| p.remote_id()) else {
return;
};
let close = PathClose {
path_id: remote_id,
reason,
};
if let Err(e) = self
.send_encrypted_link_message(node_addr, &close.encode())
.await
{
trace!(
peer = %self.peer_display_name(node_addr),
%transport_id,
error = %e,
"Path close send failed"
);
}
}
/// The peer is closing the path it calls `path_id` (a `PathClose`
/// arrived). Withdraw our side of it as if its transport had gone: Dead
/// with history, traffic moved if it was there. Advisory: the peer's
/// next probe on it revives it.
pub(in crate::node) async fn handle_path_close(&mut self, from: &NodeAddr, payload: &[u8]) {
let close = match PathClose::decode(payload) {
Ok(c) => c,
Err(e) => {
debug!(peer = %self.peer_display_name(from), error = %e, "Malformed path close");
return;
}
};
let now_ms = crate::time::mono_ms();
let policy = self.path_policy();
let Some(peer) = self.peers.get_mut(from) else {
return;
};
let Some((transport_id, outcome)) =
peer.withdraw_path_by_local_id(close.path_id, now_ms, &policy)
else {
trace!(peer = %self.peer_display_name(from), path_id = close.path_id, "Path close named no path");
return;
};
match outcome {
PathWithdrawal::NoPath => {}
PathWithdrawal::Standby => debug!(
peer = %self.peer_display_name(from),
%transport_id,
reason = ?close.reason,
"Peer closed a standby path"
),
PathWithdrawal::Switched { from: was, to } => {
info!(
peer = %self.peer_display_name(from),
from_transport = %was.0,
to_transport = %to.0,
reason = ?close.reason,
"Peer closed our active path: traffic moved to the standby, session kept"
);
self.apply_path_switch(from, to);
}
PathWithdrawal::NoAlternative => {
// The peer says the only path we have to it is going. Leave
// the peer to the echo timeout and the liveness reaper: a
// close is advisory, and the path may outlive the warning.
debug!(
peer = %self.peer_display_name(from),
%transport_id,
reason = ?close.reason,
"Peer closed our only path; waiting for liveness to confirm"
);
}
}
}
/// Pad a link message to fill the link MTU on `transport_id` to `addr`,
/// so the frame is data-sized: outer header, inner timestamp and AEAD
/// tag are accounted for.
fn pad_to_link_mtu(
&self,
mut wire: Vec<u8>,
transport_id: TransportId,
addr: &TransportAddr,
) -> Vec<u8> {
let Some(transport) = self.transports.get(&transport_id) else {
return wire;
};
let room = usize::from(transport.link_mtu(addr))
.saturating_sub(super::session::LINK_FRAME_OVERHEAD);
if wire.len() < room {
wire.resize(room, 0);
}
wire
}
/// The kernel refused a send to the peer on `transport_id` for want of
/// a route: the path is `Suspect` now, not after an echo timeout.
pub(in crate::node) fn note_path_unreachable(
&mut self,
node_addr: &NodeAddr,
transport_id: TransportId,
) {
if let Some(peer) = self.peers.get_mut(node_addr)
&& peer.mark_path_suspect(transport_id)
{
debug!(
peer = %self.peer_display_name(node_addr),
%transport_id,
"Path suspect: send unreachable"
);
}
}
/// A transport's presence came back: clear the probe backoff on every
/// path over it so the next discovery tick may probe at once.
pub(in crate::node) fn reset_probe_backoff_on_transport(&mut self, transport_id: TransportId) {
for peer in self.peers.values_mut() {
peer.reset_probe_backoff_on(transport_id);
}
}
}
+1 -1
View File
@@ -67,7 +67,7 @@ pub(in crate::node) const PATH_MTU_RELEASE_MIN_INTERVAL: std::time::Duration =
/// import above, which is `#[cfg(unix)]`. This constant feeds `link_wire_len`,
/// whose caller `send_session_datagram` is compiled on every platform, so
/// taking the name from that import fails to build on Windows.
const LINK_FRAME_OVERHEAD: usize =
pub(in crate::node) const LINK_FRAME_OVERHEAD: usize =
crate::proto::fmp::wire::ESTABLISHED_HEADER_SIZE + 4 + crate::noise::TAG_SIZE;
/// Wire size of an encoded `SessionDatagram` of `encoded_len` bytes.
+13 -4
View File
@@ -779,6 +779,9 @@ impl Node {
// dataplane maps are unmutated, so the core's per-peer cap sees a stable
// in-flight count — the same guarantee the old collect-then-dial had.
let mut transport_neighbors: Vec<Candidate> = Vec::new();
// Live peers beaconing on a transport we hold no path to them over.
// Added after the loop, which borrows the transport table.
let mut path_candidates: Vec<(NodeAddr, TransportId, TransportAddr)> = Vec::new();
for (transport_id, transport) in &self.transports {
if !transport.is_operational() {
continue;
@@ -840,6 +843,11 @@ impl Node {
// again. What is given up is switching away from a link
// that is working, which is not a thing worth doing.
if self.active_peer_link_is_live(&node_addr) {
// A live peer beaconing on a transport we hold no
// path to it over is a path to add, not a link to
// replace: the heartbeat tick probes it under the
// existing session instead of dialling.
path_candidates.push((node_addr, candidate_transport_id, remote_addr));
continue;
}
if self.is_connecting_to_peer_on_path(
@@ -867,6 +875,10 @@ impl Node {
}
}
for (node_addr, transport_id, remote_addr) in path_candidates {
self.add_path_candidate(node_addr, transport_id, remote_addr);
}
if transport_neighbors.is_empty() {
return;
}
@@ -3434,10 +3446,7 @@ impl Node {
/// Notifies the peer, removes it locally, closes the transport connection
/// it was using, and suppresses auto-reconnect.
pub(crate) async fn api_disconnect(&mut self, npub: &str) -> Result<serde_json::Value, String> {
let peer_identity =
PeerIdentity::from_npub(npub).map_err(|e| format!("invalid npub '{npub}': {e}"))?;
let node_addr = *peer_identity.node_addr();
let node_addr = self.resolve_peer_npub(npub)?;
let Some(peer) = self.peers.get(&node_addr) else {
return Err(format!("peer not found: {npub}"));
};
+66 -9
View File
@@ -636,6 +636,9 @@ pub struct Node {
/// the peer entry itself. Pruned on insert; see
/// `EPOCH_RESTART_MIN_INTERVAL_SECS`.
restart_dampener: HashMap<NodeAddr, std::time::Instant>,
/// Last carrier reading per interface-bound transport, for the carrier
/// edge the fast path tick detects. Absent until first read.
carrier_seen: HashMap<TransportId, bool>,
// === Rate Limiting ===
/// Rate limiter for msg1 processing (DoS protection).
@@ -921,6 +924,7 @@ impl Node {
peers_by_index: HashMap::new(),
pending_outbound: HashMap::new(),
restart_dampener: HashMap::new(),
carrier_seen: HashMap::new(),
msg1_rate_limiter,
setup_rate_limiter,
icmp_rate_limiter: IcmpRateLimiter::new(),
@@ -1095,6 +1099,7 @@ impl Node {
peers_by_index: HashMap::new(),
pending_outbound: HashMap::new(),
restart_dampener: HashMap::new(),
carrier_seen: HashMap::new(),
msg1_rate_limiter,
setup_rate_limiter,
icmp_rate_limiter: IcmpRateLimiter::new(),
@@ -2279,6 +2284,7 @@ impl Node {
// (their effective_depth is `None`); during cold start (no peer has
// SRTT) every peer falls back to the default link cost of 1.0.
let any_peer_has_srtt = self.peers().any(|p| p.has_srtt());
let now_ms = crate::time::mono_ms();
let now_ms = Self::now_ms();
let peer_rows: Vec<snap::PeerRow> = self
@@ -2324,7 +2330,7 @@ impl Node {
if any_peer_has_srtt && !peer.has_srtt() {
None
} else {
Some(coords.depth() as f64 + peer.link_cost())
Some(coords.depth() as f64 + peer.link_cost(now_ms))
}
});
@@ -3741,6 +3747,41 @@ impl Node {
node_addr: &NodeAddr,
plaintext: &[u8],
ce_flag: bool,
) -> Result<(), NodeError> {
self.send_encrypted_link_message_via(node_addr, plaintext, ce_flag, None)
.await
}
/// Like `send_encrypted_link_message` but on a chosen path rather than
/// the peer's active one.
///
/// The path probe exchange uses this to reach a peer over a transport it
/// is not (yet) sending on. Same session, same counter, same key: only
/// the transport and address differ.
pub(super) async fn send_encrypted_link_message_on_path(
&mut self,
node_addr: &NodeAddr,
plaintext: &[u8],
transport_id: TransportId,
remote_addr: TransportAddr,
) -> Result<(), NodeError> {
self.send_encrypted_link_message_via(
node_addr,
plaintext,
false,
Some((transport_id, remote_addr)),
)
.await
}
/// The one send path for encrypted link messages. `via` picks the
/// transport and address; `None` means the peer's active path.
async fn send_encrypted_link_message_via(
&mut self,
node_addr: &NodeAddr,
plaintext: &[u8],
ce_flag: bool,
via: Option<(TransportId, TransportAddr)>,
) -> Result<(), NodeError> {
let peer = self
.peers
@@ -3751,17 +3792,24 @@ impl Node {
node_addr: *node_addr,
reason: "no their_index".into(),
})?;
let on_active_path = via.is_none();
let (transport_id, remote_addr) = match via {
Some(target) => target,
None => {
let transport_id = peer.transport_id().ok_or_else(|| NodeError::SendFailed {
node_addr: *node_addr,
reason: "no transport_id".into(),
})?;
let remote_addr = peer
.current_addr()
let remote_addr =
peer.current_addr()
.cloned()
.ok_or_else(|| NodeError::SendFailed {
node_addr: *node_addr,
reason: "no current_addr".into(),
})?;
(transport_id, remote_addr)
}
};
// Prepend 4-byte session-relative timestamp (inner header)
let timestamp_ms = peer.session_elapsed_ms();
@@ -3779,8 +3827,16 @@ impl Node {
// Snapshot the per-peer connect()-ed UDP socket BEFORE the
// session borrow so the encrypt-worker dispatch can refcount-
// clone the Arc without re-borrowing self.peers later.
// The connected socket is pinned to the active path's 5-tuple, so a
// send on any other path must go through the listen socket.
#[cfg(any(target_os = "linux", target_os = "macos"))]
let connected_socket = peer.connected_udp();
let connected_socket = if on_active_path {
peer.connected_udp()
} else {
None
};
#[cfg(not(any(target_os = "linux", target_os = "macos")))]
let _ = on_active_path;
let session = peer
.noise_session_mut()
@@ -3918,10 +3974,11 @@ impl Node {
}
}
let bytes_sent = transport
.send(&remote_addr, &wire_packet)
.await
.map_err(|e| match e {
let sent = transport.send(&remote_addr, &wire_packet).await;
if sent.as_ref().is_err_and(|e| e.is_unreachable()) {
self.note_path_unreachable(node_addr, transport_id);
}
let bytes_sent = sent.map_err(|e| match e {
TransportError::MtuExceeded { packet_size, mtu } => NodeError::MtuExceeded {
node_addr: *node_addr,
packet_size,
@@ -4006,7 +4063,7 @@ impl routing::RoutingView for NodeRoutingView<'_> {
}
fn peer_link_cost<'a>(&'a self, peer: Self::Peer<'a>) -> f64 {
peer.1.link_cost()
peer.1.link_cost(crate::time::mono_ms())
}
fn peer_coords<'a>(&'a self, peer: Self::Peer<'a>) -> Option<&'a TreeCoordinate> {
+2
View File
@@ -2212,6 +2212,7 @@ async fn a_transient_msg2_failure_keeps_the_link_for_the_retry() {
// `InterfaceUnavailable` — the real error, from the real code path,
// rather than a stub that merely returns something transient.
let config = EthernetConfig {
role: None,
interface: "fips-absent-x0".to_string(),
ethertype: None,
mtu: None,
@@ -2294,6 +2295,7 @@ async fn a_transient_msg2_failure_on_the_restart_path_leaves_the_fresh_leg_pendi
// An interface no host has, so every send off this transport reports
// `InterfaceUnavailable` — the real error from the real code path.
let config = EthernetConfig {
role: None,
interface: "fips-absent-x0".to_string(),
ethertype: None,
mtu: None,
+2 -2
View File
@@ -368,7 +368,7 @@ async fn a_detached_transport_withdraws_the_peers_that_needed_it() {
.transport_id()
.expect("an established peer names its transport");
let reaped = nodes[0].node.reap_peers_on_transport(transport_id).await;
let reaped = nodes[0].node.withdraw_transport(transport_id).await;
assert_eq!(reaped, 1);
assert!(
@@ -399,7 +399,7 @@ async fn a_detached_transport_leaves_other_transports_peers_alone() {
// A transport this peer was never reachable through.
let unrelated = TransportId::new(peer_transport.as_u32() + 100);
let reaped = nodes[0].node.reap_peers_on_transport(unrelated).await;
let reaped = nodes[0].node.withdraw_transport(unrelated).await;
assert_eq!(reaped, 0, "an unrelated transport withdraws nothing");
assert!(
File diff suppressed because it is too large Load Diff
+20 -5
View File
@@ -1717,8 +1717,14 @@ fn test_seam_link_cost_etx_orders_bloom_candidates() {
seam_set_cost(&mut node, &low, 3.0, 1_000);
seam_set_cost(&mut node, &high, 1.0, 1_000);
let cost_low = node.get_peer(&low).unwrap().link_cost();
let cost_high = node.get_peer(&high).unwrap().link_cost();
let cost_low = node
.get_peer(&low)
.unwrap()
.link_cost(crate::time::mono_ms());
let cost_high = node
.get_peer(&high)
.unwrap()
.link_cost(crate::time::mono_ms());
assert!(
cost_high < cost_low,
"fixture: ETX alone must make high cheaper ({cost_high} vs {cost_low})"
@@ -1749,8 +1755,14 @@ fn test_seam_link_cost_srtt_orders_bloom_candidates() {
seam_set_cost(&mut node, &low, 1.0, 50_000); // 50 ms -> cost 1.5
seam_set_cost(&mut node, &high, 1.0, 1_000); // 1 ms -> cost 1.01
let cost_low = node.get_peer(&low).unwrap().link_cost();
let cost_high = node.get_peer(&high).unwrap().link_cost();
let cost_low = node
.get_peer(&low)
.unwrap()
.link_cost(crate::time::mono_ms());
let cost_high = node
.get_peer(&high)
.unwrap()
.link_cost(crate::time::mono_ms());
assert!(
cost_high < cost_low,
"fixture: SRTT alone must make high cheaper ({cost_high} vs {cost_low})"
@@ -1951,7 +1963,10 @@ fn test_seam_routing_view_reads_match_live_peer_state() {
let addr = view.peer_addr(*peer);
let live = node.peers.get(&addr).unwrap();
assert_eq!(view.peer_may_reach(*peer, &dest), live.may_reach(&dest));
assert_eq!(view.peer_link_cost(*peer), live.link_cost());
assert_eq!(
view.peer_link_cost(*peer),
live.link_cost(crate::time::mono_ms())
);
assert_eq!(
view.peer_coords(*peer),
node.tree_state().peer_coords(&addr)
+2 -2
View File
@@ -19,13 +19,13 @@ static LARGE_NETWORK_TEST_LOCK: std::sync::LazyLock<tokio::sync::Mutex<()>> =
/// address. Each node gets a unique synthetic address (`loopback:{n}`) from
/// `LOOPBACK_ADDR_COUNTER`, so addresses never collide across concurrently
/// running tests and stale entries from finished tests are harmless.
static LOOPBACK_REGISTRY: std::sync::LazyLock<LoopbackRegistry> =
pub(super) static LOOPBACK_REGISTRY: std::sync::LazyLock<LoopbackRegistry> =
std::sync::LazyLock::new(new_registry);
static LOOPBACK_ADDR_COUNTER: std::sync::atomic::AtomicU64 = std::sync::atomic::AtomicU64::new(0);
/// Allocate the next globally-unique loopback address.
fn next_loopback_addr() -> TransportAddr {
pub(super) fn next_loopback_addr() -> TransportAddr {
let n = LOOPBACK_ADDR_COUNTER.fetch_add(1, std::sync::atomic::Ordering::Relaxed);
TransportAddr::from_string(&format!("loopback:{}", n))
}
+6 -3
View File
@@ -329,11 +329,12 @@ impl Node {
// Re-evaluate parent selection with current link costs.
// Exclude peers without MMP RTT data — they are not yet eligible
// as parent candidates (prevents oscillation from optimistic defaults).
let now_ms = crate::time::mono_ms();
let peer_costs: BTreeMap<NodeAddr, f64> = self
.peers
.iter()
.filter(|(_, peer)| peer.has_srtt())
.map(|(addr, peer)| (*addr, peer.link_cost()))
.map(|(addr, peer)| (*addr, peer.link_cost(now_ms)))
.collect();
// No peers are excluded from parent candidacy on this branch; the
// non-full/leaf skip is a next-only shell refinement.
@@ -600,11 +601,12 @@ impl Node {
self.last_parent_reeval = Some(now);
let now_ms = crate::time::mono_ms();
let peer_costs: BTreeMap<NodeAddr, f64> = self
.peers
.iter()
.filter(|(_, peer)| peer.has_srtt())
.map(|(addr, peer)| (*addr, peer.link_cost()))
.map(|(addr, peer)| (*addr, peer.link_cost(now_ms)))
.collect();
// No peers are excluded from parent candidacy on this branch; the
// non-full/leaf skip is a next-only shell refinement.
@@ -752,11 +754,12 @@ impl Node {
// before the recovery mutation — exactly as before.
self.metrics().tree.parent_losses.inc();
let now_ms = crate::time::mono_ms();
let peer_costs: BTreeMap<NodeAddr, f64> = self
.peers
.iter()
.filter(|(_, peer)| peer.has_srtt())
.map(|(addr, peer)| (*addr, peer.link_cost()))
.map(|(addr, peer)| (*addr, peer.link_cost(now_ms)))
.collect();
// Wall-clock seconds stamped onto the new declaration; monotonic ms for
+986 -13
View File
File diff suppressed because it is too large Load Diff
+4 -1
View File
@@ -8,7 +8,10 @@
mod active;
pub(crate) mod machine;
pub use active::{ActivePeer, ConnectivityState, PeerPath};
pub use active::{
ActivePeer, ConnectivityState, HeartbeatPlan, HeartbeatSend, HeartbeatTiming, PathPolicy,
PathState, PathSwitch, PathWithdrawal, PeerPath, SwitchReason,
};
use crate::NodeAddr;
use crate::transport::LinkId;
+167
View File
@@ -47,6 +47,16 @@ pub enum LinkMessageType {
/// Periodic heartbeat for link liveness detection.
/// No payload — the msg_type byte alone is sufficient.
Heartbeat = 0x51,
/// Probe a candidate path to a peer under the existing session.
/// Payload is a [`PathMessage`].
PathProbe = 0x52,
/// Answer to a [`LinkMessageType::PathProbe`], sent back on the path
/// the probe arrived on. Payload is a [`PathMessage`].
PathAck = 0x53,
/// "I am closing this path": sent on any other path when the sender
/// knows a path is going (interface gone, carrier lost). Payload is a
/// [`PathClose`].
PathClose = 0x54,
}
impl LinkMessageType {
@@ -62,6 +72,9 @@ impl LinkMessageType {
0x31 => Some(LinkMessageType::LookupResponse),
0x50 => Some(LinkMessageType::Disconnect),
0x51 => Some(LinkMessageType::Heartbeat),
0x52 => Some(LinkMessageType::PathProbe),
0x53 => Some(LinkMessageType::PathAck),
0x54 => Some(LinkMessageType::PathClose),
_ => None,
}
}
@@ -84,11 +97,165 @@ impl fmt::Display for LinkMessageType {
LinkMessageType::LookupResponse => "LookupResponse",
LinkMessageType::Disconnect => "Disconnect",
LinkMessageType::Heartbeat => "Heartbeat",
LinkMessageType::PathProbe => "PathProbe",
LinkMessageType::PathAck => "PathAck",
LinkMessageType::PathClose => "PathClose",
};
write!(f, "{}", name)
}
}
// ============================================================================
// Path Probe / Path Ack
// ============================================================================
/// Payload shared by `PathProbe` (0x52) and `PathAck` (0x53).
///
/// A probe is an ordinary encrypted frame under the current session, sent on
/// a candidate transport. The receiver, having decrypted it against the
/// session found by index, has proof the peer is reachable there: it adds
/// the path and answers with an ack **on that same path**. The prober's
/// receipt of the ack proves the reverse direction. One round trip, no
/// handshake, no new key material, no index allocation.
///
/// ## Wire Format
///
/// | Offset | Field | Size | Notes |
/// |--------|---------------|---------|-----------------------------------------|
/// | 0 | msg_type | 1 byte | 0x52 or 0x53 |
/// | 1 | probe_id | 4 bytes | LE; the ack echoes the probe's |
/// | 5 | flags | 1 byte | bit 0: `remote_active` |
/// | 6 | path_id | 4 bytes | LE; the sender's id for its path |
/// | 10 | padding | any | ignored; a full-size probe pads to MTU |
///
/// Trailing bytes are ignored on decode, so a probe may be padded to the
/// link MTU: a path that forwards small frames and drops large ones then
/// never proves itself.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub struct PathMessage {
/// Per-path sequence chosen by the prober; the ack carries it back.
pub probe_id: u32,
/// "This path is where I currently send." Costs one bit and is a free
/// detection signal: a peer that stops sending here may have stopped
/// hearing us here too.
pub remote_active: bool,
/// The sender's own identifier for the path this travels on. Transport
/// ids are local to each node; each side learns the other's id from its
/// probes and acks, and a later [`PathClose`] names the path by the
/// *receiver's* id, so the receiver needs no lookup table.
pub path_id: u32,
}
impl PathMessage {
/// Encoded size including the msg_type byte.
pub const WIRE_SIZE: usize = 10;
const FLAG_REMOTE_ACTIVE: u8 = 0x01;
/// Encode as a `PathProbe` link message (msg_type included).
pub fn encode_probe(&self) -> [u8; Self::WIRE_SIZE] {
self.encode(LinkMessageType::PathProbe)
}
/// Encode as a `PathAck` link message (msg_type included).
pub fn encode_ack(&self) -> [u8; Self::WIRE_SIZE] {
self.encode(LinkMessageType::PathAck)
}
fn encode(&self, kind: LinkMessageType) -> [u8; Self::WIRE_SIZE] {
let mut out = [0u8; Self::WIRE_SIZE];
out[0] = kind.to_byte();
out[1..5].copy_from_slice(&self.probe_id.to_le_bytes());
if self.remote_active {
out[5] |= Self::FLAG_REMOTE_ACTIVE;
}
out[6..10].copy_from_slice(&self.path_id.to_le_bytes());
out
}
/// Decode from the link-layer payload (after the msg_type byte).
pub fn decode(payload: &[u8]) -> Result<Self, Error> {
let mut reader = crate::proto::codec::Reader::new(payload);
let probe_id = reader.read_u32_le()?;
let flags = reader.read_u8()?;
let path_id = reader.read_u32_le()?;
Ok(Self {
probe_id,
remote_active: flags & Self::FLAG_REMOTE_ACTIVE != 0,
path_id,
})
}
}
/// Why a path is being closed.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
#[repr(u8)]
pub enum PathCloseReason {
/// No particular reason given.
Unspecified = 0,
/// The interface under the path went away.
InterfaceGone = 1,
/// The interface lost carrier.
CarrierLost = 2,
/// An operator asked for it.
Operator = 3,
}
impl PathCloseReason {
fn from_byte(b: u8) -> Self {
match b {
1 => Self::InterfaceGone,
2 => Self::CarrierLost,
3 => Self::Operator,
_ => Self::Unspecified,
}
}
}
/// `PathClose` (0x54): the sender is closing the path it identifies.
///
/// Sent on any path that still works, so the peer learns at once rather
/// than after an echo timeout. The peer withdraws its side of that path
/// (keeping its history) and moves its traffic if it was on it. Advisory:
/// a later probe on the path revives it.
///
/// ## Wire Format
///
/// | Offset | Field | Size | Notes |
/// |--------|----------|---------|-------------------------------------------|
/// | 0 | msg_type | 1 byte | 0x54 |
/// | 1 | path_id | 4 bytes | LE; the *receiver's* id for the path |
/// | 5 | reason | 1 byte | [`PathCloseReason`] |
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub struct PathClose {
/// The receiver's identifier for the path, as it carried in its own
/// probes and acks on it.
pub path_id: u32,
pub reason: PathCloseReason,
}
impl PathClose {
/// Encoded size including the msg_type byte.
pub const WIRE_SIZE: usize = 6;
/// Encode as a link message (msg_type included).
pub fn encode(&self) -> [u8; Self::WIRE_SIZE] {
let mut out = [0u8; Self::WIRE_SIZE];
out[0] = LinkMessageType::PathClose.to_byte();
out[1..5].copy_from_slice(&self.path_id.to_le_bytes());
out[5] = self.reason as u8;
out
}
/// Decode from the link-layer payload (after the msg_type byte).
pub fn decode(payload: &[u8]) -> Result<Self, Error> {
let mut reader = crate::proto::codec::Reader::new(payload);
let path_id = reader.read_u32_le()?;
let reason = PathCloseReason::from_byte(reader.read_u8()?);
Ok(Self { path_id, reason })
}
}
// ============================================================================
// Session Datagram (Link-Layer Encapsulation)
// ============================================================================
+4
View File
@@ -790,6 +790,10 @@ impl<I: BleIo> BleTransport<I> {
}
impl<I: BleIo> Transport for BleTransport<I> {
fn role(&self) -> crate::config::TransportRole {
self.config.role()
}
fn transport_id(&self) -> TransportId {
self.transport_id
}
+7
View File
@@ -553,6 +553,10 @@ impl Drop for EthernetTransport {
}
impl Transport for EthernetTransport {
fn role(&self) -> crate::config::TransportRole {
self.config.role()
}
fn transport_id(&self) -> TransportId {
self.transport_id
}
@@ -1637,6 +1641,7 @@ mod tests {
// A name no host has. `fips` is not a valid netdev prefix anywhere and
// the suffix keeps it clear of the test harness's own veth pairs.
let config = EthernetConfig {
role: None,
interface: "fips-absent-x0".to_string(),
ethertype: None,
mtu: None,
@@ -1821,6 +1826,7 @@ mod tests {
let can_open = PacketSocket::open(loopback, 0x2121).is_ok();
let config = EthernetConfig {
role: None,
interface: loopback.to_string(),
ethertype: None,
mtu: None,
@@ -2401,6 +2407,7 @@ mod tests {
}
let config = EthernetConfig {
role: None,
interface: iface.to_string(),
ethertype: None,
mtu: None,
+34 -1
View File
@@ -50,6 +50,13 @@ pub struct LoopbackTransport {
mtu: u16,
/// Shared address-to-receiver registry.
registry: LoopbackRegistry,
/// Beacons a test has queued for the next `discover()` drain, standing
/// in for the transport-neighbor beacons a real transport hears.
discovered: Mutex<Vec<DiscoveredPeer>>,
/// Carrier a test has set, standing in for an interface-bound
/// transport's `IFF_RUNNING`. `None`: not interface-bound, no presence
/// reported, which is the default.
carrier: Mutex<Option<bool>>,
}
impl LoopbackTransport {
@@ -75,9 +82,30 @@ impl LoopbackTransport {
my_addr,
mtu,
registry,
discovered: Mutex::new(Vec::new()),
carrier: Mutex::new(None),
}
}
/// Pretend this transport is bound to an interface with (or without)
/// carrier; `None` returns it to reporting no presence at all.
pub fn set_carrier(&self, carrier: Option<bool>) {
*self.carrier.lock().unwrap() = carrier;
}
/// The carrier a test set, if any.
pub fn carrier(&self) -> Option<bool> {
*self.carrier.lock().unwrap()
}
/// Queue a beacon for the next `discover()` drain, as if this transport
/// had heard `addr` announce `pubkey_hint`.
pub fn inject_discovered(&self, addr: TransportAddr, pubkey_hint: secp256k1::XOnlyPublicKey) {
let mut peer = DiscoveredPeer::new(self.transport_id, addr);
peer.pubkey_hint = Some(pubkey_hint);
self.discovered.lock().unwrap().push(peer);
}
/// This transport's synthetic loopback address.
pub fn my_addr(&self) -> &TransportAddr {
&self.my_addr
@@ -169,6 +197,11 @@ impl Transport for LoopbackTransport {
}
fn discover(&self) -> Result<Vec<DiscoveredPeer>, TransportError> {
Ok(Vec::new())
Ok(std::mem::take(&mut *self.discovered.lock().unwrap()))
}
/// Beacons a test injects are meant to be acted on.
fn auto_connect(&self) -> bool {
true
}
}
+54
View File
@@ -268,6 +268,20 @@ impl TransportError {
/// which is a statement about the peer rather than about this node's
/// ability to transmit, and the existing retry paths for them already sit
/// at a different layer.
/// Whether the kernel refused the send for want of a route: the
/// interface is up but nothing is reachable through it. A hard signal
/// that the path is gone (`ENETUNREACH`, `EHOSTUNREACH`), distinct from
/// `is_transient`: the binder is not going to fix this.
pub fn is_unreachable(&self) -> bool {
match self {
Self::Io(e) => matches!(
e.kind(),
std::io::ErrorKind::NetworkUnreachable | std::io::ErrorKind::HostUnreachable
),
_ => false,
}
}
pub fn is_transient(&self) -> bool {
match self {
// The interface is absent or mid-rebind. The binder is polling for
@@ -565,6 +579,15 @@ impl Link {
link
}
/// Point the link at another transport and address.
///
/// A peer whose active path switched keeps its link (the control machine
/// is keyed on it); the record follows the traffic.
pub fn rebind(&mut self, transport_id: TransportId, remote_addr: TransportAddr) {
self.transport_id = transport_id;
self.remote_addr = remote_addr;
}
/// Get the link ID.
pub fn link_id(&self) -> LinkId {
self.link_id
@@ -748,6 +771,12 @@ pub trait Transport {
true
}
/// The transport's path-selection role. Default: normal. Concrete
/// transports read from their own config.
fn role(&self) -> crate::config::TransportRole {
crate::config::TransportRole::Normal
}
/// Close a specific connection (connection-oriented transports only).
///
/// For connectionless transports (UDP, Ethernet), this is a no-op.
@@ -1026,6 +1055,15 @@ impl TransportHandle {
failed_attempts: state.attempts(),
})
}
#[cfg(test)]
TransportHandle::Loopback(t) => t.carrier().map(|carrier| InterfacePresence {
presence: "present",
carrier,
policy: "optional",
since_secs: 0,
binds: 1,
failed_attempts: 0,
}),
_ => None,
}
}
@@ -1107,6 +1145,22 @@ impl TransportHandle {
}
}
/// The transport's path-selection role.
pub fn role(&self) -> crate::config::TransportRole {
match self {
TransportHandle::Udp(t) => t.role(),
#[cfg(any(target_os = "linux", target_os = "macos"))]
TransportHandle::Ethernet(t) => t.role(),
TransportHandle::Tcp(t) => t.role(),
TransportHandle::Tor(t) => t.role(),
TransportHandle::Nym(t) => t.role(),
#[cfg(ble_available)]
TransportHandle::Ble(t) => t.role(),
#[cfg(test)]
TransportHandle::Loopback(t) => t.role(),
}
}
/// Whether this transport auto-connects to discovered peers.
pub fn auto_connect(&self) -> bool {
match self {
+4
View File
@@ -611,6 +611,10 @@ impl NymTransport {
}
impl Transport for NymTransport {
fn role(&self) -> crate::config::TransportRole {
self.config.role()
}
fn transport_id(&self) -> TransportId {
self.transport_id
}
+4
View File
@@ -786,6 +786,10 @@ impl TcpTransport {
}
impl Transport for TcpTransport {
fn role(&self) -> crate::config::TransportRole {
self.config.role()
}
fn transport_id(&self) -> TransportId {
self.transport_id
}
+4
View File
@@ -1047,6 +1047,10 @@ impl TorTransport {
}
impl Transport for TorTransport {
fn role(&self) -> crate::config::TransportRole {
self.config.role()
}
fn transport_id(&self) -> TransportId {
self.transport_id
}
+4
View File
@@ -437,6 +437,10 @@ impl UdpTransport {
}
impl Transport for UdpTransport {
fn role(&self) -> crate::config::TransportRole {
self.config.role()
}
fn transport_id(&self) -> TransportId {
self.transport_id
}
+14
View File
@@ -101,11 +101,25 @@ Explicit topologies exercising non-UDP transports.
| ethernet-only | 4 | Ethernet | Ring | 30s | yes | -- | AF_PACKET transport with beacon discovery |
| ethernet-mesh | 6 | UDP + Ethernet | Mesh | 120s | yes | yes | Mixed UDP/Ethernet, netem mutation + flaps |
| tcp-mesh | 6 | UDP + TCP | Mesh | 120s | yes | yes | Mixed UDP/TCP, netem mutation + flaps |
| dual-path-flap | 2 | Ethernet + UDP | Pair | 180s | yes | yes | One session, two paths; cable flaps, traffic moves without re-peering |
| dual-udp-flap | 2 | UDP + UDP | Pair | 180s | yes | yes | Same, all-IP: two interface-bound UDP instances as two paths |
- **ethernet-only**: 4-node ring on raw Ethernet (AF_PACKET). Peers discovered
via beacons, not static config. Minimal netem (1-5ms delay).
- **ethernet-mesh**: Mirrors `tcp-mesh` topology but with Ethernet instead of
TCP. UDP edges use static config; Ethernet edges use beacon discovery.
- **dual-path-flap**: two nodes joined twice, by a raw-Ethernet veth and by
UDP over the bridge (`[n01, n02, ethernet+udp]`). The cable flaps
(`link_flaps.only_transport: ethernet`); traffic must move to the UDP path
under the same session and never re-peer. The calibration scenario for
`node.path.*`; see the file header for what to read from a run.
- **dual-udp-flap**: the all-IP twin (`[n01, n02, udp-veth+udp]`): a veth
carrying IP with an interface-bound UDP instance at each end, plus UDP
over the bridge. The veth half is handed to the daemon by the runner
(`fipsctl connect ... udp/<veth>`) once the pair has peered, and becomes a
path under the existing session. Exercises `udp.interface` on the listen
and per-peer connected sockets, and a configured address on a new
transport becoming a path.
- **tcp-mesh**: 6-node mesh with 4 UDP and 3 TCP edges. Both transports use
static peer config. Netem mutation (30% fraction, every 20-40s) and link
flaps (1 link max, 10-20s down).
+123
View File
@@ -0,0 +1,123 @@
# One peer, two paths: switchover without a second handshake
#
# The calibration scenario for `node.path.*`
# (docs/design/fips-multi-path-switchover.md, "Calibration"). Two nodes joined
# twice: a raw-Ethernet veth (the cable) and the Docker bridge over UDP (the
# wifi). The UDP half is dialled from static config, the Ethernet half is
# found by beacon and added as a second path under the same session by the
# probe exchange. iperf runs across the pair throughout while the cable is
# taken down and brought back.
#
# What "down" means here: `link_flaps` sets 100% loss on the veth, so the
# interface stays IFF_UP and keeps carrier. That is deliberate. Presence and
# carrier are covered by unit tests; the thing only a network can exercise
# is the heartbeat echo timeout — the detector every medium falls back to —
# and the discretionary fail-back when the cable returns.
#
# What to read from a run (sim-results/<run>/analysis.txt and node logs):
# - "Path switched" / "traffic moved to the standby": one per node per
# flap edge, so two flaps give at least four; the ceiling below allows
# one extra per flap for a discretionary fail-back and nothing more.
# - the timestamps between "Link DOWN" in runner.log and the first switch
# on each node: the switch latency. The target is under one second.
# - "Peer promoted to active" more than once on a node: a re-peering,
# which is the failure this scenario exists to catch. `max_promotions`
# fails the run on it; `switch_latency` and `max_stall` fail it on a
# slow switch and on a switch that did not carry traffic.
#
# n01 ===eth=== n02 (flapped)
# n01 ---udp--- n02 (never touched)
scenario:
name: "dual-path-flap"
seed: 42
duration_secs: 180
topology:
algorithm: explicit
num_nodes: 2
default_transport: udp
params:
adjacency:
- [n01, n02, ethernet+udp]
# The cable is mild; the UDP half carries 40 ms each way. Score is
# etx × (1 + min_rtt/100): about 1.0 on the cable against about 1.8 over
# UDP, well past the 1.3 margin, so the pair moves onto the cable soon after
# both paths are live (the UDP static dial lands first) and returns to it
# after each flap once the cable is proven again. Every flap therefore hits
# the *active* path, and the fail-back is exercised too.
netem:
enabled: true
default_policy:
delay_ms: [1, 3]
jitter_ms: [0, 1]
loss_pct: [0, 0.2]
dual_udp_policy:
delay_ms: [40, 40]
jitter_ms: [0, 1]
loss_pct: [0, 0.2]
# The mechanism. Only the cable flaps; the UDP path is the standby that
# has to be warm when it does. `protect_connectivity` is off because the
# harness counts the dual edge once and would refuse to flap the "only"
# link; the UDP half keeps the pair connected throughout.
link_flaps:
enabled: true
only_transport: ethernet
interval_secs: {min: 35, max: 45}
max_down_links: 1
down_duration_secs: {min: 12, max: 18}
protect_connectivity: false
traffic:
enabled: true
max_concurrent: 1
interval_secs: {min: 5, max: 10}
duration_secs: {min: 20, max: 30}
parallel_streams: 2
node_churn:
enabled: false
assertions:
baseline:
min_nodes_reporting: 2
max_roots: 1
min_nodes_parented: 1
# Traffic crossed the pair while the cable was being taken away.
min_traffic:
min_sessions_ok: 2
# Three to four flaps in 180 s at the interval above. Each flap moves both
# nodes off the cable (2) and its return moves them back (2): four per
# flap, plus the initial move onto the cable (2). Floor: the first flap
# moved something. Ceiling: four flaps, both directions, both nodes, the
# initial move, and nothing else.
path_switches:
min_total: 4
max_total: 18
# The detectors that can fail on a switchover that did not carry. A
# promotion past the first is a re-peering: the session was lost to the
# link-dead reaper and rebuilt, which standby probes alone cannot show.
# The switch must land inside a second of the runner taking the link
# down. And no iperf3 session may sit at zero bytes for longer than two
# of its one-second intervals: min_traffic passes on bytes either side
# of a hole, this fails on the hole.
max_promotions:
per_node: 1
switch_latency:
max_ms: 1000
max_stall:
max_secs: 2
max_errors:
max_total: 0
logging:
# The path handler at debug: "Path suspect" and "Sent path probe" are what
# the switch latency is read from.
rust_log: "info,fips::node::handlers::path=debug"
output_dir: "./sim-results"
+121
View File
@@ -0,0 +1,121 @@
# One peer, two UDP paths: switchover between two interface-bound UDP instances
#
# The all-IP twin of dual-path-flap. Two nodes joined twice, both by UDP: a
# dedicated veth carrying IP (`10.222.0.0/24`), on which each node runs a
# UDP instance bound to that interface (`transports.udp.<veth>.interface`),
# and the Docker bridge on the ordinary wildcard instance. Both are static
# peer addresses on the dial owner (`udp/main` first, `udp/<veth>` second),
# so both handshakes run at startup; the first to complete makes the
# session, the second settles as a cross-connection and creates no path
# state — the dialled address is a candidate, and the probe exchange under
# the session proves it as a path at both ends.
#
# What this exercises that dual-path-flap cannot: two UDP instances as two
# paths, `udp.interface` binding on the listen socket *and* on the per-peer
# connected sockets (Linux `SO_BINDTODEVICE`), and a configured address on a
# new transport becoming a path. The flap and fail-back mechanics are the
# same: `link_flaps` sets 100% loss on the veth, the interface stays up, and
# the heartbeat echo timeout is the detector.
#
# What to read from a run: as dual-path-flap. "Path live: the peer answers
# on this transport" in both node logs is the second instance being proven
# as a path; `fipsctl path show` on either node should list both instances
# live before the first flap, and "Peer promoted to active" appears once
# per node (`max_promotions` fails the run otherwise).
#
# n01 ===udp/veth=== n02 (flapped)
# n01 ----udp/main--- n02 (never touched)
scenario:
name: "dual-udp-flap"
seed: 42
duration_secs: 180
topology:
algorithm: explicit
num_nodes: 2
default_transport: udp
params:
adjacency:
- [n01, n02, udp-veth+udp]
# The veth is mild; the bridge half carries 40 ms each way. Score is
# etx × (1 + min_rtt/100): about 1.0 on the veth against about 1.8 over the
# bridge, well past the 1.3 margin, so the pair moves onto the veth soon
# after the runner adds it and returns to it after each flap once it is
# proven again. Every flap therefore hits the *active* path, and the
# fail-back is exercised too.
netem:
enabled: true
default_policy:
delay_ms: [1, 3]
jitter_ms: [0, 1]
loss_pct: [0, 0.2]
dual_udp_policy:
delay_ms: [40, 40]
jitter_ms: [0, 1]
loss_pct: [0, 0.2]
# The mechanism. Only the veth flaps; the bridge path is the standby that
# has to be warm when it does. `protect_connectivity` is off because the
# harness counts the dual edge once and would refuse to flap the "only"
# link; the bridge half keeps the pair connected throughout.
link_flaps:
enabled: true
only_transport: udp-veth
interval_secs: {min: 35, max: 45}
max_down_links: 1
down_duration_secs: {min: 12, max: 18}
protect_connectivity: false
traffic:
enabled: true
max_concurrent: 1
interval_secs: {min: 5, max: 10}
duration_secs: {min: 20, max: 30}
parallel_streams: 2
node_churn:
enabled: false
assertions:
baseline:
min_nodes_reporting: 2
max_roots: 1
min_nodes_parented: 1
# Traffic crossed the pair while the cable was being taken away.
min_traffic:
min_sessions_ok: 2
# Three to four flaps in 180 s at the interval above. Each flap moves both
# nodes off the veth (2) and its return moves them back (2): four per
# flap, plus the initial move onto the veth (2). Floor: the first flap
# moved something. Ceiling: four flaps, both directions, both nodes, the
# initial move, and nothing else.
path_switches:
min_total: 4
max_total: 18
# The detectors that can fail on a switchover that did not carry. A
# promotion past the first is a re-peering: the session was lost to the
# link-dead reaper and rebuilt, which standby probes alone cannot show.
# The switch must land inside a second of the runner taking the link
# down. And no iperf3 session may sit at zero bytes for longer than two
# of its one-second intervals: min_traffic passes on bytes either side
# of a hole, this fails on the hole.
max_promotions:
per_node: 1
switch_latency:
max_ms: 1000
max_stall:
max_secs: 2
max_errors:
max_total: 0
logging:
# The path handler at debug: "Path suspect" and "Sent path probe" are what
# the switch latency is read from.
rust_log: "info,fips::node::handlers::path=debug,fips::node::lifecycle=debug"
output_dir: "./sim-results"
+240
View File
@@ -25,7 +25,10 @@ from .scenario import (
CongestionSignalsAssertion,
MaxErrorsAssertion,
MaxParentSwitchesAssertion,
MaxPromotionsAssertion,
MaxStallAssertion,
MinParentSwitchesAssertion,
SwitchLatencyAssertion,
TreeParentsAssertion,
)
from .topology import SimTopology
@@ -501,3 +504,240 @@ def evaluate_min_traffic(
f"means the data path did not survive what the scenario did to it."
),
)
def evaluate_path_switches(cfg, count: int) -> AssertionOutcome:
"""Band on path switches (traffic moving between transports under one
session) over the run.
``min_total`` catches a harness that flapped a link nothing was
switching over: a dual-path scenario in which no switch happened
tested nothing. ``max_total`` is the stability ceiling: a healthy
dual-path pair should switch only when a link goes and comes back,
never on its own.
"""
if cfg.min_total is not None and count < cfg.min_total:
return AssertionOutcome(
name="path_switches",
passed=False,
detail=(
f"FAIL path_switches: {count} switches < floor {cfg.min_total} "
f"— the flaps did not move traffic between paths. Check that "
f"both paths came up (fipsctl path show) before the first flap."
),
)
if cfg.max_total is not None and count > cfg.max_total:
return AssertionOutcome(
name="path_switches",
passed=False,
detail=(
f"FAIL path_switches: {count} switches > ceiling {cfg.max_total} "
f"— traffic is moving between paths more than the flaps "
f"account for. Look for discretionary switches on a healthy "
f"pair: the margin or the dwell is too small."
),
)
return AssertionOutcome(
name="path_switches",
passed=True,
detail=(
f"PASS path_switches: {count} switches within "
f"[{cfg.min_total if cfg.min_total is not None else 0}, "
f"{cfg.max_total if cfg.max_total is not None else 'inf'}]"
),
)
def evaluate_max_promotions(
cfg: MaxPromotionsAssertion,
promotions: list[tuple[str, str]],
) -> AssertionOutcome:
"""Per-node ceiling on "Peer promoted to active" lines.
``promotions`` is ``AnalysisResult.peers_promoted``: ``(source, line)``
pairs, one per handshake that completed on that node. The first per
peer is the pair meeting; any beyond the ceiling is a re-peering — the
session was torn down and rebuilt — which a switchover scenario exists
to prove does not happen. Liveness is no backstop here: standby probes
keep a peer alive while its data path is blackholed, so this counts
the one event a blackhole that lasts to the reaper cannot avoid.
"""
per_node: dict[str, int] = {}
for source, _line in promotions:
per_node[source] = per_node.get(source, 0) + 1
over = {src: n for src, n in per_node.items() if n > cfg.per_node}
if not over:
return AssertionOutcome(
name="max_promotions",
passed=True,
detail=(
f"PASS max_promotions: every node promoted at most "
f"{cfg.per_node} time(s) ({len(promotions)} total)"
),
)
breakdown = ", ".join(f"{src}={n}" for src, n in sorted(over.items()))
samples = "\n".join(
f" [{src}] {line.strip()}"
for src, line in promotions
if src in over
)
return AssertionOutcome(
name="max_promotions",
passed=False,
detail=(
f"FAIL max_promotions: {breakdown} exceed(s) the per-node ceiling "
f"of {cfg.per_node}. A second promotion is a re-peering: the "
f"session was lost and rebuilt, so a switchover did not carry.\n"
f"{samples}"
),
)
def _line_epoch(line: str) -> float | None:
"""Epoch seconds of a node log line's leading RFC 3339 timestamp, if any.
``tracing`` writes ``2026-09-13T13:39:01.123456Z`` first on every line.
Anything else (a bare stderr line, a runner line) is not timed.
"""
from datetime import datetime, timezone
head = line.strip().split(" ", 1)[0]
if not head.endswith("Z"):
return None
try:
return datetime.fromisoformat(head.replace("Z", "+00:00")).timestamp()
except ValueError:
return None
def evaluate_switch_latency(
cfg: SwitchLatencyAssertion,
flap_events: list[tuple[float, str, str, str]],
switches: list[tuple[str, str]],
) -> AssertionOutcome:
"""Ceiling on the time from each link-down to the first switch on
either endpoint.
``flap_events`` is the link manager's record: ``(epoch, "down" | "up",
a, b)``. ``switches`` is ``AnalysisResult.path_switches``: ``(source,
line)``, where ``source`` is the node id and the line carries its own
timestamp. For each down edge, the latency is the earliest switch line
on ``a`` or ``b`` stamped at or after the down; a down with no switch
inside ``max_ms`` fails. The worst flap is what is reported.
"""
downs = [(t, a, b) for t, kind, a, b in flap_events if kind == "down"]
if not downs:
return AssertionOutcome(
name="switch_latency",
passed=False,
detail="FAIL switch_latency: no link was taken down, nothing measured",
)
timed: list[tuple[str, float]] = []
for source, line in switches:
t = _line_epoch(line)
if t is not None:
timed.append((source, t))
limit = cfg.max_ms / 1000.0
worst: tuple[float, str, str] | None = None
missing: list[str] = []
for down_at, a, b in downs:
after = [
t - down_at
for src, t in timed
if src in (a, b) and t >= down_at and t - down_at <= limit
]
if not after:
from datetime import datetime, timezone
when = datetime.fromtimestamp(down_at, timezone.utc).strftime("%H:%M:%S")
missing.append(f"{a}--{b} down at {when}Z")
continue
latency = min(after)
if worst is None or latency > worst[0]:
worst = (latency, a, b)
if missing:
return AssertionOutcome(
name="switch_latency",
passed=False,
detail=(
f"FAIL switch_latency: {len(missing)} of {len(downs)} link-down(s) "
f"had no path switch on either endpoint within {cfg.max_ms} ms: "
+ "; ".join(missing)
),
)
assert worst is not None
return AssertionOutcome(
name="switch_latency",
passed=True,
detail=(
f"PASS switch_latency: worst {worst[0] * 1000:.0f} ms "
f"({worst[1]}--{worst[2]}) over {len(downs)} link-down(s), "
f"ceiling {cfg.max_ms} ms"
),
)
def _longest_stall_secs(result: dict) -> float:
"""Longest run of consecutive zero-byte intervals in one iperf3 result,
in seconds. 0 for a result with no intervals."""
intervals = result.get("intervals") if isinstance(result, dict) else None
if not isinstance(intervals, list):
return 0.0
longest = 0.0
run = 0.0
for iv in intervals:
summary = iv.get("sum") if isinstance(iv, dict) else None
if not isinstance(summary, dict):
continue
seconds = summary.get("seconds", 1.0)
if not isinstance(seconds, (int, float)) or seconds <= 0:
seconds = 1.0
if summary.get("bytes", 0) == 0:
run += seconds
longest = max(longest, run)
else:
run = 0.0
return longest
def evaluate_max_stall(
cfg: MaxStallAssertion,
results: list[dict],
) -> AssertionOutcome:
"""Ceiling on the longest zero-byte run inside any iperf3 session.
``min_traffic`` cannot see a hole: a session that stalls for ten
seconds mid-run still moves bytes before and after. This reads the
per-interval totals iperf3 records and fails on the longest run of
zeros across every session, which is the stall a switchover leaves
when it does not carry.
"""
if not results:
return AssertionOutcome(
name="max_stall",
passed=False,
detail="FAIL max_stall: no iperf3 session ran, nothing measured",
)
stalls = [(_longest_stall_secs(r), r) for r in results]
worst_secs, worst = max(stalls, key=lambda pair: pair[0])
if worst_secs <= cfg.max_secs:
return AssertionOutcome(
name="max_stall",
passed=True,
detail=(
f"PASS max_stall: longest zero-byte run {worst_secs:.0f} s "
f"across {len(results)} session(s), ceiling {cfg.max_secs:g} s"
),
)
start = worst.get("start", {}) if isinstance(worst, dict) else {}
when = start.get("timestamp", {}).get("time", "?") if isinstance(start, dict) else "?"
return AssertionOutcome(
name="max_stall",
passed=False,
detail=(
f"FAIL max_stall: a session starting {when} moved nothing for "
f"{worst_secs:.0f} s, over the ceiling of {cfg.max_secs:g} s. Bytes "
f"either side of the hole satisfied min_traffic; the hole is a "
f"switchover that did not carry."
),
)
+51 -5
View File
@@ -7,7 +7,7 @@ from copy import deepcopy
import yaml
from .topology import SimTopology
from .topology import UDP_VETH, SimTopology
def _deep_merge(base: dict, override: dict) -> dict:
@@ -33,6 +33,7 @@ def _load_template() -> str:
_TRANSPORT_PORTS = {
"udp": 2121,
"udp-veth": 2122,
"tcp": 443,
}
@@ -49,16 +50,32 @@ def generate_peers_block(
if not outbound_peers:
return " []"
# With interface-bound UDP instances the UDP transport is named, and a
# bare ``udp`` address would resolve to the lowest instance id, which
# may be the one bound to a veth: qualify the bridge half.
bridge = "udp/main" if topology.udp_veth_links(node_id) else "udp"
lines = []
for peer_id in sorted(outbound_peers):
peer = topology.nodes[peer_id]
transport = topology.transport_for_edge(node_id, peer_id)
port = _TRANSPORT_PORTS.get(transport, 2121)
if topology.is_dual_udp_edge(node_id, peer_id):
# The veth half of a dual edge is found by beacon (Ethernet) or
# added as a path by the runner after the pair has peered
# (udp-veth); the bridge half is dialled from here.
transport = "udp"
if transport == UDP_VETH:
link = next(l for l in topology.udp_veth_links(node_id) if l.peer_id == peer_id)
transport = f"udp/{link.instance}"
addr = link.peer_addr
else:
addr = f"{peer.docker_ip}:{_TRANSPORT_PORTS.get(transport, 2121)}"
if transport == "udp":
transport = bridge
lines.append(f' - npub: "{peer.npub}"')
lines.append(f' alias: "{peer_id}"')
lines.append(f" addresses:")
lines.append(f" - transport: {transport}")
lines.append(f' addr: "{peer.docker_ip}:{port}"')
lines.append(f' addr: "{addr}"')
lines.append(f" connect_policy: auto_connect")
return "\n".join(lines)
@@ -110,6 +127,28 @@ def _inject_ethernet_transports(parsed: dict, eth_ifaces: list[str]):
}
def _inject_udp_instances(parsed: dict, topology: SimTopology, node_id: str, has_udp: bool):
"""Turn the template's single UDP transport into named instances: ``main``
(the bridge, kept only if the node has bridge-UDP peers) plus one
interface-bound instance per ``udp-veth`` edge, named after its veth.
"""
links = topology.udp_veth_links(node_id)
if not links:
return
transports = parsed.setdefault("transports", {})
main = transports.pop("udp", None) or {"bind_addr": "0.0.0.0:2121"}
instances = {}
if has_udp:
instances["main"] = main
for link in links:
instances[link.instance] = {
"bind_addr": f"0.0.0.0:{_TRANSPORT_PORTS['udp-veth']}",
"interface": link.iface,
"mtu": main.get("mtu", 1472),
}
transports["udp"] = instances
def _inject_tcp_transport(parsed: dict):
"""Inject TCP transport config into a parsed FIPS config."""
transports = parsed.setdefault("transports", {})
@@ -152,9 +191,12 @@ def generate_node_config(
eth_ifaces = topology.ethernet_interfaces(node_id)
has_tcp = bool(topology.tcp_peers(node_id))
has_udp = _has_transport_peers(topology, node_id, "udp")
has_udp_veth = bool(topology.udp_veth_links(node_id))
# Inject non-UDP transport configs and handle pure-transport nodes
needs_yaml_rewrite = eth_ifaces or has_tcp or not has_udp or fips_overrides
needs_yaml_rewrite = (
eth_ifaces or has_tcp or has_udp_veth or not has_udp or fips_overrides
)
if needs_yaml_rewrite:
parsed = yaml.safe_load(config)
@@ -164,7 +206,9 @@ def generate_node_config(
_inject_ethernet_transports(parsed, eth_ifaces)
if has_tcp:
_inject_tcp_transport(parsed)
if not has_udp:
if has_udp_veth:
_inject_udp_instances(parsed, topology, node_id, has_udp)
elif not has_udp:
# No UDP edges: remove UDP transport
transports = parsed.get("transports", {})
transports.pop("udp", None)
@@ -179,6 +223,8 @@ def _has_transport_peers(topology: SimTopology, node_id: str, transport: str) ->
edge = (min(node_id, peer_id), max(node_id, peer_id))
if topology.edge_transport.get(edge, "udp") == transport:
return True
if transport == "udp" and edge in topology.dual_udp_edges:
return True
return False
+12 -1
View File
@@ -52,6 +52,11 @@ class LinkManager:
self.link_states: dict[tuple[str, str], LinkState] = {
edge: LinkState(edge=edge) for edge in topology.edges
}
# Every edge taken down or restored, as (epoch seconds, "down" |
# "up", a, b). The switch-latency assertion reads the down edges
# against the nodes' own log timestamps, so this is wall-clock time
# (the containers share the host clock).
self.flap_events: list[tuple[float, str, str, str]] = []
@property
def down_count(self) -> int:
@@ -68,6 +73,10 @@ class LinkManager:
up_links = [
e for e, ls in self.link_states.items()
if not ls.is_down and e[0] not in down and e[1] not in down
and (
self.config.only_transport is None
or self.topology.transport_for_edge(*e) == self.config.only_transport
)
]
if not up_links:
return
@@ -115,6 +124,7 @@ class LinkManager:
state.is_down = True
state.down_since = now
state.restore_at = now + duration
self.flap_events.append((now, "down", a, b))
log.info("Link DOWN: %s -- %s (restore in %.0fs)", a, b, duration)
@@ -130,6 +140,7 @@ class LinkManager:
self._set_held(b, a, False)
down_for = time.time() - state.down_since if state.down_since else 0
self.flap_events.append((time.time(), "up", a, b))
state.is_down = False
state.down_since = None
state.restore_at = None
@@ -155,7 +166,7 @@ class LinkManager:
container = self.topology.container_name(src_node)
transport = self.topology.transport_for_edge(src_node, dst_node)
if transport == "ethernet":
if self.topology.is_veth_transport(transport):
iface = veth_interface_name(src_node, dst_node)
state = self.netem_mgr.veth_states.get(container, {}).get(iface)
if state is None:
+16 -4
View File
@@ -210,8 +210,12 @@ class NetemManager:
eth_peers = []
for peer_id in sorted(node.peers):
transport = self.topology.transport_for_edge(node_id, peer_id)
if transport == "ethernet":
if self.topology.is_veth_transport(transport):
eth_peers.append(peer_id)
# A dual edge also has a UDP half over the bridge, which
# gets its own HTB class like any IP peer.
if self.topology.is_dual_udp_edge(node_id, peer_id):
ip_peers[peer_id] = self.topology.nodes[peer_id].docker_ip
else:
ip_peers[peer_id] = self.topology.nodes[peer_id].docker_ip
@@ -232,6 +236,11 @@ class NetemManager:
netem_handle = f"{idx + 10}:"
policy = self._policy_for_edge(node_id, peer_id)
if (
self.topology.is_dual_udp_edge(node_id, peer_id)
and self.config.dual_udp_policy is not None
):
policy = self.config.dual_udp_policy
params = self._sample_policy(policy)
rate = self._htb_rate(node_id, peer_id)
@@ -399,7 +408,9 @@ class NetemManager:
for peer_id in sorted(self.topology.nodes[node_id].peers):
if peer_id in self.down_nodes:
continue
if self.topology.transport_for_edge(node_id, peer_id) != "ethernet":
if not self.topology.is_veth_transport(
self.topology.transport_for_edge(node_id, peer_id)
):
continue
peer_container = self.topology.container_name(peer_id)
state = self.veth_states.get(peer_container, {}).get(
@@ -477,8 +488,9 @@ class NetemManager:
self.down_nodes.add(src)
continue
if transport == "ethernet":
# Ethernet: simple netem replace on veth
if self.topology.is_veth_transport(transport):
# Ethernet, or a UDP instance bound to a veth: simple netem
# replace on the veth
iface = veth_interface_name(src, dst)
veth_states = self.veth_states.get(container, {})
state = veth_states.get(iface)
+107 -5
View File
@@ -20,6 +20,10 @@ from .assertions import (
evaluate_max_errors,
evaluate_max_parent_switches,
evaluate_min_parent_switches,
evaluate_path_switches,
evaluate_max_promotions,
evaluate_switch_latency,
evaluate_max_stall,
evaluate_min_traffic,
evaluate_tree_parents,
)
@@ -318,9 +322,9 @@ class SimRunner:
# The entrypoint script waits for configured Ethernet interfaces
# to appear before starting FIPS, so we just need to create the
# veth pairs promptly after containers are running.
if self.topology.has_ethernet():
if self.topology.has_veth():
self.veth_mgr = VethManager(self.topology)
log.info("Setting up Ethernet veth pairs...")
log.info("Setting up veth pairs...")
self.veth_mgr.setup_all()
# 7. Initialize managers
@@ -385,6 +389,11 @@ class SimRunner:
self._sleep(wait)
self._take_snapshot("warmup")
# The veth half of every udp-veth+udp edge: UDP has no beacon, so
# the runner hands the daemon the address once the pair has peered
# over the bridge, and it becomes a path under that session.
self._add_udp_veth_paths()
# Populate npub cache after convergence (nodes must be running)
if self.peer_churn_mgr:
self.peer_churn_mgr.refresh_all_npubs()
@@ -395,13 +404,67 @@ class SimRunner:
if self.link_swap_mgr:
self.link_swap_mgr.setup_initial()
def _add_udp_veth_paths(self, only_node: str | None = None):
"""Give each dual udp-veth edge its veth path.
Sent from the edge's dial owner (the side whose static config holds
the bridge address) as a control-socket ``connect`` naming the
interface-bound instance: to a peer it already holds a session with,
the daemon adds that as a path rather than dialling. Waits for the
bridge session first, so the command cannot become the first dial.
"""
from .control import send_command
outbound = self.topology.directed_outbound()
for node_id in sorted(self.topology.nodes):
if only_node is not None and node_id != only_node:
continue
for link in self.topology.udp_veth_links(node_id):
if not self.topology.is_dual_udp_edge(node_id, link.peer_id):
continue
if link.peer_id not in outbound.get(node_id, []):
continue
if node_id in self._down_nodes or link.peer_id in self._down_nodes:
continue
container = self.topology.container_name(node_id)
npub = self.topology.nodes[link.peer_id].npub
params = {
"npub": npub,
"address": link.peer_addr,
"transport": f"udp/{link.instance}",
}
added = None
for _ in range(30):
if send_command(container, "path_show", {"npub": npub}) is None:
time.sleep(1) # not peered over the bridge yet
continue
added = send_command(container, "connect", params)
break
if added is None:
log.warning(
"udp-veth path %s -> %s via %s not added",
node_id, link.peer_id, link.instance,
)
else:
log.info(
"udp-veth path %s -> %s via %s (%s)",
node_id, link.peer_id, link.instance, link.peer_addr,
)
def _handle_node_restart(self, node_id: str):
"""Called after a node container is restarted.
For ephemeral identity nodes, waits briefly for the daemon to
start, then queries its new npub and updates the peer churn
manager's cache.
Re-adds the node's udp-veth paths once it has re-peered, and for
ephemeral identity nodes waits briefly for the daemon to start,
then queries its new npub and updates the peer churn manager's
cache.
"""
if any(
self.topology.is_dual_udp_edge(node_id, link.peer_id)
for link in self.topology.udp_veth_links(node_id)
):
time.sleep(2)
self._add_udp_veth_paths(only_node=node_id)
if not self.peer_churn_mgr:
return
if node_id not in self.peer_churn_mgr.ephemeral_nodes:
@@ -769,6 +832,45 @@ class SimRunner:
else:
log.error("%s", outcome.detail)
ps_cfg = self.scenario.assertions.path_switches
if ps_cfg is not None:
outcome = evaluate_path_switches(ps_cfg, len(result.path_switches))
self.assertion_outcomes.append(outcome)
if outcome.passed:
log.info("%s", outcome.detail)
else:
log.error("%s", outcome.detail)
mp_cfg = self.scenario.assertions.max_promotions
if mp_cfg is not None:
outcome = evaluate_max_promotions(mp_cfg, result.peers_promoted)
self.assertion_outcomes.append(outcome)
if outcome.passed:
log.info("%s", outcome.detail)
else:
log.error("%s", outcome.detail)
sl_cfg = self.scenario.assertions.switch_latency
if sl_cfg is not None:
flap_events = self.link_mgr.flap_events if self.link_mgr else []
outcome = evaluate_switch_latency(
sl_cfg, flap_events, result.path_switches
)
self.assertion_outcomes.append(outcome)
if outcome.passed:
log.info("%s", outcome.detail)
else:
log.error("%s", outcome.detail)
stall_cfg = self.scenario.assertions.max_stall
if stall_cfg is not None:
outcome = evaluate_max_stall(stall_cfg, iperf_results)
self.assertion_outcomes.append(outcome)
if outcome.passed:
log.info("%s", outcome.detail)
else:
log.error("%s", outcome.detail)
# Applied to every scenario by default, so this is the one
# assertion that is present even when the YAML declares no
# assertions block at all.
+136 -7
View File
@@ -28,7 +28,10 @@ class Range:
raise ValueError(f"{name}: min ({self.min}) must be >= 0")
VALID_TRANSPORTS = ("udp", "ethernet", "tcp")
VALID_TRANSPORTS = ("udp", "ethernet", "tcp", "udp-veth")
# Dual edges: a veth half and a UDP-over-the-bridge half between the same
# two nodes (see topology.py).
DUAL_TRANSPORTS = ("ethernet+udp", "udp-veth+udp")
@dataclass
@@ -100,6 +103,10 @@ class NetemConfig:
default_policy: NetemPolicy = field(default_factory=NetemPolicy)
link_policies: list[LinkPolicyOverride] = field(default_factory=list)
mutation: NetemMutationConfig = field(default_factory=NetemMutationConfig)
# Policy for the UDP half of ``ethernet+udp`` dual edges. The Ethernet
# half takes the edge's ordinary policy. ``None`` means the UDP half
# gets the default policy too.
dual_udp_policy: NetemPolicy | None = None
@dataclass
@@ -109,6 +116,9 @@ class LinkFlapsConfig:
max_down_links: int = 2
down_duration_secs: Range = field(default_factory=lambda: Range(10, 30))
protect_connectivity: bool = True
# Only flap edges of this transport type (``ethernet``, ``udp``, ...).
# ``None`` flaps any edge.
only_transport: str | None = None
@dataclass
@@ -226,6 +236,62 @@ class MinParentSwitchesAssertion:
min_total: int = 1
@dataclass
class PathSwitchesAssertion:
"""Band on path switches: a peer's traffic moving to another
transport under the same session. ``min_total`` proves the flaps moved
something; ``max_total`` is the stability ceiling on a pair that should
switch only when a link goes and comes back.
"""
min_total: int | None = None
max_total: int | None = None
@dataclass
class MaxPromotionsAssertion:
"""Per-node ceiling on "Peer promoted to active".
A promotion is a handshake completing: the first one per peer is how a
pair meets, every later one is a re-peering — the session was lost and
rebuilt, which is the failure a switchover scenario exists to catch.
Counted per node, because a pair re-peering shows up on both sides and
a mesh-wide total would hide which one started it.
"""
per_node: int = 1
@dataclass
class SwitchLatencyAssertion:
"""Ceiling on the time from a link going down to the first path switch.
Read from the runner's own record of when each edge was taken down
against the timestamp of the first "Path switched" line on either of
its endpoints after that moment. This is the number the design puts a
target on ("under one second"); without it a scenario could switch
thirty seconds after the flap and still pass ``path_switches``.
``max_ms`` bounds the worst flap; a flap with no switch at all within
``max_ms`` fails outright.
"""
max_ms: int = 1000
@dataclass
class MaxStallAssertion:
"""Ceiling on the longest run of iperf3 intervals that moved no bytes.
``min_traffic`` passes on any session that moved any data, and a
switchover that blackholes traffic for ten seconds still leaves plenty
of bytes on either side of the hole. iperf3 reports per-interval
totals (one second by default); this reads the longest run of zeros
inside any session, in seconds, and fails if it exceeds ``max_secs``.
"""
max_secs: float = 2.0
@dataclass
class MaxParentSwitchesAssertion:
"""Stability ceiling: fail if parent switches exceed ``max_total``.
@@ -351,6 +417,10 @@ class AssertionsConfig:
bloom_send_rate: BloomSendRateAssertion | None = None
min_parent_switches: MinParentSwitchesAssertion | None = None
max_parent_switches: MaxParentSwitchesAssertion | None = None
path_switches: PathSwitchesAssertion | None = None
max_promotions: MaxPromotionsAssertion | None = None
switch_latency: SwitchLatencyAssertion | None = None
max_stall: MaxStallAssertion | None = None
max_errors: MaxErrorsAssertion | None = None
congestion_signals: CongestionSignalsAssertion | None = None
tree_parents: TreeParentsAssertion | None = None
@@ -413,12 +483,12 @@ _SECTION_KEYS = {
"num_nodes", "algorithm", "params", "ensure_connected", "subnet",
"ip_start", "default_transport", "transport_mix", "pin_root",
},
"netem": {"enabled", "default_policy", "link_policies", "mutation"},
"netem": {"enabled", "default_policy", "link_policies", "mutation", "dual_udp_policy"},
"netem.link_policies[]": {"edges", "policy", "policy_name"},
"netem.mutation": {"interval_secs", "fraction", "policies", "exclude_edges"},
"link_flaps": {
"enabled", "interval_secs", "max_down_links", "down_duration_secs",
"protect_connectivity",
"protect_connectivity", "only_transport",
},
"traffic": {
"enabled", "max_concurrent", "interval_secs", "duration_secs",
@@ -435,8 +505,9 @@ _SECTION_KEYS = {
"link_swap.edges[]": {"edge", "policy"},
"assertions": {
"bloom_send_rate", "min_parent_switches", "max_parent_switches",
"max_errors", "congestion_signals", "tree_parents", "baseline",
"min_traffic",
"path_switches", "max_promotions", "switch_latency", "max_stall",
"max_errors", "congestion_signals", "tree_parents",
"baseline", "min_traffic",
},
"logging": {"rust_log", "output_dir"},
}
@@ -444,6 +515,10 @@ _ASSERTION_KEYS = {
"bloom_send_rate": {"window_secs", "max_per_node"},
"min_parent_switches": {"min_total"},
"max_parent_switches": {"max_total", "node"},
"path_switches": {"min_total", "max_total"},
"max_promotions": {"per_node"},
"switch_latency": {"max_ms"},
"max_stall": {"max_secs"},
"max_errors": {"max_total"},
"min_traffic": {"min_sessions_ok", "min_bytes_total"},
"congestion_signals": {
@@ -556,6 +631,10 @@ def load_scenario(path: str) -> Scenario:
s.netem.default_policy = _parse_netem_policy(
nc["default_policy"], "netem.default_policy"
)
if "dual_udp_policy" in nc:
s.netem.dual_udp_policy = _parse_netem_policy(
nc["dual_udp_policy"], "netem.dual_udp_policy"
)
if "link_policies" in nc:
for lp_data in nc["link_policies"]:
_reject_unknown(
@@ -597,6 +676,8 @@ def load_scenario(path: str) -> Scenario:
lf["down_duration_secs"], "link_flaps.down_duration_secs"
)
s.link_flaps.protect_connectivity = lf.get("protect_connectivity", True)
only = lf.get("only_transport")
s.link_flaps.only_transport = str(only) if only is not None else None
# Traffic section
tf = raw.get("traffic", {})
@@ -696,6 +777,54 @@ def load_scenario(path: str) -> Scenario:
s.assertions.min_parent_switches = MinParentSwitchesAssertion(
min_total=int(mps.get("min_total", 1)),
)
if "path_switches" in asrt:
ps = asrt["path_switches"]
_reject_unknown(ps, _ASSERTION_KEYS["path_switches"], "assertions.path_switches")
if "min_total" not in ps and "max_total" not in ps:
raise ValueError(
"assertions.path_switches: give min_total, max_total or both "
"(an empty band asserts nothing)"
)
for key in ("min_total", "max_total"):
if key in ps and (isinstance(ps[key], bool) or not isinstance(ps[key], int) or ps[key] < 0):
raise ValueError(
f"assertions.path_switches: {key} must be a non-negative "
f"integer, got {ps[key]!r}"
)
s.assertions.path_switches = PathSwitchesAssertion(
min_total=ps.get("min_total"),
max_total=ps.get("max_total"),
)
if "max_promotions" in asrt:
mp = asrt["max_promotions"]
_reject_unknown(mp, _ASSERTION_KEYS["max_promotions"], "assertions.max_promotions")
per_node = mp.get("per_node", 1)
if isinstance(per_node, bool) or not isinstance(per_node, int) or per_node < 1:
raise ValueError(
"assertions.max_promotions: per_node must be a positive integer, "
f"got {per_node!r} — a pair has to meet once"
)
s.assertions.max_promotions = MaxPromotionsAssertion(per_node=per_node)
if "switch_latency" in asrt:
sl = asrt["switch_latency"]
_reject_unknown(sl, _ASSERTION_KEYS["switch_latency"], "assertions.switch_latency")
max_ms = sl.get("max_ms", 1000)
if isinstance(max_ms, bool) or not isinstance(max_ms, int) or max_ms < 1:
raise ValueError(
"assertions.switch_latency: max_ms must be a positive integer, "
f"got {max_ms!r}"
)
s.assertions.switch_latency = SwitchLatencyAssertion(max_ms=max_ms)
if "max_stall" in asrt:
ms = asrt["max_stall"]
_reject_unknown(ms, _ASSERTION_KEYS["max_stall"], "assertions.max_stall")
max_secs = ms.get("max_secs", 2.0)
if isinstance(max_secs, bool) or not isinstance(max_secs, (int, float)) or max_secs <= 0:
raise ValueError(
"assertions.max_stall: max_secs must be a positive number, "
f"got {max_secs!r}"
)
s.assertions.max_stall = MaxStallAssertion(max_secs=float(max_secs))
if "max_parent_switches" in asrt:
xps = asrt["max_parent_switches"]
_reject_unknown(
@@ -1012,10 +1141,10 @@ def _validate(s: Scenario):
node_ids.update(str(p) for p in entry[:2])
if len(entry) == 3:
transport = str(entry[2])
if transport not in VALID_TRANSPORTS:
if transport not in VALID_TRANSPORTS and transport not in DUAL_TRANSPORTS:
raise ValueError(
f"explicit adjacency[{i}]: transport '{transport}' "
f"not in {VALID_TRANSPORTS}"
f"not in {VALID_TRANSPORTS} or {DUAL_TRANSPORTS}"
)
if len(node_ids) != s.topology.num_nodes:
raise ValueError(
+113 -6
View File
@@ -11,6 +11,41 @@ from .keys import derive_full
from .naming import name_suffix, veth_token
from .scenario import TopologyConfig
# An edge carried by UDP over a dedicated veth pair, each end an
# interface-bound UDP instance (``transports.udp.<iface>.interface``). The
# harness's stand-in for "wifi and cable, both IP": two UDP instances on
# two interfaces, so a peer reachable over both holds two paths.
UDP_VETH = "udp-veth"
# Port of the interface-bound UDP instances. Not 2121: the bridge instance
# binds the wildcard on that port, and a second wildcard bind on the same
# port would conflict.
UDP_VETH_PORT = 2122
# Second octet of the /24s the veth pairs carry. Clear of docker's default
# pool (172.17-31), the sim's claimed 10.30.x ranges and sidecar's 10.40.x.
_UDP_VETH_NET = "10.222"
@dataclass(frozen=True)
class UdpVethLink:
"""One end of a ``udp-veth`` edge, as a node sees it."""
peer_id: str
# The veth interface in this node's container, and the UDP instance
# name bound to it.
iface: str
local_ip: str
peer_ip: str
@property
def instance(self) -> str:
return self.iface
@property
def peer_addr(self) -> str:
return f"{self.peer_ip}:{UDP_VETH_PORT}"
@dataclass
class SimNode:
@@ -29,6 +64,11 @@ class SimTopology:
edges: set[tuple[str, str]] = field(default_factory=set)
# Per-edge transport type; edges not in this dict default to "udp"
edge_transport: dict[tuple[str, str], str] = field(default_factory=dict)
# Edges declared ``ethernet+udp``: an Ethernet veth (found by beacon)
# *and* a UDP static-peer entry over the bridge, so the pair holds two
# paths under one session. ``edge_transport`` says ``ethernet`` for
# these, which is what netem and link flaps act on.
dual_udp_edges: set[tuple[str, str]] = field(default_factory=set)
# Suffix scoping globally-visible names to this run and scenario; empty
# outside the CI harness, which keeps a bare run's names unchanged.
name_suffix: str = ""
@@ -42,6 +82,60 @@ class SimTopology:
"""
return veth_token(self.name_suffix)
def is_dual_udp_edge(self, a: str, b: str) -> bool:
"""Whether the edge also carries a UDP static-peer link."""
return _make_edge(a, b) in self.dual_udp_edges
def is_veth_transport(self, transport: str) -> bool:
"""Whether edges of this transport run over a dedicated veth pair."""
return transport in ("ethernet", UDP_VETH)
def veth_edges(self) -> list[tuple[str, str]]:
"""Every edge that needs a veth pair: Ethernet and ``udp-veth``."""
return sorted(
e for e, t in self.edge_transport.items() if self.is_veth_transport(t)
)
def has_veth(self) -> bool:
return bool(self.veth_edges())
def udp_veth_edges(self) -> list[tuple[str, str]]:
"""Edges carried by UDP over a veth, in canonical order. The index
of an edge here is what its /24 is numbered by."""
return sorted(e for e, t in self.edge_transport.items() if t == UDP_VETH)
def udp_veth_links(self, node_id: str) -> list[UdpVethLink]:
"""This node's ends of its ``udp-veth`` edges, with addressing.
Edge ``k`` (in ``udp_veth_edges`` order) is ``10.222.k.0/24``: the
lower node id is ``.1``, the higher ``.2``.
"""
links = []
for k, (a, b) in enumerate(self.udp_veth_edges()):
if node_id not in (a, b):
continue
if k > 255:
raise ValueError("more than 256 udp-veth edges are not addressable")
local, peer = (a, b) if node_id == a else (b, a)
local_ip = f"{_UDP_VETH_NET}.{k}.{1 if node_id == a else 2}"
peer_ip = f"{_UDP_VETH_NET}.{k}.{2 if node_id == a else 1}"
links.append(
UdpVethLink(
peer_id=peer,
iface=veth_interface_name(local, peer),
local_ip=local_ip,
peer_ip=peer_ip,
)
)
return links
def udp_veth_ip(self, node_id: str, peer_id: str) -> str | None:
"""The veth IP ``node_id`` has on its ``udp-veth`` edge to ``peer_id``."""
for link in self.udp_veth_links(node_id):
if link.peer_id == peer_id:
return link.local_ip
return None
def transport_for_edge(self, a: str, b: str) -> str:
"""Get the transport type for an edge (defaults to 'udp')."""
edge = _make_edge(a, b)
@@ -161,6 +255,7 @@ class SimTopology:
static_edges = {
e for e in self.edges
if self.edge_transport.get(e, "udp") != "ethernet"
or e in self.dual_udp_edges
}
outbound: dict[str, list[str]] = {nid: [] for nid in self.nodes}
@@ -249,7 +344,7 @@ def generate_topology(
adjacency = config.params.get("adjacency")
if not adjacency:
raise ValueError("explicit topology requires params.adjacency")
edges, edge_transport = _generate_explicit(
edges, edge_transport, dual_udp_edges = _generate_explicit(
adjacency, config.default_transport
)
# Validate all referenced nodes exist
@@ -264,6 +359,7 @@ def generate_topology(
# Assign transport types to edges
if config.algorithm != "explicit":
edge_transport = _assign_edge_transports(edges, config, rng)
dual_udp_edges = set()
# Build peer lists from edges
for a, b in edges:
@@ -276,6 +372,7 @@ def generate_topology(
nodes=nodes,
edges=edges,
edge_transport=edge_transport,
dual_udp_edges=dual_udp_edges,
name_suffix=name_suffix(),
)
@@ -354,17 +451,24 @@ def _generate_erdos_renyi(
def _generate_explicit(
adjacency: list, default_transport: str = "udp"
) -> tuple[set[tuple[str, str]], dict[tuple[str, str], str]]:
) -> tuple[set[tuple[str, str]], dict[tuple[str, str], str], set[tuple[str, str]]]:
"""Build edges from an explicit adjacency list.
Each entry is a 2-element list ``[nodeA, nodeB]`` (uses default
transport) or a 3-element list ``[nodeA, nodeB, transport]``.
transport) or a 3-element list ``[nodeA, nodeB, transport]``. The
transport ``ethernet+udp`` declares a dual edge: an Ethernet veth and
a UDP static-peer link between the same two nodes, so the pair holds
two paths under one session. ``udp-veth+udp`` is the all-IP dual edge:
UDP over a dedicated veth (an interface-bound UDP instance at each
end) and UDP over the bridge.
Returns ``(edges, edge_transport)`` where ``edge_transport`` maps
each edge to its transport type.
Returns ``(edges, edge_transport, dual_udp_edges)`` where
``edge_transport`` maps each edge to its transport type (``ethernet``
for a dual edge) and ``dual_udp_edges`` is the set of dual edges.
"""
edges = set()
edge_transport: dict[tuple[str, str], str] = {}
dual_udp_edges: set[tuple[str, str]] = set()
for i, entry in enumerate(adjacency):
if not isinstance(entry, (list, tuple)) or len(entry) not in (2, 3):
raise ValueError(
@@ -374,8 +478,11 @@ def _generate_explicit(
edge = _make_edge(str(entry[0]), str(entry[1]))
edges.add(edge)
transport = str(entry[2]) if len(entry) == 3 else default_transport
if transport in ("ethernet+udp", f"{UDP_VETH}+udp"):
transport = transport[: -len("+udp")]
dual_udp_edges.add(edge)
edge_transport[edge] = transport
return edges, edge_transport
return edges, edge_transport, dual_udp_edges
def _assign_edge_transports(
+20 -11
View File
@@ -37,7 +37,7 @@ import subprocess
import time
from .docker_exec import DockerExecError, docker_exec, docker_exec_quiet
from .topology import SimTopology, veth_interface_name
from .topology import UDP_VETH, SimTopology, veth_interface_name
log = logging.getLogger(__name__)
@@ -111,14 +111,14 @@ class VethManager:
4. Rename to final names and bring up
5. Query MACs and store in SimNode.ethernet_macs
"""
eth_edges = self.topology.ethernet_edges()
if not eth_edges:
edges = self.topology.veth_edges()
if not edges:
return
image = self._get_image()
log.info("Setting up %d Ethernet veth pairs (helper image: %s)...", len(eth_edges), image)
log.info("Setting up %d veth pairs (helper image: %s)...", len(edges), image)
for a, b in eth_edges:
for a, b in edges:
self._create_veth_pair(a, b, image)
log.info(
@@ -135,7 +135,7 @@ class VethManager:
stopped, whose pairs are left for their own restart.
"""
image = self._get_image()
for a, b in self.topology.ethernet_edges():
for a, b in self.topology.veth_edges():
if a != node_id and b != node_id:
continue
# Remove existing pair if any (host-side might still exist)
@@ -216,8 +216,15 @@ class VethManager:
_require_host(["ip", "link", "set", host_a, "netns", str(pid_a)], image)
_require_host(["ip", "link", "set", host_b, "netns", str(pid_b)], image)
_raise_link(container_a, host_a, final_a)
_raise_link(container_b, host_b, final_b)
# A udp-veth edge carries IP: address each end before it comes up,
# so the daemon's interface wait never sees the interface without
# its address.
ip_a = ip_b = None
if self.topology.transport_for_edge(node_a, node_b) == UDP_VETH:
ip_a = self.topology.udp_veth_ip(node_a, node_b)
ip_b = self.topology.udp_veth_ip(node_b, node_a)
_raise_link(container_a, host_a, final_a, ip_a)
_raise_link(container_b, host_b, final_b, ip_b)
_await_up(container_a, final_a)
_await_up(container_b, final_b)
@@ -272,11 +279,13 @@ def _require_host(cmd: list[str], image: str):
raise VethSetupError(f"host command failed: {' '.join(cmd)}")
def _raise_link(container: str, temp: str, final: str):
"""Rename a moved veth end to its final name and set it up."""
def _raise_link(container: str, temp: str, final: str, ip: str | None = None):
"""Rename a moved veth end to its final name, address it if asked,
and set it up."""
addr = f" && ip addr add {ip}/24 dev {final}" if ip else ""
_in_container(
container,
f"ip link set {temp} name {final} && ip link set {final} up",
f"ip link set {temp} name {final}{addr} && ip link set {final} up",
f"renaming {temp} to {final}",
)
+3 -1
View File
@@ -31,7 +31,7 @@
# nat-lan, nostr-publish-consume, stun-faults,
# chaos-churn-mixed-10, chaos-ethernet-mesh,
# chaos-ethernet-only, chaos-ethernet-churn, chaos-tcp-mesh,
# chaos-congestion-stress,
# chaos-congestion-stress, dual-path-flap, dual-udp-flap,
# sidecar, dns-resolver, deb-install, medium-change
#
# Opt-in (require --with-tor; depend on live Tor network):
@@ -166,6 +166,8 @@ CHAOS_SUITES=(
"ethernet-mesh ethernet-mesh"
"ethernet-only ethernet-only"
"ethernet-churn ethernet-churn"
"dual-path-flap dual-path-flap"
"dual-udp-flap dual-udp-flap"
"tcp-mesh tcp-mesh"
"congestion-stress congestion-stress"
)
+6 -6
View File
@@ -34,14 +34,14 @@ enable_ecn() {
}
wait_for_ethernet() {
# If config references ethernet transports, wait for interfaces to appear.
# Veth pairs are created from the host after the container starts.
# If config binds any transport to an interface (Ethernet, or a UDP
# instance with `interface:`), wait for those interfaces to appear. Veth
# pairs are created from the host after the container starts, and a UDP
# instance bound to an interface that is not there yet fails to start.
local eth_ifaces=""
if grep -q 'ethernet:' "$CONFIG" 2>/dev/null; then
eth_ifaces=$(grep '^\s*interface:' "$CONFIG" \
eth_ifaces=$(grep '^\s*interface:' "$CONFIG" 2>/dev/null \
| sed 's/.*interface:\s*//' \
| tr -d ' ' || true)
fi
| tr -d ' "' || true)
if [ -n "$eth_ifaces" ]; then
echo "Waiting for Ethernet interfaces: $eth_ifaces"
+9
View File
@@ -40,6 +40,7 @@ class AnalysisResult:
peers_promoted: list[tuple[str, str]] = field(default_factory=list)
peer_removals: list[tuple[str, str]] = field(default_factory=list)
parent_switches: list[tuple[str, str]] = field(default_factory=list)
path_switches: list[tuple[str, str]] = field(default_factory=list)
mmp_link_metrics: list[tuple[str, str]] = field(default_factory=list)
mmp_session_metrics: list[tuple[str, str]] = field(default_factory=list)
handshake_timeouts: list[tuple[str, str]] = field(default_factory=list)
@@ -70,6 +71,7 @@ class AnalysisResult:
f"Peers promoted: {len(self.peers_promoted)}",
f"Peer removals: {len(self.peer_removals)}",
f"Parent switches: {len(self.parent_switches)}",
f"Path switches: {len(self.path_switches)}",
f"Handshake timeouts: {len(self.handshake_timeouts)}",
f"MMP link samples: {len(self.mmp_link_metrics)}",
f"MMP session samples: {len(self.mmp_session_metrics)}",
@@ -160,6 +162,13 @@ def _analyze_lines(result: AnalysisResult, source: str, log_text: str):
# Parent switches
if "Parent switched" in line:
result.parent_switches.append((source, line))
# Path switches: a peer's traffic moved to another transport under
# the same session. Three emitters, one per trigger (selection, the
# presence edge, a peer's PathClose); all say "session kept".
if "session kept" in line and (
"Path switched" in line or "traffic moved to the standby" in line
):
result.path_switches.append((source, line))
# Handshake timeouts
if "timed out" in line and ("handshake" in line.lower() or "Handshake" in line):
result.handshake_timeouts.append((source, line))