Merge master into next, forward-porting each side onto the other

Five files conflicted and each needed a different call, because the two
lines had rewritten different halves of the same code.

The handshake handler takes next's version whole. master's entire change
there was three comment blocks and one widened debug_assert, and the
assert names HandshakePhase::ReceivedMsg1, a variant next's XX rewrite does
not have. Nothing semantic was dropped.

The peer reaper takes master's: reap_peers_on_transport is new and its
route_link_dead doc now describes both callers, which is true on this line
too.

The ethernet transport takes master's binder rewrite with next's wire
format re-applied on top. The send path, the receive path and the frame
tests merged to the 4-byte header on their own, but three sites are new in
master's rewrite and had never seen it: the Binding default and both arms
of the binder's MTU calculation still subtracted 3. The transports snapshot
fixture moved with them, 1499 to 1496, and that single field was the whole
diff. Beacons carry no pubkey here, so local_pubkey leaves the transport,
its binder context and the node's transport construction with it.

The changelog keeps both sides' entries, with master's Added subsection
lifted back out of Changed where the merge had left it.

Two tests do not come across. a_transient_msg2_failure_keeps_the_link_for_
the_retry and its restart-path sibling assert that the machine rests at
ReceivedMsg1. This line's nearest state is SentMsg2, and it means something
else: the inbound leg parks there awaiting msg3, where on the other line
that phase was the last stop before promotion. Renaming it would produce a
test that passes without exercising the deferral. The behaviour they guard
did merge and sits in the transient arm of the msg2 send failure; what is
missing is coverage shaped for this handshake, which is tracked separately.

The two connected-socket tests did come across. Their helper took the
responder's session straight after msg2, which is an IK assumption; it now
runs msg3 as well. Both pass here and both go red when the clear is removed
or made unconditional.

The test-harness fixes arrive through master rather than as follow-ups
here, so this line never carries the versions that failed: the
interface-binding suite's veth naming, and the chaos veth restore, random
streams, settle wait, netem restore and shared down-node set.
This commit is contained in:
Johnathan Corgan
2026-09-10 19:19:27 +00:00
56 changed files with 8376 additions and 743 deletions
+184
View File
@@ -284,6 +284,30 @@ jobs:
- name: Install system dependencies
run: sudo apt-get update && sudo apt-get install -y libdbus-1-dev
# The same address-less interface the musl leg creates. Pinning the
# contract on both libcs is what turns "glibc and musl agree about
# `getifaddrs`" from an assumption into a checked fact — and makes this
# leg fail first if glibc is the one that changes.
- name: Create an address-less interface for the presence probe
run: |
sudo ip link add fips-probe0 type dummy
# `addrgenmode none` before bringing it up: the kernel hands an IPv6
# link-local to any interface that comes up, and an interface with a
# link-local is not address-less — the fixture would have quietly
# tested nothing.
sudo ip link set fips-probe0 addrgenmode none
sudo ip link set fips-probe0 up
ip addr show fips-probe0
# Fail rather than test the wrong thing if it acquired one anyway.
if ip addr show fips-probe0 | grep -qE "inet6? "; then
echo "fips-probe0 has an address; it cannot test the address-less case" >&2
exit 1
fi
echo "FIPS_TEST_ADDRLESS_IFACE=fips-probe0" >> "$GITHUB_ENV"
# Declare that this runner has fixtures, so a test that depends on
# one fails when the fixture is missing instead of skipping silently.
echo "FIPS_TEST_REQUIRE_FIXTURES=1" >> "$GITHUB_ENV"
- name: Install Rust toolchain
uses: actions-rust-lang/setup-rust-toolchain@166cdcfd11aee3cb47222f9ddb555ce30ddb9659 # v1
with:
@@ -304,9 +328,39 @@ jobs:
- name: Install cargo-nextest
uses: taiki-e/install-action@nextest
- name: Run unit tests
run: cargo nextest run --all --profile ci
# The bind-success half. Every other unit-test leg runs unprivileged, so
# `PacketSocket::open` cannot succeed on any of them and everything past
# a successful bind — the post-store shutdown check, the `Present` arm of
# the binder loop, `bind_now` itself — runs nowhere in CI.
#
# Built as the runner user and only *executed* under sudo: `cargo` run as
# root would use root's CARGO_HOME and discard the cache this job just
# restored.
#
# `FIPS_TEST_PRIVILEGED` is what makes this leg honest. A test that needs
# a raw socket skips quietly without it; with it set, a test that cannot
# open one fails and says so, so a runner that stops granting the
# capability shows up as a red leg rather than as silence.
- name: Run interface-binding tests with privilege
run: |
cargo test --lib --no-run
BIN=$(cargo test --lib --no-run --message-format=json \
| jq -r 'select(.reason == "compiler-artifact")
| select(.executable != null)
| select(.target.kind[0] == "lib")
| .executable' \
| tail -1)
if [ -z "$BIN" ] || [ ! -x "$BIN" ]; then
echo "could not locate the lib test binary" >&2
exit 1
fi
echo "running $BIN as root"
sudo -E env FIPS_TEST_PRIVILEGED=1 "$BIN" transport::ethernet --test-threads=1
- name: Publish test report (Checks tab)
uses: dorny/test-reporter@4a2e97665d5fa767581ef38eca97b9694bd4eef4 # v2
if: always()
@@ -363,9 +417,116 @@ jobs:
- name: Install cargo-nextest
uses: taiki-e/install-action@nextest
# The Darwin half of the address-less presence contract. The Linux legs
# pin that `getifaddrs` reports an interface with no addresses as
# present, on both glibc and musl; without this the same claim on the
# BSD-derived implementation the macOS backend actually calls was
# untested, and the test skipped itself silently on this runner.
#
# `feth` is macOS's fake-Ethernet pseudo-interface. It is created
# address-less, and the check below fails the leg rather than testing the
# wrong thing if this runner hands it one anyway — the same shape as the
# Linux fixture step, which needs `addrgenmode none` for exactly that
# reason.
- name: Create an address-less interface for the presence probe
run: |
sudo ifconfig feth0 create
sudo ifconfig feth0 up
ifconfig feth0
if ifconfig feth0 | grep -qE "^[[:space:]]*inet6? "; then
echo "feth0 has an address; it cannot test the address-less case" >&2
exit 1
fi
echo "FIPS_TEST_ADDRLESS_IFACE=feth0" >> "$GITHUB_ENV"
# Declare that this runner has fixtures, so a test that depends on
# one fails when the fixture is missing instead of skipping silently.
echo "FIPS_TEST_REQUIRE_FIXTURES=1" >> "$GITHUB_ENV"
- name: Run unit tests
run: cargo nextest run --all --profile ci
# ─────────────────────────────────────────────────────────────────────────────
# Job 2bb – Unit tests (musl)
#
# OpenWrt — the platform the Ethernet transport's dynamic interface binding
# exists for — is musl, and musl reimplements the libc calls that binding is
# built on rather than sharing glibc's. `interface_present` reads `ifa_flags`
# out of `getifaddrs`, and the interfaces it has to see (`fips-mesh0`,
# `fips-ap0`) are deliberately unbridged with no IP address at all, which is
# exactly where getifaddrs implementations differ. Every other leg is glibc, so
# without this one the presence probe is asserted on a libc no test has ever
# run it against, on the target it was written for.
#
# Built for the musl target on a glibc host rather than inside an Alpine
# container. The test binary links musl statically and runs natively on the
# runner, so musl's `getifaddrs` is the one under test — while the build
# scripts stay host artifacts, which keeps rustables' bindgen on the same
# libclang the glibc leg already builds with. Building inside Alpine put
# bindgen on a musl toolchain it does not work on: statically linked build
# scripts cannot `dlopen` libclang, and turning the static CRT off then left
# it loading libclang but unable to parse. None of that is anything this leg
# is trying to test.
# ─────────────────────────────────────────────────────────────────────────────
test-musl:
name: Unit tests (musl)
runs-on: ubuntu-latest
needs: [build]
steps:
- uses: actions/checkout@d23441a48e516b6c34aea4fa41551a30e30af803 # v6
- name: Set SOURCE_DATE_EPOCH from git
run: echo "SOURCE_DATE_EPOCH=$(git log -1 --format=%ct)" >> "$GITHUB_ENV"
# libdbus for the host build scripts; musl-tools for the musl C
# toolchain the `cc`-driven dependencies link against. BLE is excluded on
# musl by a Cargo.toml cfg, so bluer is not in this build at all.
- name: Install system dependencies
run: sudo apt-get update && sudo apt-get install -y libdbus-1-dev musl-tools
# The address-less interface the presence probe has to be tested
# against; see the matching step on the glibc leg for why loopback
# cannot stand in for it.
- name: Create an address-less interface for the presence probe
run: |
sudo ip link add fips-probe0 type dummy
# `addrgenmode none` before bringing it up: the kernel hands an IPv6
# link-local to any interface that comes up, and an interface with a
# link-local is not address-less — the fixture would have quietly
# tested nothing.
sudo ip link set fips-probe0 addrgenmode none
sudo ip link set fips-probe0 up
ip addr show fips-probe0
# Fail rather than test the wrong thing if it acquired one anyway.
if ip addr show fips-probe0 | grep -qE "inet6? "; then
echo "fips-probe0 has an address; it cannot test the address-less case" >&2
exit 1
fi
echo "FIPS_TEST_ADDRLESS_IFACE=fips-probe0" >> "$GITHUB_ENV"
# Declare that this runner has fixtures, so a test that depends on
# one fails when the fixture is missing instead of skipping silently.
echo "FIPS_TEST_REQUIRE_FIXTURES=1" >> "$GITHUB_ENV"
- name: Install Rust toolchain
uses: actions-rust-lang/setup-rust-toolchain@166cdcfd11aee3cb47222f9ddb555ce30ddb9659 # v1
with:
cache: false
rustflags: ''
target: x86_64-unknown-linux-musl
- name: Cache Cargo registry + build
uses: actions/cache@caa296126883cff596d87d8935842f9db880ef25 # v5
with:
path: |
~/.cargo/registry
~/.cargo/git
target
key: musl-cargo-${{ hashFiles('**/Cargo.lock') }}
restore-keys: |
musl-cargo-
- name: Run library tests
run: cargo test --lib --target x86_64-unknown-linux-musl
# ─────────────────────────────────────────────────────────────────────────────
# Job 2c – Unit tests (Windows)
# ─────────────────────────────────────────────────────────────────────────────
@@ -395,6 +556,7 @@ jobs:
- name: Install cargo-nextest
uses: taiki-e/install-action@nextest
- name: Run unit tests
run: cargo nextest run --all --profile ci
@@ -456,6 +618,9 @@ jobs:
# ── Firewall baseline (fips0 nftables default-deny) ────────────
- suite: firewall
type: firewall
# ── Dynamic interface binding (absent → present → absent) ──────
- suite: iface-binding
type: iface-binding
# ── Outbound LAN gateway integration test ──────────────────────
- suite: gateway
type: gateway
@@ -471,6 +636,9 @@ jobs:
- suite: ethernet-only
type: chaos
scenario: ethernet-only
- suite: ethernet-churn
type: chaos
scenario: ethernet-churn
- suite: tcp-mesh
type: chaos
scenario: tcp-mesh
@@ -606,6 +774,22 @@ jobs:
run: |
docker compose -f testing/firewall/docker-compose.yml down --volumes --remove-orphans
# ── Dynamic interface binding integration test ─────────────────────────
- name: Run interface binding integration test
if: matrix.type == 'iface-binding'
run: bash testing/iface-binding/test.sh --skip-build --keep-up
- name: Collect logs on failure (iface-binding)
if: matrix.type == 'iface-binding' && failure()
run: |
docker compose -f testing/iface-binding/docker-compose.yml logs --no-color
docker exec fips-ifb-node-a fipsctl show transports || true
- name: Stop containers (iface-binding)
if: matrix.type == 'iface-binding' && always()
run: |
docker compose -f testing/iface-binding/docker-compose.yml down --volumes --remove-orphans
# ── Chaos simulation ───────────────────────────────────────────────────
- name: Install Python deps (chaos)
if: matrix.type == 'chaos'
+254
View File
@@ -117,6 +117,71 @@ with v0.5.x or earlier peers.
- The receive-path `RejectReason` classification (shipped in 0.4.0) is
additionally wired into the Noise XX handshake cluster
(msg1/msg2/msg3) and the rekey-initiator outbound sites on `next`.
- Dynamic interface binding for the Ethernet transport. An interface-bound
transport is now a long-lived object that is *sometimes bound*: the interface
it names need not exist when the daemon starts, may appear minutes later, and
may vanish and return mid-operation. `start_async` returns `Ok` with the
transport **absent** rather than failing, and a per-transport binder task
binds when the interface appears, unbinds when it goes away, and rebinds when
it returns. Start-time absence and runtime detach are one code path.
Detection is event-driven where the kernel offers a source — netlink
`RTNLGRP_LINK` on Linux, `PF_ROUTE` on macOS and FreeBSD — with a 1 s
`getifaddrs` poll underneath as a backstop. Presence means
`IFF_UP` — the interface exists and is administratively up — and
deliberately not `IFF_RUNNING`. Binding needs no carrier and a socket
outlives a carrier flap, so a bridge with nothing plugged into it (`br-lan`
on a wifi-only router) is bound and healthy rather than permanently
`Degraded`, and starts carrying traffic the moment a port comes up. Whether
an interface has carrier is reported separately as `interface.carrier` in
`show_transports`, never acted on.
This closes the OpenWrt boot race (procd starts `fips` before wifi has
created `fips-mesh0` / `fips-ap0`; both transports were skipped for the life
of the process while the 802.11s peer link formed anyway, so the node looked
healthy and reached nothing), the intermittent-adapter case, and the
mid-operation `wifi reload` that destroyed and recreated an interface under a
live socket.
- `transports.ethernet.*.optional` (bool, default `false`). Naming an interface
in configuration is a statement that you expect it, so the default is to
complain: while a required interface is missing the node reports `Degraded`
and logs the edge, at a severity that follows how long the absence lasts
(see below). `optional: true` makes absence silent (`info` on the edge, no
health impact) for hardware that is legitimately not always there. It describes the interface's *presence*, not the transport's
importance — an optional interface that is present is used exactly as hard as
any other — and no value of it makes a missing interface fatal at startup.
- `fipsctl show transports` reports interface presence per transport under a new
`interface` block: `name`, `presence` (`absent` / `binding` / `present`),
`carrier`, `policy` (`required` / `optional`), `since_secs`, `binds` and
`failed_attempts`. The original boot-race bug was expensive precisely because
nothing an operator could see said the node was deaf.
- `fipstop`'s transports view carries the same interface presence. The State
column shows an interface-bound transport's presence rather than its
lifecycle state — `up` is true from the moment the transport starts and stays
true while its interface is missing, which is precisely the wrong answer in
the one case someone is scanning that column for — and a new Policy column
reads `required` or `optional` beside it, with an absent required interface
red and an absent optional one yellow: the same split the daemon makes
between staying `Full` and reporting `Degraded`. The instance name and the
thing a transport is bound to are now separate columns, so netdev names line
up down the list instead of trailing ragged inside a packed label. The detail
pane gains an Interface block: netdev, presence and how long it has been
held, carrier, what the absence policy means rather than which key sets it,
bind count (flagged once it has rebound) and failed binds when there are any.
The table fits an 80-column terminal — the OpenWrt serial console and the
xterm and tmux default — dropping the byte counters below 100 columns and
stacking the detail pane below 110, rather than shrinking every column until
none of them can be read.
- `testing/iface-binding/` integration suite (`ci-local.sh --only
iface-binding`, and a GitHub matrix leg): two daemons whose only transports
are interface-bound, run against a veth pair the harness creates, downs,
deletes and recreates underneath them. Asserts the boot race, the late
attach and peering over it, the flap in both directions,
destroy-and-recreate, that an `optional` interface never moves node health,
and that absence is logged once on the edge rather than once per retry.
#### Node lifecycle
@@ -174,8 +239,23 @@ with v0.5.x or earlier peers.
stay under it, or set `node.netmon.enabled: false`. A
`node.link_dead_timeout_secs` of 0 is exempt from the check.
#### Library surface and internals
- `TransportError::InterfaceUnavailable { interface }`. A missing interface and
a typo'd interface name were previously the same flat
`StartFailed(String)`; nothing downstream could branch on absence.
### Changed
- The lockfile moves `chacha20` from 0.10.1 to 0.10.2, because 0.10.1 is yanked.
It arrives through `rand`, a direct dependency,
so it sits on the built path rather than off to one side. The requirement in
`Cargo.toml` already admitted 0.10.2, so this is a lockfile change and no code
changed with it. **This is not a security fix**: `cargo audit` reports nothing
against `chacha20` at either version, and 0.10.1 was withdrawn by its
maintainer rather than flagged by an advisory. What it buys is that a fresh
checkout can resolve the lockfile without reaching for a yanked version.
- `node.rekey.enabled` now means "initiate rekeys" and nothing else. The
responder half of the establish decision was also gated on it, and once the
rekey is declared in the msg3 negotiation payload that flag was the only
@@ -254,6 +334,127 @@ with v0.5.x or earlier peers.
shared-media legs and every inbound leg, are unaffected and still promote
whoever answers.
- `Degraded` is now a level rather than a latch. The supervisor's reason set
was monotonic, which was correct while no child could recover; with recovery
it would have meant "something broke at some point since boot" rather than
"something is broken now". Interface absence is tracked in its own reversible
set and node health is recomputed on every transition **in both directions**,
so plugging the WAN back in clears `Degraded` without a restart. A transport
whose interface is absent still counts as up, so a single-interface node that
boots before its wifi degrades rather than exiting on "no transports".
- The OpenWrt package ships the `mesh0`/`mesh1` and `ap0`/`ap1` Ethernet
transports **enabled** with `optional: true`, instead of commented out.
`fips-mesh-setup` and `fips-ap-setup` no longer comment-toggle blocks in
`fips.yaml`, and no longer tell the operator to restart the daemon after
creating an interface — the daemon binds it on its own. `phy0-sta0` (`wwan`)
is marked `optional: true` for the same reason: it only exists while a radio
is in station mode.
**Upgrade note: an existing `/etc/fips/fips.yaml` is preserved and does not
gain the new key.** It is a package conffile, so on a router where
`fips-mesh-setup` or `fips-ap-setup` had already uncommented a block, that
block stays as it was, with no `optional` key — and `optional` defaults to
false. Such a block is therefore `required`, so an absent `fips-mesh0` keeps
the node `Degraded` and is reported once at `error` ten seconds in, where the
same block in the shipped file is silent. Add `optional: true` to the block
to match what the package now ships.
- The Ethernet receive loop backs off and exits on a dead socket instead of
spinning on `Err` with a `warn!` per iteration, and the ad-hoc ENXIO
socket-reopen in the beacon sender is gone. Both hand recovery to the
presence machine: one mechanism for every cause rather than one hack per
symptom. Beacons pause while an interface is absent.
- The absence edge is not itself an error, and there is exactly one deadline
after it. An interface missing when the daemon starts logs at `info` — that
is the boot race the mechanism exists to absorb, not a fault — and a runtime
detach at `warn`, because a link coming and going is ordinary weather for a
mesh daemon. Ten seconds is the whole grace: past it, absence is no longer a
race against a radio or a container, so a **required** interface still
missing is reported once at `error`. Start-time absence and a runtime detach
share that one deadline rather than getting one each. An `optional`
interface never reaches `error`. Node health does not wait for any of it,
publishing `Degraded` on the first edge either way.
- The TUN boundary's TCP MSS clamp now tracks the node's egress MTU at
runtime instead of freezing it at startup. `transport_mtu()` is the minimum
across *bound* transports, so a transport that binds minutes after start can
be the narrow one — but the TUN reader and writer were handed a `u16`
computed once when they spawned, while every other consumer
(`show_status`, the control-socket snapshot, the session-layer fragmentation
check) read it live. A node could therefore report one effective IPv6 MTU
and clamp to another. The ceiling is now shared with those threads and
recomputed whenever the bound set changes, in both directions: a narrow
interface appearing tightens it, and its departure releases it. MSS is
negotiated per connection, so a change applies to connections opened after
it; existing ones are not disturbed.
- Rebinds that keep succeeding into a socket that dies moments later are
damped: consecutive bindings shorter than ten seconds back off on the
1 s → 30 s curve, and past three of them the binder stops announcing each
bind as a recovery until one lasts. Undamped, a persistently broken socket
behind a healthy interface produced a log pair and a `Degraded`→`Running`
health flap every second.
- A bind failure that is **not** absence — no `CAP_NET_RAW`, no readable
`/dev/bpf*`, a buffer the kernel refused — fails the daemon's start as it
always has, rather than being waited out. It is a fault, not a state, and
will not resolve on its own; only a missing interface is retried at start.
A non-absence failure during a later rebind still backs off, since the node
is serving by then.
- The binder cannot outlive its transport, and a teardown that races a bind
cannot leave a live receive loop on a socket nothing owns. A shared stop flag
is raised before teardown and checked by the binder after it stores a
binding, so whichever order the two interleave exactly one of them cleans up;
`EthernetTransport` gained a `Drop` that raises the flag, aborts the binder
and releases the socket, for handles dropped without `stop_async`.
- Presence edges are published with `try_send` and retried on the next tick
rather than awaited. A bounded channel could previously park the binder
mid-publish — a health channel able to deadlock the machine whose health it
carries — freezing the interface in whatever state it held.
- `TransportHandle::is_bound()` joins `is_operational()`: the latter means the
transport was *started*, which for an interface-bound transport no longer
implies a live socket. `Node::transport_mtu` now filters on the former,
because an interface that has never existed was clamping the whole node's
IPv6 MTU to a number derived from absent hardware.
- An interface deleted and recreated under the same name is detected as a
detach. Both backends bind by device rather than by name, so the old socket
is attached to nothing while the name still resolves — and a stale
`AF_PACKET` socket never becomes readable, so nothing errors and nothing
exits. Detection previously rested entirely on the beacon sender failing,
which a node with `announce: false` does not have. The bound interface index
is now captured at bind and compared on every poll.
- The link-event watcher distinguishes a genuine receive error from
`WouldBlock`. `try_io` clears readiness only on the latter, so a persistent
error — `ENOBUFS` after a burst of link events overflows the socket buffer —
span a core flat with nothing logged. Errors are now counted, logged once,
backed off, and after five the source is abandoned for the presence poll.
- Presence probes are coalesced to at most ten a second. Linux netlink is
filtered to `RTNLGRP_LINK`, but `PF_ROUTE` has no group filter, so the macOS
source delivers every routing message on the host — route churn, ARP, DHCP
renewals, a VPN going up and down — and each would otherwise drive a full
`getifaddrs` walk.
- CI runs the library tests on musl (Alpine) as well as glibc. Presence is
built on `getifaddrs` and `ifa_flags`, musl reimplements both independently,
and the interfaces this feature exists for (`fips-mesh0`, `fips-ap0` on
OpenWrt) are unbridged with no IP address at all — the case where
implementations most plausibly differ. It was previously asserted on a libc
no test had ever run it against, on the target it was written for.
- Interface presence state ignores lock poisoning. Treating a poisoned lock as
a failure meant reading "no socket, tasks dead", which is the destructive
direction: a transport reporting itself present while every send fails, or a
binder tearing down and rebinding every second while teardown silently
declined to abort anything.
### Fixed
- A leaf-profile node no longer self-elects as tree root. A leaf holding the
@@ -360,6 +561,59 @@ with v0.5.x or earlier peers.
#### Node lifecycle
- Losing an interface no longer leaves its peers in the routing table. The
peers stayed in the registry, the routes through them stayed selectable, and
the node kept advertising reachability it no longer had — so transit traffic
was dropped in silence and other nodes kept routing toward this one for those
destinations, until the liveness reaper noticed up to
`node.link_dead_timeout_secs` later. Measured on real hardware, a detached
dongle took the node's parent with it and no new parent was chosen for
twenty-seven seconds, with four alternative peers available the whole time. A
transport's detach edge now withdraws every peer whose active link runs over
it, on the same path the liveness reaper uses, so sessions, path MTU, session
indices, the link, the control machine, tree cleanup and re-announce, and
bloom withdrawal unwind exactly as they already did. It is not filtered by
`optional`: whether an interface's absence is normal is a statement about
node health, and says nothing about whether the routes over it still work.
The trade is that an absence shorter than the dead timeout that then recovers
now costs a re-peer where it previously cost nothing, accepted because
black-holing is silent, poisons other nodes' routing and takes the full
timeout to clear, where a re-peer is bounded, visible and self-healing.
- A local interface flap during a handshake is no longer charged to the remote.
A msg2 send refused because the interface is absent or mid-rebind was treated
as a failed handshake: the link was removed, the reverse-address entry
dropped, the session index freed, the control machine torn down, and the
whole thing recorded under the reject reason that means "the remote sent
something invalid", which is what an operator reading the rejects would have
concluded. The initiator meanwhile resent msg1 into a link that no longer
existed and had to rebuild from nothing. A transport error the daemon is
already working to clear now leaves the half-built link exactly where it is
for that resend to land on; only a terminal error still tears down, and a
link nobody resends to is reaped at `node.rate_limit.handshake_timeout_secs`
like every other abandoned handshake. The rekey msg1 send site keeps its
teardown, which was already benign, and stops reporting a local self-clearing
condition at `warn`.
- A peer reachable over two interfaces is no longer re-dialled on the path it
is not using. Beacon discovery skipped only a candidate naming the peer's
*current* path, which is the one case that could not churn anything, so the
alternate path was dialled every discovery tick; each dial that completed
promoted and displaced a healthy incumbent, tore down the session, and, when
that peer was the parent, switched parents and re-announced mesh-wide.
Measured on real hardware, seventeen dials to one peer in fifteen minutes,
alternating wifi and cable, displacing a link reporting etx 1.0 and loss 0.0.
Discovery now asks whether the link it already holds is answering rather than
which path the candidate names. Failover is unchanged: a peer that goes quiet
for longer than `node.heartbeat_interval_secs` is dialled again on every
path, alternate included. **One behaviour goes with it.** A peer held on an
adopted NAT-traversal transport that is *also* reachable by Ethernet or BLE
beacon used to drift onto the local path on the next discovery tick, and now
stays on the traversed path for as long as that path answers. Migrating it is
still done by the configured-peer refresh (a config reload, a runtime peer
update, or `fipsctl connect`), and a traversed link that goes quiet still
releases the peer to every path.
- A heartbeat whose send failed no longer counts as one that was delivered.
The peer's "last heartbeat" timestamp was stamped before the send and left
alone whatever came back, so a failure suppressed the next attempt for a
Generated
+3 -3
View File
@@ -498,9 +498,9 @@ dependencies = [
[[package]]
name = "chacha20"
version = "0.10.1"
version = "0.10.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "d524456ba66e72eb8b115ff89e01e497f8e6d11d78b70b1aa13c0fbd97540a81"
checksum = "65c35e4b699c7e15ccbe7ee35c005e4fc0a278d22238a2857e6ce2dadeda1b06"
dependencies = [
"cfg-if",
"cpufeatures 0.3.0",
@@ -2614,7 +2614,7 @@ version = "0.10.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "c7f5fa3a058cd35567ef9bfa5e75732bee0f9e4c55fa90477bef2dfcdbc4be80"
dependencies = [
"chacha20 0.10.1",
"chacha20 0.10.2",
"getrandom 0.4.3",
"rand_core 0.10.1",
]
+314
View File
@@ -1008,6 +1008,320 @@ Transports begin in `Configured` state with all parameters set. `start()`
transitions through `Starting` to `Up` (operational). `stop()` moves to
`Down`. Transport failures move to `Failed`.
## Interface Presence
`Up` describes the *transport*, not the socket. An interface-bound transport
(today: Ethernet) carries a second, orthogonal state — whether it is bound
right now — and the two are independent: a transport is `Up` from the moment
it starts, whether or not the interface it names exists.
### The Gap This Closes
Three deployment scenarios exercise one missing mechanism:
- **Boot ordering.** On OpenWrt, procd starts `fips` before wifi has created
`fips-mesh0` / `fips-ap0`. Both transports were skipped and never retried,
while the 802.11s peer link formed anyway — that is mac80211, not the
daemon — so the node looked healthy and reached nothing. The failure was
expensive precisely because nothing an operator could see said the node was
deaf.
- **Intermittent hardware.** A USB ethernet adapter named in `fips.yaml` is
plugged in some days and not others. Its absence is normal and must be
silent; its arrival must bind without operator action.
- **Mid-operation restart.** `wifi reload` for a channel change destroys and
recreates the mesh interface within a couple of seconds. The socket dies,
the receive loop spun on `Err` with no backoff and no exit, and nothing
rebound.
These are not three features. They are one presence machine plus one policy
field. Before it existed, the first observation was final: an interface
missing at start was logged once and skipped for the life of the process, and
one that disappeared at runtime published a health change but was never
rebound.
### The Presence Machine
```text
Absent ──attach──> Binding ──ok──> Present
^ │ │
└──── fail/backoff ─┘ │
└──────────── detach ──────────────┘
```
`start_async` binds if it can and otherwise returns `Ok` with the transport
`Up` and `Absent`; a per-transport binder task then binds when the interface
appears, tears the socket down when it goes away, and rebinds when it
returns. Two invariants do the work:
- **The transport object survives detach.** Config, `TransportId`,
statistics and the neighbor buffer persist; only the file descriptor and
its loops go. A transport is never destroyed because its interface went
away.
- **Start-time absence and runtime detach are the same transition.** A node
that boots before its wifi and a node whose wifi reloads at 03:00 take one
code path. The old asymmetry — skip forever at start, publish health at
runtime — is gone.
`TransportError::InterfaceUnavailable` is what makes absence branchable. A
missing interface and a typo'd interface name were the same flat
`StartFailed(String)`, so nothing downstream could tell a state from a fault.
### What Counts as Present
Presence means `IFF_UP` — the interface exists and the operator has enabled
it — and deliberately **not** `IFF_RUNNING`.
Carrier and bindability are different questions, and only the second belongs
in a bind gate. An `AF_PACKET` socket on a carrier-less bridge is valid and
starts carrying traffic the instant a member port comes up, with no rebind:
the socket outlives the carrier. Gating on `IFF_RUNNING` bought nothing and
cost three things:
- `br-lan` on a router with nothing in its LAN ports is `UP` with
`NO-CARRIER`, so a healthy wifi-only router reported `Degraded` forever and
errored for a fault it did not have;
- every carrier flap the socket would have survived became an unbind/rebind
cycle — churn the presence machine then has to damp, a mechanism
compensating for a policy error;
- an 802.11s interface that reports `RUNNING` only once it has peered cannot
peer, because peering needs beacons, beacons need a bound socket, and the
gate refuses to bind. A deadlock reachable on the hardware this mechanism
was written for.
The signal `IFF_RUNNING` carries is not lost: `show_transports` reports
`interface.carrier` beside presence, so an operator can still tell a bound
transport carrying nothing from a working one. It is reported rather than
obeyed.
The probe is `getifaddrs` plus `ifa_flags` rather than an `SIOCGIFFLAGS`
ioctl: it needs no socket, so the watcher can probe before any file
descriptor exists, and it is spelled the same on Linux and the BSDs.
### Interface Identity
The configured name is the key, but a name is not a device. Both backends
bind by *device* — `AF_PACKET` stores `sll_ifindex`, a BPF descriptor follows
the interface it was attached to — so an interface deleted and recreated
under the same name leaves the socket attached to something that no longer
exists while the name resolves perfectly well.
Nothing else notices. A stale `AF_PACKET` socket never becomes readable, so
the receive loop neither errors nor exits, and send failures go to the caller
rather than to the binder. A listen-only node (`announce: false`, so no
beacon sender to fail) therefore sat `present` and deaf indefinitely after a
`wifi reload` — the original bug wearing a different hat. The bound index is
captured at bind and compared on every poll; a mismatch is a detach.
Hardware can also change underneath a name. If the name reappears with a MAC
other than the one last bound, that is a different device, so the cached
neighbor entries for that transport are dropped rather than resumed onto, and
the swap is logged at `warn`. Richer selectors (`match: { name | mac |
id_path }`) are deliberately deferred; the requirement here is only that FIPS
never silently resumes onto different hardware.
### Detection
| Platform | Source |
| -------- | ------ |
| Linux | netlink `RTNLGRP_LINK` (`RTM_NEWLINK` / `RTM_DELLINK`) |
| macOS, FreeBSD† | `PF_ROUTE` socket, `RTM_IFINFO` |
| Fallback | poll `getifaddrs` + flags, 1 s |
† Aspirational: the Ethernet transport is
`cfg(any(target_os = "linux", target_os = "macos"))`, so FreeBSD has no
interface-bound transport for a watcher to serve. The `PF_ROUTE` branch
compiles for the BSD family, but only macOS reaches it.
Where an event source exists, detection is sub-second. The poll stays
underneath as a backstop rather than as the mechanism, and must stay at ~1 s:
the probe is cheap, and letting the interval drift to tens of seconds
reintroduces exactly the latency the event source was added to remove.
Construction is best-effort — a kernel or sandbox that refuses the socket
yields a watcher that never fires, and the binder degrades to its poll.
Link-event payloads are **not parsed**. An event is a hint to re-run the
presence probe, which is cheap and authoritative; decoding
`nlmsghdr`/`ifinfomsg` to reach the same answer would add a parser whose bugs
would be presence bugs.
Two rate limits protect the binder from its own event source. Probes are
coalesced to ten a second, because `PF_ROUTE` has no group filter and
delivers every routing message on the host — route churn, ARP, DHCP renewals,
a VPN going up and down — each of which would otherwise drive a full
`getifaddrs` walk. And a persistently failing event source is counted, logged
once, backed off, and after five consecutive errors abandoned for the poll:
losing events is survivable because the poll is the backstop, but burning a
core on a socket that is readable-but-erroring is not.
Bind failures that are *not* absence back off 1 s → 30 s. Absence itself does
not back off; there is nothing to poll but the probe.
### Policy: `optional`
One field per transport, `transports.ethernet.*.optional`, default `false`:
| | absence | log | retries |
| --- | --- | --- | --- |
| `optional: false` (default) | node reports `Degraded` | `info` at boot / `warn` on a runtime detach, then `error` once if it lasts past 10 s | forever |
| `optional: true` | no health impact | `info`, and nothing after | forever |
Naming an interface in configuration is a statement that you expect it, so
the default is to complain; silence is opted into.
`optional` describes **the interface's presence, not the transport's
importance**. An optional interface that is present is used exactly as hard
as any other. No value of it makes a missing interface fatal at startup: the
only fatal case remains "no transports at all came up". If a deployment ever
needs absence to abort startup, that arrives as an explicit `on_absent: exit`
— never as a second meaning for `optional`.
A bind failure that is not absence — no `CAP_NET_RAW`, no readable
`/dev/bpf*`, a buffer the kernel refused — is a fault, not a state, and still
fails the daemon's start. Retrying those forever would convert a hard,
actionable deployment error into a daemon that retries a socket it can never
open behind a `Degraded` nobody is watching. Only absence is waited out at
start; a non-absence failure during a later *rebind* does back off, since by
then the node is serving and killing it would be the worse answer.
### Health
Node health is recomputed on every presence transition **in both
directions**, so a returning interface clears `Degraded` without a restart.
That makes `Degraded` a level rather than a latch: the supervisor's reason set
was monotonic, which was correct while no child could recover, but with
recovery it would have come to mean "something broke at some point since boot"
rather than "something is broken now". Absence lives in its own reversible
set, separate from the one-way `failed` set a start failure enters.
An absent transport still counts as *up*. It came up — `start_async` returned
`Ok` — so it does not push a single-transport node into the fatal
`NoTransports`, which would make a node that merely booted before its wifi
exit instead of waiting. Absence degrades; it never kills.
There is deliberately **no restart action in the supervisor FSM.** The
presence watcher and the rebind loop live inside the transport, next to the
file descriptor they manage, and once that exists a supervisor-authored retry
has nothing left to do — it would be a second mechanism racing the first for
the same socket. The supervisor learns about presence
(`Event::ChildAbsent` / `Event::ChildPresent`) and republishes health; it does
not drive rebinding.
Peer state gets no grace period, and needs none. A transport's detach edge
withdraws every peer whose active link runs over it, on the same path the
liveness reaper uses, so those peers and the routes through them are gone at
the edge rather than up to `link_dead_timeout_secs` later — there is nothing
left for a linger timer to bound. A send over an absent interface still
returns `InterfaceUnavailable`, and a *half-built* link is still held on that
error, because the binder is already working to bring the interface back and
the initiator's resend has somewhere to land; an established peer is not held.
The trade is deliberate: an absence shorter than the dead timeout that then
recovers now costs a re-peer where it previously cost nothing, and that is
accepted because black-holing is silent, poisons other nodes' routing and
takes the full timeout to clear, where a re-peer is bounded, visible and
self-healing. A recreated mesh interface comes back with the same MAC (it is
derived from the phy), so the local address peers hold is unchanged across the
rebind.
### Logging
Edges, never attempts. A loop that logs per attempt reproduces the hot log
spin this mechanism removed, at 1–30 s intervals forever on any router with
an unplugged WAN — and operators learn to filter it, which is how the next
real failure gets missed.
The edge itself is not an error. An interface missing when the daemon starts
and bound a moment later is the ordinary case the mechanism exists to absorb,
so it is `info`; calling it an error at t=0 and "recovered" at t=0.2 s is the
cry-wolf failure this rule exists to prevent. A runtime detach is `warn` — a
link coming and going is ordinary weather for a mesh daemon.
There is exactly one deadline. Ten seconds is the window in which absence
could still be a race — a radio, a container, a veth arriving late. Past it a
**required** interface is a fault an operator has to fix, and it is reported
once at `error`. Start-time absence and a runtime detach share that deadline
rather than getting one each, for the same reason they share a code path
everywhere else here. An `optional` interface never reaches `error`; that is
what `optional` means.
Once, not repeated. This was a 1 m / 10 m / 1 h ladder that re-announced the
same fact at rising severity and then went permanently quiet after an hour,
which got both halves wrong: it used the log as a store for something already
published continuously as state, and it stopped mentioning a fault that was
still live. Duration belongs in `interface.since_secs` and in how long
`Degraded` has been held, where a monitor can threshold it per deployment
instead of the daemon compiling one in.
Node health does not wait for the deadline. `Degraded` publishes on the first
edge, which is the signal an operator actually watches.
Successful rebinds are damped. Backoff covers failed binds; the opposite and
nastier case is binds that keep *succeeding* into a socket that dies moments
later, which a receive loop giving up on a persistent error while the
interface stays `UP` produces once per second, forever. The binder counts
consecutive bindings that die inside ten seconds, backs off on the same
1 s → 30 s curve, and past three of them stops announcing each bind as a
recovery — holding health where it is until a binding lasts.
### Egress MTU
`transport_mtu()` is the minimum across *bound* transports — `is_bound()`,
not `is_operational()`, because an interface-bound transport is operational
from the moment it starts whether or not it holds a socket. Filtering on the
weaker predicate let a transport whose interface had never appeared set the
whole node's IPv6 MTU from hardware that was not present.
Since a transport can now bind long after start, that minimum moves at
runtime, and every consumer has to read it live. `show_status`, the
control-socket snapshot and the session-layer fragmentation check always did.
The TUN reader and writer did not: they were handed a `u16` at spawn, so a
narrow interface binding later never tightened the TCP MSS clamp and the node
reported one effective MTU while clamping to another. The ceiling is now
shared with those threads — an atomic beside the per-destination
`path_mtu_lookup` they already read on the same packet — and recomputed on
every change to the bound set.
Both directions, for the same reason `Degraded` is a level rather than a
latch: a narrow interface arriving must tighten the clamp or traffic
egressing over it is clamped too loose, and that interface leaving must
release it or unplugging a low-MTU adapter leaves the node over-clamped until
it restarts. MSS is negotiated per connection at SYN time, so a change binds
connections opened after it and leaves established ones alone.
### Observability
`show_transports` carries an `interface` block per interface-bound transport:
netdev name, `presence` (`absent` / `binding` / `present`), `carrier`,
`policy` (`required` / `optional`), `since_secs`, `binds` and
`failed_attempts`. The two counters separate an interface that is flapping
from one that is there and refusing to bind, and `since_secs` measures the
absence *episode* rather than the phase — a bind that fails walks
`Absent → Binding → Absent`, and restarting the clock on those edges would
report a permanently unbindable interface as one second old forever.
`fipstop`'s transports view names the netdev and the absence policy in their
own columns and shows presence in the State column for these transports,
because `state` reads `up` from the moment the transport starts and is
therefore precisely the wrong answer in the one case someone is scanning that
column for. Both render sites sort by ascending transport id — creation
order, and so grouped by transport type — rather than by `HashMap` iteration
order, which was arbitrary and differed on every daemon restart.
### What This Retires
- The `hotplug.d/net` rule that restarted the daemon when the FIPS radio
interfaces appeared, and the `wifi down; wifi up; sleep` dance provisioning
performed to sequence around the race.
- The YAML comment-toggling in `fips-mesh-setup` / `fips-ap-setup` — the mesh
and AP blocks ship enabled with `optional: true` and simply wait.
- The ad-hoc ENXIO socket reopen in the beacon sender: beacons now pause while
absent because the task does not exist then, and recovery is the presence
machine's job. One mechanism for every cause rather than one hack per
symptom.
- The start-time versus runtime asymmetry in the supervisor.
It also covers the case none of those workarounds did: an interface that flaps
while the daemon is running.
## Implementation Status
| Transport | Status | Notes |
+33 -17
View File
@@ -72,9 +72,9 @@ fips-mesh-setup radio1
This creates an open 802.11s interface with mesh ID `fips-mesh` and
HWMP forwarding off, attaches it to an unmanaged netifd interface (no
IP configuration — none is needed), uncomments the matching `meshN`
transport entry in `/etc/fips/fips.yaml` (see Step 2), and reloads the
radio. Interfaces are named by radio index: `radio0` → `fips-mesh0`,
IP configuration — none is needed), and reloads the radio. It does not
touch `/etc/fips/fips.yaml`: the matching `meshN` transport ships
enabled and the daemon binds the interface once it exists (see Step 2). Interfaces are named by radio index: `radio0` → `fips-mesh0`,
`radio1` → `fips-mesh1`. Pass a second argument to use a different
mesh ID.
@@ -133,18 +133,20 @@ wifi reload
## Step 2 — check the FIPS transport binding
The `fips.yaml` shipped in the OpenWrt package carries one transport
entry per radio, but **commented out** — so a stock install that never
runs this helper logs no per-boot "interface missing" warning.
`fips-mesh-setup` uncommented the matching `meshN` entry in Step 1, so
there is normally nothing to do here. If you maintain your own config
(or ran the manual UCI above instead of the helper), make sure the
entries are present and uncommented:
entry per radio, **enabled** and marked `optional: true`. The daemon
treats a named interface that is not there as absent rather than as a
failure, and `optional: true` is what keeps a stock install that never
runs this helper quiet and un-`Degraded` about a radio it was never
going to have. There is normally nothing to do here. If you maintain
your own config (or ran the manual UCI above instead of the helper),
make sure the entries are present:
```yaml
transports:
ethernet:
mesh0:
interface: "fips-mesh0"
optional: true
listen: true
announce: true
auto_connect: true
@@ -161,19 +163,33 @@ transports:
parses as an alias, so an existing config keeps working (see
[../reference/configuration.md](../reference/configuration.md)).
## Step 3 — restart the daemon (order matters)
## Step 3 — no restart needed
The daemon binds an interface when it appears. A transport whose
interface is missing is *absent*, not skipped: it waits, binds within
a second of the interface coming up, unbinds if it goes away, and
rebinds when it returns. Order does not matter, and neither
`/etc/init.d/fips restart` nor any hotplug rule is part of this
procedure.
Watch it happen:
```sh
fipsctl show transports
```
The transport's `interface` block reports `presence` (`absent` /
`binding` / `present`), `policy` (`required` / `optional`) and how
long it has held that state.
If you *changed a config value* above rather than only creating an
interface, that does need a restart — configuration is read at
startup, interfaces are not:
```sh
/etc/init.d/fips restart
```
Restart fips **after** the mesh interface is up. A transport whose
interface is missing at startup is logged and skipped, not retried —
so if the daemon comes up before the radio, the mesh transport stays
dead until the next restart. (An interface that *vanishes and
returns* after startup is recovered automatically; only the missing-
at-startup case needs this ordering.)
## Verify
L2 first — the 802.11s peering, with a second configured router in
+34 -17
View File
@@ -161,15 +161,17 @@ TCP 8443).
## Step 2 — check the FIPS transport binding
The `fips.yaml` shipped in the OpenWrt package carries one transport
entry per access interface, but **commented out** — so a stock install
that never runs this helper logs no per-boot "interface missing"
warning. `fips-ap-setup` uncommented the matching `apN` entry in Step 1,
and also enabled `node.rendezvous.lan` (the daemon's mDNS/DNS-SD
rendezvous — phone FIPS apps cannot see raw-Ethernet beacons, so mDNS
is how they find the daemon; the switch is daemon-wide and stays on if
you later remove the AP). So there is normally nothing to do here. If
you maintain your own config (or ran the manual UCI above instead of
the helper), make sure both are present and uncommented:
entry per access interface, **enabled** and marked `optional: true` —
the daemon binds the interface once `fips-ap-setup` creates it, and
`optional: true` keeps a stock install that never runs the helper quiet
and un-`Degraded`. `fips-ap-setup` does still edit one thing: it enables
`node.rendezvous.lan` (the daemon's mDNS/DNS-SD rendezvous — phone FIPS
apps cannot see raw-Ethernet beacons, so mDNS is how they find the
daemon; the switch is daemon-wide and stays on if you later remove the
AP). That one is a config value rather than an interface, so it needs a
restart to take effect. Otherwise there is normally nothing to do here.
If you maintain your own config (or ran the manual UCI above instead of
the helper), make sure both are present:
```yaml
node:
@@ -183,6 +185,7 @@ transports:
ethernet:
ap0:
interface: "fips-ap0"
optional: true
listen: true
announce: true
auto_connect: true
@@ -199,19 +202,33 @@ transports:
parses as an alias, so an existing config keeps working (see
[../reference/configuration.md](../reference/configuration.md)).
## Step 3 — restart the daemon (order matters)
## Step 3 — no restart needed
The daemon binds an interface when it appears. A transport whose
interface is missing is *absent*, not skipped: it waits, binds within
a second of the interface coming up, unbinds if it goes away, and
rebinds when it returns. Order does not matter, and neither
`/etc/init.d/fips restart` nor any hotplug rule is part of this
procedure.
Watch it happen:
```sh
fipsctl show transports
```
The transport's `interface` block reports `presence` (`absent` /
`binding` / `present`), `policy` (`required` / `optional`) and how
long it has held that state.
If you *changed a config value* above rather than only creating an
interface, that does need a restart — configuration is read at
startup, interfaces are not:
```sh
/etc/init.d/fips restart
```
Restart fips **after** the AP interface is up. A transport whose
interface is missing at startup is logged and skipped, not retried —
so if the daemon comes up before the radio, the access transport
stays dead until the next restart. (An interface that *vanishes and
returns* after startup is recovered automatically; only the missing-
at-startup case needs this ordering.)
## Verify
L2 and addressing first, with a phone or laptop connected to `!FIPS`:
+1 -1
View File
@@ -52,7 +52,7 @@ prints the response's `data` object as pretty JSON.
| `show mmp` | `show_mmp` | MMP metrics summary: per-peer link-layer metrics and per-session session-layer metrics. |
| `show cache` | `show_cache` | Coordinate cache: TTL, fill ratio, per-destination coords and path MTU. |
| `show connections` | `show_connections` | Pending handshake connections: state, idle time, resend count. |
| `show transports` | `show_transports` | Transport instances: type, state, MTU, local address, per-transport stats. |
| `show transports` | `show_transports` | Transport instances: type, state, MTU, local address, per-transport stats, and — for interface-bound transports — interface presence (`absent` / `binding` / `present`), carrier, absence policy (`required` / `optional`), time in the current phase, and bind/failed-attempt counts. |
| `show routing` | `show_routing` | Routing summary: pending lookups, retry state, forwarding/discovery/error/congestion counters. |
| `show identity-cache` | `show_identity_cache` | Cached `(node_addr → npub)` entries with last-seen timestamps. |
| `show native-flows` | `show_native_flows` | Native datagram API: open and pending flows with their ports, queue depth and age, bound listeners with their backlog, and the `native` counters. |
+21 -1
View File
@@ -41,7 +41,7 @@ query on its first activation and on every refresh tick while active.
| --- | ----- | ----- |
| **Node** | `show_status` (+ `show_listening_sockets`) | Identity, version, uptime, peer/link/session counts, sparklines for mesh size, tree depth, peer count, bytes, loss. The Traffic block on this tab is split: TUN counters on the left, the **Listening on fips0** panel on the right (see below). |
| **Peers** | `show_peers` (+ `show_links`, `show_transports` cross-refs) | Authenticated peers in a table. Selecting a row and pressing Enter opens a detail view. |
| **Transports** | `show_transports` (+ `show_links`, `show_peers` cross-refs) | Tree of transport instances with per-link children when expanded. |
| **Transports** | `show_transports` (+ `show_links`, `show_peers` cross-refs) | Tree of transport instances with per-link children when expanded. An interface-bound transport also carries an interface block; see [Interface block](#interface-block-transports-tab). |
| **Sessions** | `show_sessions` | End-to-end FSP sessions. |
| **Tree** | `show_tree` | Spanning-tree state and per-peer coordinates. |
| **Filters** | `show_bloom` | Per-peer Bloom-filter state. |
@@ -128,6 +128,26 @@ the keys the current context accepts.
| `e` | Expand all transports. |
| `c` | Collapse all transports. |
### Interface block (Transports tab)
A transport bound to a named interface carries an extra detail block.
It exists because the observability data shipped as JSON before it
reached this view, which left the live view reporting `up` for a
transport bound to nothing. The original OpenWrt failure was expensive
for that reason: the 802.11s link formed regardless, so nothing an
operator could see said the node was deaf.
| Field | Meaning |
| ----- | ------- |
| Interface | The interface name the instance is bound to. |
| Presence | Whether the interface is present, and for how long. |
| Carrier | Whether a present interface has carrier. Present without carrier is a distinct state. |
| Bound to | The address bound now, or nothing while absent. |
| On absence | The policy: `required` degrades the node, `optional` does not. |
| Binds, Failed binds | Counts over the instance's life, so a flapping interface reads as churn. |
**A transport with no interface has no block**, rather than an empty one.
### Multi-pane scrolling tabs (Tree, Filters, Routing)
Each lays out stacked panes that scroll independently.
+74
View File
@@ -681,6 +681,80 @@ root; on macOS it requires read/write access to a `/dev/bpf*` device.
| `auto_connect` | bool | `false` | Auto-connect to discovered peers |
| `accept_connections` | bool | `false` | Accept incoming connection attempts from discovered peers |
| `beacon_interval_secs` | u64 | `30` | Announcement beacon interval in seconds (minimum 10) |
| `optional` | bool | `false` | Whether absence of the interface is normal. See below |
**Dynamic binding.** The interface does not have to exist when the daemon
starts. A transport whose interface is missing comes up *absent*: it is not a
start failure, it is not skipped, and it binds on its own the moment the
interface appears — sub-second where the kernel offers link events (netlink on
Linux, `PF_ROUTE` on the BSDs), within a second otherwise. An interface that
goes away at runtime unbinds and rebinds by the same path, so a `wifi reload`
or an unplugged adapter needs no restart. Presence means `IFF_UP` — the
interface exists and is administratively up — and deliberately not
`IFF_RUNNING`: binding needs no carrier, and the socket keeps working across a
carrier flap without rebinding. A bridge with nothing plugged into it, such as
`br-lan` on a wifi-only router, is therefore bound and healthy rather than
permanently `Degraded`, and starts carrying traffic the moment a port comes up.
Whether an interface has carrier is reported separately, as `interface.carrier`
in `show_transports`.
`optional` selects how that absence is reported:
| | absence | log | retries |
| --- | --- | --- | --- |
| `optional: false` (default) | node reports `Degraded` | `info` at boot / `warn` on a runtime detach, then `error` once if it lasts past 10 s | forever |
| `optional: true` | no health impact | `info`, and nothing after | forever |
Naming an interface in configuration is a statement that you expect it, so the
default is to complain; silence is opted into. Set `optional: true` for
hardware that is legitimately not always there — a dock adapter, a radio only
some boards carry.
`optional` describes **the interface's presence, not the transport's
importance**. An optional interface that is present is used exactly as hard as
any other. No value of it makes a missing interface fatal at startup: the only
fatal case remains "no transports at all came up".
Either way the edge is logged once, on entering absence and on recovery —
never once per retry.
The edge itself is not an error. An interface missing when the daemon starts
and bound a moment later is the ordinary boot race this mechanism exists to
absorb, so it is `info`; a runtime detach is `warn`, because a link coming and
going is ordinary weather for a mesh daemon and a cable unplugged for two
seconds does not need a human.
Ten seconds is the whole grace. Past that it is no longer a race against a
radio or a container coming up, so a **required** interface still missing is
reported once at `error` and stays `Degraded` until it returns. Start-time
absence and a runtime detach share the one deadline — they are the same
transition throughout this mechanism. An `optional` interface never reaches
`error`; that is what `optional` means.
Said once, not repeated. How long the absence has lasted is a *state*, and it
is published as one: `interface.since_secs` in `show_transports`, and
`Degraded` for as long as it holds. Re-announcing it on a timer would put a
second, lossier copy of that in the log.
Node health does not wait for the ten seconds — `Degraded` is published on the
first edge, and that is the signal to watch.
If a binding keeps dying moments after it is established (a socket that errors
persistently while the interface stays up), the binder stops treating each
bind as a recovery: it backs off on the same 1 s → 30 s curve, holds node
health at its degraded reading, and stays quiet until a binding survives ten
seconds. Without that damping a broken socket produces a health flap and a log
pair every second, which is the same cry-wolf failure the edge-only logging
rule exists to prevent.
A bind failure that is **not** absence — no `CAP_NET_RAW`, no readable
`/dev/bpf*`, a buffer the kernel refused — is a fault, not a state, and fails
the daemon's start as it always has. Only a missing interface is waited out.
`fipsctl show transports` reports the current state per transport under
`interface`: `presence` (`absent` / `binding` / `present`), `carrier`,
`policy` (`required` / `optional`), `since_secs`, `binds`, and
`failed_attempts`.
**Named instances.** Multiple Ethernet interfaces can be configured by
using named sub-keys instead of flat parameters:
+1 -1
View File
@@ -124,7 +124,7 @@ table below lists every command currently registered.
| `show_mmp` | — | `peers[]` (link-layer per peer), `sessions[]` (session-layer per session). Each entry includes loss/RTT/ETX/goodput, smoothed values, trends. |
| `show_cache` | — | `count`, `max_entries`, `fill_ratio`, `default_ttl_ms`, `expired`, `avg_age_ms`, `entries[]` — per-destination coords, depth, age, last-used, optional `path_mtu`. |
| `show_connections` | — | `connections[]` — pending handshakes: `link_id`, `direction`, `handshake_state`, `started_at_ms`, `idle_ms`, `resend_count`, optional `expected_peer`. |
| `show_transports` | — | `transports[]` — `transport_id`, `type`, `state`, `mtu`, `name`, `local_addr`, optional `tor_mode`, `onion_address`, `tor_monitoring`, `stats`. |
| `show_transports` | — | `transports[]` — `transport_id`, `type`, `state`, `mtu`, `name`, `local_addr`, optional `tor_mode`, `onion_address`, `tor_monitoring`, `stats`, and `interface` for interface-bound transports. `interface` carries `name` (the configured netdev), `presence` (`absent` / `binding` / `present`), `carrier` (whether the link has `IFF_RUNNING` — reported only, never acted on: presence is `IFF_UP`, so a bound interface with no carrier is normal), `policy` (`required` / `optional`), `since_secs` (how long the current presence phase has been held), `binds` (successful binds since the transport was created — `1` after a clean start, more means it has rebound) and `failed_attempts` (failed binds since the last success). Absent entirely for transports that are not bound to a named interface, rather than reported as a permanently-`present` interface named `""`. Note that `state` describes the *transport* (`up` once started) and `interface.presence` describes the *socket*: an `up` transport whose interface is `absent` is started and waiting, which is a normal state and not a failure. |
| `show_routing` | — | `coord_cache_entries`, `identity_cache_entries`, `pending_lookups[]`, `pending_tun_destinations`, `pending_tun_packets`, `recent_requests`, `retries[]`, `forwarding`, `discovery` (request/response sub-counters; includes `req_deduplicated` — requests suppressed as recent duplicates — and `req_dedup_cache_full` — requests admitted because the dedup cache was full), `error_signals`, `congestion`. |
| `show_identity_cache` | — | `entries[]`, `count`, `max_entries`. Each entry: `node_addr`, `npub`, `display_name`, `ipv6_addr`, `last_seen_ms`, `age_ms`. |
| `show_native_flows` | — | `flows[]`, `listeners[]`, `stats` (the `native` counter family). Each flow: `flow_id`, `peer` (the peer's npub, which is its address; always present, because the flow carries the key its client named or its session authenticated), `peer_addr` (the 16-byte node address in hex — a truncated hash of the same key, kept because it is what `show_sessions` and `show_routing` key on), `local_port`, `remote_port`, `state` (`established` / `pending_accept`), `queued` (datagrams the node is holding for the flow), `age_ms` (time since the flow reached its current state: opened for a flow this node opened, accepted for one taken off a listener, announced for one still pending — accepting a pending flow restarts the clock). Each listener: `local_port`, `backlog`. |
+11
View File
@@ -223,6 +223,7 @@ is "all four flags on both ends."
> accept_connections: true
> dongle:
> interface: "enx00aabbccddee"
> optional: true
> announce: true
> # ...
> ```
@@ -231,6 +232,16 @@ is "all four flags on both ends."
> A single ground-up link only needs the flat form shown
> first; named instances become useful when the same node
> bridges multiple physical segments.
>
> `optional: true` on the dongle says its absence is normal —
> a USB adapter that is plugged in some days and not others.
> Without it, naming an interface is a statement that you
> expect it, and while it is missing the node reports
> `Degraded` and logs at `error`. Either way the interface
> does not have to exist when the daemon starts: a transport
> whose interface is missing waits and binds when it appears,
> and rebinds if it later goes away. Watch that with
> `fipsctl show transports`.
## Step 3: Grant the daemon permission to open raw sockets
+51 -47
View File
@@ -101,6 +101,9 @@ transports:
accept_connections: true
wwan:
interface: "phy0-sta0"
# Only exists while a radio is in station mode. Absence is normal, so
# it must not report the router Degraded.
optional: true
listen: true
announce: true
auto_connect: true
@@ -113,54 +116,55 @@ transports:
accept_connections: true
# 802.11s mesh backhaul between FIPS routers. These entries ship
# commented out so a stock install that never creates fips-mesh*
# logs no per-boot "interface missing" bind warning. Running
# 'fips-mesh-setup <radio>' creates the interface AND uncomments the
# matching block here (once per radio; radio0 -> fips-mesh0, radio1 ->
# fips-mesh1); 'fips-mesh-setup remove' re-comments it. Restart fips
# after — a transport whose interface is missing at startup is skipped,
# not retried. Dual-band routers can mesh on both bands at once —
# failover, not multipath: FIPS keeps one active link per peer, the
# other band stands by. The mesh runs OPEN (no SAE) with 802.11s
# forwarding off: FIPS's Noise handshake is the encryption and
# authentication, and FIPS is the routing layer. See
# docs/how-to/set-up-80211s-mesh-backhaul.md.
# mesh0:
# interface: "fips-mesh0"
# listen: true
# announce: true
# auto_connect: true
# accept_connections: true
# mesh1:
# interface: "fips-mesh1"
# listen: true
# announce: true
# auto_connect: true
# accept_connections: true
# ENABLED and marked 'optional: true'. The daemon treats a named
# interface that is not there as absent rather than as a failure: it
# waits, binds the moment 'fips-mesh-setup <radio>' creates the
# interface, and unbinds again if it goes away — no config edit and no
# restart, and 'optional: true' keeps a stock install that never runs
# that script quiet and un-Degraded.
#
# Dual-band routers can mesh on both bands at once — failover, not
# multipath: FIPS keeps one active link per peer, the other band stands
# by. The mesh runs OPEN (no SAE) with 802.11s forwarding off: FIPS's
# Noise handshake is the encryption and authentication, and FIPS is the
# routing layer. See docs/how-to/set-up-80211s-mesh-backhaul.md.
mesh0:
interface: "fips-mesh0"
optional: true
listen: true
announce: true
auto_connect: true
accept_connections: true
mesh1:
interface: "fips-mesh1"
optional: true
listen: true
announce: true
auto_connect: true
accept_connections: true
# Open "!FIPS" access SSID for phones and laptops running FIPS. These
# entries ship commented out so a stock install that never creates
# fips-ap* logs no per-boot "interface missing" bind warning. Running
# 'fips-ap-setup <radio>' creates the interface AND uncomments the
# matching block here (once per radio; radio0 -> fips-ap0, radio1 ->
# fips-ap1); 'fips-ap-setup remove' re-comments it. Restart fips after
# — a transport whose interface is missing at startup is skipped, not
# retried. The SSID is OPEN and isolated on purpose: FIPS's Noise
# handshake is the only security layer, and associated clients reach
# nothing but the FIPS handshake surface. See
# docs/how-to/set-up-open-access-ssid.md.
# ap0:
# interface: "fips-ap0"
# listen: true
# announce: true
# auto_connect: true
# accept_connections: true
# ap1:
# interface: "fips-ap1"
# listen: true
# announce: true
# auto_connect: true
# accept_connections: true
# Open "!FIPS" access SSID for phones and laptops running FIPS. Ships
# ENABLED and marked 'optional: true', for the same reason as the mesh
# blocks above: 'fips-ap-setup <radio>' creates the interface and the
# daemon binds it when it appears, with no config edit and no restart.
#
# The SSID is OPEN and isolated on purpose: FIPS's Noise handshake is the
# only security layer, and associated clients reach nothing but the FIPS
# handshake surface. See docs/how-to/set-up-open-access-ssid.md.
ap0:
interface: "fips-ap0"
optional: true
listen: true
announce: true
auto_connect: true
accept_connections: true
ap1:
interface: "fips-ap1"
optional: true
listen: true
announce: true
auto_connect: true
accept_connections: true
# Bluetooth Low Energy transport — requires BlueZ and the 'ble' feature.
# ble:
@@ -43,14 +43,17 @@
# access channels freely.
#
# The shipped /etc/fips/fips.yaml carries 'ap0' and 'ap1' entries under
# 'transports.ethernet' bound to these names, but commented out — a stock
# install that never creates fips-ap* then logs no bind warning. This
# helper uncomments the matching entry when it creates an interface and
# re-comments it on remove, so the daemon binds the transport without a
# manual config edit. It also uncomments the node.rendezvous.lan block
# (mDNS/DNS-SD — how phone FIPS apps discover the daemon); that switch is
# daemon-wide and stays on at remove. After an interface is up, restart
# fips.
# 'transports.ethernet' bound to these names, enabled and marked
# 'optional: true'. The daemon treats a named interface that is not there as
# ABSENT rather than as a start failure: it waits, binds the moment the
# interface appears, and unbinds again when it goes away — so creating the
# interface is all this helper has to do for the transport, with no config
# rewrite and no daemon restart.
#
# It does still edit one thing: the node.rendezvous.lan block (mDNS/DNS-SD —
# how phone FIPS apps discover the daemon). That is a config value rather than
# an interface, so it does need a restart to take effect. The switch is
# daemon-wide and stays on at remove.
# See docs/how-to/set-up-open-access-ssid.md for the full guide.
DEFAULT_SSID="!FIPS"
@@ -63,19 +66,14 @@ ap_config_write() {
chmod 600 "$CONFIG.tmp" && mv "$CONFIG.tmp" "$CONFIG"
}
# Uncomment the 'ap<idx>' transports.ethernet block in $CONFIG (created by
# 'fips-ap-setup'). Reversible with ap_config_disable. Returns:
# 0 enabled (or already active) 1 no config file 2 no such block
ap_config_enable() {
# Whether $CONFIG carries an enabled 'ap<idx>' transports.ethernet block.
# Reports only; the daemon owns the binding. Returns:
# 0 present 1 no config file 2 no such block
ap_config_present() {
idx="$1"
[ -f "$CONFIG" ] || return 1
grep -q "^ ap$idx:" "$CONFIG" && return 0
grep -q "^ # ap$idx:" "$CONFIG" || return 2
awk -v idx="$idx" '
$0 ~ ("^ # ap" idx ":[ \t]*$") { blk = 1; sub(/^ # /, " "); print; next }
blk && /^ # / { sub(/^ # /, " "); print; next }
{ blk = 0; print }
' "$CONFIG" > "$CONFIG.tmp" && ap_config_write
grep -q "^ ap$idx:" "$CONFIG" || return 2
return 0
}
# Uncomment the 'lan' block under node.rendezvous in $CONFIG — the daemon's
@@ -107,19 +105,6 @@ lan_rendezvous_enable() {
' "$CONFIG" > "$CONFIG.tmp" && ap_config_write
}
# Re-comment the 'ap<idx>' block so the daemon stops binding it (and stops
# warning about the now-missing interface). Inverse of ap_config_enable.
ap_config_disable() {
idx="$1"
[ -f "$CONFIG" ] || return 1
grep -q "^ ap$idx:" "$CONFIG" || return 0
awk -v idx="$idx" '
$0 ~ ("^ ap" idx ":[ \t]*$") { blk = 1; sub(/^ /, " # "); print; next }
blk && /^ / { sub(/^ /, " # "); print; next }
{ blk = 0; print }
' "$CONFIG" > "$CONFIG.tmp" && ap_config_write
}
usage() {
echo "Usage: fips-ap-setup <radio> [ssid]" >&2
echo " fips-ap-setup remove [radio]" >&2
@@ -154,10 +139,9 @@ if [ "$1" = "remove" ]; then
uci -q delete "network.$section"
uci -q delete "dhcp.$section"
uci -q del_list "firewall.fips_ap.network=$section"
# Re-comment the matching ap<N> transport in fips.yaml so the
# daemon stops warning about the interface we just removed.
idx="$(printf '%s' "$ifname" | sed -n 's/.*[^0-9]\([0-9]\{1,\}\)$/\1/p')"
[ -n "$idx" ] && ap_config_disable "$idx"
# The fips.yaml block stays as it is. The daemon notices the
# interface going away, unbinds, and waits for it — silently,
# because the block is marked 'optional: true'.
echo "Removed ${ifname:-$section}."
done
# Drop the shared zone and its rules once the last instance is gone.
@@ -356,12 +340,12 @@ wifi reload
/etc/init.d/odhcpd reload
/etc/init.d/firewall reload
# Enable the matching ap<N> transport in the shipped fips.yaml (it ships
# commented out). Tailor the restart hint to what we could do.
ap_config_enable "$IDX"
# Report whether the shipped fips.yaml still carries the matching transport.
# It ships enabled, so this is a check, not an edit.
ap_config_present "$IDX"
case $? in
0) TRANSPORT_NOTE="The ap$IDX transport in $CONFIG that binds '$AP_IFNAME' is
now uncommented and enabled." ;;
0) TRANSPORT_NOTE="The ap$IDX transport in $CONFIG binds '$AP_IFNAME'; the
daemon picks the interface up on its own, with no restart." ;;
1) TRANSPORT_NOTE="No $CONFIG found — add a transports.ethernet entry binding
interface '$AP_IFNAME' by hand." ;;
*) TRANSPORT_NOTE="No 'ap$IDX' entry in $CONFIG — add a transports.ethernet
@@ -398,9 +382,9 @@ per SSID, so it covers every FIPS router.
Next steps:
1. $TRANSPORT_NOTE
Watch it bind with: fipsctl show transports
2. $MDNS_NOTE
Restart the daemon AFTER the interface is up — a transport whose
interface is missing at startup is skipped, not retried:
mDNS is a config value, not an interface, so it needs a restart:
/etc/init.d/fips restart
3. Associate a phone or laptop running FIPS and verify:
iw dev $AP_IFNAME station dump
@@ -27,49 +27,26 @@
# Ethernet transport binds each directly and runs discovery beacons over it.
#
# The shipped /etc/fips/fips.yaml carries 'mesh0' and 'mesh1' entries under
# 'transports.ethernet' bound to these names, but commented out — a stock
# install that never creates fips-mesh* then logs no bind warning. This
# helper uncomments the matching entry when it creates an interface and
# re-comments it on remove, so the daemon binds the transport without a
# manual config edit. After an interface is up, restart fips.
# 'transports.ethernet' bound to these names, enabled and marked
# 'optional: true'. The daemon treats a named interface that is not there as
# ABSENT rather than as a start failure: it waits, binds the moment the
# interface appears, and unbinds again when it goes away. So this helper
# creates the interface and nothing else — no config rewrite, and no daemon
# restart. 'optional: true' is what keeps a stock install that never runs this
# script from reporting Degraded for an interface it was never going to have.
# See docs/how-to/set-up-80211s-mesh-backhaul.md for the full guide.
DEFAULT_MESH_ID="fips-mesh"
CONFIG="/etc/fips/fips.yaml"
# Replace $CONFIG with the rewritten $CONFIG.tmp. Force mode 0600 first: the
# package installs fips.yaml 0600 (it may hold an inline 'nsec' private key),
# and a fresh tmp file would otherwise land world-readable after the move.
mesh_config_write() {
chmod 600 "$CONFIG.tmp" && mv "$CONFIG.tmp" "$CONFIG"
}
# Uncomment the 'mesh<idx>' transports.ethernet block in $CONFIG (created by
# 'fips-mesh-setup'). Reversible with mesh_config_disable. Returns:
# 0 enabled (or already active) 1 no config file 2 no such block
mesh_config_enable() {
# Whether $CONFIG carries an enabled 'mesh<idx>' transports.ethernet block.
# Reports only; the daemon owns the binding. Returns:
# 0 present 1 no config file 2 no such block
mesh_config_present() {
idx="$1"
[ -f "$CONFIG" ] || return 1
grep -q "^ mesh$idx:" "$CONFIG" && return 0
grep -q "^ # mesh$idx:" "$CONFIG" || return 2
awk -v idx="$idx" '
$0 ~ ("^ # mesh" idx ":[ \t]*$") { blk = 1; sub(/^ # /, " "); print; next }
blk && /^ # / { sub(/^ # /, " "); print; next }
{ blk = 0; print }
' "$CONFIG" > "$CONFIG.tmp" && mesh_config_write
}
# Re-comment the 'mesh<idx>' block so the daemon stops binding it (and stops
# warning about the now-missing interface). Inverse of mesh_config_enable.
mesh_config_disable() {
idx="$1"
[ -f "$CONFIG" ] || return 1
grep -q "^ mesh$idx:" "$CONFIG" || return 0
awk -v idx="$idx" '
$0 ~ ("^ mesh" idx ":[ \t]*$") { blk = 1; sub(/^ /, " # "); print; next }
blk && /^ / { sub(/^ /, " # "); print; next }
{ blk = 0; print }
' "$CONFIG" > "$CONFIG.tmp" && mesh_config_write
grep -q "^ mesh$idx:" "$CONFIG" || return 2
return 0
}
usage() {
@@ -103,10 +80,9 @@ if [ "$1" = "remove" ]; then
ifname="$(uci -q get "wireless.$section.ifname")"
uci -q delete "wireless.$section"
uci -q delete "network.$section"
# Re-comment the matching mesh<N> transport in fips.yaml so the
# daemon stops warning about the interface we just removed.
idx="$(printf '%s' "$ifname" | sed -n 's/.*[^0-9]\([0-9]\{1,\}\)$/\1/p')"
[ -n "$idx" ] && mesh_config_disable "$idx"
# The fips.yaml block stays as it is. The daemon notices the
# interface going away, unbinds, and waits for it — silently,
# because the block is marked 'optional: true'.
echo "Removed ${ifname:-$section}."
done
uci commit wireless
@@ -114,7 +90,7 @@ if [ "$1" = "remove" ]; then
# 'wifi reload' re-applies the whole wireless config, so it briefly drops
# every client AP on all radios (a few seconds) — expected on remove.
wifi reload
echo "Restart fips: /etc/init.d/fips restart"
echo "No fips restart needed — the daemon unbinds the interface itself."
exit 0
fi
@@ -226,12 +202,12 @@ uci commit network
# every client AP on all radios (a few seconds) — expected when adding a mesh.
wifi reload
# Enable the matching mesh<N> transport in the shipped fips.yaml (it ships
# commented out). Tailor the restart hint to what we could do.
mesh_config_enable "$IDX"
# Report whether the shipped fips.yaml still carries the matching transport.
# It ships enabled, so this is a check, not an edit.
mesh_config_present "$IDX"
case $? in
0) TRANSPORT_NOTE="The mesh$IDX transport in $CONFIG that binds '$MESH_IFNAME' is
now uncommented and enabled." ;;
0) TRANSPORT_NOTE="The mesh$IDX transport in $CONFIG binds '$MESH_IFNAME'; the
daemon picks the interface up on its own." ;;
1) TRANSPORT_NOTE="No $CONFIG found — add a transports.ethernet entry binding
interface '$MESH_IFNAME' by hand." ;;
*) TRANSPORT_NOTE="No 'mesh$IDX' entry in $CONFIG — add a transports.ethernet
@@ -248,9 +224,9 @@ radio too — second band is a standby path (failover, not multipath).
Next steps:
1. $TRANSPORT_NOTE
Restart the daemon AFTER the interface is up — a transport whose
interface is missing at startup is skipped, not retried:
/etc/init.d/fips restart
No restart: the daemon binds an interface when it appears and
rebinds it if it goes away. Watch it happen with:
fipsctl show transports
2. Verify L2 peering with a second FIPS router in range:
iw dev $MESH_IFNAME station dump
and the FIPS link on top of it:
+374
View File
@@ -1313,3 +1313,377 @@ fn help_overlay_lists_keys() {
assert!(testkit::contains_row(&buf, "quit"));
assert!(testkit::contains_row(&buf, "Press ? or Esc to close"));
}
/// An interface-bound transport names its netdev in the table, so an operator
/// reading the list sees `br-lan` rather than an instance label that means
/// nothing outside the config file.
#[test]
fn transports_row_names_the_interface() {
let data = json!({
"transports": [{
"transport_id": 1,
"type": "ethernet",
"state": "up",
"mtu": 1497,
"name": "lan",
"interface": {
"name": "br-lan",
"presence": "present",
"carrier": true,
"policy": "required",
"since_secs": 3600,
"binds": 1,
"failed_attempts": 0
},
"stats": {}
}]
});
let mut app = app_with(Tab::Transports, data);
let buf = testkit::render(80, 20, |frame, area| {
super::transports::draw(frame, &mut app, area);
});
assert!(testkit::contains_row(&buf, "br-lan"));
// A bound interface reads as the transport state; only the exceptional
// case displaces it, or the annotation stops being read.
assert!(testkit::contains_row(&buf, "up"));
assert!(!testkit::contains_row(&buf, "absent"));
}
/// The narrow layout keeps the columns an operator is scanning for.
///
/// 80x24 is the OpenWrt serial console and the xterm/tmux default. The full
/// layout's fixed columns sum to 96 against ~77 usable, and ratatui resolves
/// an over-subscribed layout by shrinking *every* column — so the overflow
/// does not clip the rightmost one, it clips all of them, and
/// `mesh0 (optional)` became `mesh0 (optio`. The marker is the one thing on
/// that row worth reading.
///
/// This is the gap that let the regression through: the test asserting the
/// marker rendered at width 110, and the one rendering at 80 asserted only the
/// netdev name.
#[test]
fn the_optional_marker_survives_an_80_column_terminal() {
let data = json!({
"transports": [{
"transport_id": 2,
"type": "ethernet",
"state": "up",
"mtu": 1499,
"name": "mesh0",
"interface": {
"name": "fips-mesh0",
"presence": "absent",
"carrier": false,
"policy": "optional",
"since_secs": 252,
"binds": 0,
"failed_attempts": 0
},
"stats": {}
}]
});
let mut app = app_with(Tab::Transports, data);
let buf = testkit::render(80, 20, |frame, area| {
super::transports::draw(frame, &mut app, area);
});
assert!(
testkit::contains_row(&buf, "mesh0 (optional)"),
"the policy marker must not be the thing that clips at 80 columns"
);
assert!(testkit::contains_row(&buf, "fips-mesh0"));
assert!(testkit::contains_row(&buf, "absent"));
}
/// An absent interface displaces the transport state in the State column, and
/// its severity follows the absence policy.
///
/// `state` reads `up` from the moment the transport starts, whether or not it
/// is bound to anything, so the row would otherwise look identical to a
/// working one — the "looks healthy, reaches nothing" failure dynamic binding
/// exists to make visible.
///
/// Optional is a warning: a dock adapter that is not plugged in, or a radio
/// this board never had, is the case `optional: true` was added to describe,
/// and the daemon stays `Full` for it. Red there would train the operator to
/// ignore red.
#[test]
fn an_absent_optional_interface_warns() {
let data = json!({
"transports": [{
"transport_id": 2,
"type": "ethernet",
"state": "up",
"mtu": 1499,
"name": "mesh0",
"interface": {
"name": "fips-mesh0",
"presence": "absent",
"carrier": false,
"policy": "optional",
"since_secs": 252,
"binds": 0,
"failed_attempts": 0
},
"stats": {}
}]
});
let mut app = app_with(Tab::Transports, data);
let buf = testkit::render(110, 20, |frame, area| {
super::transports::draw(frame, &mut app, area);
});
assert!(testkit::contains_row(&buf, "fips-mesh0"));
// Policy rides with the instance name, and only the exception prints.
assert!(testkit::contains_row(&buf, "mesh0 (optional)"));
// The State column carries presence, because `up` is what it would
// otherwise say about an interface that has never existed.
assert!(testkit::contains_row(&buf, "absent"));
assert_eq!(
testkit::fg_at(&buf, "fips-mesh0"),
Some(ratatui::style::Color::Yellow),
"an interface whose absence is normal is a warning, not an error"
);
}
/// A required interface being absent is an error.
///
/// Naming an interface without `optional: true` is a statement that you expect
/// it, and the daemon reports `Degraded` while it is missing. The view has to
/// agree, or the colour stops carrying the same meaning as the health state.
#[test]
fn an_absent_required_interface_errors() {
let data = json!({
"transports": [{
"transport_id": 2,
"type": "ethernet",
"state": "up",
"mtu": 1499,
"name": "wan",
"interface": {
"name": "eth0",
"presence": "absent",
"carrier": false,
"policy": "required",
"since_secs": 30,
"binds": 0,
"failed_attempts": 0
},
"stats": {}
}]
});
let mut app = app_with(Tab::Transports, data);
let buf = testkit::render(110, 20, |frame, area| {
super::transports::draw(frame, &mut app, area);
});
// `required` is the default and is spelled by omission — a column of it
// on nearly every row would be a column of noise.
assert!(!testkit::contains_row(&buf, "required"));
assert!(!testkit::contains_row(&buf, "(optional)"));
assert!(testkit::contains_row(&buf, "wan"));
assert_eq!(
testkit::fg_at(&buf, "eth0"),
Some(ratatui::style::Color::Red),
"an interface the config says to expect is an error while it is gone"
);
}
/// The detail pane reports presence, carrier, policy and bind counts —
/// everything `state` cannot say.
#[test]
fn transport_detail_reports_interface_presence() {
let data = json!({
"transports": [{
"transport_id": 3,
"type": "ethernet",
"state": "up",
"mtu": 1497,
"name": "wan",
"interface": {
"name": "eth0",
"presence": "present",
"carrier": false,
"policy": "required",
"since_secs": 90,
"binds": 3,
"failed_attempts": 2
},
"stats": {}
}]
});
let mut app = app_with(Tab::Transports, data);
app.detail_view = Some(crate::app::DetailView { scroll: 0 });
let buf = testkit::render(120, 30, |frame, area| {
super::transports::draw(frame, &mut app, area);
});
assert!(testkit::contains_row(&buf, "Interface"));
assert!(testkit::contains_row(&buf, "eth0"));
assert!(testkit::contains_row(&buf, "present"));
// Carrier is reported rather than acted on, so "bound with no carrier" —
// a bridge with nothing plugged into it — has to be legible as its own
// state rather than inferred from silence.
assert!(testkit::contains_row(&buf, "Carrier"));
// A rebind count and a failed-bind count distinguish an interface that is
// flapping from one that is there and refusing.
assert!(testkit::contains_row(&buf, "rebound"));
assert!(testkit::contains_row(&buf, "Failed binds"));
}
/// Instance names and interfaces line up down the list.
///
/// Packed into one label they did not: `ethernet dongle2 en25` over
/// `ethernet wifi en0` left the netdev names ragged, which is the column an
/// operator scans to find the interface they are looking for. Separate facts,
/// separate columns.
#[test]
fn transports_columns_align_across_mixed_types() {
let data = json!({
"transports": [
{
"transport_id": 1, "type": "udp", "state": "up", "mtu": 1472,
"local_addr": "0.0.0.0:2121", "stats": {}
},
{
"transport_id": 2, "type": "tcp", "state": "up", "mtu": 1400,
"local_addr": "0.0.0.0:8443", "stats": {}
},
{
"transport_id": 3, "type": "ethernet", "state": "up", "mtu": 1497,
"name": "dongle2",
"interface": {
"name": "en25", "presence": "absent", "carrier": false,
"policy": "required", "since_secs": 12, "binds": 0,
"failed_attempts": 0
},
"stats": {}
},
{
"transport_id": 4, "type": "ethernet", "state": "up", "mtu": 1497,
"name": "wifi",
"interface": {
"name": "en0", "presence": "present", "carrier": true,
"policy": "optional", "since_secs": 900, "binds": 1,
"failed_attempts": 0
},
"stats": {}
}
]
});
let mut app = app_with(Tab::Transports, data);
let buf = testkit::render(100, 20, |frame, area| {
super::transports::draw(frame, &mut app, area);
});
let col_of = |needle: &str| testkit::find(&buf, needle).map(|(x, _)| x);
// The instance column starts at one x for every row that has one.
assert_eq!(col_of("dongle2"), col_of("wifi"));
// ... and so does the interface column, across long and short names and
// across transports that name a netdev and ones that name a socket.
assert_eq!(col_of("en25"), col_of("en0"));
assert_eq!(col_of("en25"), col_of("0.0.0.0:2121"));
assert_eq!(col_of("0.0.0.0:2121"), col_of("0.0.0.0:8443"));
// The header sits over the column it names.
assert_eq!(col_of("Bound to"), col_of("en25"));
assert_eq!(col_of("Instance"), col_of("dongle2"));
}
/// A link's remote address shares the `Bound to` column with its parent's
/// interface.
///
/// Same question — what is this attached to — so the same column: a netdev for
/// the transport, a remote endpoint for the link. Keeping a full MAC out of
/// the first column is also what lets the three identifying columns sit
/// against the left edge rather than being pushed right by the widest link
/// row.
#[test]
fn a_link_puts_its_remote_address_in_the_bound_to_column() {
let transports = json!({
"transports": [{
"transport_id": 3, "type": "ethernet", "state": "up", "mtu": 1497,
"name": "wifi",
"interface": {
"name": "en0", "presence": "present", "carrier": true,
"policy": "optional", "since_secs": 60, "binds": 1,
"failed_attempts": 0
},
"stats": {}
}]
});
let links = json!({
"links": [{
"link_id": 1, "transport_id": 3, "direction": "Outbound",
"remote_addr": "aa:bb:cc:dd:ee:ff", "state": "connected"
}]
});
let mut app = app_with(Tab::Transports, transports);
app.data.insert(Tab::Links, links);
app.expanded_transports.insert(3);
let buf = testkit::render(110, 20, |frame, area| {
super::transports::draw(frame, &mut app, area);
});
let col_of = |needle: &str| testkit::find(&buf, needle).map(|(x, _)| x);
// A right-aligned State clips from the left on overflow, so `connected`
// reading as `onnected` is a width bug rather than an honest truncation.
assert!(testkit::contains_row(&buf, "connected"));
// The MAC is not truncated and sits under the same header as the netdev.
assert!(testkit::contains_row(&buf, "aa:bb:cc:dd:ee:ff"));
assert_eq!(col_of("aa:bb:cc:dd:ee:ff"), col_of("en0"));
assert_eq!(col_of("Bound to"), col_of("en0"));
// The link keeps its direction and tree glyph in the first column, which
// is therefore narrow enough to leave the identifying columns at the left.
assert!(testkit::contains_row(&buf, "Out"));
// Packed at the left rather than floating in the middle. With the first
// column flexible, as it was, it absorbed every spare column of a wide
// terminal and pushed these two past the halfway mark.
assert!(col_of("Instance").unwrap() <= 20);
assert!(
col_of("Bound to").unwrap() < 45,
"the identifying columns must stay against the left edge"
);
}
/// `State` is right-aligned, so the values share a right edge and the column
/// reads as a status strip rather than as ragged text.
#[test]
fn the_state_column_is_right_aligned() {
let data = json!({
"transports": [
{
"transport_id": 1, "type": "udp", "state": "up", "mtu": 1472,
"local_addr": "0.0.0.0:2121", "stats": {}
},
{
"transport_id": 2, "type": "ethernet", "state": "up", "mtu": 1499,
"name": "mesh0",
"interface": {
"name": "fips-mesh0", "presence": "absent", "carrier": false,
"policy": "optional", "since_secs": 10, "binds": 0,
"failed_attempts": 0
},
"stats": {}
}
]
});
let mut app = app_with(Tab::Transports, data);
let buf = testkit::render(110, 20, |frame, area| {
super::transports::draw(frame, &mut app, area);
});
let end_of = |needle: &str| testkit::find(&buf, needle).map(|(x, _)| x + needle.len() as u16);
// Same right edge for a two-character value and a six-character one, and
// the header shares it.
assert_eq!(end_of("up"), end_of("absent"));
assert_eq!(end_of("up"), end_of("State"));
}
+280 -48
View File
@@ -1,5 +1,5 @@
use ratatui::Frame;
use ratatui::layout::{Constraint, Layout, Rect};
use ratatui::layout::{Alignment, Constraint, Layout, Rect};
use ratatui::style::{Color, Modifier, Style};
use ratatui::text::{Line, Span};
use ratatui::widgets::{
@@ -23,6 +23,21 @@ enum TreeRow {
},
}
/// Below this width the table drops its byte counters and keeps its
/// identifying columns.
///
/// The full layout's fixed columns sum to 90, plus six single-column gaps —
/// 96 against the ~77 usable inside an 80-column terminal's border and
/// scrollbar. Ratatui resolves an over-subscribed layout by shrinking every
/// column, so the overflow does not clip the rightmost column, it clips *all*
/// of them: `mesh0 (optional)` becomes `mesh0 (optio`, losing the one marker
/// on the row worth reading.
const NARROW_TABLE_WIDTH: u16 = 100;
/// Below this width the detail view stacks above/below the table instead of
/// beside it.
const SIDE_BY_SIDE_MIN_WIDTH: u16 = 110;
pub fn draw(frame: &mut Frame, app: &mut App, area: Rect) {
let transports = get_transports(app);
let links = get_links(app);
@@ -33,8 +48,18 @@ pub fn draw(frame: &mut Frame, app: &mut App, area: Rect) {
update_selected_tree_item(app, &tree_rows);
if app.detail_view.is_some() {
let chunks = Layout::horizontal([Constraint::Percentage(40), Constraint::Percentage(60)])
.split(area);
// Side by side only where the table half can still show its
// identifying columns. At 80 columns — the OpenWrt serial console, and
// the xterm/tmux default — a 40% split leaves the table 32 columns for
// a layout that needs 65 even in its narrow form, and ratatui resolves
// that by giving the trailing columns everything and rendering
// Transport, Instance, Bound-to and State at width zero. Stacking
// keeps both panes readable instead of keeping both unreadable.
let chunks = if area.width < SIDE_BY_SIDE_MIN_WIDTH {
Layout::vertical([Constraint::Percentage(50), Constraint::Percentage(50)]).split(area)
} else {
Layout::horizontal([Constraint::Percentage(40), Constraint::Percentage(60)]).split(area)
};
draw_table(frame, app, chunks[0], &transports, &links, &tree_rows);
draw_detail(frame, app, chunks[1], &transports, &links, &tree_rows);
@@ -113,6 +138,19 @@ fn update_selected_tree_item(app: &mut App, tree_rows: &[TreeRow]) {
};
}
/// Build a row, dropping the trailing byte counters in the narrow layout.
///
/// Both row shapes carry the same seven cells in the same order, so the
/// narrow variant is the same list with its tail cut — keeping one place
/// where the column count is decided, rather than two that must agree.
fn table_row<'a>(cells: Vec<Cell<'a>>, narrow: bool) -> Row<'a> {
let mut cells = cells;
if narrow {
cells.truncate(5);
}
Row::new(cells)
}
fn draw_table(
frame: &mut Frame,
app: &mut App,
@@ -121,14 +159,24 @@ fn draw_table(
links: &[serde_json::Value],
tree_rows: &[TreeRow],
) {
let header = Row::new(vec![
// Tx/Rx are the first thing to go when width is short: they are the only
// columns whose absence costs nothing an operator is scanning this table
// to find, and the detail pane carries them in full.
let narrow = area.width < NARROW_TABLE_WIDTH;
let mut header_cells = vec![
Cell::from("Transport / Link"),
Cell::from("State"),
Cell::from("Instance"),
Cell::from("Bound to"),
Cell::from(Line::from("State").alignment(Alignment::Right)),
Cell::from("Peer"),
Cell::from("Tx"),
Cell::from("Rx"),
])
.style(
];
if narrow {
header_cells.truncate(5);
}
let header = Row::new(header_cells).style(
Style::default()
.fg(Color::Yellow)
.add_modifier(Modifier::BOLD),
@@ -153,29 +201,72 @@ fn draw_table(
let typ = helpers::str_field(t, "type");
let name = t.get("name").and_then(|v| v.as_str()).unwrap_or("");
let addr = t.get("local_addr").and_then(|v| v.as_str()).unwrap_or("");
let label = if !name.is_empty() {
format!("{indicator}{typ} {name}")
} else if typ == "tor" {
// An interface-bound transport is identified by the netdev it
// names, not by its instance label: "ethernet lan" tells an
// operator nothing, "ethernet lan br-lan" tells them where to
// look. The presence marker is what makes the row honest —
// `state` reads `up` from the moment the transport starts,
// whether or not it is bound to anything.
let iface = t.get("interface");
let iface_name = iface
.and_then(|i| i.get("name"))
.and_then(|v| v.as_str())
.unwrap_or("");
let presence = iface
.and_then(|i| i.get("presence"))
.and_then(|v| v.as_str())
.unwrap_or("");
let policy = iface
.and_then(|i| i.get("policy"))
.and_then(|v| v.as_str())
.unwrap_or("");
// Three cells, not one packed string. The instance name and
// the thing the transport is bound to are separate facts about
// separate columns of a table, and running them together left
// the netdev names ragged down the list — the column an
// operator scans to find the interface they are looking for.
let label = if typ == "tor" {
let mode = t
.get("tor_mode")
.and_then(|v| v.as_str())
.unwrap_or("socks5");
let onion_hint = t
.get("onion_address")
format!("{indicator}tor({mode})")
} else {
format!("{indicator}{typ}")
};
// What this transport is attached to: a netdev for the
// interface-bound ones, the bound socket address for IP
// transports, an onion for tor. Different answers, one
// question, so one column.
let bound_to = if !iface_name.is_empty() {
iface_name.to_string()
} else if typ == "tor" {
t.get("onion_address")
.and_then(|v| v.as_str())
.map(|a| {
let short = if a.len() > 16 { &a[..16] } else { a };
format!(" {short}..")
format!("{short}..")
})
.unwrap_or_default();
format!("{indicator}tor({mode}){onion_hint}")
.unwrap_or_default()
} else if !addr.is_empty() {
format!("{indicator}{typ} {addr}")
addr.to_string()
} else {
format!("{indicator}{typ} #{transport_id}")
format!("#{transport_id}")
};
let state = helpers::str_field(t, "state");
// The State column carries presence for an interface-bound
// transport, not the lifecycle state. `up` is true from the
// moment the transport starts and stays true while its
// interface is missing, so it is precisely the wrong answer in
// the one case an operator is scanning this column for. There
// is no room to show both, and only one of them is news.
let state = if presence.is_empty() || presence == "present" {
helpers::str_field(t, "state")
} else {
presence
};
let tx = t
.get("stats")
.and_then(|s| s.get("packets_sent").or_else(|| s.get("frames_sent")))
@@ -189,14 +280,45 @@ fn draw_table(
.map(|n| n.to_string())
.unwrap_or_else(|| "-".into());
Row::new(vec![
Cell::from(label),
Cell::from(state.to_string()),
Cell::from(""),
Cell::from(tx),
Cell::from(rx),
])
.style(Style::default().fg(Color::White))
// Colour follows bindability, not lifecycle: an absent
// interface is the case the operator most needs to spot, and
// it is precisely the one `state` cannot show. Severity then
// follows the absence policy, because that is what the policy
// means — a dock adapter that is not plugged in is a warning,
// an interface the config says to expect is an error. Same
// split the daemon makes between `Degraded` and silence.
let row_style = match presence {
"" | "present" => Style::default().fg(Color::White),
"binding" => Style::default().fg(Color::Yellow),
_ if policy == "optional" => Style::default().fg(Color::Yellow),
_ => Style::default().fg(Color::Red),
};
// Policy rides with the instance name rather than owning a
// column: `required` is the default and appears on nearly
// every row, so a column of it is a column of noise. Only the
// exception is worth printing, and its absence then means the
// rule.
let instance = match (name.is_empty(), policy == "optional") {
(true, true) => "(optional)".to_string(),
(true, false) => String::new(),
(false, true) => format!("{name} (optional)"),
(false, false) => name.to_string(),
};
table_row(
vec![
Cell::from(label),
Cell::from(instance),
Cell::from(bound_to),
Cell::from(Line::from(state.to_string()).alignment(Alignment::Right)),
Cell::from(""),
Cell::from(tx),
Cell::from(rx),
],
narrow,
)
.style(row_style)
}
TreeRow::Link { index, is_last } => {
let link = &links[*index];
@@ -214,28 +336,42 @@ fn draw_table(
// Wide enough to render a full MAC (~17) or `hci0/MAC`
// (~22) without chopping mid-octet; the link detail view
// shows the untruncated address.
// The remote address goes in `Bound to`, not in the label. A
// link is bound to a remote endpoint exactly as a transport is
// bound to a netdev or a socket — same question, same column —
// and keeping a full MAC out of the first column is what lets
// the three left columns sit against the left edge instead of
// being pushed right by the widest link row.
let addr = helpers::truncate_hex(helpers::str_field(link, "remote_addr"), 24);
let label = format!(" {tree_char} {dir_short} {addr}");
let label = format!(" {tree_char} {dir_short}");
let state = helpers::str_field(link, "state");
let peer_name = lookup_peer_for_link(app, link)
.map(|p| helpers::str_field(&p, "display_name").to_string())
.unwrap_or_default();
Row::new(vec![
Cell::from(Span::styled(
label,
Style::default().fg(if dir == "Outbound" {
Color::Cyan
} else {
Color::Green
}),
)),
Cell::from(state.to_string()),
Cell::from(peer_name),
Cell::from(""),
Cell::from(""),
])
table_row(
vec![
Cell::from(Span::styled(
label,
Style::default().fg(if dir == "Outbound" {
Color::Cyan
} else {
Color::Green
}),
)),
// A link has no instance name of its own — it inherits its
// parent transport's, shown one row up — and no absence
// policy, which is a property of an interface.
Cell::from(""),
Cell::from(addr),
Cell::from(Line::from(state.to_string()).alignment(Alignment::Right)),
Cell::from(peer_name),
Cell::from(""),
Cell::from(""),
],
narrow,
)
}
})
.collect();
@@ -251,15 +387,44 @@ fn draw_table(
format!(" Transports ({transport_count}) ")
};
let widths = [
Constraint::Min(28),
Constraint::Length(12),
Constraint::Length(14),
Constraint::Length(9),
Constraint::Length(9),
// Every identifying column is fixed-width and packed against the left
// edge; `Peer` takes the slack. The first column used to be `Min`, which
// meant it absorbed all spare width and shoved Instance and Bound-to into
// the middle of the terminal, away from the names an operator is scanning.
//
// It is sized for the widest label that lives in it — a link's
// ` └─ Out` — rather than for a full MAC, because the address moved to
// `Bound to` where it belongs. `Bound to` is sized for a MAC (17), which
// also covers every netdev name and socket address that shares it.
// Narrow: the same columns, sized down to what still reads. Instance keeps
// 18 because `mesh0 (optional)` is 16 and the marker is the point; `Bound
// to` keeps 17 because that is a full MAC.
let widths: &[Constraint] = if narrow {
&[
Constraint::Length(12), // Transport / Link
Constraint::Length(18), // Instance, plus "(optional)"
Constraint::Length(17), // Bound to: a full MAC
Constraint::Length(9), // State, right-aligned
Constraint::Min(6), // Peer — takes the slack
]
} else {
&FULL_WIDTHS
};
const FULL_WIDTHS: [Constraint; 7] = [
Constraint::Length(18), // Transport / Link
Constraint::Length(20), // Instance, plus "(optional)" where it applies
Constraint::Length(18), // Bound to: netdev, socket addr, onion, MAC
// Wide enough for the longest value that lands here — `connected` (9)
// and `binding` — because a right-aligned cell clips from the LEFT,
// so an overflow reads as `onnected` rather than as a truncation.
Constraint::Length(10), // State, right-aligned
Constraint::Min(10), // Peer — takes the slack
Constraint::Length(7), // Tx
Constraint::Length(7), // Rx
];
let table = Table::new(rows, widths)
let table = Table::new(rows, widths.to_vec())
.header(header)
.block(Block::default().borders(Borders::ALL).title(title))
.row_highlight_style(
@@ -341,6 +506,73 @@ fn draw_transport_detail(frame: &mut Frame, app: &App, area: Rect, t: &serde_jso
lines.push(helpers::kv_line("Local Addr", addr));
}
// Interface presence, for the transports that are bound to a netdev.
//
// `State` above answers a lifecycle question — was this transport started
// — and reads `up` for an interface that has never existed. That gap is
// the whole reason interface binding is observable at all: the original
// OpenWrt bug was expensive because the 802.11s link formed regardless, so
// nothing an operator could see said the node was deaf. This is where they
// see it.
if let Some(iface) = t.get("interface") {
lines.push(Line::from(""));
lines.push(helpers::section_header("Interface"));
lines.push(helpers::kv_line(
"Interface",
helpers::str_field(iface, "name"),
));
let presence = helpers::str_field(iface, "presence");
let since = iface
.get("since_secs")
.and_then(|v| v.as_u64())
.map(|secs| helpers::format_duration_ms(secs.saturating_mul(1000)))
.unwrap_or_else(|| "-".into());
lines.push(helpers::kv_line(
"Presence",
&format!("{presence} for {since}"),
));
// Carrier is reported, never acted on: presence is IFF_UP, so a bound
// interface with no carrier is normal (a bridge with nothing plugged
// into it) rather than a fault. Saying so beats an operator inferring
// it from silence.
let carrier = iface
.get("carrier")
.and_then(|v| v.as_bool())
.map(|c| if c { "yes" } else { "no" })
.unwrap_or("-");
lines.push(helpers::kv_line("Carrier", carrier));
// The list marks only the exception, `(optional)`, beside the
// instance name. The detail pane has room to spell out both, as the
// consequence rather than the config key: `optional` is a statement
// about what absence *means*, and someone who has opened this pane
// wants the meaning.
let policy = helpers::str_field(iface, "policy");
let absence = if policy == "optional" {
"optional (absence is normal)"
} else {
"required (absence degrades the node)"
};
lines.push(helpers::kv_line("On absence", absence));
// Binds past the first are rebinds, and a climbing failed-attempt
// count is an interface that is there and refusing — a different
// problem from one that is missing, and invisible without this.
let binds = iface.get("binds").and_then(|v| v.as_u64()).unwrap_or(0);
if binds > 1 {
lines.push(helpers::kv_line("Binds", &format!("{binds} (rebound)")));
} else {
lines.push(helpers::kv_line("Binds", &binds.to_string()));
}
if let Some(failed) = iface.get("failed_attempts").and_then(|v| v.as_u64())
&& failed > 0
{
lines.push(helpers::kv_line("Failed binds", &failed.to_string()));
}
}
// Tor-specific info
if let Some(mode) = t.get("tor_mode").and_then(|v| v.as_str()) {
lines.push(helpers::kv_line("Tor Mode", mode));
+104 -60
View File
@@ -1057,6 +1057,68 @@ impl Config {
/// Validate cross-field configuration invariants.
pub fn validate(&self) -> Result<(), ConfigError> {
self.validate_ethernet_interfaces()?;
self.validate_rendezvous()
}
/// Reject interface names no kernel could ever hand back.
///
/// The presence machine deliberately cannot tell a typo from an interface
/// that has not been created yet — both are simply absent, and waiting is
/// the right answer for the second. That is what makes this check worth
/// having: a name that is *impossible* is the one case still separable
/// from "not there yet", and without it a typo costs a permanently
/// `Degraded` node whose only symptom is an interface that never arrives.
///
/// Syntax only. Whether a well-formed name exists is the binder's
/// question, asked once a second, forever.
fn validate_ethernet_interfaces(&self) -> Result<(), ConfigError> {
// Kernel limit: `IFNAMSIZ` is 16 including the terminating NUL, on
// both Linux and the BSDs.
const MAX_INTERFACE_NAME: usize = 15;
let mut seen: std::collections::HashMap<&str, &str> = std::collections::HashMap::new();
for (name, cfg) in self.transports.ethernet.iter() {
let label = name.unwrap_or("ethernet");
let iface = cfg.interface.as_str();
if iface.is_empty() {
return Err(ConfigError::Validation(format!(
"transport `{label}` has an empty `interface`"
)));
}
if iface.len() > MAX_INTERFACE_NAME {
return Err(ConfigError::Validation(format!(
"transport `{label}` interface `{iface}` is {} bytes; \
the kernel limit is {MAX_INTERFACE_NAME}, so no such \
interface can exist",
iface.len()
)));
}
if iface.contains('/') || iface.chars().any(char::is_whitespace) {
return Err(ConfigError::Validation(format!(
"transport `{label}` interface `{iface}` contains a \
character no interface name may hold"
)));
}
// Two transports on one netdev means two sockets on the same
// device at the same ethertype, each receiving every frame the
// other does.
if let Some(prior) = seen.insert(iface, label) {
return Err(ConfigError::Validation(format!(
"transports `{prior}` and `{label}` both bind interface \
`{iface}`"
)));
}
}
Ok(())
}
/// Cross-checks between transports, peers and the Nostr rendezvous.
fn validate_rendezvous(&self) -> Result<(), ConfigError> {
let nostr = &self.node.rendezvous.nostr;
let any_transport_advertises_on_nostr = self
@@ -1407,83 +1469,65 @@ node:
}
/// The fips.yaml shipped in the OpenWrt package must keep parsing as the
/// config schema evolves. Both the 802.11s mesh backhaul entries
/// config schema evolves, and must keep the absence policy it depends on.
///
/// The 802.11s mesh backhaul entries
/// (docs/how-to/set-up-80211s-mesh-backhaul.md) and the open-access SSID
/// entries (docs/how-to/set-up-open-access-ssid.md) ship commented out —
/// one per radio, so dual-band routers can run either on both bands — so
/// a stock install that never creates fips-mesh*/fips-ap* logs no
/// per-boot bind warning; `fips-mesh-setup`/`fips-ap-setup` uncomment the
/// matching block when they create the interface. Verify both states
/// parse: as shipped (both inactive), and after the uncomment the helpers
/// perform.
/// entries (docs/how-to/set-up-open-access-ssid.md) ship **enabled** — one
/// per radio, so dual-band routers can run either on both bands — and
/// marked `optional: true`. They used to ship commented out, with
/// `fips-mesh-setup`/`fips-ap-setup` uncommenting the matching block; that
/// was a workaround for a transport whose missing interface was skipped
/// for the life of the process. The daemon now waits for the interface and
/// binds it when it appears, and `optional: true` is what keeps a stock
/// install that never runs those helpers quiet and un-`Degraded` about a
/// radio it was never going to have.
#[test]
fn shipped_openwrt_config_parses() {
let yaml = include_str!("../../packaging/openwrt-ipk/files/etc/fips/fips.yaml");
// As shipped: parses, and the mesh/ap entries are commented out (a
// running daemon binds no fips-mesh*/fips-ap* transport, no warning).
let config: Config = serde_yaml::from_str(yaml).expect("shipped OpenWrt fips.yaml");
for name in ["mesh0", "mesh1", "ap0", "ap1"] {
assert!(
!config
.transports
.ethernet
.iter()
.any(|(n, _)| n == Some(name)),
"{name} must ship commented out, not active, in fips.yaml"
);
}
// What `fips-mesh-setup`/`fips-ap-setup` produce: uncomment each
// block, which must still parse into a transport bound to the right
// netdev.
let uncommented =
uncomment_transport_blocks(&uncomment_transport_blocks(yaml, "mesh"), "ap");
let config: Config = serde_yaml::from_str(&uncommented)
.expect("fips.yaml with mesh and ap transports uncommented");
let eth = |name: &str| {
config
.transports
.ethernet
.iter()
.find(|(n, _)| *n == Some(name))
.map(|(_, cfg)| cfg.clone())
};
// The interfaces the setup helpers create: present, bound to the right
// netdev, and optional. A regression to `optional: false` here would
// report every stock router `Degraded` for a mesh it never configured.
for (name, interface) in [
("mesh0", "fips-mesh0"),
("mesh1", "fips-mesh1"),
("ap0", "fips-ap0"),
("ap1", "fips-ap1"),
] {
let cfg = eth(name).unwrap_or_else(|| panic!("{name} entry missing from fips.yaml"));
assert_eq!(cfg.interface, interface, "{name} binds the wrong netdev");
assert!(cfg.optional(), "{name} must ship optional");
}
// `phy0-sta0` only exists while a radio is in station mode.
assert!(
eth("wwan").expect("wwan entry").optional(),
"wwan must ship optional"
);
// The wired ports exist on every supported board, so their absence is
// a real fault and must stay loud. Marking these optional too would
// make the whole ethernet block silent, which is the failure mode the
// presence mechanism exists to stop hiding.
for name in ["wan", "lan"] {
assert!(
config
.transports
.ethernet
.iter()
.any(|(n, eth)| n == Some(name) && eth.interface == interface),
"{name} entry missing after uncommenting shipped fips.yaml"
!eth(name).expect("wired entry").optional(),
"{name} must stay required"
);
}
}
/// Mirror the setup helpers' block uncomment: strip the ` # ` prefix
/// from each `# <prefix><N>:` header and its ` # ` continuation
/// lines, leaving every other comment untouched.
fn uncomment_transport_blocks(yaml: &str, prefix: &str) -> String {
let header = format!(" # {prefix}");
let mut out = String::new();
let mut in_block = false;
for line in yaml.lines() {
let is_header = line
.strip_prefix(&header)
.and_then(|r| r.strip_suffix(':'))
.is_some_and(|n| !n.is_empty() && n.bytes().all(|b| b.is_ascii_digit()));
if is_header {
in_block = true;
out.push_str(&line.replacen(" # ", " ", 1));
} else if in_block && line.starts_with(" # ") {
out.push_str(&line.replacen(" # ", " ", 1));
} else {
in_block = false;
out.push_str(line);
}
out.push('\n');
}
out
}
#[test]
fn test_parse_yaml_with_hex() {
let yaml = r#"
+60
View File
@@ -302,6 +302,25 @@ pub struct EthernetConfig {
/// Announcement beacon interval in seconds. Default: 30.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub beacon_interval_secs: Option<u64>,
/// Whether the absence of this interface is normal. Default: false.
///
/// Naming an interface in configuration is a statement that you expect it,
/// so the default is to complain: while the interface is missing the node
/// reports `Degraded`, the edge is logged (`info` at startup, `warn` on a
/// runtime detach), and an absence that outlasts the bring-up window — ten
/// seconds, past which it is no longer a race against a radio or a
/// container — is reported once at `error`. Set `optional: true` for
/// hardware that is legitimately not always there — a dock adapter, a
/// radio that only some boards carry — and its absence becomes silent
/// (`info` on the edge, no health impact, no error).
///
/// This describes **the interface's presence, not the transport's
/// importance**. An optional interface that is present is used exactly as
/// hard as any other; setting it does not deprioritize the transport, and
/// no value of this field makes a missing interface fatal at startup.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub optional: Option<bool>,
}
impl EthernetConfig {
@@ -340,6 +359,11 @@ impl EthernetConfig {
self.accept_connections.unwrap_or(false)
}
/// Whether absence of the interface is normal. Default: false.
pub fn optional(&self) -> bool {
self.optional.unwrap_or(false)
}
/// Get the beacon interval, clamped to minimum. Default: 30s.
pub fn beacon_interval_secs(&self) -> u64 {
self.beacon_interval_secs
@@ -1071,4 +1095,40 @@ mod tests {
serde_yaml::from_str("interface: eth0\nbogus: true\n");
assert!(bogus.is_err());
}
#[test]
fn ethernet_absence_is_an_error_unless_opted_out() {
// Naming an interface is a statement that you expect it, so the
// default has to be the loud one. A default of `true` here would make
// every missing interface silent, which is the failure mode the
// whole presence mechanism exists to stop hiding.
let bare: EthernetConfig = serde_yaml::from_str("interface: eth0\n").unwrap();
assert_eq!(bare.optional, None);
assert!(!bare.optional(), "absence must default to an error");
let opted: EthernetConfig =
serde_yaml::from_str("interface: enx00e04c680001\noptional: true\n").unwrap();
assert!(opted.optional());
let explicit: EthernetConfig =
serde_yaml::from_str("interface: eth0\noptional: false\n").unwrap();
assert!(!explicit.optional());
}
#[test]
fn ethernet_optional_survives_a_round_trip() {
// The packaged OpenWrt config ships `optional: true` on the mesh and
// access blocks; a serializer that dropped it would silently turn
// every stock router Degraded.
let opted: EthernetConfig =
serde_yaml::from_str("interface: fips-mesh0\noptional: true\n").unwrap();
let round: EthernetConfig =
serde_yaml::from_str(&serde_yaml::to_string(&opted).unwrap()).unwrap();
assert!(round.optional());
// ... and the default stays absent from the output rather than being
// written back as an explicit `false`.
let bare: EthernetConfig = serde_yaml::from_str("interface: eth0\n").unwrap();
assert!(!serde_yaml::to_string(&bare).unwrap().contains("optional"));
}
}
+90 -2
View File
@@ -1374,8 +1374,17 @@ pub(crate) fn show_connections_from_handle(
/// `show_transports` — Transport instances.
pub fn show_transports(node: &Node) -> Value {
let transports: Vec<Value> = node
.transport_ids()
// Ascending id, which is creation order: UDP, then Ethernet, then TCP,
// Tor, Nym, BLE, with each type's instances in config order. The map
// behind `transport_ids` is a `HashMap`, so without this the array order
// is whatever the hash seed produced — arbitrary, and different on every
// daemon restart. Anything scripting against this output, and every view
// rendering it, inherits that. Sorting by id groups the list by transport
// type for free, because the ids were handed out that way.
let mut ids: Vec<_> = node.transport_ids().copied().collect();
ids.sort_by_key(|id| id.as_u32());
let transports: Vec<Value> = ids
.iter()
.map(|id| {
let handle = node.get_transport(id).unwrap();
let mut t_json = json!({
@@ -1403,6 +1412,21 @@ pub fn show_transports(node: &Node) -> Value {
t_json["tor_monitoring"] = serde_json::to_value(&monitoring).unwrap_or_default();
}
// Interface presence, for the transports that have an interface.
// Absent from the payload entirely for the ones that do not, rather
// than reported as a permanently-`present` interface named "".
if let Some(p) = handle.interface_presence() {
t_json["interface"] = json!({
"name": handle.interface_name().unwrap_or_default(),
"presence": p.presence,
"carrier": p.carrier,
"policy": p.policy,
"since_secs": p.since_secs,
"binds": p.binds,
"failed_attempts": p.failed_attempts,
});
}
t_json["stats"] = handle.transport_stats();
t_json
@@ -1447,6 +1471,18 @@ pub(crate) fn show_transports_from_handle(handle: &super::read_handle::ControlRe
t_json["tor_monitoring"] = monitoring.clone();
}
if let Some(iface) = &t.interface {
t_json["interface"] = json!({
"name": iface.name,
"presence": iface.presence,
"carrier": iface.carrier,
"policy": iface.policy,
"since_secs": iface.since_secs,
"binds": iface.binds,
"failed_attempts": iface.failed_attempts,
});
}
t_json["stats"] = t.stats.clone();
t_json
@@ -2568,6 +2604,8 @@ mod tests {
"idle_ms",
"first_seen_secs_ago",
"last_contact_secs_ago",
// Interface presence: elapsed since the current phase began.
"since_secs",
];
/// Build a Node with a fixed identity, default config, and empty
@@ -2699,6 +2737,56 @@ mod tests {
// ---- 19 handler snapshot tests --------------------------------------
/// The `interface` block, which `build_test_node` cannot produce: it
/// keeps every transport list empty, so the nineteen snapshots above pin
/// `show_transports` only in its empty form. The block is emitted by two
/// hand-duplicated sites (the live handler and the read-handle variant)
/// that agree today with nothing enforcing it, and the control-socket
/// reference states the response schema is pinned by these snapshots.
///
/// A separate node rather than a richer `build_test_node`, so the other
/// snapshots keep their empty-state determinism.
#[cfg(any(target_os = "linux", target_os = "macos"))]
#[tokio::test]
async fn snapshot_show_transports_with_interface() {
use crate::config::EthernetConfig;
use crate::transport::ethernet::EthernetTransport;
use crate::transport::{TransportHandle, TransportId};
let mut node = build_test_node();
// An interface no host has, so presence is deterministically absent
// and carrier deterministically false on every machine this runs on.
let config = EthernetConfig {
interface: "fips-absent-x0".to_string(),
ethertype: None,
mtu: None,
recv_buf_size: None,
send_buf_size: None,
listen: Some(true),
announce: Some(false),
auto_connect: None,
accept_connections: None,
beacon_interval_secs: None,
optional: Some(false),
};
let (tx, _rx) = crate::transport::packet_channel(8);
let mut eth = EthernetTransport::new(TransportId::new(1), Some("lab".into()), config, tx);
// Started, because "up with its interface absent" is the state an
// operator actually meets — and the one whose shape is new here.
eth.start_async()
.await
.expect("absence is not a start failure");
node.insert_transport_for_test(TransportId::new(1), TransportHandle::Ethernet(eth));
assert_snapshot(
"show_transports_with_interface",
&render(show_transports(&node)),
);
}
#[test]
fn snapshot_show_status() {
let node = build_test_node();
+15
View File
@@ -778,6 +778,21 @@ pub(crate) struct TransportRow {
pub onion_address: Option<String>,
pub tor_monitoring: Option<serde_json::Value>,
pub stats: serde_json::Value,
/// Interface presence for interface-bound transports; `None` for the rest.
pub interface: Option<InterfaceRow>,
}
/// Interface name, presence and policy for an interface-bound transport, as
/// `show_transports` renders it.
#[derive(Clone, PartialEq)]
pub(crate) struct InterfaceRow {
pub name: String,
pub presence: &'static str,
pub carrier: bool,
pub policy: &'static str,
pub since_secs: u64,
pub binds: u64,
pub failed_attempts: u32,
}
/// MMP trend labels for a peer's link-layer block in `show_mmp` (each present
@@ -0,0 +1,36 @@
{
"data": {
"transports": [
{
"interface": {
"binds": 0,
"carrier": false,
"failed_attempts": 0,
"name": "fips-absent-x0",
"policy": "required",
"presence": "absent",
"since_secs": "<redacted>"
},
"mtu": 1496,
"name": "lab",
"state": "up",
"stats": {
"beacons_dropped": 0,
"beacons_recv": 0,
"beacons_sent": 0,
"bytes_recv": 0,
"bytes_sent": 0,
"frames_recv": 0,
"frames_sent": 0,
"frames_too_long": 0,
"frames_too_short": 0,
"recv_errors": 0,
"send_errors": 0
},
"transport_id": 1,
"type": "ethernet"
}
]
},
"status": "ok"
}
+25
View File
@@ -195,6 +195,31 @@ impl Node {
None => None,
};
if let Some(e) = send_err {
// A transient refusal is not a failed handshake. The
// interface under this transport is absent or
// mid-rebind, and the binder is already working to
// bring it back — so the half-built link is left
// exactly as it is for the initiator's msg1 resend to
// land on, rather than being torn down and rebuilt.
//
// Tearing down here charged a *local* interface flap
// to the remote: the reject counter it recorded means
// "the peer sent something invalid", which is a
// different thing entirely and one an operator reads
// as the peer's fault.
//
// Nothing leaks by staying. An initiator that never
// resends leaves a stale connection, which
// `check_timeouts` reaps at `handshake_timeout_secs`
// exactly as it reaps every other abandoned handshake.
if e.is_transient() {
debug!(
link_id = %link,
error = %e,
"Deferred msg2: the transport is between interfaces"
);
return;
}
// Restored pre-refactor msg2-send-failure warn!
// (`handle_msg1` L665): the send error text is surfaced
// at the executor point where the failure is now handled.
+75
View File
@@ -109,6 +109,16 @@ impl Node {
}
};
// Interface-presence receiver, or a dummy channel — same pattern and
// same reason as the child-liveness receiver above.
let (mut presence_rx, _presence_guard) = match self.transport_presence_rx.take() {
Some(rx) => (rx, None),
None => {
let (tx, rx) = tokio::sync::mpsc::channel(1);
(rx, Some(tx))
}
};
let tick_period = Duration::from_secs(self.config().node.tick_interval_secs);
let mut tick = tokio::time::interval(tick_period);
@@ -346,6 +356,71 @@ impl Node {
self.supervisor.state = ns;
}
}
// A transport child exiting leaves the bound set, so
// it can be the one that was holding the node's egress
// MTU down. `is_bound()` is `is_operational()` plus the
// presence refinement, and this moves the first half.
self.refresh_tun_mss_ceiling();
}
}
// Interface presence. An interface-bound transport's binder
// reports attach and detach; the FSM folds it into health.
// Unlike `ChildExited` this is reversible in both directions —
// the interface coming back republishes `Running` — which is
// the whole point of `Degraded` being a level rather than a
// latch.
maybe_presence = presence_rx.recv() => {
if let Some(edge) = maybe_presence {
// Health is policy-filtered; the MTU floor below is
// not. An `optional` interface's absence is normal and
// must not move the node off `Full`, but it changes
// the bound set all the same.
if edge.health_relevant {
let child = crate::node::lifecycle::supervisor::Child::Transport(
edge.transport_id,
);
let event = if edge.present {
crate::node::lifecycle::supervisor::Event::ChildPresent { child }
} else {
crate::node::lifecycle::supervisor::Event::ChildAbsent { child }
};
let actions = self.supervisor.fsm.step(event);
for action in actions {
if let crate::node::lifecycle::supervisor::Action::PublishState(
ns,
) = action
{
self.supervisor.state = ns;
}
}
}
// The bound set just changed, so the node's egress MTU
// floor may have. Both directions: an interface that
// binds can be the narrow one, and one that detaches
// can be the reason the clamp was tight.
self.refresh_tun_mss_ceiling();
// A peer reachable only through an interface that has
// gone is not reachable. Withdraw it now rather than
// leaving the liveness reaper to notice up to
// `link_dead_timeout_secs` later, during which this
// node both drops transit traffic in silence and keeps
// advertising reachability it does not have.
//
// Not policy-filtered: whether an interface's absence
// is normal is a statement about node *health*, not
// about whether the routes over it still work.
if !edge.present {
let reaped =
self.reap_peers_on_transport(edge.transport_id).await;
if reaped > 0 {
info!(
transport_id = %edge.transport_id,
peers = reaped,
"Withdrew peers whose interface went away"
);
}
}
}
}
Some(ipv6_packet) = tun_outbound_rx.recv() => {
+70 -7
View File
@@ -601,10 +601,72 @@ impl Node {
}
}
/// Route a link-dead liveness reap through the peer machine + executor.
/// The shell already decided (the tick sweep's `plan_heartbeats` batch
/// emitted this `ReapPeer` in phase order), so the machine only CONSUMES the
/// decision via [`PeerEvent::LinkDeadSuspected`]. The resulting executor arms
/// Reap every active peer reachable only through `transport_id`.
///
/// Called on a transport's detach edge. Until this existed, losing an
/// interface withdrew nothing: the peers stayed in the registry, the
/// routes through them stayed selectable, and this node kept advertising
/// reachability it no longer had — so transit traffic was dropped in
/// silence, and other nodes kept routing toward us for those destinations,
/// until the liveness reaper noticed up to `link_dead_timeout_secs` later.
/// Measured on real hardware that was 27 seconds of routing through a link
/// that had already gone, with four alternative peers available the whole
/// time.
///
/// The detach edge is both earlier and more certain than inactivity, so it
/// is the better trigger. This routes through the same
/// [`Self::route_link_dead`] the liveness reaper uses rather than
/// open-coding a second teardown — every consequence of losing a peer
/// (sessions, path MTU, session indices, the link, the control machine,
/// tree cleanup and re-announce, bloom withdrawal) already hangs off that
/// one path, and a parallel one would drift from it.
///
/// Deliberately undamped. A flapping interface cannot drive a reap storm
/// through here, because `ChurnGuard` stops publishing presence edges
/// after three short-lived bindings and does not resume until one lasts —
/// so the edges this reacts to are already rate-limited at the source, and
/// a second damper here would only add a way for the two to disagree.
///
/// Returns how many peers were reaped.
pub(in crate::node) async fn reap_peers_on_transport(
&mut self,
transport_id: TransportId,
) -> usize {
let doomed: Vec<NodeAddr> = self
.peers
.iter()
.filter(|(_, peer)| peer.transport_id() == Some(transport_id))
.map(|(node_addr, _)| *node_addr)
.collect();
if doomed.is_empty() {
return 0;
}
let now_ms = std::time::SystemTime::now()
.duration_since(std::time::UNIX_EPOCH)
.map(|d| d.as_millis() as u64)
.unwrap_or(0);
let reaped = doomed.len();
for node_addr in doomed {
debug!(
peer = %self.peer_display_name(&node_addr),
%transport_id,
"Removing peer: its interface went away"
);
self.route_link_dead(node_addr, now_ms).await;
}
reaped
}
/// Route a link-dead reap through the peer machine + executor. Two callers
/// decide: the tick sweep's `plan_heartbeats` batch emits a `ReapPeer` for a
/// peer that has gone quiet, and [`Self::reap_peers_on_transport`] withdraws
/// a transport's peers when its interface goes away. Mirrors
/// [`route_rekey_cadence`](Node::route_rekey_cadence): the shell has already
/// decided by the time this runs, so the machine only CONSUMES the decision
/// via [`PeerEvent::LinkDeadSuspected`]. The resulting executor arms
/// (`InvalidateSendState` → `remove_active_peer`, `ReportLost` →
/// `note_link_dead`) reproduce the pre-refactor inline reap body exactly, in
/// that order.
@@ -614,9 +676,10 @@ impl Node {
/// no-op, so we return; if the machine is absent (which should be impossible)
/// we fall back to the byte-identical inline body under a `debug_assert`.
///
/// `now_ms` is the sweep's hoisted wall-clock ms (the same value the old reap
/// fed `note_link_dead`); it flows to the executor `ReportLost` arm via
/// `ambient.now_ms`.
/// `now_ms` is the caller's hoisted wall-clock ms (the same value the old
/// reap fed `note_link_dead`); it flows to the executor `ReportLost` arm via
/// `ambient.now_ms`. Both callers hoist it once per batch, so every peer
/// removed in one pass carries the same instant.
async fn route_link_dead(&mut self, node_addr: NodeAddr, now_ms: u64) {
let link = match self.peers.get(&node_addr) {
Some(peer) => peer.link_id(),
+20 -5
View File
@@ -385,11 +385,26 @@ impl Node {
);
}
Err(e) => {
warn!(
peer = %self.peer_display_name(node_addr),
error = %e,
"Failed to send rekey msg1"
);
// The teardown here is already benign — this returns before
// `set_rekey_state`, so the cycle simply does not start and
// is retried when rekey next comes due, with nothing torn
// down and nothing charged to the peer. Only the severity
// is wrong for a transport between interfaces, which is a
// local and self-clearing condition the presence machine
// has already reported.
if e.is_transient() {
debug!(
peer = %self.peer_display_name(node_addr),
error = %e,
"Deferred rekey msg1: the transport is between interfaces"
);
} else {
warn!(
peer = %self.peer_display_name(node_addr),
error = %e,
"Failed to send rekey msg1"
);
}
let _ = self.index_allocator.free(our_index);
self.stats_mut()
.record_reject(RejectReason::Handshake(HandshakeReject::BadState));
+129 -38
View File
@@ -1000,15 +1000,37 @@ impl Node {
let connected = self.peers.contains_key(&node_addr);
if connected {
// Active peer: skip a candidate whose path is already the
// current, still-fresh one (avoid churning a healthy link).
let transport_name = transport.transport_type().name;
let peer_addr_candidate =
PeerAddress::new(transport_name, remote_addr.to_string());
if self.active_peer_candidate_is_fresh_enough_to_skip(
&node_addr,
std::slice::from_ref(&peer_addr_candidate),
) {
// Active peer: skip every candidate while the link we
// already hold is live — the current path *and* any
// alternate one.
//
// Only the same-path case used to be skipped, which left
// the stated intent ("avoid churning a healthy link")
// covering exactly the case that could not churn anything.
// A peer reachable twice — the ordinary result of two
// machines sharing a LAN and a cable, since each beacons on
// both — was therefore re-dialled on its alternate path
// every discovery tick, forever. Each dial that completed
// promoted and displaced the incumbent, so the peer's link
// migrated back and forth on a fixed cadence, tearing down
// and re-establishing its session each time. Measured on
// real hardware: seventeen dials to one peer in fifteen
// minutes, alternating wifi and cable, displacing a link
// reporting `etx = 1.0` and `loss = 0.0`.
//
// When that peer is the parent — which the best path
// usually is — every migration also switched parents,
// invalidating the downstream coordinate cache and
// re-announcing to every peer. The cost of the churn was
// therefore mesh-wide while the benefit was nil: the link
// being replaced was already perfect.
//
// Failover is unaffected. Liveness is the gate, so a peer
// that stops answering goes stale within a heartbeat
// interval and every path, alternate included, is dialled
// again. What is given up is switching away from a link
// that is working, which is not a thing worth doing.
if self.active_peer_link_is_live(&node_addr) {
continue;
}
if self.is_connecting_to_peer_on_path(
@@ -1617,6 +1639,17 @@ impl Node {
self.child_exit_tx = Some(child_exit_tx);
self.child_exit_rx = Some(child_exit_rx);
// Interface-presence channel. Created before `create_transports` so
// every interface-bound transport gets the sender at construction and
// its very first bind attempt — the one `start_async` makes inline —
// is already reportable. A boot race therefore reaches the FSM while
// it is still `Starting`, and start-completion health resolves to
// `Degraded` on the first publish rather than publishing `Full` and
// correcting it a moment later.
let (presence_tx, presence_rx) = tokio::sync::mpsc::channel(16);
self.transport_presence_tx = Some(presence_tx);
self.transport_presence_rx = Some(presence_rx);
// Initialize transports first (before TUN, before Nostr discovery).
// Creation allocates each transport's id; the supervisor FSM authors
// the start order over those ids.
@@ -1878,12 +1911,20 @@ impl Node {
info!(" address: {}", device.address());
info!(" mtu: {}", mtu);
// Calculate max MSS for TCP clamping
// Seed the shared MSS ceiling from whatever is bound
// right now. Both TUN threads read it live from here
// on, so a transport binding or unbinding later moves
// the clamp instead of leaving it at this instant's
// value — see `crate::upper::tun::MssCeiling`.
self.refresh_tun_mss_ceiling();
let max_mss = self.tun_mss_ceiling.clone();
let effective_mtu = self.effective_ipv6_mtu();
let max_mss = effective_mtu.saturating_sub(40).saturating_sub(20); // IPv6 + TCP headers
info!("effective MTU: {} bytes", effective_mtu);
debug!(" max TCP MSS: {} bytes", max_mss);
debug!(
" max TCP MSS: {} bytes",
max_mss.load(std::sync::atomic::Ordering::Relaxed)
);
// On macOS and FreeBSD, create a shutdown pipe. Writing to it
// unblocks the reader thread's select() loop without closing
@@ -1907,8 +1948,8 @@ impl Node {
// Create writer (dups the fd for independent write access).
// Pass path_mtu_lookup so inbound SYN-ACK clamp can read
// per-destination path MTU learned via discovery.
let (writer, tun_tx) =
device.create_writer(max_mss, self.path_mtu_lookup.clone())?;
let (writer, tun_tx) = device
.create_writer(max_mss.clone(), self.path_mtu_lookup.clone())?;
// Spawn writer thread. On exit it self-reports
// `Child::Tun` (sync context → `blocking_send`); TUN
@@ -1934,7 +1975,6 @@ impl Node {
// self-reports `Child::Tun` on exit (sync context →
// `blocking_send`). Exactly one cfg variant compiles,
// so the single clone is moved into that closure.
let transport_mtu = self.transport_mtu();
let path_mtu_lookup = self.path_mtu_lookup.clone();
let reader_child_tx = self.child_exit_tx.clone();
#[cfg(any(target_os = "macos", target_os = "freebsd"))]
@@ -1945,7 +1985,7 @@ impl Node {
our_addr,
reader_tun_tx,
outbound_tx,
transport_mtu,
max_mss,
path_mtu_lookup,
shutdown_read_fd,
);
@@ -1961,7 +2001,7 @@ impl Node {
our_addr,
reader_tun_tx,
outbound_tx,
transport_mtu,
max_mss,
path_mtu_lookup,
);
if let Some(tx) = &reader_child_tx {
@@ -2092,6 +2132,16 @@ impl Node {
}
};
// Drain any presence edges this child's start produced *before*
// reporting the child itself. An interface-bound transport whose
// interface is missing reports absence from inside `start_async`
// and then reports `SubstrateUp` (absence is a state, not a start
// failure), so ordering the drain first means start-completion
// health already knows about the absence when `pending` empties.
// Otherwise a boot race publishes `Full` and corrects itself a
// moment later, and every consumer sees a spurious transition.
let _ = self.drain_transport_presence();
let feedback_actions = self.supervisor.fsm.step(feedback);
for action in &feedback_actions {
if let Action::PublishState(ns) = action {
@@ -2100,6 +2150,12 @@ impl Node {
}
}
// Late edges: a transport that bound after its `SubstrateUp` was
// reported, or one that detached during a later child's bring-up.
if let Some(ns) = self.drain_transport_presence() {
start_outcome = Some(ns);
}
// Seams that never triggered inside the loop: the "Transports
// initialized" info! when there was no non-transport child, and the
// peer-connect when there was no Tun/Dns child (today it still runs,
@@ -2138,8 +2194,10 @@ impl Node {
// children. Enumerate them for the operator, then proceed —
// a degraded node serves traffic.
warn!(
degraded_children = ?self.supervisor.fsm.failed(),
"Node started DEGRADED: one or more configured optional children failed to start"
degraded_children = ?self.supervisor.fsm.degraded_children(),
absent_interfaces = ?self.supervisor.fsm.absent(),
"Node started DEGRADED: one or more configured optional children failed to \
start, or a configured interface is absent"
);
}
_ => {}
@@ -2498,6 +2556,50 @@ impl Node {
}
}
/// Feed every queued interface-presence edge to the supervisor FSM,
/// returning the last [`NodeState`] it asked to publish (if any).
///
/// Non-blocking: it drains what is already queued and returns. Used during
/// bring-up, where the rx_loop's presence arm is not running yet — from
/// then on that arm owns the same translation.
pub(in crate::node) fn drain_transport_presence(&mut self) -> Option<NodeState> {
let mut edges = Vec::new();
if let Some(rx) = self.transport_presence_rx.as_mut() {
while let Ok(edge) = rx.try_recv() {
edges.push(edge);
}
}
let mut published = None;
let saw_edge = !edges.is_empty();
for edge in edges {
// Health is policy-filtered; the MTU refresh below is not. See
// `saw_edge`.
if !edge.health_relevant {
continue;
}
let child = Child::Transport(edge.transport_id);
let event = if edge.present {
Event::ChildPresent { child }
} else {
Event::ChildAbsent { child }
};
for action in self.supervisor.fsm.step(event) {
if let Action::PublishState(ns) = action {
published = Some(ns);
}
}
}
if saw_edge {
// A bind or unbind changes which transports are bound, and so the
// node's egress MTU floor. During bring-up this runs before the
// TUN threads exist, which is exactly when it must: they read the
// ceiling this leaves behind.
self.refresh_tun_mss_ceiling();
}
published
}
/// Reconstruct the supervised up-set from observed runtime presence, so the
/// FSM authors the teardown order regardless of how the node reached
/// `Running`. Worker pools are deliberately excluded: today's teardown never
@@ -3425,14 +3527,13 @@ impl Node {
candidates
}
pub(in crate::node) fn active_peer_candidate_is_fresh_enough_to_skip(
&self,
peer_node_addr: &NodeAddr,
candidates: &[PeerAddress],
) -> bool {
if !self.active_peer_matches_any_candidate(peer_node_addr, candidates) {
return false;
}
/// Whether the link we already hold to this peer is answering.
///
/// The gate on dialling an active peer at all. Phrased as liveness rather
/// than as a property of the candidate, because the candidate's path is
/// not the question: a live link should not be replaced by *any* path,
/// and a dead one should be replaced by whichever path answers.
pub(in crate::node) fn active_peer_link_is_live(&self, peer_node_addr: &NodeAddr) -> bool {
!self.active_peer_needs_same_path_refresh(peer_node_addr)
}
@@ -3449,17 +3550,7 @@ impl Node {
peer.idle_time(Self::now_ms()) > stale_after_ms
}
fn active_peer_matches_any_candidate(
&self,
peer_node_addr: &NodeAddr,
candidates: &[PeerAddress],
) -> bool {
candidates
.iter()
.any(|candidate| self.active_peer_matches_candidate(peer_node_addr, candidate))
}
fn active_peer_matches_candidate(
pub(in crate::node) fn active_peer_matches_candidate(
&self,
peer_node_addr: &NodeAddr,
candidate: &PeerAddress,
+434 -12
View File
@@ -67,6 +67,39 @@
//! when a task/thread dies at runtime) is **deferred**: start-completion health
//! resolution is start-framed, and liveness monitoring is a substantial unbuilt
//! mechanism. This commit is start-time health only.
//!
//! ## Scope: interface presence, and `Degraded` as a level (this commit)
//!
//! Interface-bound transports are now *sometimes bound*: a transport whose
//! interface is missing at start comes up [`Absent`] and binds later, and one
//! whose interface goes away at runtime unbinds and rebinds when it returns.
//! Two things follow for this machine.
//!
//! - **A second reason set.** [`Event::ChildAbsent`] / [`Event::ChildPresent`]
//! move a child in and out of `absent`, which feeds `Degraded` exactly like
//! `failed` does. It is kept separate because it is *reversible* and `failed`
//! is not: a child that failed to start stays failed for the bring-up, while
//! an absent interface is expected to come back.
//! - **`Degraded` is a level, not a latch.** Nothing ever removed from `failed`,
//! which was correct while no child could recover — a monotonic set accurately
//! described a one-way door. Once recovery exists the assumption inverts: plug
//! the WAN back in and the node would stay `Degraded` until the process
//! restarted, and `Degraded` would come to mean "something broke at some point
//! since boot" rather than "something is broken now".
//! [`Self::classify_health`](SupervisorFsm::classify_health) is therefore
//! recomputed on every transition **in both directions**.
//!
//! An absent transport still counts as *up*. It came up — `start_async` returns
//! `Ok` with the transport absent — so it does not push a single-transport node
//! into the fatal [`FailReason::NoTransports`], which would make a node that
//! merely booted before its wifi exit instead of waiting. Absence degrades; it
//! never kills.
//!
//! There is deliberately **no restart action**. Rebinding is owned by the
//! transport's own binder task, which is where the file descriptor and the
//! presence watcher live; the FSM is told what happened and republishes health.
//!
//! [`Absent`]: crate::transport::ethernet::Presence::Absent
use std::collections::HashSet;
use std::sync::Arc;
@@ -174,6 +207,22 @@ pub(crate) enum Event {
/// The child whose task/thread exited.
child: Child,
},
/// An interface-bound child lost its interface — it was never there at
/// start, or it went away at runtime. The child stays *up* (the transport
/// object survives detach, only its socket goes) but contributes
/// `Degraded`. Valid while `Running`; while `Starting` the edge is recorded
/// so start-completion health already reflects it.
ChildAbsent {
/// The child whose interface is absent.
child: Child,
},
/// An interface-bound child's interface came back and it rebound. Clears
/// the absence and republishes health, which is how `Degraded` becomes
/// reversible.
ChildPresent {
/// The child whose interface is present again.
child: Child,
},
}
/// A driver-scheduled timer the supervisor can arm. Only the
@@ -320,7 +369,16 @@ pub(crate) struct SupervisorFsm {
up: HashSet<Child>,
/// Configured children that failed to start during the current bring-up.
/// Feeds the `Degraded` health determination when `pending` empties.
///
/// One-way within a bring-up: a start failure is not recoverable, so
/// nothing removes from this set until the next `Start`.
failed: HashSet<Child>,
/// Children that are up but whose network interface is currently absent.
///
/// Reversible, unlike [`Self::failed`] — that is the whole reason it is a
/// second set rather than more entries in the first. Feeds `Degraded` the
/// same way, and empties as interfaces come back.
absent: HashSet<Child>,
}
impl SupervisorFsm {
@@ -330,6 +388,7 @@ impl SupervisorFsm {
state: SupState::Created,
up: HashSet::new(),
failed: HashSet::new(),
absent: HashSet::new(),
}
}
@@ -348,6 +407,7 @@ impl SupervisorFsm {
},
up: up.into_iter().collect(),
failed: HashSet::new(),
absent: HashSet::new(),
}
}
@@ -357,13 +417,26 @@ impl SupervisorFsm {
&self.state
}
/// The configured children that failed to start during bring-up. The driver
/// reads this on the `Degraded` start outcome to enumerate the degraded
/// children in an operator-visible `warn!`.
/// The configured children that failed to start during bring-up. Kept for
/// the tests that pin the failure-vs-absence split; the driver reports
/// [`Self::degraded_children`], which is the union of the two.
#[cfg(test)]
pub(in crate::node) fn failed(&self) -> &HashSet<Child> {
&self.failed
}
/// Children whose interface is currently absent.
pub(in crate::node) fn absent(&self) -> &HashSet<Child> {
&self.absent
}
/// Every child currently contributing `Degraded` — the ones that failed to
/// start plus the ones whose interface is away. This is what an operator
/// wants named when the node reports `Degraded`.
pub(in crate::node) fn degraded_children(&self) -> HashSet<Child> {
self.failed.union(&self.absent).copied().collect()
}
/// Whether the machine is in the bounded-drain window. The driver uses this
/// after the rx loop returns to decide between the drain-teardown path and
/// the immediate-`stop()` fallback.
@@ -398,6 +471,8 @@ impl SupervisorFsm {
Event::DrainDeadlineElapsed => self.on_drain_deadline_elapsed(),
Event::ChildStopped { child } => self.on_child_stopped(child),
Event::ChildExited { child } => self.on_child_exited(child),
Event::ChildAbsent { child } => self.on_child_absent(child),
Event::ChildPresent { child } => self.on_child_present(child),
}
}
@@ -444,6 +519,7 @@ impl SupervisorFsm {
self.up.clear();
self.failed.clear();
self.absent.clear();
// A node with no children at all resolves health immediately. Zero
// transports up → `Failed` (this is the behavioral
@@ -497,19 +573,30 @@ impl SupervisorFsm {
self.classify_health()
}
/// Classify health from the current `up` / `failed` sets and set the
/// resulting state, returning the [`NodeState`] the driver should publish.
/// Shared by start-completion ([`Self::resolve_start_health`]) and runtime
/// child-exit ([`Self::on_child_exited`]):
/// Classify health from the current `up` / `failed` / `absent` sets and set
/// the resulting state, returning the [`NodeState`] the driver should
/// publish. Shared by start-completion ([`Self::resolve_start_health`]),
/// runtime child-exit ([`Self::on_child_exited`]) and the presence edges
/// ([`Self::on_child_absent`] / [`Self::on_child_present`]):
///
/// - zero transports up → [`SupState::Failed`] / [`NodeState::Failed`];
/// - ≥1 transport up but some child in `failed` → [`Health::Degraded`] /
/// [`NodeState::Degraded`];
/// - everything up and nothing failed → [`Health::Full`] / [`NodeState::Running`].
/// - ≥1 transport up but some child in `failed` or `absent` →
/// [`Health::Degraded`] / [`NodeState::Degraded`];
/// - everything up, nothing failed, nothing absent → [`Health::Full`] /
/// [`NodeState::Running`].
///
/// Worker-pool failures are captured in `failed` like any other optional
/// child, so they contribute `Degraded` (never `Failed`) — the inline crypto
/// fallback keeps the node correct without the pools.
///
/// A transport whose interface is absent is still counted among
/// `transports_up`: it came up, it is simply not bound. Excluding it would
/// make a single-ethernet node that booted before its wifi resolve to the
/// fatal [`FailReason::NoTransports`] and exit — which is the failure this
/// whole mechanism exists to remove. Absence degrades; it never kills.
///
/// Recomputed in **both** directions. This function is the reason
/// `Degraded` is a level rather than a latch.
fn classify_health(&mut self) -> NodeState {
let transports_up = self
.up
@@ -521,10 +608,10 @@ impl SupervisorFsm {
reason: FailReason::NoTransports,
};
NodeState::Failed
} else if !self.failed.is_empty() {
} else if !self.failed.is_empty() || !self.absent.is_empty() {
self.state = SupState::Running {
health: Health::Degraded {
reasons: self.failed.clone(),
reasons: self.degraded_children(),
},
};
NodeState::Degraded
@@ -536,6 +623,53 @@ impl SupervisorFsm {
}
}
/// An interface-bound child lost its interface.
///
/// The child stays in the up-set: the transport object survives detach —
/// config, id, statistics and neighbor buffer persist, only the descriptor
/// and its loops go. Only the reason set changes.
///
/// While `Starting` the edge is recorded silently; start-completion health
/// picks it up when `pending` empties, so a boot race resolves to
/// `Degraded` on the first publish rather than publishing `Full` and
/// immediately correcting it. Inert outside `Starting` / `Running`: a
/// teardown in flight owns its own bookkeeping.
fn on_child_absent(&mut self, child: Child) -> Vec<Action> {
match self.state {
SupState::Starting { .. } => {
self.absent.insert(child);
Vec::new()
}
SupState::Running { .. } => {
if !self.absent.insert(child) {
return Vec::new();
}
vec![Action::PublishState(self.classify_health())]
}
_ => Vec::new(),
}
}
/// An interface-bound child's interface came back.
///
/// Clears the absence and republishes. A child that is not currently
/// recorded absent produces nothing — a duplicate edge is not an event.
fn on_child_present(&mut self, child: Child) -> Vec<Action> {
match self.state {
SupState::Starting { .. } => {
self.absent.remove(&child);
Vec::new()
}
SupState::Running { .. } => {
if !self.absent.remove(&child) {
return Vec::new();
}
vec![Action::PublishState(self.classify_health())]
}
_ => Vec::new(),
}
}
fn on_stop(&mut self) -> Vec<Action> {
if !matches!(self.state, SupState::Running { .. }) {
return Vec::new();
@@ -588,6 +722,7 @@ impl SupervisorFsm {
if let SupState::Stopping { pending } = &mut self.state {
pending.remove(&child);
self.up.remove(&child);
self.absent.remove(&child);
if pending.is_empty() {
self.state = SupState::Stopped;
}
@@ -617,6 +752,8 @@ impl SupervisorFsm {
if !self.up.remove(&child) {
return Vec::new();
}
// An exit supersedes an absence: the child is gone, not waiting.
self.absent.remove(&child);
self.failed.insert(child);
vec![Action::PublishState(self.classify_health())]
}
@@ -1393,4 +1530,289 @@ mod tests {
);
assert!(s.failed().is_empty());
}
// ── Interface presence ────────────────────────────────────────────────
//
// The absence set is separate from `failed` because it is reversible, and
// reversibility is what turns `Degraded` from a latch into a level.
/// Helper: bring a full node up cleanly and leave it `Running{Full}`.
fn running_node() -> SupervisorFsm {
let mut s = SupervisorFsm::new();
s.step(start_full());
for child in [
Child::Transport(tid(1)),
Child::Transport(tid(2)),
Child::EncryptWorkers,
Child::DecryptWorkers,
Child::Nostr,
Child::Mdns,
Child::Tun,
Child::Dns,
] {
s.step(Event::SubstrateUp { child });
}
assert_eq!(
s.state(),
&SupState::Running {
health: Health::Full
}
);
s
}
#[test]
fn an_absent_interface_degrades_a_running_node() {
let mut s = running_node();
assert_eq!(
s.step(Event::ChildAbsent {
child: Child::Transport(tid(1))
}),
vec![Action::PublishState(NodeState::Degraded)]
);
assert!(s.absent().contains(&Child::Transport(tid(1))));
// Absence is not failure: the two sets stay distinct.
assert!(s.failed().is_empty());
}
#[test]
fn a_returning_interface_clears_degraded() {
// The regression this whole design turns on: plug the WAN back in and
// the node must leave `Degraded`, not carry it until the process
// restarts. `Degraded` is a level, not a latch.
let mut s = running_node();
s.step(Event::ChildAbsent {
child: Child::Transport(tid(1)),
});
assert_eq!(
s.step(Event::ChildPresent {
child: Child::Transport(tid(1))
}),
vec![Action::PublishState(NodeState::Running)]
);
assert_eq!(
s.state(),
&SupState::Running {
health: Health::Full
}
);
assert!(s.absent().is_empty());
}
#[test]
fn health_reflects_the_last_interface_still_away() {
// Two interfaces away, one returns: still Degraded. Only the empty
// absence set publishes Full.
let mut s = running_node();
s.step(Event::ChildAbsent {
child: Child::Transport(tid(1)),
});
s.step(Event::ChildAbsent {
child: Child::Transport(tid(2)),
});
assert_eq!(
s.step(Event::ChildPresent {
child: Child::Transport(tid(1))
}),
vec![Action::PublishState(NodeState::Degraded)]
);
assert_eq!(
s.step(Event::ChildPresent {
child: Child::Transport(tid(2))
}),
vec![Action::PublishState(NodeState::Running)]
);
}
#[test]
fn a_duplicate_presence_edge_is_not_an_event() {
let mut s = running_node();
s.step(Event::ChildAbsent {
child: Child::Transport(tid(1)),
});
assert_eq!(
s.step(Event::ChildAbsent {
child: Child::Transport(tid(1))
}),
vec![],
"re-reporting the same absence must not republish"
);
s.step(Event::ChildPresent {
child: Child::Transport(tid(1)),
});
assert_eq!(
s.step(Event::ChildPresent {
child: Child::Transport(tid(1))
}),
vec![],
"re-reporting the same return must not republish"
);
}
#[test]
fn absence_during_bringup_resolves_degraded_on_the_first_publish() {
// The boot race. The transport reports absence from inside its own
// start, then reports up. Start-completion health must already know,
// so the node never publishes `Full` and corrects itself.
let mut s = SupervisorFsm::new();
s.step(start_full());
s.step(Event::ChildAbsent {
child: Child::Transport(tid(1)),
});
for child in [
Child::Transport(tid(1)),
Child::Transport(tid(2)),
Child::EncryptWorkers,
Child::DecryptWorkers,
Child::Nostr,
Child::Mdns,
Child::Tun,
] {
assert_eq!(s.step(Event::SubstrateUp { child }), vec![]);
}
assert_eq!(
s.step(Event::SubstrateUp { child: Child::Dns }),
vec![Action::PublishState(NodeState::Degraded)]
);
let mut reasons = HashSet::new();
reasons.insert(Child::Transport(tid(1)));
assert_eq!(
s.state(),
&SupState::Running {
health: Health::Degraded { reasons }
}
);
}
#[test]
fn a_lone_absent_transport_degrades_rather_than_fails() {
// A single-ethernet node that boots before its wifi. The transport
// came up — absence is a state, not a start failure — so it counts
// among the transports up and the node serves. Resolving to `Failed`
// here would make the daemon exit on the very race this mechanism
// exists to survive.
let mut s = SupervisorFsm::new();
s.step(Event::Start {
transports: vec![tid(1)],
encrypt_workers: false,
decrypt_workers: false,
nostr: false,
mdns: false,
tun: false,
dns: false,
});
s.step(Event::ChildAbsent {
child: Child::Transport(tid(1)),
});
assert_eq!(
s.step(Event::SubstrateUp {
child: Child::Transport(tid(1))
}),
vec![Action::PublishState(NodeState::Degraded)]
);
assert!(
!matches!(s.state(), SupState::Failed { .. }),
"an absent interface must never be the fatal no-transports case"
);
}
#[test]
fn an_exit_supersedes_an_absence() {
// A transport that is away and then exits is gone, not waiting: it
// leaves the up-set and moves from `absent` to `failed`, so a later
// spurious `ChildPresent` cannot resurrect it into `Full`.
let mut s = running_node();
s.step(Event::ChildAbsent {
child: Child::Transport(tid(1)),
});
s.step(Event::ChildExited {
child: Child::Transport(tid(1)),
});
assert!(s.absent().is_empty());
assert!(s.failed().contains(&Child::Transport(tid(1))));
assert_eq!(
s.step(Event::ChildPresent {
child: Child::Transport(tid(1))
}),
vec![]
);
assert!(matches!(
s.state(),
SupState::Running {
health: Health::Degraded { .. }
}
));
}
#[test]
fn presence_edges_are_inert_outside_starting_and_running() {
// Teardown and drain own their own bookkeeping; a late edge from a
// binder that has not noticed the stop must not author actions.
let mut created = SupervisorFsm::new();
assert_eq!(
created.step(Event::ChildAbsent {
child: Child::Transport(tid(1))
}),
vec![]
);
let mut draining = running_node();
draining.step(Event::Drain { deadline_ms: 1 });
assert_eq!(
draining.step(Event::ChildAbsent {
child: Child::Transport(tid(1))
}),
vec![]
);
assert!(
draining.is_draining(),
"a presence edge must not end a drain"
);
}
#[test]
fn a_new_start_clears_the_previous_absences() {
let mut s = running_node();
s.step(Event::ChildAbsent {
child: Child::Transport(tid(1)),
});
s.step(Event::Stop);
for child in [
Child::Dns,
Child::Nostr,
Child::Mdns,
Child::Transport(tid(1)),
Child::Transport(tid(2)),
Child::Tun,
] {
s.step(Event::ChildStopped { child });
}
s.step(start_full());
assert!(s.absent().is_empty());
}
#[test]
fn degraded_children_is_the_union_of_both_reason_sets() {
let mut s = SupervisorFsm::new();
s.step(start_full());
s.step(Event::ChildAbsent {
child: Child::Transport(tid(1)),
});
s.step(Event::SubstrateFailed { child: Child::Mdns });
for child in [
Child::Transport(tid(1)),
Child::Transport(tid(2)),
Child::EncryptWorkers,
Child::DecryptWorkers,
Child::Nostr,
Child::Tun,
Child::Dns,
] {
s.step(Event::SubstrateUp { child });
}
let degraded = s.degraded_children();
assert!(degraded.contains(&Child::Mdns));
assert!(degraded.contains(&Child::Transport(tid(1))));
assert_eq!(degraded.len(), 2);
}
}
+156 -4
View File
@@ -154,6 +154,17 @@ pub enum NodeError {
#[error("send failed to {node_addr}: {reason}")]
SendFailed { node_addr: NodeAddr, reason: String },
/// A send refused by a condition that is expected to clear on its own.
///
/// Distinct from [`Self::SendFailed`] because the right response differs:
/// the state built around the send — a half-finished handshake, a route,
/// a queued packet — is worth keeping across a transient refusal and
/// worth tearing down after a terminal one. Carries the transport's own
/// classification ([`TransportError::is_transient`]) rather than a
/// re-derivation of it.
#[error("send to {node_addr} unavailable: {reason}")]
SendUnavailable { node_addr: NodeAddr, reason: String },
#[error("mtu exceeded forwarding to {node_addr}: packet {packet_size} > mtu {mtu}")]
MtuExceeded {
node_addr: NodeAddr,
@@ -186,6 +197,32 @@ pub enum NodeError {
NoOperationalTransports,
}
impl Node {
/// Test-only: place a transport into the node's map directly.
///
/// The snapshot tests live in `crate::control` and so cannot reach the
/// private `transports` field, but the interface-presence block they need
/// to pin only exists on a real interface-bound transport. Mirrors
/// `isolate_peer_acl_for_test`: a narrow hook, so the fixture stays honest
/// rather than the snapshot being hand-authored JSON that nothing
/// produces.
#[cfg(all(test, any(target_os = "linux", target_os = "macos")))]
pub(crate) fn insert_transport_for_test(&mut self, id: TransportId, handle: TransportHandle) {
self.transports.insert(id, handle);
}
}
impl NodeError {
/// Whether this failure is expected to clear on its own.
///
/// Mirrors [`TransportError::is_transient`] across the node boundary, so
/// a caller holding a `NodeError` can ask the same question a caller
/// holding a `TransportError` can, and get the same answer.
pub fn is_transient(&self) -> bool {
matches!(self, Self::SendUnavailable { .. })
}
}
/// Node operational state.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub enum NodeState {
@@ -394,6 +431,16 @@ pub struct Node {
/// SYN/SYN-ACK clamp can use the smaller of the local-egress floor
/// and the learned per-destination path MTU.
path_mtu_lookup: crate::upper::tun::PathMtuLookup,
/// Node-global TCP MSS ceiling, shared live with the TUN reader and writer
/// threads and recomputed whenever the set of *bound* transports changes.
///
/// Sits beside `path_mtu_lookup` because it answers the other half of the
/// same question at the same moment: that map supplies the per-destination
/// ceiling, this supplies the local-egress one, and the clamp takes the
/// smaller. Both have to be read live — a transport that binds after start
/// can be the narrow one, and one that unbinds can be the reason the node
/// was clamped at all.
tun_mss_ceiling: crate::upper::tun::MssCeiling,
/// Which transport last supplied a *link seed* into `path_mtu_lookup`,
/// per destination.
///
@@ -438,6 +485,20 @@ pub struct Node {
/// rx_loop select arm that feeds `Event::ChildExited` to the supervisor FSM.
child_exit_rx: Option<tokio::sync::mpsc::Receiver<crate::node::lifecycle::supervisor::Child>>,
// === Interface Presence Channel ===
/// Sender half of the interface-presence channel, cloned into every
/// interface-bound transport so its binder task can report attach and
/// detach. Held on `self` for the rx_loop's lifetime as the keep-alive
/// sender, exactly like [`Self::child_exit_tx`].
///
/// Separate from the child-exit channel because presence is *reversible*:
/// an exit is one-way, an interface comes back.
transport_presence_tx: Option<crate::transport::PresenceTx>,
/// Receiver half of the interface-presence channel, `take()`-en by the
/// rx_loop select arm that feeds `Event::ChildAbsent` / `Event::ChildPresent`
/// to the supervisor FSM.
transport_presence_rx: Option<crate::transport::PresenceRx>,
// === Per-Peer Control Machines ===
/// Per-peer lifecycle control FSMs, keyed by the stable `LinkId` that spans
/// the handshake→active lifetime. Each machine owns its handshake crypto
@@ -840,6 +901,8 @@ impl Node {
packet_rx: None,
child_exit_tx: None,
child_exit_rx: None,
transport_presence_tx: None,
transport_presence_rx: None,
peer_machines: HashMap::new(),
peer_timers: HashMap::new(),
peers: HashMap::new(),
@@ -908,6 +971,12 @@ impl Node {
peer_acl,
host_map,
path_mtu_lookup: Arc::new(std::sync::RwLock::new(HashMap::new())),
// Seeded at the IPv6 minimum, which is what `transport_mtu()`
// itself falls back to when nothing is bound. Refreshed before
// the TUN threads start and on every change to the bound set.
tun_mss_ceiling: Arc::new(std::sync::atomic::AtomicU16::new(
crate::upper::icmp::mss_ceiling(crate::upper::tun::IPV6_MIN_MTU),
)),
path_mtu_seeded_by: Arc::new(std::sync::RwLock::new(HashMap::new())),
#[cfg(unix)]
decrypt_registered_sessions: std::collections::HashSet::new(),
@@ -1012,6 +1081,8 @@ impl Node {
packet_rx: None,
child_exit_tx: None,
child_exit_rx: None,
transport_presence_tx: None,
transport_presence_rx: None,
peer_machines: HashMap::new(),
peer_timers: HashMap::new(),
peers: HashMap::new(),
@@ -1077,6 +1148,12 @@ impl Node {
peer_acl,
host_map,
path_mtu_lookup: Arc::new(std::sync::RwLock::new(HashMap::new())),
// Seeded at the IPv6 minimum, which is what `transport_mtu()`
// itself falls back to when nothing is bound. Refreshed before
// the TUN threads start and on every change to the bound set.
tun_mss_ceiling: Arc::new(std::sync::atomic::AtomicU16::new(
crate::upper::icmp::mss_ceiling(crate::upper::tun::IPV6_MIN_MTU),
)),
path_mtu_seeded_by: Arc::new(std::sync::RwLock::new(HashMap::new())),
#[cfg(unix)]
decrypt_registered_sessions: std::collections::HashSet::new(),
@@ -1133,7 +1210,13 @@ impl Node {
.collect();
for (name, eth_config) in eth_instances {
let transport_id = self.allocate_transport_id();
let eth = EthernetTransport::new(transport_id, name, eth_config, packet_tx.clone());
let mut eth =
EthernetTransport::new(transport_id, name, eth_config, packet_tx.clone());
// The binder task reports attach and detach here, so node
// health tracks the interface in both directions.
if let Some(tx) = self.transport_presence_tx.clone() {
eth.set_presence_tx(tx);
}
transports.push(TransportHandle::Ethernet(eth));
}
}
@@ -1445,6 +1528,46 @@ impl Node {
crate::upper::icmp::effective_ipv6_mtu(self.transport_mtu())
}
/// The TCP MSS ceiling the TUN threads are currently clamping to.
#[cfg(test)]
pub(crate) fn tun_mss_ceiling(&self) -> u16 {
self.tun_mss_ceiling
.load(std::sync::atomic::Ordering::Relaxed)
}
/// Recompute the shared TUN MSS ceiling from the currently bound
/// transports, and log it if it moved.
///
/// Called wherever the bound set can change — the presence edges that
/// bind and unbind an interface-bound transport, and a child exiting —
/// so the clamp the TUN threads apply keeps agreeing with the
/// `effective_ipv6_mtu` this node reports in `show_status`.
///
/// Moves in **both** directions, deliberately. A narrow interface
/// appearing has to tighten the ceiling or the clamp is wrong for
/// traffic that will egress over it; that same interface going away has
/// to release it, or unplugging a low-MTU adapter leaves the node
/// over-clamped until it restarts. It is the same argument that makes
/// `Degraded` a level rather than a latch: nothing here is one-way once
/// a transport can come back.
///
/// Existing flows are not re-clamped — MSS is negotiated per connection
/// at SYN time, so a change applies to connections opened after it.
pub(crate) fn refresh_tun_mss_ceiling(&self) {
use std::sync::atomic::Ordering;
let ceiling = crate::upper::icmp::mss_ceiling(self.transport_mtu());
let previous = self.tun_mss_ceiling.swap(ceiling, Ordering::Relaxed);
if previous != ceiling {
tracing::info!(
previous_max_mss = previous,
max_mss = ceiling,
effective_ipv6_mtu = self.effective_ipv6_mtu(),
"Node egress MTU changed; TCP MSS ceiling updated for new connections"
);
}
}
/// Get the transport MTU governing the global TUN-boundary MSS clamp.
///
/// Returns the **minimum** MTU across all operational transports, or
@@ -1460,10 +1583,17 @@ impl Node {
/// to vary across HashMap iteration order + async-startup race) makes
/// the clamp deterministic across daemon restarts.
pub fn transport_mtu(&self) -> u16 {
// `is_bound`, not `is_operational`. An interface-bound transport is
// "operational" from the moment it starts, whether or not its
// interface exists — so filtering on that let a transport whose
// interface has never appeared clamp the whole node's IPv6 MTU to a
// number derived from hardware that is not present. Before dynamic
// binding a transport that could not bind was never inserted here at
// all, so the distinction did not exist to get wrong.
let min_operational = self
.transports
.values()
.filter(|h| h.is_operational())
.filter(|h| h.is_bound())
.map(|h| h.mtu())
.min();
if let Some(mtu) = min_operational {
@@ -2374,8 +2504,13 @@ impl Node {
.collect();
// --- transports (show_transports) ---
let transport_rows: Vec<snap::TransportRow> = self
.transport_ids()
// Ascending id, matching `show_transports`; see the note there. The
// off-loop renderer reads this table verbatim, so the two paths would
// otherwise disagree about ordering as well as being arbitrary.
let mut transport_ids: Vec<_> = self.transport_ids().copied().collect();
transport_ids.sort_by_key(|id| id.as_u32());
let transport_rows: Vec<snap::TransportRow> = transport_ids
.iter()
.map(|id| {
let handle = self.get_transport(id).unwrap();
snap::TransportRow {
@@ -2391,6 +2526,15 @@ impl Node {
.tor_monitoring()
.map(|m| serde_json::to_value(&m).unwrap_or_default()),
stats: handle.transport_stats(),
interface: handle.interface_presence().map(|p| snap::InterfaceRow {
name: handle.interface_name().unwrap_or_default().to_string(),
presence: p.presence,
carrier: p.carrier,
policy: p.policy,
since_secs: p.since_secs,
binds: p.binds,
failed_attempts: p.failed_attempts,
}),
}
})
.collect();
@@ -3817,6 +3961,14 @@ impl Node {
packet_size,
mtu,
},
// Preserve the transport's own classification instead of
// flattening every non-MTU failure into one string. A caller
// that wants to keep its half-built state across an interface
// flap can only do that if the distinction survives to it.
other if other.is_transient() => NodeError::SendUnavailable {
node_addr: *node_addr,
reason: format!("transport send: {}", other),
},
other => NodeError::SendFailed {
node_addr: *node_addr,
reason: format!("transport send: {}", other),
+203
View File
@@ -0,0 +1,203 @@
//! Per-peer connected UDP sockets against the in-line decrypt path.
//!
//! A `connect(2)`-ed UDP socket is pinned to one 5-tuple. When the peer
//! moves, the address the socket was opened against is gone, but the
//! socket is still installed and the send path prefers it over the
//! wildcard listen socket. `ActivePeer::set_current_addr` returns
//! whether the address actually changed precisely so the caller can
//! drop the stale socket, and both post-decrypt paths have to act on
//! that return: the decrypt-worker completion path
//! (`process_authentic_fmp_plaintext`) and the in-line one
//! (`handle_encrypted_frame`). These tests cover the in-line path,
//! which is the one the worker path's own coverage does not reach.
//!
//! `bool` carries no `#[must_use]`, so discarding the return here is
//! silent under `-D warnings`; the assertions below are what makes the
//! difference between binding it and dropping it observable.
use super::*;
use crate::noise::NoiseSession;
use crate::proto::fmp::wire::{build_encrypted, build_established_header, prepend_inner_header};
/// The address `seed_completed_connection` promotes a peer on, and so
/// the peer's `current_addr` before anything rotates it.
const PROMOTED_ADDR: &str = "127.0.0.1:5000";
/// The address the peer is made to move to.
const ROAMED_ADDR: &str = "127.0.0.1:5001";
/// Build a promoted peer and hand back the far side's Noise session.
///
/// [`seed_completed_connection`] runs every leg of the handshake and
/// then drops the responder, so nothing outside it can produce a frame
/// the node will actually authenticate. This is the same seeding with
/// the responder's session kept, which is what lets these tests reach
/// the post-decrypt side effects rather than stopping at the AEAD.
///
/// Returns the node, the peer's `NodeAddr`, the session index an
/// inbound frame must name to be routed to that peer, and the session
/// to encrypt those frames with.
fn promoted_peer_with_the_far_side_session(
transport_id: TransportId,
) -> (Node, NodeAddr, SessionIndex, NoiseSession) {
let mut node = make_node();
let link_id = LinkId::new(1);
let peer_identity_full = Identity::generate();
// from_pubkey_full, not from_pubkey: the ECDH needs the parity bit.
let peer_identity = PeerIdentity::from_pubkey_full(peer_identity_full.pubkey_full());
let our_index = node.index_allocator.allocate().unwrap();
node.seed_handshake_machine(
HandshakeSeed::outbound(link_id, peer_identity, 1_000)
.with_our_index(our_index)
.with_their_index(SessionIndex::new(42))
.with_transport_id(transport_id)
.with_source_addr(TransportAddr::from_string(PROMOTED_ADDR)),
)
.unwrap();
let our_keypair = node.identity().keypair();
let startup_epoch = node.startup_epoch();
let msg1 = node
.peer_machines
.get_mut(&link_id)
.unwrap()
.start_handshake(our_keypair, startup_epoch, 1_000)
.unwrap();
let mut responder = inbound_leg(LinkId::new(999), 1_000);
let mut responder_epoch = [0u8; 8];
rand::Rng::fill_bytes(&mut rand::rng(), &mut responder_epoch);
let msg2 = responder
.receive_handshake_init(
peer_identity_full.keypair(),
responder_epoch,
&msg1,
None,
1_000,
)
.unwrap();
// XX, so msg2 does not finish the responder: it holds no session until it
// has processed msg3. The line this came from ran IK, where the responder
// was done at msg2 and the third leg did not exist.
let (msg3, _negotiation) = node
.peer_machines
.get_mut(&link_id)
.unwrap()
.complete_handshake(&msg2, None, 1_000)
.unwrap();
responder
.complete_handshake_msg3(&msg3, 1_000)
.expect("the responder must accept a genuine msg3");
let far_side_session = responder
.take_session()
.expect("the responder holds a session once it has processed msg3");
node.promote_connection(link_id, peer_identity, 2_000)
.unwrap();
let node_addr = *peer_identity.node_addr();
let our_index = node
.get_peer(&node_addr)
.and_then(|p| p.our_index())
.expect("a promoted peer carries the index it was allocated");
(node, node_addr, our_index, far_side_session)
}
/// Encrypt one well-formed established frame from the far side.
///
/// The link message is a heartbeat (`0x51`), which the dispatcher
/// handles as a no-op — these tests are about the side effects that run
/// before the dispatch, so the message must not have any of its own.
fn far_side_frame(session: &mut NoiseSession, receiver_idx: SessionIndex) -> Vec<u8> {
let inner = prepend_inner_header(0, &[0x51]);
let counter = session.current_send_counter();
let header = build_established_header(receiver_idx, counter, 0, inner.len() as u16);
let ciphertext = session.encrypt_with_aad(&inner, &header).unwrap();
build_encrypted(&header, &ciphertext)
}
/// **The defect.**
///
/// The in-line decrypt path called `set_current_addr` as a bare
/// statement and dropped its return, so a peer could roam, have its
/// `current_addr` updated, and keep a connected socket pinned to the
/// 5-tuple it had just left. The send path prefers that socket while it
/// is installed, so every frame after the move goes out to an address
/// the peer is no longer at.
#[cfg(any(target_os = "linux", target_os = "macos"))]
#[tokio::test]
async fn a_peer_that_roams_loses_the_connected_socket_pinned_to_the_address_it_left() {
let transport_id = TransportId::new(1);
let (mut node, node_addr, our_index, mut far_side) =
promoted_peer_with_the_far_side_session(transport_id);
install_connected_udp(&mut node, &node_addr, transport_id);
assert!(
node.get_peer(&node_addr).unwrap().connected_udp().is_some(),
"precondition: the peer holds a connected socket before it moves"
);
let frame = far_side_frame(&mut far_side, our_index);
node.handle_encrypted_frame(ReceivedPacket::new(
transport_id,
TransportAddr::from_string(ROAMED_ADDR),
frame,
))
.await;
let peer = node
.get_peer(&node_addr)
.expect("the peer survives an authentic frame");
assert_eq!(
peer.current_addr(),
Some(&TransportAddr::from_string(ROAMED_ADDR)),
"precondition for the assertion below: the frame must have been \
authenticated and the rotation recorded, or the test proves nothing"
);
assert!(
peer.connected_udp().is_none(),
"a socket pinned to the address the peer has left must not survive \
the rotation"
);
}
/// **The healthy path.**
///
/// A frame from the address the peer is already on changes nothing, so
/// the connected socket has to stay. A fix that cleared unconditionally
/// would tear down and reopen the socket on every single frame.
#[cfg(any(target_os = "linux", target_os = "macos"))]
#[tokio::test]
async fn a_frame_from_the_address_the_peer_is_already_on_keeps_the_connected_socket() {
let transport_id = TransportId::new(1);
let (mut node, node_addr, our_index, mut far_side) =
promoted_peer_with_the_far_side_session(transport_id);
install_connected_udp(&mut node, &node_addr, transport_id);
let frame = far_side_frame(&mut far_side, our_index);
node.handle_encrypted_frame(ReceivedPacket::new(
transport_id,
TransportAddr::from_string(PROMOTED_ADDR),
frame,
))
.await;
let peer = node
.get_peer(&node_addr)
.expect("the peer survives an authentic frame");
assert_eq!(
peer.current_addr(),
Some(&TransportAddr::from_string(PROMOTED_ADDR)),
"the peer has not moved"
);
assert!(
peer.connected_udp().is_some(),
"a frame from the address already in use must leave the socket alone"
);
}
+71
View File
@@ -375,3 +375,74 @@ async fn a_failing_peer_is_retried_after_the_gap_and_not_before() {
cleanup_nodes(&mut nodes).await;
}
// ---------------------------------------------------------------------------
// Reaping peers when their interface goes away
//
// The detach edge is both earlier and more certain than inactivity, so it is
// the better trigger for withdrawing what the interface carried. These drive
// the same real two-node peering the liveness tests use, because a peer only
// reaches the established context the reap acts on by actually peering.
// ---------------------------------------------------------------------------
/// A peer reachable only through an interface that has gone is withdrawn on
/// the detach edge, without waiting out `link_dead_timeout_secs`.
///
/// Note what is *not* set here: the link-dead timeout keeps its default, and
/// no time is advanced. The peer is live by every liveness measure and is
/// still withdrawn, because the transport under it is gone — which is the
/// whole distinction this adds.
#[tokio::test]
async fn a_detached_transport_withdraws_the_peers_that_needed_it() {
let mut nodes = run_tree_test(2, &[(0, 1)], false).await;
verify_tree_convergence(&nodes);
let addr_1 = *nodes[1].node.node_addr();
let transport_id = nodes[0]
.node
.get_peer(&addr_1)
.expect("peer present")
.transport_id()
.expect("an established peer names its transport");
let reaped = nodes[0].node.reap_peers_on_transport(transport_id).await;
assert_eq!(reaped, 1);
assert!(
nodes[0].node.get_peer(&addr_1).is_none(),
"a peer must not outlive the interface it was reachable through"
);
cleanup_nodes(&mut nodes).await;
}
/// The reap is scoped to the transport that detached.
///
/// The failure this guards is the one that would make the feature worse than
/// the defect: an interface going away must not withdraw the peers that were
/// never reachable through it, which on a mesh router is most of them.
#[tokio::test]
async fn a_detached_transport_leaves_other_transports_peers_alone() {
let mut nodes = run_tree_test(2, &[(0, 1)], false).await;
verify_tree_convergence(&nodes);
let addr_1 = *nodes[1].node.node_addr();
let peer_transport = nodes[0]
.node
.get_peer(&addr_1)
.expect("peer present")
.transport_id()
.expect("an established peer names its transport");
// A transport this peer was never reachable through.
let unrelated = TransportId::new(peer_transport.as_u32() + 100);
let reaped = nodes[0].node.reap_peers_on_transport(unrelated).await;
assert_eq!(reaped, 0, "an unrelated transport withdraws nothing");
assert!(
nodes[0].node.get_peer(&addr_1).is_some(),
"a peer on a healthy transport must survive another one detaching"
);
cleanup_nodes(&mut nodes).await;
}
+34
View File
@@ -11,6 +11,7 @@ mod ble;
mod bloom;
mod bloom_poison;
mod bootstrap;
mod connected_udp;
mod control;
mod decrypt_failure;
mod disconnect;
@@ -57,6 +58,39 @@ pub(super) fn make_node_with(config: Config) -> Node {
Node::new(config).unwrap()
}
/// Install a real `connect()`-ed UDP socket on a peer, the way the tick-driven
/// activation in `dataplane::connected_udp` does.
///
/// The socket is opened against the loopback discard port: nothing is ever sent
/// through it, and the callers only care whether the handle is still installed
/// afterwards.
#[cfg(any(target_os = "linux", target_os = "macos"))]
pub(super) fn install_connected_udp(
node: &mut Node,
addr: &NodeAddr,
transport_id: crate::transport::TransportId,
) {
let local: std::net::SocketAddr = "0.0.0.0:0".parse().unwrap();
let peer_sa: std::net::SocketAddr = "127.0.0.1:9".parse().unwrap();
let owned = crate::transport::udp::open_connected_fd(local, peer_sa, 65_536, 65_536)
.expect("open a connected UDP socket");
let bound = crate::transport::udp::ConnectedPeerSocket::from_fd(owned, peer_sa, local);
let socket = std::sync::Arc::new(bound);
let (packet_tx, _packet_rx) = packet_channel(8);
let drain = crate::transport::udp::PeerRecvDrain::spawn(
socket.clone(),
transport_id,
peer_sa,
packet_tx,
)
.expect("spawn the peer recv drain");
node.get_peer_mut(addr)
.expect("peer present")
.set_connected_udp(socket, drain);
}
/// Build a test node with an explicit `max_peers` limit (replaces the removed
/// `set_max_peers` setter; resource limits are immutable post-construction).
pub(super) fn make_node_with_max_peers(max_peers: usize) -> Node {
-29
View File
@@ -31,35 +31,6 @@ fn identity_of(nodes: &[TestNode], j: usize) -> PeerIdentity {
PeerIdentity::from_pubkey_full(nodes[j].node.identity().pubkey_full())
}
/// Install a real `connect()`-ed UDP socket on a peer, the way the tick-driven
/// activation does.
///
/// The socket is opened against a discard port on loopback: nothing is ever
/// sent through it, and the test only cares whether the handle survives a
/// medium change.
#[cfg(any(target_os = "linux", target_os = "macos"))]
fn install_connected_udp(node: &mut Node, addr: &NodeAddr, transport_id: TransportId) {
let local: std::net::SocketAddr = "0.0.0.0:0".parse().unwrap();
let peer_sa: std::net::SocketAddr = "127.0.0.1:9".parse().unwrap();
let owned = crate::transport::udp::open_connected_fd(local, peer_sa, 65_536, 65_536)
.expect("open a connected UDP socket");
let bound = crate::transport::udp::ConnectedPeerSocket::from_fd(owned, peer_sa, local);
let socket = std::sync::Arc::new(bound);
let (packet_tx, _packet_rx) = crate::transport::packet_channel(8);
let drain = crate::transport::udp::PeerRecvDrain::spawn(
socket.clone(),
transport_id,
peer_sa,
packet_tx,
)
.expect("spawn the peer recv drain");
node.get_peer_mut(addr)
.expect("peer present")
.set_connected_udp(socket, drain);
}
/// **The defect this feature exists for.**
///
/// Established UDP peers get a per-peer `connect()`-ed socket. `open_connected_fd`
+172 -16
View File
@@ -1258,7 +1258,7 @@ async fn test_try_peer_addresses_skips_connecting_peer() {
}
#[test]
fn active_peer_same_path_discovery_skips_fresh_peer() {
fn a_peer_heard_from_within_the_heartbeat_interval_has_a_live_link() {
let mut node = make_node();
let peer_full = Identity::generate();
let peer_identity = PeerIdentity::from_pubkey_full(peer_full.pubkey_full());
@@ -1268,16 +1268,14 @@ fn active_peer_same_path_discovery_skips_fresh_peer() {
let mut active_peer = ActivePeer::new(peer_identity, LinkId::new(7), Node::now_ms());
active_peer.set_current_addr(transport_id, current_addr.clone());
node.peers.insert(peer_node_addr, active_peer);
let candidate = crate::config::PeerAddress::new("udp", "127.0.0.1:9");
assert!(node.active_peer_candidate_is_fresh_enough_to_skip(
&peer_node_addr,
std::slice::from_ref(&candidate),
));
// A link heard from just now is live, so discovery must not dial this
// peer at all — on this path or on any other.
assert!(node.active_peer_link_is_live(&peer_node_addr));
}
#[test]
fn active_peer_same_path_discovery_refreshes_stale_peer() {
fn a_peer_quiet_past_the_heartbeat_interval_no_longer_has_a_live_link() {
let mut node = make_node();
let peer_full = Identity::generate();
let peer_identity = PeerIdentity::from_pubkey_full(peer_full.pubkey_full());
@@ -1294,12 +1292,62 @@ fn active_peer_same_path_discovery_refreshes_stale_peer() {
let mut active_peer = ActivePeer::new(peer_identity, LinkId::new(7), stale_at);
active_peer.set_current_addr(transport_id, current_addr.clone());
node.peers.insert(peer_node_addr, active_peer);
let candidate = crate::config::PeerAddress::new("udp", "127.0.0.1:9");
assert!(!node.active_peer_candidate_is_fresh_enough_to_skip(
&peer_node_addr,
std::slice::from_ref(&candidate),
));
// Gone quiet past the heartbeat interval: every path is dialable again,
// which is what keeps failover working now that a live link is never
// displaced.
assert!(!node.active_peer_link_is_live(&peer_node_addr));
}
/// A bootstrap-held peer is never its own configured candidate.
///
/// `adopt_established_traversal` refuses a peer that is already connected, so
/// an adopted NAT-traversal transport is the way *on* to a traversed path and
/// not the way off. Now that beacon discovery asks only whether the link it
/// holds is answering, the configured-peer refresh is the one automatic
/// off-ramp left, and it works only because a bootstrap-held peer is refused
/// as a match for its own address: a configured `udp` address can be
/// byte-identical to the traversal's remote address and the transport kinds
/// match, so without this carve-out `has_alternative` would be false and the
/// peer could never be moved off the traversed socket at all.
#[test]
fn a_bootstrap_held_peer_is_never_its_own_configured_candidate() {
let mut node = make_node();
let peer_full = Identity::generate();
let peer_identity = PeerIdentity::from_pubkey_full(peer_full.pubkey_full());
let peer_node_addr = *peer_identity.node_addr();
let npub = peer_identity.npub();
let transport_id = TransportId::new(1);
let current_addr = TransportAddr::from_string("203.0.113.5:41234");
let (packet_tx, _packet_rx) = packet_channel(8);
let udp = UdpTransport::new(
transport_id,
Some("main".to_string()),
crate::config::UdpConfig {
bind_addr: Some("127.0.0.1:0".to_string()),
..Default::default()
},
packet_tx,
);
node.transports
.insert(transport_id, TransportHandle::Udp(udp));
let mut active_peer = ActivePeer::new(peer_identity, LinkId::new(7), Node::now_ms());
active_peer.set_current_addr(transport_id, current_addr);
node.peers.insert(peer_node_addr, active_peer);
node.supervisor
.nostr_rendezvous
.insert_bootstrap_transport(transport_id, npub);
assert!(
!node.active_peer_matches_candidate(
&peer_node_addr,
&crate::config::PeerAddress::new("udp", "203.0.113.5:41234")
),
"a bootstrap-held peer must not count its own traversal address as its \
current path, or the configured-peer refresh can never migrate it off"
);
}
/// An instance-qualified candidate is the peer's *current* path only when it
@@ -1341,12 +1389,11 @@ async fn an_instance_qualified_candidate_matches_only_its_own_instance() {
active_peer.set_current_addr(main_id, TransportAddr::from_string("127.0.0.1:9"));
node.peers.insert(peer_node_addr, active_peer);
// Path matching, which still gates the *configured-peer* refresh even
// though beacon discovery now gates on liveness alone.
let matches = |transport: &str| {
let candidate = crate::config::PeerAddress::new(transport, "127.0.0.1:9");
node.active_peer_candidate_is_fresh_enough_to_skip(
&peer_node_addr,
std::slice::from_ref(&candidate),
)
node.active_peer_matches_candidate(&peer_node_addr, &candidate)
};
assert!(
@@ -1845,6 +1892,115 @@ async fn test_transport_mtu_returns_min_across_operational() {
}
}
#[tokio::test]
async fn the_tun_mss_ceiling_follows_a_transport_arriving_and_leaving() {
// The TUN reader and writer used to be handed a `u16` computed once at
// spawn. Every other consumer of `transport_mtu()` reads it live, so once
// a transport could bind minutes after start the daemon reported one
// effective MTU in `show_status` and clamped to another.
//
// Both directions. A narrow transport arriving has to tighten the ceiling
// or traffic egressing over it is clamped too loose; the same transport
// leaving has to release it, or unplugging a low-MTU adapter leaves the
// node over-clamped until it restarts.
let mut node = make_node();
let (packet_tx, packet_rx) = packet_channel(64);
node.supervisor.packet_tx = Some(packet_tx);
node.packet_rx = Some(packet_rx);
let wide = make_udp_transport_with_mtu(1, 1452).await;
node.transports.insert(TransportId::new(1), wide);
node.refresh_tun_mss_ceiling();
let wide_ceiling = node.tun_mss_ceiling();
assert_eq!(
wide_ceiling,
crate::upper::icmp::mss_ceiling(1452),
"the seeded ceiling must match the only bound transport"
);
// A narrower transport arrives after the TUN threads would already be
// running. The shared ceiling has to tighten.
let narrow = make_udp_transport_with_mtu(2, 1280).await;
node.transports.insert(TransportId::new(2), narrow);
node.refresh_tun_mss_ceiling();
let narrow_ceiling = node.tun_mss_ceiling();
assert_eq!(narrow_ceiling, crate::upper::icmp::mss_ceiling(1280));
assert!(
narrow_ceiling < wide_ceiling,
"a narrower transport must tighten the clamp, not be ignored"
);
assert_eq!(
narrow_ceiling,
crate::upper::icmp::mss_ceiling(node.transport_mtu()),
"the clamp and the reported MTU must not disagree"
);
// ...and leaving has to release it again.
if let Some(mut gone) = node.transports.remove(&TransportId::new(2)) {
gone.stop().await.ok();
}
node.refresh_tun_mss_ceiling();
assert_eq!(
node.tun_mss_ceiling(),
wide_ceiling,
"the ceiling must rise again when the narrow transport goes away"
);
for transport in node.transports.values_mut() {
transport.stop().await.ok();
}
}
#[tokio::test]
async fn a_presence_edge_refreshes_the_tun_mss_ceiling_without_being_asked() {
// The one above pins the arithmetic; this pins the wiring. A presence
// edge arriving has to refresh the ceiling on its own — if the refresh is
// dropped from the edge handlers the value silently stops tracking, which
// is the defect in its original form.
let mut node = make_node();
let (packet_tx, packet_rx) = packet_channel(64);
node.supervisor.packet_tx = Some(packet_tx);
node.packet_rx = Some(packet_rx);
let (presence_tx, presence_rx) = tokio::sync::mpsc::channel(4);
node.transport_presence_tx = Some(presence_tx.clone());
node.transport_presence_rx = Some(presence_rx);
// Nothing bound: the conservative seed.
node.refresh_tun_mss_ceiling();
let seeded = node.tun_mss_ceiling();
assert_eq!(seeded, crate::upper::icmp::mss_ceiling(1280));
// A wide transport appears, and an edge announces it. No explicit
// refresh call here — draining the edge is the whole trigger.
let wide = make_udp_transport_with_mtu(1, 1452).await;
node.transports.insert(TransportId::new(1), wide);
presence_tx
.send(crate::transport::TransportPresence {
transport_id: TransportId::new(1),
present: true,
health_relevant: true,
})
.await
.expect("presence edge queued");
node.drain_transport_presence();
assert_eq!(
node.tun_mss_ceiling(),
crate::upper::icmp::mss_ceiling(1452),
"draining a presence edge must refresh the ceiling on its own"
);
assert_ne!(
node.tun_mss_ceiling(),
seeded,
"the ceiling stayed at its seed, so the edge did not refresh it"
);
for transport in node.transports.values_mut() {
transport.stop().await.ok();
}
}
#[tokio::test]
async fn test_transport_mtu_fallback_when_no_operational_transports() {
// No transports configured at all → falls back to 1280 (IPv6 minimum).
+143
View File
@@ -9,6 +9,120 @@ use crate::transport::TransportError;
/// Broadcast MAC address.
pub const ETHERNET_BROADCAST: [u8; 6] = [0xff; 6];
/// Whether the named interface exists and is administratively up.
///
/// Presence is `IFF_UP` — the interface exists and the operator has enabled
/// it — and deliberately **not** `IFF_RUNNING`.
///
/// Carrier is a different question from bindability, and only the second one
/// belongs in a bind gate. An `AF_PACKET` socket on a carrier-less bridge is
/// perfectly valid and starts carrying traffic the instant a member port comes
/// up, with no rebind: the socket outlives the carrier. Gating on `IFF_RUNNING`
/// bought nothing and cost three things —
///
/// - `br-lan` on a router with nothing plugged into its LAN ports is `UP` with
/// `NO-CARRIER`, so a perfectly healthy wifi-only router reported `Degraded`
/// forever;
/// - every carrier flap the socket would have survived became an unbind /
/// rebind cycle, which is churn the presence machine then has to damp;
/// - an 802.11s mesh interface that reports `RUNNING` only once it has peered
/// cannot peer, because peering needs beacons, which need a bound socket,
/// which the gate refuses. A deadlock reachable on shipped hardware.
///
/// The signal `IFF_RUNNING` does carry — "is anything plugged in" — is not
/// lost; it is reported alongside presence by [`interface_carrier`] and
/// surfaced in `show_transports`, where an operator can read it without it
/// steering the daemon.
///
/// `getifaddrs` rather than an `SIOCGIFFLAGS` ioctl: it needs no socket, so
/// the presence watcher can poll before any file descriptor exists, and it is
/// spelled the same on Linux and the BSDs.
#[cfg(unix)]
pub fn interface_present(interface: &str) -> bool {
interface_present_probe(interface).unwrap_or(false)
}
/// [`interface_present`], keeping "the probe failed" distinct from "absent".
///
/// `None` means the kernel would not answer. A caller deciding whether to
/// *bind* can treat that as absence and retry on the next tick, which is what
/// [`interface_present`] does. A caller deciding whether to *unbind* must not:
/// see [`interface_has_flags`].
#[cfg(unix)]
pub fn interface_present_probe(interface: &str) -> Option<bool> {
interface_has_flags(interface, libc::IFF_UP as u32)
}
/// The kernel's index for the named interface, or `None` if it does not exist.
///
/// A name is not a device, and neither is a name that is still there. Both
/// backends bind by index — `AF_PACKET` stores `sll_ifindex`, and a BPF
/// descriptor follows the device it was attached to — so an interface deleted
/// and recreated under the same name leaves the socket attached to a device
/// that no longer exists while the *name* resolves perfectly well. Comparing
/// the live index against the one captured at bind is what tells those apart.
#[cfg(unix)]
pub fn interface_index(interface: &str) -> Option<u32> {
let c_name = std::ffi::CString::new(interface).ok()?;
// Cheaper than `getifaddrs`: one syscall, no allocation, no walk.
match unsafe { libc::if_nametoindex(c_name.as_ptr()) } {
0 => None,
idx => Some(idx),
}
}
/// Whether the named interface currently has carrier (`IFF_RUNNING`).
///
/// Reported, never acted on — see [`interface_present`]. `false` for an
/// interface that does not exist, which keeps "no carrier" and "no interface"
/// from being told apart here; presence answers that.
#[cfg(unix)]
pub fn interface_carrier(interface: &str) -> bool {
// Report-only, so a probe failure reads the same as no carrier.
interface_has_flags(interface, (libc::IFF_UP | libc::IFF_RUNNING) as u32).unwrap_or(false)
}
/// Whether the named interface exists and has every flag in `wanted` set, or
/// `None` if the question could not be asked.
///
/// The `None` matters. `getifaddrs` is a netlink dump on Linux and it does
/// fail for reasons that have nothing to do with the interface — `ENOBUFS`
/// under memory pressure or a busy netlink socket, `EMFILE`/`ENFILE` under fd
/// exhaustion, since it opens a socket of its own. Answering `false` there
/// reports a present interface as gone, and a caller holding a live binding
/// would tear a working socket down over a transient syscall failure. Callers
/// that can tell the two apart should.
#[cfg(unix)]
fn interface_has_flags(interface: &str, wanted: u32) -> Option<bool> {
let Ok(c_name) = std::ffi::CString::new(interface) else {
// An interior NUL is not a probe failure — no such interface can
// exist, and no retry will change that.
return Some(false);
};
let mut addrs: *mut libc::ifaddrs = std::ptr::null_mut();
if unsafe { libc::getifaddrs(&mut addrs) } != 0 {
return None;
}
let mut matched = false;
let mut cur = addrs;
while !cur.is_null() {
let entry = unsafe { &*cur };
if !entry.ifa_name.is_null()
&& unsafe { libc::strcmp(entry.ifa_name, c_name.as_ptr()) } == 0
&& entry.ifa_flags & wanted == wanted
{
matched = true;
break;
}
cur = entry.ifa_next;
}
unsafe { libc::freeifaddrs(addrs) };
Some(matched)
}
// Platform-specific PacketSocket implementation.
#[cfg(target_os = "linux")]
#[path = "io_linux.rs"]
@@ -193,6 +307,12 @@ mod async_impl {
/// A received frame: (payload, source_mac).
type Frame = (Vec<u8>, [u8; 6]);
/// Consecutive failed BPF reads before the reader thread gives up.
///
/// Mirrors the receive loop's own error threshold: the point is not to
/// tolerate errors but to end the task so the binder can rebind.
const READ_ERROR_EXIT_THRESHOLD: u32 = 5;
pub struct AsyncPacketSocket {
inner: Arc<PacketSocket>,
/// `None` once shutdown has taken the receiver, which is what makes
@@ -219,6 +339,9 @@ mod async_impl {
let mut parse_buf = vec![0u8; bpf_buflen];
let mut parse_offset: usize = 0;
let mut parse_len: usize = 0;
// Consecutive failed reads, to bound a socket whose
// interface went away underneath it.
let mut read_errors: u32 = 0;
let nfds = bpf_fd.max(shutdown_fd) + 1;
loop {
@@ -286,11 +409,31 @@ mod async_impl {
if err.raw_os_error() == Some(libc::EBADF) {
break;
}
if err.kind() == std::io::ErrorKind::Interrupted {
continue;
}
}
// Anything else — `ENXIO` is the one that matters,
// which is what BPF answers once the interface it
// was attached to is torn away — used to loop here
// forever. That mattered beyond the spin: the
// binder's detach check asks whether this thread is
// still running, so a thread that never returns
// reports a dead socket as a live one, and the
// transport sits `present` and deaf until the name
// or index happens to change too. Give up after a
// streak and let the return close the channel,
// which fails `recv_from`, which ends the tokio
// task the binder is actually watching.
read_errors += 1;
if read_errors >= READ_ERROR_EXIT_THRESHOLD {
break;
}
parse_len = 0;
parse_offset = 0;
continue;
}
read_errors = 0;
parse_len = ret as usize;
parse_offset = 0;
}
+13 -5
View File
@@ -59,6 +59,14 @@ impl PacketSocket {
if ret < 0 {
let err = std::io::Error::last_os_error();
unsafe { libc::close(fd) };
// The interface can disappear between the index lookup and the
// bind. That is absence arriving a few microseconds late, not a
// configuration fault, so it reports as absence.
if matches!(err.raw_os_error(), Some(libc::ENODEV) | Some(libc::ENXIO)) {
return Err(TransportError::InterfaceUnavailable {
interface: interface.to_string(),
});
}
return Err(TransportError::StartFailed(format!(
"bind(AF_PACKET, {}) failed: {}",
interface, err
@@ -231,11 +239,11 @@ fn get_if_index(_fd: RawFd, interface: &str) -> Result<i32, TransportError> {
let idx = unsafe { libc::if_nametoindex(c_name.as_ptr()) };
if idx == 0 {
return Err(TransportError::StartFailed(format!(
"interface not found: {} ({})",
interface,
std::io::Error::last_os_error()
)));
// Absence, not a fault: the caller's presence watcher rebinds when the
// interface shows up. See `TransportError::InterfaceUnavailable`.
return Err(TransportError::InterfaceUnavailable {
interface: interface.to_string(),
});
}
Ok(idx as i32)
}
+15 -7
View File
@@ -393,10 +393,18 @@ fn bind_to_interface(fd: RawFd, interface: &str) -> Result<(), TransportError> {
let ret = unsafe { libc::ioctl(fd, BIOCSETIF, ifreq.as_ptr()) };
if ret < 0 {
let err = std::io::Error::last_os_error();
// BIOCSETIF answers ENXIO for an interface that is not there. The
// interface can also vanish between the index lookup and this ioctl,
// so absence is reported as absence rather than as a bind fault.
if matches!(err.raw_os_error(), Some(libc::ENXIO) | Some(libc::ENODEV)) {
return Err(TransportError::InterfaceUnavailable {
interface: interface.to_string(),
});
}
return Err(TransportError::StartFailed(format!(
"BIOCSETIF({}) failed: {}",
interface,
std::io::Error::last_os_error()
interface, err
)));
}
Ok(())
@@ -498,11 +506,11 @@ fn get_if_index(interface: &str) -> Result<i32, TransportError> {
let idx = unsafe { libc::if_nametoindex(c_name.as_ptr()) };
if idx == 0 {
return Err(TransportError::StartFailed(format!(
"interface not found: {} ({})",
interface,
std::io::Error::last_os_error()
)));
// Absence, not a fault — see the Linux twin and
// `TransportError::InterfaceUnavailable`.
return Err(TransportError::InterfaceUnavailable {
interface: interface.to_string(),
});
}
Ok(idx as i32)
}
File diff suppressed because it is too large Load Diff
+211
View File
@@ -25,6 +25,23 @@ pub mod ethernet;
#[cfg(unix)]
pub(crate) mod watcher;
/// Presence lifecycle for a transport bound to a local resource that can
/// disappear and come back: the phase machine, the absence policy, and the
/// damping that keeps a flapping resource from flapping node health with it.
///
/// Transport-agnostic on purpose. Only the *probe* — "is my thing there, and
/// is it still the same one?" — is specific to what is bound, and that stays
/// with the transport that knows how to ask.
///
/// Crate-internal on purpose, for the same reason as `watcher` above: it is a
/// mechanism the crate's own transports share, not a surface an embedder
/// builds against. `ethernet` re-exports the two types it used to own, so the
/// published path stays `transport::ethernet::{AbsencePolicy, Presence}`.
/// Gated with the one transport that binds through it today; widen the gate
/// when a second binder arrives.
#[cfg(any(target_os = "linux", target_os = "macos"))]
pub(crate) mod presence;
#[cfg(ble_available)]
pub mod ble;
@@ -112,6 +129,62 @@ pub fn packet_channel(buffer: usize) -> (PacketTx, PacketRx) {
tokio::sync::mpsc::channel(buffer)
}
/// Operator-visible interface presence, rendered by `show_transports`.
///
/// Worth as much as the retry itself. The original boot-race bug was expensive
/// precisely because the 802.11s peer link formed regardless of the daemon, so
/// nothing an operator could see said the node was deaf.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub struct InterfacePresence {
/// `absent`, `binding`, or `present`.
pub presence: &'static str,
/// Whether the interface currently has carrier (`IFF_RUNNING`).
///
/// Reported, never acted on. Presence is `IFF_UP`, because binding does
/// not need carrier and a socket outlives a carrier flap — but "is
/// anything plugged in" is still what an operator wants to know when a
/// bound transport is carrying nothing, so it is reported here instead of
/// steering the daemon.
pub carrier: bool,
/// `required` or `optional`.
pub policy: &'static str,
/// How long the current phase has been held.
pub since_secs: u64,
/// Successful binds since the transport was created (`1` after a clean
/// start; more means it has rebound).
pub binds: u64,
/// Failed bind attempts since the last successful bind.
pub failed_attempts: u32,
}
/// A presence edge published by an interface-bound transport.
///
/// Absence and return are the same transition seen from two sides, so one
/// event type carries both: `present: false` on detach (including a start
/// where the interface was never there), `present: true` on every successful
/// bind after the first observation.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub struct TransportPresence {
/// The transport whose interface changed presence.
pub transport_id: TransportId,
/// Whether the interface is now bound.
pub present: bool,
/// Whether this edge should move node health.
///
/// `false` for an `optional` interface, whose absence is normal and must
/// not take the node off `Full`. The edge is still published, because the
/// bound set changed either way and the node's egress MTU floor is derived
/// from it — health and MTU are two different questions riding one
/// channel, and only the first one is policy-filtered.
pub health_relevant: bool,
}
/// Channel sender for transport presence edges.
pub type PresenceTx = tokio::sync::mpsc::Sender<TransportPresence>;
/// Channel receiver for transport presence edges.
pub type PresenceRx = tokio::sync::mpsc::Receiver<TransportPresence>;
// ============================================================================
// Errors
// ============================================================================
@@ -128,6 +201,20 @@ pub enum TransportError {
#[error("transport failed to start: {0}")]
StartFailed(String),
/// The named network interface is not usable right now: it does not exist,
/// or it exists but is administratively down (no `IFF_UP`).
///
/// Distinct from [`TransportError::StartFailed`] because absence is a
/// *state*, not a fault. Interface-bound transports treat it as "not bound
/// yet" and keep a presence watcher running; a `StartFailed` carrying the
/// same text could not be told apart from a typo'd interface name or a
/// missing capability.
#[error("interface unavailable: {interface}")]
InterfaceUnavailable {
/// The configured interface name.
interface: String,
},
#[error("transport shutdown failed: {0}")]
ShutdownFailed(String),
@@ -159,6 +246,49 @@ pub enum TransportError {
Io(#[from] std::io::Error),
}
impl TransportError {
/// Whether this failure is expected to clear on its own.
///
/// The distinction callers need is not *what* went wrong but whether
/// waiting fixes it. A transient failure means the operation was refused
/// by a condition the daemon is already working to resolve, so the state
/// built up around it — a half-finished handshake, a route, a queued
/// packet — is worth keeping. A terminal one means the state is worth
/// tearing down.
///
/// This lives here, on the error, rather than being re-derived at each
/// call site: `InterfaceUnavailable` used to be flattened into a
/// formatted string on its way out of the transport layer, so every
/// caller downstream saw a generic send failure and could only treat a
/// two-second interface flap exactly as it treated a permanent fault.
///
/// Deliberately narrow. [`Self::Timeout`] and [`Self::ConnectionRefused`]
/// are *not* transient here: they describe a remote that did not answer,
/// which is a statement about the peer rather than about this node's
/// ability to transmit, and the existing retry paths for them already sit
/// at a different layer.
pub fn is_transient(&self) -> bool {
match self {
// The interface is absent or mid-rebind. The binder is polling for
// it and will bind it the moment it returns.
Self::InterfaceUnavailable { .. } => true,
Self::NotStarted
| Self::AlreadyStarted
| Self::StartFailed(_)
| Self::ShutdownFailed(_)
| Self::LinkFailed(_)
| Self::SendFailed(_)
| Self::RecvFailed(_)
| Self::InvalidAddress(_)
| Self::MtuExceeded { .. }
| Self::Timeout
| Self::ConnectionRefused
| Self::NotSupported(_)
| Self::Io(_) => false,
}
}
}
// ============================================================================
// Transport Type Metadata
// ============================================================================
@@ -874,6 +1004,52 @@ impl TransportHandle {
}
}
/// Interface presence for interface-bound transports: the phase label, the
/// absence policy, how long the phase has been held, and the failed-bind
/// count since the last successful bind.
///
/// `None` for transports that are not bound to a named interface — for
/// those, presence is not a concept and an operator should not be shown an
/// always-`present` column.
pub fn interface_presence(&self) -> Option<InterfacePresence> {
match self {
#[cfg(any(target_os = "linux", target_os = "macos"))]
TransportHandle::Ethernet(t) => {
let state = t.presence_state();
Some(InterfacePresence {
presence: t.presence().as_str(),
carrier: t.has_carrier(),
policy: t.absence_policy().as_str(),
since_secs: state.since().as_secs(),
binds: state.binds(),
failed_attempts: state.attempts(),
})
}
_ => None,
}
}
/// Whether this transport can actually put a frame on the wire *now*.
///
/// [`Self::is_operational`] answers a different question: it means the
/// transport was started, which for an interface-bound transport no longer
/// implies a live socket — that is the whole point of presence. Callers
/// that are choosing a transport to use, or deriving a value from one,
/// want this; callers reasoning about lifecycle want `is_operational`.
///
/// `true` for every transport that is not interface-bound, so this is
/// `is_operational` with the presence refinement applied where it exists.
pub fn is_bound(&self) -> bool {
if !self.is_operational() {
return false;
}
match self {
#[cfg(any(target_os = "linux", target_os = "macos"))]
TransportHandle::Ethernet(t) => t.presence() == ethernet::Presence::Present,
_ => true,
}
}
/// Get the interface name (Ethernet only, returns None for other transports).
pub fn interface_name(&self) -> Option<&str> {
match self {
@@ -1571,4 +1747,39 @@ mod tests {
assert_eq!(handle.link_mtu(&addr), expected_mtu);
assert_eq!(handle.link_mtu(&addr), handle.mtu());
}
#[test]
fn only_an_absent_interface_classifies_as_transient() {
// The whole point of the classification is that it is narrow. An
// interface the binder is already polling for will come back; nothing
// else on this list resolves itself by waiting, and treating one of
// them as transient would mean holding state open for a fault that is
// never going to clear.
assert!(
TransportError::InterfaceUnavailable {
interface: "eth0".into()
}
.is_transient()
);
for terminal in [
TransportError::NotStarted,
TransportError::AlreadyStarted,
TransportError::StartFailed("no CAP_NET_RAW".into()),
TransportError::SendFailed("ENOBUFS".into()),
TransportError::MtuExceeded {
packet_size: 2000,
mtu: 1500,
},
// Deliberately terminal: both describe a remote that did not
// answer, not this node's inability to transmit.
TransportError::Timeout,
TransportError::ConnectionRefused,
] {
assert!(
!terminal.is_transient(),
"{terminal:?} must not be classified transient"
);
}
}
}
+851
View File
@@ -0,0 +1,851 @@
//! Interface presence state for interface-bound transports.
//!
//! A transport bound to a network interface is a long-lived object that is
//! *sometimes bound*. The interface it names may not exist when the daemon
//! starts, may appear minutes later, may vanish and return mid-operation, and
//! may never appear at all. This module holds the state that makes that
//! observable and drives the rebind loop in [`super`].
//!
//! ```text
//! Absent ──attach──> Binding ──ok──> Present
//! ^ │ │
//! └──── fail/backoff ─┘ │
//! └──────────── detach ──────────────┘
//! ```
//!
//! Two invariants do the work:
//!
//! - **The transport object survives detach.** Config, `TransportId`,
//! statistics, and the neighbor buffer persist; only the file descriptor and
//! its loops go. A transport is never destroyed because its interface went
//! away.
//! - **Start-time absence and runtime detach are the same transition.** A node
//! that boots before wifi and a node whose wifi reloads at 03:00 take one
//! code path.
//!
//! Presence tracks `IFF_UP` — the interface exists and the operator has
//! enabled it — and deliberately not `IFF_RUNNING`: binding needs no carrier,
//! and a socket outlives a carrier flap. Carrier is reported alongside it
//! rather than steering it. See
//! the ethernet transport's `interface_present` for why.
use std::sync::atomic::{AtomicU8, AtomicU32, AtomicU64, Ordering};
use std::sync::{PoisonError, RwLock, RwLockReadGuard, RwLockWriteGuard};
use std::time::{Duration, Instant};
/// Read a lock, ignoring poisoning.
///
/// Every value guarded in this module is plain data — an `Instant`, an
/// `Option<[u8; 6]>` — that a panic mid-write cannot leave logically
/// inconsistent, so poisoning carries no information worth propagating.
/// Treating it as a failure is what would hurt: the callers here are on the
/// presence path, and "assume the worst" there means a transport that reports
/// itself bound while every send fails, or a binder that tears down and
/// rebinds every second forever. A stuck state is a worse outcome than
/// reading a byte written by a thread that later panicked.
fn read<T>(lock: &RwLock<T>) -> RwLockReadGuard<'_, T> {
lock.read().unwrap_or_else(PoisonError::into_inner)
}
/// Write a lock, ignoring poisoning. See [`read`].
fn write<T>(lock: &RwLock<T>) -> RwLockWriteGuard<'_, T> {
lock.write().unwrap_or_else(PoisonError::into_inner)
}
/// How absence of the configured interface is reported.
///
/// Describes *the interface's presence*, not the transport's importance: an
/// optional interface that is present is used exactly as hard as any other.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub enum AbsencePolicy {
/// Naming an interface in configuration is a statement that you expect it,
/// so the default is to complain: absence degrades node health from the
/// first edge, and if it outlasts [`ABSENCE_ERROR_AFTER`] — the window in
/// which it could still have been an ordinary bring-up race — it is
/// reported once at `error`.
Required,
/// Absence is normal for this interface (a dock adapter, a radio that only
/// exists on some hardware): no health impact, `info` on the edge.
Optional,
}
impl AbsencePolicy {
/// `optional: true` in configuration selects [`AbsencePolicy::Optional`].
pub fn from_optional(optional: bool) -> Self {
if optional {
Self::Optional
} else {
Self::Required
}
}
/// Whether absence should be hidden from node health.
pub fn is_optional(self) -> bool {
matches!(self, Self::Optional)
}
/// Operator-facing label, used by `show_transports`.
pub fn as_str(self) -> &'static str {
match self {
Self::Required => "required",
Self::Optional => "optional",
}
}
}
/// Where a transport sits in the presence cycle.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub enum Presence {
/// The interface is not there (or is there without carrier). No socket, no
/// loops; the watcher is waiting.
Absent,
/// The interface appeared and a bind is in flight, or a non-absence bind
/// failure is backing off.
Binding,
/// Bound, with a live socket and running loops.
Present,
}
impl Presence {
fn from_u8(v: u8) -> Self {
match v {
1 => Self::Binding,
2 => Self::Present,
_ => Self::Absent,
}
}
fn as_u8(self) -> u8 {
match self {
Self::Absent => 0,
Self::Binding => 1,
Self::Present => 2,
}
}
/// Operator-facing label, used by `show_transports`.
pub fn as_str(self) -> &'static str {
match self {
Self::Absent => "absent",
Self::Binding => "binding",
Self::Present => "present",
}
}
}
impl std::fmt::Display for Presence {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
f.write_str(self.as_str())
}
}
/// Shared, lock-light presence state.
///
/// Written by the transport's binder task, read by `send`, by the control
/// plane, and by tests. Held behind an `Arc` so the binder task can outlive
/// any particular borrow of the transport.
#[derive(Debug)]
pub struct PresenceState {
phase: AtomicU8,
/// Successful binds since the transport was created. `1` after a clean
/// start; every increment past that is a rebind.
binds: AtomicU64,
/// Failed bind attempts since the last successful bind. Reset on bind.
attempts: AtomicU32,
/// When the current phase was entered.
since: RwLock<Instant>,
/// MAC observed at the last successful bind. A name reappearing with a
/// different MAC is different hardware, not the same device returning.
last_mac: RwLock<Option<[u8; 6]>>,
}
impl Default for PresenceState {
fn default() -> Self {
Self::new()
}
}
impl PresenceState {
/// A fresh tracker in [`Presence::Absent`].
pub fn new() -> Self {
Self {
phase: AtomicU8::new(Presence::Absent.as_u8()),
binds: AtomicU64::new(0),
attempts: AtomicU32::new(0),
since: RwLock::new(Instant::now()),
last_mac: RwLock::new(None),
}
}
/// Current phase.
pub fn presence(&self) -> Presence {
Presence::from_u8(self.phase.load(Ordering::Acquire))
}
/// Whether the transport currently holds a bound socket.
pub fn is_present(&self) -> bool {
self.presence() == Presence::Present
}
/// How long the current presence *episode* has been held.
///
/// Episode, not phase: `Binding` is part of the absence episode until it
/// succeeds. An interface that has been gone for a week while a bind is
/// retried and refused every second must report a week, not one second —
/// otherwise [`ABSENCE_ERROR_AFTER`] is never reached and the
/// operator-facing `since_secs` reads as a healthy young absence forever.
pub fn since(&self) -> Duration {
read(&self.since).elapsed()
}
/// Restart the episode clock, for a transport that is about to start.
///
/// `new()` stamps the clock at construction, but construction and
/// `start_async` need not be adjacent — config load and supervisor staging
/// sit between them. Left alone, a transport staged for longer than
/// [`ABSENCE_ERROR_AFTER`] logs the sustained-absence error on its very
/// first binder tick, having given the interface no bring-up window at
/// all. The window is supposed to absorb exactly that race.
pub fn mark_starting(&self) {
*write(&self.since) = Instant::now();
}
/// Test-only: age the episode clock, so deadline behaviour can be tested
/// without sleeping through it.
///
/// `ABSENCE_ERROR_AFTER` is ten seconds. A test that waited it out would
/// be ten seconds of nothing, which is how deadline logic ends up
/// untested.
#[cfg(test)]
pub(crate) fn backdate_for_test(&self, by: Duration) {
let mut since = write(&self.since);
*since = since.checked_sub(by).unwrap_or(*since);
}
/// Successful binds since creation (`1` after a clean start).
pub fn binds(&self) -> u64 {
self.binds.load(Ordering::Relaxed)
}
/// Failed bind attempts since the last successful bind.
pub fn attempts(&self) -> u32 {
self.attempts.load(Ordering::Relaxed)
}
/// MAC observed at the last successful bind, if any.
pub fn last_mac(&self) -> Option<[u8; 6]> {
*read(&self.last_mac)
}
/// Move to `phase`, returning `true` if this was an actual edge.
///
/// Edge-vs-level is what the logging policy keys on: logged once on
/// entering absence and once on recovery, never per retry attempt.
///
/// The episode clock ([`Self::since`]) restarts only when *boundness*
/// changes — Present↔not-Present. A failed bind walks
/// `Absent → Binding → Absent`, and resetting the clock on those would
/// hide a permanent absence behind a timer that never gets past one
/// second.
pub fn transition(&self, phase: Presence) -> bool {
let prev = Presence::from_u8(self.phase.swap(phase.as_u8(), Ordering::AcqRel));
if prev == phase {
return false;
}
if (prev == Presence::Present) != (phase == Presence::Present) {
*write(&self.since) = Instant::now();
}
true
}
/// Record a successful bind at `mac`.
///
/// Returns `true` when the interface came back as *different hardware* —
/// the name reappeared with a MAC other than the one last bound. The
/// caller drops cached neighbor state rather than silently resuming onto
/// a different adapter.
pub fn record_bind(&self, mac: [u8; 6]) -> bool {
let changed = match *read(&self.last_mac) {
Some(prev) => prev != mac,
None => false,
};
*write(&self.last_mac) = Some(mac);
self.binds.fetch_add(1, Ordering::Relaxed);
self.attempts.store(0, Ordering::Relaxed);
self.transition(Presence::Present);
changed
}
/// Record a failed bind attempt, returning the new attempt count.
pub fn record_attempt(&self) -> u32 {
self.attempts.fetch_add(1, Ordering::Relaxed) + 1
}
}
/// Backoff for bind failures that are *not* absence — permission denied,
/// buffer sizing, a BPF device shortage. Absence itself does not back off
/// where an event source is available: there is nothing to poll.
///
/// 1 s doubling to a 30 s ceiling.
pub fn bind_backoff(attempts: u32) -> Duration {
const BASE_SECS: u64 = 1;
const CEILING_SECS: u64 = 30;
let shift = attempts.saturating_sub(1).min(5);
Duration::from_secs((BASE_SECS << shift).min(CEILING_SECS))
}
/// How long a *required* interface may be absent before it is an error.
///
/// Absence is a state the presence machine handles, so it is not an error for
/// happening — a daemon that wins the race against its own radio, or a cable
/// out for two seconds, is the ordinary case this mechanism exists to absorb,
/// and calling that an error at t=0 and "recovered" at t=0.2 s is cry-wolf.
/// Past this window it is no longer a race: something an operator has to fix
/// is wrong, and the log should say so once.
///
/// One window for both shapes of absence. A node that boots before its wifi
/// and a node whose wifi reloads at 03:00 take one code path everywhere else
/// in this module; giving them different deadlines would reintroduce exactly
/// the start-versus-runtime asymmetry the presence machine removed.
///
/// Tuned against the platforms this exists for: comfortably past a veth or a
/// container coming up, short enough that a mesh radio which never appears is
/// named while somebody is still watching the boot. Raising it hides a real
/// fault for longer; lowering it starts reporting ordinary bring-up races.
pub const ABSENCE_ERROR_AFTER: Duration = Duration::from_secs(10);
/// Minimum lifetime for a binding to count as a real recovery.
///
/// A socket that dies sooner than this never really came back.
pub const MIN_STABLE_BINDING: Duration = Duration::from_secs(10);
/// Consecutive short-lived bindings before the binder stops treating a
/// successful bind as a recovery.
pub const CHURN_THRESHOLD: u32 = 3;
/// Damping for the rebind loop.
///
/// Backoff covers *failed* binds; this covers the opposite and nastier case —
/// binds that keep **succeeding** into a socket that dies moments later. A
/// receive loop that gives up on a persistent error while the interface stays
/// `UP` produces exactly that: tear down, rebind, succeed, fail again, once
/// per second, forever. Undamped it is an `error!`/`info!` pair and a
/// `Degraded`→`Running` health flap every cycle, which defeats both the
/// "log edges, not attempts" rule and the meaning of `Degraded`.
///
/// So: count consecutive bindings that die young, back off between them on
/// the same 1 s → 30 s curve, and once the streak reaches
/// [`CHURN_THRESHOLD`] stop announcing each bind as a recovery — hold the
/// node at its degraded reading until a binding actually survives
/// [`MIN_STABLE_BINDING`]. A binding that holds ends the streak.
///
/// Pure state, driven by an injected clock, so the policy is testable without
/// a network interface.
#[derive(Debug, Default)]
pub struct ChurnGuard {
/// Consecutive bindings that died younger than [`MIN_STABLE_BINDING`].
streak: u32,
/// When the current binding was established.
bound_at: Option<Instant>,
/// Whether the current binding was announced as a recovery.
announced: bool,
}
/// What the caller should do about a successful bind.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub struct BindOutcome {
/// Log the recovery and publish presence now. `false` while churning:
/// the bind is held back until it proves it will last.
pub announce: bool,
}
/// What the caller should do about a detach.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub struct DetachOutcome {
/// Publish absence. `false` when the binding that just died was never
/// announced, so there is nothing to retract.
pub retract: bool,
/// Log the detach edge. `false` once churning — the fact has been said.
pub log_edge: bool,
/// This detach is the one that crossed [`CHURN_THRESHOLD`]; say so once.
pub entered_churn: bool,
/// Wait this long before trying to bind again.
pub backoff: Option<Duration>,
}
impl ChurnGuard {
/// A guard with no history.
pub fn new() -> Self {
Self::default()
}
/// Consecutive short-lived bindings, for logging and tests.
pub fn streak(&self) -> u32 {
self.streak
}
/// Record a successful bind.
pub fn bound(&mut self, now: Instant) -> BindOutcome {
self.bound_at = Some(now);
let announce = self.streak < CHURN_THRESHOLD;
if announce {
self.announced = true;
}
BindOutcome { announce }
}
/// Called on every tick while bound. Returns `true` exactly once, at the
/// moment a held-back binding has proved stable and should be announced.
pub fn stabilized(&mut self, now: Instant) -> bool {
let Some(bound_at) = self.bound_at else {
return false;
};
if now.duration_since(bound_at) < MIN_STABLE_BINDING {
return false;
}
let newly_announced = !self.announced;
self.streak = 0;
self.announced = true;
newly_announced
}
/// Record a detach.
pub fn detached(&mut self, now: Instant) -> DetachOutcome {
let young = self
.bound_at
.is_some_and(|t| now.duration_since(t) < MIN_STABLE_BINDING);
self.bound_at = None;
if young {
self.streak += 1;
} else {
self.streak = 0;
}
DetachOutcome {
retract: std::mem::take(&mut self.announced),
log_edge: self.streak < CHURN_THRESHOLD,
entered_churn: self.streak == CHURN_THRESHOLD,
backoff: (self.streak > 0).then(|| bind_backoff(self.streak)),
}
}
}
#[cfg(test)]
mod tests {
use super::*;
use std::sync::Arc;
#[test]
fn presence_starts_absent() {
let p = PresenceState::new();
assert_eq!(p.presence(), Presence::Absent);
assert_eq!(p.binds(), 0);
assert_eq!(p.attempts(), 0);
assert!(p.last_mac().is_none());
}
#[test]
fn transition_reports_edges_only() {
let p = PresenceState::new();
assert!(p.transition(Presence::Binding));
assert!(!p.transition(Presence::Binding));
assert!(p.transition(Presence::Present));
}
#[test]
fn bind_records_mac_and_clears_attempts() {
let p = PresenceState::new();
p.record_attempt();
p.record_attempt();
assert_eq!(p.attempts(), 2);
let mac = [0x02, 0, 0, 0, 0, 1];
assert!(!p.record_bind(mac), "first bind is not a hardware change");
assert_eq!(p.attempts(), 0);
assert_eq!(p.binds(), 1);
assert_eq!(p.presence(), Presence::Present);
assert_eq!(p.last_mac(), Some(mac));
}
#[test]
fn rebind_on_same_mac_is_not_a_hardware_change() {
let p = PresenceState::new();
let mac = [0x02, 0, 0, 0, 0, 1];
p.record_bind(mac);
p.transition(Presence::Absent);
assert!(!p.record_bind(mac));
assert_eq!(p.binds(), 2);
}
#[test]
fn rebind_on_different_mac_is_a_hardware_change() {
let p = PresenceState::new();
p.record_bind([0x02, 0, 0, 0, 0, 1]);
p.transition(Presence::Absent);
assert!(p.record_bind([0x02, 0, 0, 0, 0, 2]));
}
#[test]
fn backoff_climbs_to_a_ceiling() {
assert_eq!(bind_backoff(0), Duration::from_secs(1));
assert_eq!(bind_backoff(1), Duration::from_secs(1));
assert_eq!(bind_backoff(2), Duration::from_secs(2));
assert_eq!(bind_backoff(3), Duration::from_secs(4));
assert_eq!(bind_backoff(6), Duration::from_secs(30));
assert_eq!(bind_backoff(u32::MAX), Duration::from_secs(30));
}
#[test]
fn the_error_deadline_outlasts_an_ordinary_bring_up_race() {
// The window has to clear the races it exists to absorb — a veth
// arriving a fraction of a second late, a container starting — while
// staying short enough that a radio which never appears is named
// during the boot somebody is watching. It is also the one deadline:
// start-time absence and a runtime detach share it.
assert!(
ABSENCE_ERROR_AFTER >= Duration::from_secs(5),
"shorter than a bring-up race would report the ordinary case"
);
assert!(
ABSENCE_ERROR_AFTER <= Duration::from_secs(60),
"longer and a required interface that never appears goes unsaid \
for the whole boot"
);
}
// ── The absence clock measures an episode, not a phase ────────────────
#[test]
fn a_failed_bind_does_not_restart_the_absence_clock() {
// An interface that is present but refuses to bind walks
// Absent → Binding → Absent on every retry. If those edges reset the
// clock, `since_secs` reads as a one-second-old absence forever and
// the error deadline is never reached — so a permission error would
// sit silently behind a healthy-looking counter.
let p = PresenceState::new();
std::thread::sleep(Duration::from_millis(30));
let before = p.since();
assert!(p.transition(Presence::Binding));
assert!(p.transition(Presence::Absent));
assert!(p.transition(Presence::Binding));
assert!(p.transition(Presence::Absent));
assert!(
p.since() >= before,
"the absence clock ran backwards across failed binds"
);
}
#[test]
fn the_clock_restarts_only_when_boundness_changes() {
let p = PresenceState::new();
std::thread::sleep(Duration::from_millis(30));
// Absent → Present restarts it: a new episode began.
p.record_bind([0x02, 0, 0, 0, 0, 1]);
assert!(p.since() < Duration::from_millis(30));
std::thread::sleep(Duration::from_millis(30));
let bound_for = p.since();
// Present → Present is not an edge at all.
assert!(!p.transition(Presence::Present));
assert!(p.since() >= bound_for);
// Present → Absent restarts it: the episode ended.
assert!(p.transition(Presence::Absent));
assert!(p.since() < Duration::from_millis(30));
}
// ── Poisoning must not be a stuck state ───────────────────────────────
#[test]
fn a_poisoned_lock_still_reports_presence() {
// `.ok()`-style handling would make a poisoned lock read as "no MAC,
// no socket, tasks dead" — a transport reporting itself present while
// every send fails, and a binder rebinding once a second forever.
// Poisoning carries no information about plain data, so it is ignored.
let p = Arc::new(PresenceState::new());
p.record_bind([0x02, 0, 0, 0, 0, 7]);
let poisoner = Arc::clone(&p);
let panicked = std::thread::spawn(move || {
let _guard = poisoner.last_mac.write().unwrap();
panic!("poison the lock while holding it");
})
.join();
assert!(panicked.is_err(), "the helper thread was supposed to panic");
assert!(
p.last_mac.is_poisoned(),
"the lock was supposed to be poisoned"
);
assert_eq!(
p.last_mac(),
Some([0x02, 0, 0, 0, 0, 7]),
"a poisoned lock must not erase the binding"
);
// And the clock still answers rather than collapsing to zero.
let _ = p.since();
}
// ── Rebind churn ──────────────────────────────────────────────────────
#[test]
fn a_healthy_bind_and_detach_is_not_churn() {
let mut g = ChurnGuard::new();
let t0 = Instant::now();
assert!(g.bound(t0).announce, "a first bind is a recovery");
// Held well past the stability floor, then lost.
let out = g.detached(t0 + MIN_STABLE_BINDING + Duration::from_secs(60));
assert!(out.retract, "an announced binding must be retracted");
assert!(out.log_edge, "an isolated detach is worth a line");
assert!(!out.entered_churn);
assert_eq!(out.backoff, None, "one clean outage must not back off");
assert_eq!(g.streak(), 0);
}
#[test]
fn an_unseeded_guard_retracts_nothing_and_never_repairs_itself() {
// Why `binder_loop` seeds the guard when it inherits a binding from
// `start_async`, rather than leaving it fresh.
//
// A guard that was never told about a bind believes it has announced
// nothing, so it asks for no retraction — and health, which learned
// `present: true` from the inline bind, would keep reading `Full` with
// the interface gone. `stabilized` cannot rescue it either: with no
// `bound_at` there is nothing for it to judge stable.
let mut g = ChurnGuard::new();
let t0 = Instant::now();
assert!(
!g.stabilized(t0 + MIN_STABLE_BINDING + Duration::from_secs(60)),
"a guard with no recorded bind has nothing to stabilize"
);
let out = g.detached(t0 + Duration::from_secs(60));
assert!(
!out.retract,
"an unseeded guard retracts nothing — which is exactly why the \
binder must seed it from the inline bind"
);
}
#[test]
fn a_seeded_guard_retracts_the_edge_the_inline_bind_published() {
// The fix, from the binder's angle: seeding with `bound` is what makes
// the first detach after a clean start reach node health.
let mut g = ChurnGuard::new();
let t0 = Instant::now();
g.bound(t0);
let out = g.detached(t0 + MIN_STABLE_BINDING + Duration::from_secs(60));
assert!(
out.retract,
"the edge `start_async` published must be retracted on detach"
);
assert!(out.log_edge);
}
#[test]
fn short_lived_bindings_back_off() {
// The failure this guards: a receive loop that gives up on a
// persistent error while the interface stays UP. Bind succeeds, dies,
// rebinds, dies — once per second, forever, undamped.
let mut g = ChurnGuard::new();
let mut t = Instant::now();
for expected in [1u64, 2, 4] {
g.bound(t);
t += Duration::from_secs(1);
let out = g.detached(t);
assert_eq!(
out.backoff,
Some(Duration::from_secs(expected)),
"streak {} should back off {expected}s",
g.streak()
);
}
assert_eq!(g.streak(), 3);
}
#[test]
fn churn_stops_announcing_and_stops_logging() {
let mut g = ChurnGuard::new();
let mut t = Instant::now();
// The detaches below the threshold are still news and still logged.
for _ in 0..CHURN_THRESHOLD - 1 {
assert!(g.bound(t).announce);
t += Duration::from_secs(1);
let out = g.detached(t);
assert!(out.log_edge, "the first few detaches are still news");
assert!(!out.entered_churn);
}
// The detach that crosses the threshold reports the churn instead of
// the edge: one line saying "this keeps happening", not two saying
// "it happened" and "it keeps happening".
assert!(g.bound(t).announce);
t += Duration::from_secs(1);
let crossing = g.detached(t);
assert!(crossing.entered_churn, "crossing must be announced once");
assert!(!crossing.log_edge, "the churn line replaces the edge line");
// Past the threshold: bindings are no longer announced as recoveries,
// so node health stays put instead of flapping every second, and the
// edges stop being logged.
for _ in 0..5 {
assert!(!g.bound(t).announce, "a churning bind is not a recovery");
t += Duration::from_secs(1);
let out = g.detached(t);
assert!(!out.log_edge, "churn must not log per cycle");
assert!(!out.retract, "nothing was announced, so nothing to retract");
assert!(
!out.entered_churn,
"the threshold is crossed once, not repeatedly"
);
}
}
#[test]
fn backoff_during_churn_is_capped() {
let mut g = ChurnGuard::new();
let mut t = Instant::now();
let mut last = None;
for _ in 0..12 {
g.bound(t);
t += Duration::from_secs(1);
last = g.detached(t).backoff;
}
assert_eq!(
last,
Some(Duration::from_secs(30)),
"churn backoff must climb to the ceiling and stop"
);
}
#[test]
fn a_binding_that_lasts_ends_the_streak_and_announces_once() {
let mut g = ChurnGuard::new();
let mut t = Instant::now();
// Churn into the held-back state.
for _ in 0..CHURN_THRESHOLD + 1 {
g.bound(t);
t += Duration::from_secs(1);
g.detached(t);
}
assert!(!g.bound(t).announce);
// Not yet stable: still nothing to say.
assert!(!g.stabilized(t + Duration::from_secs(1)));
// Survived the floor: announce exactly once, and the streak is over.
let stable_at = t + MIN_STABLE_BINDING;
assert!(
g.stabilized(stable_at),
"a binding that lasts is a recovery"
);
assert!(
!g.stabilized(stable_at + Duration::from_secs(60)),
"recovery is announced once, not on every tick"
);
assert_eq!(g.streak(), 0);
// And the next detach behaves like an ordinary one again.
let out = g.detached(stable_at + Duration::from_secs(60));
assert!(out.retract);
assert!(out.log_edge);
assert_eq!(out.backoff, None);
}
#[test]
fn stabilized_is_silent_for_an_ordinary_binding() {
// A bind that was announced immediately must not be announced again
// when it passes the stability floor.
let mut g = ChurnGuard::new();
let t = Instant::now();
assert!(g.bound(t).announce);
assert!(!g.stabilized(t + MIN_STABLE_BINDING + Duration::from_secs(1)));
}
#[test]
fn policy_labels() {
assert_eq!(AbsencePolicy::from_optional(true), AbsencePolicy::Optional);
assert_eq!(AbsencePolicy::from_optional(false), AbsencePolicy::Required);
assert!(AbsencePolicy::Optional.is_optional());
assert!(!AbsencePolicy::Required.is_optional());
assert_eq!(AbsencePolicy::Required.as_str(), "required");
// Both labels, not just one. `show_transports` renders this string and
// fipstop's severity split keys on it, so a swapped pair would paint
// every expected interface as the tolerated kind and vice versa —
// while a test that checks only `Required` stays green through it.
assert_eq!(AbsencePolicy::Optional.as_str(), "optional");
}
#[test]
fn a_first_bind_is_not_a_hardware_change() {
// The boundary the flush hangs off. `record_bind` returns "different
// hardware", and on the very first bind there is no previous MAC to
// differ from — so it must answer false, or every clean start would
// drop a neighbour cache it had just built and log a hardware swap
// that never happened.
let state = PresenceState::new();
assert!(
!state.record_bind([1, 2, 3, 4, 5, 6]),
"the first bind has nothing to differ from"
);
assert_eq!(state.binds(), 1);
assert_eq!(state.presence(), Presence::Present);
}
#[test]
fn a_rebind_on_new_hardware_reports_the_change_once() {
// And it reports the change once, not on every subsequent bind: the
// caller drops its cached neighbours on a `true`, so a sticky answer
// would flush the cache on every rebind forever.
let state = PresenceState::new();
state.record_bind([1, 2, 3, 4, 5, 6]);
assert!(
state.record_bind([0xaa, 0xbb, 0xcc, 0xdd, 0xee, 0xff]),
"a name returning on a different MAC is different hardware"
);
assert!(
!state.record_bind([0xaa, 0xbb, 0xcc, 0xdd, 0xee, 0xff]),
"the same hardware rebinding is not a change"
);
assert_eq!(state.binds(), 3);
}
#[test]
fn every_presence_label_is_distinct_and_round_trips() {
// The labels are the operator-facing vocabulary — `show_transports`
// emits them and the fipstop State column renders them — and
// `Presence::as_str` had no test at all, so `binding` in particular was
// never observed by anything.
use std::collections::HashSet;
let all = [Presence::Absent, Presence::Binding, Presence::Present];
let labels: HashSet<&str> = all.iter().map(|p| p.as_str()).collect();
assert_eq!(labels.len(), 3, "each phase needs its own label");
assert!(labels.contains("binding"));
for phase in all {
assert_eq!(
Presence::from_u8(phase.as_u8()),
phase,
"{phase} must survive the atomic round trip the state uses"
);
assert_eq!(phase.to_string(), phase.as_str(), "Display must agree");
}
// Anything outside the enum reads as absent rather than panicking: the
// byte comes out of an AtomicU8 that a torn write could leave at any
// value, and the safe answer there is "not bound".
assert_eq!(Presence::from_u8(99), Presence::Absent);
}
}
+81 -42
View File
@@ -80,6 +80,35 @@ impl PathMtuEntry {
/// address).
pub type PathMtuLookup = Arc<RwLock<HashMap<FipsAddress, PathMtuEntry>>>;
/// The node-global TCP MSS ceiling, shared live with the TUN reader and
/// writer threads.
///
/// Shared rather than copied because the value it is derived from moves at
/// runtime. `Node::transport_mtu()` is the minimum across *bound* transports,
/// and since dynamic interface binding a transport can bind minutes after
/// start or unbind mid-operation — so a narrow interface appearing must
/// tighten the clamp, and its departure must release it. Every other consumer
/// of `transport_mtu()` already reads it live (`show_status`, the snapshot,
/// the session-layer fragmentation check); these two threads captured a `u16`
/// at spawn and were the only place left where the daemon could report one
/// effective MTU and clamp to another.
///
/// A relaxed load per packet, beside the `PathMtuLookup` read that already
/// happens on the same packet — strictly the cheaper of the two. Ordering is
/// irrelevant: this is a clamp, and a packet that reads the previous value
/// during the store is clamped by the ceiling that was correct a microsecond
/// earlier. The per-flow ceiling has always had that property.
/// The IPv6 minimum link MTU (RFC 8200): every compliant path accepts a packet
/// this large, so an MSS derived from it fits anywhere.
///
/// Two callers, and they must not disagree: the cold-flow fallback in
/// [`per_flow_max_mss`], and the seed for the node's [`MssCeiling`] before any
/// transport has bound — which is the same value `Node::transport_mtu()` falls
/// back to when nothing is bound.
pub const IPV6_MIN_MTU: u16 = 1280;
pub type MssCeiling = Arc<std::sync::atomic::AtomicU16>;
/// Compute the effective TCP MSS ceiling for a packet given its peer
/// address bytes (a 16-byte IPv6 destination on outbound, source on
/// inbound). Returns `min(global_max_mss, learned_path_max_mss)` when
@@ -115,7 +144,6 @@ pub(crate) fn per_flow_max_mss(
// RFC 8200 IPv6-minimum MTU (1280) → effective FIPS-encapsulated
// payload (1203) → TCP segment after IPv6+TCP headers (1143).
// Used as the conservative ceiling for empty-lookup destinations.
const IPV6_MIN_MTU: u16 = 1280;
let conservative_max_mss = mss_ceiling(IPV6_MIN_MTU);
let empty_lookup_ceiling = std::cmp::min(global_max_mss, conservative_max_mss);
@@ -435,13 +463,14 @@ impl TunDevice {
/// a channel sender for submitting packets to be written.
///
/// `max_mss` is the global TCP MSS ceiling derived from the local
/// `transport_mtu()` floor. `path_mtu_lookup` is a read-only handle to
/// the per-destination path MTU map populated by discovery; the writer
/// reads it on each inbound SYN-ACK to compute a per-flow ceiling that
/// honors learned narrow paths through the mesh.
/// `transport_mtu()` floor, shared live so a transport that binds or
/// unbinds after start moves it (see [`MssCeiling`]). `path_mtu_lookup`
/// is a read-only handle to the per-destination path MTU map populated by
/// discovery; the writer reads both on each inbound SYN-ACK to compute a
/// per-flow ceiling that honors learned narrow paths through the mesh.
pub fn create_writer(
&self,
max_mss: u16,
max_mss: MssCeiling,
path_mtu_lookup: PathMtuLookup,
) -> Result<(TunWriter, TunTx), TunError> {
let fd = self.device.as_raw_fd();
@@ -517,7 +546,7 @@ pub struct TunWriter {
file: File,
rx: mpsc::Receiver<Vec<u8>>,
name: String,
max_mss: u16,
max_mss: MssCeiling,
path_mtu_lookup: PathMtuLookup,
}
@@ -531,16 +560,23 @@ impl TunWriter {
pub fn run(mut self) {
use super::tcp_mss::clamp_tcp_mss;
debug!(name = %self.name, max_mss = self.max_mss, "TUN writer starting");
debug!(
name = %self.name,
max_mss = self.max_mss.load(std::sync::atomic::Ordering::Relaxed),
"TUN writer starting"
);
for mut packet in self.rx {
// Read per packet, not once: a transport binding or unbinding
// moves the node's egress floor at runtime. See `MssCeiling`.
let global_max_mss = self.max_mss.load(std::sync::atomic::Ordering::Relaxed);
// Per-destination clamp: peer IPv6 source address (bytes 8..24)
// identifies the flow's remote end. If discovery has learned a
// smaller path MTU for that peer, tighten the ceiling.
let effective_max_mss = if packet.len() >= 24 {
per_flow_max_mss(&self.path_mtu_lookup, &packet[8..24], self.max_mss)
per_flow_max_mss(&self.path_mtu_lookup, &packet[8..24], global_max_mss)
} else {
self.max_mss
global_max_mss
};
// Clamp TCP MSS on inbound SYN-ACK packets
if clamp_tcp_mss(&mut packet, effective_max_mss) {
@@ -622,17 +658,17 @@ pub fn run_tun_reader(
our_addr: FipsAddress,
tun_tx: TunTx,
outbound_tx: TunOutboundTx,
transport_mtu: u16,
max_mss: MssCeiling,
path_mtu_lookup: PathMtuLookup,
) {
let (name, mut buf, max_mss) = tun_reader_setup(device.name(), mtu, transport_mtu);
let (name, mut buf) = tun_reader_setup(device.name(), mtu, &max_mss);
loop {
match device.read_packet(&mut buf) {
Ok(n) if n > 0 => {
if !handle_tun_packet(
&mut buf[..n],
max_mss,
max_mss.load(std::sync::atomic::Ordering::Relaxed),
&name,
our_addr,
&tun_tx,
@@ -685,13 +721,13 @@ pub fn run_tun_reader(
our_addr: FipsAddress,
tun_tx: TunTx,
outbound_tx: TunOutboundTx,
transport_mtu: u16,
max_mss: MssCeiling,
path_mtu_lookup: PathMtuLookup,
shutdown_fd: std::os::unix::io::RawFd,
) {
let _shutdown_fd = ShutdownFd(shutdown_fd);
let tun_fd = device.device().as_raw_fd();
let (name, mut buf, max_mss) = tun_reader_setup(device.name(), mtu, transport_mtu);
let (name, mut buf) = tun_reader_setup(device.name(), mtu, &max_mss);
// Set TUN fd to non-blocking so we can use select + read without blocking
// past the point where select returns readable.
@@ -741,7 +777,7 @@ pub fn run_tun_reader(
Ok(n) if n > 0 => {
if !handle_tun_packet(
&mut buf[..n],
max_mss,
max_mss.load(std::sync::atomic::Ordering::Relaxed),
&name,
our_addr,
&tun_tx,
@@ -774,30 +810,25 @@ pub fn run_tun_reader(
// _shutdown_fd closes on drop
}
/// Common setup for TUN reader: allocates buffer, computes max MSS.
fn tun_reader_setup(device_name: &str, mtu: u16, transport_mtu: u16) -> (String, Vec<u8>, u16) {
use super::icmp::effective_ipv6_mtu;
/// Common setup for TUN reader: allocates the buffer and names the device.
///
/// The MSS ceiling is deliberately *not* returned. It is read from the shared
/// [`MssCeiling`] on every packet, because a transport binding or unbinding
/// moves it after this function has run; returning it here is what let the
/// reader clamp to a floor derived from the transports that happened to be
/// bound at spawn.
fn tun_reader_setup(device_name: &str, mtu: u16, max_mss: &MssCeiling) -> (String, Vec<u8>) {
let name = device_name.to_string();
let buf = vec![0u8; mtu as usize + 100];
const IPV6_HEADER: u16 = 40;
const TCP_HEADER: u16 = 20;
let effective_mtu = effective_ipv6_mtu(transport_mtu);
let max_mss = effective_mtu
.saturating_sub(IPV6_HEADER)
.saturating_sub(TCP_HEADER);
debug!(
name = %name,
tun_mtu = mtu,
transport_mtu = transport_mtu,
effective_mtu = effective_mtu,
max_mss = max_mss,
max_mss = max_mss.load(std::sync::atomic::Ordering::Relaxed),
"TUN reader starting"
);
(name, buf, max_mss)
(name, buf)
}
/// Process a single TUN packet. Returns `false` if the reader should exit.
@@ -1076,12 +1107,13 @@ mod windows_tun {
/// packets independently. Returns the writer and a channel sender for
/// submitting packets to be written.
///
/// `max_mss` is the global TCP MSS ceiling. `path_mtu_lookup` is a
/// read-only handle to per-destination path MTU learned via
/// discovery.
/// `max_mss` is the global TCP MSS ceiling, shared live so a transport
/// that binds or unbinds after start moves it (see [`MssCeiling`]).
/// `path_mtu_lookup` is a read-only handle to per-destination path MTU
/// learned via discovery.
pub fn create_writer(
&self,
max_mss: u16,
max_mss: MssCeiling,
path_mtu_lookup: PathMtuLookup,
) -> Result<(TunWriter, TunTx), TunError> {
let (tx, rx) = mpsc::channel();
@@ -1119,7 +1151,7 @@ mod windows_tun {
session: Arc<wintun::Session>,
rx: mpsc::Receiver<Vec<u8>>,
name: String,
max_mss: u16,
max_mss: MssCeiling,
path_mtu_lookup: PathMtuLookup,
}
@@ -1132,14 +1164,21 @@ mod windows_tun {
use super::per_flow_max_mss;
use crate::upper::tcp_mss::clamp_tcp_mss;
debug!(name = %self.name, max_mss = self.max_mss, "TUN writer starting");
debug!(
name = %self.name,
max_mss = self.max_mss.load(std::sync::atomic::Ordering::Relaxed),
"TUN writer starting"
);
for mut packet in self.rx {
// Read per packet, not once: a transport binding or unbinding
// moves the node's egress floor. See `MssCeiling`.
let global_max_mss = self.max_mss.load(std::sync::atomic::Ordering::Relaxed);
// Per-destination clamp (peer source IPv6 = bytes 8..24)
let effective_max_mss = if packet.len() >= 24 {
per_flow_max_mss(&self.path_mtu_lookup, &packet[8..24], self.max_mss)
per_flow_max_mss(&self.path_mtu_lookup, &packet[8..24], global_max_mss)
} else {
self.max_mss
global_max_mss
};
// Clamp TCP MSS on inbound SYN-ACK packets
if clamp_tcp_mss(&mut packet, effective_max_mss) {
@@ -1188,17 +1227,17 @@ mod windows_tun {
our_addr: FipsAddress,
tun_tx: TunTx,
outbound_tx: TunOutboundTx,
transport_mtu: u16,
max_mss: MssCeiling,
path_mtu_lookup: PathMtuLookup,
) {
let (name, mut buf, max_mss) = super::tun_reader_setup(device.name(), mtu, transport_mtu);
let (name, mut buf) = super::tun_reader_setup(device.name(), mtu, &max_mss);
loop {
match device.read_packet(&mut buf) {
Ok(n) if n > 0 => {
if !super::handle_tun_packet(
&mut buf[..n],
max_mss,
max_mss.load(std::sync::atomic::Ordering::Relaxed),
&name,
our_addr,
&tun_tx,
+26
View File
@@ -85,6 +85,15 @@ End-to-end exercise of the production `fips0` nftables baseline at
`packaging/common/fips.nft`, covering the default-deny, conntrack and
drop-in semantics.
### [iface-binding/](iface-binding/) -- Dynamic Interface Binding
Two nodes whose only transports are interface-bound, started before the
interface they name exists. Asserts the boot race (the daemon comes up
`Degraded` and binds when the interface appears, with no restart), the flap
(down/up in both directions), destroy-and-recreate, that an `optional`
interface's absence never moves node health, and that absence is logged once
on the edge rather than once per retry.
### [acl-allowlist/](acl-allowlist/) -- Peer ACL Enforcement
Six nodes with per-node allowlist files mounted at the runtime ACL
@@ -146,6 +155,23 @@ matrix; a divergence fails the run. GitHub
runs the same check as its own `ci-parity` job. `--check-parity` runs it
alone (see [check-ci-parity.sh](check-ci-parity.sh)).
Note that `ci-local.sh` covers the integration suites and the glibc unit
tests. GitHub additionally runs the library tests on macOS, Windows and
**musl** (built for the musl target and run natively), and a `--features
profiling` pass; the musl
leg exists because interface presence is built on `getifaddrs`/`ifa_flags`,
which musl reimplements independently, and OpenWrt is a musl target. A local
green run does not certify those four.
The Linux and musl legs also create an address-less dummy interface
(`fips-probe0`) and pass its name to the tests as
`FIPS_TEST_ADDRLESS_IFACE`. That is the one assumption the interface-binding
mechanism rests on that no ordinary test can reach: loopback has addresses, so
probing it asks whether `getifaddrs` works rather than whether it reports an
interface that has none — which is exactly what `fips-mesh0` and `fips-ap0`
are on OpenWrt. Set the variable by hand to run the assertion locally against
an interface you have created; leave it unset and the assertion does not run.
### Per-run isolation and the `FIPS_CI_RUN_ID` override
Every invocation derives a **run id** and scopes all of its Docker
+127
View File
@@ -0,0 +1,127 @@
# Ethernet rebind under active traffic
#
# The one case dynamic interface binding has no coverage for anywhere:
# traffic crossing an Ethernet link while the interface underneath it goes
# away and comes back.
#
# The other Ethernet scenarios cannot reach it. `ethernet-only` and
# `ethernet-mesh` both run with `traffic.enabled: false`, so no datagram
# crosses an Ethernet link in either — framing, the length field that trims
# NIC minimum-frame padding, and AEAD over Ethernet are all control-plane
# assumptions there. And `link_flaps` cannot produce a rebind whatever it is
# pointed at: it simulates a down link with netem 100% loss, so the interface
# stays IFF_UP and the presence machine never sees an edge.
#
# `node_churn` is what actually moves an interface. Stopping a container
# destroys its network namespace, which deletes every veth in it — and
# deleting one end of a veth deletes its peer — so a *surviving* node watches
# its Ethernet interface disappear outright. On restart the harness recreates
# the pair (`NodeChurnManager._start_node`), and the survivor watches it come
# back. That is a real detach and a real rebind, driven from outside the
# daemon, with iperf3 running across the mesh throughout.
#
# Topology: a 4-node ring, so removing any single node leaves the remaining
# three connected in a line and `protect_connectivity` has something to
# protect.
#
# n01 ---eth--- n02
# | |
# eth eth
# | |
# n04 ---eth--- n03
scenario:
name: "ethernet-churn"
seed: 42
duration_secs: 240
topology:
algorithm: explicit
num_nodes: 4
default_transport: ethernet
params:
adjacency:
- [n01, n02]
- [n02, n03]
- [n03, n04]
- [n04, n01]
# Mild, and deliberately so. The variable under test is the interface going
# away, not the link being bad while it is there; heavy loss here would make a
# traffic shortfall ambiguous between the two.
netem:
enabled: true
default_policy:
delay_ms: [1, 5]
jitter_ms: [0, 1]
loss_pct: [0, 0.5]
# Off on purpose. A netem-simulated down link would add outage that never
# reaches the presence machine, which is the opposite of what this isolates.
link_flaps:
enabled: false
traffic:
enabled: true
max_concurrent: 2
interval_secs: {min: 10, max: 20}
duration_secs: {min: 15, max: 25}
parallel_streams: 2
# The mechanism. One node down at a time, long enough to outlast the
# ten-second bring-up window so the absence is a real one rather than a race,
# and to give traffic time to run against the reduced mesh before it returns.
node_churn:
enabled: true
interval_secs: {min: 45, max: 60}
max_down_nodes: 1
down_duration_secs: {min: 20, max: 35}
protect_connectivity: true
assertions:
# The mesh re-forms after each interface comes back, fully.
#
# Calibrated 2026-09-10 against sixteen runs, eight on master-line code and
# eight on next-line code, on a harness that fails a veth restore it cannot
# complete and waits for the tree to settle before the final snapshot. Twelve
# ran beside the other five CI chaos scenarios and four CPU busy loops; seeds
# 42 (twelve runs), 7 and 1234:
#
# nodes answering 4 in all sixteen
# distinct roots 1 in all sixteen
# nodes parented 3 in all sixteen
# settle 11-22 s
# traffic 4-11 sessions per run, 3.8-11.7 GB
#
# So the thresholds are full convergence: every node answers, one root, three
# parented. A red here is a mesh that did not re-form.
#
# READ BEFORE RETUNING: the 2026-09-02 calibration recorded 2 roots and 2
# parented in all four of its runs, and that was not the daemon. The veth
# restore then failed silently whenever a stopped node's namespace outlived
# the stop, and in every later run that was inspected, the node islanded at
# the snapshot was one whose links the harness had left down. Before widening these, read the red run's runner.log for a
# harness fault; a failed restore now aborts the run rather than reaching
# this assertion.
baseline:
min_nodes_reporting: 4
max_roots: 1
min_nodes_parented: 3
# The point of the scenario. Traffic must actually have crossed Ethernet
# links while interfaces were being taken away underneath it — a green tree
# with zero bytes moved is the failure this catches, and is exactly what a
# control-plane-only assertion would have called a pass.
min_traffic:
min_sessions_ok: 2
# Chaos Ethernet transports are `optional: true` (see config_gen), precisely
# because a neighbour's interface disappearing is the scenario here rather
# than a fault. So absence must stay silent: any ERROR means something other
# than the churn went wrong.
max_errors:
max_total: 0
logging:
rust_log: "info"
output_dir: "./sim-results"
+59
View File
@@ -19,6 +19,7 @@ from dataclasses import dataclass
from .control import snapshot_all_bloom
from .scenario import (
MinTrafficAssertion,
BaselineAssertion,
BloomSendRateAssertion,
CongestionSignalsAssertion,
@@ -442,3 +443,61 @@ def evaluate_min_parent_switches(
f"node win the root election?"
),
)
def _session_bytes(result: dict) -> int:
"""Bytes actually received in one iperf3 session, or 0 if it failed.
iperf3 reports a failed run as a top-level ``error`` string with no
``end`` block, and a killed-but-partial run still carries whatever it
managed. Both are handled by reading the received total and treating
anything missing as zero, so a session only counts when it moved bytes.
"""
if not isinstance(result, dict) or result.get("error"):
return 0
end = result.get("end")
if not isinstance(end, dict):
return 0
summary = end.get("sum_received") or end.get("sum_sent")
if not isinstance(summary, dict):
return 0
value = summary.get("bytes", 0)
return value if isinstance(value, int) and value > 0 else 0
def evaluate_min_traffic(
cfg: MinTrafficAssertion,
results: list[dict],
) -> AssertionOutcome:
"""Floor on iperf3 sessions that actually carried data.
Without this the traffic generator is decoration: the results were
written to disk and never read, so a scenario whose every session
failed still passed on a healthy control plane. A rebind under load is
exactly the case a tree snapshot cannot see.
"""
per_session = [_session_bytes(r) for r in results]
ok = [b for b in per_session if b > 0]
total = sum(ok)
if len(ok) >= cfg.min_sessions_ok and total >= cfg.min_bytes_total:
return AssertionOutcome(
name="min_traffic",
passed=True,
detail=(
f"PASS min_traffic: {len(ok)}/{len(results)} session(s) moved "
f"data (need {cfg.min_sessions_ok}); {total} byte(s) total "
f"(need {cfg.min_bytes_total})"
),
)
return AssertionOutcome(
name="min_traffic",
passed=False,
detail=(
f"FAIL min_traffic: {len(ok)}/{len(results)} session(s) moved data "
f"(need {cfg.min_sessions_ok}); {total} byte(s) total (need "
f"{cfg.min_bytes_total}). A green control plane with no traffic "
f"means the data path did not survive what the scenario did to it."
),
)
+17 -1
View File
@@ -64,7 +64,22 @@ def generate_peers_block(
def _build_ethernet_config(iface: str) -> dict:
"""Build an Ethernet transport config dict for a single interface."""
"""Build an Ethernet transport config dict for a single interface.
``optional: True`` because in this harness a neighbour's interface
disappearing is the scenario, not a fault. ``node_churn`` stops a
container, which destroys its netns and with it both ends of every veth
pair it held (see ``nodes.py``: the veths are recreated on restart), so a
surviving node watches a *required* interface vanish for the 30-90s the
neighbour is down -- once per churn event, on every neighbour. The daemon
reports a required interface absent past its bring-up window at ERROR,
which is correct for a deployment and wrong for a harness that tears the
interface down on purpose; the mesh-wide zero-ERROR ceiling would fail on
injected chaos rather than on a defect.
Absence behaviour itself is asserted in ``testing/iface-binding/``, which
exists for it and drives both policies deliberately.
"""
return {
"interface": iface,
"listen": True,
@@ -72,6 +87,7 @@ def _build_ethernet_config(iface: str) -> dict:
"auto_connect": True,
"accept_connections": True,
"beacon_interval_secs": 10,
"optional": True,
}
+30 -10
View File
@@ -370,22 +370,42 @@ class NetemManager:
# Re-apply Ethernet veth netem
veth_states = self.veth_states.get(container)
if veth_states:
for iface, state in veth_states.items():
cmd = (
f"tc qdisc del dev {iface} root 2>/dev/null || true && "
f"tc qdisc add dev {iface} root netem {state.params.to_tc_args()}"
)
result = docker_exec_quiet(container, cmd, timeout=10)
if result is not None:
log.debug("Re-applied veth netem on %s:%s", container, iface)
else:
log.warning("Failed to re-apply veth netem on %s:%s", container, iface)
for state in veth_states.values():
self._apply_veth(state)
log.info(
"Re-applied veth netem on %s (%d Ethernet peers)",
container,
len(veth_states),
)
# And on each running neighbour's end of those links. The restore
# recreates the whole pair, so the survivor's end is a new interface
# with no qdisc, and without this that direction of every restored
# link ran unshaped for the rest of the run.
for peer_id in sorted(self.topology.nodes[node_id].peers):
if peer_id in self.down_nodes:
continue
if self.topology.transport_for_edge(node_id, peer_id) != "ethernet":
continue
peer_container = self.topology.container_name(peer_id)
state = self.veth_states.get(peer_container, {}).get(
veth_interface_name(peer_id, node_id)
)
if state is not None:
self._apply_veth(state)
def _apply_veth(self, state: VethNetemState):
"""Install a veth end's current netem parameters as its root qdisc."""
cmd = (
f"tc qdisc del dev {state.iface} root 2>/dev/null || true && "
f"tc qdisc add dev {state.iface} root netem {state.params.to_tc_args()}"
)
result = docker_exec_quiet(state.container, cmd, timeout=10)
if result is not None:
log.debug("Re-applied veth netem on %s:%s", state.container, state.iface)
else:
log.warning("Failed to re-apply veth netem on %s:%s", state.container, state.iface)
def mutate(self):
"""Randomly mutate netem params on a fraction of links."""
if not self.config.mutation.policies:
+9 -2
View File
@@ -50,7 +50,10 @@ class NodeManager:
self.rng = rng
self.netem_mgr = netem_mgr
self.veth_mgr = veth_mgr
self.down_nodes = down_nodes or set()
# `is not None`, not `or`: the runner passes its shared set while it
# is still empty, and an empty set is falsy, so `or` replaced it with
# a private one and no other manager ever saw a node go down.
self.down_nodes = down_nodes if down_nodes is not None else set()
self.on_node_restart = on_node_restart
self.node_states: dict[str, NodeState] = {
nid: NodeState(node_id=nid) for nid in topology.nodes
@@ -173,7 +176,11 @@ class NodeManager:
# Re-create veth pairs (container restart destroys netns)
if self.veth_mgr:
time.sleep(1)
self.veth_mgr.setup_node(node_id)
# Churn's own record, not the shared set: netem's safety net also
# adds to that set, on a docker hiccup or a crash churn did not
# cause, and a pair deferred on that basis would never be rebuilt.
stopped = {nid for nid, ns in self.node_states.items() if ns.is_down}
self.veth_mgr.setup_node(node_id, stopped)
# Re-apply netem after a brief delay for the container to initialize
if self.netem_mgr:
+100 -20
View File
@@ -20,6 +20,7 @@ from .assertions import (
evaluate_max_errors,
evaluate_max_parent_switches,
evaluate_min_parent_switches,
evaluate_min_traffic,
evaluate_tree_parents,
)
from .compose import generate_compose
@@ -41,11 +42,21 @@ from .veth import VethManager
log = logging.getLogger(__name__)
# The final snapshot waits for this many identical consecutive tree reads,
# taken this far apart, so the tree must hold still for two intervals.
SETTLE_READS = 3
SETTLE_INTERVAL_SECS = 5
SETTLE_TIMEOUT_SECS = 90
class SimRunner:
def __init__(self, scenario: Scenario):
self.scenario = scenario
# Setup draws only: the topology and the ephemeral node choice, made
# once and in a fixed order. Everything drawn while the simulation
# runs comes from `_stream` instead.
self.rng = random.Random(scenario.seed)
self._streams: dict[str, random.Random] = {}
self.topology: SimTopology | None = None
self.compose_file: str | None = None
# Claimed in _setup; the compose file refers to it as external, so
@@ -316,14 +327,16 @@ class SimRunner:
if s.netem.enabled:
bw = s.bandwidth if s.bandwidth.enabled else None
ig = s.ingress if s.ingress.enabled else None
self.netem_mgr = NetemManager(self.topology, s.netem, self.rng, bandwidth=bw, ingress=ig)
self.netem_mgr = NetemManager(
self.topology, s.netem, self._stream("netem"), bandwidth=bw, ingress=ig
)
self.netem_mgr.down_nodes = self._down_nodes
log.info("Applying initial per-link netem...")
self.netem_mgr.setup_initial()
if s.link_flaps.enabled:
self.link_mgr = LinkManager(
self.topology, s.link_flaps, self.rng, netem_mgr=self.netem_mgr
self.topology, s.link_flaps, self._stream("flaps"), netem_mgr=self.netem_mgr
)
if s.link_swap.enabled:
@@ -332,7 +345,7 @@ class SimRunner:
"link_swap requires netem.enabled (depends on per-link tc state)"
)
self.link_swap_mgr = LinkSwapManager(
self.topology, s.link_swap, self.netem_mgr, self.rng,
self.topology, s.link_swap, self.netem_mgr, self._stream("swap"),
)
if s.assertions.bloom_send_rate is not None:
@@ -342,12 +355,12 @@ class SimRunner:
if s.traffic.enabled:
self.traffic_mgr = TrafficManager(
self.topology, s.traffic, self.rng, down_nodes=self._down_nodes
self.topology, s.traffic, self._stream("traffic"), down_nodes=self._down_nodes
)
if s.node_churn.enabled:
self.node_mgr = NodeManager(
self.topology, s.node_churn, self.rng,
self.topology, s.node_churn, self._stream("churn"),
netem_mgr=self.netem_mgr, down_nodes=self._down_nodes,
veth_mgr=self.veth_mgr,
on_node_restart=self._handle_node_restart,
@@ -355,7 +368,7 @@ class SimRunner:
if s.peer_churn.enabled:
self.peer_churn_mgr = PeerChurnManager(
self.topology, s.peer_churn, self.rng,
self.topology, s.peer_churn, self._stream("peer-churn"),
down_nodes=self._down_nodes,
ephemeral_nodes=self._ephemeral_nodes,
)
@@ -410,11 +423,11 @@ class SimRunner:
log.info("Simulation running for %ds...", duration)
# Schedule first events
next_netem = self._schedule_next(start, s.netem.mutation.interval_secs) if self.netem_mgr else float("inf")
next_flap = self._schedule_next(start, s.link_flaps.interval_secs) if self.link_mgr else float("inf")
next_traffic = self._schedule_next(start, s.traffic.interval_secs) if self.traffic_mgr else float("inf")
next_churn = self._schedule_next(start, s.node_churn.interval_secs) if self.node_mgr else float("inf")
next_peer_churn = self._schedule_next(start, s.peer_churn.interval_secs) if self.peer_churn_mgr else float("inf")
next_netem = self._schedule_next(start, s.netem.mutation.interval_secs, "netem") if self.netem_mgr else float("inf")
next_flap = self._schedule_next(start, s.link_flaps.interval_secs, "flaps") if self.link_mgr else float("inf")
next_traffic = self._schedule_next(start, s.traffic.interval_secs, "traffic") if self.traffic_mgr else float("inf")
next_churn = self._schedule_next(start, s.node_churn.interval_secs, "churn") if self.node_mgr else float("inf")
next_peer_churn = self._schedule_next(start, s.peer_churn.interval_secs, "peer-churn") if self.peer_churn_mgr else float("inf")
# Bloom-send-rate assertion: sample at window_secs before end.
bloom_window_start_at = float("inf")
@@ -451,34 +464,34 @@ class SimRunner:
# Netem mutation
if self.netem_mgr and now >= next_netem:
self.netem_mgr.mutate()
next_netem = self._schedule_next(now, s.netem.mutation.interval_secs)
next_netem = self._schedule_next(now, s.netem.mutation.interval_secs, "netem")
# Link flaps
if self.link_mgr:
if now >= next_flap:
self.link_mgr.maybe_flap()
next_flap = self._schedule_next(now, s.link_flaps.interval_secs)
next_flap = self._schedule_next(now, s.link_flaps.interval_secs, "flaps")
self.link_mgr.restore_expired()
# Traffic generation
if self.traffic_mgr:
if now >= next_traffic:
self.traffic_mgr.maybe_spawn()
next_traffic = self._schedule_next(now, s.traffic.interval_secs)
next_traffic = self._schedule_next(now, s.traffic.interval_secs, "traffic")
self.traffic_mgr.cleanup_expired()
# Node churn
if self.node_mgr:
if now >= next_churn:
self.node_mgr.maybe_kill()
next_churn = self._schedule_next(now, s.node_churn.interval_secs)
next_churn = self._schedule_next(now, s.node_churn.interval_secs, "churn")
self.node_mgr.restore_expired()
# Peer churn (topology mutation)
if self.peer_churn_mgr:
if now >= next_peer_churn:
self.peer_churn_mgr.maybe_churn()
next_peer_churn = self._schedule_next(now, s.peer_churn.interval_secs)
next_peer_churn = self._schedule_next(now, s.peer_churn.interval_secs, "peer-churn")
# Status line
down_links = self.link_mgr.down_count if self.link_mgr else 0
@@ -699,6 +712,7 @@ class SimRunner:
self.node_mgr.restore_all()
# Collect iperf3 throughput results before containers stop
iperf_results: list[dict] = []
if self.traffic_mgr:
iperf_results = self.traffic_mgr.collect_results()
if iperf_results:
@@ -707,7 +721,11 @@ class SimRunner:
json.dump(iperf_results, f, indent=2)
log.info("Saved %d iperf3 results to %s", len(iperf_results), iperf_path)
# Take final tree snapshot while nodes are still running
# Take final tree snapshot while nodes are still running, once the
# tree has stopped moving. A node restored a moment ago is its own
# root until it re-parents, so a snapshot taken straight after the
# restore reads a mesh still converging.
self._settle_tree()
self._take_snapshot("final")
# Collect logs before stopping containers
@@ -758,6 +776,13 @@ class SimRunner:
if err_cfg is not None:
outcome = evaluate_max_errors(err_cfg, result.errors)
self.assertion_outcomes.append(outcome)
# Traffic. Evaluated even when no session completed, because
# "nothing ran" is the failure this exists to catch.
traffic_cfg = self.scenario.assertions.min_traffic
if traffic_cfg is not None:
outcome = evaluate_min_traffic(traffic_cfg, iperf_results)
self.assertion_outcomes.append(outcome)
if outcome.passed:
log.info("%s", outcome.detail)
else:
@@ -838,6 +863,46 @@ class SimRunner:
return result
def _settle_tree(self):
"""Wait until consecutive tree reads agree, or the settle time runs out.
Compares each answering node's root and parent. A fixed delay would
either waste time on a mesh that settled at once or cut off one that
had not. Running out is logged and is not a failure in itself: the
final snapshot is taken anyway, and the assertions judge what it
shows.
A read that no node answered never counts toward agreement. It does
not catch a node that stays its own root for longer than the reads
span, which is a tree that is stable and wrong, and is left to the
assertions.
"""
started = time.time()
previous = None
agreeing = 0
while not self._interrupted:
trees = snapshot_all_trees(self.topology)
shape = {
nid: (data.get("root"), data.get("parent"))
for nid, data in trees.items()
}
if not shape:
agreeing = 0
else:
agreeing = agreeing + 1 if shape == previous else 1
previous = shape
waited = time.time() - started
if agreeing >= SETTLE_READS:
log.info("Tree settled after %.0fs", waited)
return
if waited >= SETTLE_TIMEOUT_SECS:
log.warning(
"Tree still changing after %.0fs; taking the final snapshot anyway",
waited,
)
return
self._sleep(SETTLE_INTERVAL_SECS)
def _take_snapshot(self, label: str):
"""Query all nodes via control socket and save tree/MMP/congestion snapshots."""
if not self.topology:
@@ -874,9 +939,24 @@ class SimRunner:
len(self.topology.nodes),
)
def _schedule_next(self, now: float, interval) -> float:
"""Schedule the next event using a Range interval."""
return now + self.rng.uniform(interval.min, interval.max)
def _schedule_next(self, now: float, interval, kind: str) -> float:
"""Schedule the next event of one kind using a Range interval."""
return now + self._stream(f"{kind}-schedule").uniform(interval.min, interval.max)
def _stream(self, name: str) -> random.Random:
"""Return the random stream for one consumer, derived from the seed.
One stream shared by every manager was drawn in wall-clock order, and
a manager that returns early draws nothing, so host load changed
which node the churn stopped next. Under heavy host load that walked
the stops around the ring. A stream per consumer means one manager's
draws no longer shift another's. It does not make a schedule a
function of the seed alone: a manager's own draws can still depend on
mesh state at the tick, such as which nodes are down.
"""
if name not in self._streams:
self._streams[name] = random.Random(f"{self.scenario.seed}:{name}")
return self._streams[name]
def _sleep(self, seconds: float):
"""Sleep in small increments so SIGINT can break out."""
+58 -6
View File
@@ -323,6 +323,27 @@ class MaxErrorsAssertion:
max_total: int = 0
@dataclass
class MinTrafficAssertion:
"""Floor on how much iperf3 traffic actually completed.
Traffic has always been generated and its results saved to
``iperf3-results.json``, but nothing read them: a scenario could carry
``traffic.enabled: true``, have every single session fail, and still
exit 0 on a green control plane. That gap matters most for the
scenarios where traffic is the point — a datagram crossing an Ethernet
link while its interface rebinds underneath is not observable in the
tree snapshot at all.
``min_sessions_ok`` counts sessions that finished with bytes actually
received. ``min_bytes_total`` is the aggregate floor across them; 0
disables it and leaves the session count as the only gate.
"""
min_sessions_ok: int = 1
min_bytes_total: int = 0
@dataclass
class AssertionsConfig:
"""Optional post-run assertions evaluated against control-socket data."""
@@ -334,6 +355,7 @@ class AssertionsConfig:
congestion_signals: CongestionSignalsAssertion | None = None
tree_parents: TreeParentsAssertion | None = None
baseline: BaselineAssertion | None = None
min_traffic: MinTrafficAssertion | None = None
@dataclass
@@ -414,6 +436,7 @@ _SECTION_KEYS = {
"assertions": {
"bloom_send_rate", "min_parent_switches", "max_parent_switches",
"max_errors", "congestion_signals", "tree_parents", "baseline",
"min_traffic",
},
"logging": {"rust_log", "output_dir"},
}
@@ -422,6 +445,7 @@ _ASSERTION_KEYS = {
"min_parent_switches": {"min_total"},
"max_parent_switches": {"max_total", "node"},
"max_errors": {"max_total"},
"min_traffic": {"min_sessions_ok", "min_bytes_total"},
"congestion_signals": {
"min_nodes_detected", "min_nodes_ce_forwarded", "min_nodes_ce_received",
},
@@ -739,6 +763,27 @@ def load_scenario(path: str) -> Scenario:
f"got {err_total}"
)
s.assertions.max_errors = MaxErrorsAssertion(max_total=err_total)
if "min_traffic" in asrt:
mt = asrt["min_traffic"]
_reject_unknown(
mt, _ASSERTION_KEYS["min_traffic"], "assertions.min_traffic",
)
sessions = mt.get("min_sessions_ok", 1)
if not isinstance(sessions, int) or isinstance(sessions, bool) or sessions < 1:
raise ValueError(
"assertions.min_traffic: min_sessions_ok must be a positive "
f"integer, got {sessions!r} — a floor of zero asserts nothing"
)
min_bytes = mt.get("min_bytes_total", 0)
if not isinstance(min_bytes, int) or isinstance(min_bytes, bool) or min_bytes < 0:
raise ValueError(
"assertions.min_traffic: min_bytes_total must be a "
f"non-negative integer, got {min_bytes!r}"
)
s.assertions.min_traffic = MinTrafficAssertion(
min_sessions_ok=sessions, min_bytes_total=min_bytes,
)
else:
# Default-on. See MaxErrorsAssertion for why this one assertion is
# applied without being asked for: it is the floor on what a green
@@ -835,14 +880,21 @@ def load_scenario(path: str) -> Scenario:
+ ", ".join(sorted(_ASSERTION_KEYS["baseline"]))
+ "; a block with none asserts nothing"
)
if vals.get("min_nodes_parented", 0) > 0 or vals.get("max_roots"):
if vals.get("min_nodes_parented", 0) > 0:
# Every root is its own parent, so a mesh with R roots has at most
# n - R nodes parented. Checked against the roots ceiling rather
# than against one root: a floor above n - max_roots reds a run
# that sits exactly at the ceiling the file itself allows, so the
# two thresholds contradict each other there.
n = s.topology.num_nodes
if vals.get("min_nodes_parented", 0) > n - 1:
roots = vals.get("max_roots", 1)
floor = vals["min_nodes_parented"]
if floor > n - roots:
raise ValueError(
f"assertions.baseline.min_nodes_parented: "
f"{vals['min_nodes_parented']} exceeds {n - 1}, the most a "
f"{n}-node mesh can reach — the root is its own parent, so "
f"this could never pass"
f"assertions.baseline.min_nodes_parented: {floor} exceeds "
f"{n - roots}, the most a {n}-node mesh with {roots} "
f"root(s) can reach — each root is its own parent, so a "
f"run at the max_roots ceiling could never pass"
)
s.assertions.baseline = BaselineAssertion(**vals)
+4 -1
View File
@@ -44,7 +44,10 @@ class TrafficManager:
self.topology = topology
self.config = config
self.rng = rng
self.down_nodes = down_nodes or set()
# `is not None`, not `or`: the runner passes its shared set while it
# is still empty, and an empty set is falsy, so `or` replaced it with
# a private one and no other manager ever saw a node go down.
self.down_nodes = down_nodes if down_nodes is not None else set()
self.npub_cache = npub_cache or {}
self.active_sessions: list[TrafficSession] = []
self.completed_results: list[dict] = []
+177 -52
View File
@@ -34,12 +34,36 @@ from __future__ import annotations
import logging
import subprocess
import time
from .docker_exec import docker_exec_quiet
from .docker_exec import DockerExecError, docker_exec, docker_exec_quiet
from .topology import SimTopology, veth_interface_name
log = logging.getLogger(__name__)
# How long a freshly raised veth end may take to report operstate up. The
# kernel publishes the carrier change through linkwatch, which is deferred, so
# a read straight after `ip link set up` can still see the old state. Long
# enough for at least two reads under host load, since one slow `docker exec`
# must not abort a run whose link is up.
OPERSTATE_WAIT_SECS = 15
# Timeout for each `docker exec` a restore makes. The run blocks on these, and
# a timeout aborts it, so this errs long: a loaded host is the normal case.
EXEC_TIMEOUT_SECS = 30
class VethSetupError(RuntimeError):
"""A veth pair could not be created, placed, renamed or raised.
Raised rather than logged. A pair that fails part way leaves a ring link
down while the run carries on, and what reaches the verdict is then a tree
that did not converge: a daemon failure in every respect a reader can see.
Letting this propagate aborts the run, so the fault is reported as the
harness's own.
"""
class VethManager:
"""Manages veth pairs for Ethernet-transport edges."""
@@ -102,12 +126,13 @@ class VethManager:
len(self._host_pairs),
)
def setup_node(self, node_id: str):
def setup_node(self, node_id: str, down_nodes: set[str] | None = None):
"""Re-create veth endpoints for a single node after container restart.
When a container restarts (node churn), its network namespace is
destroyed. We re-create the veth pairs for all Ethernet edges
involving this node.
involving this node. ``down_nodes`` names the neighbours churn has
stopped, whose pairs are left for their own restart.
"""
image = self._get_image()
for a, b in self.topology.ethernet_edges():
@@ -117,7 +142,7 @@ class VethManager:
host_a = self.topology.veth_host_name(a, b, "a")
_run_host(["ip", "link", "delete", host_a], image, check=False)
# Re-create
self._create_veth_pair(a, b, image)
self._create_veth_pair(a, b, image, down_nodes or set())
def teardown_all(self):
"""Clean up all veth pairs."""
@@ -126,18 +151,40 @@ class VethManager:
_run_host(["ip", "link", "delete", host_a], image, check=False)
self._host_pairs.clear()
def _create_veth_pair(self, node_a: str, node_b: str, image: str):
"""Create a single veth pair between two containers."""
def _create_veth_pair(
self, node_a: str, node_b: str, image: str, down_nodes: set[str] | None = None
):
"""Create a single veth pair between two containers.
``down_nodes`` is None at first setup, when every container must be
running. On a restore it names the nodes churn has stopped.
"""
container_a = self.topology.container_name(node_a)
container_b = self.topology.container_name(node_b)
# Get container PIDs
pid_a = _get_container_pid(container_a)
pid_b = _get_container_pid(container_b)
pid_a = _container_pid(container_a)
pid_b = _container_pid(container_b)
if pid_a is None or pid_b is None:
log.warning(
"Cannot create veth %s--%s: container PID not found", node_a, node_b
)
stopped = [n for n, pid in ((node_a, pid_a), (node_b, pid_b)) if pid is None]
names = ", ".join(stopped)
if down_nodes is None:
raise VethSetupError(
f"veth {node_a}--{node_b}: {names} not running at setup"
)
if all(n in down_nodes for n in stopped):
# Stopped by churn: it has no namespace to join, and its own
# restart recreates this pair.
log.info("Veth %s--%s deferred: %s down", node_a, node_b, names)
else:
# Not raised: a container that exited on its own is a daemon
# failure, and aborting here would report it as the harness's.
# Nothing recreates this pair, so say so loudly.
log.warning(
"Veth %s--%s not recreated: %s not running, and churn did "
"not stop all of them; the link stays absent for the rest "
"of the run",
node_a, node_b, names,
)
return
# Generate names
@@ -146,68 +193,146 @@ class VethManager:
final_a = veth_interface_name(node_a, node_b)
final_b = veth_interface_name(node_b, node_a)
# Clear both final names and both temporary names out of the
# containers first. A stopped node's network namespace can outlive
# the stop by minutes, and while it does the survivor still holds its
# old `ve-X-Y`, whose peer sits in that namespace; the rename below
# then fails with "File exists". Deleting the survivor's end removes
# its peer too. A temporary name is left behind only by a restore
# that failed after the move, and it blocks the next move the same
# way.
_purge_links(container_a, [final_a, host_a])
_purge_links(container_b, [final_b, host_b])
# Clean up a stale pair left by this scenario. The token makes the
# name unique to this run, so a pair orphaned by an earlier run is
# no longer reclaimed here — `ci-cleanup.sh` reaps those instead.
_run_host(["ip", "link", "delete", host_a], image, check=False)
# Create veth pair on host
ok = _run_host([
"ip", "link", "add", host_a, "type", "veth", "peer", "name", host_b,
], image)
if not ok:
log.warning("Failed to create veth pair %s/%s", host_a, host_b)
return
# Move into container namespaces
_run_host(["ip", "link", "set", host_a, "netns", str(pid_a)], image)
_run_host(["ip", "link", "set", host_b, "netns", str(pid_b)], image)
# Rename and bring up inside containers
docker_exec_quiet(
container_a,
f"ip link set {host_a} name {final_a} && ip link set {final_a} up",
timeout=10,
)
docker_exec_quiet(
container_b,
f"ip link set {host_b} name {final_b} && ip link set {final_b} up",
timeout=10,
_require_host(
["ip", "link", "add", host_a, "type", "veth", "peer", "name", host_b],
image,
)
_require_host(["ip", "link", "set", host_a, "netns", str(pid_a)], image)
_require_host(["ip", "link", "set", host_b, "netns", str(pid_b)], image)
# Query MAC addresses
_raise_link(container_a, host_a, final_a)
_raise_link(container_b, host_b, final_b)
_await_up(container_a, final_a)
_await_up(container_b, final_b)
# Read only after both ends are proven renamed and up. Before, a
# failed rename left the old interface under the final name, and this
# read its MAC and reported the restore as a success.
mac_a = _get_mac_in_container(container_a, final_a)
mac_b = _get_mac_in_container(container_b, final_b)
if mac_a:
self.topology.nodes[node_a].ethernet_macs[node_b] = mac_a
if mac_b:
self.topology.nodes[node_b].ethernet_macs[node_a] = mac_b
if not mac_a or not mac_b:
raise VethSetupError(
f"veth {node_a}--{node_b} is up but its MAC could not be read "
f"({final_a}: {mac_a or '?'}, {final_b}: {mac_b or '?'})"
)
self.topology.nodes[node_a].ethernet_macs[node_b] = mac_a
self.topology.nodes[node_b].ethernet_macs[node_a] = mac_b
self._host_pairs.append((node_a, node_b, host_a, host_b))
log.info(
"Veth %s(%s) -- %s(%s) MAC: %s / %s",
node_a, final_a, node_b, final_b,
mac_a or "?", mac_b or "?",
node_a, final_a, node_b, final_b, mac_a, mac_b,
)
def _get_container_pid(container: str) -> int | None:
"""Get the PID of a running Docker container."""
def _in_container(container: str, cmd: str, what: str, timeout: int = EXEC_TIMEOUT_SECS) -> str:
"""Run a command inside a container, raising `VethSetupError` on failure."""
try:
return docker_exec(container, cmd, timeout=timeout)
except (DockerExecError, subprocess.TimeoutExpired) as e:
raise VethSetupError(f"{what} in {container} failed: {e}") from e
def _purge_links(container: str, names: list[str]):
"""Delete each named interface in a container if it exists, in one exec.
An absent name is the ordinary case and is skipped. A delete that fails
raises, since the rename that follows would then fail on the name too,
unless the name is gone by then: a lingering namespace can be reaped
between the check and the delete, taking the survivor's end with it.
"""
script = "; ".join(
f"if ip link show {name} >/dev/null 2>&1; then "
f"ip link delete {name} || ! ip link show {name} >/dev/null 2>&1 || exit 1; fi"
for name in names
)
_in_container(container, script, f"deleting stale {', '.join(names)}")
def _require_host(cmd: list[str], image: str):
"""Run a host-namespace ``ip`` command, raising `VethSetupError` on failure."""
if not _run_host(cmd, image):
raise VethSetupError(f"host command failed: {' '.join(cmd)}")
def _raise_link(container: str, temp: str, final: str):
"""Rename a moved veth end to its final name and set it up."""
_in_container(
container,
f"ip link set {temp} name {final} && ip link set {final} up",
f"renaming {temp} to {final}",
)
def _await_up(container: str, iface: str):
"""Wait for an interface's operstate to read ``up``, else raise.
A veth end reports up only once both ends are up, so this proves the
pair is joined as well as that this end was raised.
"""
deadline = time.monotonic() + OPERSTATE_WAIT_SECS
state = None
reads = 0
while True:
state = docker_exec_quiet(
container, f"cat /sys/class/net/{iface}/operstate", timeout=EXEC_TIMEOUT_SECS
)
state = state.strip() if state is not None else None
reads += 1
if state == "up":
return
if reads >= 2 and time.monotonic() >= deadline:
raise VethSetupError(
f"{iface} in {container} reads operstate {state or '?'}, "
f"not up, {OPERSTATE_WAIT_SECS}s after it was raised"
)
time.sleep(0.2)
def _container_pid(container: str) -> int | None:
"""Return a container's PID, or None when docker reports it not running.
Raises `VethSetupError` when docker cannot be asked. A timed-out or failed
inspect of a running container used to read as "not running", so under
host load a live link was skipped and never recreated.
"""
try:
result = subprocess.run(
["docker", "inspect", "-f", "{{.State.Pid}}", container],
capture_output=True,
text=True,
timeout=10,
timeout=EXEC_TIMEOUT_SECS,
)
if result.returncode == 0:
pid = int(result.stdout.strip())
return pid if pid > 0 else None
except (subprocess.TimeoutExpired, ValueError):
pass
return None
except subprocess.TimeoutExpired as e:
raise VethSetupError(f"docker inspect {container} timed out") from e
if result.returncode != 0:
raise VethSetupError(
f"docker inspect {container} failed: {result.stderr.strip()}"
)
try:
pid = int(result.stdout.strip())
except ValueError as e:
raise VethSetupError(
f"docker inspect {container} returned no PID: {result.stdout.strip()!r}"
) from e
return pid if pid > 0 else None
def _get_mac_in_container(container: str, iface: str) -> str | None:
@@ -215,7 +340,7 @@ def _get_mac_in_container(container: str, iface: str) -> str | None:
result = docker_exec_quiet(
container,
f"cat /sys/class/net/{iface}/address",
timeout=5,
timeout=EXEC_TIMEOUT_SECS,
)
if result is not None:
return result.strip()
+13 -7
View File
@@ -204,13 +204,19 @@ VETH_NODE_ID='(0[0-9]|[1-9][0-9]+)'
# from the simulation itself rather than repeated here, so widening it cannot
# leave this matching the old width. Empty output means "reap nothing".
#
# Two producers, two shapes. The chaos simulation makes vh{token}{NN}{MM}{a,b};
# the NAT lab (nat/scripts/setup-topology.sh) makes vn{a,b}{token}{0,1}, using
# the RUN-wide suffix rather than any chaos scenario's. Widening this regex is
# Three producers, three shapes. The chaos simulation makes
# vh{token}{NN}{MM}{a,b}; the NAT lab (nat/scripts/setup-topology.sh) makes
# vn{a,b}{token}{0,1}; the interface-binding suite
# (iface-binding/test.sh) makes vhifb{token}{a,b,c,d}. The last two use the
# RUN-wide suffix rather than any chaos scenario's. Widening this regex is
# only half the fix: the token set is derived separately below, so a suffix
# list carrying no NAT suffix leaves the NAT half matching nothing while
# list carrying no run-wide suffix leaves those halves matching nothing while
# looking correct. ci-local.sh's ci_teardown therefore appends the run-wide
# suffix to --veth-suffixes.
#
# vhifb is listed before the bare vh shape in each alternation only for
# readability; the shapes are anchored and cannot overlap, because a node id
# is digits and `ifb` is not.
veth_pattern() {
# vh{token}{NN}{MM}{a,b}: the token is 4 hex or wholly absent — never a
# part of one — and the two node ids follow. Anchored and shaped this
@@ -219,7 +225,7 @@ veth_pattern() {
# single run, no missing or empty suffix list may widen this back out to
# every run.
if [[ -z "$RUN_ID" && -z "$VETH_SUFFIXES" ]]; then
printf '^vh([0-9a-f]{4})?%s%s[ab]$|^vn[ab]([0-9a-f]{4})?[01]$' \
printf '^vhifb([0-9a-f]{4})?[a-d]$|^vh([0-9a-f]{4})?%s%s[ab]$|^vn[ab]([0-9a-f]{4})?[01]$' \
"$VETH_NODE_ID" "$VETH_NODE_ID"
return 0
fi
@@ -253,8 +259,8 @@ veth_pattern() {
veth_warn "no interface tokens derived from: ${sfx[*]}"
return 0
fi
printf '^vh(%s)%s%s[ab]$|^vn[ab](%s)[01]$' \
"$alt" "$VETH_NODE_ID" "$VETH_NODE_ID" "$alt"
printf '^vhifb(%s)[a-d]$|^vh(%s)%s%s[ab]$|^vn[ab](%s)[01]$' \
"$alt" "$alt" "$VETH_NODE_ID" "$VETH_NODE_ID" "$alt"
}
# ip(8) in a privileged --net=host container, matching how the simulation
+81 -4
View File
@@ -30,7 +30,8 @@
# firewall, nat-cone, nat-symmetric,
# nat-lan, nostr-publish-consume, stun-faults,
# chaos-churn-mixed-10, chaos-ethernet-mesh,
# chaos-ethernet-only, chaos-tcp-mesh, chaos-congestion-stress,
# chaos-ethernet-only, chaos-ethernet-churn, chaos-tcp-mesh,
# chaos-congestion-stress,
# sidecar, dns-resolver, deb-install, medium-change
#
# Opt-in (require --with-tor; depend on live Tor network):
@@ -164,6 +165,7 @@ CHAOS_SUITES=(
"churn-mixed-10 churn-mixed --nodes 10 --duration 120"
"ethernet-mesh ethernet-mesh"
"ethernet-only ethernet-only"
"ethernet-churn ethernet-churn"
"tcp-mesh tcp-mesh"
"congestion-stress congestion-stress"
)
@@ -208,6 +210,7 @@ CHAOS_SUITES=(
GATEWAY_SUITES=(gateway)
SIDECAR_SUITES=(sidecar)
FIREWALL_SUITES=(firewall)
IFACE_BINDING_SUITES=(iface-binding)
NAT_SUITES=(cone symmetric lan)
NOSTR_RELAY_SUITES=(nostr-publish-consume)
STUN_FAULTS_SUITES=(stun-faults)
@@ -247,6 +250,9 @@ list_suites() {
echo " Firewall baseline:"
for s in "${FIREWALL_SUITES[@]}"; do echo " $s"; done
echo ""
echo " Dynamic interface binding:"
for s in "${IFACE_BINDING_SUITES[@]}"; do echo " $s"; done
echo ""
echo " NAT scenarios:"
for s in "${NAT_SUITES[@]}"; do echo " nat-$s"; done
echo ""
@@ -669,6 +675,47 @@ run_static() {
record "static-$topology" $rc
}
# Lines kept from each node's log when a red scenario's results are printed.
CHAOS_DUMP_LINES=300
# Print a red chaos scenario's results into the run log.
#
# A CI worker that runs each job in a worktree it deletes afterwards, under a
# private /tmp, loses the results directory when the run ends, and the log is
# the only record that survives. Without this such a red cannot be traced node
# by node. The runner's own log is not repeated here: it already reached the run
# log as the scenario's output. Each node log is capped, so the total grows with
# node count; the largest CI scenario has ten nodes.
chaos_dump() {
local name="$1" base="$2" dir f
dir="$(find "$base" -mindepth 1 -maxdepth 1 -type d 2>/dev/null | sort | tail -n 1)"
if [[ -z "$dir" ]]; then
echo "[chaos/$name] no results directory under $base"
return 0
fi
echo "===== chaos-$name results from $dir ====="
for f in status.txt assertions.txt; do
[[ -f "$dir/$f" ]] || continue
echo "----- $f -----"
cat "$dir/$f"
done
if [[ -f "$dir/tree-snapshot-final.json" ]]; then
echo "----- final tree: node, root, parent -----"
python3 -c '
import json, sys
for node, tree in sorted(json.load(open(sys.argv[1])).items()):
print(node, tree.get("root"), tree.get("parent"))
' "$dir/tree-snapshot-final.json" || true
fi
for f in "$dir"/fips-node-*.log; do
[[ -f "$f" ]] || continue
echo "----- $(basename "$f"), last $CHAOS_DUMP_LINES lines -----"
tail -n "$CHAOS_DUMP_LINES" "$f"
done
echo "===== end of chaos-$name results ====="
return 0
}
# Run a chaos scenario
run_chaos() {
local name="$1"
@@ -686,11 +733,17 @@ run_chaos() {
suffix="$(ci_chaos_suffix "$name")"
local -x FIPS_CI_NAME_SUFFIX="$suffix"
# The same scoping for the results, so a red can find the directory this
# run wrote rather than guess among earlier runs' timestamps.
local results="$SCRIPT_DIR/chaos/sim-results/ci$suffix"
local -x FIPS_SIM_OUTPUT="$results"
info "[chaos/$name] Running simulation"
if bash testing/chaos/scripts/chaos.sh "$@" 2>&1; then
rc=0
else
rc=1
chaos_dump "$name" "$results"
fi
record "chaos-$name" $rc
@@ -796,6 +849,22 @@ run_sidecar() {
record "sidecar" $rc
}
# Run the dynamic interface binding integration test.
#
# Creates a veth pair from the host namespace after the daemons are up, so it
# needs the same privileged ip(8) helper the chaos harness uses. Scoped by
# COMPOSE_PROJECT_NAME like every other suite; the veth names carry the run
# suffix themselves (see testing/iface-binding/test.sh).
run_iface_binding() {
export COMPOSE_PROJECT_NAME="$(ci_project iface_binding)"
info "[iface-binding] Running integration test"
if bash testing/iface-binding/test.sh --skip-build 2>&1; then
record "iface-binding" 0
else
record "iface-binding" 1
fi
}
# Run firewall baseline integration test
run_firewall() {
export COMPOSE_PROJECT_NAME="$(ci_project firewall)"
@@ -1262,6 +1331,9 @@ run_integration() {
# Firewall baseline
run_firewall
# Dynamic interface binding
run_iface_binding
# NAT scenarios (sequential — each owns its compose project)
for scenario in "${NAT_SUITES[@]}"; do
run_nat "$scenario"
@@ -1334,9 +1406,12 @@ run_integration() {
record "chaos-$scenario" 0
else
record "chaos-$scenario" 1
# Show tail of failure log
echo "--- chaos-$scenario output (last 20 lines) ---"
tail -20 "$logfile" 2>/dev/null || true
# The whole log, not its tail: the child printed the results
# directory into it, and this is the only place that reaches
# the run log. The sed keeps the last frame of each line the
# progress display redraws with carriage returns.
echo "--- chaos-$scenario output ---"
sed 's/.*\r//' "$logfile" 2>/dev/null || true
echo "---"
fi
rm -f "$logfile"
@@ -1375,6 +1450,8 @@ run_suite() {
run_gateway ;;
firewall)
run_firewall ;;
iface-binding)
run_iface_binding ;;
nat-cone|nat-symmetric|nat-lan)
run_nat "${suite#nat-}" ;;
nostr-publish-consume)
+57
View File
@@ -0,0 +1,57 @@
# Dynamic Interface Binding
Two FIPS daemons whose **only** transports are bound to network interfaces,
exercised against a veth pair the harness creates, downs, deletes and recreates
underneath them while they run.
```
node-a node-b
lab ve-lab0 required ── veth ── ve-lab0 required
dock fips-dock0 optional fips-dock0 optional
```
`ve-lab0` does not exist when the daemons start. `fips-dock0` never exists at
all, on any host, ever — it is the negative control for `optional: true`.
## What it asserts
| | Behavior |
| - | -------- |
| (a) | A daemon whose only interface is missing **starts**, reports the transport `absent`, and reports `Degraded` — it does not exit on `NoTransports`, and it does not skip the transport for the life of the process |
| (b) | The interface appears; both daemons bind it with no restart, `Degraded` clears, and they discover and peer over it |
| (c) | The interface goes down and comes back; presence and health follow it in **both** directions, and the rebind is counted |
| (d) | The interface is deleted outright and recreated; both daemons rebind and re-peer — the case the old ENXIO beacon-socket reopen half-covered |
| (e) | An `optional` interface that never appears logs at `info` and never moves node health |
| | Absence is logged **once on the edge**, not once per retry |
Health is asserted through `fipsctl show status` (`state`), presence through
`fipsctl show transports` (the per-transport `interface` block: `presence`,
`policy`, `binds`, `since_secs`).
## Running
```sh
./test.sh # builds the image first
./test.sh --skip-build # reuse an existing image
./test.sh --keep-up # leave the containers running for inspection
```
Via the local CI runner:
```sh
./testing/ci-local.sh --only iface-binding
```
## Notes
The containers run under `FIPS_TEST_MODE=default`, **not** `chaos`. The chaos
entrypoint waits up to 30 s for every configured Ethernet interface before
starting the daemon — which is exactly the workaround this mechanism retires.
The daemon has to do its own waiting here or the suite proves nothing.
Every `ip link` operation on the host network stack runs inside a short-lived
privileged container sharing the host network and PID namespaces, for the
reason [chaos/sim/veth.py](../chaos/sim/veth.py) documents: on macOS the
containers live in the Docker VM, so ip(8) run on the macOS host could never
reach them, while on Linux the shared namespaces make it identical to running
ip(8) directly.
+80
View File
@@ -0,0 +1,80 @@
networks:
# Management bridge only. The FIPS transport under test is raw Ethernet on a
# veth pair the harness creates *after* the daemons are already running —
# that is the whole point of the suite — so no FIPS traffic crosses this
# network. No subnet is requested, so two concurrent runs cannot collide on
# one address range.
#
# The compose project name is still fixed, so two runs that do not set
# COMPOSE_PROJECT_NAME share a project and the second `up` recreates the
# first's containers. The local CI runner scopes it externally
# (run_iface_binding in ci-local.sh); a bare hand run does not.
ifb-net:
driver: bridge
labels:
- "com.corganlabs.fips-ci=1"
x-fips-common: &fips-common
build:
# The harness scopes its build context per run and passes it here; the
# shared directory is the hand-run default. Compose resolves a relative
# value against THIS file's directory, so the harness must export an
# absolute path.
context: ${FIPS_BUILD_CONTEXT:-../docker}
image: ${FIPS_TEST_IMAGE:-fips-test:latest}
entrypoint: ["/usr/local/bin/entrypoint.sh"]
cap_add:
- NET_ADMIN
- NET_RAW
restart: "no"
environment:
# `default`, deliberately — NOT `chaos`. The chaos entrypoint waits up to
# 30 s for every configured Ethernet interface to appear before it starts
# the daemon, which is precisely the workaround this mechanism retires. The
# daemon must do its own waiting here or the suite proves nothing.
- FIPS_TEST_MODE=default
- RUST_LOG=info,fips::transport::ethernet=debug,fips::node=debug
networks:
- ifb-net
services:
node-a:
<<: *fips-common
container_name: fips-ifb-node-a${FIPS_CI_NAME_SUFFIX:-}
hostname: host-a
volumes:
- ../docker/resolv.conf:/etc/resolv.conf:ro
- ./generated-configs${FIPS_CI_NAME_SUFFIX:-}/node-a/fips.yaml:/etc/fips/fips.yaml:ro
- ./generated-configs${FIPS_CI_NAME_SUFFIX:-}/node-a/fips.key:/etc/fips/fips.key:ro
node-b:
<<: *fips-common
container_name: fips-ifb-node-b${FIPS_CI_NAME_SUFFIX:-}
hostname: host-b
volumes:
- ../docker/resolv.conf:/etc/resolv.conf:ro
- ./generated-configs${FIPS_CI_NAME_SUFFIX:-}/node-b/fips.yaml:/etc/fips/fips.yaml:ro
- ./generated-configs${FIPS_CI_NAME_SUFFIX:-}/node-b/fips.key:/etc/fips/fips.key:ro
# The clean-start case. Its interface exists before its daemon does, which is
# the ordinary state of a booted router and the one ordering node-a and
# node-b cannot produce: their interface is created after they are already
# running, so they can only ever bind through the binder loop.
#
# The gate is what buys that ordering. The harness needs a running container
# to have a netns to move a veth into, but the daemon must not start until
# after the move — so the container comes up, parks on this file, and the
# harness releases it once the interface is in place.
node-c:
<<: *fips-common
container_name: fips-ifb-node-c${FIPS_CI_NAME_SUFFIX:-}
hostname: host-c
entrypoint: ["/bin/sh", "-c"]
command:
- |
while [ ! -e /tmp/fips-go ]; do sleep 0.2; done
exec /usr/local/bin/entrypoint.sh
volumes:
- ../docker/resolv.conf:/etc/resolv.conf:ro
- ./generated-configs${FIPS_CI_NAME_SUFFIX:-}/node-c/fips.yaml:/etc/fips/fips.yaml:ro
- ./generated-configs${FIPS_CI_NAME_SUFFIX:-}/node-c/fips.key:/etc/fips/fips.key:ro
+136
View File
@@ -0,0 +1,136 @@
#!/bin/bash
# Generate fixtures for the dynamic interface binding integration test.
#
# Two FIPS nodes, each with two Ethernet transports and nothing else:
#
# lab ve-lab0 required — does not exist when the daemon starts; the
# harness creates the veth pair afterwards
# dock fips-dock0 optional — never exists, on any host, ever
#
# Plus a third node whose single required interface exists *before* its daemon
# starts — the one ordering the other two cannot produce, and the one that
# `start_async`'s inline bind takes. See node-c in test.sh case (f).
#
# There is deliberately no UDP transport. A node whose only transports are
# interface-bound is the case that used to be unrecoverable: every transport
# skipped at start, nothing retried, and the node up and deaf.
set -euo pipefail
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
# Scoped by the per-run suffix because this directory is wiped and rewritten
# below: two runs sharing one output directory would delete each other's
# fixtures out from under running containers.
GENERATED_DIR="$SCRIPT_DIR/generated-configs${FIPS_CI_NAME_SUFFIX:-}"
# Deterministic test identities (mirrors the firewall/acl-allowlist style).
KEY_A="0102030405060708090a0b0c0d0e0f101112131415161718191a1b1c1d1e1f20"
KEY_B="b102030405060708090a0b0c0d0e0f101112131415161718191a1b1c1d1e1fb0"
KEY_C="c102030405060708090a0b0c0d0e0f101112131415161718191a1b1c1d1e1fc0"
write_file() {
local path="$1"
mkdir -p "$(dirname "$path")"
cat > "$path"
}
# Peers are found by beacon, not configured: a MAC address that does not exist
# until the harness creates the veth pair cannot be written into a config file
# ahead of time. Discovery over the late-bound interface is part of what the
# suite asserts.
write_node_config() {
write_file "$GENERATED_DIR/$1/fips.yaml" <<EOF
node:
identity:
persistent: true
# No TUN and no DNS: this suite is about transport binding, and every extra
# child is another way for a failure to be misattributed.
tun:
enabled: false
dns:
enabled: false
transports:
ethernet:
lab:
interface: "$LAB_IFACE"
listen: true
announce: true
auto_connect: true
accept_connections: true
# Fast beacons so peering after a late bind is observed in seconds
# rather than in the 30 s production default.
beacon_interval_secs: 2
dock:
interface: "$DOCK_IFACE"
# Absence is normal for this one, so it must never move node health.
optional: true
listen: true
announce: true
auto_connect: true
accept_connections: true
beacon_interval_secs: 2
peers: []
EOF
}
# node-c: one required interface, present at daemon start.
#
# The other two nodes can only ever reach `Present` through the binder loop,
# because their interface does not exist until the harness makes it. That left
# the inline bind in `start_async` — the ordinary case on a booted router —
# with no coverage at all, which is exactly where the churn guard went unseeded
# and the first detach stopped reaching node health.
write_boot_node_config() {
write_file "$GENERATED_DIR/node-c/fips.yaml" <<EOF
node:
identity:
persistent: true
tun:
enabled: false
dns:
enabled: false
transports:
ethernet:
boot:
interface: "$BOOT_IFACE"
listen: true
announce: true
auto_connect: true
accept_connections: true
beacon_interval_secs: 2
peers: []
EOF
}
LAB_IFACE="${LAB_IFACE:-ve-lab0}"
DOCK_IFACE="${DOCK_IFACE:-fips-dock0}"
BOOT_IFACE="${BOOT_IFACE:-ve-boot0}"
echo "Generating interface-binding fixtures..."
rm -rf "$GENERATED_DIR"
write_node_config node-a
write_file "$GENERATED_DIR/node-a/fips.key" <<EOF
$KEY_A
EOF
write_node_config node-b
write_file "$GENERATED_DIR/node-b/fips.key" <<EOF
$KEY_B
EOF
write_boot_node_config
write_file "$GENERATED_DIR/node-c/fips.key" <<EOF
$KEY_C
EOF
echo "Interface-binding fixtures written to $GENERATED_DIR"
+632
View File
@@ -0,0 +1,632 @@
#!/bin/bash
# Integration test for dynamic interface binding.
#
# Asserts the five behaviors the presence machine exists to provide, against
# real daemons and a real veth pair:
#
# (a) boot race — a daemon whose interface does not exist yet starts,
# reports the transport ABSENT, and reports Degraded
# (b) late attach — the interface appears; the daemon binds it with no
# restart, clears Degraded, and peers over it
# (c) flap — the interface goes down and comes back; presence and
# health follow it in BOTH directions
# (d) destroy/recreate — the interface is deleted outright and recreated; the
# daemon rebinds (the case the old ENXIO beacon hack
# half-covered)
# (e) optional — an interface that never appears logs at info and
# never moves node health
#
# plus the log-hygiene property the design is explicit about: absence is logged
# once on the edge, never once per retry.
#
# The daemons run under FIPS_TEST_MODE=default, NOT chaos: the chaos entrypoint
# waits for Ethernet interfaces before starting the daemon, which is exactly
# the workaround being retired. The daemon must wait for itself here.
#
# Usage: ./test.sh [--skip-build] [--keep-up]
set -euo pipefail
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
TESTING_DIR="$(cd "$SCRIPT_DIR/.." && pwd)"
COMPOSE_FILE="$SCRIPT_DIR/docker-compose.yml"
NODE_A="fips-ifb-node-a${FIPS_CI_NAME_SUFFIX:-}"
NODE_B="fips-ifb-node-b${FIPS_CI_NAME_SUFFIX:-}"
NODE_C="fips-ifb-node-c${FIPS_CI_NAME_SUFFIX:-}"
# The interface each node binds. Same name on both sides: they are in separate
# network namespaces, and using one name keeps the fixtures identical.
LAB_IFACE="ve-lab0"
# The interface that never exists. `optional: true` in both configs.
DOCK_IFACE="fips-dock0"
# node-c's interface, present before its daemon starts. See case (f).
BOOT_IFACE="ve-boot0"
# Host-side veth names, scoped per run: these live in the host (or Docker VM)
# namespace for the moment between creation and the move into the containers,
# where two concurrent runs would otherwise collide on one name.
#
# The scope comes from a four-hex hash of the run suffix, not from the suffix
# itself. An interface name gets 15 characters and the suffix alone can spend
# 24 of them (`-20260910t025205-2802512`), so interpolating it produced
# `vhifb-20260910t025205-2802512b` and ip(8) refused the name outright. The
# chaos simulation had already met this and answers it in `sim.naming`, which
# the NAT topology script also calls; this uses the same token so the reaper
# in `ci-cleanup.sh` can match these interfaces by the same rule.
#
# Empty for an empty suffix, so a bare run keeps short unscoped names.
veth_token() {
local suffix="${FIPS_CI_NAME_SUFFIX:-}"
if [ -z "$suffix" ]; then
echo ""
return 0
fi
PYTHONPATH="$TESTING_DIR/chaos" python3 -m sim.naming "$suffix"
}
VETH_TOKEN="$(veth_token)"
HOST_VETH_A="vhifb${VETH_TOKEN}a"
HOST_VETH_B="vhifb${VETH_TOKEN}b"
HOST_VETH_C="vhifb${VETH_TOKEN}c"
HOST_VETH_D="vhifb${VETH_TOKEN}d"
SKIP_BUILD=false
KEEP_UP=false
while [ $# -gt 0 ]; do
case "$1" in
--skip-build) SKIP_BUILD=true; shift ;;
--keep-up) KEEP_UP=true; shift ;;
*) echo "Unknown option: $1" >&2; exit 1 ;;
esac
done
log() { echo "=== $*"; }
pass() { echo "PASS: $*"; }
fail() { echo "FAIL: $*" >&2; dump_diagnostics; exit 1; }
dump_diagnostics() {
echo "--- node-a transports ---" >&2
docker exec "$NODE_A" fipsctl show transports >&2 2>&1 || true
echo "--- node-a status ---" >&2
docker exec "$NODE_A" fipsctl show status >&2 2>&1 || true
echo "--- node-a log (tail) ---" >&2
docker logs --tail 80 "$NODE_A" >&2 2>&1 || true
echo "--- node-b log (tail) ---" >&2
docker logs --tail 80 "$NODE_B" >&2 2>&1 || true
return 0
}
cleanup() {
# Remove the veth pair wherever it survived: inside a container if the move
# succeeded, on the host if the run died between creation and the move.
docker exec "$NODE_A" ip link del "$LAB_IFACE" >/dev/null 2>&1 || true
ip_host "ip link del $HOST_VETH_A" >/dev/null 2>&1 || true
docker exec "$NODE_C" ip link del "$BOOT_IFACE" >/dev/null 2>&1 || true
ip_host "ip link del $HOST_VETH_C" >/dev/null 2>&1 || true
if [ "$KEEP_UP" = false ]; then
docker compose -f "$COMPOSE_FILE" down --volumes --remove-orphans >/dev/null 2>&1 || true
fi
return 0
}
# ── Host-namespace ip(8) ─────────────────────────────────────────────────
#
# Every `ip link` operation that touches the host network stack runs inside a
# short-lived privileged container sharing the host network and PID namespaces,
# for the reason testing/chaos/sim/veth.py documents at length: on macOS the
# containers live in the Docker VM, so running ip(8) on the macOS host could
# never reach them, while on Linux the shared namespaces make it identical to
# running ip(8) directly.
ip_host() {
docker run --rm --privileged --network host --pid host \
--entrypoint /bin/sh "$IMAGE" -c "$1"
}
# Docker's own view of a container, so the gate case can wait for a netns
# without implying the daemon inside it has started.
container_state() {
docker inspect -f '{{.State.Status}}' "$1" 2>/dev/null || true
}
container_pid() {
docker inspect -f '{{.State.Pid}}' "$1"
}
# Create the veth pair and move one end into each container.
create_veth() {
local pid_a pid_b
pid_a="$(container_pid "$NODE_A")"
pid_b="$(container_pid "$NODE_B")"
# One invocation, not three. The host-side names exist only between the
# `add` and the two `netns` moves, and ci-cleanup.sh's host-veth sweep is
# deliberately shaped to the chaos simulation's names and does not cover
# these — so the window in which a hard kill could strand them is kept to
# a single command, with the EXIT trap covering the rest.
ip_host "set -e
ip link add $HOST_VETH_A type veth peer name $HOST_VETH_B
ip link set $HOST_VETH_A netns $pid_a name $LAB_IFACE
ip link set $HOST_VETH_B netns $pid_b name $LAB_IFACE" >/dev/null
# A moved link arrives down. Presence is IFF_UP, so the daemon correctly
# does not bind until this runs — which is also why (c) can flap it with
# nothing but `ip link set down`.
docker exec "$NODE_A" ip link set "$LAB_IFACE" up
docker exec "$NODE_B" ip link set "$LAB_IFACE" up
return 0
}
# ── Daemon introspection ─────────────────────────────────────────────────
# Field of a named transport's `interface` block, or "" if the transport, the
# block, or the daemon is not there.
iface_field() {
docker exec "$1" fipsctl show transports 2>/dev/null | python3 -c '
import json, sys
try:
data = json.load(sys.stdin)
except Exception:
print(""); raise SystemExit
for t in data.get("transports", []):
if t.get("name") == sys.argv[1]:
print(t.get("interface", {}).get(sys.argv[2], ""))
break
else:
print("")
' "$2" "$3"
}
node_state() {
docker exec "$1" fipsctl show status 2>/dev/null | python3 -c '
import json, sys
try:
print(json.load(sys.stdin).get("state", ""))
except Exception:
print("")
'
}
peer_count() {
docker exec "$1" fipsctl show peers 2>/dev/null | python3 -c '
import json, sys
try:
data = json.load(sys.stdin)
except Exception:
print(0); raise SystemExit
print(len(data.get("peers", [])))
'
}
# Count of a literal in a container log. Used for the log-hygiene assertion.
log_count() {
docker logs "$1" 2>&1 | grep -c -- "$2" || true
}
# Poll `expr` until it prints `want`, up to `timeout` seconds.
# Usage: wait_for <timeout> <want> <command...>
wait_for() {
local timeout="$1" want="$2"; shift 2
local i got
for i in $(seq 1 "$timeout"); do
got="$("$@" || true)"
if [ "$got" = "$want" ]; then
return 0
fi
sleep 1
done
echo " (last value: '${got:-}', wanted '$want')" >&2
return 1
}
# Poll until the command prints a value no greater than `want`.
wait_for_at_most() {
local timeout="$1" want="$2"; shift 2
local i got
for i in $(seq 1 "$timeout"); do
got="$("$@" || true)"
if [ -n "$got" ] && [ "$got" -le "$want" ] 2>/dev/null; then
return 0
fi
sleep 1
done
echo " (last value: '${got:-}', wanted <= '$want')" >&2
return 1
}
# Poll until the command prints a value that is at least `want`.
wait_for_at_least() {
local timeout="$1" want="$2"; shift 2
local i got
for i in $(seq 1 "$timeout"); do
got="$("$@" || true)"
if [ -n "$got" ] && [ "$got" -ge "$want" ] 2>/dev/null; then
return 0
fi
sleep 1
done
echo " (last value: '${got:-}', wanted >= '$want')" >&2
return 1
}
# ── Run ──────────────────────────────────────────────────────────────────
IMAGE="${FIPS_TEST_IMAGE:-fips-test:latest}"
trap cleanup EXIT
if [ "$SKIP_BUILD" = false ]; then
log "Building test image"
bash "$TESTING_DIR/scripts/build.sh"
fi
log "Generating fixtures"
LAB_IFACE="$LAB_IFACE" DOCK_IFACE="$DOCK_IFACE" BOOT_IFACE="$BOOT_IFACE" \
bash "$SCRIPT_DIR/generate-configs.sh"
log "Starting nodes with $LAB_IFACE absent"
docker compose -f "$COMPOSE_FILE" up -d
# The daemon must reach a serving state without the interface. Waiting on the
# control socket answering at all is the first half of assertion (a): a daemon
# that exited on `NoTransports` never answers.
if ! wait_for 40 "degraded" node_state "$NODE_A"; then
fail "(a) node-a did not come up Degraded with $LAB_IFACE absent"
fi
pass "(a) daemon started and serves with its only interface absent"
# ── (a) presence and policy are visible to an operator ───────────────────
[ "$(iface_field "$NODE_A" lab presence)" = "absent" ] \
|| fail "(a) lab transport is not reported ABSENT"
[ "$(iface_field "$NODE_A" lab policy)" = "required" ] \
|| fail "(a) lab transport is not reported required"
[ "$(iface_field "$NODE_A" lab name)" = "$LAB_IFACE" ] \
|| fail "(a) lab transport does not name $LAB_IFACE"
[ "$(iface_field "$NODE_A" dock presence)" = "absent" ] \
|| fail "(e) dock transport is not reported ABSENT"
[ "$(iface_field "$NODE_A" dock policy)" = "optional" ] \
|| fail "(e) dock transport is not reported optional"
pass "(a) presence and policy are visible in show_transports"
# ── log hygiene, measured across the absence ─────────────────────────────
#
# The required interface's absence must be logged once, on the edge — not once
# per retry. Measured over an interval long enough for many retries (the binder
# polls every second). The 12 s also carries the next assertion past the 10 s
# bring-up window, so both the edge rule and the deadline are covered by one
# wait.
# Matched on the edge line specifically. The deadline error below is a
# different line by design, and counting both here would read the deadline
# firing during the sleep as a repeated edge.
EDGE_LINE="Ethernet interface absent; waiting"
absent_before="$(log_count "$NODE_A" "$EDGE_LINE")"
sleep 12
absent_after="$(log_count "$NODE_A" "$EDGE_LINE")"
if [ "$absent_after" -ne "$absent_before" ]; then
fail "absence was logged $((absent_after - absent_before)) more times over 12s of retries; \
the edge must be logged once, not per attempt"
fi
# Two edges total: one for the required lab interface, one for the optional
# dock interface. More would mean the edge is not an edge.
[ "$absent_before" -eq 2 ] \
|| fail "expected exactly 2 absence edges at start, saw $absent_before"
pass "absence is logged once on the edge, not per retry"
# ── a sustained absence errors exactly once ──────────────────────────────
#
# The edge is not an error: an interface missing for a moment at boot and bound
# a moment later is the ordinary case this mechanism exists to absorb. Past the
# 10 s bring-up window it is no longer a race, and a *required* interface says
# so — once. The 12 s above put us on the far side of that window.
#
# Exactly one line, from lab. The optional dock interface has been absent just
# as long and must be silent, which is what `optional` means; two here would
# mean the policy is not being consulted.
startup_errors="$(log_count "$NODE_A" " ERROR ")"
if [ "$startup_errors" -ne 1 ]; then
docker logs "$NODE_A" 2>&1 | grep -- " ERROR " >&2 || true
fail "expected exactly 1 ERROR line for the required interface past the \
bring-up window, saw $startup_errors (an optional interface must contribute none)"
fi
[ "$(log_count "$NODE_A" "still missing past the bring-up window")" -eq 1 ] \
|| fail "the ERROR line is not the sustained-absence report"
pass "a required interface absent past the window errors, an optional one does not"
# ...and does not keep saying it. Duration is state, published as
# interface.since_secs and as Degraded; re-announcing it on a timer is what
# the old 1 m / 10 m / 1 h ladder did.
sleep 12
[ "$(log_count "$NODE_A" " ERROR ")" -eq 1 ] \
|| fail "the sustained-absence error repeated; it must be said once per episode"
pass "the sustained-absence error is said once, not on a schedule"
# ── (b) late attach ──────────────────────────────────────────────────────
log "Creating the veth pair"
create_veth
if ! wait_for 30 "present" iface_field "$NODE_A" lab presence; then
fail "(b) node-a did not bind $LAB_IFACE after it appeared"
fi
if ! wait_for 30 "present" iface_field "$NODE_B" lab presence; then
fail "(b) node-b did not bind $LAB_IFACE after it appeared"
fi
pass "(b) both daemons bound the interface with no restart"
# Health must clear. This is `Degraded` behaving as a level rather than a
# latch — the property that made the old monotonic failed-set wrong.
if ! wait_for 20 "running" node_state "$NODE_A"; then
fail "(b) node-a stayed Degraded after its interface returned"
fi
pass "(b) Degraded cleared when the interface came back"
# The optional interface is still absent and must still not matter.
[ "$(iface_field "$NODE_A" dock presence)" = "absent" ] \
|| fail "(e) dock unexpectedly bound"
pass "(e) an absent optional interface does not degrade the node"
# Discovery and peering over the late-bound interface: the point of binding at
# all. Without this the suite would prove the daemon can open a socket, not
# that traffic flows over it.
if ! wait_for_at_least 45 1 peer_count "$NODE_A"; then
fail "(b) node-a found no peer over the late-bound interface"
fi
if ! wait_for_at_least 45 1 peer_count "$NODE_B"; then
fail "(b) node-b found no peer over the late-bound interface"
fi
pass "(b) nodes discovered and peered over the late-bound interface"
# ── (c) flap ─────────────────────────────────────────────────────────────
errors_before_detach="$(log_count "$NODE_A" " ERROR ")"
log "Taking $LAB_IFACE down on node-a"
docker exec "$NODE_A" ip link set "$LAB_IFACE" down
detach_start=$SECONDS
if ! wait_for 20 "absent" iface_field "$NODE_A" lab presence; then
fail "(c) node-a did not notice the interface going down"
fi
detach_elapsed=$(( SECONDS - detach_start ))
if ! wait_for 20 "degraded" node_state "$NODE_A"; then
fail "(c) node-a did not report Degraded while its interface was down"
fi
pass "(c) a link going down is observed as absence and degrades health"
# A detach is reported, at warn. The edge is not an error — a link coming and
# going is the weather in a mesh daemon, and the error is the 10 s deadline's
# to give, not the edge's.
if ! wait_for_at_least 10 1 log_count "$NODE_A" "Ethernet interface detached"; then
fail "(c) a runtime detach was not reported"
fi
# Only meaningful while we are still inside the bring-up window. Detection is
# sub-second over netlink, so this is the ordinary path; if the runner was slow
# enough that the deadline could have fired, the check has nothing to say and
# says so rather than failing on the harness's own latency.
if [ "$detach_elapsed" -lt 8 ]; then
detach_errors="$(log_count "$NODE_A" " ERROR ")"
if [ "$detach_errors" -ne "$errors_before_detach" ]; then
docker logs "$NODE_A" 2>&1 | grep -- " ERROR " >&2 || true
fail "(c) the detach edge logged an ERROR after ${detach_elapsed}s; the \
edge is a warn and only outlasting the window earns an error"
fi
pass "(c) a detach is reported without crying error"
else
echo " (skipped the edge-not-an-error check: detach took ${detach_elapsed}s,"
echo " which is inside the deadline's reach)"
fi
# The peers that interface carried must go with it, and go *now*.
#
# `link_dead_timeout_secs` is at its 30 s default here, so a withdrawal inside
# 15 s can only have come from the detach edge and not from the liveness
# reaper. That gap is the whole point: until the edge drove the teardown, this
# node kept the peer, kept selecting routes through it, and kept advertising
# reachability it no longer had — dropping transit traffic in silence for the
# whole timeout, with alternative paths sitting unused.
if ! wait_for_at_most 15 0 peer_count "$NODE_A"; then
fail "(c) node-a kept a peer that was only reachable over the downed \
interface; the detach edge did not withdraw it"
fi
pass "(c) the peers the interface carried were withdrawn on the detach edge"
log "Bringing $LAB_IFACE back up on node-a"
docker exec "$NODE_A" ip link set "$LAB_IFACE" up
if ! wait_for 30 "present" iface_field "$NODE_A" lab presence; then
fail "(c) node-a did not rebind after the interface came back"
fi
if ! wait_for 20 "running" node_state "$NODE_A"; then
fail "(c) node-a stayed Degraded after the interface came back"
fi
if ! wait_for_at_least 45 2 iface_field "$NODE_A" lab binds; then
fail "(c) the rebind was not counted"
fi
pass "(c) the interface flapped and the daemon followed it both ways"
# And the withdrawal is not a one-way door: the peer comes back over the
# rebound interface on its own, by beacon, with no operator action.
if ! wait_for_at_least 60 1 peer_count "$NODE_A"; then
fail "(c) node-a did not re-peer after its interface came back"
fi
pass "(c) peering re-established over the rebound interface"
# ── (d) destroy and recreate ─────────────────────────────────────────────
#
# Deleting the netdev outright is the case the old ENXIO beacon-socket reopen
# half-covered: the veth is gone, the socket underneath is stale, and the name
# comes back a moment later. One presence machine now covers it.
log "Deleting and recreating the veth pair"
docker exec "$NODE_A" ip link del "$LAB_IFACE"
if ! wait_for 20 "absent" iface_field "$NODE_A" lab presence; then
fail "(d) node-a did not notice the interface being deleted"
fi
if ! wait_for 20 "absent" iface_field "$NODE_B" lab presence; then
fail "(d) node-b did not notice its end of the pair disappearing"
fi
create_veth
if ! wait_for 30 "present" iface_field "$NODE_A" lab presence; then
fail "(d) node-a did not rebind the recreated interface"
fi
if ! wait_for 30 "present" iface_field "$NODE_B" lab presence; then
fail "(d) node-b did not rebind the recreated interface"
fi
if ! wait_for 20 "running" node_state "$NODE_A"; then
fail "(d) node-a stayed Degraded after the interface was recreated"
fi
pass "(d) a destroyed and recreated interface is rebound"
# Peering must re-establish over the new hardware. A recreated veth has a new
# MAC, so this also exercises the "same name, different hardware" path that
# drops cached neighbors instead of resuming onto them.
if ! wait_for_at_least 60 1 peer_count "$NODE_A"; then
fail "(d) node-a did not re-peer after the interface was recreated"
fi
pass "(d) peering re-established over the recreated interface"
# ── (f) an interface present before the daemon starts ────────────────────
#
# Everything above binds through the binder loop, because the interface does
# not exist until the harness makes it. The ordinary case on a booted router is
# the opposite one: the interface is already there and `start_async` binds it
# inline, before the loop is running.
#
# That path published its presence edge outside the churn guard, so the guard
# believed it had announced nothing and the *first* detach asked for no
# retraction. The node kept reporting Running with its only required interface
# gone, and stayed that way until a second detach happened to repair the guard.
# Nothing in cases (a)-(e) can reach it.
log "(f) starting node-c with its interface already present"
docker compose -f "$COMPOSE_FILE" up -d node-c
# The container parks on the gate, so this is the netns and not yet the daemon.
if ! wait_for 30 "running" container_state "$NODE_C"; then
fail "(f) node-c container did not start"
fi
pid_c="$(container_pid "$NODE_C")"
ip_host "set -e
ip link add $HOST_VETH_C type veth peer name $HOST_VETH_D
ip link set $HOST_VETH_C netns $pid_c name $BOOT_IFACE
ip link set $HOST_VETH_D netns $pid_c name ${BOOT_IFACE}p" >/dev/null
docker exec "$NODE_C" ip link set "$BOOT_IFACE" up
docker exec "$NODE_C" ip link set "${BOOT_IFACE}p" up
# Release the gate. The daemon now starts with the interface already up.
docker exec "$NODE_C" touch /tmp/fips-go
if ! wait_for 40 "running" node_state "$NODE_C"; then
fail "(f) node-c did not come up Running with its interface present at start"
fi
[ "$(iface_field "$NODE_C" boot presence)" = "present" ] \
|| fail "(f) node-c did not bind $BOOT_IFACE inline at start"
pass "(f) an interface present at start is bound inline and reports Running"
# The assertion. One detach, on a binding this loop did not create.
docker exec "$NODE_C" ip link set "$BOOT_IFACE" down
if ! wait_for 30 "absent" iface_field "$NODE_C" boot presence; then
fail "(f) node-c did not notice $BOOT_IFACE going down"
fi
if ! wait_for 30 "degraded" node_state "$NODE_C"; then
fail "(f) node-c stayed Running after its only required interface went \
away — the start-time bind never reached node health"
fi
pass "(f) the first detach after a clean start degrades the node"
# And it is still a level, not a latch, on this path too.
docker exec "$NODE_C" ip link set "$BOOT_IFACE" up
if ! wait_for 30 "running" node_state "$NODE_C"; then
fail "(f) node-c stayed Degraded after its interface returned"
fi
pass "(f) health clears again when the interface returns"
# ── (g) the link-event fast path is actually the one in use ──────────────
#
# The whole suite would pass with `open_link_socket()` hardcoded to Err: the
# 1 s poll is a complete fallback and covers every `wait_for` window here, so
# nothing else asserts that the netlink path exists, let alone that it is what
# detected anything. The binder says which backing it got at startup, so ask
# it directly rather than inferring from timing that the poll would also
# satisfy.
# `log_count`, not `grep -q`: under `set -o pipefail` a `grep -q` that exits on
# its first match closes the pipe, `docker logs` takes SIGPIPE, and the
# pipeline reports failure even though the line was found. `grep -c` reads the
# stream to the end.
if [ "$(log_count "$NODE_A" "event_driven=true")" -eq 0 ]; then
docker logs "$NODE_A" 2>&1 | grep -i "binder started" >&2 || true
fail "(g) the binder fell back to polling; the netlink link-event source \
did not open, and every timing assertion in this suite would still pass"
fi
pass "(g) detection is driven by netlink events, not by the poll fallback"
# ── (h) churn damping engages on a genuinely flapping interface ──────────
#
# This is load-bearing twice over. It is what stops a flapping interface
# logging a recovery per cycle, and — since the detach edge now withdraws the
# peers that interface carried — it is also the only thing bounding how often
# that withdrawal can fire. Nothing exercised it: every flap elsewhere in this
# suite is a single down/up with long settles either side, which is precisely
# the shape the damper ignores.
#
# Four bindings that each die well inside MIN_STABLE_BINDING (10 s). The
# streak crosses CHURN_THRESHOLD (3) on the third, which is the edge that
# announces itself.
log "(h) flapping $LAB_IFACE to drive the churn guard"
for _ in 1 2 3 4; do
docker exec "$NODE_A" ip link set "$LAB_IFACE" down
sleep 1
docker exec "$NODE_A" ip link set "$LAB_IFACE" up
sleep 2
done
if ! wait_for_at_least 30 1 log_count "$NODE_A" "keeps dying immediately after binding"; then
docker logs "$NODE_A" 2>&1 | grep -i "ethernet" | tail -20 >&2
fail "(h) four short-lived bindings did not engage the churn guard"
fi
pass "(h) a flapping interface engages churn damping"
# Having engaged, the guard must hold health rather than announcing each bind.
# The failure this catches is a damper that counts but does not damp.
recoveries_during_churn="$(log_count "$NODE_A" "Ethernet interface recovered")"
if [ "$recoveries_during_churn" -gt 6 ]; then
fail "(h) node-a announced $recoveries_during_churn recoveries; the guard \
counted the churn but kept announcing through it"
fi
pass "(h) churn suppressed the per-cycle recovery announcements"
# And it is not a latch: once a binding lasts, the interface is announced
# again and the node returns to Running on its own.
log "(h) letting $LAB_IFACE settle"
docker exec "$NODE_A" ip link set "$LAB_IFACE" up >/dev/null 2>&1 || true
if ! wait_for 60 "present" iface_field "$NODE_A" lab presence; then
fail "(h) node-a did not rebind after the flapping stopped"
fi
if ! wait_for 60 "running" node_state "$NODE_A"; then
fail "(h) node-a stayed Degraded after the flapping stopped"
fi
pass "(h) a settled interface is announced again after churn"
# ── final log hygiene ────────────────────────────────────────────────────
#
# Four outages happened above (start, down, delete, and node-b's end of the
# delete). A generous ceiling still catches the failure mode that matters: a
# retry loop logging per attempt would be in the hundreds by now.
edges="$(log_count "$NODE_A" "Ethernet interface")"
[ "$edges" -lt 40 ] \
|| fail "node-a logged $edges interface lines; the edges are not edges"
pass "log volume stayed proportional to edges, not to retries"
echo
echo "ALL PASSED"