Losing an interface withdrew nothing. The peers stayed in the registry, the routes through them stayed selectable, and this node kept advertising reachability it no longer had — so transit traffic was dropped in silence and other nodes kept routing toward us for those destinations, until the liveness reaper noticed up to link_dead_timeout_secs (30 s) later. Measured on real hardware: a dongle detached at 07:37:19 took the node's parent with it, and no new parent was chosen until 07:37:46. Twenty-seven seconds routing through a link that had already gone, with four alternative peers available the whole time. The alternatives are the point — a mesh that can route around a dead link should not be the last to hear the link is dead. The detach edge is both earlier and more certain than inactivity, so it is the better trigger. reap_peers_on_transport routes through the same route_link_dead the liveness reaper uses rather than open-coding a second teardown: every consequence of losing a peer — sessions, path MTU release, session indices, decrypt-worker unregistration, the link, the control machine, tree cleanup and re-announce, bloom withdrawal — already hangs off that one path, and a parallel one would drift from it. Not policy-filtered. Whether an interface's absence is normal is a statement about node *health*; it says nothing about whether the routes over it still work. An optional interface's peers are exactly as unreachable. PathBroken needs no new wiring. Once the peers are gone resolve_next_hop returns None, which takes the NoRoute path — and that one already synthesises the routing error, rate limiting included. The cure for the silent drop was to stop having a route, not to add a second error path. Deliberately undamped. A flapping interface cannot drive a reap storm through here: ChurnGuard suppresses `announce` after three short-lived bindings, which leaves `announced` false, which makes `detached()` return `retract: false` — so no presence edge is published at all during churn. The edges this reacts to are already rate-limited at the source, reaping an already-reaped transport is a no-op, and a second damper would only add a way for the two to disagree. The trade taken: immediate reaping costs a re-peer for an absence shorter than the dead timeout that then recovers — a `wifi reload` returns in ~5 s and today costs nothing, where this costs ~15 s of re-peering. Accepted, because black-holing is silent, poisons other nodes' routing and needs the full timeout to clear, where a re-peer is bounded, visible and self-healing. A grace period remains available if that proves wrong; reference/ records its shape. The integration assertion is the one that proves the wiring rather than the unit: link_dead_timeout_secs is left at its 30 s default, so a withdrawal inside 15 s can only have come from the detach edge. Verified against the defect — with the reap disabled, that assertion fails and every other case in the suite still passes.
Dynamic Interface Binding
Two FIPS daemons whose only transports are bound to network interfaces, exercised against a veth pair the harness creates, downs, deletes and recreates underneath them while they run.
node-a node-b
lab ve-lab0 required ── veth ── ve-lab0 required
dock fips-dock0 optional fips-dock0 optional
ve-lab0 does not exist when the daemons start. fips-dock0 never exists at
all, on any host, ever — it is the negative control for optional: true.
What it asserts
| Behavior | |
|---|---|
| (a) | A daemon whose only interface is missing starts, reports the transport absent, and reports Degraded — it does not exit on NoTransports, and it does not skip the transport for the life of the process |
| (b) | The interface appears; both daemons bind it with no restart, Degraded clears, and they discover and peer over it |
| (c) | The interface goes down and comes back; presence and health follow it in both directions, and the rebind is counted |
| (d) | The interface is deleted outright and recreated; both daemons rebind and re-peer — the case the old ENXIO beacon-socket reopen half-covered |
| (e) | An optional interface that never appears logs at info and never moves node health |
| Absence is logged once on the edge, not once per retry |
Health is asserted through fipsctl show status (state), presence through
fipsctl show transports (the per-transport interface block: presence,
policy, binds, since_secs).
Running
./test.sh # builds the image first
./test.sh --skip-build # reuse an existing image
./test.sh --keep-up # leave the containers running for inspection
Via the local CI runner:
./testing/ci-local.sh --only iface-binding
Notes
The containers run under FIPS_TEST_MODE=default, not chaos. The chaos
entrypoint waits up to 30 s for every configured Ethernet interface before
starting the daemon — which is exactly the workaround this mechanism retires.
The daemon has to do its own waiting here or the suite proves nothing.
Every ip link operation on the host network stack runs inside a short-lived
privileged container sharing the host network and PID namespaces, for the
reason chaos/sim/veth.py documents: on macOS the
containers live in the Docker VM, so ip(8) run on the macOS host could never
reach them, while on Linux the shared namespaces make it identical to running
ip(8) directly.