Files
fips/testing/iface-binding
Arjen ee5384f013 test(iface-binding): assert the fast path and the churn guard
Three gaps, two of them in tests that existed and asserted nothing.

**The netlink path was never asserted to be in use.** The 1 s poll is a
complete fallback and covers every wait in the suite, so the whole thing
passed with `open_link_socket()` hardcoded to Err — the fast path could have
been dead for a release and no test would have said so. The binder reports
which backing it got at startup, so case (g) asks it directly rather than
inferring from timing the poll would also satisfy, and the unit test that used
to write `let _ = w.is_event_driven();` now asserts it on Linux, where the
source is an unprivileged `AF_NETLINK` socket and falling back is a real loss
rather than a sandbox's prerogative.

**Churn damping had no end-to-end coverage**, which now matters twice over: it
bounds the recovery announcements, and since the detach edge withdraws peers
it is also the only thing bounding how often that withdrawal fires. Every flap
elsewhere in the suite is a single down/up with long settles either side —
exactly the shape the damper ignores. Case (h) drives four bindings that each
die inside `MIN_STABLE_BINDING`, asserts the guard engages, asserts it then
*suppresses* rather than merely counting, and asserts it is not a latch.

**`a_poisoned_binding_does_not_strand_the_transport` discarded its result.**
`let _ = eth.binding.tasks_alive();` left the entire point unasserted: reading
a poisoned lock as "alive" would have the binder believe a dead binding
healthy and never rebind, and treating it as an error would strand the
transport. `false` is what routes it back through detach and rebind, so say so.

`a_stop_racing_a_bind_leaves_nothing_behind` now asserts the error *kind*.
`bind_and_spawn` refuses at its presence probe long before the post-store
shutdown check, so `is_err()` alone passed on absence and would still pass
with that check deleted. The test keeps the coverage it genuinely has — stop
raises the flag before teardown, teardown leaves no socket and no loops — and
says plainly that the race it is named for needs a bind that succeeds, which
needs privilege no unit test has.

Both new cases were verified against the defect: with the netlink source
forced to Err, (g) fails; with `CHURN_THRESHOLD` raised out of reach, (h)
fails. Nothing else in the suite notices either.

One case was attempted and removed rather than shipped: `"interface
replaced"` cannot be produced deterministically, because the delete that
changes an ifindex fires a netlink event the binder acts on within
microseconds, so `gone` wins the race. It passed about one run in three.
reference/notes.md records the measurement and the two approaches that could
work.

Also fixes a real bug in the harness: `grep -q` under `set -o pipefail` exits
on its first match, `docker logs` takes SIGPIPE, and the pipeline reports
failure even though the line was found. That cost two false failures before it
was spotted; `log_count` reads the stream to the end.
2026-09-02 10:13:42 +01:00
..

Dynamic Interface Binding

Two FIPS daemons whose only transports are bound to network interfaces, exercised against a veth pair the harness creates, downs, deletes and recreates underneath them while they run.

node-a                                     node-b
  lab   ve-lab0     required  ── veth ──   ve-lab0     required
  dock  fips-dock0  optional               fips-dock0  optional

ve-lab0 does not exist when the daemons start. fips-dock0 never exists at all, on any host, ever — it is the negative control for optional: true.

What it asserts

Behavior
(a) A daemon whose only interface is missing starts, reports the transport absent, and reports Degraded — it does not exit on NoTransports, and it does not skip the transport for the life of the process
(b) The interface appears; both daemons bind it with no restart, Degraded clears, and they discover and peer over it
(c) The interface goes down and comes back; presence and health follow it in both directions, and the rebind is counted
(d) The interface is deleted outright and recreated; both daemons rebind and re-peer — the case the old ENXIO beacon-socket reopen half-covered
(e) An optional interface that never appears logs at info and never moves node health
Absence is logged once on the edge, not once per retry

Health is asserted through fipsctl show status (state), presence through fipsctl show transports (the per-transport interface block: presence, policy, binds, since_secs).

Running

./test.sh                 # builds the image first
./test.sh --skip-build    # reuse an existing image
./test.sh --keep-up       # leave the containers running for inspection

Via the local CI runner:

./testing/ci-local.sh --only iface-binding

Notes

The containers run under FIPS_TEST_MODE=default, not chaos. The chaos entrypoint waits up to 30 s for every configured Ethernet interface before starting the daemon — which is exactly the workaround this mechanism retires. The daemon has to do its own waiting here or the suite proves nothing.

Every ip link operation on the host network stack runs inside a short-lived privileged container sharing the host network and PID namespaces, for the reason chaos/sim/veth.py documents: on macOS the containers live in the Docker VM, so ip(8) run on the macOS host could never reach them, while on Linux the shared namespaces make it identical to running ip(8) directly.