Files
fips/testing/iface-binding
Johnathan Corgan 1eb0e8a34e Move the TUN and DNS child start and stop bodies into ipv6tun
Bringing the TUN device and the .fips DNS responder up and taking them
down is host-side work, but the bodies sat inline in the node's
supervisor arms, and their handles were eight loose fields on the
supervisor. Move the bodies to ipv6tun::lifecycle and gather the
handles into one Handles struct there, held by the supervisor. The
supervisor arms, their order and the child-exit reporting are
unchanged; each arm now calls into ipv6tun.

The TUN start is two calls so the node can refresh its MSS ceiling
between them, exactly where it did before: open_tun creates and logs
the device, then spawn_tun creates the macOS/FreeBSD shutdown pipe and
starts the writer and reader threads. A failure to create the device
still continues without a TUN, and a pipe or writer failure still fails
the node's start. stop_tun and stop_dns carry the teardown unchanged,
including the shutdown-pipe write that wakes the reader on macOS and
FreeBSD.

The TUN device name moves into Handles as well, so the teardown up-set
can ask ipv6tun whether each child is up. A TUN counts as up when it
has a device name, not when it has a sender, so an app-owned TUN still
produces no TUN teardown; DNS counts as up while its task handle
exists. Node::tun_name, tun_tx, dns_local_addr and
enable_app_owned_tun keep their behaviour and now read or write the
handles. Node::mesh_ifindex had no caller left outside a test and is
replaced by the same method on Handles. Tests install a TUN sender
through a test-only Node::install_tun.

The moved log lines now log under fips::ipv6tun::lifecycle instead of
fips::node::lifecycle. Add that target to the NAT harness and its trace
overlay, and to the harnesses that relied on fips::node=debug, and note
the rename in the changelog.
2026-09-24 14:45:51 +00:00
..

Dynamic Interface Binding

Two FIPS daemons whose only transports are bound to network interfaces, exercised against a veth pair the harness creates, downs, deletes and recreates underneath them while they run.

node-a                                     node-b
  lab   ve-lab0     required  ── veth ──   ve-lab0     required
  dock  fips-dock0  optional               fips-dock0  optional

ve-lab0 does not exist when the daemons start. fips-dock0 never exists at all, on any host, ever — it is the negative control for optional: true.

What it asserts

Behavior
(a) A daemon whose only interface is missing starts, reports the transport absent, and reports Degraded — it does not exit on NoTransports, and it does not skip the transport for the life of the process
(b) The interface appears; both daemons bind it with no restart, Degraded clears, and they discover and peer over it
(c) The interface goes down and comes back; presence and health follow it in both directions, and the rebind is counted
(d) The interface is deleted outright and recreated; both daemons rebind and re-peer — the case the old ENXIO beacon-socket reopen half-covered
(e) An optional interface that never appears logs at info and never moves node health
Absence is logged once on the edge, not once per retry

Health is asserted through fipsctl show status (state), presence through fipsctl show transports (the per-transport interface block: presence, policy, binds, since_secs).

Running

./test.sh                 # builds the image first
./test.sh --skip-build    # reuse an existing image
./test.sh --keep-up       # leave the containers running for inspection

Via the local CI runner:

./testing/ci-local.sh --only iface-binding

Notes

The containers run under FIPS_TEST_MODE=default, not chaos. The chaos entrypoint waits up to 30 s for every configured Ethernet interface before starting the daemon — which is exactly the workaround this mechanism retires. The daemon has to do its own waiting here or the suite proves nothing.

Every ip link operation on the host network stack runs inside a short-lived privileged container sharing the host network and PID namespaces, for the reason chaos/sim/veth.py documents: on macOS the containers live in the Docker VM, so ip(8) run on the macOS host could never reach them, while on Linux the shared namespaces make it identical to running ip(8) directly.