Files
fips/testing/iface-binding/README.md
T
Arjen 6485271ef5 test(iface-binding): cover the presence machine end to end
Two daemons whose only transports are interface-bound, run against a veth
pair the harness creates, downs, deletes and recreates underneath them.

Asserts the boot race (a daemon whose only interface is missing starts,
reports the transport absent and the node Degraded, rather than exiting on
NoTransports or skipping the transport for the life of the process), the
late attach and discovery over it, the flap in both directions,
destroy-and-recreate, and that an optional interface which never appears
never moves node health.

Also the log policy, which is the half that is easy to regress silently:
absence is logged once on the edge and not once per retry; a required
interface still absent past the ten-second bring-up window errors exactly
once, while the optional one — absent just as long — stays silent; and that
error is not repeated on a schedule. The detach edge is checked not to
error, guarded by how long detection actually took, so a slow runner skips
the check rather than failing on the harness's own latency.

The containers run FIPS_TEST_MODE=default, not chaos. The chaos entrypoint
waits up to 30 s for every configured Ethernet interface before starting the
daemon, which is precisely the workaround under test — the daemon has to do
its own waiting here or the suite proves nothing.

Host-namespace ip(8) runs in a short-lived privileged container sharing the
host network and PID namespaces, for the reason chaos/sim/veth.py documents:
on macOS the containers live in the Docker VM, so ip(8) run on the macOS
host could never reach them.

Chaos ethernet transports are marked optional: true. In that harness a
neighbour's interface disappearing is the scenario, not a fault — node_churn
stops a container, which destroys its netns and with it both ends of every
veth it held, so a surviving node watches a required interface vanish for
the 30-90 s the neighbour is down, once per churn event. Reporting that at
error is right for a deployment and wrong for a harness that tears the
interface down on purpose; the mesh-wide zero-ERROR ceiling would have
failed on injected chaos rather than on a defect.
2026-09-01 09:54:36 +01:00

58 lines
2.5 KiB
Markdown

# Dynamic Interface Binding
Two FIPS daemons whose **only** transports are bound to network interfaces,
exercised against a veth pair the harness creates, downs, deletes and recreates
underneath them while they run.
```
node-a node-b
lab ve-lab0 required ── veth ── ve-lab0 required
dock fips-dock0 optional fips-dock0 optional
```
`ve-lab0` does not exist when the daemons start. `fips-dock0` never exists at
all, on any host, ever — it is the negative control for `optional: true`.
## What it asserts
| | Behavior |
| - | -------- |
| (a) | A daemon whose only interface is missing **starts**, reports the transport `absent`, and reports `Degraded` — it does not exit on `NoTransports`, and it does not skip the transport for the life of the process |
| (b) | The interface appears; both daemons bind it with no restart, `Degraded` clears, and they discover and peer over it |
| (c) | The interface goes down and comes back; presence and health follow it in **both** directions, and the rebind is counted |
| (d) | The interface is deleted outright and recreated; both daemons rebind and re-peer — the case the old ENXIO beacon-socket reopen half-covered |
| (e) | An `optional` interface that never appears logs at `info` and never moves node health |
| | Absence is logged **once on the edge**, not once per retry |
Health is asserted through `fipsctl show status` (`state`), presence through
`fipsctl show transports` (the per-transport `interface` block: `presence`,
`policy`, `binds`, `since_secs`).
## Running
```sh
./test.sh # builds the image first
./test.sh --skip-build # reuse an existing image
./test.sh --keep-up # leave the containers running for inspection
```
Via the local CI runner:
```sh
./testing/ci-local.sh --only iface-binding
```
## Notes
The containers run under `FIPS_TEST_MODE=default`, **not** `chaos`. The chaos
entrypoint waits up to 30 s for every configured Ethernet interface before
starting the daemon — which is exactly the workaround this mechanism retires.
The daemon has to do its own waiting here or the suite proves nothing.
Every `ip link` operation on the host network stack runs inside a short-lived
privileged container sharing the host network and PID namespaces, for the
reason [chaos/sim/veth.py](../chaos/sim/veth.py) documents: on macOS the
containers live in the Docker VM, so ip(8) run on the macOS host could never
reach them, while on Linux the shared namespaces make it identical to running
ip(8) directly.