Files
fips/testing
ArjenandJohnathan Corgan 95866b7c7c test(iface-binding): cover the presence machine end to end
Two daemons whose only transports are interface-bound, run against a veth
pair the harness creates, downs, deletes and recreates underneath them.

Asserts the boot race (a daemon whose only interface is missing starts,
reports the transport absent and the node Degraded, rather than exiting on
NoTransports or skipping the transport for the life of the process), the
late attach and discovery over it, the flap in both directions,
destroy-and-recreate, and that an optional interface which never appears
never moves node health.

Also the log policy, which is the half that is easy to regress silently:
absence is logged once on the edge and not once per retry; a required
interface still absent past the ten-second bring-up window errors exactly
once, while the optional one — absent just as long — stays silent; and that
error is not repeated on a schedule. The detach edge is checked not to
error, guarded by how long detection actually took, so a slow runner skips
the check rather than failing on the harness's own latency.

The containers run FIPS_TEST_MODE=default, not chaos. The chaos entrypoint
waits up to 30 s for every configured Ethernet interface before starting the
daemon, which is precisely the workaround under test — the daemon has to do
its own waiting here or the suite proves nothing.

Host-namespace ip(8) runs in a short-lived privileged container sharing the
host network and PID namespaces, for the reason chaos/sim/veth.py documents:
on macOS the containers live in the Docker VM, so ip(8) run on the macOS
host could never reach them.

Chaos ethernet transports are marked optional: true. In that harness a
neighbour's interface disappearing is the scenario, not a fault — node_churn
stops a container, which destroys its netns and with it both ends of every
veth it held, so a surviving node watches a required interface vanish for
the 30-90 s the neighbour is down, once per churn event. Reporting that at
error is right for a deployment and wrong for a harness that tears the
interface down on purpose; the mesh-wide zero-ERROR ceiling would have
failed on injected chaos rather than on a defect.

test(iface-binding): cover an interface present before the daemon starts

Every scenario in the suite created its interface after the daemons were
already running — that ordering is the boot race the suite was written for.
But it means both nodes could only ever reach Present through binder_loop,
so the inline bind in start_async, which is the ordinary case on a booted
router, had no end-to-end coverage at all. That is where the churn guard
went unseeded and the first detach stopped reaching node health, and no
existing case could reach it: they all detach from a binding the loop
created, which seeds the guard as a side effect.

Case (f) adds a third node whose single required interface exists before its
daemon does. The gate is what buys that ordering — the harness needs a
running container to have a netns to move a veth into, but the daemon must
not start until after the move, so node-c comes up parked on a file and the
harness releases it once the interface is in place. Then one detach, on a
binding the loop did not create, and the node must degrade.

Verified against the defect rather than only against the fix: with the guard
seed reverted, cases (a) through (e) all still pass and (f) is the only
failure. A regression test that has never been seen to fail is a claim, not
a test.

It also asserts the reverse edge, so Degraded stays a level rather than a
latch on this path too.

test(iface-binding): assert the fast path and the churn guard

Three gaps, two of them in tests that existed and asserted nothing.

**The netlink path was never asserted to be in use.** The 1 s poll is a
complete fallback and covers every wait in the suite, so the whole thing
passed with `open_link_socket()` hardcoded to Err — the fast path could have
been dead for a release and no test would have said so. The binder reports
which backing it got at startup, so case (g) asks it directly rather than
inferring from timing the poll would also satisfy, and the unit test that used
to write `let _ = w.is_event_driven();` now asserts it on Linux, where the
source is an unprivileged `AF_NETLINK` socket and falling back is a real loss
rather than a sandbox's prerogative.

**Churn damping had no end-to-end coverage**, which now matters twice over: it
bounds the recovery announcements, and since the detach edge withdraws peers
it is also the only thing bounding how often that withdrawal fires. Every flap
elsewhere in the suite is a single down/up with long settles either side —
exactly the shape the damper ignores. Case (h) drives four bindings that each
die inside `MIN_STABLE_BINDING`, asserts the guard engages, asserts it then
*suppresses* rather than merely counting, and asserts it is not a latch.

**`a_poisoned_binding_does_not_strand_the_transport` discarded its result.**
`let _ = eth.binding.tasks_alive();` left the entire point unasserted: reading
a poisoned lock as "alive" would have the binder believe a dead binding
healthy and never rebind, and treating it as an error would strand the
transport. `false` is what routes it back through detach and rebind, so say so.

`a_stop_racing_a_bind_leaves_nothing_behind` now asserts the error *kind*.
`bind_and_spawn` refuses at its presence probe long before the post-store
shutdown check, so `is_err()` alone passed on absence and would still pass
with that check deleted. The test keeps the coverage it genuinely has — stop
raises the flag before teardown, teardown leaves no socket and no loops — and
says plainly that the race it is named for needs a bind that succeeds, which
needs privilege no unit test has.

Both new cases were verified against the defect: with the netlink source
forced to Err, (g) fails; with `CHURN_THRESHOLD` raised out of reach, (h)
fails. Nothing else in the suite notices either.

One case was attempted and removed rather than shipped: `"interface
replaced"` cannot be produced deterministically, because the delete that
changes an ifindex fires a netlink event the binder acts on within
microseconds, so `gone` wins the race. It passed about one run in three.
reference/notes.md records the measurement and the two approaches that could
work.

Also fixes a real bug in the harness: `grep -q` under `set -o pipefail` exits
on its first match, `docker logs` takes SIGPIPE, and the pipeline reports
failure even though the line was found. That cost two false failures before it
was spotted; `log_count` reads the stream to the end.

test(chaos): cover an Ethernet rebind under active traffic

The one case dynamic interface binding had no coverage for anywhere: a
datagram crossing an Ethernet link while the interface underneath it goes away
and comes back.

No existing scenario could reach it, for two separate reasons.

`ethernet-only` and `ethernet-mesh` both run with `traffic.enabled: false`, so
no datagram crosses an Ethernet link in any test — `ethernet-only`'s own
comment says exactly that, and names framing, the length field that trims NIC
minimum-frame padding, and AEAD over Ethernet as unexercised because of it.

And `link_flaps` cannot produce a rebind whatever it is pointed at: it
simulates a down link with netem 100% loss, so the interface stays IFF_UP and
the presence machine never sees an edge. `ethernet-mesh` has had link flaps
enabled all along without once exercising a rebind.

`node_churn` is what actually moves an interface. Stopping a container
destroys its network namespace, deleting every veth in it — and deleting one
end of a veth deletes its peer — so a *surviving* node watches its Ethernet
interface disappear outright, and watches it return when the harness recreates
the pair on restart. That is a real detach and a real rebind, driven from
outside the daemon.

The new scenario is a 4-node Ethernet ring with traffic on and one node
churned at a time, with link flaps deliberately off so the only outage is a
genuine interface removal and a traffic shortfall cannot be ambiguous between
the two. Measured across four runs: 206-388 MB moved over Ethernet links while
interfaces were being taken away underneath.

It also needed an assertion that did not exist. Traffic results have always
been written to `iperf3-results.json` and never read, so a scenario carrying
`traffic.enabled: true` could have every session fail and still exit 0 on a
green control plane — and a rebind under load is precisely what a tree
snapshot cannot see. `min_traffic` counts sessions that finished with bytes
actually received, treating iperf3's top-level `error` and a missing `end`
block as zero, so a session only counts when it moved data.

The baseline is calibrated against four runs rather than assumed: `max_roots`
starts at the observed maximum plus one, and the site records the sample, its
size, and why four runs is thin. The first draft asserted a single root and
failed every run — the harness restores stopped nodes immediately before the
final snapshot, so a just-restarted node has not re-parented yet and is
briefly its own root. That is the scenario working.

Wired into both runners, since a chaos scenario on one side only makes "local
green" and "GitHub green" stop meaning the same thing; check-ci-parity was
confirmed to fail on a one-sided addition before this was committed.

The iface-binding suite's entry in the GitHub workflow's integration matrix
moves here from the commit that introduced the presence machine. That commit
declared the suite on GitHub before testing/iface-binding/ existed and before
testing/ci-local.sh knew about it, so testing/check-ci-parity.sh failed there
and the three workflow steps named files that were not yet in the tree.
Registering both runners in the commit that adds the suite settles both.
2026-09-10 19:18:09 +00:00
..
2026-08-22 10:46:42 +01:00
2026-08-30 10:42:59 +00:00

FIPS Testing

Integration and simulation test harnesses for FIPS, using Docker containers running the full protocol stack.

Test Harnesses

static/ -- Static Docker Network

Fixed topologies with manual scripts for building, config generation, connectivity tests (ping, iperf), and network impairment (netem). Useful for deterministic debugging and validating specific topology configurations.

Topology Nodes Transport Description
mesh 5 UDP Sparse mesh, 6 links, multi-hop
chain 5 UDP Linear chain, max 4-hop paths
rekey 5 UDP Rekey integration test topology

tor/ -- Tor Transport Integration

End-to-end Tor transport testing with Docker containers running real Tor daemons. Requires internet access for Tor bootstrapping.

Scenario Description
socks5-outbound Outbound SOCKS5 connections through Tor to clearnet peer
directory-mode Inbound via HiddenServiceDir onion service (co-located)

nat/ -- NAT Traversal Lab

Real Docker NAT traversal tests for the Nostr/STUN bootstrap path, using router containers with iptables-based NAT, a local Nostr relay, and a local STUN responder.

Scenario Description
cone Two NATed peers establish a UDP traversal path
symmetric UDP traversal fails under symmetric NAT, TCP fallback wins
lan Peers on the same LAN prefer local addresses over reflexive

chaos/ -- Stochastic Simulation

Automated network testing with configurable node counts, topology algorithms (random geometric, Erdos-Renyi, chain, explicit), and fault injection (netem mutation, link flaps, traffic generation, node churn). 10 scenarios covering general stress and node churn, discovery over sparse topologies, spanning-tree and bloom-propagation regression, transport-specific validation (UDP, TCP, Ethernet), and ECN/congestion testing. Scenarios are defined in YAML and executed via a Python harness that manages the full lifecycle: topology generation, Docker orchestration, fault scheduling, log collection, and analysis.

interop/ -- Mixed-Version Interop Harness

On-demand harness that runs an N-node full mesh from a node-spec where each node can run a different build of the FIPS daemon, then attributes every FMP/FSP/rekey/connectivity failure to a specific version pair (same-version vs MIXED). Used to catch interop regressions between builds, not as a per-commit CI gate; not part of ci-local.sh.

mesh-lab/ -- Mesh Reliability Lab

On-demand harness that runs a chosen integration suite N times under a configurable host-pressure profile (idle / light / github-runner- equivalent / heavy via stress-ng), per-container netem impairment, and optional trace-level RUST_LOG, capturing per-rep diagnostics and a mechanism-match summary across the run. Used for statistical reliability characterization of known flake classes under calibrated stress, not as a per-commit gate; not part of ci-local.sh.

sidecar/ -- Network Sidecar Isolation

FIPS running as a sidecar container that owns the network namespace of a companion application container, with iptables/ip6tables rules confining the app to the mesh. scripts/test-sidecar.sh boots a three-node chain of such pairs and asserts both connectivity and isolation.

firewall/ -- nftables Baseline

End-to-end exercise of the production fips0 nftables baseline at packaging/common/fips.nft, covering the default-deny, conntrack and drop-in semantics.

iface-binding/ -- Dynamic Interface Binding

Two nodes whose only transports are interface-bound, started before the interface they name exists. Asserts the boot race (the daemon comes up Degraded and binds when the interface appears, with no restart), the flap (down/up in both directions), destroy-and-recreate, that an optional interface's absence never moves node health, and that absence is logged once on the edge rather than once per retry.

acl-allowlist/ -- Peer ACL Enforcement

Six nodes with per-node allowlist files mounted at the runtime ACL paths, exercising insiders, outsiders and allowed remotes at once to check which peer pairs are admitted and which are rejected.

native-api/ -- Native Datagram API

Checks the experimental native datagram API: a client process opens a flow to a remote pubkey over a Unix socket, receives a file descriptor, and exchanges datagrams on it with no TUN device and no IPv6 emulation.

medium-change/ -- Transport-Medium Change

A multi-homed node whose default route moves between two live access paths while mesh traffic is in flight, with the far peer reachable only through a router so the path to it actually follows that default route. Asserts the peering survives without a re-handshake (link_id and authenticated_at_ms unchanged) and that the far side re-pins to the new source address.

Includes a negative control that runs the same move with node.netmon.enabled: false and requires the outage, so a topology that has stopped exercising the bug fails rather than passing quietly.

dns-resolver/ -- fips-dns-setup Backends

Runs fips-dns-setup against each supported Linux resolver backend in systemd containers, verifying backend detection, generated config and teardown, plus an end-to-end scenario that resolves a .fips name through the configured backend.

deb-install/ -- Debian Package Install

Installs the built .deb in privileged systemd containers for each target distro and verifies unit enablement, conffile placement and end-to-end .fips resolution as a user would meet it.

boringtun/ -- WireGuard Throughput Baseline

Two userspace WireGuard peers running Cloudflare BoringTun, measured with iperf3, as a comparison baseline for FIPS tunnel throughput.

ble/ -- BLE L2CAP Spike

Standalone cargo project (ble_spike) that validates the bluer API assumptions behind the BleIo trait against real adapters on two machines. Not a Docker harness.

Running CI locally (ci-local.sh)

ci-local.sh runs the full local CI pipeline — build, clippy, unit tests, and the integration suites (including the chaos scenarios) — mirroring the GitHub ci.yml integration matrices. Run ./ci-local.sh --help for the full option list and --list for the available suites. Every run starts with a parity check that verifies the local suite set covers the same work as the GitHub matrix, per scenario for chaos and per distro for deb-install, across every job that carries a matrix; a divergence fails the run. GitHub runs the same check as its own ci-parity job. --check-parity runs it alone (see check-ci-parity.sh).

Note that ci-local.sh covers the integration suites and the glibc unit tests. GitHub additionally runs the library tests on macOS, Windows and musl (built for the musl target and run natively), and a --features profiling pass; the musl leg exists because interface presence is built on getifaddrs/ifa_flags, which musl reimplements independently, and OpenWrt is a musl target. A local green run does not certify those four.

The Linux and musl legs also create an address-less dummy interface (fips-probe0) and pass its name to the tests as FIPS_TEST_ADDRLESS_IFACE. That is the one assumption the interface-binding mechanism rests on that no ordinary test can reach: loopback has addresses, so probing it asks whether getifaddrs works rather than whether it reports an interface that has none — which is exactly what fips-mesh0 and fips-ap0 are on OpenWrt. Set the variable by hand to run the assertion locally against an interface you have created; leave it unset and the assertion does not run.

Per-run isolation and the FIPS_CI_RUN_ID override

Every invocation derives a run id and scopes all of its Docker resources to it, so two simultaneous runs on the same host (for example, one per git worktree, or an operator testing by hand while CI is in flight) never collide:

  • Compose projects are named fipsci_<run-id>_<suite>, so container, network, and volume names are all prefixed per run.
  • Build images are tagged fips-test:<run-id> and fips-test-app:<run-id>, exported as FIPS_TEST_IMAGE / FIPS_TEST_APP_IMAGE, and every compose file and suite script reads those. The run does not write fips-test:latest at all: a bridge back to that shared mutable name would let a consumer that had been missed keep working while resolving whichever concurrent run wrote the tag last. :latest stays the hand-build name, produced by testing/scripts/build.sh, and remains the default every consumer falls back to when the variables are unset.
  • The build context is a per-run copy at testing/docker-<run-id>/, exported as FIPS_BUILD_CONTEXT. It is absolute because compose resolves a relative build context against the compose file's own directory rather than the working directory. testing/docker/ is the hand-run context and a CI run does not write to it. Without this, two runs race on the contents of one directory and either can build a correctly-per-run-tagged image from the other's binaries.
  • Each parallel chaos child gets a unique, non-overlapping /24 in 10.30.x (via the sim --subnet override). 10.30.x sits outside Docker's default address pool and the fixed-subnet suites' 172.x ranges, so neither a sibling chaos child nor an auto-assigned network can swallow a pinned subnet.

By default the run id is <short-git-sha>-<random> — the SHA portion records what code a container is testing, the random suffix keeps simultaneous runs of the same SHA disjoint. Override it for a reproducible, attach-by-name debug session:

FIPS_CI_RUN_ID=mydebug ./ci-local.sh --only static-mesh
# containers are named fipsci_mydebug_static_fips-node-a, etc.

Preemption-safety and exit codes

ci-local.sh is safe to cancel mid-run. A signal trap tears down every compose project the run started (not just the current suite) and reaps any in-flight parallel chaos children, bounded by a timeout so a stuck compose down cannot wedge the trap. Exit codes distinguish a cancelled run from a failing one:

Code Meaning
0 all stages passed
1 one or more stages failed
130 interrupted by SIGINT — cancelled, not a failure
143 terminated by SIGTERM — cancelled, not a failure

A preempting CI worker (the push-triggered, CI-gated build pipeline that kills an in-flight run when a newer same-branch tip arrives) maps 130/143 → cancelled (discard, do not record a failing commit), 0 → green, any other non-zero → red.

Cleaning up leftover resources

Every CI-created container, network, and volume carries the label com.corganlabs.fips-ci=1. If a run is hard-killed (SIGKILL, OOM, crash) and leaves resources behind, reap them with:

./ci-local.sh --reap        # or: ./ci-cleanup.sh

ci-cleanup.sh force-removes everything bearing the CI label or a fipsci_ compose-project prefix; it is safe to run when there is nothing to reap and safe to run repeatedly. Pass --project-prefix to scope the sweep to a single run.

It also removes the chaos simulation's leftover host-namespace veth interfaces (vh…a/vh…b), the one resource it touches that is neither a docker object nor labelled — a host interface can carry neither a label nor a compose project, so it is matched by name shape alone. That makes the reach here asymmetric with everything above, and worth stating plainly:

  • A bare chaos.sh run's containers survive a broad reap. Its compose project is not fipsci_, and the simulation labels only the network, not the services.
  • A bare chaos.sh run's veth interfaces do not. An unscoped reap deletes them while they are in use, severing the Ethernet links of a live simulation and leaving its containers running.

So do not run a broad --reap while a bare simulation is up. Scope the interface sweep with --veth-suffixes (which is what ci-local.sh's own teardown passes) or wait for the simulation to finish. --project-prefix does not help here: it scopes only the compose-project sweep.