A CI worker may preempt an in-flight ci-local.sh run (SIGTERM, then SIGKILL after a grace period) to restart on a newer commit. For that kill to be safe, the script must clean up after itself and never let a dying run collide with its restart. It previously had no signal handling, shared the default compose project name across runs, and tore down each suite only at the suite end. - Derive a per-run id (honoring FIPS_CI_RUN_ID, else short-sha+random) and namespace every docker resource to it: a fipsci_<run>_<suite> compose project per suite and per parallel chaos child, and per-run image tags (fips-test:<run>, fips-test-app:<run>) retagged to :latest only after both builds succeed so :latest never points at a half-built image. - Install a bounded, idempotent teardown trap on SIGTERM/SIGINT (+ EXIT): reap parallel chaos children, then force-remove this run's docker resources via the new ci-cleanup.sh, wrapped in timeout so a stuck down cannot wedge it. - Exit 143 (SIGTERM) / 130 (SIGINT), distinct from 0 (pass) / 1 (failed), so a preempting worker tells a cancelled run from a real failure. - Add ci-cleanup.sh (also ci-local.sh --reap): force-removes leftover CI resources by the com.corganlabs.fips-ci=1 label and the fipsci_ project prefix, robust to however a prior run died. - Label every per-suite docker resource so the label sweep reaps it after a SIGKILL regardless of network name: direct docker run/network resources, the sidecar compose services, and every per-suite compose network (acl-allowlist, boringtun, firewall, nat, static, both tor suites, and the chaos generator template). Parametrize the static/sidecar compose image refs so the per-run tags are honored. - Give each parallel chaos child a unique /24 from 10.30.x (a new --subnet override on the sim CLI, assigned per-child in ci-local.sh) so parallel children never collide on a shared docker subnet, and a chaos net can never span a fixed-subnet suite (sidecar/static in 172.20.x). 10.30.x sits outside docker's default-address-pool range, so an auto-assigned net cannot land on it either; node IPs derive from the subnet, so no scenario config changes.
Firewall Baseline Test
End-to-end exercise of the production fips0 nftables baseline at
packaging/common/fips.nft. Closes the v0.3.0 audit gap that the
default-deny + conntrack + drop-in semantics had no integration coverage.
What this exercises
The fips.nft baseline polices ONLY the fips0 mesh interface and
implements default-deny inbound. This suite asserts the four behaviors
documented in the file's header are actually true on a live mesh:
- (a) Unallowed inbound on fips0 is dropped
- (b) Outbound-initiated flows get their reply via the
ct state established,related acceptrule - (c) ICMPv6 echo-request is accepted (ping6 reachability)
- (d) A drop-in
.nftfile under/etc/fips/fips.d/adds an allowlisted port and that port is accepted
A drop-counter check after case (a) confirms the connection was actively DROP'd by the fips chain (not silently unrouted).
Topology
Two FIPS nodes peered over UDP on a Docker bridge network:
| Container | Hostname | docker IPv4 | Firewall |
|---|---|---|---|
fips-fw-container-a |
host-a |
172.32.0.10 | none (probe) |
fips-fw-container-b |
host-b |
172.32.0.11 | fips.nft + drop-in |
node-b mounts the production packaging/common/fips.nft read-only at
/etc/fips/fips.nft, plus a drop-in at /etc/fips/fips.d/services.nft
containing tcp dport 22 accept. node-a is unfirewalled and serves
as the probe origin.
Both containers run the unified test image's default mode, which
starts dnsmasq + sshd (port 22) + iperf3 + python http.server on
port 8000 + the FIPS daemon.
fips-firewall.service activation
The production unit's ExecStart is:
ExecStart=/usr/sbin/nft -f /etc/fips/fips.nft
The unified test image does not run systemd, so test.sh invokes the
same nft -f command directly inside node-b after fips0 is up and
peering has converged. The deb-install harness covers the systemd
unit-enablement path under real systemd separately.
Run
Build the Linux binaries and test image:
./testing/scripts/build.sh --no-docker
Run the suite:
./testing/firewall/test.sh
test.sh regenerates fixtures automatically before starting Docker.
Use --skip-build to reuse the existing release binaries. Use
--keep-up to leave the containers running for inspection.
Expected output shape
=== Generating firewall fixtures
=== Starting firewall harness
=== Waiting for fips0 on both nodes
=== Waiting for peer convergence
=== Resolving fips0 addresses
node-a: fd97:...
node-b: fd97:...
=== Activating fips-firewall on fips-fw-container-b
PASS: fips-fw-container-b: fips.nft baseline + drop-in loaded
=== Case (c): ICMPv6 echo-request to firewalled node
PASS: (c) ICMPv6 ping node-a → node-b accepted
=== Case (a): unallowed inbound TCP/8000 from node-a → node-b
PASS: (a) inbound TCP/8000 blocked (curl rc=28)
=== Case (b): node-b initiates outbound TCP, expects reply via conntrack
PASS: (b) outbound from node-b got HTTP 200 via conntrack reply path
=== Case (d): drop-in allowlisted TCP/22 from node-a → node-b
PASS: (d) drop-in allowlisted TCP/22 reachable
=== Drop counter incremented (case a should have ticked it)
PASS: drop counter = N (case a was actually dropped, not just unrouted)
=== Firewall integration test passed
Inspect the loaded ruleset
docker exec fips-fw-container-b nft list table inet fips
Stop and clean up
docker compose -f testing/firewall/docker-compose.yml down
Generated fixture location
testing/firewall/generated-configs/ (gitignored).