Files
fips/testing/chaos/sim
Johnathan Corgan 7e9ad6e213 Check that teardown and node stops actually did what they report
Two commands in the chaos harness were trusted without being checked.

docker compose down exits 0 while leaving this run's containers alive, so
a partly-failed bring-up leaked named containers with nothing to detect
it. Teardown now asks whether the containers this run owns are gone. A
survivor that the forced removal clears is a warning, since nothing
outlives the run; one that survives that, or a query that could not run
at all, aborts and writes the names to an artifact. The check is scoped
to this run's own names, so a concurrent run cannot trip it.

Node churn marked a node down whether or not docker stop worked: the
return code was never inspected and the captured output was discarded.
The simulation's model of the mesh then diverged from reality, and
nodes_down, the max_down_nodes cap and the connectivity guard are all
computed from that model, so the guard written to prevent a partition
could cause one. A failed stop now warns, carries the daemon's own
message, and leaves the node out of the down set, which is the part that
matters: the next churn tick simply retries against an honest model.

Detection and remedy are kept separate in the code so a later reader can
tell which is which.
2026-08-22 09:19:04 +01:00
..