mirror of
https://github.com/jmcorgan/fips.git
synced 2026-10-05 19:18:25 +00:00
Two commands in the chaos harness were trusted without being checked. docker compose down exits 0 while leaving this run's containers alive, so a partly-failed bring-up leaked named containers with nothing to detect it. Teardown now asks whether the containers this run owns are gone. A survivor that the forced removal clears is a warning, since nothing outlives the run; one that survives that, or a query that could not run at all, aborts and writes the names to an artifact. The check is scoped to this run's own names, so a concurrent run cannot trip it. Node churn marked a node down whether or not docker stop worked: the return code was never inspected and the captured output was discarded. The simulation's model of the mesh then diverged from reality, and nodes_down, the max_down_nodes cap and the connectivity guard are all computed from that model, so the guard written to prevent a partition could cause one. A failed stop now warns, carries the daemon's own message, and leaves the node out of the down set, which is the part that matters: the next churn tick simply retries against an honest model. Detection and remedy are kept separate in the code so a later reader can tell which is which.