mirror of
https://github.com/jmcorgan/fips.git
synced 2026-10-06 03:28:24 +00:00
churn-mixed-10 failed its baseline with "9 node(s) answered (need 10)" when the host was busy. The node missing from the final snapshot had just been restarted by the teardown restore, and the snapshot queried it before its control socket was open. In every churn-mixed-10 run the scenario's seed restarts n01 at that point; its control socket opened 2.7-5.5 s after the restart across 21 runs at three load levels, while the snapshot followed the restore by 2-6 s. So the suite passed or failed on which came first, and load only made the restart slower. Reproduced once in 17 runs with 24 CPU hogs plus I/O on a 12-CPU host. Teardown now waits, after restoring nodes, until every node's control socket answers, bounded at 60 s (11 times the slowest start measured), and then takes the snapshot. The floor of 10 answering nodes is unchanged. A node that never answers is still reported absent: with one node stopped for good just before the wait, the wait gave up after 61.7 s and the baseline failed with 9 answered. Five more runs of the chaos set under the same load were green, the wait taking 1.6-3.6 s.