mirror of
https://github.com/jmcorgan/fips.git
synced 2026-10-05 11:08:25 +00:00
Several local CI reds have come from suites whose timing floors are sensitive to host contention, typically when two runs overlapped. The summary gave no way to tell such a red from any other, so each one had to be investigated from scratch. ci-local.sh now samples CPU pressure (/proc/pressure/cpu, "some" stall time) and the number of other CI runs' containers every 5 s for the whole run, and writes a begin/end window around each suite: one end line per sequential suite from record(), a begin/end pair per chaos scenario from run_chaos (which runs once per scenario on the parallel and the --only path), and a barrier after the chaos block. testing/lib/load-annotate.py reads that log and prints, under each failed suite, the stall over its window, the peak number of foreign runs, and LOAD-DEGRADED at or above 15%, not load-degraded below it, or load unknown when the window or the samples are missing. A run-level line gives the peak stall and peak foreign runs. The annotation changes no verdict and no exit code. 15% comes from measurement on the 12-CPU host: green chaos sets with no foreign load ran at 4-11% stall, loaded sets at 39-70%, and the one reproduced load-sensitive red at 63%. Setting FIPS_CI_LOAD_LOG keeps the sample log after the run; otherwise teardown removes it.