Files
Johnathan Corgan fa5c7aaadf Show the host load behind each failed suite in the local CI summary
Several local CI reds have come from suites whose timing floors are
sensitive to host contention, typically when two runs overlapped. The
summary gave no way to tell such a red from any other, so each one had to
be investigated from scratch.

ci-local.sh now samples CPU pressure (/proc/pressure/cpu, "some" stall
time) and the number of other CI runs' containers every 5 s for the whole
run, and writes a begin/end window around each suite: one end line per
sequential suite from record(), a begin/end pair per chaos scenario from
run_chaos (which runs once per scenario on the parallel and the --only
path), and a barrier after the chaos block. testing/lib/load-annotate.py
reads that log and prints, under each failed suite, the stall over its
window, the peak number of foreign runs, and LOAD-DEGRADED at or above
15%, not load-degraded below it, or load unknown when the window or the
samples are missing. A run-level line gives the peak stall and peak
foreign runs. The annotation changes no verdict and no exit code.

15% comes from measurement on the 12-CPU host: green chaos sets with no
foreign load ran at 4-11% stall, loaded sets at 39-70%, and the one
reproduced load-sensitive red at 63%. Setting FIPS_CI_LOAD_LOG keeps the
sample log after the run; otherwise teardown removes it.
2026-09-19 17:21:45 +00:00
..