Files
fips/.github/workflows
Johnathan Corgan cd1957c894 Stop using SOCK_SEQPACKET on FreeBSD, which is not a record socket there
The FreeBSD package job had stalled seven consecutive runs, the last two
killed by the thirty-minute bound and the five before it burning up to
six hours each. The job builds and gets into the unit tests; libtest
names the culprit itself in the run log, in a line the earlier reading of
that log missed:

  test native::client::tests::an_empty_datagram_is_not_reported_as_a_closed_flow
  has been running for over 60 seconds

That test sends a zero-length datagram and then reads it back with an
unbounded blocking recv. FreeBSD accepts AF_UNIX SOCK_SEQPACKET and
returns a socket that is not an atomic-record socket: seqpacketproto
carries no PR_ATOMIC and shares its send and receive handlers with
SOCK_STREAM. A zero-length send with no control data therefore queues
nothing, wakes nobody, and returns success, so the recv waits for a
message that was never delivered. Nothing bounds it: plain cargo test has
no per-test deadline, so one blocked thread holds the whole binary open.
The same kernel fact accounts for the five boundary tests that failed
within a third of a second in the same run, which needed no separate
explanation.

macOS already takes SOCK_DGRAM here because it has no AF_UNIX
SOCK_SEQPACKET at all. FreeBSD needs the same substitution for a
different reason, so it joins that arm, along with the close-detection
retry the arm carries. The cfg is written by exclusion so that a new unix
target gets the Linux arm, which fails loudly by refusing the socketpair
rather than quietly losing messages.

Two things guard against a repeat rather than fixing this instance. Every
blocking read in these tests now carries a deadline, so a platform that
swallows a message fails with an assertion naming the flow instead of
wedging the suite. And the FreeBSD job runs its tests under a wall-clock
bound, so a hang is a named step failure in fifteen minutes rather than a
job cancelled at the ceiling, with the output kept up to the kill because
that output is what names the blocked test.

What this does not establish, stated plainly because the code comments
would otherwise imply otherwise: nobody ran any of this on FreeBSD. The
mechanism is read out of the FreeBSD kernel source and inferred from a CI
log whose shape matches it. The freebsd-gated probes added alongside are
what would turn that into a measurement, and the next run of the job is
the first real test.
2026-08-22 12:55:52 +01:00
..
2026-08-11 15:56:49 +00:00