Files
fips/testing/lib/relay-verdict.sh
Johnathan Corgan 5a5faa0857 test(nat): contain strfry relay aborts in the NAT lab
The NAT-lab suites (cone, symmetric, lan, STUN faults and nostr
publish/consume) share one strfry relay, and it has aborted in several runs.
Three things made each abort cost a whole run and a misdirected diagnosis:
only the publish/consume suite said what state the relay was in, the relay
image moved with upstream's latest tag, and the relay never restarted.

State the relay's condition in every NAT-lab suite. relay_verdict moves into
testing/lib/relay-verdict.sh, taking the container as an argument, and is
called first in every NAT-lab failure dump and on each success path, where a
relay event the assertions survived is noted rather than made a failure.
Before this, the cone, symmetric and lan dumps named the nodes and the
network instead: hundreds of lines of socket tables and router counters,
with the relay's crash lines unlabelled in an 80-line log tail. The STUN
fault dump now also carries the relay's log, which it did not include.

Pin the relay's strfry build. Dockerfile.app takes the strfry image as a
build argument, defaulting to the same latest tag, so the user-facing example
is unchanged, and the NAT lab's compose file passes the current multi-arch
index digest. Which build a run exercised was never recorded before. This
stops the drift; it does not select a build that does not abort, since
upstream publishes no other tag to choose from.

Restart the relay when it aborts. The relay ran with restart "no", so one
abort left every suite sharing it without a relay for the rest of the run.
It now restarts on failure up to three times, so a relay that keeps aborting
still ends up exited and is reported rather than hidden in a crash loop. The
restart policy is the containment. The verdict says the relay restarted
under that policy and prints the fault lines from its log, which spans
restarts of the same container, so a rescued run still names the event.

The relay service also gains init: true, only so the restart can be
exercised by an injected abort. It changes the relay's PID 1 from strfry
(started with exec in the image's entrypoint) to docker's init, with strfry
as its child. Without it, a SIGABRT sent with docker kill to strfry as PID 1
logged "caught a signal: SIGABRT" and left the container running with no
restart, so that injection could not show the policy working. With the init,
strfry signalled from inside the container exits it non-zero, which is what
the real aborts did: clients saw the relay vanish. Those real aborts would
restart under the policy with or without the init.

Injected aborts (strfry signalled from inside the container the moment both
cone nodes had connected) restarted the relay once each time; in eight of
nine cone runs both nodes reconnected and peered 8 to 16 s after the abort.
In the ninth, the initiator's offer went out in the second between its
reconnect and the responder's, was lost, and the 30 s answer timeout pushed
peering past the 45 s wait, so restart shortens the outage but does not
guarantee the run. Without the restart policy the same injection left the
relay exited and the cone suite timed out waiting for its peer, with no
line in the dump naming the relay.
2026-09-19 09:11:46 +00:00

86 lines
3.4 KiB
Bash

#!/bin/bash
# Shared relay verdict for the NAT-lab suites.
#
# Source this file to get relay_verdict().
#
# Usage:
# source "$ROOT_DIR/testing/lib/relay-verdict.sh"
# relay_verdict <relay-container>
#
# The lab's Nostr relay is a third-party container (strfry), and it has died
# mid-run: once on a SIGSEGV twelve milliseconds after both nodes had
# connected, and several times since on an internal assertion that aborts it
# the instant two clients connect together. Every symptom under that is a
# correct report of a dead relay: subscriptions dropped, no advert consumed,
# empty peer lists, and a run that fails on the peer-count wait. That verdict
# is indistinguishable from a product failure unless the relay's state is
# stated, and the only evidence of the real cause was one container-log line
# far above the summary. relay_verdict says whose failure it was, in a line a
# reader of the output can find.
# State the relay's own condition, in its own words.
#
# This does not decide the run. It returns 0 when the run had a relay event
# (the relay is gone, not running, restarted, faulted, or its health could not
# be read) and 1 when the relay was running with no fault in its log. Callers
# use it first in a failure dump, and on the success path to note a relay
# event the assertions survived.
#
# A relay whose state or log cannot be read has not been shown to be healthy,
# so that case is reported as unestablished rather than as "not the relay".
relay_verdict() {
local container="$1"
local state="" status="" exit_code="" restarts="" logs=""
local faults="caught a signal|SIGSEGV|SIGABRT|terminate called"
state="$(docker inspect \
-f '{{.State.Status}} {{.State.ExitCode}} {{.RestartCount}}' \
"$container" 2>/dev/null)" || state=""
if [ -z "$state" ]; then
echo "RELAY FAILURE: $container is gone; whatever this run" \
"asserted about peering happened without a relay" >&2
return 0
fi
read -r status exit_code restarts <<<"$state"
if [ "$status" != "running" ]; then
echo "RELAY FAILURE: $container is $status (exit $exit_code);" \
"the assertions in this run are downstream of that, not of the" \
"nodes" >&2
return 0
fi
# docker logs spans every restart of the same container, so the fault
# lines of the run that died are still readable here.
if [ "${restarts:-0}" -gt 0 ]; then
echo "RELAY FAILURE: $container restarted $restarts time(s) during" \
"this run under its restart policy; the nodes lost their" \
"subscriptions each time and had to reconnect" >&2
if logs="$(docker logs "$container" 2>&1)"; then
grep -E "$faults" <<<"$logs" >&2 || true
else
echo " (its logs could not be read, so the fault lines are" \
"not shown)" >&2
fi
return 0
fi
if ! logs="$(docker logs "$container" 2>&1)"; then
echo "RELAY HEALTH NOT ESTABLISHED: could not read" \
"$container's logs, so a crash cannot be ruled in or out" >&2
return 0
fi
if grep -Eq "$faults" <<<"$logs"; then
echo "RELAY FAILURE: $container faulted during this run:" >&2
grep -E "$faults" <<<"$logs" >&2
return 0
fi
echo "relay: $container running, no fault in its log — this run's" \
"verdict is about the nodes"
return 1
}