Retry test, package and NAT lab image builds across transient registry failures

Image builds in CI pull base images from Docker Hub and ghcr.io and fetch
packages from distribution mirrors, and each has failed a leg for a few
seconds at a time: a registry dial timeout, a 502, an apt file truncated
mid-fetch. A single failed build failed the leg, and a rerun of the whole
workflow was the only recovery.

Add retry_build to testing/lib/image-build.sh: up to three attempts with
10 s and then 20 s between them, whole-build so the base is resolved
again and every package-fetching RUN step runs again. It is quiet on a
first-attempt success, prints each failed attempt and a recovery to
stderr, and on GitHub Actions also raises a warning on the run summary
when a build recovered, so recovered failures remain countable. It never
wraps a test or a container start.

Use it for the shared test images in the integration job, the
deb-install and dns-resolver runtime images, and the package builder
image. Each deb-install and dns-resolver attempt prints its captured
output on failure, so a recovered failure still shows its cause.

The NAT, nostr publish-consume and STUN fault suites built their lab
images inside `docker compose up --build`: the relay image from ghcr.io
and Alpine, the STUN server from the Python image, and the routers from
Debian with apt. Build the profile's images first through retry_build,
then bring the lab up with --no-build. The start is not retried, since a
container that fails to start is a test result; only the build is.
This commit is contained in:
Johnathan Corgan
2026-09-19 15:43:27 +00:00
parent d8574ce8de
commit 0c3c0c79c0
8 changed files with 100 additions and 15 deletions
+4 -2
View File
@@ -587,8 +587,10 @@ jobs:
chmod +x _bin/native-echo _bin/native-surface
cp _bin/native-echo testing/docker/native-echo
cp _bin/native-surface testing/docker/native-surface
docker build -t fips-test:latest testing/docker
docker build -t fips-test-app:latest -f testing/docker/Dockerfile.app testing/docker
# Retried: both builds pull from Docker Hub and the Debian mirrors.
source testing/lib/image-build.sh
retry_build "docker build fips-test" docker build -t fips-test:latest testing/docker
retry_build "docker build fips-test-app" docker build -t fips-test-app:latest -f testing/docker/Dockerfile.app testing/docker
# ── Static topology ────────────────────────────────────────────────────
- name: Generate configs (static)
+5 -1
View File
@@ -31,6 +31,8 @@ REPO_ROOT="$(cd "$SCRIPT_DIR/../.." && pwd)"
# shellcheck source=../build-floor.env
. "$REPO_ROOT/packaging/build-floor.env"
# shellcheck source=SCRIPTDIR/../../testing/lib/image-build.sh
. "$REPO_ROOT/testing/lib/image-build.sh"
DEST_DIR="$REPO_ROOT/deploy"
VERSION=""
@@ -115,7 +117,9 @@ fi
if [ "$BUILT_IMAGE" -eq 1 ]; then
echo "=== Building $IMAGE_TAG from $FIPS_BUILD_IMAGE with Rust $RUST_TOOLCHAIN ===" >&2
docker build \
# Retried because the build pulls the floor image and fetches apt packages,
# rustup and crates, any of which can fail for a few seconds at a time.
retry_build "docker build $IMAGE_TAG" docker build \
--build-arg "BASE=$FIPS_BUILD_IMAGE" \
--build-arg "RUST_TOOLCHAIN=$RUST_TOOLCHAIN" \
-t "$IMAGE_TAG" \
+3 -2
View File
@@ -83,8 +83,9 @@ cleanup_container() {
build_image() {
local tag="$1"
shift
echo "$@" | run_quiet "docker build -t $tag" \
docker build -t "$tag" -f - "$REPO_ROOT"
local dockerfile="$*"
retry_build "docker build -t $tag" build_inline "$tag" "$dockerfile" "$REPO_ROOT" || return
return 0
}
# Start the scenario's systemd container. Not privileged: see
+3 -2
View File
@@ -73,8 +73,9 @@ cleanup_container() {
build_image() {
local tag="$1"
shift
echo "$@" | run_quiet "docker build -t $tag" \
docker build -t "$tag" -f - "$REPO_ROOT"
local dockerfile="$*"
retry_build "docker build -t $tag" build_inline "$tag" "$dockerfile" "$REPO_ROOT" || return
return 0
}
# Start a systemd container in the background. Not privileged: see
+56 -3
View File
@@ -1,15 +1,22 @@
#!/bin/bash
# Shared helpers for building test images and starting test containers.
#
# Source this file to get dump_output() and run_quiet():
# Source this file to get dump_output(), run_quiet(), build_inline() and
# retry_build():
# source "$SCRIPT_DIR/../lib/image-build.sh"
# echo "$dockerfile" | run_quiet "docker build -t $tag" \
# docker build -t "$tag" -f - "$REPO_ROOT"
# retry_build "docker build $tag" docker build -t "$tag" "$context"
#
# A build or a container start that fails for a reason outside the project,
# such as a registry timeout, is indistinguishable from one the project caused
# unless its output survives. These helpers keep a command quiet when it
# succeeds and print everything it said when it fails.
# unless its output survives. dump_output and run_quiet keep a command quiet
# when it succeeds and print everything it said when it fails.
#
# Image builds reach Docker Hub, ghcr.io and distribution mirrors, and every one
# of those has timed out or served a truncated file in CI for a few seconds at a
# time. retry_build runs a whole build again after such a failure, so the base
# image is resolved again and each RUN step that fetched packages runs again.
# Emit a captured output file to stderr, delimited and labelled with the
# command it came from. Callers use this only on failure: a command that
@@ -40,3 +47,49 @@ run_quiet() {
rm -f "$out"
return "$rc"
}
# Build TAG from the Dockerfile text DOCKERFILE with context CONTEXT, output
# captured by run_quiet. The text is an argument rather than stdin, so that
# retry_build can run the build again: a pipe can be read only once.
build_inline() {
local tag="$1" dockerfile="$2" context="$3"
echo "$dockerfile" | run_quiet "docker build -t $tag" \
docker build -t "$tag" -f - "$context"
}
# Run an image build, and run it again after a failure, up to three attempts
# with 10 s and then 20 s between them. Use it for builds only: a test or a
# container start that fails is a result, and running it again would hide it.
#
# Every message goes to stderr, because some callers promise that their stdout
# holds only their result. A first-attempt success prints nothing. A success
# after a failure says so, and on GitHub Actions also raises a warning on the
# run summary, so a recovered failure is still counted. Returns the last
# attempt's exit status.
retry_build() {
local label="$1"
shift
local attempts=3 attempt=1 rc=0 wait
while true; do
rc=0
if "$@"; then
if [ "$attempt" -gt 1 ]; then
echo "image-build: $label recovered on attempt $attempt of $attempts" >&2
if [ "${GITHUB_ACTIONS:-}" = "true" ]; then
echo "::warning::image-build: $label recovered on attempt $attempt of $attempts" >&2
fi
fi
return 0
else
rc=$?
fi
if [ "$attempt" -ge "$attempts" ]; then
echo "image-build: $label failed on all $attempts attempts (last exit $rc)" >&2
return "$rc"
fi
wait=$((attempt * 10))
echo "image-build: $label failed on attempt $attempt of $attempts (exit $rc); retrying in ${wait}s" >&2
sleep "$wait"
attempt=$((attempt + 1))
done
}
+15 -3
View File
@@ -10,6 +10,7 @@ GENERATE_SCRIPT="$SCRIPT_DIR/generate-configs.sh"
TOPOLOGY_SCRIPT="$SCRIPT_DIR/setup-topology.sh"
WAIT_LIB="$ROOT_DIR/testing/lib/wait-converge.sh"
RELAY_LIB="$ROOT_DIR/testing/lib/relay-verdict.sh"
IMAGE_LIB="$ROOT_DIR/testing/lib/image-build.sh"
# Must track generate-configs.sh's OUTPUT_DIR and the compose bind-mounts: the
# npubs are read back here after the containers are up, so reading a different
# directory than the one the generator wrote pings an npub no node owns.
@@ -45,6 +46,8 @@ fi
source "$WAIT_LIB"
# shellcheck disable=SC1090
source "$RELAY_LIB"
# shellcheck disable=SC1090
source "$IMAGE_LIB"
RELAY_CONTAINER="fips-nat-relay${FIPS_CI_NAME_SUFFIX:-}"
@@ -500,7 +503,10 @@ run_cone() {
echo "=== NAT lab: cone ==="
cleanup
"$GENERATE_SCRIPT" cone
"${COMPOSE[@]}" --profile cone up -d --build --force-recreate
# Build first, with retries, because the build pulls from registries that
# time out now and then; the start is not retried, since it is the test.
retry_build "compose build (cone)" "${COMPOSE[@]}" --profile cone build
"${COMPOSE[@]}" --profile cone up -d --no-build --force-recreate
"$TOPOLOGY_SCRIPT" cone
wait_for_peers fips-nat-cone-a${FIPS_CI_NAME_SUFFIX:-} 1 45 || {
dump_cone_diagnostics
@@ -548,7 +554,10 @@ run_symmetric() {
echo "=== NAT lab: symmetric fallback ==="
cleanup
NAT_MODE_A=symmetric NAT_MODE_B=symmetric "$GENERATE_SCRIPT" symmetric
NAT_MODE_A=symmetric NAT_MODE_B=symmetric "${COMPOSE[@]}" --profile symmetric up -d --build --force-recreate
# Build first, with retries, because the build pulls from registries that
# time out now and then; the start is not retried, since it is the test.
NAT_MODE_A=symmetric NAT_MODE_B=symmetric retry_build "compose build (symmetric)" "${COMPOSE[@]}" --profile symmetric build
NAT_MODE_A=symmetric NAT_MODE_B=symmetric "${COMPOSE[@]}" --profile symmetric up -d --no-build --force-recreate
"$TOPOLOGY_SCRIPT" symmetric
wait_for_peers fips-nat-symmetric-a${FIPS_CI_NAME_SUFFIX:-} 1 60 || {
dump_symmetric_diagnostics
@@ -598,7 +607,10 @@ run_lan() {
echo "=== NAT lab: lan preference ==="
cleanup
"$GENERATE_SCRIPT" lan
"${COMPOSE[@]}" --profile lan up -d --build --force-recreate
# Build first, with retries, because the build pulls from registries that
# time out now and then; the start is not retried, since it is the test.
retry_build "compose build (lan)" "${COMPOSE[@]}" --profile lan build
"${COMPOSE[@]}" --profile lan up -d --no-build --force-recreate
wait_for_peers fips-nat-lan-a${FIPS_CI_NAME_SUFFIX:-} 1 45 || {
dump_lan_diagnostics
return 1
+7 -1
View File
@@ -22,6 +22,7 @@ BUILD_SCRIPT="$ROOT_DIR/testing/scripts/build.sh"
GENERATE_SCRIPT="$SCRIPT_DIR/generate-configs.sh"
WAIT_LIB="$ROOT_DIR/testing/lib/wait-converge.sh"
RELAY_LIB="$ROOT_DIR/testing/lib/relay-verdict.sh"
IMAGE_LIB="$ROOT_DIR/testing/lib/image-build.sh"
# Must track generate-configs.sh's OUTPUT_DIR and the compose bind-mounts.
CONFIG_DIR="$NAT_DIR/generated-configs${FIPS_CI_NAME_SUFFIX:-}"
@@ -55,6 +56,8 @@ RELAY_CONTAINER="fips-nat-relay${FIPS_CI_NAME_SUFFIX:-}"
source "$WAIT_LIB"
# shellcheck disable=SC1090
source "$RELAY_LIB"
# shellcheck disable=SC1090
source "$IMAGE_LIB"
cleanup() {
"${COMPOSE[@]}" --profile "$PROFILE" down -v --remove-orphans \
@@ -430,7 +433,10 @@ run_test() {
cleanup
"$GENERATE_SCRIPT" "$SCENARIO"
"${COMPOSE[@]}" --profile "$PROFILE" up -d --build --force-recreate
# Build first, with retries, because the build pulls from registries that
# time out now and then; the start is not retried, since it is the test.
retry_build "compose build ($PROFILE)" "${COMPOSE[@]}" --profile "$PROFILE" build
"${COMPOSE[@]}" --profile "$PROFILE" up -d --no-build --force-recreate
# Phase 1 + Phase 2 together: each side publishes its own advert,
# subscribes for the other's, then dials. Bidirectional success
+7 -1
View File
@@ -30,6 +30,7 @@ ROOT_DIR="$(cd "$NAT_DIR/../.." && pwd)"
BUILD_SCRIPT="$ROOT_DIR/testing/scripts/build.sh"
GENERATE_SCRIPT="$SCRIPT_DIR/generate-configs.sh"
RELAY_LIB="$ROOT_DIR/testing/lib/relay-verdict.sh"
IMAGE_LIB="$ROOT_DIR/testing/lib/image-build.sh"
PROFILE="stun-faults"
SCENARIO="$PROFILE"
@@ -73,6 +74,8 @@ cleanup() {
# shellcheck disable=SC1090
source "$RELAY_LIB"
# shellcheck disable=SC1090
source "$IMAGE_LIB"
trap 'echo ""; echo "stun-faults-test interrupted"; cleanup; exit 130' INT TERM
@@ -269,7 +272,10 @@ run_test() {
echo "=== stun-faults-test: setup ==="
cleanup
"$GENERATE_SCRIPT" "$SCENARIO"
"${COMPOSE[@]}" --profile "$PROFILE" up -d --build --force-recreate
# Build first, with retries, because the build pulls from registries that
# time out now and then; the start is not retried, since it is the test.
retry_build "compose build ($PROFILE)" "${COMPOSE[@]}" --profile "$PROFILE" build
"${COMPOSE[@]}" --profile "$PROFILE" up -d --no-build --force-recreate
# Give the daemons time to come up. Both fault-node and fault-peer
# need to start, publish their adverts to the relay, and discover