Files
fips/testing/openwrt/maintainer-scripts-test.sh
T
Johnathan Corgan 87f0678ce0 Point dnsmasq at the OpenWrt gateway only while its own process is listening
start_service switched dnsmasq's .fips forwarding to the gateway's port
before the gateway ran. A gateway that then failed, on a config it could
not parse or on a NAT or route error after binding, left every LAN
client's .fips lookups going to a port nothing held, until someone
started the gateway by hand; procd gives up respawning after five
failures and says nothing.

procd now runs the init script's supervise command, which starts the
gateway, switches dnsmasq to its port once the gateway's own process holds
a socket bound there, and switches it back to the daemon's port when the
gateway exits for any reason, passing procd's SIGTERM on so the gateway can
clean up. Ownership is checked by matching a socket inode for the port in
/proc/net/udp{,6} against the gateway's open descriptors, so a port held by
something else, such as an mDNS responder on 5353, is never mistaken for the
gateway. supervise records the gateway's pid in /var/run/fips-gateway.pid,
and an exiting instance skips the swap back only when the gateway named
there, the one procd started in its place, holds the port. Every change to
the dnsmasq entry is made under one lock, so a stopping and a starting
gateway cannot interleave. stop_service still swaps back as before. The
exit status is the gateway's, so procd's respawn behaves as it did.

procd sends SIGTERM to the supervise shell and, five seconds later, SIGKILL
to that shell alone, so a gateway slow to stop would outlive it, holding its
port. supervise kills the gateway three seconds after passing the SIGTERM
on, and stops its watcher as soon as the gateway is gone.

A new scenario runs the command start_service hands procd against stub
gateways that hold their socket themselves and exit before binding, fail
after binding, are stopped by SIGTERM, or ignore it, against a port taken by
another process, and against a swap back with a process other than the
successor holding the port, and checks dnsmasq ends on the daemon's port
each time.
2026-10-01 14:21:40 +00:00

69 lines
3.1 KiB
Bash
Executable File

#!/bin/bash
# ── OpenWrt maintainer-script scenarios ─────────────────────────────────────
# Runs testing/openwrt/scenarios.sh inside a busybox container, so the package
# scripts and the fips-gateway init script are interpreted by ash rather than
# by the host's bash or dash. The scripts ship to routers and are only ever run
# under ash there; a construct bash accepts and ash does not would otherwise
# surface on a router.
#
# The container is the only reason docker is needed: the scenarios touch no
# network and no FIPS binary, and they do not use the shared test image.
#
# Exit 0 = every scenario passed. Exit 1 = at least one failed. Exit 2 = the
# harness could not run; never treated as a pass.
# ─────────────────────────────────────────────────────────────────────────────
set -uo pipefail
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
PROJECT_ROOT="$(cd "$SCRIPT_DIR/../.." && pwd)"
# Pinned by digest so the shell under test does not change under a run: the
# 1.37 tag moves with every 1.37.x rebuild. This is the multi-arch index digest
# of busybox:1.37.0 as of 2026-10-01. Bump it deliberately, reading the new
# digest with `docker buildx imagetools inspect busybox:<version>`.
# Overridable for trying another ash build.
IMAGE="${OPENWRT_ASH_IMAGE:-busybox:1.37.0@sha256:bdf57e528e45e4433820e045b29b4597825a1c9e38353532d90a01445013f82e}"
if ! command -v docker >/dev/null 2>&1; then
echo "openwrt-scripts: docker not found; cannot run the ash scenarios" >&2
exit 2
fi
if [[ ! -f "$SCRIPT_DIR/scenarios.sh" ]]; then
echo "openwrt-scripts: missing $SCRIPT_DIR/scenarios.sh" >&2
exit 2
fi
# The .apk wraps the shared bodies for its upgrade path. package-test.sh builds
# the package on the host with the real build-apk.sh, checks what it registers,
# and leaves the four scripts here so the scenarios run exactly what ships.
# The directory is bind-mounted into the container, so it must be one the docker
# daemon can see: under the checkout, not /tmp, which a service running with a
# private /tmp (as the CI workers do) does not share with the daemon.
mkdir -p "$PROJECT_ROOT/target" || { echo "openwrt-scripts: cannot create target/" >&2; exit 2; }
APK_DIR="$(mktemp -d "$PROJECT_ROOT/target/openwrt-apk.XXXXXX")" || { echo "openwrt-scripts: mktemp failed" >&2; exit 2; }
trap 'rm -rf "$APK_DIR"' EXIT
bash "$SCRIPT_DIR/package-test.sh" --keep "$APK_DIR"
rc=$?
if [[ $rc -ne 0 ]]; then
echo "openwrt-scripts: package-test.sh exited $rc" >&2
exit $rc
fi
docker run --rm --network none \
-v "$PROJECT_ROOT:/src:ro" \
-v "$APK_DIR:/apk:ro" \
-e REPO=/src \
-e APK_SCRIPTS=/apk \
-e "POSTINST=${POSTINST:-}" \
-e "PRERM=${PRERM:-}" \
-e "INIT_GATEWAY=${INIT_GATEWAY:-}" \
"$IMAGE" sh /src/testing/openwrt/scenarios.sh
rc=$?
if [[ $rc -ne 0 && $rc -ne 1 ]]; then
echo "openwrt-scripts: the container exited $rc, so the scenarios did not report" >&2
exit 2
fi
exit $rc