mirror of
https://github.com/jmcorgan/fips.git
synced 2026-10-05 11:08:25 +00:00
The NAT-lab suites (cone, symmetric, lan, STUN faults and nostr publish/consume) share one strfry relay, and it has aborted in several runs. Three things made each abort cost a whole run and a misdirected diagnosis: only the publish/consume suite said what state the relay was in, the relay image moved with upstream's latest tag, and the relay never restarted. State the relay's condition in every NAT-lab suite. relay_verdict moves into testing/lib/relay-verdict.sh, taking the container as an argument, and is called first in every NAT-lab failure dump and on each success path, where a relay event the assertions survived is noted rather than made a failure. Before this, the cone, symmetric and lan dumps named the nodes and the network instead: hundreds of lines of socket tables and router counters, with the relay's crash lines unlabelled in an 80-line log tail. The STUN fault dump now also carries the relay's log, which it did not include. Pin the relay's strfry build. Dockerfile.app takes the strfry image as a build argument, defaulting to the same latest tag, so the user-facing example is unchanged, and the NAT lab's compose file passes the current multi-arch index digest. Which build a run exercised was never recorded before. This stops the drift; it does not select a build that does not abort, since upstream publishes no other tag to choose from. Restart the relay when it aborts. The relay ran with restart "no", so one abort left every suite sharing it without a relay for the rest of the run. It now restarts on failure up to three times, so a relay that keeps aborting still ends up exited and is reported rather than hidden in a crash loop. The restart policy is the containment. The verdict says the relay restarted under that policy and prints the fault lines from its log, which spans restarts of the same container, so a rescued run still names the event. The relay service also gains init: true, only so the restart can be exercised by an injected abort. It changes the relay's PID 1 from strfry (started with exec in the image's entrypoint) to docker's init, with strfry as its child. Without it, a SIGABRT sent with docker kill to strfry as PID 1 logged "caught a signal: SIGABRT" and left the container running with no restart, so that injection could not show the policy working. With the init, strfry signalled from inside the container exits it non-zero, which is what the real aborts did: clients saw the relay vanish. Those real aborts would restart under the policy with or without the init. Injected aborts (strfry signalled from inside the container the moment both cone nodes had connected) restarted the relay once each time; in eight of nine cone runs both nodes reconnected and peered 8 to 16 s after the abort. In the ninth, the initiator's offer went out in the second between its reconnect and the responder's, was lost, and the 30 s answer timeout pushed peering past the 45 s wait, so restart shortens the outage but does not guarantee the run. Without the restart policy the same injection left the relay exited and the cone suite timed out waiting for its peer, with no line in the dump naming the relay.
3.3 KiB
3.3 KiB
NAT Lab Harness
Real Docker-based NAT traversal integration tests for the mainline FIPS Nostr/STUN bootstrap path.
This harness spins up:
- two FIPS nodes
- a local Nostr relay
- a local STUN server
- one or two Linux router containers performing NAT with
iptables
For the NAT scenarios, the node LAN interfaces are not attached to
Docker bridge networks. The harness creates explicit veth pairs and
moves them into the node and router namespaces after docker compose up
so every packet must traverse the router namespace.
It covers three scenarios:
cone: both peers behind explicit namespace/veth full-cone emulation, UDP traversal succeedssymmetric: both peers behind symmetric-style NAT, UDP traversal fails, TCP fallback succeedslan: both peers share a LAN subnet, LAN targets are preferred over reflexive addresses
NAT model notes
The harness does not rely on plain Docker MASQUERADE for the cone case.
cone- uses explicit full-cone emulation in the router namespace
- outbound UDP is
SNATed to the router WAN address while preserving the source port - inbound UDP to the router WAN address is
DNATed back to the single LAN host regardless of remote source
symmetric- uses UDP
MASQUERADE --random-fully - outbound mappings may be port-randomized and are only reopened by matching conntrack state
- uses UDP
This distinction matters because plain MASQUERADE is convenient source NAT, but it does not by itself model the "accept from any remote once mapped" behavior expected from a full-cone NAT.
Prerequisites
- Docker with Compose support
- locally built
fips-test:latest
Build the test image with:
./testing/scripts/build.sh
Run
Run all scenarios:
./testing/nat/scripts/nat-test.sh
Run one scenario:
./testing/nat/scripts/nat-test.sh cone
./testing/nat/scripts/nat-test.sh symmetric
./testing/nat/scripts/nat-test.sh lan
Layout
docker-compose.yml- relay/STUN/WAN topology plus container definitions
node/- node bootstrap wrapper that waits for the injected veth interface
router/- NAT router image and
iptablessetup
- NAT router image and
stun/- minimal STUN binding responder
relay/- local
strfryconfig. The relay image's strfry build is pinned by digest indocker-compose.yml(STRFRY_IMAGE); bump it there deliberately.
- local
scripts/generate-configs.sh- derives ephemeral identities and writes per-scenario FIPS configs
scripts/setup-topology.sh- injects and configures the NAT LAN
vethpairs in the container namespaces
- injects and configures the NAT LAN
scripts/nat-test.sh- boots the lab, waits for convergence, and asserts the resulting path
scripts/nostr-relay-test.sh- exercises the Nostr overlay advert publish/consume round-trip, including rejection of a malformed advert event
scripts/stun-faults-test.sh- cycles the daemon through STUN drop, delay and outage faults and asserts graceful behavior at each step
Assertions
-
cone- both nodes connect
- connected transport is UDP
- active link remote addresses are on the WAN NAT subnet
-
symmetric- NAT bootstrap does not establish a UDP link
- fallback converges
- connected transport is TCP via router-published WAN addresses
-
lan- both nodes connect
- connected transport is UDP
- active link remote addresses stay on the shared LAN subnet