Files
fips/testing/nat/docker-compose.yml
Johnathan Corgan 5a5faa0857 test(nat): contain strfry relay aborts in the NAT lab
The NAT-lab suites (cone, symmetric, lan, STUN faults and nostr
publish/consume) share one strfry relay, and it has aborted in several runs.
Three things made each abort cost a whole run and a misdirected diagnosis:
only the publish/consume suite said what state the relay was in, the relay
image moved with upstream's latest tag, and the relay never restarted.

State the relay's condition in every NAT-lab suite. relay_verdict moves into
testing/lib/relay-verdict.sh, taking the container as an argument, and is
called first in every NAT-lab failure dump and on each success path, where a
relay event the assertions survived is noted rather than made a failure.
Before this, the cone, symmetric and lan dumps named the nodes and the
network instead: hundreds of lines of socket tables and router counters,
with the relay's crash lines unlabelled in an 80-line log tail. The STUN
fault dump now also carries the relay's log, which it did not include.

Pin the relay's strfry build. Dockerfile.app takes the strfry image as a
build argument, defaulting to the same latest tag, so the user-facing example
is unchanged, and the NAT lab's compose file passes the current multi-arch
index digest. Which build a run exercised was never recorded before. This
stops the drift; it does not select a build that does not abort, since
upstream publishes no other tag to choose from.

Restart the relay when it aborts. The relay ran with restart "no", so one
abort left every suite sharing it without a relay for the rest of the run.
It now restarts on failure up to three times, so a relay that keeps aborting
still ends up exited and is reported rather than hidden in a crash loop. The
restart policy is the containment. The verdict says the relay restarted
under that policy and prints the fault lines from its log, which spans
restarts of the same container, so a rescued run still names the event.

The relay service also gains init: true, only so the restart can be
exercised by an injected abort. It changes the relay's PID 1 from strfry
(started with exec in the image's entrypoint) to docker's init, with strfry
as its child. Without it, a SIGABRT sent with docker kill to strfry as PID 1
logged "caught a signal: SIGABRT" and left the container running with no
restart, so that injection could not show the policy working. With the init,
strfry signalled from inside the container exits it non-zero, which is what
the real aborts did: clients saw the relay vanish. Those real aborts would
restart under the policy with or without the init.

Injected aborts (strfry signalled from inside the container the moment both
cone nodes had connected) restarted the relay once each time; in eight of
nine cone runs both nodes reconnected and peered 8 to 16 s after the abort.
In the ninth, the initiator's offer went out in the second between its
reconnect and the responder's, was lost, and the 30 s answer timeout pushed
peering past the 45 s wait, so restart shortens the outage but does not
guarantee the run. Without the restart policy the same injection left the
relay exited and the cone suite timed out waiting for its peer, with no
line in the dump naming the relay.
2026-09-19 09:11:46 +00:00

361 lines
12 KiB
YAML

networks:
wan:
driver: bridge
labels:
- "com.corganlabs.fips-ci=1"
ipam:
config:
- subnet: ${NAT_WAN_PREFIX:-172.31.254}.0/24
shared-lan:
driver: bridge
labels:
- "com.corganlabs.fips-ci=1"
ipam:
config:
- subnet: ${NAT_LAN_PREFIX:-172.31.10}.0/24
volumes:
relay-data:
x-fips-common: &fips-common
image: ${FIPS_TEST_IMAGE:-fips-test:latest}
cap_add:
- NET_ADMIN
devices:
- /dev/net/tun:/dev/net/tun
sysctls:
- net.ipv6.conf.all.disable_ipv6=0
restart: "no"
environment:
- RUST_LOG=info,fips::nostr=debug,fips::node::lifecycle=debug
services:
relay:
build:
context: ../..
dockerfile: examples/sidecar-nostr-relay/Dockerfile.app
args:
# Pinned so the relay build does not move underneath the harness.
# Upstream publishes only `latest`; this is its multi-arch index
# digest as of 2026-09-19 (amd64 and arm64 manifests built
# 2026-09-04). Bump it deliberately, reading the new digest with
# `docker buildx imagetools inspect ghcr.io/hoytech/strfry:latest`.
STRFRY_IMAGE: ghcr.io/hoytech/strfry@sha256:36f1886d185a88ca57c66ebe52e6e9e8428dac2486eea0a5d50ff934f18b60c3
container_name: fips-nat-relay${FIPS_CI_NAME_SUFFIX:-}
# A third-party relay abort should not fail a whole run, so the relay comes
# back on its own; relay_verdict still reports the restart and the fault
# lines. Bounded so a relay that aborts repeatedly ends up exited, and the
# verdict says so, rather than a crash loop hiding it.
#
# `init` is here so the restart can be tested, not for the restart itself.
# It makes docker's init PID 1, with strfry as its child, where strfry was
# PID 1 before. A test SIGABRT sent with `docker kill` to strfry as PID 1
# ran its handler and left the container running, so it never exercised
# the restart. With an init, strfry signalled from inside the container
# exits it non-zero, as the relay aborts seen in the lab did.
restart: on-failure:3
init: true
volumes:
- relay-data:/usr/src/app/strfry-db
- ./relay/strfry.conf:/usr/src/app/strfry.conf:ro
- ../docker/resolv.conf:/etc/resolv.conf:ro
networks:
wan:
ipv4_address: ${NAT_WAN_PREFIX:-172.31.254}.30
shared-lan:
ipv4_address: ${NAT_LAN_PREFIX:-172.31.10}.30
stun:
build:
context: ./stun
container_name: fips-nat-stun${FIPS_CI_NAME_SUFFIX:-}
restart: "no"
networks:
wan:
ipv4_address: ${NAT_WAN_PREFIX:-172.31.254}.40
shared-lan:
ipv4_address: ${NAT_LAN_PREFIX:-172.31.10}.40
nat-a:
build:
context: ./router
profiles: ["cone", "symmetric"]
container_name: fips-nat-router-a${FIPS_CI_NAME_SUFFIX:-}
cap_add:
- NET_ADMIN
sysctls:
- net.ipv4.ip_forward=1
restart: "no"
environment:
- NAT_MODE=${NAT_MODE_A:-cone}
- TCP_FORWARD_PORTS=8443
- LAN_IF=eth1
- WAN_IF=eth0
- LAN_HOST=172.31.1.10
- LAN_SUBNET=172.31.1.0/24
- WAN_SUBNET=${NAT_WAN_PREFIX:-172.31.254}.0/24
- WAN_GATEWAY=${NAT_WAN_PREFIX:-172.31.254}.1
networks:
wan:
ipv4_address: ${NAT_WAN_PREFIX:-172.31.254}.10
nat-b:
build:
context: ./router
profiles: ["cone", "symmetric"]
container_name: fips-nat-router-b${FIPS_CI_NAME_SUFFIX:-}
cap_add:
- NET_ADMIN
sysctls:
- net.ipv4.ip_forward=1
restart: "no"
environment:
- NAT_MODE=${NAT_MODE_B:-cone}
- TCP_FORWARD_PORTS=8443
- LAN_IF=eth1
- WAN_IF=eth0
- LAN_HOST=172.31.2.10
- LAN_SUBNET=172.31.2.0/24
- WAN_SUBNET=${NAT_WAN_PREFIX:-172.31.254}.0/24
- WAN_GATEWAY=${NAT_WAN_PREFIX:-172.31.254}.1
networks:
wan:
ipv4_address: ${NAT_WAN_PREFIX:-172.31.254}.11
cone-a:
<<: *fips-common
profiles: ["cone"]
container_name: fips-nat-cone-a${FIPS_CI_NAME_SUFFIX:-}
hostname: fips-nat-cone-a
depends_on:
- nat-a
- relay
- stun
entrypoint:
- /usr/local/bin/nat-node-entrypoint.sh
environment:
- RUST_LOG=info,fips::nostr=debug,fips::node::lifecycle=debug
- DATA_IF=eth0
- ROUTE_SUBNET=${NAT_WAN_PREFIX:-172.31.254}.0/24
- ROUTE_VIA=172.31.1.254
- RELAY_HOST=${NAT_WAN_PREFIX:-172.31.254}.30
- RELAY_PORT=7777
- STUN_HOST=${NAT_WAN_PREFIX:-172.31.254}.40
- STUN_PORT=3478
volumes:
- ../docker/resolv.conf:/etc/resolv.conf:ro
- ./node/entrypoint.sh:/usr/local/bin/nat-node-entrypoint.sh:ro
- ./generated-configs${FIPS_CI_NAME_SUFFIX:-}/cone/node-a.yaml:/etc/fips/fips.yaml:ro
network_mode: none
cone-b:
<<: *fips-common
profiles: ["cone"]
container_name: fips-nat-cone-b${FIPS_CI_NAME_SUFFIX:-}
hostname: fips-nat-cone-b
depends_on:
- nat-b
- relay
- stun
entrypoint:
- /usr/local/bin/nat-node-entrypoint.sh
environment:
- RUST_LOG=info,fips::nostr=debug,fips::node::lifecycle=debug
- DATA_IF=eth0
- ROUTE_SUBNET=${NAT_WAN_PREFIX:-172.31.254}.0/24
- ROUTE_VIA=172.31.2.254
- RELAY_HOST=${NAT_WAN_PREFIX:-172.31.254}.30
- RELAY_PORT=7777
- STUN_HOST=${NAT_WAN_PREFIX:-172.31.254}.40
- STUN_PORT=3478
volumes:
- ../docker/resolv.conf:/etc/resolv.conf:ro
- ./node/entrypoint.sh:/usr/local/bin/nat-node-entrypoint.sh:ro
- ./generated-configs${FIPS_CI_NAME_SUFFIX:-}/cone/node-b.yaml:/etc/fips/fips.yaml:ro
network_mode: none
symmetric-a:
<<: *fips-common
profiles: ["symmetric"]
container_name: fips-nat-symmetric-a${FIPS_CI_NAME_SUFFIX:-}
hostname: fips-nat-symmetric-a
depends_on:
- nat-a
- relay
- stun
entrypoint:
- /usr/local/bin/nat-node-entrypoint.sh
environment:
- RUST_LOG=info,fips::nostr=debug,fips::node::lifecycle=debug
- DATA_IF=eth0
- ROUTE_SUBNET=${NAT_WAN_PREFIX:-172.31.254}.0/24
- ROUTE_VIA=172.31.1.254
- RELAY_HOST=${NAT_WAN_PREFIX:-172.31.254}.30
- RELAY_PORT=7777
- STUN_HOST=${NAT_WAN_PREFIX:-172.31.254}.40
- STUN_PORT=3478
volumes:
- ../docker/resolv.conf:/etc/resolv.conf:ro
- ./node/entrypoint.sh:/usr/local/bin/nat-node-entrypoint.sh:ro
- ./generated-configs${FIPS_CI_NAME_SUFFIX:-}/symmetric/node-a.yaml:/etc/fips/fips.yaml:ro
network_mode: none
symmetric-b:
<<: *fips-common
profiles: ["symmetric"]
container_name: fips-nat-symmetric-b${FIPS_CI_NAME_SUFFIX:-}
hostname: fips-nat-symmetric-b
depends_on:
- nat-b
- relay
- stun
entrypoint:
- /usr/local/bin/nat-node-entrypoint.sh
environment:
- RUST_LOG=info,fips::nostr=debug,fips::node::lifecycle=debug
- DATA_IF=eth0
- ROUTE_SUBNET=${NAT_WAN_PREFIX:-172.31.254}.0/24
- ROUTE_VIA=172.31.2.254
- RELAY_HOST=${NAT_WAN_PREFIX:-172.31.254}.30
- RELAY_PORT=7777
- STUN_HOST=${NAT_WAN_PREFIX:-172.31.254}.40
- STUN_PORT=3478
volumes:
- ../docker/resolv.conf:/etc/resolv.conf:ro
- ./node/entrypoint.sh:/usr/local/bin/nat-node-entrypoint.sh:ro
- ./generated-configs${FIPS_CI_NAME_SUFFIX:-}/symmetric/node-b.yaml:/etc/fips/fips.yaml:ro
network_mode: none
lan-a:
<<: *fips-common
profiles: ["lan"]
container_name: fips-nat-lan-a${FIPS_CI_NAME_SUFFIX:-}
hostname: fips-nat-lan-a
depends_on:
- relay
- stun
volumes:
- ../docker/resolv.conf:/etc/resolv.conf:ro
- ./generated-configs${FIPS_CI_NAME_SUFFIX:-}/lan/node-a.yaml:/etc/fips/fips.yaml:ro
networks:
shared-lan:
ipv4_address: ${NAT_LAN_PREFIX:-172.31.10}.10
lan-b:
<<: *fips-common
profiles: ["lan"]
container_name: fips-nat-lan-b${FIPS_CI_NAME_SUFFIX:-}
hostname: fips-nat-lan-b
depends_on:
- relay
- stun
volumes:
- ../docker/resolv.conf:/etc/resolv.conf:ro
- ./generated-configs${FIPS_CI_NAME_SUFFIX:-}/lan/node-b.yaml:/etc/fips/fips.yaml:ro
networks:
shared-lan:
ipv4_address: ${NAT_LAN_PREFIX:-172.31.10}.11
# ── Nostr publish/consume profile ──────────────────────────────────────
# Two FIPS daemons + the existing strfry relay, exercising the overlay
# advert publish → relay → consumer round-trip end-to-end. Both nodes
# share the same LAN bridge as the relay (no NAT in the way) so the
# focus of the test is the Nostr discovery layer rather than NAT
# traversal mechanics. Phase 3 (malformed advert) is driven by a
# one-shot publish from the test runner via the relay's WebSocket.
nostr-pub-a:
<<: *fips-common
profiles: ["nostr-publish-consume"]
container_name: fips-nat-nostr-pub-a${FIPS_CI_NAME_SUFFIX:-}
hostname: fips-nat-nostr-pub-a
depends_on:
- relay
- stun
volumes:
- ../docker/resolv.conf:/etc/resolv.conf:ro
- ./generated-configs${FIPS_CI_NAME_SUFFIX:-}/nostr-publish-consume/node-a.yaml:/etc/fips/fips.yaml:ro
networks:
shared-lan:
ipv4_address: ${NAT_LAN_PREFIX:-172.31.10}.20
nostr-pub-b:
<<: *fips-common
profiles: ["nostr-publish-consume"]
container_name: fips-nat-nostr-pub-b${FIPS_CI_NAME_SUFFIX:-}
hostname: fips-nat-nostr-pub-b
depends_on:
- relay
- stun
volumes:
- ../docker/resolv.conf:/etc/resolv.conf:ro
- ./generated-configs${FIPS_CI_NAME_SUFFIX:-}/nostr-publish-consume/node-b.yaml:/etc/fips/fips.yaml:ro
networks:
shared-lan:
ipv4_address: ${NAT_LAN_PREFIX:-172.31.10}.21
# ── STUN fault-injection profile ───────────────────────────────────────
# One FIPS daemon + a netns-sharing shim that injects tc/iptables faults
# against UDP egress to the STUN service. The runner script drives the
# shim via `docker exec` (Approach A) — no scripted timing inside the
# shim itself. Three phases:
# 1. drop — 100% UDP egress drop to STUN; assert daemon notices the
# observation timeout and retries.
# 2. delay — ~5s netem delay; assert daemon recovers and STUN succeeds
# again once the rule is removed.
# 3. kill — `docker stop fips-nat-stun`; assert daemon stays up and
# continues to handle "STUN unreachable" gracefully.
# The shim shares the daemon's network namespace so `tc qdisc add dev
# eth0 ...` operates on the daemon's egress path. The shim therefore
# has its own NET_ADMIN cap; the daemon already has one for TUN.
stun-fault-node:
<<: *fips-common
profiles: ["stun-faults"]
container_name: fips-nat-stun-fault-node${FIPS_CI_NAME_SUFFIX:-}
hostname: fips-nat-stun-fault-node
depends_on:
- relay
- stun
volumes:
- ../docker/resolv.conf:/etc/resolv.conf:ro
- ./generated-configs${FIPS_CI_NAME_SUFFIX:-}/stun-faults/stun-fault-node.yaml:/etc/fips/fips.yaml:ro
networks:
shared-lan:
ipv4_address: ${NAT_LAN_PREFIX:-172.31.10}.50
# Fault-free peer that publishes a valid overlay advert, so the
# fault-node's NAT-traversal attempt actually reaches
# observe_traversal_addresses() (the STUN client). Without this peer the
# daemon would abort with "no overlay advert" and never generate the
# STUN egress that the shim's tc/iptables rules are meant to drop.
# Intentionally has NO fault shim sharing its netns; runs cleanly.
stun-fault-peer:
<<: *fips-common
profiles: ["stun-faults"]
container_name: fips-nat-stun-fault-peer${FIPS_CI_NAME_SUFFIX:-}
hostname: fips-nat-stun-fault-peer
depends_on:
- relay
- stun
volumes:
- ../docker/resolv.conf:/etc/resolv.conf:ro
- ./generated-configs${FIPS_CI_NAME_SUFFIX:-}/stun-faults/stun-fault-peer.yaml:/etc/fips/fips.yaml:ro
networks:
shared-lan:
ipv4_address: ${NAT_LAN_PREFIX:-172.31.10}.51
stun-fault-shim:
image: ${FIPS_TEST_IMAGE:-fips-test:latest}
profiles: ["stun-faults"]
container_name: fips-nat-stun-fault-shim${FIPS_CI_NAME_SUFFIX:-}
depends_on:
- stun-fault-node
cap_add:
- NET_ADMIN
- NET_RAW
network_mode: "service:stun-fault-node"
restart: "no"
entrypoint:
- /bin/sh
- -c
- "exec sleep infinity"