mirror of
https://github.com/jmcorgan/fips.git
synced 2026-10-05 11:08:25 +00:00
The NAT-lab suites (cone, symmetric, lan, STUN faults and nostr publish/consume) share one strfry relay, and it has aborted in several runs. Three things made each abort cost a whole run and a misdirected diagnosis: only the publish/consume suite said what state the relay was in, the relay image moved with upstream's latest tag, and the relay never restarted. State the relay's condition in every NAT-lab suite. relay_verdict moves into testing/lib/relay-verdict.sh, taking the container as an argument, and is called first in every NAT-lab failure dump and on each success path, where a relay event the assertions survived is noted rather than made a failure. Before this, the cone, symmetric and lan dumps named the nodes and the network instead: hundreds of lines of socket tables and router counters, with the relay's crash lines unlabelled in an 80-line log tail. The STUN fault dump now also carries the relay's log, which it did not include. Pin the relay's strfry build. Dockerfile.app takes the strfry image as a build argument, defaulting to the same latest tag, so the user-facing example is unchanged, and the NAT lab's compose file passes the current multi-arch index digest. Which build a run exercised was never recorded before. This stops the drift; it does not select a build that does not abort, since upstream publishes no other tag to choose from. Restart the relay when it aborts. The relay ran with restart "no", so one abort left every suite sharing it without a relay for the rest of the run. It now restarts on failure up to three times, so a relay that keeps aborting still ends up exited and is reported rather than hidden in a crash loop. The restart policy is the containment. The verdict says the relay restarted under that policy and prints the fault lines from its log, which spans restarts of the same container, so a rescued run still names the event. The relay service also gains init: true, only so the restart can be exercised by an injected abort. It changes the relay's PID 1 from strfry (started with exec in the image's entrypoint) to docker's init, with strfry as its child. Without it, a SIGABRT sent with docker kill to strfry as PID 1 logged "caught a signal: SIGABRT" and left the container running with no restart, so that injection could not show the policy working. With the init, strfry signalled from inside the container exits it non-zero, which is what the real aborts did: clients saw the relay vanish. Those real aborts would restart under the policy with or without the init. Injected aborts (strfry signalled from inside the container the moment both cone nodes had connected) restarted the relay once each time; in eight of nine cone runs both nodes reconnected and peered 8 to 16 s after the abort. In the ninth, the initiator's offer went out in the second between its reconnect and the responder's, was lost, and the 30 s answer timeout pushed peering past the 45 s wait, so restart shortens the outage but does not guarantee the run. Without the restart policy the same injection left the relay exited and the cone suite timed out waiting for its peer, with no line in the dump naming the relay.
54 lines
1.8 KiB
Docker
54 lines
1.8 KiB
Docker
# strfry: official multi-arch image (linux/amd64 + linux/arm64)
|
|
# replaces the nostr-rs-relay source build which was amd64-only.
|
|
#
|
|
# IMPORTANT: use alpine (same musl-based libc as the strfry source image) as
|
|
# the final stage. Copying the strfry binary into a glibc-based image such as
|
|
# debian:bookworm-slim causes "not found" at exec time because the musl dynamic
|
|
# linker (/lib/ld-musl-*.so.1) and its shared libraries are absent there.
|
|
#
|
|
# STRFRY_IMAGE lets a caller pin the strfry build. The default follows the
|
|
# upstream `latest` tag; the NAT test lab passes a digest.
|
|
ARG STRFRY_IMAGE=ghcr.io/hoytech/strfry:latest
|
|
FROM ${STRFRY_IMAGE} AS strfry
|
|
|
|
FROM alpine:3.18
|
|
|
|
RUN apk add --no-cache nginx
|
|
|
|
# nginx: reverse proxy port 80 (IPv4 + IPv6) → strfry 127.0.0.1:7777
|
|
RUN printf 'server {\n\
|
|
listen 80;\n\
|
|
listen [::]:80;\n\
|
|
location / {\n\
|
|
proxy_pass http://127.0.0.1:7777;\n\
|
|
proxy_http_version 1.1;\n\
|
|
proxy_read_timeout 1d;\n\
|
|
proxy_send_timeout 1d;\n\
|
|
proxy_set_header Upgrade $http_upgrade;\n\
|
|
proxy_set_header Connection "Upgrade";\n\
|
|
proxy_set_header Host $host;\n\
|
|
}\n\
|
|
}\n' > /etc/nginx/http.d/nostr-relay.conf && \
|
|
rm -f /etc/nginx/http.d/default.conf
|
|
|
|
COPY --from=strfry /app/strfry /usr/local/bin/strfry
|
|
# Copy musl shared libraries that strfry depends on from the source image.
|
|
COPY --from=strfry \
|
|
/usr/lib/liblmdb.so.0 \
|
|
/usr/lib/libcrypto.so.50 \
|
|
/usr/lib/libssl.so.53 \
|
|
/usr/lib/libsecp256k1.so.2 \
|
|
/usr/lib/libzstd.so.1 \
|
|
/usr/lib/libstdc++.so.6 \
|
|
/usr/lib/libgcc_s.so.1 \
|
|
/usr/lib/
|
|
COPY --from=strfry /lib/libz.so.1 /lib/
|
|
|
|
WORKDIR /usr/src/app
|
|
RUN mkdir -p strfry-db
|
|
|
|
RUN printf '#!/bin/sh\nnginx\nexec strfry relay\n' > /entrypoint-app.sh && \
|
|
chmod +x /entrypoint-app.sh
|
|
|
|
ENTRYPOINT ["/entrypoint-app.sh"]
|