mirror of
https://github.com/jmcorgan/fips.git
synced 2026-10-05 19:18:25 +00:00
A peer reachable over more than one transport keeps one Noise session and moves its traffic between transports on failure or degradation. Implements docs/design/fips-multi-path-switchover.md §4-§10 and closes the two-interface case in #143. Three inner link messages next to Heartbeat: `0x52 PathProbe` and `0x53 PathAck`, carrying a probe id, the sender's path id and a `remote_active` bit, and `0x54 PathClose`, naming the receiver's path id and a reason. A probe is an ordinary encrypted frame sent on a candidate transport; the receiver, having decrypted it against the session found by index, adds the path as `Probing`, marks it `rx_live` and answers on that same path. The prober's receipt of the ack marks the path `Live`, `tx_live`, and takes an RTT sample. No handshake, no key material, no index allocation. Old nodes drop the unknown types at debug, so a path to one stays `Probing` and never becomes eligible. The discovery gate changes shape: a live peer beaconing on a transport we hold no path to it over becomes a path candidate rather than being skipped, and the heartbeat tick probes it. The active path's first probe is small — the handshake proved it and seeded its MTU — while a standby's discovery probes are padded to the link MTU, as is one a minute on every path, so a medium that passes small frames and drops large ones never proves itself. A standby the peer never answers on is given up after eight probes; the active path never is. Detection is per path and takes each medium's own failure signal: a carrier edge, an unreachable-on-send (`ENETUNREACH`/`EHOSTUNREACH`), an interface going away, or two unanswered heartbeats on a path the peer is also silent on. Any of them marks the path `Suspect` and selection leaves it at once, because the standby is warm: heartbeats run at `node.path.active_heartbeat_ms` where either side sends and `standby_heartbeat_ms` elsewhere, both stretched by the path's own round trip so a Tor or Nym path is neither flooded nor declared dead every round trip. A peer holding one live path is not heartbeated here at all — selection has nothing to move to, and the link heartbeat keeps its liveness. Soft signals (the peer's `remote_active` flipping away, silence here while a standby hears the peer) trigger a probe, never `Suspect`: reading them as a verdict forces both sides onto one path and loops under a one-way failure. A node that loses a path tells the peer with a `PathClose` on a surviving one, so the peer moves at once rather than after its own timeout. A transport that returns inside the five-minute grace revives its dead paths as `Probing` with their RTT window and ETX intact. Selection is measured, not configured: each path scores `quality_index(etx, min_rtt)`, and traffic moves when the active path is no longer eligible (mandatory) or when a standby beats it by `switch_margin` for `switch_dwell_secs` (discretionary). Min RTT over a window rather than SRTT, because SRTT inflates under load on the path carrying traffic while an idle standby looks pristine — a ping-pong generator. A `role: backup` transport carries a peer's traffic only while no normal path is eligible, and yields outright when one becomes eligible. `fipsctl path pin` overrides both while its path is eligible. A switch re-seeds the path MTU from the new path, tightens the session MTUs and refreshes the MSS ceiling, so the first frames after a switch are not black-holed. The link record and `addr_to_link` follow the active path. The link cost the tree sees is held at its pre-switch value for the dwell, and until the two receiver reports that span the switch have arrived — the first counts every frame in flight on the old path as lost and spikes the per-report ETX for one interval, the second replaces it — so neither a short flap nor that spike ripples mesh-wide through parent selection or the next-hop order. Operator surface: `role: backup` on any transport, `node.path.*` (`switch_margin`, validated finite and at least 1.0, `switch_dwell_secs`, `min_samples`, `active_heartbeat_ms`, `standby_heartbeat_ms`), and `fipsctl path show|pin|unpin` over the `path_show`, `path_pin` and `path_unpin` control commands. Two chaos scenarios calibrate the defaults and are wired into both runners: `dual-path-flap` (a raw-Ethernet veth as the cable, the Docker bridge over UDP as the wifi) and `dual-udp-flap` (two interface-bound UDP instances). Each flaps one path under iperf and carries detectors that can fail on a switchover that did not carry traffic — a per-node ceiling on "Peer promoted to active" (a second is a re-peering), a one-second ceiling from link-down to the first switch, and a two-second ceiling on any zero-byte iperf interval run — alongside the `path_switches` band. Neither has been run to calibrate; the defaults are chosen, not derived, and the design doc says so. Refs #143
266 lines
8.5 KiB
Bash
266 lines
8.5 KiB
Bash
#!/bin/bash
|
||
# Unified entrypoint for FIPS test containers.
|
||
#
|
||
# Mode is selected via FIPS_TEST_MODE environment variable:
|
||
# default — dnsmasq + sshd + iperf3 + http server + fips
|
||
# chaos — above + TCP ECN + ethernet interface wait
|
||
# sidecar — generate config from env + iptables isolation + fips
|
||
# tor-socks5 — dnsmasq + sshd + fips (tor daemon is separate)
|
||
# tor-directory — dnsmasq + tor + wait for .onion hostname + fips
|
||
|
||
set -e
|
||
|
||
MODE="${FIPS_TEST_MODE:-default}"
|
||
CONFIG="/etc/fips/fips.yaml"
|
||
|
||
# ── Common: dnsmasq ──────────────────────────────────────────────────────
|
||
|
||
start_dnsmasq() {
|
||
dnsmasq
|
||
}
|
||
|
||
# ── Common: background services (sshd, iperf3, http) ────────────────────
|
||
|
||
start_services() {
|
||
/usr/sbin/sshd
|
||
iperf3 -s -D
|
||
python3 -m http.server 8000 -d /root -b :: &>/dev/null &
|
||
}
|
||
|
||
# ── Chaos: TCP ECN + ethernet wait ──────────────────────────────────────
|
||
|
||
enable_ecn() {
|
||
sysctl -w net.ipv4.tcp_ecn=1 >/dev/null 2>&1 || true
|
||
}
|
||
|
||
wait_for_ethernet() {
|
||
# If config binds any transport to an interface (Ethernet, or a UDP
|
||
# instance with `interface:`), wait for those interfaces to appear. Veth
|
||
# pairs are created from the host after the container starts, and a UDP
|
||
# instance bound to an interface that is not there yet fails to start.
|
||
local eth_ifaces=""
|
||
eth_ifaces=$(grep '^\s*interface:' "$CONFIG" 2>/dev/null \
|
||
| sed 's/.*interface:\s*//' \
|
||
| tr -d ' "' || true)
|
||
|
||
if [ -n "$eth_ifaces" ]; then
|
||
echo "Waiting for Ethernet interfaces: $eth_ifaces"
|
||
local deadline=$((SECONDS + 30))
|
||
local all_found=false
|
||
while [ $SECONDS -lt $deadline ]; do
|
||
all_found=true
|
||
for iface in $eth_ifaces; do
|
||
if [ ! -e "/sys/class/net/$iface" ]; then
|
||
all_found=false
|
||
break
|
||
fi
|
||
done
|
||
if $all_found; then
|
||
echo "All Ethernet interfaces ready"
|
||
break
|
||
fi
|
||
sleep 0.2
|
||
done
|
||
if ! $all_found; then
|
||
echo "WARNING: Timed out waiting for Ethernet interfaces"
|
||
fi
|
||
fi
|
||
}
|
||
|
||
# ── Sidecar: config generation + iptables isolation ─────────────────────
|
||
|
||
generate_sidecar_config() {
|
||
FIPS_NSEC="${FIPS_NSEC:?FIPS_NSEC is required}"
|
||
FIPS_UDP_BIND="${FIPS_UDP_BIND:-0.0.0.0:2121}"
|
||
FIPS_TUN_MTU="${FIPS_TUN_MTU:-1280}"
|
||
FIPS_PEER_TRANSPORT="${FIPS_PEER_TRANSPORT:-udp}"
|
||
|
||
mkdir -p /etc/fips
|
||
|
||
local peers_section=""
|
||
if [ -n "$FIPS_PEER_NPUB" ] && [ -n "$FIPS_PEER_ADDR" ]; then
|
||
FIPS_PEER_ALIAS="${FIPS_PEER_ALIAS:-peer}"
|
||
peers_section=" - npub: \"${FIPS_PEER_NPUB}\"
|
||
alias: \"${FIPS_PEER_ALIAS}\"
|
||
addresses:
|
||
- transport: ${FIPS_PEER_TRANSPORT}
|
||
addr: \"${FIPS_PEER_ADDR}\"
|
||
connect_policy: auto_connect"
|
||
fi
|
||
|
||
cat > "$CONFIG" <<EOF
|
||
node:
|
||
identity:
|
||
nsec: "${FIPS_NSEC}"
|
||
|
||
tun:
|
||
enabled: true
|
||
name: fips0
|
||
mtu: ${FIPS_TUN_MTU}
|
||
|
||
dns:
|
||
enabled: true
|
||
|
||
transports:
|
||
udp:
|
||
bind_addr: "${FIPS_UDP_BIND}"
|
||
mtu: 1472
|
||
tcp: {}
|
||
|
||
peers:
|
||
${peers_section:- []}
|
||
EOF
|
||
|
||
echo "Generated $CONFIG"
|
||
}
|
||
|
||
apply_iptables_isolation() {
|
||
# Only FIPS transport (UDP 2121, TCP 443) may use eth0.
|
||
# All other eth0 traffic is dropped. fips0 and loopback unrestricted.
|
||
iptables -A OUTPUT -o lo -j ACCEPT
|
||
iptables -A INPUT -i lo -j ACCEPT
|
||
iptables -A OUTPUT -o eth0 -p udp --dport 2121 -j ACCEPT
|
||
iptables -A OUTPUT -o eth0 -p udp --sport 2121 -j ACCEPT
|
||
iptables -A INPUT -i eth0 -p udp --dport 2121 -j ACCEPT
|
||
iptables -A INPUT -i eth0 -p udp --sport 2121 -j ACCEPT
|
||
iptables -A OUTPUT -o eth0 -p tcp --dport 443 -j ACCEPT
|
||
iptables -A INPUT -i eth0 -p tcp --sport 443 -j ACCEPT
|
||
iptables -A OUTPUT -o eth0 -j DROP
|
||
iptables -A INPUT -i eth0 -j DROP
|
||
|
||
ip6tables -A OUTPUT -o lo -j ACCEPT
|
||
ip6tables -A INPUT -i lo -j ACCEPT
|
||
ip6tables -A OUTPUT -o fips0 -j ACCEPT
|
||
ip6tables -A INPUT -i fips0 -j ACCEPT
|
||
ip6tables -A OUTPUT -o eth0 -j DROP
|
||
ip6tables -A INPUT -i eth0 -j DROP
|
||
|
||
echo "iptables isolation rules applied"
|
||
}
|
||
|
||
# ── Tor directory mode: start tor + wait for hostname ────────────────────
|
||
|
||
start_tor_directory() {
|
||
local hidden_service_dir="/var/lib/tor/fips_onion_service"
|
||
local is_directory=false
|
||
|
||
if grep -qE '^\s+mode:\s+"directory"' "$CONFIG" 2>/dev/null; then
|
||
is_directory=true
|
||
fi
|
||
|
||
if [ "$is_directory" = true ]; then
|
||
mkdir -p "$hidden_service_dir"
|
||
chmod 700 "$hidden_service_dir"
|
||
fi
|
||
|
||
echo "Starting Tor daemon..."
|
||
tor -f /etc/tor/torrc &
|
||
|
||
if [ "$is_directory" = true ]; then
|
||
local hostname_file="${hidden_service_dir}/hostname"
|
||
echo "Waiting for Tor to create ${hostname_file}..."
|
||
for i in $(seq 1 120); do
|
||
if [ -f "$hostname_file" ]; then
|
||
echo "Tor hostname file ready after ${i}s: $(cat "$hostname_file")"
|
||
break
|
||
fi
|
||
sleep 1
|
||
done
|
||
if [ ! -f "$hostname_file" ]; then
|
||
echo "FATAL: Tor did not create hostname file within 120s"
|
||
exit 1
|
||
fi
|
||
fi
|
||
}
|
||
|
||
# ── Mode dispatch ────────────────────────────────────────────────────────
|
||
|
||
case "$MODE" in
|
||
default)
|
||
start_dnsmasq
|
||
start_services
|
||
exec fips --config "$CONFIG"
|
||
;;
|
||
chaos)
|
||
enable_ecn
|
||
start_dnsmasq
|
||
start_services
|
||
wait_for_ethernet
|
||
exec fips --config "$CONFIG"
|
||
;;
|
||
sidecar)
|
||
generate_sidecar_config
|
||
apply_iptables_isolation
|
||
start_dnsmasq
|
||
exec fips --config "$CONFIG"
|
||
;;
|
||
tor-socks5)
|
||
start_dnsmasq
|
||
/usr/sbin/sshd
|
||
exec fips --config "$CONFIG"
|
||
;;
|
||
tor-directory)
|
||
start_dnsmasq
|
||
start_tor_directory
|
||
echo "Starting FIPS daemon..."
|
||
exec fips --config "$CONFIG"
|
||
;;
|
||
gateway)
|
||
# No dnsmasq — gateway DNS replaces it on port 53
|
||
start_services
|
||
|
||
# Extract LAN interface from config (gateway.lan_interface)
|
||
LAN_IF=$(grep 'lan_interface:' "$CONFIG" | head -1 | sed 's/.*: *//' | tr -d '"' | tr -d "'")
|
||
LAN_IF="${LAN_IF:-eth0}"
|
||
|
||
# Wait for LAN interface (Docker attaches second network after start)
|
||
for i in $(seq 1 15); do
|
||
[ -e "/sys/class/net/$LAN_IF" ] && break
|
||
sleep 0.5
|
||
done
|
||
|
||
# Ensure IPv6 is enabled on the LAN interface (may inherit host default)
|
||
sysctl -w "net.ipv6.conf.${LAN_IF}.disable_ipv6=0" >/dev/null 2>&1 || true
|
||
sysctl -w net.ipv6.conf.all.forwarding=1 >/dev/null 2>&1 || true
|
||
sysctl -w net.ipv6.conf.all.proxy_ndp=1 >/dev/null 2>&1 || true
|
||
|
||
# Start fips in background (gateway needs fips0)
|
||
fips --config "$CONFIG" &
|
||
# Wait for fips0 TUN device
|
||
for i in $(seq 1 30); do
|
||
[ -e /sys/class/net/fips0 ] && break
|
||
sleep 1
|
||
done
|
||
if [ ! -e /sys/class/net/fips0 ]; then
|
||
echo "FATAL: fips0 did not appear within 30s"
|
||
exit 1
|
||
fi
|
||
|
||
# Wait for the daemon's DNS responder to bind [::1]:5354 before
|
||
# exec'ing fips-gateway. The gateway binary's startup probe is
|
||
# bounded (5 attempts × 1s with retry); this harness wait is the
|
||
# belt to that suspenders so we get deterministic CI behaviour
|
||
# on slow runners. Bounded to ~30 seconds; if the daemon
|
||
# really never binds DNS, the gateway's own probe will report
|
||
# the definitive error after this wait expires.
|
||
for i in $(seq 1 30); do
|
||
if dig @::1 -p 5354 +tries=1 +time=1 test.fips >/dev/null 2>&1; then
|
||
echo "Daemon DNS ready (waited ~${i}s)"
|
||
break
|
||
fi
|
||
if [ "$i" -eq 30 ]; then
|
||
echo "WARNING: daemon DNS did not respond within ~30s; proceeding with gateway startup"
|
||
fi
|
||
sleep 1
|
||
done
|
||
|
||
echo "fips0 ready, starting gateway"
|
||
exec fips-gateway --config "$CONFIG" --log-level debug
|
||
;;
|
||
*)
|
||
echo "Unknown FIPS_TEST_MODE: $MODE"
|
||
echo "Valid modes: default, chaos, sidecar, tor-socks5, tor-directory, gateway"
|
||
exit 1
|
||
;;
|
||
esac
|