mirror of
https://github.com/jmcorgan/fips.git
synced 2026-10-05 11:08:25 +00:00
An established UDP peer gets its own `connect()`-ed socket for the send fast path. `open_connected_fd` binds the wildcard and then calls `connect(2)`, which makes the kernel resolve the route once and auto-bind the local source address to whichever interface was carrying it at that moment. It never re-evaluates. So after the host changed transport medium — a laptop between WLAN and LAN, a phone between Wi-Fi and cellular — every established peer went on transmitting from an address the routing table had abandoned. The peer, which re-pins to whatever address it last heard from, answered somewhere the node was no longer sending from. The peering stayed marked connected and carried nothing until `link_dead_timeout_secs` tore it down: 60-90s of black-holed traffic per switch on a live node, then a full re-handshake and tree re-convergence. The mirror-image case, the peer rotating its address, was already handled where the rotation is observed. This is the local half, and it had no signal to hang off, because a local move is invisible in the data plane. Medium-change detection supplies that signal. `node.netmon.*` controls it and it is on by default. The node samples a coarse fingerprint of its network attachment — the source addresses the routing table would pick for an off-link destination, plus the set of up, non-loopback interface addresses — and reports a change once the picture settles. A handover is not atomic, so a short debounce coalesces the burst into one event, and a fingerprint that settles back where it started reports nothing. Linux and Android subscribe to `NETLINK_ROUTE` multicast and macOS and FreeBSD to a `PF_ROUTE` socket, both reacting in milliseconds; every other platform samples on a timer, which also runs underneath the kernel sources as a backstop. A backend decides only when to look, so the remaining ones land behind the same seam. The reaction is two steps. Drop the stale connected sockets, which is self-healing rather than disruptive: the wildcard listen socket resolves a route per packet, so sends keep working immediately, and a correctly-bound socket is reinstalled on a later tick. Then heartbeat every peer whose send path cannot block, so the far side re-pins at once rather than waiting out its own interval. That filter is the whole point rather than an optimisation. A connectionless send completes without awaiting the wire. A connection-oriented one awaits an unbounded `write_all` on a stream that the medium change has very likely just stranded, and this reaction runs on the rx loop, so it would hold every other arm of the select for as long as that socket took to fail. A peer on such a transport keeps the periodic heartbeat it had before, with `link_dead_timeout_secs` as the backstop. Covered by unit tests, by a regression test that pins the fan-out filter, and by a new `medium-change` integration suite: a multi-homed node whose default route moves between two live access paths while mesh traffic is in flight, with the far peer off-link behind a router. The changelog entries land under Unreleased rather than in the released `0.5.1` section, since none of this is in that release.
115 lines
3.4 KiB
Bash
Executable File
115 lines
3.4 KiB
Bash
Executable File
#!/bin/bash
|
|
# Render the two node configs for the medium-change lab.
|
|
#
|
|
# Usage: generate-configs.sh [mesh-name] [netmon-enabled]
|
|
#
|
|
# `netmon-enabled` re-renders node-a with detection off. The test script uses
|
|
# that for its negative control: the same topology and the same move, with the
|
|
# only difference being the mechanism under test. A regression test that has
|
|
# never been seen to fail is a claim, not a test, so the suite proves the
|
|
# failure rather than asserting it from the changelog.
|
|
set -euo pipefail
|
|
|
|
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
|
MC_DIR="$(cd "$SCRIPT_DIR/.." && pwd)"
|
|
ROOT_DIR="$(cd "$MC_DIR/../.." && pwd)"
|
|
DERIVE_KEYS="$ROOT_DIR/testing/lib/derive_keys.py"
|
|
# Per-run, matching every other generator in the tree: compose bind-mounts
|
|
# this directory and the test script reads the npubs back after the containers
|
|
# are up, so a shared path lets a second run overwrite what a first is about
|
|
# to ping.
|
|
OUTPUT_DIR="$MC_DIR/generated-configs${FIPS_CI_NAME_SUFFIX:-}"
|
|
|
|
MESH_NAME="${1:-medium-change-$(date +%s)-$$}"
|
|
NETMON_ENABLED="${2:-true}"
|
|
|
|
primary="${MC_PRIMARY_PREFIX:-172.31.60}"
|
|
far="${MC_FAR_PREFIX:-172.31.62}"
|
|
|
|
mkdir -p "$OUTPUT_DIR"
|
|
|
|
keys_a="$(python3 "$DERIVE_KEYS" "$MESH_NAME" "a")"
|
|
keys_b="$(python3 "$DERIVE_KEYS" "$MESH_NAME" "b")"
|
|
nsec_a="$(echo "$keys_a" | awk -F= '/^nsec=/{print $2}')"
|
|
npub_a="$(echo "$keys_a" | awk -F= '/^npub=/{print $2}')"
|
|
nsec_b="$(echo "$keys_b" | awk -F= '/^nsec=/{print $2}')"
|
|
npub_b="$(echo "$keys_b" | awk -F= '/^npub=/{print $2}')"
|
|
|
|
write_config() {
|
|
local output_file="$1" nsec="$2" peer_npub="$3" peer_alias="$4" \
|
|
peer_addr="$5" netmon_block="$6"
|
|
|
|
cat > "$output_file" <<EOF
|
|
# Generated by testing/medium-change/scripts/generate-configs.sh
|
|
node:
|
|
identity:
|
|
nsec: "$nsec"
|
|
$netmon_block
|
|
retry:
|
|
max_retries: 5
|
|
base_interval_secs: 2
|
|
|
|
tun:
|
|
enabled: true
|
|
name: fips0
|
|
mtu: 1280
|
|
|
|
dns:
|
|
enabled: true
|
|
|
|
transports:
|
|
udp:
|
|
# Wildcard on purpose. A per-peer connected socket is only opened over a
|
|
# wildcard-bound transport, and that socket is what this lab exercises.
|
|
bind_addr: "0.0.0.0:2121"
|
|
mtu: 1400
|
|
|
|
peers:
|
|
- npub: "$peer_npub"
|
|
alias: "$peer_alias"
|
|
addresses:
|
|
- transport: udp
|
|
addr: "$peer_addr"
|
|
priority: 1
|
|
EOF
|
|
}
|
|
|
|
if [ "$NETMON_ENABLED" = "true" ]; then
|
|
netmon_a=$(cat <<'EOF'
|
|
netmon:
|
|
enabled: true
|
|
poll_interval_secs: 5
|
|
debounce_ms: 250
|
|
EOF
|
|
)
|
|
else
|
|
netmon_a=$(cat <<'EOF'
|
|
# Negative control: the detector is off, so nothing tells the node its
|
|
# egress path moved and the connected socket stays pinned to the old one.
|
|
netmon:
|
|
enabled: false
|
|
EOF
|
|
)
|
|
fi
|
|
|
|
# node-a dials the far node through the router, so the route to it follows
|
|
# node-a's default route — the thing the suite moves.
|
|
write_config "$OUTPUT_DIR/node-a.yaml" "$nsec_a" "$npub_b" "node-b" \
|
|
"${far}.20:2121" "$netmon_a"
|
|
|
|
# node-b is given node-a's *primary* address only. After the move that address
|
|
# is stale, which is the realistic shape: a configured peer whose path went
|
|
# away. Recovery by re-dial is therefore still possible here — and the suite
|
|
# distinguishes it from the fix by asserting session continuity, not merely
|
|
# that traffic eventually returns.
|
|
write_config "$OUTPUT_DIR/node-b.yaml" "$nsec_b" "$npub_a" "node-a" \
|
|
"${primary}.10:2121" " netmon:
|
|
enabled: true"
|
|
|
|
cat > "$OUTPUT_DIR/npubs.env" <<EOF
|
|
NPUB_A=$npub_a
|
|
NPUB_B=$npub_b
|
|
EOF
|
|
|
|
echo "Generated configs in $OUTPUT_DIR (netmon on node-a: $NETMON_ENABLED)"
|