Files
fips/testing/medium-change/docker-compose.yml
fr34akyandJohnathan Corgan b922568dca fix(node): re-pin connected UDP sockets when the host changes medium
An established UDP peer gets its own `connect()`-ed socket for the send fast
path. `open_connected_fd` binds the wildcard and then calls `connect(2)`, which
makes the kernel resolve the route once and auto-bind the local source address
to whichever interface was carrying it at that moment. It never re-evaluates.

So after the host changed transport medium — a laptop between WLAN and LAN, a
phone between Wi-Fi and cellular — every established peer went on transmitting
from an address the routing table had abandoned. The peer, which re-pins to
whatever address it last heard from, answered somewhere the node was no longer
sending from. The peering stayed marked connected and carried nothing until
`link_dead_timeout_secs` tore it down: 60-90s of black-holed traffic per switch
on a live node, then a full re-handshake and tree re-convergence.

The mirror-image case, the peer rotating its address, was already handled where
the rotation is observed. This is the local half, and it had no signal to hang
off, because a local move is invisible in the data plane.

Medium-change detection supplies that signal. `node.netmon.*` controls it and it
is on by default. The node samples a coarse fingerprint of its network
attachment — the source addresses the routing table would pick for an off-link
destination, plus the set of up, non-loopback interface addresses — and reports
a change once the picture settles. A handover is not atomic, so a short debounce
coalesces the burst into one event, and a fingerprint that settles back where it
started reports nothing. Linux and Android subscribe to `NETLINK_ROUTE`
multicast and macOS and FreeBSD to a `PF_ROUTE` socket, both reacting in
milliseconds; every other platform samples on a timer, which also runs
underneath the kernel sources as a backstop. A backend decides only when to
look, so the remaining ones land behind the same seam.

The reaction is two steps. Drop the stale connected sockets, which is
self-healing rather than disruptive: the wildcard listen socket resolves a route
per packet, so sends keep working immediately, and a correctly-bound socket is
reinstalled on a later tick. Then heartbeat every peer whose send path cannot
block, so the far side re-pins at once rather than waiting out its own interval.

That filter is the whole point rather than an optimisation. A connectionless
send completes without awaiting the wire. A connection-oriented one awaits an
unbounded `write_all` on a stream that the medium change has very likely just
stranded, and this reaction runs on the rx loop, so it would hold every other
arm of the select for as long as that socket took to fail. A peer on such a
transport keeps the periodic heartbeat it had before, with
`link_dead_timeout_secs` as the backstop.

Covered by unit tests, by a regression test that pins the fan-out filter, and by
a new `medium-change` integration suite: a multi-homed node whose default route
moves between two live access paths while mesh traffic is in flight, with the
far peer off-link behind a router.

The changelog entries land under Unreleased rather than in the released `0.5.1`
section, since none of this is in that release.
2026-09-06 22:28:43 +00:00

131 lines
4.5 KiB
YAML

# Transport-medium change lab.
#
# node-a ──┬── mc-primary ───┐
# │ ├── router ── mc-far ── node-b
# └── mc-secondary ─┘
#
# node-a is multi-homed with two equally usable paths to the router; node-b
# sits beyond it and is reachable only through the router. That last part is
# the whole design: node-b has to be off-link so the route to it follows
# node-a's *default* route, which is what the suite moves. Put node-b on a
# shared bridge instead and the directly-connected route wins, the source
# address never changes, and the bug under test cannot reproduce.
networks:
mc-primary:
driver: bridge
labels:
- "com.corganlabs.fips-ci=1"
ipam:
config:
- subnet: ${MC_PRIMARY_PREFIX:-172.31.60}.0/24
mc-secondary:
driver: bridge
labels:
- "com.corganlabs.fips-ci=1"
ipam:
config:
- subnet: ${MC_SECONDARY_PREFIX:-172.31.61}.0/24
mc-far:
driver: bridge
labels:
- "com.corganlabs.fips-ci=1"
ipam:
config:
- subnet: ${MC_FAR_PREFIX:-172.31.62}.0/24
x-fips-common: &fips-common
image: ${FIPS_TEST_IMAGE:-fips-test:latest}
cap_add:
- NET_ADMIN
devices:
- /dev/net/tun:/dev/net/tun
sysctls:
- net.ipv6.conf.all.disable_ipv6=0
restart: "no"
entrypoint:
- /usr/local/bin/mc-node-entrypoint.sh
environment:
- RUST_LOG=info,fips::node::netmon=debug,fips::node::handlers::netmon=debug
- PRIMARY_PREFIX=${MC_PRIMARY_PREFIX:-172.31.60}
- SECONDARY_PREFIX=${MC_SECONDARY_PREFIX:-172.31.61}
- FAR_PREFIX=${MC_FAR_PREFIX:-172.31.62}
- ROUTER_OCTET=254
services:
router:
build:
context: ./router
container_name: fips-mc-router${FIPS_CI_NAME_SUFFIX:-}
cap_add:
- NET_ADMIN
sysctls:
- net.ipv4.ip_forward=1
# Strict reverse-path filtering, and the suite does not work without it.
#
# It is what makes a stale source address *hurt*. Both of node-a's
# interfaces stay up and both stay routable, so a packet still sourced
# from the old path is otherwise forwarded and answered quite happily —
# the pin is stale but harmless, the bug does not bite, and the negative
# control passes, which would make every assertion in this suite vacuous.
#
# A real gateway drops that packet as spoofed, because the reverse route
# for its source points out a different interface. That is the actual
# reason a medium change black-holes traffic in the field, so it is the
# thing the lab has to model.
- net.ipv4.conf.all.rp_filter=1
- net.ipv4.conf.default.rp_filter=1
restart: "no"
networks:
mc-primary:
ipv4_address: ${MC_PRIMARY_PREFIX:-172.31.60}.254
mc-secondary:
ipv4_address: ${MC_SECONDARY_PREFIX:-172.31.61}.254
mc-far:
ipv4_address: ${MC_FAR_PREFIX:-172.31.62}.254
# The node under test. Two paths up at once, default route on the primary,
# and the suite moves it to the secondary mid-traffic.
node-a:
<<: *fips-common
container_name: fips-mc-node-a${FIPS_CI_NAME_SUFFIX:-}
hostname: fips-mc-node-a
depends_on:
- router
environment:
- RUST_LOG=info,fips::node::netmon=debug,fips::node::handlers::netmon=debug
- PRIMARY_PREFIX=${MC_PRIMARY_PREFIX:-172.31.60}
- SECONDARY_PREFIX=${MC_SECONDARY_PREFIX:-172.31.61}
- ROUTER_OCTET=254
- DEFAULT_VIA=primary
volumes:
- ../docker/resolv.conf:/etc/resolv.conf:ro
- ./node/entrypoint.sh:/usr/local/bin/mc-node-entrypoint.sh:ro
- ./generated-configs${FIPS_CI_NAME_SUFFIX:-}/node-a.yaml:/etc/fips/fips.yaml:ro
networks:
mc-primary:
ipv4_address: ${MC_PRIMARY_PREFIX:-172.31.60}.10
mc-secondary:
ipv4_address: ${MC_SECONDARY_PREFIX:-172.31.61}.10
# The far peer. Single-homed and stationary — it never moves, so anything
# the suite observes at this end is a consequence of node-a's move.
node-b:
<<: *fips-common
container_name: fips-mc-node-b${FIPS_CI_NAME_SUFFIX:-}
hostname: fips-mc-node-b
depends_on:
- router
environment:
- RUST_LOG=info
- FAR_PREFIX=${MC_FAR_PREFIX:-172.31.62}
- ROUTER_OCTET=254
- DEFAULT_VIA=far
volumes:
- ../docker/resolv.conf:/etc/resolv.conf:ro
- ./node/entrypoint.sh:/usr/local/bin/mc-node-entrypoint.sh:ro
- ./generated-configs${FIPS_CI_NAME_SUFFIX:-}/node-b.yaml:/etc/fips/fips.yaml:ro
networks:
mc-far:
ipv4_address: ${MC_FAR_PREFIX:-172.31.62}.20