Files
fips/testing/chaos/sim/veth.py
Arjen dca938c104 feat(peer): switchover between a peer's transports, with no handshake
A peer reachable over more than one transport keeps one Noise session
and moves its traffic between transports on failure or degradation.
Implements docs/design/fips-multi-path-switchover.md §4-§10 and closes
the two-interface case in #143.

Three inner link messages next to Heartbeat: `0x52 PathProbe` and
`0x53 PathAck`, carrying a probe id, the sender's path id and a
`remote_active` bit, and `0x54 PathClose`, naming the receiver's path
id and a reason. A probe is an ordinary encrypted frame sent on a
candidate transport; the receiver, having decrypted it against the
session found by index, adds the path as `Probing`, marks it `rx_live`
and answers on that same path. The prober's receipt of the ack marks
the path `Live`, `tx_live`, and takes an RTT sample. No handshake, no
key material, no index allocation. Old nodes drop the unknown types at
debug, so a path to one stays `Probing` and never becomes eligible.

The discovery gate changes shape: a live peer beaconing on a transport
we hold no path to it over becomes a path candidate rather than being
skipped, and the heartbeat tick probes it. The active path's first
probe is small — the handshake proved it and seeded its MTU — while a
standby's discovery probes are padded to the link MTU, as is one a
minute on every path, so a medium that passes small frames and drops
large ones never proves itself. A standby the peer never answers on is
given up after eight probes; the active path never is.

Detection is per path and takes each medium's own failure signal: a
carrier edge, an unreachable-on-send (`ENETUNREACH`/`EHOSTUNREACH`), an
interface going away, or two unanswered heartbeats on a path the peer
is also silent on. Any of them marks the path `Suspect` and selection
leaves it at once, because the standby is warm: heartbeats run at
`node.path.active_heartbeat_ms` where either side sends and
`standby_heartbeat_ms` elsewhere, both stretched by the path's own
round trip so a Tor or Nym path is neither flooded nor declared dead
every round trip. A peer holding one live path is not heartbeated here
at all — selection has nothing to move to, and the link heartbeat keeps
its liveness. Soft signals (the peer's `remote_active` flipping away,
silence here while a standby hears the peer) trigger a probe, never
`Suspect`: reading them as a verdict forces both sides onto one path
and loops under a one-way failure.

A node that loses a path tells the peer with a `PathClose` on a
surviving one, so the peer moves at once rather than after its own
timeout. A transport that returns inside the five-minute grace revives
its dead paths as `Probing` with their RTT window and ETX intact.

Selection is measured, not configured: each path scores
`quality_index(etx, min_rtt)`, and traffic moves when the active path
is no longer eligible (mandatory) or when a standby beats it by
`switch_margin` for `switch_dwell_secs` (discretionary). Min RTT over a
window rather than SRTT, because SRTT inflates under load on the path
carrying traffic while an idle standby looks pristine — a ping-pong
generator. A `role: backup` transport carries a peer's traffic only
while no normal path is eligible, and yields outright when one becomes
eligible. `fipsctl path pin` overrides both while its path is eligible.

A switch re-seeds the path MTU from the new path, tightens the
session MTUs and refreshes the MSS ceiling, so the first frames after a
switch are not black-holed. The link record and `addr_to_link` follow
the active path. The link cost the tree sees is held at its pre-switch
value for the dwell, and until the two receiver reports that span the
switch have arrived — the first counts every frame in flight on the old
path as lost and spikes the per-report ETX for one interval, the second
replaces it — so neither a short flap nor that spike ripples mesh-wide
through parent selection or the next-hop order.

Operator surface: `role: backup` on any transport, `node.path.*`
(`switch_margin`, validated finite and at least 1.0, `switch_dwell_secs`,
`min_samples`, `active_heartbeat_ms`, `standby_heartbeat_ms`), and
`fipsctl path show|pin|unpin` over the `path_show`, `path_pin` and
`path_unpin` control commands.

Two chaos scenarios calibrate the defaults and are wired into both
runners: `dual-path-flap` (a raw-Ethernet veth as the cable, the Docker
bridge over UDP as the wifi) and `dual-udp-flap` (two interface-bound
UDP instances). Each flaps one path under iperf and carries detectors
that can fail on a switchover that did not carry traffic — a per-node
ceiling on "Peer promoted to active" (a second is a re-peering), a
one-second ceiling from link-down to the first switch, and a two-second
ceiling on any zero-byte iperf interval run — alongside the
`path_switches` band. Neither has been run to calibrate; the defaults
are chosen, not derived, and the design doc says so.

Refs #143
2026-09-23 11:48:33 -03:00

403 lines
16 KiB
Python
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""Veth pair management for Ethernet transport edges.
Creates veth pairs between Docker containers for Ethernet-transport
edges. Each Ethernet edge gets a veth pair with one end moved into
each container's network namespace. Naming:
Host (temporary): vh{token}{NN}{MM}a / vh{token}{NN}{MM}b
(via SimTopology.veth_host_name())
Container: ve-{local}-{peer} (via veth_interface_name())
After creation, the container-side MAC addresses are queried and
stored in SimNode.ethernet_macs for use in config generation.
Implementation note
-------------------
All ``ip link`` operations that manipulate the host network stack are
executed inside a short-lived privileged Docker container that shares
the host network and PID namespaces (``--net=host --pid=host``). This
works on both Linux and macOS:
* **Linux** – the helper container shares the real host network/PID
namespaces, so ``ip link set ... netns <pid>`` behaves identically to
running ``ip`` directly on the host.
* **macOS** – Docker containers run inside a Linux VM; the helper
container shares *that* VM's namespaces, which is exactly where the
simulation containers live. Running ``ip`` on the macOS host would
never work because the container PIDs are in the VM, not macOS.
The helper image is resolved from the running simulation containers
(which already have ``iproute2`` from the chaos Dockerfile).
"""
from __future__ import annotations
import logging
import subprocess
import time
from .docker_exec import DockerExecError, docker_exec, docker_exec_quiet
from .topology import UDP_VETH, SimTopology, veth_interface_name
log = logging.getLogger(__name__)
# How long a freshly raised veth end may take to report operstate up. The
# kernel publishes the carrier change through linkwatch, which is deferred, so
# a read straight after `ip link set up` can still see the old state. Long
# enough for at least two reads under host load, since one slow `docker exec`
# must not abort a run whose link is up.
OPERSTATE_WAIT_SECS = 15
# Timeout for each `docker exec` a restore makes. The run blocks on these, and
# a timeout aborts it, so this errs long: a loaded host is the normal case.
EXEC_TIMEOUT_SECS = 30
class VethSetupError(RuntimeError):
"""A veth pair could not be created, placed, renamed or raised.
Raised rather than logged. A pair that fails part way leaves a ring link
down while the run carries on, and what reaches the verdict is then a tree
that did not converge: a daemon failure in every respect a reader can see.
Letting this propagate aborts the run, so the fault is reported as the
harness's own.
"""
class VethManager:
"""Manages veth pairs for Ethernet-transport edges."""
def __init__(self, topology: SimTopology):
self.topology = topology
# Track created host-side temp names for cleanup
self._host_pairs: list[tuple[str, str, str, str]] = []
# (node_a, node_b, host_name_a, host_name_b)
self._ip_image: str | None = None
def _get_image(self) -> str:
"""Resolve the Docker image to use for ip(8) helper containers.
Uses the image of the first simulation container (which already has
iproute2 installed via the chaos Dockerfile). Must be called after
containers are started.
"""
if self._ip_image is not None:
return self._ip_image
first_node = next(iter(sorted(self.topology.nodes)))
container = self.topology.container_name(first_node)
result = subprocess.run(
["docker", "inspect", "-f", "{{.Config.Image}}", container],
capture_output=True,
text=True,
timeout=10,
)
if result.returncode != 0 or not result.stdout.strip():
raise RuntimeError(
f"Cannot determine Docker image for ip(8) helper "
f"(docker inspect {container} failed): {result.stderr.strip()}"
)
self._ip_image = result.stdout.strip()
log.debug("Using sim image %s for ip(8) helper", self._ip_image)
return self._ip_image
def setup_all(self):
"""Create veth pairs for all Ethernet edges.
For each Ethernet edge:
1. Get container PIDs
2. Create veth pair on host with temp names
3. Move ends into container network namespaces
4. Rename to final names and bring up
5. Query MACs and store in SimNode.ethernet_macs
"""
edges = self.topology.veth_edges()
if not edges:
return
image = self._get_image()
log.info("Setting up %d veth pairs (helper image: %s)...", len(edges), image)
for a, b in edges:
self._create_veth_pair(a, b, image)
log.info(
"Veth setup complete: %d pairs",
len(self._host_pairs),
)
def setup_node(self, node_id: str, down_nodes: set[str] | None = None):
"""Re-create veth endpoints for a single node after container restart.
When a container restarts (node churn), its network namespace is
destroyed. We re-create the veth pairs for all Ethernet edges
involving this node. ``down_nodes`` names the neighbours churn has
stopped, whose pairs are left for their own restart.
"""
image = self._get_image()
for a, b in self.topology.veth_edges():
if a != node_id and b != node_id:
continue
# Remove existing pair if any (host-side might still exist)
host_a = self.topology.veth_host_name(a, b, "a")
_run_host(["ip", "link", "delete", host_a], image, check=False)
# Re-create
self._create_veth_pair(a, b, image, down_nodes or set())
def teardown_all(self):
"""Clean up all veth pairs."""
image = self._get_image()
for _, _, host_a, _ in self._host_pairs:
_run_host(["ip", "link", "delete", host_a], image, check=False)
self._host_pairs.clear()
def _create_veth_pair(
self, node_a: str, node_b: str, image: str, down_nodes: set[str] | None = None
):
"""Create a single veth pair between two containers.
``down_nodes`` is None at first setup, when every container must be
running. On a restore it names the nodes churn has stopped.
"""
container_a = self.topology.container_name(node_a)
container_b = self.topology.container_name(node_b)
pid_a = _container_pid(container_a)
pid_b = _container_pid(container_b)
if pid_a is None or pid_b is None:
stopped = [n for n, pid in ((node_a, pid_a), (node_b, pid_b)) if pid is None]
names = ", ".join(stopped)
if down_nodes is None:
raise VethSetupError(
f"veth {node_a}--{node_b}: {names} not running at setup"
)
if all(n in down_nodes for n in stopped):
# Stopped by churn: it has no namespace to join, and its own
# restart recreates this pair.
log.info("Veth %s--%s deferred: %s down", node_a, node_b, names)
else:
# Not raised: a container that exited on its own is a daemon
# failure, and aborting here would report it as the harness's.
# Nothing recreates this pair, so say so loudly.
log.warning(
"Veth %s--%s not recreated: %s not running, and churn did "
"not stop all of them; the link stays absent for the rest "
"of the run",
node_a, node_b, names,
)
return
# Generate names
host_a = self.topology.veth_host_name(node_a, node_b, "a")
host_b = self.topology.veth_host_name(node_a, node_b, "b")
final_a = veth_interface_name(node_a, node_b)
final_b = veth_interface_name(node_b, node_a)
# Clear both final names and both temporary names out of the
# containers first. A stopped node's network namespace can outlive
# the stop by minutes, and while it does the survivor still holds its
# old `ve-X-Y`, whose peer sits in that namespace; the rename below
# then fails with "File exists". Deleting the survivor's end removes
# its peer too. A temporary name is left behind only by a restore
# that failed after the move, and it blocks the next move the same
# way.
_purge_links(container_a, [final_a, host_a])
_purge_links(container_b, [final_b, host_b])
# Clean up a stale pair left by this scenario. The token makes the
# name unique to this run, so a pair orphaned by an earlier run is
# no longer reclaimed here — `ci-cleanup.sh` reaps those instead.
_run_host(["ip", "link", "delete", host_a], image, check=False)
_require_host(
["ip", "link", "add", host_a, "type", "veth", "peer", "name", host_b],
image,
)
_require_host(["ip", "link", "set", host_a, "netns", str(pid_a)], image)
_require_host(["ip", "link", "set", host_b, "netns", str(pid_b)], image)
# A udp-veth edge carries IP: address each end before it comes up,
# so the daemon's interface wait never sees the interface without
# its address.
ip_a = ip_b = None
if self.topology.transport_for_edge(node_a, node_b) == UDP_VETH:
ip_a = self.topology.udp_veth_ip(node_a, node_b)
ip_b = self.topology.udp_veth_ip(node_b, node_a)
_raise_link(container_a, host_a, final_a, ip_a)
_raise_link(container_b, host_b, final_b, ip_b)
_await_up(container_a, final_a)
_await_up(container_b, final_b)
# Read only after both ends are proven renamed and up. Before, a
# failed rename left the old interface under the final name, and this
# read its MAC and reported the restore as a success.
mac_a = _get_mac_in_container(container_a, final_a)
mac_b = _get_mac_in_container(container_b, final_b)
if not mac_a or not mac_b:
raise VethSetupError(
f"veth {node_a}--{node_b} is up but its MAC could not be read "
f"({final_a}: {mac_a or '?'}, {final_b}: {mac_b or '?'})"
)
self.topology.nodes[node_a].ethernet_macs[node_b] = mac_a
self.topology.nodes[node_b].ethernet_macs[node_a] = mac_b
self._host_pairs.append((node_a, node_b, host_a, host_b))
log.info(
"Veth %s(%s) -- %s(%s) MAC: %s / %s",
node_a, final_a, node_b, final_b, mac_a, mac_b,
)
def _in_container(container: str, cmd: str, what: str, timeout: int = EXEC_TIMEOUT_SECS) -> str:
"""Run a command inside a container, raising `VethSetupError` on failure."""
try:
return docker_exec(container, cmd, timeout=timeout)
except (DockerExecError, subprocess.TimeoutExpired) as e:
raise VethSetupError(f"{what} in {container} failed: {e}") from e
def _purge_links(container: str, names: list[str]):
"""Delete each named interface in a container if it exists, in one exec.
An absent name is the ordinary case and is skipped. A delete that fails
raises, since the rename that follows would then fail on the name too,
unless the name is gone by then: a lingering namespace can be reaped
between the check and the delete, taking the survivor's end with it.
"""
script = "; ".join(
f"if ip link show {name} >/dev/null 2>&1; then "
f"ip link delete {name} || ! ip link show {name} >/dev/null 2>&1 || exit 1; fi"
for name in names
)
_in_container(container, script, f"deleting stale {', '.join(names)}")
def _require_host(cmd: list[str], image: str):
"""Run a host-namespace ``ip`` command, raising `VethSetupError` on failure."""
if not _run_host(cmd, image):
raise VethSetupError(f"host command failed: {' '.join(cmd)}")
def _raise_link(container: str, temp: str, final: str, ip: str | None = None):
"""Rename a moved veth end to its final name, address it if asked,
and set it up."""
addr = f" && ip addr add {ip}/24 dev {final}" if ip else ""
_in_container(
container,
f"ip link set {temp} name {final}{addr} && ip link set {final} up",
f"renaming {temp} to {final}",
)
def _await_up(container: str, iface: str):
"""Wait for an interface's operstate to read ``up``, else raise.
A veth end reports up only once both ends are up, so this proves the
pair is joined as well as that this end was raised.
"""
deadline = time.monotonic() + OPERSTATE_WAIT_SECS
state = None
reads = 0
while True:
state = docker_exec_quiet(
container, f"cat /sys/class/net/{iface}/operstate", timeout=EXEC_TIMEOUT_SECS
)
state = state.strip() if state is not None else None
reads += 1
if state == "up":
return
if reads >= 2 and time.monotonic() >= deadline:
raise VethSetupError(
f"{iface} in {container} reads operstate {state or '?'}, "
f"not up, {OPERSTATE_WAIT_SECS}s after it was raised"
)
time.sleep(0.2)
def _container_pid(container: str) -> int | None:
"""Return a container's PID, or None when docker reports it not running.
Raises `VethSetupError` when docker cannot be asked. A timed-out or failed
inspect of a running container used to read as "not running", so under
host load a live link was skipped and never recreated.
"""
try:
result = subprocess.run(
["docker", "inspect", "-f", "{{.State.Pid}}", container],
capture_output=True,
text=True,
timeout=EXEC_TIMEOUT_SECS,
)
except subprocess.TimeoutExpired as e:
raise VethSetupError(f"docker inspect {container} timed out") from e
if result.returncode != 0:
raise VethSetupError(
f"docker inspect {container} failed: {result.stderr.strip()}"
)
try:
pid = int(result.stdout.strip())
except ValueError as e:
raise VethSetupError(
f"docker inspect {container} returned no PID: {result.stdout.strip()!r}"
) from e
return pid if pid > 0 else None
def _get_mac_in_container(container: str, iface: str) -> str | None:
"""Query the MAC address of an interface inside a container."""
result = docker_exec_quiet(
container,
f"cat /sys/class/net/{iface}/address",
timeout=EXEC_TIMEOUT_SECS,
)
if result is not None:
return result.strip()
return None
def _run_host(cmd: list[str], image: str, check: bool = True) -> bool:
"""Run an ``ip`` command via a privileged Docker container.
Uses ``--net=host --pid=host --privileged`` so the container shares
the Docker host's (or Docker Desktop VM's) network and PID
namespaces. This makes ``ip link set ... netns <pid>`` work
correctly on both Linux and macOS.
``image`` should be a Docker image that has ``iproute2`` installed
(e.g. the simulation's own image built from the chaos Dockerfile).
``--entrypoint ip`` overrides the image's default entrypoint so the
simulation entrypoint script is not executed.
"""
docker_cmd = [
"docker", "run", "--rm",
"--privileged",
"--net=host",
"--pid=host",
"--entrypoint", "ip",
image,
] + cmd[1:] # cmd[0] is "ip", skip it since it's now the entrypoint
try:
result = subprocess.run(
docker_cmd,
capture_output=True,
text=True,
timeout=30,
)
if check and result.returncode != 0:
# Warning, not debug: the runner logs at INFO unless asked for
# -v, so at debug this never reached runner.log and the callers
# below report only that a pair could not be created. The
# check=False deletes are expected to fail and stay silent.
log.warning(
"ip cmd failed: %s -> %s",
" ".join(cmd),
result.stderr.strip(),
)
return False
return result.returncode == 0
except subprocess.TimeoutExpired:
log.warning("ip cmd timed out: %s", " ".join(cmd))
return False