Files
fips/testing/chaos/scenarios/ethernet-churn.yaml
T
Arjen a8716ca974 test(chaos): cover an Ethernet rebind under active traffic
The one case dynamic interface binding had no coverage for anywhere: a
datagram crossing an Ethernet link while the interface underneath it goes away
and comes back.

No existing scenario could reach it, for two separate reasons.

`ethernet-only` and `ethernet-mesh` both run with `traffic.enabled: false`, so
no datagram crosses an Ethernet link in any test — `ethernet-only`'s own
comment says exactly that, and names framing, the length field that trims NIC
minimum-frame padding, and AEAD over Ethernet as unexercised because of it.

And `link_flaps` cannot produce a rebind whatever it is pointed at: it
simulates a down link with netem 100% loss, so the interface stays IFF_UP and
the presence machine never sees an edge. `ethernet-mesh` has had link flaps
enabled all along without once exercising a rebind.

`node_churn` is what actually moves an interface. Stopping a container
destroys its network namespace, deleting every veth in it — and deleting one
end of a veth deletes its peer — so a *surviving* node watches its Ethernet
interface disappear outright, and watches it return when the harness recreates
the pair on restart. That is a real detach and a real rebind, driven from
outside the daemon.

The new scenario is a 4-node Ethernet ring with traffic on and one node
churned at a time, with link flaps deliberately off so the only outage is a
genuine interface removal and a traffic shortfall cannot be ambiguous between
the two. Measured across four runs: 206-388 MB moved over Ethernet links while
interfaces were being taken away underneath.

It also needed an assertion that did not exist. Traffic results have always
been written to `iperf3-results.json` and never read, so a scenario carrying
`traffic.enabled: true` could have every session fail and still exit 0 on a
green control plane — and a rebind under load is precisely what a tree
snapshot cannot see. `min_traffic` counts sessions that finished with bytes
actually received, treating iperf3's top-level `error` and a missing `end`
block as zero, so a session only counts when it moved data.

The baseline is calibrated against four runs rather than assumed: `max_roots`
starts at the observed maximum plus one, and the site records the sample, its
size, and why four runs is thin. The first draft asserted a single root and
failed every run — the harness restores stopped nodes immediately before the
final snapshot, so a just-restarted node has not re-parented yet and is
briefly its own root. That is the scenario working.

Wired into both runners, since a chaos scenario on one side only makes "local
green" and "GitHub green" stop meaning the same thing; check-ci-parity was
confirmed to fail on a one-sided addition before this was committed.
2026-09-02 10:50:41 +01:00

129 lines
4.7 KiB
YAML

# Ethernet rebind under active traffic
#
# The one case dynamic interface binding has no coverage for anywhere:
# traffic crossing an Ethernet link while the interface underneath it goes
# away and comes back.
#
# The other Ethernet scenarios cannot reach it. `ethernet-only` and
# `ethernet-mesh` both run with `traffic.enabled: false`, so no datagram
# crosses an Ethernet link in either — framing, the length field that trims
# NIC minimum-frame padding, and AEAD over Ethernet are all control-plane
# assumptions there. And `link_flaps` cannot produce a rebind whatever it is
# pointed at: it simulates a down link with netem 100% loss, so the interface
# stays IFF_UP and the presence machine never sees an edge.
#
# `node_churn` is what actually moves an interface. Stopping a container
# destroys its network namespace, which deletes every veth in it — and
# deleting one end of a veth deletes its peer — so a *surviving* node watches
# its Ethernet interface disappear outright. On restart the harness recreates
# the pair (`NodeChurnManager._start_node`), and the survivor watches it come
# back. That is a real detach and a real rebind, driven from outside the
# daemon, with iperf3 running across the mesh throughout.
#
# Topology: a 4-node ring, so removing any single node leaves the remaining
# three connected in a line and `protect_connectivity` has something to
# protect.
#
# n01 ---eth--- n02
# | |
# eth eth
# | |
# n04 ---eth--- n03
scenario:
name: "ethernet-churn"
seed: 42
duration_secs: 240
topology:
algorithm: explicit
num_nodes: 4
default_transport: ethernet
params:
adjacency:
- [n01, n02]
- [n02, n03]
- [n03, n04]
- [n04, n01]
# Mild, and deliberately so. The variable under test is the interface going
# away, not the link being bad while it is there; heavy loss here would make a
# traffic shortfall ambiguous between the two.
netem:
enabled: true
default_policy:
delay_ms: [1, 5]
jitter_ms: [0, 1]
loss_pct: [0, 0.5]
# Off on purpose. A netem-simulated down link would add outage that never
# reaches the presence machine, which is the opposite of what this isolates.
link_flaps:
enabled: false
traffic:
enabled: true
max_concurrent: 2
interval_secs: {min: 10, max: 20}
duration_secs: {min: 15, max: 25}
parallel_streams: 2
# The mechanism. One node down at a time, long enough to outlast the
# ten-second bring-up window so the absence is a real one rather than a race,
# and to give traffic time to run against the reduced mesh before it returns.
node_churn:
enabled: true
interval_secs: {min: 45, max: 60}
max_down_nodes: 1
down_duration_secs: {min: 20, max: 35}
protect_connectivity: true
assertions:
# The mesh re-forms after each interface comes back.
#
# Calibrated 2026-09-02 against four runs, all at this file's fixed seed 42
# and therefore an identical churn schedule, so the spread is container
# timing rather than differing scenarios:
#
# nodes answering 4 in all four runs
# distinct roots 2 in all four runs
# nodes parented 2 in all four runs
# traffic 3-5 sessions, 206-388 MB
#
# `max_roots: 1` was wrong and failed every run: the harness restores
# stopped nodes immediately before the final snapshot, so a node that has
# just restarted has not re-parented yet and is briefly its own root. That
# is the scenario working, not failing.
#
# The ceiling is 3, one step beyond the observed maximum of 2, with the
# parented floor at its complement. A mesh that genuinely collapsed — every
# node islanded — still fails, which is all this assertion is for. Do not
# read a pass as convergence.
#
# READ BEFORE RETUNING: four runs is a thin sample. `churn-mixed` documents
# what happens when a threshold is set one step outside a small one — it
# ends up inside the real distribution and reddens runs whatever the daemon
# does. Widen on evidence; tighten only against a much larger sample.
baseline:
min_nodes_reporting: 3
max_roots: 3
min_nodes_parented: 2
# The point of the scenario. Traffic must actually have crossed Ethernet
# links while interfaces were being taken away underneath it — a green tree
# with zero bytes moved is the failure this catches, and is exactly what a
# control-plane-only assertion would have called a pass.
min_traffic:
min_sessions_ok: 2
# Chaos Ethernet transports are `optional: true` (see config_gen), precisely
# because a neighbour's interface disappearing is the scenario here rather
# than a fault. So absence must stay silent: any ERROR means something other
# than the churn went wrong.
max_errors:
max_total: 0
logging:
rust_log: "info"
output_dir: "./sim-results"