test(chaos): cover an Ethernet rebind under active traffic

The one case dynamic interface binding had no coverage for anywhere: a
datagram crossing an Ethernet link while the interface underneath it goes away
and comes back.

No existing scenario could reach it, for two separate reasons.

`ethernet-only` and `ethernet-mesh` both run with `traffic.enabled: false`, so
no datagram crosses an Ethernet link in any test — `ethernet-only`'s own
comment says exactly that, and names framing, the length field that trims NIC
minimum-frame padding, and AEAD over Ethernet as unexercised because of it.

And `link_flaps` cannot produce a rebind whatever it is pointed at: it
simulates a down link with netem 100% loss, so the interface stays IFF_UP and
the presence machine never sees an edge. `ethernet-mesh` has had link flaps
enabled all along without once exercising a rebind.

`node_churn` is what actually moves an interface. Stopping a container
destroys its network namespace, deleting every veth in it — and deleting one
end of a veth deletes its peer — so a *surviving* node watches its Ethernet
interface disappear outright, and watches it return when the harness recreates
the pair on restart. That is a real detach and a real rebind, driven from
outside the daemon.

The new scenario is a 4-node Ethernet ring with traffic on and one node
churned at a time, with link flaps deliberately off so the only outage is a
genuine interface removal and a traffic shortfall cannot be ambiguous between
the two. Measured across four runs: 206-388 MB moved over Ethernet links while
interfaces were being taken away underneath.

It also needed an assertion that did not exist. Traffic results have always
been written to `iperf3-results.json` and never read, so a scenario carrying
`traffic.enabled: true` could have every session fail and still exit 0 on a
green control plane — and a rebind under load is precisely what a tree
snapshot cannot see. `min_traffic` counts sessions that finished with bytes
actually received, treating iperf3's top-level `error` and a missing `end`
block as zero, so a session only counts when it moved data.

The baseline is calibrated against four runs rather than assumed: `max_roots`
starts at the observed maximum plus one, and the site records the sample, its
size, and why four runs is thin. The first draft asserted a single root and
failed every run — the harness restores stopped nodes immediately before the
final snapshot, so a just-restarted node has not re-parented yet and is
briefly its own root. That is the scenario working.

Wired into both runners, since a chaos scenario on one side only makes "local
green" and "GitHub green" stop meaning the same thing; check-ci-parity was
confirmed to fail on a one-sided addition before this was committed.
This commit is contained in:
Arjen
2026-09-02 10:50:41 +01:00
parent ee5384f013
commit a8716ca974
6 changed files with 247 additions and 1 deletions
+3
View File
@@ -573,6 +573,9 @@ jobs:
- suite: ethernet-only
type: chaos
scenario: ethernet-only
- suite: ethernet-churn
type: chaos
scenario: ethernet-churn
- suite: tcp-mesh
type: chaos
scenario: tcp-mesh
+128
View File
@@ -0,0 +1,128 @@
# Ethernet rebind under active traffic
#
# The one case dynamic interface binding has no coverage for anywhere:
# traffic crossing an Ethernet link while the interface underneath it goes
# away and comes back.
#
# The other Ethernet scenarios cannot reach it. `ethernet-only` and
# `ethernet-mesh` both run with `traffic.enabled: false`, so no datagram
# crosses an Ethernet link in either — framing, the length field that trims
# NIC minimum-frame padding, and AEAD over Ethernet are all control-plane
# assumptions there. And `link_flaps` cannot produce a rebind whatever it is
# pointed at: it simulates a down link with netem 100% loss, so the interface
# stays IFF_UP and the presence machine never sees an edge.
#
# `node_churn` is what actually moves an interface. Stopping a container
# destroys its network namespace, which deletes every veth in it — and
# deleting one end of a veth deletes its peer — so a *surviving* node watches
# its Ethernet interface disappear outright. On restart the harness recreates
# the pair (`NodeChurnManager._start_node`), and the survivor watches it come
# back. That is a real detach and a real rebind, driven from outside the
# daemon, with iperf3 running across the mesh throughout.
#
# Topology: a 4-node ring, so removing any single node leaves the remaining
# three connected in a line and `protect_connectivity` has something to
# protect.
#
# n01 ---eth--- n02
# | |
# eth eth
# | |
# n04 ---eth--- n03
scenario:
name: "ethernet-churn"
seed: 42
duration_secs: 240
topology:
algorithm: explicit
num_nodes: 4
default_transport: ethernet
params:
adjacency:
- [n01, n02]
- [n02, n03]
- [n03, n04]
- [n04, n01]
# Mild, and deliberately so. The variable under test is the interface going
# away, not the link being bad while it is there; heavy loss here would make a
# traffic shortfall ambiguous between the two.
netem:
enabled: true
default_policy:
delay_ms: [1, 5]
jitter_ms: [0, 1]
loss_pct: [0, 0.5]
# Off on purpose. A netem-simulated down link would add outage that never
# reaches the presence machine, which is the opposite of what this isolates.
link_flaps:
enabled: false
traffic:
enabled: true
max_concurrent: 2
interval_secs: {min: 10, max: 20}
duration_secs: {min: 15, max: 25}
parallel_streams: 2
# The mechanism. One node down at a time, long enough to outlast the
# ten-second bring-up window so the absence is a real one rather than a race,
# and to give traffic time to run against the reduced mesh before it returns.
node_churn:
enabled: true
interval_secs: {min: 45, max: 60}
max_down_nodes: 1
down_duration_secs: {min: 20, max: 35}
protect_connectivity: true
assertions:
# The mesh re-forms after each interface comes back.
#
# Calibrated 2026-09-02 against four runs, all at this file's fixed seed 42
# and therefore an identical churn schedule, so the spread is container
# timing rather than differing scenarios:
#
# nodes answering 4 in all four runs
# distinct roots 2 in all four runs
# nodes parented 2 in all four runs
# traffic 3-5 sessions, 206-388 MB
#
# `max_roots: 1` was wrong and failed every run: the harness restores
# stopped nodes immediately before the final snapshot, so a node that has
# just restarted has not re-parented yet and is briefly its own root. That
# is the scenario working, not failing.
#
# The ceiling is 3, one step beyond the observed maximum of 2, with the
# parented floor at its complement. A mesh that genuinely collapsed — every
# node islanded — still fails, which is all this assertion is for. Do not
# read a pass as convergence.
#
# READ BEFORE RETUNING: four runs is a thin sample. `churn-mixed` documents
# what happens when a threshold is set one step outside a small one — it
# ends up inside the real distribution and reddens runs whatever the daemon
# does. Widen on evidence; tighten only against a much larger sample.
baseline:
min_nodes_reporting: 3
max_roots: 3
min_nodes_parented: 2
# The point of the scenario. Traffic must actually have crossed Ethernet
# links while interfaces were being taken away underneath it — a green tree
# with zero bytes moved is the failure this catches, and is exactly what a
# control-plane-only assertion would have called a pass.
min_traffic:
min_sessions_ok: 2
# Chaos Ethernet transports are `optional: true` (see config_gen), precisely
# because a neighbour's interface disappearing is the scenario here rather
# than a fault. So absence must stay silent: any ERROR means something other
# than the churn went wrong.
max_errors:
max_total: 0
logging:
rust_log: "info"
output_dir: "./sim-results"
+59
View File
@@ -19,6 +19,7 @@ from dataclasses import dataclass
from .control import snapshot_all_bloom
from .scenario import (
MinTrafficAssertion,
BaselineAssertion,
BloomSendRateAssertion,
CongestionSignalsAssertion,
@@ -442,3 +443,61 @@ def evaluate_min_parent_switches(
f"node win the root election?"
),
)
def _session_bytes(result: dict) -> int:
"""Bytes actually received in one iperf3 session, or 0 if it failed.
iperf3 reports a failed run as a top-level ``error`` string with no
``end`` block, and a killed-but-partial run still carries whatever it
managed. Both are handled by reading the received total and treating
anything missing as zero, so a session only counts when it moved bytes.
"""
if not isinstance(result, dict) or result.get("error"):
return 0
end = result.get("end")
if not isinstance(end, dict):
return 0
summary = end.get("sum_received") or end.get("sum_sent")
if not isinstance(summary, dict):
return 0
value = summary.get("bytes", 0)
return value if isinstance(value, int) and value > 0 else 0
def evaluate_min_traffic(
cfg: MinTrafficAssertion,
results: list[dict],
) -> AssertionOutcome:
"""Floor on iperf3 sessions that actually carried data.
Without this the traffic generator is decoration: the results were
written to disk and never read, so a scenario whose every session
failed still passed on a healthy control plane. A rebind under load is
exactly the case a tree snapshot cannot see.
"""
per_session = [_session_bytes(r) for r in results]
ok = [b for b in per_session if b > 0]
total = sum(ok)
if len(ok) >= cfg.min_sessions_ok and total >= cfg.min_bytes_total:
return AssertionOutcome(
name="min_traffic",
passed=True,
detail=(
f"PASS min_traffic: {len(ok)}/{len(results)} session(s) moved "
f"data (need {cfg.min_sessions_ok}); {total} byte(s) total "
f"(need {cfg.min_bytes_total})"
),
)
return AssertionOutcome(
name="min_traffic",
passed=False,
detail=(
f"FAIL min_traffic: {len(ok)}/{len(results)} session(s) moved data "
f"(need {cfg.min_sessions_ok}); {total} byte(s) total (need "
f"{cfg.min_bytes_total}). A green control plane with no traffic "
f"means the data path did not survive what the scenario did to it."
),
)
+9
View File
@@ -20,6 +20,7 @@ from .assertions import (
evaluate_max_errors,
evaluate_max_parent_switches,
evaluate_min_parent_switches,
evaluate_min_traffic,
evaluate_tree_parents,
)
from .compose import generate_compose
@@ -699,6 +700,7 @@ class SimRunner:
self.node_mgr.restore_all()
# Collect iperf3 throughput results before containers stop
iperf_results: list[dict] = []
if self.traffic_mgr:
iperf_results = self.traffic_mgr.collect_results()
if iperf_results:
@@ -758,6 +760,13 @@ class SimRunner:
if err_cfg is not None:
outcome = evaluate_max_errors(err_cfg, result.errors)
self.assertion_outcomes.append(outcome)
# Traffic. Evaluated even when no session completed, because
# "nothing ran" is the failure this exists to catch.
traffic_cfg = self.scenario.assertions.min_traffic
if traffic_cfg is not None:
outcome = evaluate_min_traffic(traffic_cfg, iperf_results)
self.assertion_outcomes.append(outcome)
if outcome.passed:
log.info("%s", outcome.detail)
else:
+45
View File
@@ -323,6 +323,27 @@ class MaxErrorsAssertion:
max_total: int = 0
@dataclass
class MinTrafficAssertion:
"""Floor on how much iperf3 traffic actually completed.
Traffic has always been generated and its results saved to
``iperf3-results.json``, but nothing read them: a scenario could carry
``traffic.enabled: true``, have every single session fail, and still
exit 0 on a green control plane. That gap matters most for the
scenarios where traffic is the point — a datagram crossing an Ethernet
link while its interface rebinds underneath is not observable in the
tree snapshot at all.
``min_sessions_ok`` counts sessions that finished with bytes actually
received. ``min_bytes_total`` is the aggregate floor across them; 0
disables it and leaves the session count as the only gate.
"""
min_sessions_ok: int = 1
min_bytes_total: int = 0
@dataclass
class AssertionsConfig:
"""Optional post-run assertions evaluated against control-socket data."""
@@ -334,6 +355,7 @@ class AssertionsConfig:
congestion_signals: CongestionSignalsAssertion | None = None
tree_parents: TreeParentsAssertion | None = None
baseline: BaselineAssertion | None = None
min_traffic: MinTrafficAssertion | None = None
@dataclass
@@ -414,6 +436,7 @@ _SECTION_KEYS = {
"assertions": {
"bloom_send_rate", "min_parent_switches", "max_parent_switches",
"max_errors", "congestion_signals", "tree_parents", "baseline",
"min_traffic",
},
"logging": {"rust_log", "output_dir"},
}
@@ -422,6 +445,7 @@ _ASSERTION_KEYS = {
"min_parent_switches": {"min_total"},
"max_parent_switches": {"max_total", "node"},
"max_errors": {"max_total"},
"min_traffic": {"min_sessions_ok", "min_bytes_total"},
"congestion_signals": {
"min_nodes_detected", "min_nodes_ce_forwarded", "min_nodes_ce_received",
},
@@ -739,6 +763,27 @@ def load_scenario(path: str) -> Scenario:
f"got {err_total}"
)
s.assertions.max_errors = MaxErrorsAssertion(max_total=err_total)
if "min_traffic" in asrt:
mt = asrt["min_traffic"]
_reject_unknown(
mt, _ASSERTION_KEYS["min_traffic"], "assertions.min_traffic",
)
sessions = mt.get("min_sessions_ok", 1)
if not isinstance(sessions, int) or isinstance(sessions, bool) or sessions < 1:
raise ValueError(
"assertions.min_traffic: min_sessions_ok must be a positive "
f"integer, got {sessions!r} — a floor of zero asserts nothing"
)
min_bytes = mt.get("min_bytes_total", 0)
if not isinstance(min_bytes, int) or isinstance(min_bytes, bool) or min_bytes < 0:
raise ValueError(
"assertions.min_traffic: min_bytes_total must be a "
f"non-negative integer, got {min_bytes!r}"
)
s.assertions.min_traffic = MinTrafficAssertion(
min_sessions_ok=sessions, min_bytes_total=min_bytes,
)
else:
# Default-on. See MaxErrorsAssertion for why this one assertion is
# applied without being asked for: it is the floor on what a green
+3 -1
View File
@@ -30,7 +30,8 @@
# firewall, nat-cone, nat-symmetric,
# nat-lan, nostr-publish-consume, stun-faults,
# chaos-churn-mixed-10, chaos-ethernet-mesh,
# chaos-ethernet-only, chaos-tcp-mesh, chaos-congestion-stress,
# chaos-ethernet-only, chaos-ethernet-churn, chaos-tcp-mesh,
# chaos-congestion-stress,
# sidecar, dns-resolver, deb-install
#
# Opt-in (require --with-tor; depend on live Tor network):
@@ -149,6 +150,7 @@ CHAOS_SUITES=(
"churn-mixed-10 churn-mixed --nodes 10 --duration 120"
"ethernet-mesh ethernet-mesh"
"ethernet-only ethernet-only"
"ethernet-churn ethernet-churn"
"tcp-mesh tcp-mesh"
"congestion-stress congestion-stress"
)