Files
fips/testing/chaos/sim/assertions.py
ArjenandJohnathan Corgan 95866b7c7c test(iface-binding): cover the presence machine end to end
Two daemons whose only transports are interface-bound, run against a veth
pair the harness creates, downs, deletes and recreates underneath them.

Asserts the boot race (a daemon whose only interface is missing starts,
reports the transport absent and the node Degraded, rather than exiting on
NoTransports or skipping the transport for the life of the process), the
late attach and discovery over it, the flap in both directions,
destroy-and-recreate, and that an optional interface which never appears
never moves node health.

Also the log policy, which is the half that is easy to regress silently:
absence is logged once on the edge and not once per retry; a required
interface still absent past the ten-second bring-up window errors exactly
once, while the optional one — absent just as long — stays silent; and that
error is not repeated on a schedule. The detach edge is checked not to
error, guarded by how long detection actually took, so a slow runner skips
the check rather than failing on the harness's own latency.

The containers run FIPS_TEST_MODE=default, not chaos. The chaos entrypoint
waits up to 30 s for every configured Ethernet interface before starting the
daemon, which is precisely the workaround under test — the daemon has to do
its own waiting here or the suite proves nothing.

Host-namespace ip(8) runs in a short-lived privileged container sharing the
host network and PID namespaces, for the reason chaos/sim/veth.py documents:
on macOS the containers live in the Docker VM, so ip(8) run on the macOS
host could never reach them.

Chaos ethernet transports are marked optional: true. In that harness a
neighbour's interface disappearing is the scenario, not a fault — node_churn
stops a container, which destroys its netns and with it both ends of every
veth it held, so a surviving node watches a required interface vanish for
the 30-90 s the neighbour is down, once per churn event. Reporting that at
error is right for a deployment and wrong for a harness that tears the
interface down on purpose; the mesh-wide zero-ERROR ceiling would have
failed on injected chaos rather than on a defect.

test(iface-binding): cover an interface present before the daemon starts

Every scenario in the suite created its interface after the daemons were
already running — that ordering is the boot race the suite was written for.
But it means both nodes could only ever reach Present through binder_loop,
so the inline bind in start_async, which is the ordinary case on a booted
router, had no end-to-end coverage at all. That is where the churn guard
went unseeded and the first detach stopped reaching node health, and no
existing case could reach it: they all detach from a binding the loop
created, which seeds the guard as a side effect.

Case (f) adds a third node whose single required interface exists before its
daemon does. The gate is what buys that ordering — the harness needs a
running container to have a netns to move a veth into, but the daemon must
not start until after the move, so node-c comes up parked on a file and the
harness releases it once the interface is in place. Then one detach, on a
binding the loop did not create, and the node must degrade.

Verified against the defect rather than only against the fix: with the guard
seed reverted, cases (a) through (e) all still pass and (f) is the only
failure. A regression test that has never been seen to fail is a claim, not
a test.

It also asserts the reverse edge, so Degraded stays a level rather than a
latch on this path too.

test(iface-binding): assert the fast path and the churn guard

Three gaps, two of them in tests that existed and asserted nothing.

**The netlink path was never asserted to be in use.** The 1 s poll is a
complete fallback and covers every wait in the suite, so the whole thing
passed with `open_link_socket()` hardcoded to Err — the fast path could have
been dead for a release and no test would have said so. The binder reports
which backing it got at startup, so case (g) asks it directly rather than
inferring from timing the poll would also satisfy, and the unit test that used
to write `let _ = w.is_event_driven();` now asserts it on Linux, where the
source is an unprivileged `AF_NETLINK` socket and falling back is a real loss
rather than a sandbox's prerogative.

**Churn damping had no end-to-end coverage**, which now matters twice over: it
bounds the recovery announcements, and since the detach edge withdraws peers
it is also the only thing bounding how often that withdrawal fires. Every flap
elsewhere in the suite is a single down/up with long settles either side —
exactly the shape the damper ignores. Case (h) drives four bindings that each
die inside `MIN_STABLE_BINDING`, asserts the guard engages, asserts it then
*suppresses* rather than merely counting, and asserts it is not a latch.

**`a_poisoned_binding_does_not_strand_the_transport` discarded its result.**
`let _ = eth.binding.tasks_alive();` left the entire point unasserted: reading
a poisoned lock as "alive" would have the binder believe a dead binding
healthy and never rebind, and treating it as an error would strand the
transport. `false` is what routes it back through detach and rebind, so say so.

`a_stop_racing_a_bind_leaves_nothing_behind` now asserts the error *kind*.
`bind_and_spawn` refuses at its presence probe long before the post-store
shutdown check, so `is_err()` alone passed on absence and would still pass
with that check deleted. The test keeps the coverage it genuinely has — stop
raises the flag before teardown, teardown leaves no socket and no loops — and
says plainly that the race it is named for needs a bind that succeeds, which
needs privilege no unit test has.

Both new cases were verified against the defect: with the netlink source
forced to Err, (g) fails; with `CHURN_THRESHOLD` raised out of reach, (h)
fails. Nothing else in the suite notices either.

One case was attempted and removed rather than shipped: `"interface
replaced"` cannot be produced deterministically, because the delete that
changes an ifindex fires a netlink event the binder acts on within
microseconds, so `gone` wins the race. It passed about one run in three.
reference/notes.md records the measurement and the two approaches that could
work.

Also fixes a real bug in the harness: `grep -q` under `set -o pipefail` exits
on its first match, `docker logs` takes SIGPIPE, and the pipeline reports
failure even though the line was found. That cost two false failures before it
was spotted; `log_count` reads the stream to the end.

test(chaos): cover an Ethernet rebind under active traffic

The one case dynamic interface binding had no coverage for anywhere: a
datagram crossing an Ethernet link while the interface underneath it goes away
and comes back.

No existing scenario could reach it, for two separate reasons.

`ethernet-only` and `ethernet-mesh` both run with `traffic.enabled: false`, so
no datagram crosses an Ethernet link in any test — `ethernet-only`'s own
comment says exactly that, and names framing, the length field that trims NIC
minimum-frame padding, and AEAD over Ethernet as unexercised because of it.

And `link_flaps` cannot produce a rebind whatever it is pointed at: it
simulates a down link with netem 100% loss, so the interface stays IFF_UP and
the presence machine never sees an edge. `ethernet-mesh` has had link flaps
enabled all along without once exercising a rebind.

`node_churn` is what actually moves an interface. Stopping a container
destroys its network namespace, deleting every veth in it — and deleting one
end of a veth deletes its peer — so a *surviving* node watches its Ethernet
interface disappear outright, and watches it return when the harness recreates
the pair on restart. That is a real detach and a real rebind, driven from
outside the daemon.

The new scenario is a 4-node Ethernet ring with traffic on and one node
churned at a time, with link flaps deliberately off so the only outage is a
genuine interface removal and a traffic shortfall cannot be ambiguous between
the two. Measured across four runs: 206-388 MB moved over Ethernet links while
interfaces were being taken away underneath.

It also needed an assertion that did not exist. Traffic results have always
been written to `iperf3-results.json` and never read, so a scenario carrying
`traffic.enabled: true` could have every session fail and still exit 0 on a
green control plane — and a rebind under load is precisely what a tree
snapshot cannot see. `min_traffic` counts sessions that finished with bytes
actually received, treating iperf3's top-level `error` and a missing `end`
block as zero, so a session only counts when it moved data.

The baseline is calibrated against four runs rather than assumed: `max_roots`
starts at the observed maximum plus one, and the site records the sample, its
size, and why four runs is thin. The first draft asserted a single root and
failed every run — the harness restores stopped nodes immediately before the
final snapshot, so a just-restarted node has not re-parented yet and is
briefly its own root. That is the scenario working.

Wired into both runners, since a chaos scenario on one side only makes "local
green" and "GitHub green" stop meaning the same thing; check-ci-parity was
confirmed to fail on a one-sided addition before this was committed.

The iface-binding suite's entry in the GitHub workflow's integration matrix
moves here from the commit that introduced the presence machine. That commit
declared the suite on GitHub before testing/iface-binding/ existed and before
testing/ci-local.sh knew about it, so testing/check-ci-parity.sh failed there
and the three workflow steps named files that were not yet in the tree.
Registering both runners in the commit that adds the suite settles both.
2026-09-10 19:18:09 +00:00

504 lines
17 KiB
Python

"""Post-run scenario assertions evaluated via control-socket data.
Assertions are declared in the scenario YAML under ``assertions:`` and
are evaluated near the end of the simulation, before teardown begins.
Each failing assertion is recorded with a clear pass/fail message; the
runner exits non-zero when any assertion fails.
Currently supported assertions:
- ``bloom_send_rate``: per-node trailing-window ceiling on
``stats.bloom.sent`` delta. Calibrated for the bloom-storm
regression scenario but generally usable.
"""
from __future__ import annotations
import logging
from dataclasses import dataclass
from .control import snapshot_all_bloom
from .scenario import (
MinTrafficAssertion,
BaselineAssertion,
BloomSendRateAssertion,
CongestionSignalsAssertion,
MaxErrorsAssertion,
MaxParentSwitchesAssertion,
MinParentSwitchesAssertion,
TreeParentsAssertion,
)
from .topology import SimTopology
log = logging.getLogger(__name__)
@dataclass
class AssertionOutcome:
name: str
passed: bool
detail: str
def _bloom_sent_total(node_data: dict) -> int | None:
"""Extract stats.bloom.sent from a show_bloom response."""
stats = node_data.get("stats") or {}
sent = stats.get("sent")
if sent is None:
return None
try:
return int(sent)
except (TypeError, ValueError):
return None
class BloomSendRateMonitor:
"""Samples per-node ``stats.bloom.sent`` to evaluate a trailing-window
ceiling assertion at end-of-run.
Usage:
m = BloomSendRateMonitor(topology, cfg)
m.sample_window_start() # called window_secs before scenario end
...
m.sample_end() # called at scenario end
outcome = m.evaluate()
"""
def __init__(self, topology: SimTopology, cfg: BloomSendRateAssertion):
self.topology = topology
self.cfg = cfg
self.window_start: dict[str, int] = {}
self.window_end: dict[str, int] = {}
def sample_window_start(self) -> None:
snap = snapshot_all_bloom(self.topology)
for nid, data in snap.items():
v = _bloom_sent_total(data)
if v is not None:
self.window_start[nid] = v
def sample_end(self) -> None:
snap = snapshot_all_bloom(self.topology)
for nid, data in snap.items():
v = _bloom_sent_total(data)
if v is not None:
self.window_end[nid] = v
def evaluate(self) -> AssertionOutcome:
max_per_node = self.cfg.max_per_node
window_secs = self.cfg.window_secs
if not self.window_start or not self.window_end:
return AssertionOutcome(
name="bloom_send_rate",
passed=False,
detail=(
f"FAIL bloom_send_rate: failed to sample window endpoints "
f"(start={len(self.window_start)} nodes, "
f"end={len(self.window_end)} nodes)"
),
)
per_node_deltas: dict[str, int] = {}
for nid, end_v in self.window_end.items():
start_v = self.window_start.get(nid)
if start_v is None:
continue
per_node_deltas[nid] = end_v - start_v
offenders = {
nid: d for nid, d in per_node_deltas.items() if d > max_per_node
}
max_obs = max(per_node_deltas.values()) if per_node_deltas else 0
if offenders:
sorted_off = sorted(offenders.items(), key=lambda kv: -kv[1])
details = ", ".join(f"{nid}={d}" for nid, d in sorted_off)
detail = (
f"FAIL bloom_send_rate: {len(offenders)} node(s) exceeded "
f"ceiling of {max_per_node} bloom_sent over trailing "
f"{window_secs}s — offenders: {details} "
f"(all per-node deltas: "
f"{', '.join(f'{n}={v}' for n, v in sorted(per_node_deltas.items()))})"
)
return AssertionOutcome(
name="bloom_send_rate",
passed=False,
detail=detail,
)
detail = (
f"PASS bloom_send_rate: max per-node delta {max_obs} <= "
f"ceiling {max_per_node} over trailing {window_secs}s "
f"(per-node: "
f"{', '.join(f'{n}={v}' for n, v in sorted(per_node_deltas.items()))})"
)
return AssertionOutcome(
name="bloom_send_rate",
passed=True,
detail=detail,
)
def evaluate_max_parent_switches(
cfg: MaxParentSwitchesAssertion,
parent_switch_count: int,
scope: str,
) -> AssertionOutcome:
"""Stability ceiling on parent switches over the run.
``scope`` describes what was counted and appears in the message, so a
reader can tell a per-node result from a mesh-wide one. Resolving a
per-node scope to a real node is the caller's job, and so is failing
loudly when it cannot: an unresolvable node would count zero switches
and sail under any ceiling without having observed anything.
"""
if parent_switch_count <= cfg.max_total:
return AssertionOutcome(
name="max_parent_switches",
passed=True,
detail=(
f"PASS max_parent_switches: {parent_switch_count} switches "
f"({scope}) <= ceiling {cfg.max_total}"
),
)
return AssertionOutcome(
name="max_parent_switches",
passed=False,
detail=(
f"FAIL max_parent_switches: {parent_switch_count} switches "
f"({scope}) > ceiling {cfg.max_total} — the tree is reparenting "
f"more than the hysteresis band should allow. Check whether a "
f"cost change smaller than the hysteresis margin is still "
f"triggering a switch."
),
)
def evaluate_baseline(
cfg: BaselineAssertion,
snapshot: dict | None,
sessions: int,
) -> AssertionOutcome:
"""Floor on the mesh having formed: nodes answered, agreed a root, took parents."""
if not snapshot:
return AssertionOutcome(
name="baseline",
passed=False,
detail=(
"FAIL baseline: no final tree snapshot was taken, so nothing "
"about the mesh was observed. This is a harness failure."
),
)
reporting = len(snapshot)
roots = {v.get("root") for v in snapshot.values() if v.get("root")}
parented = sum(
1 for v in snapshot.values()
if v.get("parent") and v.get("parent") != v.get("my_node_addr")
)
parts, failures = [], []
def note(ok, text):
parts.append(text)
if not ok:
failures.append(text)
if cfg.min_nodes_reporting is not None:
note(reporting >= cfg.min_nodes_reporting,
f"{reporting} node(s) answered (need {cfg.min_nodes_reporting})")
if cfg.max_roots is not None:
note(len(roots) <= cfg.max_roots and len(roots) >= 1,
f"{len(roots)} distinct root(s) (allowed {cfg.max_roots})")
if cfg.min_nodes_parented is not None:
note(parented >= cfg.min_nodes_parented,
f"{parented} node(s) have a parent (need {cfg.min_nodes_parented})")
if cfg.min_sessions is not None:
note(sessions >= cfg.min_sessions,
f"{sessions} session(s) established (need {cfg.min_sessions})")
summary = "; ".join(parts)
if failures:
return AssertionOutcome(
name="baseline",
passed=False,
detail=(
f"FAIL baseline: {'; '.join(failures)}. Full: {summary}"
),
)
return AssertionOutcome(
name="baseline", passed=True, detail=f"PASS baseline: {summary}"
)
def evaluate_tree_parents(
cfg: TreeParentsAssertion,
snapshot: dict | None,
) -> AssertionOutcome:
"""Check each node's parent in the final tree snapshot.
Parents are compared by node address, resolved from the snapshot's own
``my_node_addr`` fields, so the check does not depend on the display
name a node happened to publish.
Every way of not knowing the answer is a failure: no snapshot, the
node absent from it, the expected parent absent from it, or the node
still claiming to be its own root. Each of those produces the same
"no match" that a genuinely wrong parent does, and only saying so
separately keeps a harness problem from reading as a routing verdict.
"""
if not snapshot:
return AssertionOutcome(
name="tree_parents",
passed=False,
detail=(
"FAIL tree_parents: no final tree snapshot was taken, so no "
"node's parent was observed. This is a harness failure, not "
"a statement about the tree."
),
)
addr_of = {
nid: data.get("my_node_addr")
for nid, data in snapshot.items()
if data.get("my_node_addr")
}
id_of = {addr: nid for nid, addr in addr_of.items()}
good, bad = [], []
for child, want_parent in sorted(cfg.expected.items()):
entry = snapshot.get(child)
if entry is None:
bad.append(
f"{child} is absent from the snapshot ({len(snapshot)} node(s) "
f"present: {', '.join(sorted(snapshot))})"
)
continue
want_addr = addr_of.get(want_parent)
if want_addr is None:
bad.append(
f"{child}: expected parent {want_parent} is absent from the "
f"snapshot, so its address cannot be resolved"
)
continue
got_addr = entry.get("parent")
if got_addr == entry.get("my_node_addr"):
bad.append(
f"{child} is its own parent — it still believes it is root, "
f"so the tree never converged around it (wanted {want_parent})"
)
continue
if got_addr == want_addr:
good.append(f"{child}->{want_parent}")
continue
got_id = id_of.get(got_addr) or entry.get("parent_display_name") or got_addr
bad.append(f"{child} chose {got_id}, wanted {want_parent}")
if bad:
detail = f"FAIL tree_parents: {'; '.join(bad)}"
if good:
detail += f". Correct: {', '.join(good)}"
return AssertionOutcome(name="tree_parents", passed=False, detail=detail)
return AssertionOutcome(
name="tree_parents",
passed=True,
detail=f"PASS tree_parents: {', '.join(good)}",
)
_CONGESTION_FLOORS = (
("min_nodes_detected", "congestion_detected"),
("min_nodes_ce_forwarded", "ce_forwarded"),
("min_nodes_ce_received", "ce_received"),
)
def evaluate_congestion_signals(
cfg: CongestionSignalsAssertion,
snapshot: dict | None,
) -> AssertionOutcome:
"""Floors on how many nodes observed each congestion counter.
``snapshot`` is the final congestion snapshot keyed by node id. None
means the snapshot never ran, which fails: a missing snapshot and a
mesh that observed no congestion produce the same zero counts, and
treating them alike is how an assertion comes to pass on an absence
of evidence.
"""
if not snapshot:
return AssertionOutcome(
name="congestion_signals",
passed=False,
detail=(
"FAIL congestion_signals: no final congestion snapshot was "
"taken, so no node was observed at all. This is a harness "
"failure, not a statement about congestion."
),
)
parts, failures = [], []
for attr, counter in _CONGESTION_FLOORS:
floor = getattr(cfg, attr)
if floor is None:
continue
hits = sorted(
nid for nid, data in snapshot.items()
if (data.get("congestion") or {}).get(counter, 0) > 0
)
parts.append(f"{counter}: {len(hits)} node(s) >0 (floor {floor})")
if len(hits) < floor:
failures.append(
f"{counter} non-zero on {len(hits)} node(s), need {floor}"
)
else:
parts[-1] += f" [{', '.join(hits)}]"
summary = "; ".join(parts)
if failures:
return AssertionOutcome(
name="congestion_signals",
passed=False,
detail=(
f"FAIL congestion_signals: {'; '.join(failures)} — across "
f"{len(snapshot)} node(s) sampled. Full counts: {summary}"
),
)
return AssertionOutcome(
name="congestion_signals",
passed=True,
detail=f"PASS congestion_signals: {summary}",
)
def evaluate_max_errors(
cfg: MaxErrorsAssertion,
errors: list[tuple[str, str]],
) -> AssertionOutcome:
"""Ceiling on ERROR-level log lines across the whole mesh.
``errors`` is the ``AnalysisResult.errors`` list of ``(source, line)``
pairs rather than a bare count, so a failure can name the nodes and
quote the lines. A ceiling breach that only reports a number sends the
reader back to the logs it was supposed to save them reading.
"""
count = len(errors)
if count <= cfg.max_total:
return AssertionOutcome(
name="max_errors",
passed=True,
detail=(
f"PASS max_errors: {count} ERROR line(s) mesh-wide <= "
f"ceiling {cfg.max_total}"
),
)
per_node: dict[str, int] = {}
for source, _line in errors:
per_node[source] = per_node.get(source, 0) + 1
worst = sorted(per_node.items(), key=lambda kv: -kv[1])
breakdown = ", ".join(f"{src}={n}" for src, n in worst)
samples = "\n".join(
f" [{src}] {line.strip()}" for src, line in errors[:5]
)
return AssertionOutcome(
name="max_errors",
passed=False,
detail=(
f"FAIL max_errors: {count} ERROR line(s) mesh-wide > ceiling "
f"{cfg.max_total} — per node: {breakdown}. First "
f"{min(5, count)}:\n{samples}"
),
)
def evaluate_min_parent_switches(
cfg: MinParentSwitchesAssertion,
parent_switch_count: int,
) -> AssertionOutcome:
"""Sanity guard: fail the scenario if the harness-induced flap did
not produce at least ``cfg.min_total`` parent switches across the
run. Detects misconfiguration (e.g., wrong root election) where
the bloom-rate assertion would otherwise trivially pass on any
binary including the regressed one.
"""
if parent_switch_count >= cfg.min_total:
return AssertionOutcome(
name="min_parent_switches",
passed=True,
detail=(
f"PASS min_parent_switches: {parent_switch_count} switches "
f"(mesh-wide) >= floor {cfg.min_total}"
),
)
return AssertionOutcome(
name="min_parent_switches",
passed=False,
detail=(
f"FAIL min_parent_switches: {parent_switch_count} switches "
f"(mesh-wide) < floor {cfg.min_total} — harness did not induce "
f"sufficient "
f"parent flapping; bloom-rate assertion would be trivially "
f"true. Check tree-snapshot-warmup.json: did the expected "
f"node win the root election?"
),
)
def _session_bytes(result: dict) -> int:
"""Bytes actually received in one iperf3 session, or 0 if it failed.
iperf3 reports a failed run as a top-level ``error`` string with no
``end`` block, and a killed-but-partial run still carries whatever it
managed. Both are handled by reading the received total and treating
anything missing as zero, so a session only counts when it moved bytes.
"""
if not isinstance(result, dict) or result.get("error"):
return 0
end = result.get("end")
if not isinstance(end, dict):
return 0
summary = end.get("sum_received") or end.get("sum_sent")
if not isinstance(summary, dict):
return 0
value = summary.get("bytes", 0)
return value if isinstance(value, int) and value > 0 else 0
def evaluate_min_traffic(
cfg: MinTrafficAssertion,
results: list[dict],
) -> AssertionOutcome:
"""Floor on iperf3 sessions that actually carried data.
Without this the traffic generator is decoration: the results were
written to disk and never read, so a scenario whose every session
failed still passed on a healthy control plane. A rebind under load is
exactly the case a tree snapshot cannot see.
"""
per_session = [_session_bytes(r) for r in results]
ok = [b for b in per_session if b > 0]
total = sum(ok)
if len(ok) >= cfg.min_sessions_ok and total >= cfg.min_bytes_total:
return AssertionOutcome(
name="min_traffic",
passed=True,
detail=(
f"PASS min_traffic: {len(ok)}/{len(results)} session(s) moved "
f"data (need {cfg.min_sessions_ok}); {total} byte(s) total "
f"(need {cfg.min_bytes_total})"
),
)
return AssertionOutcome(
name="min_traffic",
passed=False,
detail=(
f"FAIL min_traffic: {len(ok)}/{len(results)} session(s) moved data "
f"(need {cfg.min_sessions_ok}); {total} byte(s) total (need "
f"{cfg.min_bytes_total}). A green control plane with no traffic "
f"means the data path did not survive what the scenario did to it."
),
)