Files
fips/testing/chaos
Johnathan Corgan 7cbe1d3d4e Retire the smoke-10 chaos scenario; convergence is covered in-process
smoke-10 was a no-stressor 10-node tree-convergence sanity check (netem
off, no ping). Its subject, spanning-tree convergence and root election,
is now covered in-process, faster and deterministically, by the loopback
spanning-tree harness (src/node/tests/spanning_tree.rs: ring, star,
chain, 100-node and disconnected-component convergence) plus end-to-end
datagram delivery (src/node/tests/forwarding.rs). Real-UDP convergence
smoke still runs via static-mesh and the other scenarios' baseline
assertions, so no Docker coverage is lost.

Drop it from both runners in lockstep (the CHAOS_SUITES list and the
GitHub chaos matrix) so the parity guard stays green, delete the scenario
YAML, and update the chaos README.
2026-07-24 01:43:20 +00:00
..

Stochastic Network Simulation

Automated network testing for FIPS. Generates random or explicit topologies, spins up Docker containers, and applies configurable stressors (network impairment, link flaps, traffic generation, node churn) over a timed simulation run. Scenarios cover general stress testing, cost-based parent selection, mixed link technologies (fiber/Bluetooth/WiFi), and transport-specific validation (UDP, TCP, Ethernet). Logs are collected and analyzed automatically.

Prerequisites

  • Docker with the compose plugin
  • Rust toolchain (for building the FIPS binary)
  • Python 3 with pyyaml and jinja2 packages

Quick Start

./testing/chaos/scripts/build.sh
./testing/chaos/scripts/chaos.sh churn-mixed

Available Scenarios

General stress tests

Random topologies with increasing stressor intensity.

Scenario Nodes Topology Duration Netem Link Flaps Traffic Node Churn Bandwidth
chaos-10 10 random_geometric 120s yes yes yes -- --
churn-10 10 random_geometric 600s yes yes yes yes --
churn-20 20 erdos_renyi 600s yes yes yes yes yes
  • chaos-10: Network degradation (5-50ms delay, 0-2% loss), link flaps (max 2 down, 10-30s), and iperf traffic (max 3 concurrent). Netem mutates 30% of links every 15-30s between normal and degraded policies.
  • churn-10: Extended run with node churn (1 node down at a time, 30-90s). Tests tree re-convergence after node departure/rejoin.
  • churn-20: Aggressive scale test. Erdos-Renyi topology, up to 5 nodes down simultaneously, bandwidth tiers (1/10/100/1000 Mbps), protect_connectivity disabled (partitions allowed).

Cost-based parent selection — retired, now sans-IO unit tests

The cost-selection scenarios (cost-avoidance, depth-vs-cost, bottleneck-parent, cost-reeval, cost-stability, mixed-technology) were retired on 2026-07-23. Their subject was the pure TreeState::evaluate_parent decision — which parent wins on effective_depth = depth + link_cost, when periodic re-evaluation switches, and when hysteresis suppresses a flap. A Docker mesh could not test that reliably: the root is whichever node holds the smallest NodeAddr, MMP costs take several measurement windows to settle, and hold-down plus hysteresis timing all confound the outcome (a deterministic link_swap attempt still produced zero periodic switches in a full run).

That logic is now covered by deterministic sans-IO unit tests in src/tree/tests.rs (test_evaluate_parent_cost_*, ..._hysteresis_*, ..._effective_depth_*), which run in the cargo quartet on every commit and can each be shown to fail by breaking the cost or hysteresis logic.

Transport-specific

Explicit topologies exercising non-UDP transports.

Scenario Nodes Transport Shape Duration Netem Link Flaps What it tests
ethernet-only 4 Ethernet Ring 90s yes -- AF_PACKET transport with beacon discovery
ethernet-mesh 6 UDP + Ethernet Mesh 120s yes yes Mixed UDP/Ethernet, netem mutation + flaps
tcp-only 4 TCP Ring 90s yes -- TCP transport with static peer config
tcp-chain 4 TCP Chain 90s yes -- TCP multi-hop routing through chain
tcp-mesh 6 UDP + TCP Mesh 120s yes yes Mixed UDP/TCP, netem mutation + flaps
  • ethernet-only: 4-node ring on raw Ethernet (AF_PACKET). Peers discovered via beacons, not static config. Minimal netem (1-5ms delay).
  • ethernet-mesh: Mirrors tcp-mesh topology but with Ethernet instead of TCP. UDP edges use static config; Ethernet edges use beacon discovery.
  • tcp-only: 4-node ring using TCP on port 8443. Tests connect-on-send, FMP framing over TCP, and reconnection. Netem enabled (1-10ms delay, 0-1% loss).
  • tcp-chain: 4-node linear chain, all TCP. Tests multi-hop routing over TCP-only mesh.
  • tcp-mesh: 6-node mesh with 4 UDP and 3 TCP edges. Both transports use static peer config. Netem mutation (30% fraction, every 20-40s) and link flaps (1 link max, 10-20s down).

Congestion and ECN

Scenarios testing ECN congestion signaling and transport-level congestion detection.

Scenario Nodes Topology Duration What it tests
congestion-stress 10 Tree 120s CE marking under kernel drops and MMP loss detection
ecn-ab-on / ecn-ab-off 6 Tree 120s A/B throughput comparison: ECN enabled vs disabled
  • congestion-stress: 10-node tree with 1 Mbps egress bandwidth caps, 5-10% netem loss, and heavy iperf3 traffic. Ingress policing (1000 kbps) and small recv_buf_size (4 KB) trigger both MMP loss detection and SO_RXQ_OVFL kernel socket drops. Validates end-to-end CE propagation: transit nodes detect congestion, set CE flag, destinations receive CE-marked packets, ecn_ce_count reported in MMP.
  • ecn-ab-on / ecn-ab-off: Paired scenarios with identical conditions (6-node tree, 10 Mbps egress, 1000 kbps ingress policing, 10ms link delay, 8 KB recv buffer) differing only in ecn.enabled. ecn-ab-compare.sh runs both and prints a side-by-side of throughput and congestion counters. It is a manual tool, not a test: it asserts nothing and no runner invokes it. The "+10.2% recv throughput with ECN enabled" figure once recorded here is not reproducible from anything on disk — the script read a fixed sim-results/ecn-ab-on/ path while the runner has written timestamped directories since 2026-03-20, and no ecn-ab result directory survives. The path bug is fixed; the figure is left out until a run produces one.

Ingress Traffic Control

Scenarios can include ingress configuration to simulate upstream bandwidth bottlenecks using tc ingress policing:

ingress:
  enabled: true
  tiers_kbps: [1000]         # per-peer rate limit in kbps
  burst_bytes: 10000         # policer burst allowance

Per-peer u32 filters on the ingress qdisc (parent ffff:) rate-limit inbound packets. Combined with small recv_buf_size, this reliably triggers SO_RXQ_OVFL kernel socket drops for congestion detection testing.

iperf3 JSON Capture

Traffic sessions capture iperf3 results using --json output. Results are collected per-session from containers and saved as iperf3-results.json in the scenario output directory, enabling automated throughput analysis across scenario runs.

CLI Options

Option Description
-v, --verbose Enable debug logging
--seed N Override the scenario's random seed
--duration secs Override the scenario's duration
--list List available scenarios

The scenario argument accepts either a name (churn-10) or a file path (scenarios/churn-10.yaml).

Scenario YAML Format

Annotated example based on churn-10.yaml:

scenario:
  name: "churn-10"
  seed: 42                          # deterministic RNG seed
  duration_secs: 600                # total simulation time

topology:
  num_nodes: 10
  algorithm: random_geometric       # or erdos_renyi, chain
  params:
    radius: 0.5                     # algorithm-specific parameter
  ensure_connected: true            # retry until graph is connected
  subnet: "172.20.0.0/24"
  ip_start: 10                      # first node gets .10

netem:
  enabled: true
  default_policy:
    delay_ms: { min: 5, max: 50 }
    jitter_ms: { min: 1, max: 10 }
    loss_pct: { min: 0, max: 2 }
  mutation:
    interval_secs: { min: 20, max: 45 }  # re-roll interval
    fraction: 0.3                         # fraction of links mutated
    policies:                             # named policy profiles
      normal:
        delay_ms: [5, 20]
        loss_pct: [0, 1]
      degraded:
        delay_ms: [50, 100]
        jitter_ms: [10, 30]
        loss_pct: [3, 8]

link_flaps:
  enabled: true
  interval_secs: { min: 30, max: 60 }
  max_down_links: 2
  down_duration_secs: { min: 10, max: 30 }
  protect_connectivity: true        # never partition the graph

traffic:
  enabled: true
  max_concurrent: 3
  interval_secs: { min: 10, max: 30 }
  duration_secs: { min: 5, max: 15 }
  parallel_streams: 4

node_churn:
  enabled: true
  interval_secs: { min: 60, max: 180 }
  max_down_nodes: 1
  down_duration_secs: { min: 30, max: 90 }
  protect_connectivity: true        # never kill the last path

bandwidth:
  enabled: false                    # per-link HTB rate limiting
  tiers_mbps: [1, 10, 100, 1000]   # each link randomly assigned a tier

logging:
  rust_log: "debug"
  output_dir: "./sim-results"

Topology Algorithms

Algorithm Parameters Description
random_geometric radius (default 0.5) Place nodes in unit square, connect pairs within radius
erdos_renyi p (default 0.3) Include each edge independently with probability p
chain -- Linear chain: n01--n02--...--nN
explicit adjacency list Hardcoded edges with optional per-edge transport type

When ensure_connected is true (default), the generator retries up to 50 times to produce a connected graph.

Directed Outbound Configs

The config generator assigns each static-config edge (UDP or TCP) to exactly one node for outbound connection using a BFS spanning tree rooted at the lowest node ID. Tree edges are assigned parent-to-child; non-tree edges are assigned from the lower node ID to the higher. This eliminates the dual-connect race condition where both sides initiate simultaneously, and creates a clear "owning side" for each link — relevant for auto-reconnect testing. Ethernet edges are excluded from static config since they use beacon discovery.

Output

Results written to sim-results/ (configurable via logging.output_dir):

  • status.txt -- How the run ended, plus the scenario, the seed and the container names it used; one key=value per line
  • analysis.txt -- Summary: panics, errors, sessions, metrics
  • metadata.txt -- Seed, node count, edges, adjacency list
  • runner.log -- Orchestration events (topology, netem, churn, traffic) with timestamps
  • fips-node-nXX.log -- Per-node log output

The status field reads:

  • completed -- ran for its configured duration
  • interrupted -- a signal cut the run short, so the artifacts are real but describe less time than the scenario asked for
  • aborted -- the run raised part way through; same caveat, and runner.log carries the traceback
  • setup-failed -- the containers never started
  • teardown-failed -- the mesh ran but its logs or analysis could not be produced

A setup-failed directory holds runner.log and status.txt and nothing else. Nothing is harvested, because container names are global to the host and reading them after a failed setup describes whichever run holds them now. So analysis.txt in a result directory is proof that this scenario's own mesh existed. A directory with no status.txt was written before this was the case and says nothing either way.

Exit codes:

  • 0 -- Ran to completion, no panics, every assertion passed
  • 1 -- The scenario file could not be loaded, or a second interrupt arrived while the first was being handled
  • 2 -- Panics found in the collected node logs. Also what the argument parser exits with when it rejects the command line, before any run starts
  • 3 -- A post-run assertion failed
  • 4 -- Setup, warmup, the simulation loop or teardown raised, so the run did not complete; runner.log carries the traceback

Codes 2 and 3 describe what a mesh that ran did. Code 4 says there is nothing to describe, and takes precedence over both. Code 2 is dual-use: a run that never started cannot have panicked, so read it together with whether runner.log exists.

A run stopped by a signal exits on this same ladder rather than one of its own: what it collected before stopping is still worth reporting, and status.txt says it was cut short. chaos.sh reports 130 for a Ctrl-C of its own accord.

Creating Custom Scenarios

  1. Copy an existing scenario from scenarios/.
  2. Adjust topology size, algorithm, and stressor parameters.
  3. Run with ./testing/chaos/scripts/chaos.sh path/to/custom.yaml.