Merge master into next

Brings up the two release lines' work: the maint harness and guard
fixes, the documentation corrections, master's probe, onion and epoch
fixes, and both rebuilt changelog blocks.

Three conflicts.

The readme conflicted on the badge pair. Resolved by taking the Rust
badge that no longer asserts a version, since rust-toolchain.toml is the
only place that states one, and keeping this line's own v0.6.0-dev
status badge.

The peer machine conflicted, and the resolution is an adaptation rather
than a pick. Master deleted PeerMachine.remote_epoch on the grounds that
nothing read it, and that reasoning had to be re-derived here because
this line's machine is the XX rewrite and shares almost no text with it.
It holds. The shadow's only production write is inbound_msg3, which is
where XX crystallizes identity, so it is inbound-only exactly as the msg1
write was on the other lines; conn carries the same value written from
both legs, complete_handshake on the outbound one and
complete_handshake_msg3 on the inbound; the only read is the cutover
action payload, whose executor arm binds nothing; and the live consumer
reads conn_remote_epoch. So an initiator cutover, which runs on an
outbound machine, carried a zeroed epoch here too.

One thing differs and needed handling. This line has an `established`
constructor the others do not, and it writes the shadow and conn from the
same argument, which would have made the field direction-correct. Its
only caller is in the test module and its own doc comment calls the
machine inert, so it is a seam that is not wired yet rather than a
production path, and it does not rescue the field. Its assignment goes
with the rest; the parameter stays, because conn still needs the value.

The adaptation is folded into this merge rather than left to a follow-on,
because master's half of the same change reached peer_actions.rs through
a clean auto-merge. Keeping this line's field while accepting that
auto-merge would have left the machine emitting a payload field the
executor no longer has, which is a break that only the test build shows.

The changelog conflicted because both lines had rebuilt their unreleased
block. The Breaking section stays at the top untouched; Unreleased now
holds the ten entries that are this line's own; and the other two lines'
work sits below under 0.5.0 and 0.4.2 headings, neither dated, matching
how master already carries 0.4.2. Four entries existed on both sides in
branch-adapted form and were merged rather than picked, so each keeps the
rework's wording and this line's accuracy: the OpenWrt entry drops its IK
reference, the msg1 classifier keeps the promotion-state paragraph, the
SessionAck entry keeps the two XX-only exits, and the msg3 epoch entry
counts six sites here against master's five.
This commit is contained in:
Johnathan Corgan
2026-08-22 11:04:58 +01:00
24 changed files with 2254 additions and 922 deletions
+7
View File
@@ -58,6 +58,13 @@ jobs:
# major (FreeBSD:15:amd64) — pkg on other majors refuses the package.
runs-on: ubuntu-latest
needs: determine-versioning
# Successful runs of this job take 8 to 11 minutes. Five consecutive runs
# in August 2026 instead sat in the VM step for 70, 190, 360, 360 and 360
# minutes and ended cancelled, the last three at GitHub's own six-hour job
# ceiling. Nothing here bounded them. This bound is deliberately loose
# enough that a slow-but-working run still passes, and tight enough that a
# stall fails in half an hour instead of burning a runner for six.
timeout-minutes: 30
steps:
- uses: actions/checkout@d23441a48e516b6c34aea4fa41551a30e30af803 # v6
+1141 -783
View File
File diff suppressed because it is too large Load Diff
+1 -1
View File
@@ -2,7 +2,7 @@
![banner](docs/logos/fips_banner.png)
[![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE)
[![Rust](https://img.shields.io/badge/rust-1.85%2B-orange.svg)](https://www.rust-lang.org/)
[![Rust](https://img.shields.io/badge/rust-orange.svg)](https://www.rust-lang.org/)
[![Status](https://img.shields.io/badge/status-v0.6.0--dev-green.svg)](#status--roadmap)
A self-organizing encrypted mesh network built on Nostr identities,
+4 -2
View File
@@ -20,8 +20,10 @@ locations, lowest to highest priority:
All found files are loaded and merged in priority order. Values from higher
priority files override those from lower priority files. This allows a system
administrator to set site-wide defaults in `/etc/fips/fips.yaml` while
individual deployments override specific values in `./fips.yaml`.
administrator to set site-wide defaults in the priority 1 path above,
`/usr/local/etc/fips/fips.yaml` on macOS and `/etc/fips/fips.yaml` on other
Unix systems, while individual deployments override specific values in
`./fips.yaml`.
### CLI Option
+45 -9
View File
@@ -28,6 +28,11 @@ The whole exercise should take about ten minutes.
your stable nsec / npub
```
The diagram shows the Linux layout. On macOS the same three files
live under `/usr/local/etc/fips/`; read
[Where these files live](#where-these-files-live) before running any
command below.
After this tutorial your node will have:
- A keypair on disk that the daemon reuses across restarts.
@@ -36,6 +41,28 @@ After this tutorial your node will have:
- A clear understanding of which file holds the secret and how
to keep it that way.
## Where these files live
Every path in this tutorial is written in its Linux form. The macOS
package (`.pkg`) installs config and keys under
`/usr/local/etc/fips/` instead of `/etc/fips/`, so on macOS
substitute as you go:
| Linux / other Unix | macOS |
| --- | --- |
| `/etc/fips/fips.yaml` | `/usr/local/etc/fips/fips.yaml` |
| `/etc/fips/fips.key` | `/usr/local/etc/fips/fips.key` |
| `/etc/fips/fips.pub` | `/usr/local/etc/fips/fips.pub` |
`fipsctl keygen` writes to `/usr/local/etc/fips/` by default on
macOS. The daemon still probes `/etc/fips/fips.yaml` as a fallback,
so an existing install is not broken by an upgrade, but the macOS
packaging only installs files under `/usr/local/etc/fips/`. If a
macOS host already carries key files at the old `/etc/fips/` path,
the daemon uses the old key and warns rather than minting a new
identity; the migration recipe is in the
[how-to guide](../how-to/persistent-identity.md).
## Why a stable identity matters
In FIPS your Nostr keypair *is* your node's identity in the most
@@ -68,7 +95,8 @@ The daemon supports two ways of holding that keypair:
> identity unless you explicitly ask for one.
> - *Persistent*: the daemon reads (or, on first start,
> generates and writes) a keypair stored at
> `/etc/fips/fips.key`. The npub stays the same across
> `/etc/fips/fips.key` (`/usr/local/etc/fips/fips.key` on
> macOS). The npub stays the same across
> restarts, reboots, and reinstalls as long as that file is
> preserved. You take on the cost of protecting an on-disk
> secret in exchange for being addressable by a stable name.
@@ -117,9 +145,9 @@ Make a note of it. We expect this to change.
## Step 2: Enable persistent identity in the config
Open `/etc/fips/fips.yaml` and find the `node:` block. The
shipped default has the relevant fragment commented out; make it
look like this:
Open `/etc/fips/fips.yaml` (`/usr/local/etc/fips/fips.yaml` on
macOS) and find the `node:` block. The shipped default has the
relevant fragment commented out; make it look like this:
```yaml
node:
@@ -137,6 +165,9 @@ The daemon's behavior on the next restart:
`/etc/fips/fips.{key,pub}` with the correct file modes, and
use that.
The daemon derives the key directory from whichever config file it
loaded, so on macOS both files land in `/usr/local/etc/fips/`.
## Step 3: Restart the daemon
```sh
@@ -164,6 +195,9 @@ The daemon wrote two files:
sudo ls -l /etc/fips/fips.key /etc/fips/fips.pub
```
On macOS, list `/usr/local/etc/fips/fips.key` and
`/usr/local/etc/fips/fips.pub` instead.
Expect:
```text
@@ -239,7 +273,7 @@ old one will be stale.
persistent identity.
- **Where it lives.** `/etc/fips/fips.key` and
`/etc/fips/fips.pub`, mode `0600` and `0644`, owned
`root:root`.
`root:root`; under `/usr/local/etc/fips/` on macOS.
- **What to share.** `fips.pub` is public; `fips.key` is not.
- **What it buys you.** A npub other operators can add to their
`peers:` list once, and that addresses the services your node
@@ -250,11 +284,13 @@ old one will be stale.
If the post-restart npub does not match `fips.pub`:
- **Check file permissions.**
`sudo ls -l /etc/fips/fips.key`. If the mode is not `0600` or
the owner is not `root:root`, the daemon may have refused to
read it. Restore with
`sudo ls -l /etc/fips/fips.key`, or
`sudo ls -l /usr/local/etc/fips/fips.key` on macOS. If the mode
is not `0600` or the owner is not `root:root`, the daemon may
have refused to read it. Restore with
`sudo chmod 0600 /etc/fips/fips.key && sudo chown root:root
/etc/fips/fips.key`.
/etc/fips/fips.key`, substituting the macOS path where it
applies.
- **Check the journal.** `sudo journalctl -u fips -n 100` after
the restart will show one of:
- `Loaded persistent identity from key file path=...` — good.
+127 -2
View File
@@ -1088,6 +1088,15 @@ fn settled_text(name: &str, stage: &serde_json::Value) -> String {
Some(n) => format!("no reply to {n} requests"),
None => "no reply".to_string(),
},
"bloom_unconfirmed" => match stage.get("attempts").and_then(|v| v.as_u64()) {
Some(1) => {
"a peer filter claimed this address; 1 request went unanswered".to_string()
}
Some(n) => {
format!("a peer filter claimed this address; {n} requests went unanswered")
}
None => "a peer filter claimed this address; nothing answered for it".to_string(),
},
"already_pending" => "joined a lookup already in flight, which failed".to_string(),
other => other.replace('_', " "),
},
@@ -1287,9 +1296,34 @@ fn rtt_text(rtt: &serde_json::Value) -> String {
/// The closing paragraph, which must not claim reachability it did not observe.
fn overall_note(report: &serde_json::Value) -> Option<&'static str> {
if field(report, "overall") != "partial" {
return None;
match field(report, "overall") {
"partial" => partial_note(report),
"failed" => failed_note(report),
_ => None,
}
}
/// A failed probe's closing paragraph. Only the unconfirmed-claim case earns
/// one: it is the reading the operator cannot make from the reason alone, and
/// it is the one that otherwise reads as a network fault.
fn failed_note(report: &serde_json::Value) -> Option<&'static str> {
match field(report.get("discovery")?, "reason") {
"bloom_unconfirmed" => Some(
" A peer's bloom filter claimed this address and then no lookup\n\
\x20 answered for it. A filter cannot miss a key that IS on the\n\
\x20 mesh, so an absent address produces exactly this whenever some\n\
\x20 peer's filter false-positives on it -- the same finding the\n\
\x20 fast bloom_miss reports, reached the slow way. An address that\n\
\x20 is present but whose lookups are being lost produces it too.\n\
\x20 The wait says nothing about which; it is the ladder running to\n\
\x20 the end.",
),
_ => None,
}
}
/// The partial-verdict paragraph, keyed on what the rtt stage settled on.
fn partial_note(report: &serde_json::Value) -> Option<&'static str> {
match field(report.get("rtt")?, "reason") {
"no_report" => Some(
" The handshake completed, so our packets reached them. Nothing has\n\
@@ -1598,6 +1632,97 @@ mod tests {
assert_eq!(rows[0].text, "no peer filter holds this address");
}
/// A finished probe of an address that is not on the mesh, under a given
/// gate outcome. `bloom_miss` is the run where no filter claimed it;
/// `bloom_unconfirmed` is the run where one did and no lookup answered.
fn absent_key_report(claimed: bool) -> serde_json::Value {
if claimed {
serde_json::json!({
"overall": "failed",
"elapsed_ms": 17000,
"bloom": {"verdict": "ok", "reason": null, "elapsed_ms": 1000, "fanout": 1},
"discovery": {"verdict": "failed", "reason": "bloom_unconfirmed",
"elapsed_ms": 16000, "attempts": 4,
"attempt_timeouts_secs": [1, 2, 4, 8]},
"path": {"verdict": "skipped", "reason": "not_reached"},
"session": {"verdict": "skipped", "reason": "not_reached"},
"rtt": {"verdict": "skipped", "reason": "not_reached"},
})
} else {
serde_json::json!({
"overall": "failed",
"elapsed_ms": 1900,
"bloom": {"verdict": "failed", "reason": "bloom_miss", "elapsed_ms": 1900,
"fanout": null},
"discovery": {"verdict": "skipped", "reason": "not_reached"},
"path": {"verdict": "skipped", "reason": "not_reached"},
"session": {"verdict": "skipped", "reason": "not_reached"},
"rtt": {"verdict": "skipped", "reason": "not_reached"},
})
}
}
#[test]
fn the_two_absent_key_paths_read_differently_to_the_operator() {
let clean = absent_key_report(false);
let claimed = absent_key_report(true);
let clean_rows = stage_rows(&clean);
let clean_text = clean_rows[row_at(&clean_rows, "bloom")].text.clone();
let claimed_rows = stage_rows(&claimed);
let claimed_text = claimed_rows[row_at(&claimed_rows, "discovery")]
.text
.clone();
assert_eq!(clean_text, "no peer filter holds this address");
assert_ne!(
clean_text, claimed_text,
"the two paths must not print the same line"
);
assert!(
claimed_text.contains("claimed this address"),
"the slow line must say a filter claimed it: {claimed_text}"
);
assert!(
claimed_text.contains('4'),
"and how many requests went unanswered: {claimed_text}"
);
// The line `no_response` used to print. It says nothing about why a
// request went out, which is the whole finding here.
assert_ne!(
claimed_text, "no reply to 4 requests",
"the claimed path must not fall back to the bare no-reply wording"
);
// The closing paragraph is where the operator is told what the wait
// meant. Only the claimed path earns one.
let note = overall_note(&claimed).unwrap_or_default();
assert!(
note.contains("false-positive"),
"the note must name the mechanism: {note}"
);
assert!(
note.contains("bloom_miss"),
"and tie it to the fast answer: {note}"
);
assert!(overall_note(&clean).is_none());
}
#[test]
fn a_bare_no_response_still_reads_as_a_missing_reply() {
// The reason survives for the case it still describes: the gate's
// answer never arrived, so no claim was ever made.
let mut report = absent_key_report(true);
report["discovery"]["reason"] = serde_json::json!("no_response");
let rows = stage_rows(&report);
assert_eq!(
rows[row_at(&rows, "discovery")].text,
"no reply to 4 requests"
);
assert!(overall_note(&report).is_none());
}
#[test]
fn a_failed_path_keeps_the_stages_that_ran_after_it() {
// The path preview names no hop and the session succeeds anyway,
+1 -1
View File
@@ -483,7 +483,7 @@ impl Node {
handler and must never reach the executor"
);
}
PeerAction::SwapSendState { .. } => {
PeerAction::SwapSendState => {
// Initiator cutover: the live authoritative rekey-cadence
// path, routed here from `check_rekey` via
// `route_rekey_cadence` → `PeerEvent::RekeyConsume`; the
+20 -19
View File
@@ -448,8 +448,9 @@ pub(crate) enum PeerAction {
/// resolution; it must never reach the action executor.
ResolveCrossConnection { swap: bool },
/// Initiator-side rekey cutover: swap the published send-state to the pending
/// epoch.
SwapSendState { epoch: [u8; 8] },
/// epoch. The remote epoch is not carried here: `conn` is its sole carrier
/// and promotion reads it from there via `conn_remote_epoch`.
SwapSendState,
/// Complete an initiator-side rekey drain: retire the previous session slot
/// (drop its `peers_by_index`/decrypt-worker entry, free its index). The
/// executor reads the REAL previous index from `ActivePeer::complete_drain`
@@ -576,8 +577,6 @@ pub(crate) struct PeerMachine {
/// Pure handshake-phase bookkeeping (link/direction/indices/transport/
/// stored handshake bytes/epoch). Reused verbatim from the FMP state core.
conn: ConnectionState,
/// Remote startup epoch (establish-path-only; NOT in send-state).
remote_epoch: Option<[u8; 8]>,
/// A stored-handshake send failure was observed on this leg. The failure is
/// carried as a flag (not a `PeerState::Failed` transition) so retransmit
/// eligibility (`is_handshaking_sent_msg1`) survives until the
@@ -635,7 +634,6 @@ impl PeerMachine {
Some(id) => ConnectionState::outbound(link, id, now),
None => ConnectionState::outbound_anonymous(link, now),
},
remote_epoch: None,
send_failed: false,
rekey_in_progress: false,
rekey_our_index: None,
@@ -663,7 +661,6 @@ impl PeerMachine {
node_addr: None,
leg: None,
conn: ConnectionState::inbound(link, now),
remote_epoch: None,
send_failed: false,
rekey_in_progress: false,
rekey_our_index: None,
@@ -739,7 +736,6 @@ impl PeerMachine {
node_addr: Some(addr),
leg: None,
conn,
remote_epoch,
send_failed: false,
rekey_in_progress: false,
rekey_our_index: None,
@@ -1578,7 +1574,6 @@ impl PeerMachine {
// Identity crystallizes at msg3 on XX (WireOutcome carries only the node
// address + epoch; the full static key stays shell-side).
self.node_addr = Some(wire.peer_node_addr);
self.remote_epoch = wire.remote_epoch;
// Seed the leg's index from the event. On a fresh classification machine
// the index was allocated at msg1 shell-side and is not otherwise known
// here; the cross-connection and rekey-responder decisions read it back to
@@ -1887,9 +1882,7 @@ impl PeerMachine {
addr: peer,
kind: MaintainKind::Rekey(RekeyPhase::Draining),
};
let mut actions = vec![PeerAction::SwapSendState {
epoch: self.remote_epoch.unwrap_or_default(),
}];
let mut actions = vec![PeerAction::SwapSendState];
if let Some(idx) = self.conn.our_index() {
actions.push(PeerAction::RegisterDecryptSession { index: idx });
}
@@ -2527,7 +2520,7 @@ mod tests {
link: LinkId::new(7),
},
PeerAction::ResolveCrossConnection { swap: true },
PeerAction::SwapSendState { epoch: [1u8; 8] },
PeerAction::SwapSendState,
PeerAction::CompleteDrain { peer },
PeerAction::InvalidateSendState,
PeerAction::RegisterDecryptSession {
@@ -2573,7 +2566,7 @@ mod tests {
| PeerAction::SendLinkMessage { .. }
| PeerAction::PromoteToActive { .. }
| PeerAction::ResolveCrossConnection { .. }
| PeerAction::SwapSendState { .. }
| PeerAction::SwapSendState
| PeerAction::CompleteDrain { .. }
| PeerAction::InvalidateSendState
| PeerAction::RegisterDecryptSession { .. }
@@ -2627,7 +2620,12 @@ mod tests {
};
m.rekey_our_index = Some(SessionIndex::new(0x2222));
m.conn.set_our_index(SessionIndex::new(0x1111));
m.remote_epoch = Some([9u8; 8]);
// The remote startup epoch lives on the surviving carrier, written
// there by BOTH handshake legs (`complete_handshake` from msg2 on the
// outbound leg, `complete_handshake_msg3` from msg3 on the inbound one)
// through this same setter. Seed it the way production does, so the
// value asserted below is one an outbound machine can actually hold.
m.conn.set_remote_epoch(Some([9u8; 8]));
m.session_established_at_ms = 0;
let actions = m.step(
@@ -2641,7 +2639,7 @@ mod tests {
assert_eq!(
actions,
vec![
PeerAction::SwapSendState { epoch: [9u8; 8] },
PeerAction::SwapSendState,
PeerAction::RegisterDecryptSession {
index: SessionIndex::new(0x2222)
},
@@ -2673,6 +2671,9 @@ mod tests {
vec![PeerAction::CompleteDrain { peer: addr }]
);
assert_eq!(m.state(), PeerState::Active { addr });
// The cutover carries no epoch of its own; `conn` is the sole carrier
// and the cutover must leave it exactly as the handshake wrote it.
assert_eq!(m.conn_remote_epoch(), Some([9u8; 8]));
}
// ---- Test 2: responder cutover (data-plane owned) ---------------------
@@ -2701,7 +2702,7 @@ mod tests {
assert!(
!actions
.iter()
.any(|a| matches!(a, PeerAction::SwapSendState { .. }))
.any(|a| matches!(a, PeerAction::SwapSendState))
);
assert_eq!(m.state(), PeerState::Active { addr });
}
@@ -3672,7 +3673,8 @@ mod tests {
};
m.rekey_our_index = Some(SessionIndex::new(0x2222));
m.conn.set_our_index(SessionIndex::new(0x1111));
m.remote_epoch = Some([9u8; 8]);
// Seeded through the setter both handshake legs use; see Test 1.
m.conn.set_remote_epoch(Some([9u8; 8]));
// Consume the shell-decided Cutover.
let cut = m.step(
@@ -3685,7 +3687,7 @@ mod tests {
assert_eq!(
cut,
vec![
PeerAction::SwapSendState { epoch: [9u8; 8] },
PeerAction::SwapSendState,
PeerAction::RegisterDecryptSession {
index: SessionIndex::new(0x2222)
},
@@ -4045,7 +4047,6 @@ mod tests {
};
m.rekey_our_index = Some(SessionIndex::new(0x2222));
m.conn.set_our_index(SessionIndex::new(0x1111));
m.remote_epoch = Some([9u8; 8]);
m.session_established_at_ms = 0;
let actions = m.step(
+19 -7
View File
@@ -408,12 +408,23 @@ impl Probe {
let gave_up = self.lookup_was_pending && !obs.lookup_pending;
if gave_up || obs.now_ms.saturating_sub(self.stage_started_ms) >= self.budgets.discovery_ms
{
let reason = if self.lookup_outcome == Some(LookupOutcomeKind::Deduplicated) {
FailKind::AlreadyPending
} else {
FailKind::NoResponse
};
self.fail_stage(reason, obs);
self.fail_stage(self.discovery_fail_kind(), obs);
}
}
/// Which "the lookup did not answer" finding this probe has earned.
///
/// The gate proceeds only when some peer's filter claimed the target, so
/// a `Sent` lookup that goes unanswered is a claim nobody could confirm.
/// Reporting that as a bare `NoResponse` reads as a network fault and
/// hides the one fact the resolver does know: a filter put the key on the
/// mesh and the mesh disagreed. A joined lookup is its own finding again,
/// because this probe never chose to issue it.
fn discovery_fail_kind(&self) -> FailKind {
match self.lookup_outcome {
Some(LookupOutcomeKind::Deduplicated) => FailKind::AlreadyPending,
Some(LookupOutcomeKind::Sent) => FailKind::BloomUnconfirmed,
_ => FailKind::NoResponse,
}
}
@@ -644,7 +655,8 @@ impl Probe {
/// the reason its own budget would have given.
fn expire_running_stage(&mut self) {
let kind = match self.stage {
Stage::Bloom | Stage::Discovery => FailKind::NoResponse,
Stage::Bloom => FailKind::NoResponse,
Stage::Discovery => self.discovery_fail_kind(),
Stage::Path => FailKind::NoNextHop,
Stage::Session => {
if self.owns_session {
+7
View File
@@ -64,6 +64,12 @@ pub(crate) enum FailKind {
NoTreePeers,
/// discovery: the attempt ladder was exhausted, or the budget expired.
NoResponse,
/// discovery: the gate proceeded on a peer filter's claim and the lookup
/// was never answered, so the claim was never confirmed. Distinct from
/// `NoResponse` because the reason a request went out at all is part of
/// the finding: a filter cannot miss a key that is present, so on an
/// absent key this is that filter false-positiving.
BloomUnconfirmed,
/// path: the two coordinates have different spanning-tree roots.
DisjointTrees,
/// path: no send-ready peer is strictly closer to the target.
@@ -96,6 +102,7 @@ impl FailKind {
FailKind::AlreadyPending => "already_pending",
FailKind::NoTreePeers => "no_tree_peers",
FailKind::NoResponse => "no_response",
FailKind::BloomUnconfirmed => "bloom_unconfirmed",
FailKind::DisjointTrees => "disjoint_trees",
FailKind::NoNextHop => "no_next_hop",
FailKind::Preexisting => "preexisting",
+109 -4
View File
@@ -4,8 +4,8 @@
use crate::NodeAddr;
use crate::proto::probe::core::{Budgets, Observation, Probe, ProbeAction};
use crate::proto::probe::state::{
FailKind, LeftIntact, LookupOutcomeKind, NextHopFacts, PathFacts, Preflight, ResolveSource,
RttCounters, StageVerdict,
FailKind, LeftIntact, LookupOutcomeKind, NextHopFacts, Overall, PathFacts, Preflight,
ResolveSource, RttCounters, StageVerdict,
};
use crate::proto::routing::RouteClass;
@@ -166,7 +166,7 @@ fn discovery_times_out_at_its_own_budget_measured_from_the_request() {
assert_eq!(probe.step(&o), vec![ProbeAction::Finish]);
let snap = probe.snapshot();
assert_eq!(snap.discovery.verdict, StageVerdict::Failed);
assert_eq!(snap.discovery.reason, Some(FailKind::NoResponse));
assert_eq!(snap.discovery.reason, Some(FailKind::BloomUnconfirmed));
assert_eq!(
snap.bloom.verdict,
StageVerdict::Ok,
@@ -198,7 +198,7 @@ fn discovery_fails_early_when_the_pending_entry_clears() {
assert_eq!(probe.step(&o), vec![ProbeAction::Finish]);
assert_eq!(
probe.snapshot().discovery.reason,
Some(FailKind::NoResponse)
Some(FailKind::BloomUnconfirmed)
);
assert!(T0 + 2_000 < T0 + b.discovery_ms);
}
@@ -250,6 +250,111 @@ fn each_gate_decision_lands_on_the_stage_that_owns_it() {
}
}
/// Drive a probe for a key that is **not** on the mesh under a given gate
/// decision, and return the finished snapshot.
///
/// `BloomMiss` is the run where no peer's filter claimed the key. `Sent` is
/// the run where at least one did — which, for a genuinely absent key, can
/// only be a false positive, since a bloom filter has no false negatives.
/// The mesh then answers nothing and the ladder runs out.
fn absent_key_probe(gate: LookupOutcomeKind) -> crate::proto::probe::ProbeSnapshot {
let mut pre = preflight();
pre.coords_cached = false;
let mut probe = Probe::new(T0, budgets(), pre);
let claimed = gate == LookupOutcomeKind::Sent;
for tick in 0..40u64 {
let mut o = obs(T0 + tick * 1_000);
o.coords_cached = false;
if tick > 0 {
o.lookup_outcome = Some(gate);
}
// The claimed run holds a pending entry while the ladder retries, then
// clears it with no coordinates: nobody answered for the address.
o.lookup_pending = claimed && tick < 16;
if probe.step(&o).contains(&ProbeAction::Finish) {
break;
}
}
probe.snapshot()
}
#[test]
fn absent_key_names_the_filter_claim_instead_of_collapsing_into_no_response() {
let clean = absent_key_probe(LookupOutcomeKind::BloomMiss);
let claimed = absent_key_probe(LookupOutcomeKind::Sent);
assert_eq!(
clean.bloom.verdict,
StageVerdict::Failed,
"no filter claimed it, so the bloom stage is where it ended"
);
assert_eq!(
claimed.bloom.verdict,
StageVerdict::Ok,
"a filter claimed it, so a request did go out"
);
assert_eq!(clean.bloom.reason, Some(FailKind::BloomMiss));
assert_eq!(
claimed.discovery.reason,
Some(FailKind::BloomUnconfirmed),
"a lookup issued on a filter claim and left unanswered is its own finding"
);
// The defect this guards: the claimed run must not land on any reason a
// run that never issued a request can also produce. `no_response` was
// exactly such a reason — the bloom stage reaches it too — so reporting
// it here told the operator nothing about which case they were in.
for shared in [
FailKind::BloomMiss,
FailKind::NoResponse,
FailKind::BackoffSuppressed,
FailKind::NoTreePeers,
] {
assert_ne!(
claimed.discovery.reason,
Some(shared),
"{} cannot distinguish a claimed lookup from one never issued",
shared.name()
);
}
// The verdict is the answer and the answer was already right: absent
// either way. Only the reason changes.
assert_eq!(clean.overall, Overall::Failed);
assert_eq!(
clean.overall, claimed.overall,
"the key is absent on both paths; the verdict must not move"
);
}
#[test]
fn cancelling_mid_discovery_still_names_the_filter_claim() {
// `expire_running_stage` writes the reason for a stage that never got to
// settle. It must give the same finding the budget path would.
let mut pre = preflight();
pre.coords_cached = false;
let mut probe = Probe::new(T0, budgets(), pre);
let mut o = obs(T0);
o.coords_cached = false;
probe.step(&o);
let mut o = obs(T0 + 1_000);
o.coords_cached = false;
o.lookup_pending = true;
o.lookup_outcome = Some(LookupOutcomeKind::Sent);
probe.step(&o);
assert_eq!(probe.snapshot().discovery.verdict, StageVerdict::Running);
probe.cancel(T0 + 2_000);
assert_eq!(
probe.snapshot().discovery.reason,
Some(FailKind::BloomUnconfirmed)
);
}
// ---- ownership ----------------------------------------------------------
#[test]
+10
View File
@@ -642,6 +642,15 @@ fn parse_target_addr(addr: &TransportAddr) -> Result<SocksTarget, TransportError
/// counters, so its teardown hook is a no-op. Emits the terminal
/// "receive loop stopped" debug (without a `direction` field) that the
/// shared loop deliberately leaves to each transport.
///
/// The first-frame deadline is `None` on every nym connection. Nym is
/// outbound-only (`accept_connections()` is `false` and no listener is ever
/// bound), so no nym connection is admitted before a byte is read and none
/// occupies a capped slot: the pool carries `()` metadata and the teardown
/// hook decrements nothing. There is no resource for a silent remote to
/// exhaust, and the mixnet's Sphinx routing makes a first frame legitimately
/// slow, so a TCP-scale deadline here would drop good connections to defend
/// a cap that does not exist.
async fn nym_receive_loop(
reader: tokio::net::tcp::OwnedReadHalf,
transport_id: TransportId,
@@ -660,6 +669,7 @@ async fn nym_receive_loop(
mtu,
stats,
"Nym",
None,
|_stats, _meta| {},
)
.await;
+36 -1
View File
@@ -8,6 +8,7 @@
use std::collections::HashMap;
use std::sync::Arc;
use std::time::Duration;
use futures::FutureExt;
use tokio::net::TcpStream;
@@ -141,6 +142,13 @@ pub(crate) trait ProxiedStats: Send + Sync + 'static {
/// The terminal "receive loop stopped" log is **not** emitted here — it is
/// hoisted into each per-transport wrapper (tor carries a `direction` field
/// nym lacks), so this loop is silent on exit.
///
/// `first_frame_timeout` bounds the wait for the *first* complete frame only.
/// It is `Some` for a connection that takes a capped inbound slot from accept
/// — today only tor's onion listener — and `None` everywhere else, which
/// covers every outbound connection and the whole of the nym transport (nym
/// is outbound-only and keeps no counted slots). A deadline expiry is not a
/// receive error and is deliberately not recorded as one.
#[allow(clippy::too_many_arguments)]
pub(crate) async fn proxied_receive_loop<S: ProxiedStats, M>(
mut reader: OwnedReadHalf,
@@ -151,6 +159,7 @@ pub(crate) async fn proxied_receive_loop<S: ProxiedStats, M>(
mtu: u16,
stats: Arc<S>,
label: &'static str,
first_frame_timeout: Option<Duration>,
on_remove: impl Fn(&S, &M),
) {
debug!(
@@ -160,8 +169,34 @@ pub(crate) async fn proxied_receive_loop<S: ProxiedStats, M>(
label
);
let mut first = true;
loop {
match read_fmp_packet(&mut reader, mtu).await {
let read = match first_frame_timeout {
// Bound the first read only. A silent remote otherwise holds its
// inbound slot for as long as it keeps the socket open.
Some(d) if first => {
match tokio::time::timeout(d, read_fmp_packet(&mut reader, mtu)).await {
Ok(result) => result,
Err(_) => {
// Not a recv error: `record_recv_error` means framing
// or I/O failure, and folding deadline expiries into
// it corrupts that counter.
debug!(
transport_id = %transport_id,
remote_addr = %remote_addr,
timeout_secs = d.as_secs_f64(),
"No complete frame within the first-frame deadline, dropping inbound {} connection",
label
);
break;
}
}
}
_ => read_fmp_packet(&mut reader, mtu).await,
};
first = false;
match read {
Ok(data) => {
stats.record_recv(data.len());
+158
View File
@@ -33,6 +33,7 @@ use crate::transport::socks5::{
ConnectingEntry, ConnectingPool, DialError, ProxiedConnection, ProxiedPool, Socks5Auth,
Socks5Dialer, SocksTarget, poll_connecting, proxied_receive_loop,
};
use crate::transport::tcp::INBOUND_FIRST_FRAME_TIMEOUT;
use control::{ControlAuth, TorControlClient, TorMonitoringInfo};
use stats::TorStats;
@@ -367,6 +368,7 @@ impl TorTransport {
pool,
mtu,
max_inbound,
INBOUND_FIRST_FRAME_TIMEOUT,
stats,
)
.await;
@@ -769,6 +771,7 @@ impl TorTransport {
mtu,
recv_stats,
Direction::Outbound,
None,
)
.await;
});
@@ -934,6 +937,7 @@ impl TorTransport {
mtu,
recv_stats,
Direction::Outbound,
None,
)
.await;
});
@@ -1046,6 +1050,10 @@ impl Transport for TorTransport {
/// actually removing it (so a concurrent close/stop never drives the counter
/// below zero). `direction` is retained for the terminal "receive loop
/// stopped" debug field the shared loop deliberately leaves to each transport.
///
/// `first_frame_timeout` is `Some` for an inbound connection, which holds a
/// capped pool slot from the moment it is accepted, and `None` for an
/// outbound one, which holds no such slot.
#[allow(clippy::too_many_arguments)]
async fn tor_receive_loop(
reader: tokio::net::tcp::OwnedReadHalf,
@@ -1056,6 +1064,7 @@ async fn tor_receive_loop(
mtu: u16,
stats: Arc<TorStats>,
direction: Direction,
first_frame_timeout: Option<Duration>,
) {
proxied_receive_loop(
reader,
@@ -1066,6 +1075,7 @@ async fn tor_receive_loop(
mtu,
stats,
"Tor",
first_frame_timeout,
|stats, meta| match meta {
Direction::Inbound => stats.record_pool_inbound_removed(),
Direction::Outbound => stats.record_pool_outbound_removed(),
@@ -1091,6 +1101,13 @@ async fn tor_receive_loop(
/// connections to a local TCP listener; we accept them, configure
/// socket options, split the stream, and spawn a per-connection
/// receive task.
///
/// `first_frame_timeout` is the deadline from accept to the first complete
/// inbound frame, handed to each spawned receive loop. An accepted socket
/// takes an inbound slot against `max_inbound` before any byte is read, so
/// without it a remote that connects and stays silent holds that slot for as
/// long as it keeps the socket open.
#[allow(clippy::too_many_arguments)]
async fn tor_accept_loop(
listener: TcpListener,
transport_id: TransportId,
@@ -1098,6 +1115,7 @@ async fn tor_accept_loop(
pool: ProxiedPool<Direction>,
mtu: u16,
max_inbound: usize,
first_frame_timeout: Duration,
stats: Arc<TorStats>,
) {
debug!(
@@ -1185,6 +1203,7 @@ async fn tor_accept_loop(
mtu,
recv_stats,
Direction::Inbound,
Some(first_frame_timeout),
)
.await;
});
@@ -1893,4 +1912,143 @@ mod tests {
let err = format!("{}", result.unwrap_err());
assert!(err.contains("directory"));
}
// ========================================================================
// Inbound first-frame deadline (onion listener)
// ========================================================================
/// Poll `f` every 10ms until it holds or `limit` elapses.
async fn wait_until<F: FnMut() -> bool>(mut f: F, limit: Duration) -> bool {
let deadline = Instant::now() + limit;
loop {
if f() {
return true;
}
if Instant::now() >= deadline {
return false;
}
tokio::time::sleep(Duration::from_millis(10)).await;
}
}
/// Drives `tor_accept_loop` directly: the only production path to it is
/// `start_directory_mode`, which needs a Tor-managed hostname file and a
/// running daemon, so it is not reachable from a unit test.
fn spawn_onion_accept_loop(
listener: TcpListener,
packet_tx: PacketTx,
first_frame_timeout: Duration,
) -> (ProxiedPool<Direction>, Arc<TorStats>, JoinHandle<()>) {
let pool: ProxiedPool<Direction> = Arc::new(Mutex::new(HashMap::new()));
let stats = Arc::new(TorStats::new());
let handle = tokio::spawn(tor_accept_loop(
listener,
TransportId::new(1),
packet_tx,
pool.clone(),
1400,
64,
first_frame_timeout,
stats.clone(),
));
(pool, stats, handle)
}
/// Mirror of the TCP case: a silent onion-side socket must lose its
/// inbound slot at the deadline. Break-check: with the timeout wrapper
/// removed from the shared loop the count stays at 1 and the second
/// assertion fails.
#[tokio::test]
async fn idle_inbound_onion_socket_releases_its_slot() {
let (tx, _rx) = packet_channel(32);
let listener = TcpListener::bind("127.0.0.1:0").await.unwrap();
let listen = listener.local_addr().unwrap();
let (pool, stats, accept) =
spawn_onion_accept_loop(listener, tx, Duration::from_millis(200));
// Held open for the whole test: any release is the deadline's doing.
let squatter = TcpStream::connect(listen).await.unwrap();
assert!(
wait_until(|| stats.pool_inbound_count() == 1, Duration::from_secs(2)).await,
"an accepted onion socket should take an inbound slot"
);
assert!(
wait_until(|| stats.pool_inbound_count() == 0, Duration::from_secs(2)).await,
"a silent onion socket should lose its slot at the first-frame deadline"
);
assert!(pool.lock().await.is_empty());
drop(squatter);
accept.abort();
}
/// The deadline covers a *complete* first frame, not merely the first
/// byte: a remote that dribbles a prefix inside the deadline and the
/// remainder after it must still lose its slot, and the late frame must
/// not be delivered.
#[tokio::test]
async fn byte_dripped_first_onion_frame_past_deadline_is_dropped() {
let (tx, mut rx) = packet_channel(32);
let listener = TcpListener::bind("127.0.0.1:0").await.unwrap();
let listen = listener.local_addr().unwrap();
let (_pool, stats, accept) =
spawn_onion_accept_loop(listener, tx, Duration::from_millis(300));
let frame = build_msg1_frame();
let mut peer = TcpStream::connect(listen).await.unwrap();
// Prefix inside the deadline, remainder well past it.
peer.write_all(&frame[..4]).await.unwrap();
tokio::time::sleep(Duration::from_millis(600)).await;
let _ = peer.write_all(&frame[4..]).await;
assert!(
tokio::time::timeout(Duration::from_millis(500), rx.recv())
.await
.is_err(),
"a first onion frame completing after the deadline must not be delivered"
);
assert!(
wait_until(|| stats.pool_inbound_count() == 0, Duration::from_secs(2)).await,
"the dripped onion connection should have released its slot"
);
drop(peer);
accept.abort();
}
/// The healthy path, and a regression guard as for TCP: the deadline is
/// scoped to the first iteration, so an established onion connection that
/// then goes quiet keeps its slot. It exists so a future general idle
/// deadline cannot start reaping quiet onion links without a test going
/// red.
#[tokio::test]
async fn established_onion_connection_survives_long_idle() {
let (tx, mut rx) = packet_channel(32);
let listener = TcpListener::bind("127.0.0.1:0").await.unwrap();
let listen = listener.local_addr().unwrap();
let (pool, stats, accept) =
spawn_onion_accept_loop(listener, tx, Duration::from_millis(200));
let mut peer = TcpStream::connect(listen).await.unwrap();
peer.write_all(&build_msg1_frame()).await.unwrap();
let packet = tokio::time::timeout(Duration::from_secs(2), rx.recv())
.await
.expect("timeout")
.expect("packet channel closed");
assert_eq!(packet.data, build_msg1_frame());
// Four deadlines' worth of silence after the first frame.
tokio::time::sleep(Duration::from_millis(800)).await;
assert_eq!(
stats.pool_inbound_count(),
1,
"an established onion connection must not be dropped by the first-frame deadline"
);
assert!(!pool.lock().await.is_empty());
drop(peer);
accept.abort();
}
}
+2 -2
View File
@@ -45,8 +45,8 @@ and a local STUN responder.
Automated network testing with configurable node counts, topology
algorithms (random geometric, Erdos-Renyi, chain, explicit), and fault
injection (netem mutation, link flaps, traffic generation, node
churn). 20 scenarios covering general stress testing, cost-based parent
selection, mixed link technologies (fiber/Bluetooth/WiFi),
churn). 10 scenarios covering general stress and node churn, discovery
over sparse topologies, spanning-tree and bloom-propagation regression,
transport-specific validation (UDP, TCP, Ethernet), and ECN/congestion
testing. Scenarios are
defined in YAML and executed via a Python harness that manages the full
+85 -36
View File
@@ -3,10 +3,11 @@
Automated network testing for FIPS. Generates random or explicit
topologies, spins up Docker containers, and applies configurable
stressors (network impairment, link flaps, traffic generation, node
churn) over a timed simulation run. Scenarios cover general stress
testing, cost-based parent selection, mixed link technologies
(fiber/Bluetooth/WiFi), and transport-specific validation (UDP, TCP,
Ethernet). Logs are collected and analyzed automatically.
churn) over a timed simulation run. Scenarios cover general stress and
node churn, discovery over sparse topologies, spanning-tree and
bloom-propagation regression, transport-specific validation (UDP, TCP,
Ethernet), and ECN/congestion testing. Logs are collected and analyzed
automatically.
## Prerequisites
@@ -23,24 +24,56 @@ Ethernet). Logs are collected and analyzed automatically.
## Available Scenarios
### General stress tests
### General stress and churn
Random topologies with increasing stressor intensity.
Random topologies with increasing stressor intensity. All three enable
netem mutation, link flaps, iperf traffic, node churn and bandwidth
tiers, and differ in transport mix, density, and whether the peer set
itself churns. Each takes a `--nodes N` override, so the node counts
below are defaults rather than fixed sizes.
| Scenario | Nodes | Topology | Duration | Netem | Link Flaps | Traffic | Node Churn | Bandwidth |
| -------- | ----- | ---------------- | -------- | ----- | ---------- | ------- | ---------- | --------- |
| chaos-10 | 10 | random_geometric | 120s | yes | yes | yes | -- | -- |
| churn-10 | 10 | random_geometric | 600s | yes | yes | yes | yes | -- |
| churn-20 | 20 | erdos_renyi | 600s | yes | yes | yes | yes | yes |
| Scenario | Nodes | Topology | Duration | Peer churn |
| ---------------- | ----- | ---------------- | -------- | ---------- |
| churn-mixed | 20 | erdos_renyi | 600s | -- |
| maelstrom | 20 | erdos_renyi | 600s | yes |
| maelstrom-sparse | 50 | random_geometric | 600s | yes |
- **chaos-10**: Network degradation (5-50ms delay, 0-2% loss), link flaps (max 2
down, 10-30s), and iperf traffic (max 3 concurrent). Netem mutates 30% of
links every 15-30s between normal and degraded policies.
- **churn-10**: Extended run with node churn (1 node down at a time, 30-90s).
Tests tree re-convergence after node departure/rejoin.
- **churn-20**: Aggressive scale test. Erdos-Renyi topology, up to 5 nodes down
simultaneously, bandwidth tiers (1/10/100/1000 Mbps), `protect_connectivity`
disabled (partitions allowed).
- **churn-mixed**: Mixed transports on one mesh (60% UDP, 20% Ethernet, 20%
TCP). Netem mutates 30% of links every 20-45s between normal and degraded
policies; link flaps (max 3 down, 10-30s, connectivity protected); node
churn (max 5 down, 30-90s, partitions allowed); bandwidth tiers
(1/10/100/1000 Mbps). Carries baseline assertions, so a run in which the
mesh never formed cannot report success. Local CI runs it as
`churn-mixed --nodes 10 --duration 120`, which is the invocation its
thresholds are calibrated for.
- **maelstrom**: The same stressors plus peer-level topology mutation
(connect/disconnect every 8-12s) and ephemeral identities on half the
nodes, with `coord_ttl_secs: 10` so coordinate cache entries expire during
the run. Tests re-convergence when the peer set and the identities behind
it both move.
- **maelstrom-sparse**: 50-node sparse random geometric graph (radius 0.20,
roughly 3-4 peers per node), which forces multi-hop routing and heavy
discovery use. The short coordinate TTL expires transit-warmed entries, so
nodes must rediscover rather than coast on the cache.
### Spanning-tree and bloom propagation
Explicit topology with an induced parent flap. **No runner invokes this
scenario.** It was retired from both the local and the cloud runner and is
hand-run only; the files remain in the tree and the retirement is recorded as a
coverage gap rather than as a migration to other tests.
| Scenario | Nodes | Topology | Duration | What it tests |
| ----------- | ----- | -------- | -------- | ------------------------------- |
| bloom-storm | 6 | explicit | 180s | Bloom rate under sustained flap |
- **bloom-storm**: Six-node depth-4 mesh. The two candidate uplinks at depth 2
swap netem delay (5ms against 100ms) every 4s with parent-flap dampening
disabled, so the node switches parents each round. Asserts a ceiling on the
`stats.bloom.sent` delta per node over the trailing 30s, and a floor of 10
parent switches so a harness that never produced a real switch cannot pass
trivially. `scenarios/bloom-storm.README.md` carries the bug-class
description and the threshold derivation.
### Cost-based parent selection — retired, now sans-IO unit tests
@@ -65,7 +98,7 @@ Explicit topologies exercising non-UDP transports.
| Scenario | Nodes | Transport | Shape | Duration | Netem | Link Flaps | What it tests |
| ------------- | ----- | -------------- | ----- | -------- | ----- | ---------- | ------------------------------------------ |
| ethernet-only | 4 | Ethernet | Ring | 90s | yes | -- | AF_PACKET transport with beacon discovery |
| ethernet-only | 4 | Ethernet | Ring | 30s | yes | -- | AF_PACKET transport with beacon discovery |
| ethernet-mesh | 6 | UDP + Ethernet | Mesh | 120s | yes | yes | Mixed UDP/Ethernet, netem mutation + flaps |
| tcp-mesh | 6 | UDP + TCP | Mesh | 120s | yes | yes | Mixed UDP/TCP, netem mutation + flaps |
@@ -137,27 +170,32 @@ scenario runs.
| `--duration secs` | Override the scenario's duration |
| `--list` | List available scenarios |
The scenario argument accepts either a name (`churn-10`) or a file
path (`scenarios/churn-10.yaml`).
The scenario argument accepts either a name (`churn-mixed`) or a file
path (`scenarios/churn-mixed.yaml`). `--list` prints the names that
resolve.
## Scenario YAML Format
Annotated example based on `churn-10.yaml`:
Annotated example based on `churn-mixed.yaml`:
```yaml
scenario:
name: "churn-10"
name: "churn-mixed"
seed: 42 # deterministic RNG seed
duration_secs: 600 # total simulation time
topology:
num_nodes: 10
algorithm: random_geometric # or erdos_renyi, chain
num_nodes: 20
algorithm: erdos_renyi # or random_geometric, chain, explicit
params:
radius: 0.5 # algorithm-specific parameter
p: 0.3 # algorithm-specific parameter
ensure_connected: true # retry until graph is connected
subnet: "172.20.0.0/24"
subnet: "172.20.0.0/16"
ip_start: 10 # first node gets .10
transport_mix: # fraction of edges per transport
udp: 0.6
ethernet: 0.2
tcp: 0.2
netem:
enabled: true
@@ -180,33 +218,44 @@ netem:
link_flaps:
enabled: true
interval_secs: { min: 30, max: 60 }
max_down_links: 2
max_down_links: 3
down_duration_secs: { min: 10, max: 30 }
protect_connectivity: true # never partition the graph
traffic:
enabled: true
max_concurrent: 3
interval_secs: { min: 10, max: 30 }
duration_secs: { min: 5, max: 15 }
max_concurrent: 10
interval_secs: { min: 0, max: 30 }
duration_secs: { min: 5, max: 90 }
parallel_streams: 4
node_churn:
enabled: true
interval_secs: { min: 60, max: 180 }
max_down_nodes: 1
interval_secs: { min: 60, max: 90 }
max_down_nodes: 5
down_duration_secs: { min: 30, max: 90 }
protect_connectivity: true # never kill the last path
protect_connectivity: false # partitions allowed
bandwidth:
enabled: false # per-link HTB rate limiting
enabled: true # per-link HTB rate limiting
tiers_mbps: [1, 10, 100, 1000] # each link randomly assigned a tier
assertions: # evaluated after the run
baseline:
min_nodes_reporting: 10
max_roots: 6
min_nodes_parented: 4
min_sessions: 10
logging:
rust_log: "debug"
output_dir: "./sim-results"
```
The assertion thresholds in the shipped file are calibrated against
recorded runs at the invocation CI uses, and the file's own comments say
what they were derived from. Read those before retuning them.
## Topology Algorithms
| Algorithm | Parameters | Description |
+51
View File
@@ -133,3 +133,54 @@ def is_container_running(container: str) -> bool:
return result.returncode == 0 and result.stdout.strip() == "true"
except subprocess.TimeoutExpired:
return False
def existing_containers(names: list[str], timeout: int = 30) -> list[str] | None:
"""Return which of `names` docker still knows about, running or not.
Deliberately scoped to the names the caller passes in. Enumerating by
compose project label would also sweep up whatever a concurrent run or
a different project owns, and under the parallel CI matrix that turns a
leak check into a flake generator.
Returns `None` when docker could not be asked -- a query that failed has
observed nothing, and reporting "no survivors" for it would recreate the
silence this check exists to break.
"""
try:
result = subprocess.run(
["docker", "ps", "-a", "--format", "{{.Names}}"],
capture_output=True,
text=True,
timeout=timeout,
)
except subprocess.TimeoutExpired:
log.error("docker ps timed out after %ds", timeout)
return None
if result.returncode != 0:
log.error("docker ps exited %d\nstderr: %s", result.returncode, _tail(result.stderr))
return None
present = set(result.stdout.split())
return [name for name in names if name in present]
def force_remove(names: list[str], timeout: int = 60) -> None:
"""Remove the named containers outright, whatever state they are in.
Best effort and never raises: this runs on a teardown path that has
already reported a leak, and the report is the part that must survive.
"""
if not names:
return
try:
result = subprocess.run(
["docker", "rm", "-f"] + names,
capture_output=True,
text=True,
timeout=timeout,
)
except subprocess.TimeoutExpired:
log.error("docker rm -f timed out after %ds", timeout)
return
if result.returncode != 0:
log.error("docker rm -f exited %d\nstderr: %s", result.returncode, _tail(result.stderr))
+19 -2
View File
@@ -104,17 +104,34 @@ class NodeManager:
self._start_node(nid)
def _stop_node(self, node_id: str, duration: float):
"""Stop a container."""
"""Stop a container, and record it as down only if it stopped.
A `docker stop` that exits non-zero -- the container already gone,
the daemon refusing -- once left the node marked down anyway. Every
figure the scenario reasons with is derived from that mark:
`down_count` and so the `max_down_nodes` cap, `_would_disconnect`
and so the `protect_connectivity` guard, and the shared
`down_nodes` set that netem, links and traffic all skip. Returning
before the mutation keeps the model equal to the mesh and lets the
next churn tick retry the node.
"""
container = self.topology.container_name(node_id)
docker_exec_quiet(container, "kill 1", timeout=5) # SIGTERM to PID 1
# Use docker stop with a short grace period
import subprocess
subprocess.run(
result = subprocess.run(
["docker", "stop", "-t", "2", container],
capture_output=True,
text=True,
timeout=15,
)
if result.returncode != 0:
log.warning(
"Failed to stop %s: %s; leaving %s marked up",
container, result.stderr.strip(), node_id,
)
return
now = time.time()
state = self.node_states[node_id]
+82 -9
View File
@@ -25,7 +25,7 @@ from .assertions import (
from .compose import generate_compose
from .config_gen import write_configs
from .control import snapshot_all_congestion, snapshot_all_mmp, snapshot_all_trees
from .docker_exec import docker_compose
from .docker_exec import docker_compose, existing_containers, force_remove
from .link_swap import LinkSwapManager
from .links import LinkManager
from .logs import AnalysisResult, analyze_logs, collect_logs, write_sim_metadata
@@ -569,6 +569,85 @@ class SimRunner:
except Exception:
log.exception("Could not write status file")
def _own_containers(self) -> list[str]:
"""Name every container this run asked compose to create.
Empty before the topology exists, which is the only window in which
a teardown can run with nothing of its own on the host.
"""
if not self.topology:
return []
return [self.topology.container_name(nid) for nid in self.topology.nodes]
def _stop_mesh(self) -> None:
"""Take the containers down, then check that they actually went.
`docker compose down` exits 0 while leaving containers behind when it
races an `up -d` that failed part way through: compose enumerates the
project before the stragglers have registered, finds nothing to
remove, and says so successfully. Nothing downstream looks at what
survived, so the leak is invisible at the one moment it is cheap to
see.
"""
log.info("Stopping containers...")
docker_compose(self.compose_file, ["down"], check=False)
self._check_teardown()
def _check_teardown(self) -> None:
"""Report, then clear, any of this run's containers that outlived `down`.
Scoped to names this run generated. A check that reasoned about the
compose project label, or about container names in general, would
answer for whatever a concurrent scenario happens to own, and the CI
matrix runs plenty of those at once.
"""
wanted = self._own_containers()
if not wanted:
return
survivors = existing_containers(wanted)
if survivors is None:
# Detection failed, which is not the same as detecting nothing.
log.error("Could not ask docker what survived `down`; leak undetectable")
return
if not survivors:
return
log.warning(
"`compose down` exited 0 but left %d of this run's %d containers: %s",
len(survivors), len(wanted), ", ".join(survivors),
)
# Remedy, kept apart from the detection above on purpose: deleting
# everything from here down leaves the report intact.
force_remove(survivors)
leaked = existing_containers(survivors)
if not leaked:
return
# Either docker could not be asked a second time, or the containers
# are still there after a forced removal. Neither is the transient
# race, and neither clears itself, so this is the run's verdict.
detail = "unknown" if leaked is None else ", ".join(leaked)
log.error("Containers survived forced removal, leaked to the host: %s", detail)
self._write_leak(leaked if leaked is not None else wanted)
self.aborted = True
def _write_leak(self, names: list[str]) -> None:
"""Record the leaked names beside the run's other artifacts.
Never raises, for the same reason `_write_status` does not: this runs
on a teardown path that has already gone wrong. Written alongside
status.txt rather than into it: the status is the simulation's own
outcome, and a leak on the way out does not retract it.
"""
try:
os.makedirs(self.output_dir, exist_ok=True)
path = os.path.join(self.output_dir, "leaked-containers.txt")
with open(path, "w") as f:
for name in names:
f.write(name + "\n")
except Exception:
log.exception("Could not write leaked-containers file")
def _release_network(self) -> None:
"""Give this run's claimed /24 back.
@@ -593,8 +672,7 @@ class SimRunner:
if self.compose_file:
# `up -d` can fail part way through, so this run may own
# containers or a network even with no mesh to speak of.
log.info("Stopping containers...")
docker_compose(self.compose_file, ["down"], check=False)
self._stop_mesh()
self._release_network()
return None
@@ -755,12 +833,7 @@ class SimRunner:
# Status first: it is the one artifact that must exist whatever
# else happens, and stopping containers can still time out.
self._write_status(status)
log.info("Stopping containers...")
docker_compose(
self.compose_file,
["down"],
check=False,
)
self._stop_mesh()
self._release_network()
return result
+73 -8
View File
@@ -55,50 +55,115 @@ cleanup_container() {
docker rm -f "$name" >/dev/null 2>&1 || true
}
# Emit a captured output file to stderr, delimited and labelled with the
# command it came from. Callers use this only on failure: a command that
# succeeds leaves no trace, so the suite stays quiet when it is green.
dump_output() {
local label="$1" file="$2"
{
echo " --- $label failed; captured output follows ---"
if [ -s "$file" ]; then
cat "$file"
else
echo " (no output)"
fi
echo " --- end captured output ---"
} >&2
}
# Run a command with both streams captured. Discard the capture on success;
# on failure emit it, so the reason a build or a container start died is not
# thrown away. Stdin is inherited, so a caller may pipe into it.
run_quiet() {
local label="$1"
shift
local out rc=0
out=$(mktemp)
"$@" >"$out" 2>&1 || rc=$?
[ "$rc" -eq 0 ] || dump_output "$label" "$out"
rm -f "$out"
return "$rc"
}
# Build an image from an inline Dockerfile.
build_image() {
local tag="$1"
shift
echo "$@" | docker build -t "$tag" -f - "$REPO_ROOT" >/dev/null 2>&1
echo "$@" | run_quiet "docker build -t $tag" \
docker build -t "$tag" -f - "$REPO_ROOT"
}
# Start a systemd container in the background.
start_systemd_container() {
local name="$1" image="$2"
cleanup_container "$name"
docker run -d --name "$name" \
run_quiet "docker run $name" \
docker run -d --name "$name" \
--label com.corganlabs.fips-ci=1 \
--privileged \
--cgroupns=host \
-v /sys/fs/cgroup:/sys/fs/cgroup:rw \
--tmpfs /run --tmpfs /run/lock \
"$image" >/dev/null 2>&1
"$image"
}
# Same, but with TUN device for the e2e scenario.
start_systemd_container_with_tun() {
local name="$1" image="$2"
cleanup_container "$name"
docker run -d --name "$name" \
run_quiet "docker run $name (with tun)" \
docker run -d --name "$name" \
--label com.corganlabs.fips-ci=1 \
--privileged \
--cgroupns=host \
--device /dev/net/tun \
-v /sys/fs/cgroup:/sys/fs/cgroup:rw \
--tmpfs /run --tmpfs /run/lock \
"$image" >/dev/null 2>&1
"$image"
}
# Report why systemd never reached a running state. Every probe is
# best-effort and captured with its own stderr: the container may have
# exited, or never have been created at all, and the docker error text
# saying so is itself the diagnosis.
dump_systemd_state() {
local name="$1" out
out=$(mktemp)
{
echo "== docker ps -a"
docker ps -a --filter "name=^${name}$" 2>&1
echo "== systemctl is-system-running"
docker exec "$name" systemctl is-system-running 2>&1
echo "== systemctl list-units --failed"
docker exec "$name" systemctl list-units --failed --no-pager 2>&1
echo "== journalctl -b (last 100 lines)"
docker exec "$name" journalctl -b --no-pager -n 100 2>&1
echo "== docker logs (last 100 lines)"
docker logs --tail 100 "$name" 2>&1
} >"$out" 2>&1
dump_output "systemd boot of $name" "$out"
rm -f "$out"
}
# Wait for systemd to reach a bootable state inside the container.
wait_for_systemd() {
local name="$1"
local state
for _i in $(seq 1 "$BOOT_TIMEOUT"); do
if docker exec "$name" systemctl is-system-running --wait 2>/dev/null | grep -qE 'running|degraded'; then
return 0
fi
# Read the state rather than piping it into grep. systemctl exits
# non-zero for "degraded", and pipefail turns that into a failed
# pipeline even when grep matched, so the piped form could never
# accept a degraded boot: in a container systemd-modules-load
# always fails, so every scenario burned the full timeout and
# warned about a container that had in fact booted.
state=$(docker exec "$name" systemctl is-system-running --wait 2>/dev/null)
case "$state" in
*running*|*degraded*) return 0 ;;
esac
sleep 1
done
echo " WARNING: systemd did not reach running state in ${BOOT_TIMEOUT}s (may still work)"
dump_systemd_state "$name"
return 0
}
+73
View File
@@ -111,8 +111,31 @@ ping_backcompat_hold() {
fi
}
# Case 6 trace: converges quickly, so the verdict case that needs a
# converged run does not spend the near-converged hold's twelve seconds.
ping_quick_converge() {
set_pt; local t=$PT
if (( t < 2 )); then
PASSED=18; FAILED=2
else
PASSED=20; FAILED=0
fi
}
HOLD_MSG="holding for full budget"
STUCK_MSG="STUCK"
NOCONV_MSG="tree did not converge"
# Run the gate in THIS shell (not a subshell) so the CONVERGE_* verdict
# globals survive, capturing its output to a file instead. `$(...)` runs the
# gate in a fork, which discards those assignments — that is why cases 1-4
# can only assert on text.
VERDICT_OUT=$(mktemp)
run_gate() {
reset_ping
CONVERGE_OUTCOME=""; CONVERGE_REACHED=-1; CONVERGE_PENDING=-1
wait_until_connected "$@" >"$VERDICT_OUT" 2>&1
}
# --- Case 1: near-converged hold --------------------------------------
echo
@@ -221,6 +244,56 @@ check "case5: floor of 1 polled its full budget" "$c5_polled_ok" "elapsed=${elap
unset -f docker
# --- Case 6: the verdict discriminates non-convergence from connectivity --
#
# This is the break-what-it-guards check for the verdict itself. The recorded
# failure exited 1 while reporting "20 passed, 0 failed": every connectivity
# pair passed and only the tree fell short, and nothing in the summary told
# the two apart. The gate now names its verdict, so drive it into each
# outcome and assert the verdict is the one that outcome deserves.
echo
echo "== Case 6: verdict names which condition failed =="
echo "-- Case 6a: genuinely unconverged tree, hard cap --"
run_gate ping_never_converges 6 3 1 2; rc=$?
cat "$VERDICT_OUT"
c6a_rc_ok=1; [ "$rc" -ne 0 ] && c6a_rc_ok=0
check "case6a: unconverged tree still reds" "$c6a_rc_ok" "rc=$rc"
c6a_out_ok=1; [ "$CONVERGE_OUTCOME" = "timeout" ] && c6a_out_ok=0
check "case6a: verdict is timeout" "$c6a_out_ok" "CONVERGE_OUTCOME=$CONVERGE_OUTCOME"
c6a_cnt_ok=1
[ "$CONVERGE_REACHED" -eq 19 ] && [ "$CONVERGE_PENDING" -eq 1 ] && c6a_cnt_ok=0
check "case6a: verdict carries the shortfall" "$c6a_cnt_ok" \
"reached=$CONVERGE_REACHED pending=$CONVERGE_PENDING"
c6a_msg_ok=1; grep -q "$NOCONV_MSG" "$VERDICT_OUT" && c6a_msg_ok=0
check "case6a: message says the tree did not converge" "$c6a_msg_ok"
echo "-- Case 6b: wedged far from convergence, stall bail --"
run_gate ping_far_stall 30 4 1 2; rc=$?
cat "$VERDICT_OUT"
c6b_rc_ok=1; [ "$rc" -ne 0 ] && c6b_rc_ok=0
check "case6b: wedged tree still reds" "$c6b_rc_ok" "rc=$rc"
c6b_out_ok=1; [ "$CONVERGE_OUTCOME" = "stalled" ] && c6b_out_ok=0
check "case6b: verdict is stalled" "$c6b_out_ok" "CONVERGE_OUTCOME=$CONVERGE_OUTCOME"
c6b_msg_ok=1; grep -q "$NOCONV_MSG" "$VERDICT_OUT" && c6b_msg_ok=0
check "case6b: message says the tree did not converge" "$c6b_msg_ok"
echo "-- Case 6c: converged tree, and the verdict does not cry non-convergence --"
run_gate ping_quick_converge 20 4 1 2; rc=$?
cat "$VERDICT_OUT"
c6c_rc_ok=1; [ "$rc" -eq 0 ] && c6c_rc_ok=0
check "case6c: converged tree still passes" "$c6c_rc_ok" "rc=$rc"
c6c_out_ok=1; [ "$CONVERGE_OUTCOME" = "converged" ] && c6c_out_ok=0
check "case6c: verdict is converged" "$c6c_out_ok" "CONVERGE_OUTCOME=$CONVERGE_OUTCOME"
c6c_cnt_ok=1
[ "$CONVERGE_REACHED" -eq 20 ] && [ "$CONVERGE_PENDING" -eq 0 ] && c6c_cnt_ok=0
check "case6c: verdict carries a clean tree" "$c6c_cnt_ok" \
"reached=$CONVERGE_REACHED pending=$CONVERGE_PENDING"
c6c_quiet_ok=0; grep -q "$NOCONV_MSG" "$VERDICT_OUT" && c6c_quiet_ok=1
check "case6c: no non-convergence message on a clean run" "$c6c_quiet_ok"
rm -f "$VERDICT_OUT"
# --- Summary ----------------------------------------------------------
echo
echo "=============================================="
+32 -2
View File
@@ -9,6 +9,9 @@
# wait_until_connected <ping_fn> <max_secs> <stall_secs> [poll_secs] \
# [near_converged_slack]
#
# wait_until_connected also sets CONVERGE_OUTCOME / CONVERGE_REACHED /
# CONVERGE_PENDING; see the block above it.
#
# There was a wait_for_links() here. It was removed rather than kept for
# symmetry: it had no caller anywhere in the tree on any branch, and its
# reader carried the same failure-to-zero fallback wait_for_peers does. An
@@ -56,6 +59,30 @@ wait_for_peers() {
return 1
}
# Verdict of the most recent wait_until_connected() call, so a caller can
# report WHICH condition failed rather than only that one did:
# CONVERGE_OUTCOME converged | stalled | timeout
# CONVERGE_REACHED reachable pairs at the moment of the verdict
# CONVERGE_PENDING unreachable pairs at that moment
#
# These exist because the gate's own probe is strictly harsher than the
# assertion it guards, so a run can fail the gate at 18/20 and then pass
# the strict all-pairs assertion 20/20. Without them the caller's summary
# line reads "20 passed, 0 failed" on a non-convergence exit, which a
# reader cannot tell from a connectivity failure.
CONVERGE_OUTCOME=""
CONVERGE_REACHED=0
CONVERGE_PENDING=0
# Record the verdict of a wait_until_connected() return.
#
# shellcheck disable=SC2034 # read by sourcing suites, not within this file
_converge_verdict() {
CONVERGE_OUTCOME="$1"
CONVERGE_REACHED="$PASSED"
CONVERGE_PENDING="$FAILED"
}
# Wait until a connectivity check reports every pair reachable, using a
# progress-aware deadline instead of a fixed one.
#
@@ -102,6 +129,7 @@ wait_until_connected() {
while (( SECONDS - start_secs < max_secs )); do
"$ping_fn"
if (( FAILED == 0 )); then
_converge_verdict converged
echo " converge: all $PASSED pair(s) reachable after $((SECONDS - start_secs))s"
return 0
fi
@@ -111,7 +139,8 @@ wait_until_connected() {
echo " converge: $PASSED reachable, $FAILED pending (progressing) after $((SECONDS - start_secs))s"
elif (( SECONDS - last_progress >= stall_secs )); then
if (( FAILED > near_converged_slack )); then
echo " converge: STUCK at $PASSED reachable / $FAILED pending — no progress for ${stall_secs}s (after $((SECONDS - start_secs))s)"
_converge_verdict stalled
echo " converge: STUCK — tree did not converge: $PASSED reachable / $FAILED pending, no progress for ${stall_secs}s (after $((SECONDS - start_secs))s)"
return 1
fi
if (( held_for_budget == 0 )); then
@@ -122,6 +151,7 @@ wait_until_connected() {
sleep "$poll_secs"
done
echo " converge: TIMEOUT at $PASSED reachable / $FAILED pending after ${max_secs}s"
_converge_verdict timeout
echo " converge: TIMEOUT — tree did not converge: $PASSED reachable / $FAILED pending after ${max_secs}s"
return 1
}
+121 -29
View File
@@ -306,41 +306,97 @@ require_docker_daemon() {
fi
}
# The two path assertions below decide whether a converged mesh took the path
# the scenario expects. Every failure branch names the container, what was
# expected and what was read instead, and callers wrap them in the same
# `|| { dump_*_diagnostics; return 1; }` shape the convergence waits use, so the
# container state behind a mismatch is captured with it. Both stay silent on
# success: this suite passes most of the time.
assert_peer_path() {
local container="$1"
local expected_transport="$2"
local expected_prefix="$3"
docker exec "$container" fipsctl show peers \
| python3 -c "
local peers errfile rc=0
# stdout and stderr are kept apart deliberately: any warning fipsctl writes
# to stderr would otherwise be spliced into the JSON and turn a healthy read
# into a parse failure.
errfile="$(mktemp)"
peers="$(docker exec "$container" fipsctl show peers 2>"$errfile")" || rc=$?
if [ "$rc" != 0 ]; then
echo "ASSERT FAIL: peer path $container: expected transport ${expected_transport} to a peer at ${expected_prefix}*, but 'fipsctl show peers' exited ${rc}:" >&2
cat "$errfile" >&2
rm -f "$errfile"
return 1
fi
rm -f "$errfile"
if ! python3 -c "
import json, sys
data = json.load(sys.stdin)
peers = [p for p in data.get('peers', []) if p.get('connectivity') == 'connected']
container, want_transport, want_prefix = sys.argv[1], sys.argv[2], sys.argv[3]
raw = sys.stdin.read()
want = f'transport {want_transport!r} to a peer at {want_prefix}*'
head = f'ASSERT FAIL: peer path {container}: expected {want}, '
try:
data = json.loads(raw)
except ValueError as exc:
raise SystemExit(head + f'but the peer JSON did not parse: {exc}; read: {raw!r}')
reported = data.get('peers', [])
seen = [
{k: p.get(k) for k in
('npub', 'connectivity', 'transport_type', 'transport_addr', 'direction', 'last_seen_ms')}
for p in reported
]
peers = [p for p in reported if p.get('connectivity') == 'connected']
if not peers:
raise SystemExit(1)
raise SystemExit(head + f'observed no connected peer; {len(reported)} peer(s) reported: {seen}')
peer = peers[0]
transport = peer.get('transport_type', '')
addr = peer.get('transport_addr', '')
if transport != sys.argv[1]:
raise SystemExit(f'transport mismatch: expected {sys.argv[1]!r}, got {transport!r}')
if not addr.startswith(sys.argv[2]):
raise SystemExit(f'addr mismatch: expected prefix {sys.argv[2]!r}, got {addr!r}')
" "$expected_transport" "$expected_prefix"
if transport != want_transport:
raise SystemExit(head + f'observed transport {transport!r} at {addr!r}; peers: {seen}')
if not addr.startswith(want_prefix):
raise SystemExit(head + f'observed addr {addr!r} on transport {transport!r}; peers: {seen}')
" "$container" "$expected_transport" "$expected_prefix" <<<"$peers"; then
return 1
fi
}
assert_link_path() {
local container="$1"
local expected_prefix="$2"
docker exec "$container" fipsctl show links \
| python3 -c "
local links errfile rc=0
# See assert_peer_path: stderr is kept out of the JSON on purpose.
errfile="$(mktemp)"
links="$(docker exec "$container" fipsctl show links 2>"$errfile")" || rc=$?
if [ "$rc" != 0 ]; then
echo "ASSERT FAIL: link path $container: expected a link to ${expected_prefix}*, but 'fipsctl show links' exited ${rc}:" >&2
cat "$errfile" >&2
rm -f "$errfile"
return 1
fi
rm -f "$errfile"
if ! python3 -c "
import json, sys
data = json.load(sys.stdin)
container, want_prefix = sys.argv[1], sys.argv[2]
raw = sys.stdin.read()
want = f'a link to {want_prefix}*'
head = f'ASSERT FAIL: link path {container}: expected {want}, '
try:
data = json.loads(raw)
except ValueError as exc:
raise SystemExit(head + f'but the link JSON did not parse: {exc}; read: {raw!r}')
links = data.get('links', [])
seen = [
{k: link.get(k) for k in ('link_id', 'remote_addr', 'direction', 'state')}
for link in links
]
if not links:
raise SystemExit(1)
raise SystemExit(head + 'observed no links at all')
addr = links[0].get('remote_addr', '')
if not addr.startswith(sys.argv[1]):
raise SystemExit(f'link addr mismatch: expected prefix {sys.argv[1]!r}, got {addr!r}')
" "$expected_prefix"
if not addr.startswith(want_prefix):
raise SystemExit(head + f'observed {addr!r}; links: {seen}')
" "$container" "$expected_prefix" <<<"$links"; then
return 1
fi
}
require_bootstrap_activity() {
@@ -373,10 +429,22 @@ run_cone() {
dump_cone_diagnostics
return 1
}
assert_peer_path fips-nat-cone-a${FIPS_CI_NAME_SUFFIX:-} udp ${NAT_WAN}.
assert_peer_path fips-nat-cone-b${FIPS_CI_NAME_SUFFIX:-} udp ${NAT_WAN}.
assert_link_path fips-nat-cone-a${FIPS_CI_NAME_SUFFIX:-} ${NAT_WAN}.
assert_link_path fips-nat-cone-b${FIPS_CI_NAME_SUFFIX:-} ${NAT_WAN}.
assert_peer_path fips-nat-cone-a${FIPS_CI_NAME_SUFFIX:-} udp ${NAT_WAN}. || {
dump_cone_diagnostics
return 1
}
assert_peer_path fips-nat-cone-b${FIPS_CI_NAME_SUFFIX:-} udp ${NAT_WAN}. || {
dump_cone_diagnostics
return 1
}
assert_link_path fips-nat-cone-a${FIPS_CI_NAME_SUFFIX:-} ${NAT_WAN}. || {
dump_cone_diagnostics
return 1
}
assert_link_path fips-nat-cone-b${FIPS_CI_NAME_SUFFIX:-} ${NAT_WAN}. || {
dump_cone_diagnostics
return 1
}
# shellcheck disable=SC1090
source "$CONFIG_DIR/cone/npubs.env"
ping_peer fips-nat-cone-a${FIPS_CI_NAME_SUFFIX:-} "$NPUB_B"
@@ -398,10 +466,22 @@ run_symmetric() {
dump_symmetric_diagnostics
return 1
}
assert_peer_path fips-nat-symmetric-a${FIPS_CI_NAME_SUFFIX:-} tcp ${NAT_WAN}.11:
assert_peer_path fips-nat-symmetric-b${FIPS_CI_NAME_SUFFIX:-} tcp ${NAT_WAN}.10:
assert_link_path fips-nat-symmetric-a${FIPS_CI_NAME_SUFFIX:-} ${NAT_WAN}.11:
assert_link_path fips-nat-symmetric-b${FIPS_CI_NAME_SUFFIX:-} ${NAT_WAN}.10:
assert_peer_path fips-nat-symmetric-a${FIPS_CI_NAME_SUFFIX:-} tcp ${NAT_WAN}.11: || {
dump_symmetric_diagnostics
return 1
}
assert_peer_path fips-nat-symmetric-b${FIPS_CI_NAME_SUFFIX:-} tcp ${NAT_WAN}.10: || {
dump_symmetric_diagnostics
return 1
}
assert_link_path fips-nat-symmetric-a${FIPS_CI_NAME_SUFFIX:-} ${NAT_WAN}.11: || {
dump_symmetric_diagnostics
return 1
}
assert_link_path fips-nat-symmetric-b${FIPS_CI_NAME_SUFFIX:-} ${NAT_WAN}.10: || {
dump_symmetric_diagnostics
return 1
}
require_bootstrap_activity fips-nat-symmetric-a${FIPS_CI_NAME_SUFFIX:-}
require_bootstrap_activity fips-nat-symmetric-b${FIPS_CI_NAME_SUFFIX:-}
# shellcheck disable=SC1090
@@ -424,10 +504,22 @@ run_lan() {
dump_lan_diagnostics
return 1
}
assert_peer_path fips-nat-lan-a${FIPS_CI_NAME_SUFFIX:-} udp ${NAT_LAN}.
assert_peer_path fips-nat-lan-b${FIPS_CI_NAME_SUFFIX:-} udp ${NAT_LAN}.
assert_link_path fips-nat-lan-a${FIPS_CI_NAME_SUFFIX:-} ${NAT_LAN}.
assert_link_path fips-nat-lan-b${FIPS_CI_NAME_SUFFIX:-} ${NAT_LAN}.
assert_peer_path fips-nat-lan-a${FIPS_CI_NAME_SUFFIX:-} udp ${NAT_LAN}. || {
dump_lan_diagnostics
return 1
}
assert_peer_path fips-nat-lan-b${FIPS_CI_NAME_SUFFIX:-} udp ${NAT_LAN}. || {
dump_lan_diagnostics
return 1
}
assert_link_path fips-nat-lan-a${FIPS_CI_NAME_SUFFIX:-} ${NAT_LAN}. || {
dump_lan_diagnostics
return 1
}
assert_link_path fips-nat-lan-b${FIPS_CI_NAME_SUFFIX:-} ${NAT_LAN}. || {
dump_lan_diagnostics
return 1
}
# shellcheck disable=SC1090
source "$CONFIG_DIR/lan/npubs.env"
ping_peer fips-nat-lan-a${FIPS_CI_NAME_SUFFIX:-} "$NPUB_B"
+31 -5
View File
@@ -196,6 +196,11 @@ PASSED=0
FAILED=0
TOTAL_PASSED=0
TOTAL_FAILED=0
# Counted separately from TOTAL_FAILED on purpose: a convergence-gate
# failure and a connectivity failure are different outcomes, and folding
# the first into the second is what made the recorded transcript report
# "20 passed, 0 failed" on an exit-1 run.
TOTAL_UNCONVERGED=0
# Node identities
ENV_FILE="$SCRIPT_DIR/../generated-configs${FIPS_CI_NAME_SUFFIX:-}/npubs.env"
@@ -274,6 +279,24 @@ _baseline_ping() {
ping_all quiet "$CONVERGENCE_PING_TIMEOUT"
}
# Emit the one-line run summary.
#
# When the convergence gate is what failed, the line says so and names the
# shortfall. The gate's probe is strictly harsher than the strict all-pairs
# assertion it guards, so a run can fail the gate at 18/20 relationships and
# still pass the assertion 20/20 — which is exactly the recorded failure this
# discriminator exists for. Without it both outcomes print the same counts.
results_line() {
local line="=== Results: $TOTAL_PASSED passed, $TOTAL_FAILED failed"
if [ "$TOTAL_UNCONVERGED" -ne 0 ]; then
line+=", tree did not converge"
line+=" ($CONVERGE_REACHED/$((CONVERGE_REACHED + CONVERGE_PENDING))"
line+=" relationships, $CONVERGE_OUTCOME)"
fi
echo "$line ==="
return 0
}
phase_result() {
local phase="$1"
TOTAL_PASSED=$((TOTAL_PASSED + PASSED))
@@ -404,16 +427,19 @@ if wait_until_connected _baseline_ping "$BASELINE_CONVERGENCE_TIMEOUT" 20; then
if [ "$FAILED" -ne 0 ]; then
echo ""
dump_peer_connectivity
echo "=== Results: $TOTAL_PASSED passed, $TOTAL_FAILED failed ==="
results_line
exit 1
fi
else
echo " Mesh did not reach a converged tree before timeout"
TOTAL_UNCONVERGED=$((TOTAL_UNCONVERGED + 1))
echo " Mesh did not reach a converged tree before timeout" \
"($CONVERGE_OUTCOME at $CONVERGE_REACHED reachable /" \
"$CONVERGE_PENDING pending)"
ping_all quiet "$CONVERGENCE_PING_TIMEOUT"
phase_result "Pre-rekey baseline (all 20 pairs)"
echo ""
dump_peer_connectivity
echo "=== Results: $TOTAL_PASSED passed, $TOTAL_FAILED failed ==="
results_line
exit 1
fi
echo ""
@@ -570,9 +596,9 @@ phase_result "Log analysis"
echo ""
# ── Summary ────────────────────────────────────────────────────────────
echo "=== Results: $TOTAL_PASSED passed, $TOTAL_FAILED failed ==="
results_line
if [ "$TOTAL_FAILED" -eq 0 ]; then
if [ "$TOTAL_FAILED" -eq 0 ] && [ "$TOTAL_UNCONVERGED" -eq 0 ]; then
exit 0
else
# Dump logs on failure for diagnostics.