mirror of
https://github.com/jmcorgan/fips.git
synced 2026-10-05 19:18:25 +00:00
Merge master into next, forward-porting each side onto the other
Five files conflicted and each needed a different call, because the two lines had rewritten different halves of the same code. The handshake handler takes next's version whole. master's entire change there was three comment blocks and one widened debug_assert, and the assert names HandshakePhase::ReceivedMsg1, a variant next's XX rewrite does not have. Nothing semantic was dropped. The peer reaper takes master's: reap_peers_on_transport is new and its route_link_dead doc now describes both callers, which is true on this line too. The ethernet transport takes master's binder rewrite with next's wire format re-applied on top. The send path, the receive path and the frame tests merged to the 4-byte header on their own, but three sites are new in master's rewrite and had never seen it: the Binding default and both arms of the binder's MTU calculation still subtracted 3. The transports snapshot fixture moved with them, 1499 to 1496, and that single field was the whole diff. Beacons carry no pubkey here, so local_pubkey leaves the transport, its binder context and the node's transport construction with it. The changelog keeps both sides' entries, with master's Added subsection lifted back out of Changed where the merge had left it. Two tests do not come across. a_transient_msg2_failure_keeps_the_link_for_ the_retry and its restart-path sibling assert that the machine rests at ReceivedMsg1. This line's nearest state is SentMsg2, and it means something else: the inbound leg parks there awaiting msg3, where on the other line that phase was the last stop before promotion. Renaming it would produce a test that passes without exercising the deferral. The behaviour they guard did merge and sits in the transient arm of the msg2 send failure; what is missing is coverage shaped for this handshake, which is tracked separately. The two connected-socket tests did come across. Their helper took the responder's session straight after msg2, which is an IK assumption; it now runs msg3 as well. Both pass here and both go red when the clear is removed or made unconditional. The test-harness fixes arrive through master rather than as follow-ups here, so this line never carries the versions that failed: the interface-binding suite's veth naming, and the chaos veth restore, random streams, settle wait, netem restore and shared down-node set.
This commit is contained in:
@@ -284,6 +284,30 @@ jobs:
|
||||
- name: Install system dependencies
|
||||
run: sudo apt-get update && sudo apt-get install -y libdbus-1-dev
|
||||
|
||||
# The same address-less interface the musl leg creates. Pinning the
|
||||
# contract on both libcs is what turns "glibc and musl agree about
|
||||
# `getifaddrs`" from an assumption into a checked fact — and makes this
|
||||
# leg fail first if glibc is the one that changes.
|
||||
- name: Create an address-less interface for the presence probe
|
||||
run: |
|
||||
sudo ip link add fips-probe0 type dummy
|
||||
# `addrgenmode none` before bringing it up: the kernel hands an IPv6
|
||||
# link-local to any interface that comes up, and an interface with a
|
||||
# link-local is not address-less — the fixture would have quietly
|
||||
# tested nothing.
|
||||
sudo ip link set fips-probe0 addrgenmode none
|
||||
sudo ip link set fips-probe0 up
|
||||
ip addr show fips-probe0
|
||||
# Fail rather than test the wrong thing if it acquired one anyway.
|
||||
if ip addr show fips-probe0 | grep -qE "inet6? "; then
|
||||
echo "fips-probe0 has an address; it cannot test the address-less case" >&2
|
||||
exit 1
|
||||
fi
|
||||
echo "FIPS_TEST_ADDRLESS_IFACE=fips-probe0" >> "$GITHUB_ENV"
|
||||
# Declare that this runner has fixtures, so a test that depends on
|
||||
# one fails when the fixture is missing instead of skipping silently.
|
||||
echo "FIPS_TEST_REQUIRE_FIXTURES=1" >> "$GITHUB_ENV"
|
||||
|
||||
- name: Install Rust toolchain
|
||||
uses: actions-rust-lang/setup-rust-toolchain@166cdcfd11aee3cb47222f9ddb555ce30ddb9659 # v1
|
||||
with:
|
||||
@@ -304,9 +328,39 @@ jobs:
|
||||
- name: Install cargo-nextest
|
||||
uses: taiki-e/install-action@nextest
|
||||
|
||||
|
||||
- name: Run unit tests
|
||||
run: cargo nextest run --all --profile ci
|
||||
|
||||
# The bind-success half. Every other unit-test leg runs unprivileged, so
|
||||
# `PacketSocket::open` cannot succeed on any of them and everything past
|
||||
# a successful bind — the post-store shutdown check, the `Present` arm of
|
||||
# the binder loop, `bind_now` itself — runs nowhere in CI.
|
||||
#
|
||||
# Built as the runner user and only *executed* under sudo: `cargo` run as
|
||||
# root would use root's CARGO_HOME and discard the cache this job just
|
||||
# restored.
|
||||
#
|
||||
# `FIPS_TEST_PRIVILEGED` is what makes this leg honest. A test that needs
|
||||
# a raw socket skips quietly without it; with it set, a test that cannot
|
||||
# open one fails and says so, so a runner that stops granting the
|
||||
# capability shows up as a red leg rather than as silence.
|
||||
- name: Run interface-binding tests with privilege
|
||||
run: |
|
||||
cargo test --lib --no-run
|
||||
BIN=$(cargo test --lib --no-run --message-format=json \
|
||||
| jq -r 'select(.reason == "compiler-artifact")
|
||||
| select(.executable != null)
|
||||
| select(.target.kind[0] == "lib")
|
||||
| .executable' \
|
||||
| tail -1)
|
||||
if [ -z "$BIN" ] || [ ! -x "$BIN" ]; then
|
||||
echo "could not locate the lib test binary" >&2
|
||||
exit 1
|
||||
fi
|
||||
echo "running $BIN as root"
|
||||
sudo -E env FIPS_TEST_PRIVILEGED=1 "$BIN" transport::ethernet --test-threads=1
|
||||
|
||||
- name: Publish test report (Checks tab)
|
||||
uses: dorny/test-reporter@4a2e97665d5fa767581ef38eca97b9694bd4eef4 # v2
|
||||
if: always()
|
||||
@@ -363,9 +417,116 @@ jobs:
|
||||
- name: Install cargo-nextest
|
||||
uses: taiki-e/install-action@nextest
|
||||
|
||||
# The Darwin half of the address-less presence contract. The Linux legs
|
||||
# pin that `getifaddrs` reports an interface with no addresses as
|
||||
# present, on both glibc and musl; without this the same claim on the
|
||||
# BSD-derived implementation the macOS backend actually calls was
|
||||
# untested, and the test skipped itself silently on this runner.
|
||||
#
|
||||
# `feth` is macOS's fake-Ethernet pseudo-interface. It is created
|
||||
# address-less, and the check below fails the leg rather than testing the
|
||||
# wrong thing if this runner hands it one anyway — the same shape as the
|
||||
# Linux fixture step, which needs `addrgenmode none` for exactly that
|
||||
# reason.
|
||||
- name: Create an address-less interface for the presence probe
|
||||
run: |
|
||||
sudo ifconfig feth0 create
|
||||
sudo ifconfig feth0 up
|
||||
ifconfig feth0
|
||||
if ifconfig feth0 | grep -qE "^[[:space:]]*inet6? "; then
|
||||
echo "feth0 has an address; it cannot test the address-less case" >&2
|
||||
exit 1
|
||||
fi
|
||||
echo "FIPS_TEST_ADDRLESS_IFACE=feth0" >> "$GITHUB_ENV"
|
||||
# Declare that this runner has fixtures, so a test that depends on
|
||||
# one fails when the fixture is missing instead of skipping silently.
|
||||
echo "FIPS_TEST_REQUIRE_FIXTURES=1" >> "$GITHUB_ENV"
|
||||
|
||||
- name: Run unit tests
|
||||
run: cargo nextest run --all --profile ci
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Job 2bb – Unit tests (musl)
|
||||
#
|
||||
# OpenWrt — the platform the Ethernet transport's dynamic interface binding
|
||||
# exists for — is musl, and musl reimplements the libc calls that binding is
|
||||
# built on rather than sharing glibc's. `interface_present` reads `ifa_flags`
|
||||
# out of `getifaddrs`, and the interfaces it has to see (`fips-mesh0`,
|
||||
# `fips-ap0`) are deliberately unbridged with no IP address at all, which is
|
||||
# exactly where getifaddrs implementations differ. Every other leg is glibc, so
|
||||
# without this one the presence probe is asserted on a libc no test has ever
|
||||
# run it against, on the target it was written for.
|
||||
#
|
||||
# Built for the musl target on a glibc host rather than inside an Alpine
|
||||
# container. The test binary links musl statically and runs natively on the
|
||||
# runner, so musl's `getifaddrs` is the one under test — while the build
|
||||
# scripts stay host artifacts, which keeps rustables' bindgen on the same
|
||||
# libclang the glibc leg already builds with. Building inside Alpine put
|
||||
# bindgen on a musl toolchain it does not work on: statically linked build
|
||||
# scripts cannot `dlopen` libclang, and turning the static CRT off then left
|
||||
# it loading libclang but unable to parse. None of that is anything this leg
|
||||
# is trying to test.
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
test-musl:
|
||||
name: Unit tests (musl)
|
||||
runs-on: ubuntu-latest
|
||||
needs: [build]
|
||||
steps:
|
||||
- uses: actions/checkout@d23441a48e516b6c34aea4fa41551a30e30af803 # v6
|
||||
|
||||
- name: Set SOURCE_DATE_EPOCH from git
|
||||
run: echo "SOURCE_DATE_EPOCH=$(git log -1 --format=%ct)" >> "$GITHUB_ENV"
|
||||
|
||||
# libdbus for the host build scripts; musl-tools for the musl C
|
||||
# toolchain the `cc`-driven dependencies link against. BLE is excluded on
|
||||
# musl by a Cargo.toml cfg, so bluer is not in this build at all.
|
||||
- name: Install system dependencies
|
||||
run: sudo apt-get update && sudo apt-get install -y libdbus-1-dev musl-tools
|
||||
|
||||
# The address-less interface the presence probe has to be tested
|
||||
# against; see the matching step on the glibc leg for why loopback
|
||||
# cannot stand in for it.
|
||||
- name: Create an address-less interface for the presence probe
|
||||
run: |
|
||||
sudo ip link add fips-probe0 type dummy
|
||||
# `addrgenmode none` before bringing it up: the kernel hands an IPv6
|
||||
# link-local to any interface that comes up, and an interface with a
|
||||
# link-local is not address-less — the fixture would have quietly
|
||||
# tested nothing.
|
||||
sudo ip link set fips-probe0 addrgenmode none
|
||||
sudo ip link set fips-probe0 up
|
||||
ip addr show fips-probe0
|
||||
# Fail rather than test the wrong thing if it acquired one anyway.
|
||||
if ip addr show fips-probe0 | grep -qE "inet6? "; then
|
||||
echo "fips-probe0 has an address; it cannot test the address-less case" >&2
|
||||
exit 1
|
||||
fi
|
||||
echo "FIPS_TEST_ADDRLESS_IFACE=fips-probe0" >> "$GITHUB_ENV"
|
||||
# Declare that this runner has fixtures, so a test that depends on
|
||||
# one fails when the fixture is missing instead of skipping silently.
|
||||
echo "FIPS_TEST_REQUIRE_FIXTURES=1" >> "$GITHUB_ENV"
|
||||
|
||||
- name: Install Rust toolchain
|
||||
uses: actions-rust-lang/setup-rust-toolchain@166cdcfd11aee3cb47222f9ddb555ce30ddb9659 # v1
|
||||
with:
|
||||
cache: false
|
||||
rustflags: ''
|
||||
target: x86_64-unknown-linux-musl
|
||||
|
||||
- name: Cache Cargo registry + build
|
||||
uses: actions/cache@caa296126883cff596d87d8935842f9db880ef25 # v5
|
||||
with:
|
||||
path: |
|
||||
~/.cargo/registry
|
||||
~/.cargo/git
|
||||
target
|
||||
key: musl-cargo-${{ hashFiles('**/Cargo.lock') }}
|
||||
restore-keys: |
|
||||
musl-cargo-
|
||||
|
||||
- name: Run library tests
|
||||
run: cargo test --lib --target x86_64-unknown-linux-musl
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Job 2c – Unit tests (Windows)
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
@@ -395,6 +556,7 @@ jobs:
|
||||
- name: Install cargo-nextest
|
||||
uses: taiki-e/install-action@nextest
|
||||
|
||||
|
||||
- name: Run unit tests
|
||||
run: cargo nextest run --all --profile ci
|
||||
|
||||
@@ -456,6 +618,9 @@ jobs:
|
||||
# ── Firewall baseline (fips0 nftables default-deny) ────────────
|
||||
- suite: firewall
|
||||
type: firewall
|
||||
# ── Dynamic interface binding (absent → present → absent) ──────
|
||||
- suite: iface-binding
|
||||
type: iface-binding
|
||||
# ── Outbound LAN gateway integration test ──────────────────────
|
||||
- suite: gateway
|
||||
type: gateway
|
||||
@@ -471,6 +636,9 @@ jobs:
|
||||
- suite: ethernet-only
|
||||
type: chaos
|
||||
scenario: ethernet-only
|
||||
- suite: ethernet-churn
|
||||
type: chaos
|
||||
scenario: ethernet-churn
|
||||
- suite: tcp-mesh
|
||||
type: chaos
|
||||
scenario: tcp-mesh
|
||||
@@ -606,6 +774,22 @@ jobs:
|
||||
run: |
|
||||
docker compose -f testing/firewall/docker-compose.yml down --volumes --remove-orphans
|
||||
|
||||
# ── Dynamic interface binding integration test ─────────────────────────
|
||||
- name: Run interface binding integration test
|
||||
if: matrix.type == 'iface-binding'
|
||||
run: bash testing/iface-binding/test.sh --skip-build --keep-up
|
||||
|
||||
- name: Collect logs on failure (iface-binding)
|
||||
if: matrix.type == 'iface-binding' && failure()
|
||||
run: |
|
||||
docker compose -f testing/iface-binding/docker-compose.yml logs --no-color
|
||||
docker exec fips-ifb-node-a fipsctl show transports || true
|
||||
|
||||
- name: Stop containers (iface-binding)
|
||||
if: matrix.type == 'iface-binding' && always()
|
||||
run: |
|
||||
docker compose -f testing/iface-binding/docker-compose.yml down --volumes --remove-orphans
|
||||
|
||||
# ── Chaos simulation ───────────────────────────────────────────────────
|
||||
- name: Install Python deps (chaos)
|
||||
if: matrix.type == 'chaos'
|
||||
|
||||
+254
@@ -117,6 +117,71 @@ with v0.5.x or earlier peers.
|
||||
- The receive-path `RejectReason` classification (shipped in 0.4.0) is
|
||||
additionally wired into the Noise XX handshake cluster
|
||||
(msg1/msg2/msg3) and the rekey-initiator outbound sites on `next`.
|
||||
- Dynamic interface binding for the Ethernet transport. An interface-bound
|
||||
transport is now a long-lived object that is *sometimes bound*: the interface
|
||||
it names need not exist when the daemon starts, may appear minutes later, and
|
||||
may vanish and return mid-operation. `start_async` returns `Ok` with the
|
||||
transport **absent** rather than failing, and a per-transport binder task
|
||||
binds when the interface appears, unbinds when it goes away, and rebinds when
|
||||
it returns. Start-time absence and runtime detach are one code path.
|
||||
Detection is event-driven where the kernel offers a source — netlink
|
||||
`RTNLGRP_LINK` on Linux, `PF_ROUTE` on macOS and FreeBSD — with a 1 s
|
||||
`getifaddrs` poll underneath as a backstop. Presence means
|
||||
`IFF_UP` — the interface exists and is administratively up — and
|
||||
deliberately not `IFF_RUNNING`. Binding needs no carrier and a socket
|
||||
outlives a carrier flap, so a bridge with nothing plugged into it (`br-lan`
|
||||
on a wifi-only router) is bound and healthy rather than permanently
|
||||
`Degraded`, and starts carrying traffic the moment a port comes up. Whether
|
||||
an interface has carrier is reported separately as `interface.carrier` in
|
||||
`show_transports`, never acted on.
|
||||
|
||||
This closes the OpenWrt boot race (procd starts `fips` before wifi has
|
||||
created `fips-mesh0` / `fips-ap0`; both transports were skipped for the life
|
||||
of the process while the 802.11s peer link formed anyway, so the node looked
|
||||
healthy and reached nothing), the intermittent-adapter case, and the
|
||||
mid-operation `wifi reload` that destroyed and recreated an interface under a
|
||||
live socket.
|
||||
|
||||
- `transports.ethernet.*.optional` (bool, default `false`). Naming an interface
|
||||
in configuration is a statement that you expect it, so the default is to
|
||||
complain: while a required interface is missing the node reports `Degraded`
|
||||
and logs the edge, at a severity that follows how long the absence lasts
|
||||
(see below). `optional: true` makes absence silent (`info` on the edge, no
|
||||
health impact) for hardware that is legitimately not always there. It describes the interface's *presence*, not the transport's
|
||||
importance — an optional interface that is present is used exactly as hard as
|
||||
any other — and no value of it makes a missing interface fatal at startup.
|
||||
|
||||
- `fipsctl show transports` reports interface presence per transport under a new
|
||||
`interface` block: `name`, `presence` (`absent` / `binding` / `present`),
|
||||
`carrier`, `policy` (`required` / `optional`), `since_secs`, `binds` and
|
||||
`failed_attempts`. The original boot-race bug was expensive precisely because
|
||||
nothing an operator could see said the node was deaf.
|
||||
|
||||
- `fipstop`'s transports view carries the same interface presence. The State
|
||||
column shows an interface-bound transport's presence rather than its
|
||||
lifecycle state — `up` is true from the moment the transport starts and stays
|
||||
true while its interface is missing, which is precisely the wrong answer in
|
||||
the one case someone is scanning that column for — and a new Policy column
|
||||
reads `required` or `optional` beside it, with an absent required interface
|
||||
red and an absent optional one yellow: the same split the daemon makes
|
||||
between staying `Full` and reporting `Degraded`. The instance name and the
|
||||
thing a transport is bound to are now separate columns, so netdev names line
|
||||
up down the list instead of trailing ragged inside a packed label. The detail
|
||||
pane gains an Interface block: netdev, presence and how long it has been
|
||||
held, carrier, what the absence policy means rather than which key sets it,
|
||||
bind count (flagged once it has rebound) and failed binds when there are any.
|
||||
The table fits an 80-column terminal — the OpenWrt serial console and the
|
||||
xterm and tmux default — dropping the byte counters below 100 columns and
|
||||
stacking the detail pane below 110, rather than shrinking every column until
|
||||
none of them can be read.
|
||||
|
||||
- `testing/iface-binding/` integration suite (`ci-local.sh --only
|
||||
iface-binding`, and a GitHub matrix leg): two daemons whose only transports
|
||||
are interface-bound, run against a veth pair the harness creates, downs,
|
||||
deletes and recreates underneath them. Asserts the boot race, the late
|
||||
attach and peering over it, the flap in both directions,
|
||||
destroy-and-recreate, that an `optional` interface never moves node health,
|
||||
and that absence is logged once on the edge rather than once per retry.
|
||||
|
||||
#### Node lifecycle
|
||||
|
||||
@@ -174,8 +239,23 @@ with v0.5.x or earlier peers.
|
||||
stay under it, or set `node.netmon.enabled: false`. A
|
||||
`node.link_dead_timeout_secs` of 0 is exempt from the check.
|
||||
|
||||
#### Library surface and internals
|
||||
|
||||
- `TransportError::InterfaceUnavailable { interface }`. A missing interface and
|
||||
a typo'd interface name were previously the same flat
|
||||
`StartFailed(String)`; nothing downstream could branch on absence.
|
||||
|
||||
### Changed
|
||||
|
||||
- The lockfile moves `chacha20` from 0.10.1 to 0.10.2, because 0.10.1 is yanked.
|
||||
It arrives through `rand`, a direct dependency,
|
||||
so it sits on the built path rather than off to one side. The requirement in
|
||||
`Cargo.toml` already admitted 0.10.2, so this is a lockfile change and no code
|
||||
changed with it. **This is not a security fix**: `cargo audit` reports nothing
|
||||
against `chacha20` at either version, and 0.10.1 was withdrawn by its
|
||||
maintainer rather than flagged by an advisory. What it buys is that a fresh
|
||||
checkout can resolve the lockfile without reaching for a yanked version.
|
||||
|
||||
- `node.rekey.enabled` now means "initiate rekeys" and nothing else. The
|
||||
responder half of the establish decision was also gated on it, and once the
|
||||
rekey is declared in the msg3 negotiation payload that flag was the only
|
||||
@@ -254,6 +334,127 @@ with v0.5.x or earlier peers.
|
||||
shared-media legs and every inbound leg, are unaffected and still promote
|
||||
whoever answers.
|
||||
|
||||
- `Degraded` is now a level rather than a latch. The supervisor's reason set
|
||||
was monotonic, which was correct while no child could recover; with recovery
|
||||
it would have meant "something broke at some point since boot" rather than
|
||||
"something is broken now". Interface absence is tracked in its own reversible
|
||||
set and node health is recomputed on every transition **in both directions**,
|
||||
so plugging the WAN back in clears `Degraded` without a restart. A transport
|
||||
whose interface is absent still counts as up, so a single-interface node that
|
||||
boots before its wifi degrades rather than exiting on "no transports".
|
||||
|
||||
- The OpenWrt package ships the `mesh0`/`mesh1` and `ap0`/`ap1` Ethernet
|
||||
transports **enabled** with `optional: true`, instead of commented out.
|
||||
`fips-mesh-setup` and `fips-ap-setup` no longer comment-toggle blocks in
|
||||
`fips.yaml`, and no longer tell the operator to restart the daemon after
|
||||
creating an interface — the daemon binds it on its own. `phy0-sta0` (`wwan`)
|
||||
is marked `optional: true` for the same reason: it only exists while a radio
|
||||
is in station mode.
|
||||
|
||||
**Upgrade note: an existing `/etc/fips/fips.yaml` is preserved and does not
|
||||
gain the new key.** It is a package conffile, so on a router where
|
||||
`fips-mesh-setup` or `fips-ap-setup` had already uncommented a block, that
|
||||
block stays as it was, with no `optional` key — and `optional` defaults to
|
||||
false. Such a block is therefore `required`, so an absent `fips-mesh0` keeps
|
||||
the node `Degraded` and is reported once at `error` ten seconds in, where the
|
||||
same block in the shipped file is silent. Add `optional: true` to the block
|
||||
to match what the package now ships.
|
||||
|
||||
- The Ethernet receive loop backs off and exits on a dead socket instead of
|
||||
spinning on `Err` with a `warn!` per iteration, and the ad-hoc ENXIO
|
||||
socket-reopen in the beacon sender is gone. Both hand recovery to the
|
||||
presence machine: one mechanism for every cause rather than one hack per
|
||||
symptom. Beacons pause while an interface is absent.
|
||||
|
||||
- The absence edge is not itself an error, and there is exactly one deadline
|
||||
after it. An interface missing when the daemon starts logs at `info` — that
|
||||
is the boot race the mechanism exists to absorb, not a fault — and a runtime
|
||||
detach at `warn`, because a link coming and going is ordinary weather for a
|
||||
mesh daemon. Ten seconds is the whole grace: past it, absence is no longer a
|
||||
race against a radio or a container, so a **required** interface still
|
||||
missing is reported once at `error`. Start-time absence and a runtime detach
|
||||
share that one deadline rather than getting one each. An `optional`
|
||||
interface never reaches `error`. Node health does not wait for any of it,
|
||||
publishing `Degraded` on the first edge either way.
|
||||
|
||||
- The TUN boundary's TCP MSS clamp now tracks the node's egress MTU at
|
||||
runtime instead of freezing it at startup. `transport_mtu()` is the minimum
|
||||
across *bound* transports, so a transport that binds minutes after start can
|
||||
be the narrow one — but the TUN reader and writer were handed a `u16`
|
||||
computed once when they spawned, while every other consumer
|
||||
(`show_status`, the control-socket snapshot, the session-layer fragmentation
|
||||
check) read it live. A node could therefore report one effective IPv6 MTU
|
||||
and clamp to another. The ceiling is now shared with those threads and
|
||||
recomputed whenever the bound set changes, in both directions: a narrow
|
||||
interface appearing tightens it, and its departure releases it. MSS is
|
||||
negotiated per connection, so a change applies to connections opened after
|
||||
it; existing ones are not disturbed.
|
||||
|
||||
- Rebinds that keep succeeding into a socket that dies moments later are
|
||||
damped: consecutive bindings shorter than ten seconds back off on the
|
||||
1 s → 30 s curve, and past three of them the binder stops announcing each
|
||||
bind as a recovery until one lasts. Undamped, a persistently broken socket
|
||||
behind a healthy interface produced a log pair and a `Degraded`→`Running`
|
||||
health flap every second.
|
||||
|
||||
- A bind failure that is **not** absence — no `CAP_NET_RAW`, no readable
|
||||
`/dev/bpf*`, a buffer the kernel refused — fails the daemon's start as it
|
||||
always has, rather than being waited out. It is a fault, not a state, and
|
||||
will not resolve on its own; only a missing interface is retried at start.
|
||||
A non-absence failure during a later rebind still backs off, since the node
|
||||
is serving by then.
|
||||
|
||||
- The binder cannot outlive its transport, and a teardown that races a bind
|
||||
cannot leave a live receive loop on a socket nothing owns. A shared stop flag
|
||||
is raised before teardown and checked by the binder after it stores a
|
||||
binding, so whichever order the two interleave exactly one of them cleans up;
|
||||
`EthernetTransport` gained a `Drop` that raises the flag, aborts the binder
|
||||
and releases the socket, for handles dropped without `stop_async`.
|
||||
|
||||
- Presence edges are published with `try_send` and retried on the next tick
|
||||
rather than awaited. A bounded channel could previously park the binder
|
||||
mid-publish — a health channel able to deadlock the machine whose health it
|
||||
carries — freezing the interface in whatever state it held.
|
||||
|
||||
- `TransportHandle::is_bound()` joins `is_operational()`: the latter means the
|
||||
transport was *started*, which for an interface-bound transport no longer
|
||||
implies a live socket. `Node::transport_mtu` now filters on the former,
|
||||
because an interface that has never existed was clamping the whole node's
|
||||
IPv6 MTU to a number derived from absent hardware.
|
||||
|
||||
- An interface deleted and recreated under the same name is detected as a
|
||||
detach. Both backends bind by device rather than by name, so the old socket
|
||||
is attached to nothing while the name still resolves — and a stale
|
||||
`AF_PACKET` socket never becomes readable, so nothing errors and nothing
|
||||
exits. Detection previously rested entirely on the beacon sender failing,
|
||||
which a node with `announce: false` does not have. The bound interface index
|
||||
is now captured at bind and compared on every poll.
|
||||
|
||||
- The link-event watcher distinguishes a genuine receive error from
|
||||
`WouldBlock`. `try_io` clears readiness only on the latter, so a persistent
|
||||
error — `ENOBUFS` after a burst of link events overflows the socket buffer —
|
||||
span a core flat with nothing logged. Errors are now counted, logged once,
|
||||
backed off, and after five the source is abandoned for the presence poll.
|
||||
|
||||
- Presence probes are coalesced to at most ten a second. Linux netlink is
|
||||
filtered to `RTNLGRP_LINK`, but `PF_ROUTE` has no group filter, so the macOS
|
||||
source delivers every routing message on the host — route churn, ARP, DHCP
|
||||
renewals, a VPN going up and down — and each would otherwise drive a full
|
||||
`getifaddrs` walk.
|
||||
|
||||
- CI runs the library tests on musl (Alpine) as well as glibc. Presence is
|
||||
built on `getifaddrs` and `ifa_flags`, musl reimplements both independently,
|
||||
and the interfaces this feature exists for (`fips-mesh0`, `fips-ap0` on
|
||||
OpenWrt) are unbridged with no IP address at all — the case where
|
||||
implementations most plausibly differ. It was previously asserted on a libc
|
||||
no test had ever run it against, on the target it was written for.
|
||||
|
||||
- Interface presence state ignores lock poisoning. Treating a poisoned lock as
|
||||
a failure meant reading "no socket, tasks dead", which is the destructive
|
||||
direction: a transport reporting itself present while every send fails, or a
|
||||
binder tearing down and rebinding every second while teardown silently
|
||||
declined to abort anything.
|
||||
|
||||
### Fixed
|
||||
|
||||
- A leaf-profile node no longer self-elects as tree root. A leaf holding the
|
||||
@@ -360,6 +561,59 @@ with v0.5.x or earlier peers.
|
||||
|
||||
#### Node lifecycle
|
||||
|
||||
- Losing an interface no longer leaves its peers in the routing table. The
|
||||
peers stayed in the registry, the routes through them stayed selectable, and
|
||||
the node kept advertising reachability it no longer had — so transit traffic
|
||||
was dropped in silence and other nodes kept routing toward this one for those
|
||||
destinations, until the liveness reaper noticed up to
|
||||
`node.link_dead_timeout_secs` later. Measured on real hardware, a detached
|
||||
dongle took the node's parent with it and no new parent was chosen for
|
||||
twenty-seven seconds, with four alternative peers available the whole time. A
|
||||
transport's detach edge now withdraws every peer whose active link runs over
|
||||
it, on the same path the liveness reaper uses, so sessions, path MTU, session
|
||||
indices, the link, the control machine, tree cleanup and re-announce, and
|
||||
bloom withdrawal unwind exactly as they already did. It is not filtered by
|
||||
`optional`: whether an interface's absence is normal is a statement about
|
||||
node health, and says nothing about whether the routes over it still work.
|
||||
The trade is that an absence shorter than the dead timeout that then recovers
|
||||
now costs a re-peer where it previously cost nothing, accepted because
|
||||
black-holing is silent, poisons other nodes' routing and takes the full
|
||||
timeout to clear, where a re-peer is bounded, visible and self-healing.
|
||||
|
||||
- A local interface flap during a handshake is no longer charged to the remote.
|
||||
A msg2 send refused because the interface is absent or mid-rebind was treated
|
||||
as a failed handshake: the link was removed, the reverse-address entry
|
||||
dropped, the session index freed, the control machine torn down, and the
|
||||
whole thing recorded under the reject reason that means "the remote sent
|
||||
something invalid", which is what an operator reading the rejects would have
|
||||
concluded. The initiator meanwhile resent msg1 into a link that no longer
|
||||
existed and had to rebuild from nothing. A transport error the daemon is
|
||||
already working to clear now leaves the half-built link exactly where it is
|
||||
for that resend to land on; only a terminal error still tears down, and a
|
||||
link nobody resends to is reaped at `node.rate_limit.handshake_timeout_secs`
|
||||
like every other abandoned handshake. The rekey msg1 send site keeps its
|
||||
teardown, which was already benign, and stops reporting a local self-clearing
|
||||
condition at `warn`.
|
||||
|
||||
- A peer reachable over two interfaces is no longer re-dialled on the path it
|
||||
is not using. Beacon discovery skipped only a candidate naming the peer's
|
||||
*current* path, which is the one case that could not churn anything, so the
|
||||
alternate path was dialled every discovery tick; each dial that completed
|
||||
promoted and displaced a healthy incumbent, tore down the session, and, when
|
||||
that peer was the parent, switched parents and re-announced mesh-wide.
|
||||
Measured on real hardware, seventeen dials to one peer in fifteen minutes,
|
||||
alternating wifi and cable, displacing a link reporting etx 1.0 and loss 0.0.
|
||||
Discovery now asks whether the link it already holds is answering rather than
|
||||
which path the candidate names. Failover is unchanged: a peer that goes quiet
|
||||
for longer than `node.heartbeat_interval_secs` is dialled again on every
|
||||
path, alternate included. **One behaviour goes with it.** A peer held on an
|
||||
adopted NAT-traversal transport that is *also* reachable by Ethernet or BLE
|
||||
beacon used to drift onto the local path on the next discovery tick, and now
|
||||
stays on the traversed path for as long as that path answers. Migrating it is
|
||||
still done by the configured-peer refresh (a config reload, a runtime peer
|
||||
update, or `fipsctl connect`), and a traversed link that goes quiet still
|
||||
releases the peer to every path.
|
||||
|
||||
- A heartbeat whose send failed no longer counts as one that was delivered.
|
||||
The peer's "last heartbeat" timestamp was stamped before the send and left
|
||||
alone whatever came back, so a failure suppressed the next attempt for a
|
||||
|
||||
Generated
+3
-3
@@ -498,9 +498,9 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "chacha20"
|
||||
version = "0.10.1"
|
||||
version = "0.10.2"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "d524456ba66e72eb8b115ff89e01e497f8e6d11d78b70b1aa13c0fbd97540a81"
|
||||
checksum = "65c35e4b699c7e15ccbe7ee35c005e4fc0a278d22238a2857e6ce2dadeda1b06"
|
||||
dependencies = [
|
||||
"cfg-if",
|
||||
"cpufeatures 0.3.0",
|
||||
@@ -2614,7 +2614,7 @@ version = "0.10.2"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "c7f5fa3a058cd35567ef9bfa5e75732bee0f9e4c55fa90477bef2dfcdbc4be80"
|
||||
dependencies = [
|
||||
"chacha20 0.10.1",
|
||||
"chacha20 0.10.2",
|
||||
"getrandom 0.4.3",
|
||||
"rand_core 0.10.1",
|
||||
]
|
||||
|
||||
@@ -1008,6 +1008,320 @@ Transports begin in `Configured` state with all parameters set. `start()`
|
||||
transitions through `Starting` to `Up` (operational). `stop()` moves to
|
||||
`Down`. Transport failures move to `Failed`.
|
||||
|
||||
## Interface Presence
|
||||
|
||||
`Up` describes the *transport*, not the socket. An interface-bound transport
|
||||
(today: Ethernet) carries a second, orthogonal state — whether it is bound
|
||||
right now — and the two are independent: a transport is `Up` from the moment
|
||||
it starts, whether or not the interface it names exists.
|
||||
|
||||
### The Gap This Closes
|
||||
|
||||
Three deployment scenarios exercise one missing mechanism:
|
||||
|
||||
- **Boot ordering.** On OpenWrt, procd starts `fips` before wifi has created
|
||||
`fips-mesh0` / `fips-ap0`. Both transports were skipped and never retried,
|
||||
while the 802.11s peer link formed anyway — that is mac80211, not the
|
||||
daemon — so the node looked healthy and reached nothing. The failure was
|
||||
expensive precisely because nothing an operator could see said the node was
|
||||
deaf.
|
||||
- **Intermittent hardware.** A USB ethernet adapter named in `fips.yaml` is
|
||||
plugged in some days and not others. Its absence is normal and must be
|
||||
silent; its arrival must bind without operator action.
|
||||
- **Mid-operation restart.** `wifi reload` for a channel change destroys and
|
||||
recreates the mesh interface within a couple of seconds. The socket dies,
|
||||
the receive loop spun on `Err` with no backoff and no exit, and nothing
|
||||
rebound.
|
||||
|
||||
These are not three features. They are one presence machine plus one policy
|
||||
field. Before it existed, the first observation was final: an interface
|
||||
missing at start was logged once and skipped for the life of the process, and
|
||||
one that disappeared at runtime published a health change but was never
|
||||
rebound.
|
||||
|
||||
### The Presence Machine
|
||||
|
||||
```text
|
||||
Absent ──attach──> Binding ──ok──> Present
|
||||
^ │ │
|
||||
└──── fail/backoff ─┘ │
|
||||
└──────────── detach ──────────────┘
|
||||
```
|
||||
|
||||
`start_async` binds if it can and otherwise returns `Ok` with the transport
|
||||
`Up` and `Absent`; a per-transport binder task then binds when the interface
|
||||
appears, tears the socket down when it goes away, and rebinds when it
|
||||
returns. Two invariants do the work:
|
||||
|
||||
- **The transport object survives detach.** Config, `TransportId`,
|
||||
statistics and the neighbor buffer persist; only the file descriptor and
|
||||
its loops go. A transport is never destroyed because its interface went
|
||||
away.
|
||||
- **Start-time absence and runtime detach are the same transition.** A node
|
||||
that boots before its wifi and a node whose wifi reloads at 03:00 take one
|
||||
code path. The old asymmetry — skip forever at start, publish health at
|
||||
runtime — is gone.
|
||||
|
||||
`TransportError::InterfaceUnavailable` is what makes absence branchable. A
|
||||
missing interface and a typo'd interface name were the same flat
|
||||
`StartFailed(String)`, so nothing downstream could tell a state from a fault.
|
||||
|
||||
### What Counts as Present
|
||||
|
||||
Presence means `IFF_UP` — the interface exists and the operator has enabled
|
||||
it — and deliberately **not** `IFF_RUNNING`.
|
||||
|
||||
Carrier and bindability are different questions, and only the second belongs
|
||||
in a bind gate. An `AF_PACKET` socket on a carrier-less bridge is valid and
|
||||
starts carrying traffic the instant a member port comes up, with no rebind:
|
||||
the socket outlives the carrier. Gating on `IFF_RUNNING` bought nothing and
|
||||
cost three things:
|
||||
|
||||
- `br-lan` on a router with nothing in its LAN ports is `UP` with
|
||||
`NO-CARRIER`, so a healthy wifi-only router reported `Degraded` forever and
|
||||
errored for a fault it did not have;
|
||||
- every carrier flap the socket would have survived became an unbind/rebind
|
||||
cycle — churn the presence machine then has to damp, a mechanism
|
||||
compensating for a policy error;
|
||||
- an 802.11s interface that reports `RUNNING` only once it has peered cannot
|
||||
peer, because peering needs beacons, beacons need a bound socket, and the
|
||||
gate refuses to bind. A deadlock reachable on the hardware this mechanism
|
||||
was written for.
|
||||
|
||||
The signal `IFF_RUNNING` carries is not lost: `show_transports` reports
|
||||
`interface.carrier` beside presence, so an operator can still tell a bound
|
||||
transport carrying nothing from a working one. It is reported rather than
|
||||
obeyed.
|
||||
|
||||
The probe is `getifaddrs` plus `ifa_flags` rather than an `SIOCGIFFLAGS`
|
||||
ioctl: it needs no socket, so the watcher can probe before any file
|
||||
descriptor exists, and it is spelled the same on Linux and the BSDs.
|
||||
|
||||
### Interface Identity
|
||||
|
||||
The configured name is the key, but a name is not a device. Both backends
|
||||
bind by *device* — `AF_PACKET` stores `sll_ifindex`, a BPF descriptor follows
|
||||
the interface it was attached to — so an interface deleted and recreated
|
||||
under the same name leaves the socket attached to something that no longer
|
||||
exists while the name resolves perfectly well.
|
||||
|
||||
Nothing else notices. A stale `AF_PACKET` socket never becomes readable, so
|
||||
the receive loop neither errors nor exits, and send failures go to the caller
|
||||
rather than to the binder. A listen-only node (`announce: false`, so no
|
||||
beacon sender to fail) therefore sat `present` and deaf indefinitely after a
|
||||
`wifi reload` — the original bug wearing a different hat. The bound index is
|
||||
captured at bind and compared on every poll; a mismatch is a detach.
|
||||
|
||||
Hardware can also change underneath a name. If the name reappears with a MAC
|
||||
other than the one last bound, that is a different device, so the cached
|
||||
neighbor entries for that transport are dropped rather than resumed onto, and
|
||||
the swap is logged at `warn`. Richer selectors (`match: { name | mac |
|
||||
id_path }`) are deliberately deferred; the requirement here is only that FIPS
|
||||
never silently resumes onto different hardware.
|
||||
|
||||
### Detection
|
||||
|
||||
| Platform | Source |
|
||||
| -------- | ------ |
|
||||
| Linux | netlink `RTNLGRP_LINK` (`RTM_NEWLINK` / `RTM_DELLINK`) |
|
||||
| macOS, FreeBSD† | `PF_ROUTE` socket, `RTM_IFINFO` |
|
||||
| Fallback | poll `getifaddrs` + flags, 1 s |
|
||||
|
||||
† Aspirational: the Ethernet transport is
|
||||
`cfg(any(target_os = "linux", target_os = "macos"))`, so FreeBSD has no
|
||||
interface-bound transport for a watcher to serve. The `PF_ROUTE` branch
|
||||
compiles for the BSD family, but only macOS reaches it.
|
||||
|
||||
Where an event source exists, detection is sub-second. The poll stays
|
||||
underneath as a backstop rather than as the mechanism, and must stay at ~1 s:
|
||||
the probe is cheap, and letting the interval drift to tens of seconds
|
||||
reintroduces exactly the latency the event source was added to remove.
|
||||
Construction is best-effort — a kernel or sandbox that refuses the socket
|
||||
yields a watcher that never fires, and the binder degrades to its poll.
|
||||
|
||||
Link-event payloads are **not parsed**. An event is a hint to re-run the
|
||||
presence probe, which is cheap and authoritative; decoding
|
||||
`nlmsghdr`/`ifinfomsg` to reach the same answer would add a parser whose bugs
|
||||
would be presence bugs.
|
||||
|
||||
Two rate limits protect the binder from its own event source. Probes are
|
||||
coalesced to ten a second, because `PF_ROUTE` has no group filter and
|
||||
delivers every routing message on the host — route churn, ARP, DHCP renewals,
|
||||
a VPN going up and down — each of which would otherwise drive a full
|
||||
`getifaddrs` walk. And a persistently failing event source is counted, logged
|
||||
once, backed off, and after five consecutive errors abandoned for the poll:
|
||||
losing events is survivable because the poll is the backstop, but burning a
|
||||
core on a socket that is readable-but-erroring is not.
|
||||
|
||||
Bind failures that are *not* absence back off 1 s → 30 s. Absence itself does
|
||||
not back off; there is nothing to poll but the probe.
|
||||
|
||||
### Policy: `optional`
|
||||
|
||||
One field per transport, `transports.ethernet.*.optional`, default `false`:
|
||||
|
||||
| | absence | log | retries |
|
||||
| --- | --- | --- | --- |
|
||||
| `optional: false` (default) | node reports `Degraded` | `info` at boot / `warn` on a runtime detach, then `error` once if it lasts past 10 s | forever |
|
||||
| `optional: true` | no health impact | `info`, and nothing after | forever |
|
||||
|
||||
Naming an interface in configuration is a statement that you expect it, so
|
||||
the default is to complain; silence is opted into.
|
||||
|
||||
`optional` describes **the interface's presence, not the transport's
|
||||
importance**. An optional interface that is present is used exactly as hard
|
||||
as any other. No value of it makes a missing interface fatal at startup: the
|
||||
only fatal case remains "no transports at all came up". If a deployment ever
|
||||
needs absence to abort startup, that arrives as an explicit `on_absent: exit`
|
||||
— never as a second meaning for `optional`.
|
||||
|
||||
A bind failure that is not absence — no `CAP_NET_RAW`, no readable
|
||||
`/dev/bpf*`, a buffer the kernel refused — is a fault, not a state, and still
|
||||
fails the daemon's start. Retrying those forever would convert a hard,
|
||||
actionable deployment error into a daemon that retries a socket it can never
|
||||
open behind a `Degraded` nobody is watching. Only absence is waited out at
|
||||
start; a non-absence failure during a later *rebind* does back off, since by
|
||||
then the node is serving and killing it would be the worse answer.
|
||||
|
||||
### Health
|
||||
|
||||
Node health is recomputed on every presence transition **in both
|
||||
directions**, so a returning interface clears `Degraded` without a restart.
|
||||
That makes `Degraded` a level rather than a latch: the supervisor's reason set
|
||||
was monotonic, which was correct while no child could recover, but with
|
||||
recovery it would have come to mean "something broke at some point since boot"
|
||||
rather than "something is broken now". Absence lives in its own reversible
|
||||
set, separate from the one-way `failed` set a start failure enters.
|
||||
|
||||
An absent transport still counts as *up*. It came up — `start_async` returned
|
||||
`Ok` — so it does not push a single-transport node into the fatal
|
||||
`NoTransports`, which would make a node that merely booted before its wifi
|
||||
exit instead of waiting. Absence degrades; it never kills.
|
||||
|
||||
There is deliberately **no restart action in the supervisor FSM.** The
|
||||
presence watcher and the rebind loop live inside the transport, next to the
|
||||
file descriptor they manage, and once that exists a supervisor-authored retry
|
||||
has nothing left to do — it would be a second mechanism racing the first for
|
||||
the same socket. The supervisor learns about presence
|
||||
(`Event::ChildAbsent` / `Event::ChildPresent`) and republishes health; it does
|
||||
not drive rebinding.
|
||||
|
||||
Peer state gets no grace period, and needs none. A transport's detach edge
|
||||
withdraws every peer whose active link runs over it, on the same path the
|
||||
liveness reaper uses, so those peers and the routes through them are gone at
|
||||
the edge rather than up to `link_dead_timeout_secs` later — there is nothing
|
||||
left for a linger timer to bound. A send over an absent interface still
|
||||
returns `InterfaceUnavailable`, and a *half-built* link is still held on that
|
||||
error, because the binder is already working to bring the interface back and
|
||||
the initiator's resend has somewhere to land; an established peer is not held.
|
||||
The trade is deliberate: an absence shorter than the dead timeout that then
|
||||
recovers now costs a re-peer where it previously cost nothing, and that is
|
||||
accepted because black-holing is silent, poisons other nodes' routing and
|
||||
takes the full timeout to clear, where a re-peer is bounded, visible and
|
||||
self-healing. A recreated mesh interface comes back with the same MAC (it is
|
||||
derived from the phy), so the local address peers hold is unchanged across the
|
||||
rebind.
|
||||
|
||||
### Logging
|
||||
|
||||
Edges, never attempts. A loop that logs per attempt reproduces the hot log
|
||||
spin this mechanism removed, at 1–30 s intervals forever on any router with
|
||||
an unplugged WAN — and operators learn to filter it, which is how the next
|
||||
real failure gets missed.
|
||||
|
||||
The edge itself is not an error. An interface missing when the daemon starts
|
||||
and bound a moment later is the ordinary case the mechanism exists to absorb,
|
||||
so it is `info`; calling it an error at t=0 and "recovered" at t=0.2 s is the
|
||||
cry-wolf failure this rule exists to prevent. A runtime detach is `warn` — a
|
||||
link coming and going is ordinary weather for a mesh daemon.
|
||||
|
||||
There is exactly one deadline. Ten seconds is the window in which absence
|
||||
could still be a race — a radio, a container, a veth arriving late. Past it a
|
||||
**required** interface is a fault an operator has to fix, and it is reported
|
||||
once at `error`. Start-time absence and a runtime detach share that deadline
|
||||
rather than getting one each, for the same reason they share a code path
|
||||
everywhere else here. An `optional` interface never reaches `error`; that is
|
||||
what `optional` means.
|
||||
|
||||
Once, not repeated. This was a 1 m / 10 m / 1 h ladder that re-announced the
|
||||
same fact at rising severity and then went permanently quiet after an hour,
|
||||
which got both halves wrong: it used the log as a store for something already
|
||||
published continuously as state, and it stopped mentioning a fault that was
|
||||
still live. Duration belongs in `interface.since_secs` and in how long
|
||||
`Degraded` has been held, where a monitor can threshold it per deployment
|
||||
instead of the daemon compiling one in.
|
||||
|
||||
Node health does not wait for the deadline. `Degraded` publishes on the first
|
||||
edge, which is the signal an operator actually watches.
|
||||
|
||||
Successful rebinds are damped. Backoff covers failed binds; the opposite and
|
||||
nastier case is binds that keep *succeeding* into a socket that dies moments
|
||||
later, which a receive loop giving up on a persistent error while the
|
||||
interface stays `UP` produces once per second, forever. The binder counts
|
||||
consecutive bindings that die inside ten seconds, backs off on the same
|
||||
1 s → 30 s curve, and past three of them stops announcing each bind as a
|
||||
recovery — holding health where it is until a binding lasts.
|
||||
|
||||
### Egress MTU
|
||||
|
||||
`transport_mtu()` is the minimum across *bound* transports — `is_bound()`,
|
||||
not `is_operational()`, because an interface-bound transport is operational
|
||||
from the moment it starts whether or not it holds a socket. Filtering on the
|
||||
weaker predicate let a transport whose interface had never appeared set the
|
||||
whole node's IPv6 MTU from hardware that was not present.
|
||||
|
||||
Since a transport can now bind long after start, that minimum moves at
|
||||
runtime, and every consumer has to read it live. `show_status`, the
|
||||
control-socket snapshot and the session-layer fragmentation check always did.
|
||||
The TUN reader and writer did not: they were handed a `u16` at spawn, so a
|
||||
narrow interface binding later never tightened the TCP MSS clamp and the node
|
||||
reported one effective MTU while clamping to another. The ceiling is now
|
||||
shared with those threads — an atomic beside the per-destination
|
||||
`path_mtu_lookup` they already read on the same packet — and recomputed on
|
||||
every change to the bound set.
|
||||
|
||||
Both directions, for the same reason `Degraded` is a level rather than a
|
||||
latch: a narrow interface arriving must tighten the clamp or traffic
|
||||
egressing over it is clamped too loose, and that interface leaving must
|
||||
release it or unplugging a low-MTU adapter leaves the node over-clamped until
|
||||
it restarts. MSS is negotiated per connection at SYN time, so a change binds
|
||||
connections opened after it and leaves established ones alone.
|
||||
|
||||
### Observability
|
||||
|
||||
`show_transports` carries an `interface` block per interface-bound transport:
|
||||
netdev name, `presence` (`absent` / `binding` / `present`), `carrier`,
|
||||
`policy` (`required` / `optional`), `since_secs`, `binds` and
|
||||
`failed_attempts`. The two counters separate an interface that is flapping
|
||||
from one that is there and refusing to bind, and `since_secs` measures the
|
||||
absence *episode* rather than the phase — a bind that fails walks
|
||||
`Absent → Binding → Absent`, and restarting the clock on those edges would
|
||||
report a permanently unbindable interface as one second old forever.
|
||||
|
||||
`fipstop`'s transports view names the netdev and the absence policy in their
|
||||
own columns and shows presence in the State column for these transports,
|
||||
because `state` reads `up` from the moment the transport starts and is
|
||||
therefore precisely the wrong answer in the one case someone is scanning that
|
||||
column for. Both render sites sort by ascending transport id — creation
|
||||
order, and so grouped by transport type — rather than by `HashMap` iteration
|
||||
order, which was arbitrary and differed on every daemon restart.
|
||||
|
||||
### What This Retires
|
||||
|
||||
- The `hotplug.d/net` rule that restarted the daemon when the FIPS radio
|
||||
interfaces appeared, and the `wifi down; wifi up; sleep` dance provisioning
|
||||
performed to sequence around the race.
|
||||
- The YAML comment-toggling in `fips-mesh-setup` / `fips-ap-setup` — the mesh
|
||||
and AP blocks ship enabled with `optional: true` and simply wait.
|
||||
- The ad-hoc ENXIO socket reopen in the beacon sender: beacons now pause while
|
||||
absent because the task does not exist then, and recovery is the presence
|
||||
machine's job. One mechanism for every cause rather than one hack per
|
||||
symptom.
|
||||
- The start-time versus runtime asymmetry in the supervisor.
|
||||
|
||||
It also covers the case none of those workarounds did: an interface that flaps
|
||||
while the daemon is running.
|
||||
|
||||
## Implementation Status
|
||||
|
||||
| Transport | Status | Notes |
|
||||
|
||||
@@ -72,9 +72,9 @@ fips-mesh-setup radio1
|
||||
|
||||
This creates an open 802.11s interface with mesh ID `fips-mesh` and
|
||||
HWMP forwarding off, attaches it to an unmanaged netifd interface (no
|
||||
IP configuration — none is needed), uncomments the matching `meshN`
|
||||
transport entry in `/etc/fips/fips.yaml` (see Step 2), and reloads the
|
||||
radio. Interfaces are named by radio index: `radio0` → `fips-mesh0`,
|
||||
IP configuration — none is needed), and reloads the radio. It does not
|
||||
touch `/etc/fips/fips.yaml`: the matching `meshN` transport ships
|
||||
enabled and the daemon binds the interface once it exists (see Step 2). Interfaces are named by radio index: `radio0` → `fips-mesh0`,
|
||||
`radio1` → `fips-mesh1`. Pass a second argument to use a different
|
||||
mesh ID.
|
||||
|
||||
@@ -133,18 +133,20 @@ wifi reload
|
||||
## Step 2 — check the FIPS transport binding
|
||||
|
||||
The `fips.yaml` shipped in the OpenWrt package carries one transport
|
||||
entry per radio, but **commented out** — so a stock install that never
|
||||
runs this helper logs no per-boot "interface missing" warning.
|
||||
`fips-mesh-setup` uncommented the matching `meshN` entry in Step 1, so
|
||||
there is normally nothing to do here. If you maintain your own config
|
||||
(or ran the manual UCI above instead of the helper), make sure the
|
||||
entries are present and uncommented:
|
||||
entry per radio, **enabled** and marked `optional: true`. The daemon
|
||||
treats a named interface that is not there as absent rather than as a
|
||||
failure, and `optional: true` is what keeps a stock install that never
|
||||
runs this helper quiet and un-`Degraded` about a radio it was never
|
||||
going to have. There is normally nothing to do here. If you maintain
|
||||
your own config (or ran the manual UCI above instead of the helper),
|
||||
make sure the entries are present:
|
||||
|
||||
```yaml
|
||||
transports:
|
||||
ethernet:
|
||||
mesh0:
|
||||
interface: "fips-mesh0"
|
||||
optional: true
|
||||
listen: true
|
||||
announce: true
|
||||
auto_connect: true
|
||||
@@ -161,19 +163,33 @@ transports:
|
||||
parses as an alias, so an existing config keeps working (see
|
||||
[../reference/configuration.md](../reference/configuration.md)).
|
||||
|
||||
## Step 3 — restart the daemon (order matters)
|
||||
## Step 3 — no restart needed
|
||||
|
||||
The daemon binds an interface when it appears. A transport whose
|
||||
interface is missing is *absent*, not skipped: it waits, binds within
|
||||
a second of the interface coming up, unbinds if it goes away, and
|
||||
rebinds when it returns. Order does not matter, and neither
|
||||
`/etc/init.d/fips restart` nor any hotplug rule is part of this
|
||||
procedure.
|
||||
|
||||
Watch it happen:
|
||||
|
||||
```sh
|
||||
fipsctl show transports
|
||||
```
|
||||
|
||||
The transport's `interface` block reports `presence` (`absent` /
|
||||
`binding` / `present`), `policy` (`required` / `optional`) and how
|
||||
long it has held that state.
|
||||
|
||||
If you *changed a config value* above rather than only creating an
|
||||
interface, that does need a restart — configuration is read at
|
||||
startup, interfaces are not:
|
||||
|
||||
```sh
|
||||
/etc/init.d/fips restart
|
||||
```
|
||||
|
||||
Restart fips **after** the mesh interface is up. A transport whose
|
||||
interface is missing at startup is logged and skipped, not retried —
|
||||
so if the daemon comes up before the radio, the mesh transport stays
|
||||
dead until the next restart. (An interface that *vanishes and
|
||||
returns* after startup is recovered automatically; only the missing-
|
||||
at-startup case needs this ordering.)
|
||||
|
||||
## Verify
|
||||
|
||||
L2 first — the 802.11s peering, with a second configured router in
|
||||
|
||||
@@ -161,15 +161,17 @@ TCP 8443).
|
||||
## Step 2 — check the FIPS transport binding
|
||||
|
||||
The `fips.yaml` shipped in the OpenWrt package carries one transport
|
||||
entry per access interface, but **commented out** — so a stock install
|
||||
that never runs this helper logs no per-boot "interface missing"
|
||||
warning. `fips-ap-setup` uncommented the matching `apN` entry in Step 1,
|
||||
and also enabled `node.rendezvous.lan` (the daemon's mDNS/DNS-SD
|
||||
rendezvous — phone FIPS apps cannot see raw-Ethernet beacons, so mDNS
|
||||
is how they find the daemon; the switch is daemon-wide and stays on if
|
||||
you later remove the AP). So there is normally nothing to do here. If
|
||||
you maintain your own config (or ran the manual UCI above instead of
|
||||
the helper), make sure both are present and uncommented:
|
||||
entry per access interface, **enabled** and marked `optional: true` —
|
||||
the daemon binds the interface once `fips-ap-setup` creates it, and
|
||||
`optional: true` keeps a stock install that never runs the helper quiet
|
||||
and un-`Degraded`. `fips-ap-setup` does still edit one thing: it enables
|
||||
`node.rendezvous.lan` (the daemon's mDNS/DNS-SD rendezvous — phone FIPS
|
||||
apps cannot see raw-Ethernet beacons, so mDNS is how they find the
|
||||
daemon; the switch is daemon-wide and stays on if you later remove the
|
||||
AP). That one is a config value rather than an interface, so it needs a
|
||||
restart to take effect. Otherwise there is normally nothing to do here.
|
||||
If you maintain your own config (or ran the manual UCI above instead of
|
||||
the helper), make sure both are present:
|
||||
|
||||
```yaml
|
||||
node:
|
||||
@@ -183,6 +185,7 @@ transports:
|
||||
ethernet:
|
||||
ap0:
|
||||
interface: "fips-ap0"
|
||||
optional: true
|
||||
listen: true
|
||||
announce: true
|
||||
auto_connect: true
|
||||
@@ -199,19 +202,33 @@ transports:
|
||||
parses as an alias, so an existing config keeps working (see
|
||||
[../reference/configuration.md](../reference/configuration.md)).
|
||||
|
||||
## Step 3 — restart the daemon (order matters)
|
||||
## Step 3 — no restart needed
|
||||
|
||||
The daemon binds an interface when it appears. A transport whose
|
||||
interface is missing is *absent*, not skipped: it waits, binds within
|
||||
a second of the interface coming up, unbinds if it goes away, and
|
||||
rebinds when it returns. Order does not matter, and neither
|
||||
`/etc/init.d/fips restart` nor any hotplug rule is part of this
|
||||
procedure.
|
||||
|
||||
Watch it happen:
|
||||
|
||||
```sh
|
||||
fipsctl show transports
|
||||
```
|
||||
|
||||
The transport's `interface` block reports `presence` (`absent` /
|
||||
`binding` / `present`), `policy` (`required` / `optional`) and how
|
||||
long it has held that state.
|
||||
|
||||
If you *changed a config value* above rather than only creating an
|
||||
interface, that does need a restart — configuration is read at
|
||||
startup, interfaces are not:
|
||||
|
||||
```sh
|
||||
/etc/init.d/fips restart
|
||||
```
|
||||
|
||||
Restart fips **after** the AP interface is up. A transport whose
|
||||
interface is missing at startup is logged and skipped, not retried —
|
||||
so if the daemon comes up before the radio, the access transport
|
||||
stays dead until the next restart. (An interface that *vanishes and
|
||||
returns* after startup is recovered automatically; only the missing-
|
||||
at-startup case needs this ordering.)
|
||||
|
||||
## Verify
|
||||
|
||||
L2 and addressing first, with a phone or laptop connected to `!FIPS`:
|
||||
|
||||
@@ -52,7 +52,7 @@ prints the response's `data` object as pretty JSON.
|
||||
| `show mmp` | `show_mmp` | MMP metrics summary: per-peer link-layer metrics and per-session session-layer metrics. |
|
||||
| `show cache` | `show_cache` | Coordinate cache: TTL, fill ratio, per-destination coords and path MTU. |
|
||||
| `show connections` | `show_connections` | Pending handshake connections: state, idle time, resend count. |
|
||||
| `show transports` | `show_transports` | Transport instances: type, state, MTU, local address, per-transport stats. |
|
||||
| `show transports` | `show_transports` | Transport instances: type, state, MTU, local address, per-transport stats, and — for interface-bound transports — interface presence (`absent` / `binding` / `present`), carrier, absence policy (`required` / `optional`), time in the current phase, and bind/failed-attempt counts. |
|
||||
| `show routing` | `show_routing` | Routing summary: pending lookups, retry state, forwarding/discovery/error/congestion counters. |
|
||||
| `show identity-cache` | `show_identity_cache` | Cached `(node_addr → npub)` entries with last-seen timestamps. |
|
||||
| `show native-flows` | `show_native_flows` | Native datagram API: open and pending flows with their ports, queue depth and age, bound listeners with their backlog, and the `native` counters. |
|
||||
|
||||
@@ -41,7 +41,7 @@ query on its first activation and on every refresh tick while active.
|
||||
| --- | ----- | ----- |
|
||||
| **Node** | `show_status` (+ `show_listening_sockets`) | Identity, version, uptime, peer/link/session counts, sparklines for mesh size, tree depth, peer count, bytes, loss. The Traffic block on this tab is split: TUN counters on the left, the **Listening on fips0** panel on the right (see below). |
|
||||
| **Peers** | `show_peers` (+ `show_links`, `show_transports` cross-refs) | Authenticated peers in a table. Selecting a row and pressing Enter opens a detail view. |
|
||||
| **Transports** | `show_transports` (+ `show_links`, `show_peers` cross-refs) | Tree of transport instances with per-link children when expanded. |
|
||||
| **Transports** | `show_transports` (+ `show_links`, `show_peers` cross-refs) | Tree of transport instances with per-link children when expanded. An interface-bound transport also carries an interface block; see [Interface block](#interface-block-transports-tab). |
|
||||
| **Sessions** | `show_sessions` | End-to-end FSP sessions. |
|
||||
| **Tree** | `show_tree` | Spanning-tree state and per-peer coordinates. |
|
||||
| **Filters** | `show_bloom` | Per-peer Bloom-filter state. |
|
||||
@@ -128,6 +128,26 @@ the keys the current context accepts.
|
||||
| `e` | Expand all transports. |
|
||||
| `c` | Collapse all transports. |
|
||||
|
||||
### Interface block (Transports tab)
|
||||
|
||||
A transport bound to a named interface carries an extra detail block.
|
||||
It exists because the observability data shipped as JSON before it
|
||||
reached this view, which left the live view reporting `up` for a
|
||||
transport bound to nothing. The original OpenWrt failure was expensive
|
||||
for that reason: the 802.11s link formed regardless, so nothing an
|
||||
operator could see said the node was deaf.
|
||||
|
||||
| Field | Meaning |
|
||||
| ----- | ------- |
|
||||
| Interface | The interface name the instance is bound to. |
|
||||
| Presence | Whether the interface is present, and for how long. |
|
||||
| Carrier | Whether a present interface has carrier. Present without carrier is a distinct state. |
|
||||
| Bound to | The address bound now, or nothing while absent. |
|
||||
| On absence | The policy: `required` degrades the node, `optional` does not. |
|
||||
| Binds, Failed binds | Counts over the instance's life, so a flapping interface reads as churn. |
|
||||
|
||||
**A transport with no interface has no block**, rather than an empty one.
|
||||
|
||||
### Multi-pane scrolling tabs (Tree, Filters, Routing)
|
||||
|
||||
Each lays out stacked panes that scroll independently.
|
||||
|
||||
@@ -681,6 +681,80 @@ root; on macOS it requires read/write access to a `/dev/bpf*` device.
|
||||
| `auto_connect` | bool | `false` | Auto-connect to discovered peers |
|
||||
| `accept_connections` | bool | `false` | Accept incoming connection attempts from discovered peers |
|
||||
| `beacon_interval_secs` | u64 | `30` | Announcement beacon interval in seconds (minimum 10) |
|
||||
| `optional` | bool | `false` | Whether absence of the interface is normal. See below |
|
||||
|
||||
**Dynamic binding.** The interface does not have to exist when the daemon
|
||||
starts. A transport whose interface is missing comes up *absent*: it is not a
|
||||
start failure, it is not skipped, and it binds on its own the moment the
|
||||
interface appears — sub-second where the kernel offers link events (netlink on
|
||||
Linux, `PF_ROUTE` on the BSDs), within a second otherwise. An interface that
|
||||
goes away at runtime unbinds and rebinds by the same path, so a `wifi reload`
|
||||
or an unplugged adapter needs no restart. Presence means `IFF_UP` — the
|
||||
interface exists and is administratively up — and deliberately not
|
||||
`IFF_RUNNING`: binding needs no carrier, and the socket keeps working across a
|
||||
carrier flap without rebinding. A bridge with nothing plugged into it, such as
|
||||
`br-lan` on a wifi-only router, is therefore bound and healthy rather than
|
||||
permanently `Degraded`, and starts carrying traffic the moment a port comes up.
|
||||
Whether an interface has carrier is reported separately, as `interface.carrier`
|
||||
in `show_transports`.
|
||||
|
||||
`optional` selects how that absence is reported:
|
||||
|
||||
| | absence | log | retries |
|
||||
| --- | --- | --- | --- |
|
||||
| `optional: false` (default) | node reports `Degraded` | `info` at boot / `warn` on a runtime detach, then `error` once if it lasts past 10 s | forever |
|
||||
| `optional: true` | no health impact | `info`, and nothing after | forever |
|
||||
|
||||
Naming an interface in configuration is a statement that you expect it, so the
|
||||
default is to complain; silence is opted into. Set `optional: true` for
|
||||
hardware that is legitimately not always there — a dock adapter, a radio only
|
||||
some boards carry.
|
||||
|
||||
`optional` describes **the interface's presence, not the transport's
|
||||
importance**. An optional interface that is present is used exactly as hard as
|
||||
any other. No value of it makes a missing interface fatal at startup: the only
|
||||
fatal case remains "no transports at all came up".
|
||||
|
||||
Either way the edge is logged once, on entering absence and on recovery —
|
||||
never once per retry.
|
||||
|
||||
The edge itself is not an error. An interface missing when the daemon starts
|
||||
and bound a moment later is the ordinary boot race this mechanism exists to
|
||||
absorb, so it is `info`; a runtime detach is `warn`, because a link coming and
|
||||
going is ordinary weather for a mesh daemon and a cable unplugged for two
|
||||
seconds does not need a human.
|
||||
|
||||
Ten seconds is the whole grace. Past that it is no longer a race against a
|
||||
radio or a container coming up, so a **required** interface still missing is
|
||||
reported once at `error` and stays `Degraded` until it returns. Start-time
|
||||
absence and a runtime detach share the one deadline — they are the same
|
||||
transition throughout this mechanism. An `optional` interface never reaches
|
||||
`error`; that is what `optional` means.
|
||||
|
||||
Said once, not repeated. How long the absence has lasted is a *state*, and it
|
||||
is published as one: `interface.since_secs` in `show_transports`, and
|
||||
`Degraded` for as long as it holds. Re-announcing it on a timer would put a
|
||||
second, lossier copy of that in the log.
|
||||
|
||||
Node health does not wait for the ten seconds — `Degraded` is published on the
|
||||
first edge, and that is the signal to watch.
|
||||
|
||||
If a binding keeps dying moments after it is established (a socket that errors
|
||||
persistently while the interface stays up), the binder stops treating each
|
||||
bind as a recovery: it backs off on the same 1 s → 30 s curve, holds node
|
||||
health at its degraded reading, and stays quiet until a binding survives ten
|
||||
seconds. Without that damping a broken socket produces a health flap and a log
|
||||
pair every second, which is the same cry-wolf failure the edge-only logging
|
||||
rule exists to prevent.
|
||||
|
||||
A bind failure that is **not** absence — no `CAP_NET_RAW`, no readable
|
||||
`/dev/bpf*`, a buffer the kernel refused — is a fault, not a state, and fails
|
||||
the daemon's start as it always has. Only a missing interface is waited out.
|
||||
|
||||
`fipsctl show transports` reports the current state per transport under
|
||||
`interface`: `presence` (`absent` / `binding` / `present`), `carrier`,
|
||||
`policy` (`required` / `optional`), `since_secs`, `binds`, and
|
||||
`failed_attempts`.
|
||||
|
||||
**Named instances.** Multiple Ethernet interfaces can be configured by
|
||||
using named sub-keys instead of flat parameters:
|
||||
|
||||
@@ -124,7 +124,7 @@ table below lists every command currently registered.
|
||||
| `show_mmp` | — | `peers[]` (link-layer per peer), `sessions[]` (session-layer per session). Each entry includes loss/RTT/ETX/goodput, smoothed values, trends. |
|
||||
| `show_cache` | — | `count`, `max_entries`, `fill_ratio`, `default_ttl_ms`, `expired`, `avg_age_ms`, `entries[]` — per-destination coords, depth, age, last-used, optional `path_mtu`. |
|
||||
| `show_connections` | — | `connections[]` — pending handshakes: `link_id`, `direction`, `handshake_state`, `started_at_ms`, `idle_ms`, `resend_count`, optional `expected_peer`. |
|
||||
| `show_transports` | — | `transports[]` — `transport_id`, `type`, `state`, `mtu`, `name`, `local_addr`, optional `tor_mode`, `onion_address`, `tor_monitoring`, `stats`. |
|
||||
| `show_transports` | — | `transports[]` — `transport_id`, `type`, `state`, `mtu`, `name`, `local_addr`, optional `tor_mode`, `onion_address`, `tor_monitoring`, `stats`, and `interface` for interface-bound transports. `interface` carries `name` (the configured netdev), `presence` (`absent` / `binding` / `present`), `carrier` (whether the link has `IFF_RUNNING` — reported only, never acted on: presence is `IFF_UP`, so a bound interface with no carrier is normal), `policy` (`required` / `optional`), `since_secs` (how long the current presence phase has been held), `binds` (successful binds since the transport was created — `1` after a clean start, more means it has rebound) and `failed_attempts` (failed binds since the last success). Absent entirely for transports that are not bound to a named interface, rather than reported as a permanently-`present` interface named `""`. Note that `state` describes the *transport* (`up` once started) and `interface.presence` describes the *socket*: an `up` transport whose interface is `absent` is started and waiting, which is a normal state and not a failure. |
|
||||
| `show_routing` | — | `coord_cache_entries`, `identity_cache_entries`, `pending_lookups[]`, `pending_tun_destinations`, `pending_tun_packets`, `recent_requests`, `retries[]`, `forwarding`, `discovery` (request/response sub-counters; includes `req_deduplicated` — requests suppressed as recent duplicates — and `req_dedup_cache_full` — requests admitted because the dedup cache was full), `error_signals`, `congestion`. |
|
||||
| `show_identity_cache` | — | `entries[]`, `count`, `max_entries`. Each entry: `node_addr`, `npub`, `display_name`, `ipv6_addr`, `last_seen_ms`, `age_ms`. |
|
||||
| `show_native_flows` | — | `flows[]`, `listeners[]`, `stats` (the `native` counter family). Each flow: `flow_id`, `peer` (the peer's npub, which is its address; always present, because the flow carries the key its client named or its session authenticated), `peer_addr` (the 16-byte node address in hex — a truncated hash of the same key, kept because it is what `show_sessions` and `show_routing` key on), `local_port`, `remote_port`, `state` (`established` / `pending_accept`), `queued` (datagrams the node is holding for the flow), `age_ms` (time since the flow reached its current state: opened for a flow this node opened, accepted for one taken off a listener, announced for one still pending — accepting a pending flow restarts the clock). Each listener: `local_port`, `backlog`. |
|
||||
|
||||
@@ -223,6 +223,7 @@ is "all four flags on both ends."
|
||||
> accept_connections: true
|
||||
> dongle:
|
||||
> interface: "enx00aabbccddee"
|
||||
> optional: true
|
||||
> announce: true
|
||||
> # ...
|
||||
> ```
|
||||
@@ -231,6 +232,16 @@ is "all four flags on both ends."
|
||||
> A single ground-up link only needs the flat form shown
|
||||
> first; named instances become useful when the same node
|
||||
> bridges multiple physical segments.
|
||||
>
|
||||
> `optional: true` on the dongle says its absence is normal —
|
||||
> a USB adapter that is plugged in some days and not others.
|
||||
> Without it, naming an interface is a statement that you
|
||||
> expect it, and while it is missing the node reports
|
||||
> `Degraded` and logs at `error`. Either way the interface
|
||||
> does not have to exist when the daemon starts: a transport
|
||||
> whose interface is missing waits and binds when it appears,
|
||||
> and rebinds if it later goes away. Watch that with
|
||||
> `fipsctl show transports`.
|
||||
|
||||
## Step 3: Grant the daemon permission to open raw sockets
|
||||
|
||||
|
||||
@@ -101,6 +101,9 @@ transports:
|
||||
accept_connections: true
|
||||
wwan:
|
||||
interface: "phy0-sta0"
|
||||
# Only exists while a radio is in station mode. Absence is normal, so
|
||||
# it must not report the router Degraded.
|
||||
optional: true
|
||||
listen: true
|
||||
announce: true
|
||||
auto_connect: true
|
||||
@@ -113,54 +116,55 @@ transports:
|
||||
accept_connections: true
|
||||
|
||||
# 802.11s mesh backhaul between FIPS routers. These entries ship
|
||||
# commented out so a stock install that never creates fips-mesh*
|
||||
# logs no per-boot "interface missing" bind warning. Running
|
||||
# 'fips-mesh-setup <radio>' creates the interface AND uncomments the
|
||||
# matching block here (once per radio; radio0 -> fips-mesh0, radio1 ->
|
||||
# fips-mesh1); 'fips-mesh-setup remove' re-comments it. Restart fips
|
||||
# after — a transport whose interface is missing at startup is skipped,
|
||||
# not retried. Dual-band routers can mesh on both bands at once —
|
||||
# failover, not multipath: FIPS keeps one active link per peer, the
|
||||
# other band stands by. The mesh runs OPEN (no SAE) with 802.11s
|
||||
# forwarding off: FIPS's Noise handshake is the encryption and
|
||||
# authentication, and FIPS is the routing layer. See
|
||||
# docs/how-to/set-up-80211s-mesh-backhaul.md.
|
||||
# mesh0:
|
||||
# interface: "fips-mesh0"
|
||||
# listen: true
|
||||
# announce: true
|
||||
# auto_connect: true
|
||||
# accept_connections: true
|
||||
# mesh1:
|
||||
# interface: "fips-mesh1"
|
||||
# listen: true
|
||||
# announce: true
|
||||
# auto_connect: true
|
||||
# accept_connections: true
|
||||
# ENABLED and marked 'optional: true'. The daemon treats a named
|
||||
# interface that is not there as absent rather than as a failure: it
|
||||
# waits, binds the moment 'fips-mesh-setup <radio>' creates the
|
||||
# interface, and unbinds again if it goes away — no config edit and no
|
||||
# restart, and 'optional: true' keeps a stock install that never runs
|
||||
# that script quiet and un-Degraded.
|
||||
#
|
||||
# Dual-band routers can mesh on both bands at once — failover, not
|
||||
# multipath: FIPS keeps one active link per peer, the other band stands
|
||||
# by. The mesh runs OPEN (no SAE) with 802.11s forwarding off: FIPS's
|
||||
# Noise handshake is the encryption and authentication, and FIPS is the
|
||||
# routing layer. See docs/how-to/set-up-80211s-mesh-backhaul.md.
|
||||
mesh0:
|
||||
interface: "fips-mesh0"
|
||||
optional: true
|
||||
listen: true
|
||||
announce: true
|
||||
auto_connect: true
|
||||
accept_connections: true
|
||||
mesh1:
|
||||
interface: "fips-mesh1"
|
||||
optional: true
|
||||
listen: true
|
||||
announce: true
|
||||
auto_connect: true
|
||||
accept_connections: true
|
||||
|
||||
# Open "!FIPS" access SSID for phones and laptops running FIPS. These
|
||||
# entries ship commented out so a stock install that never creates
|
||||
# fips-ap* logs no per-boot "interface missing" bind warning. Running
|
||||
# 'fips-ap-setup <radio>' creates the interface AND uncomments the
|
||||
# matching block here (once per radio; radio0 -> fips-ap0, radio1 ->
|
||||
# fips-ap1); 'fips-ap-setup remove' re-comments it. Restart fips after
|
||||
# — a transport whose interface is missing at startup is skipped, not
|
||||
# retried. The SSID is OPEN and isolated on purpose: FIPS's Noise
|
||||
# handshake is the only security layer, and associated clients reach
|
||||
# nothing but the FIPS handshake surface. See
|
||||
# docs/how-to/set-up-open-access-ssid.md.
|
||||
# ap0:
|
||||
# interface: "fips-ap0"
|
||||
# listen: true
|
||||
# announce: true
|
||||
# auto_connect: true
|
||||
# accept_connections: true
|
||||
# ap1:
|
||||
# interface: "fips-ap1"
|
||||
# listen: true
|
||||
# announce: true
|
||||
# auto_connect: true
|
||||
# accept_connections: true
|
||||
# Open "!FIPS" access SSID for phones and laptops running FIPS. Ships
|
||||
# ENABLED and marked 'optional: true', for the same reason as the mesh
|
||||
# blocks above: 'fips-ap-setup <radio>' creates the interface and the
|
||||
# daemon binds it when it appears, with no config edit and no restart.
|
||||
#
|
||||
# The SSID is OPEN and isolated on purpose: FIPS's Noise handshake is the
|
||||
# only security layer, and associated clients reach nothing but the FIPS
|
||||
# handshake surface. See docs/how-to/set-up-open-access-ssid.md.
|
||||
ap0:
|
||||
interface: "fips-ap0"
|
||||
optional: true
|
||||
listen: true
|
||||
announce: true
|
||||
auto_connect: true
|
||||
accept_connections: true
|
||||
ap1:
|
||||
interface: "fips-ap1"
|
||||
optional: true
|
||||
listen: true
|
||||
announce: true
|
||||
auto_connect: true
|
||||
accept_connections: true
|
||||
|
||||
# Bluetooth Low Energy transport — requires BlueZ and the 'ble' feature.
|
||||
# ble:
|
||||
|
||||
@@ -43,14 +43,17 @@
|
||||
# access channels freely.
|
||||
#
|
||||
# The shipped /etc/fips/fips.yaml carries 'ap0' and 'ap1' entries under
|
||||
# 'transports.ethernet' bound to these names, but commented out — a stock
|
||||
# install that never creates fips-ap* then logs no bind warning. This
|
||||
# helper uncomments the matching entry when it creates an interface and
|
||||
# re-comments it on remove, so the daemon binds the transport without a
|
||||
# manual config edit. It also uncomments the node.rendezvous.lan block
|
||||
# (mDNS/DNS-SD — how phone FIPS apps discover the daemon); that switch is
|
||||
# daemon-wide and stays on at remove. After an interface is up, restart
|
||||
# fips.
|
||||
# 'transports.ethernet' bound to these names, enabled and marked
|
||||
# 'optional: true'. The daemon treats a named interface that is not there as
|
||||
# ABSENT rather than as a start failure: it waits, binds the moment the
|
||||
# interface appears, and unbinds again when it goes away — so creating the
|
||||
# interface is all this helper has to do for the transport, with no config
|
||||
# rewrite and no daemon restart.
|
||||
#
|
||||
# It does still edit one thing: the node.rendezvous.lan block (mDNS/DNS-SD —
|
||||
# how phone FIPS apps discover the daemon). That is a config value rather than
|
||||
# an interface, so it does need a restart to take effect. The switch is
|
||||
# daemon-wide and stays on at remove.
|
||||
# See docs/how-to/set-up-open-access-ssid.md for the full guide.
|
||||
|
||||
DEFAULT_SSID="!FIPS"
|
||||
@@ -63,19 +66,14 @@ ap_config_write() {
|
||||
chmod 600 "$CONFIG.tmp" && mv "$CONFIG.tmp" "$CONFIG"
|
||||
}
|
||||
|
||||
# Uncomment the 'ap<idx>' transports.ethernet block in $CONFIG (created by
|
||||
# 'fips-ap-setup'). Reversible with ap_config_disable. Returns:
|
||||
# 0 enabled (or already active) 1 no config file 2 no such block
|
||||
ap_config_enable() {
|
||||
# Whether $CONFIG carries an enabled 'ap<idx>' transports.ethernet block.
|
||||
# Reports only; the daemon owns the binding. Returns:
|
||||
# 0 present 1 no config file 2 no such block
|
||||
ap_config_present() {
|
||||
idx="$1"
|
||||
[ -f "$CONFIG" ] || return 1
|
||||
grep -q "^ ap$idx:" "$CONFIG" && return 0
|
||||
grep -q "^ # ap$idx:" "$CONFIG" || return 2
|
||||
awk -v idx="$idx" '
|
||||
$0 ~ ("^ # ap" idx ":[ \t]*$") { blk = 1; sub(/^ # /, " "); print; next }
|
||||
blk && /^ # / { sub(/^ # /, " "); print; next }
|
||||
{ blk = 0; print }
|
||||
' "$CONFIG" > "$CONFIG.tmp" && ap_config_write
|
||||
grep -q "^ ap$idx:" "$CONFIG" || return 2
|
||||
return 0
|
||||
}
|
||||
|
||||
# Uncomment the 'lan' block under node.rendezvous in $CONFIG — the daemon's
|
||||
@@ -107,19 +105,6 @@ lan_rendezvous_enable() {
|
||||
' "$CONFIG" > "$CONFIG.tmp" && ap_config_write
|
||||
}
|
||||
|
||||
# Re-comment the 'ap<idx>' block so the daemon stops binding it (and stops
|
||||
# warning about the now-missing interface). Inverse of ap_config_enable.
|
||||
ap_config_disable() {
|
||||
idx="$1"
|
||||
[ -f "$CONFIG" ] || return 1
|
||||
grep -q "^ ap$idx:" "$CONFIG" || return 0
|
||||
awk -v idx="$idx" '
|
||||
$0 ~ ("^ ap" idx ":[ \t]*$") { blk = 1; sub(/^ /, " # "); print; next }
|
||||
blk && /^ / { sub(/^ /, " # "); print; next }
|
||||
{ blk = 0; print }
|
||||
' "$CONFIG" > "$CONFIG.tmp" && ap_config_write
|
||||
}
|
||||
|
||||
usage() {
|
||||
echo "Usage: fips-ap-setup <radio> [ssid]" >&2
|
||||
echo " fips-ap-setup remove [radio]" >&2
|
||||
@@ -154,10 +139,9 @@ if [ "$1" = "remove" ]; then
|
||||
uci -q delete "network.$section"
|
||||
uci -q delete "dhcp.$section"
|
||||
uci -q del_list "firewall.fips_ap.network=$section"
|
||||
# Re-comment the matching ap<N> transport in fips.yaml so the
|
||||
# daemon stops warning about the interface we just removed.
|
||||
idx="$(printf '%s' "$ifname" | sed -n 's/.*[^0-9]\([0-9]\{1,\}\)$/\1/p')"
|
||||
[ -n "$idx" ] && ap_config_disable "$idx"
|
||||
# The fips.yaml block stays as it is. The daemon notices the
|
||||
# interface going away, unbinds, and waits for it — silently,
|
||||
# because the block is marked 'optional: true'.
|
||||
echo "Removed ${ifname:-$section}."
|
||||
done
|
||||
# Drop the shared zone and its rules once the last instance is gone.
|
||||
@@ -356,12 +340,12 @@ wifi reload
|
||||
/etc/init.d/odhcpd reload
|
||||
/etc/init.d/firewall reload
|
||||
|
||||
# Enable the matching ap<N> transport in the shipped fips.yaml (it ships
|
||||
# commented out). Tailor the restart hint to what we could do.
|
||||
ap_config_enable "$IDX"
|
||||
# Report whether the shipped fips.yaml still carries the matching transport.
|
||||
# It ships enabled, so this is a check, not an edit.
|
||||
ap_config_present "$IDX"
|
||||
case $? in
|
||||
0) TRANSPORT_NOTE="The ap$IDX transport in $CONFIG that binds '$AP_IFNAME' is
|
||||
now uncommented and enabled." ;;
|
||||
0) TRANSPORT_NOTE="The ap$IDX transport in $CONFIG binds '$AP_IFNAME'; the
|
||||
daemon picks the interface up on its own, with no restart." ;;
|
||||
1) TRANSPORT_NOTE="No $CONFIG found — add a transports.ethernet entry binding
|
||||
interface '$AP_IFNAME' by hand." ;;
|
||||
*) TRANSPORT_NOTE="No 'ap$IDX' entry in $CONFIG — add a transports.ethernet
|
||||
@@ -398,9 +382,9 @@ per SSID, so it covers every FIPS router.
|
||||
|
||||
Next steps:
|
||||
1. $TRANSPORT_NOTE
|
||||
Watch it bind with: fipsctl show transports
|
||||
2. $MDNS_NOTE
|
||||
Restart the daemon AFTER the interface is up — a transport whose
|
||||
interface is missing at startup is skipped, not retried:
|
||||
mDNS is a config value, not an interface, so it needs a restart:
|
||||
/etc/init.d/fips restart
|
||||
3. Associate a phone or laptop running FIPS and verify:
|
||||
iw dev $AP_IFNAME station dump
|
||||
|
||||
@@ -27,49 +27,26 @@
|
||||
# Ethernet transport binds each directly and runs discovery beacons over it.
|
||||
#
|
||||
# The shipped /etc/fips/fips.yaml carries 'mesh0' and 'mesh1' entries under
|
||||
# 'transports.ethernet' bound to these names, but commented out — a stock
|
||||
# install that never creates fips-mesh* then logs no bind warning. This
|
||||
# helper uncomments the matching entry when it creates an interface and
|
||||
# re-comments it on remove, so the daemon binds the transport without a
|
||||
# manual config edit. After an interface is up, restart fips.
|
||||
# 'transports.ethernet' bound to these names, enabled and marked
|
||||
# 'optional: true'. The daemon treats a named interface that is not there as
|
||||
# ABSENT rather than as a start failure: it waits, binds the moment the
|
||||
# interface appears, and unbinds again when it goes away. So this helper
|
||||
# creates the interface and nothing else — no config rewrite, and no daemon
|
||||
# restart. 'optional: true' is what keeps a stock install that never runs this
|
||||
# script from reporting Degraded for an interface it was never going to have.
|
||||
# See docs/how-to/set-up-80211s-mesh-backhaul.md for the full guide.
|
||||
|
||||
DEFAULT_MESH_ID="fips-mesh"
|
||||
CONFIG="/etc/fips/fips.yaml"
|
||||
|
||||
# Replace $CONFIG with the rewritten $CONFIG.tmp. Force mode 0600 first: the
|
||||
# package installs fips.yaml 0600 (it may hold an inline 'nsec' private key),
|
||||
# and a fresh tmp file would otherwise land world-readable after the move.
|
||||
mesh_config_write() {
|
||||
chmod 600 "$CONFIG.tmp" && mv "$CONFIG.tmp" "$CONFIG"
|
||||
}
|
||||
|
||||
# Uncomment the 'mesh<idx>' transports.ethernet block in $CONFIG (created by
|
||||
# 'fips-mesh-setup'). Reversible with mesh_config_disable. Returns:
|
||||
# 0 enabled (or already active) 1 no config file 2 no such block
|
||||
mesh_config_enable() {
|
||||
# Whether $CONFIG carries an enabled 'mesh<idx>' transports.ethernet block.
|
||||
# Reports only; the daemon owns the binding. Returns:
|
||||
# 0 present 1 no config file 2 no such block
|
||||
mesh_config_present() {
|
||||
idx="$1"
|
||||
[ -f "$CONFIG" ] || return 1
|
||||
grep -q "^ mesh$idx:" "$CONFIG" && return 0
|
||||
grep -q "^ # mesh$idx:" "$CONFIG" || return 2
|
||||
awk -v idx="$idx" '
|
||||
$0 ~ ("^ # mesh" idx ":[ \t]*$") { blk = 1; sub(/^ # /, " "); print; next }
|
||||
blk && /^ # / { sub(/^ # /, " "); print; next }
|
||||
{ blk = 0; print }
|
||||
' "$CONFIG" > "$CONFIG.tmp" && mesh_config_write
|
||||
}
|
||||
|
||||
# Re-comment the 'mesh<idx>' block so the daemon stops binding it (and stops
|
||||
# warning about the now-missing interface). Inverse of mesh_config_enable.
|
||||
mesh_config_disable() {
|
||||
idx="$1"
|
||||
[ -f "$CONFIG" ] || return 1
|
||||
grep -q "^ mesh$idx:" "$CONFIG" || return 0
|
||||
awk -v idx="$idx" '
|
||||
$0 ~ ("^ mesh" idx ":[ \t]*$") { blk = 1; sub(/^ /, " # "); print; next }
|
||||
blk && /^ / { sub(/^ /, " # "); print; next }
|
||||
{ blk = 0; print }
|
||||
' "$CONFIG" > "$CONFIG.tmp" && mesh_config_write
|
||||
grep -q "^ mesh$idx:" "$CONFIG" || return 2
|
||||
return 0
|
||||
}
|
||||
|
||||
usage() {
|
||||
@@ -103,10 +80,9 @@ if [ "$1" = "remove" ]; then
|
||||
ifname="$(uci -q get "wireless.$section.ifname")"
|
||||
uci -q delete "wireless.$section"
|
||||
uci -q delete "network.$section"
|
||||
# Re-comment the matching mesh<N> transport in fips.yaml so the
|
||||
# daemon stops warning about the interface we just removed.
|
||||
idx="$(printf '%s' "$ifname" | sed -n 's/.*[^0-9]\([0-9]\{1,\}\)$/\1/p')"
|
||||
[ -n "$idx" ] && mesh_config_disable "$idx"
|
||||
# The fips.yaml block stays as it is. The daemon notices the
|
||||
# interface going away, unbinds, and waits for it — silently,
|
||||
# because the block is marked 'optional: true'.
|
||||
echo "Removed ${ifname:-$section}."
|
||||
done
|
||||
uci commit wireless
|
||||
@@ -114,7 +90,7 @@ if [ "$1" = "remove" ]; then
|
||||
# 'wifi reload' re-applies the whole wireless config, so it briefly drops
|
||||
# every client AP on all radios (a few seconds) — expected on remove.
|
||||
wifi reload
|
||||
echo "Restart fips: /etc/init.d/fips restart"
|
||||
echo "No fips restart needed — the daemon unbinds the interface itself."
|
||||
exit 0
|
||||
fi
|
||||
|
||||
@@ -226,12 +202,12 @@ uci commit network
|
||||
# every client AP on all radios (a few seconds) — expected when adding a mesh.
|
||||
wifi reload
|
||||
|
||||
# Enable the matching mesh<N> transport in the shipped fips.yaml (it ships
|
||||
# commented out). Tailor the restart hint to what we could do.
|
||||
mesh_config_enable "$IDX"
|
||||
# Report whether the shipped fips.yaml still carries the matching transport.
|
||||
# It ships enabled, so this is a check, not an edit.
|
||||
mesh_config_present "$IDX"
|
||||
case $? in
|
||||
0) TRANSPORT_NOTE="The mesh$IDX transport in $CONFIG that binds '$MESH_IFNAME' is
|
||||
now uncommented and enabled." ;;
|
||||
0) TRANSPORT_NOTE="The mesh$IDX transport in $CONFIG binds '$MESH_IFNAME'; the
|
||||
daemon picks the interface up on its own." ;;
|
||||
1) TRANSPORT_NOTE="No $CONFIG found — add a transports.ethernet entry binding
|
||||
interface '$MESH_IFNAME' by hand." ;;
|
||||
*) TRANSPORT_NOTE="No 'mesh$IDX' entry in $CONFIG — add a transports.ethernet
|
||||
@@ -248,9 +224,9 @@ radio too — second band is a standby path (failover, not multipath).
|
||||
|
||||
Next steps:
|
||||
1. $TRANSPORT_NOTE
|
||||
Restart the daemon AFTER the interface is up — a transport whose
|
||||
interface is missing at startup is skipped, not retried:
|
||||
/etc/init.d/fips restart
|
||||
No restart: the daemon binds an interface when it appears and
|
||||
rebinds it if it goes away. Watch it happen with:
|
||||
fipsctl show transports
|
||||
2. Verify L2 peering with a second FIPS router in range:
|
||||
iw dev $MESH_IFNAME station dump
|
||||
and the FIPS link on top of it:
|
||||
|
||||
@@ -1313,3 +1313,377 @@ fn help_overlay_lists_keys() {
|
||||
assert!(testkit::contains_row(&buf, "quit"));
|
||||
assert!(testkit::contains_row(&buf, "Press ? or Esc to close"));
|
||||
}
|
||||
|
||||
/// An interface-bound transport names its netdev in the table, so an operator
|
||||
/// reading the list sees `br-lan` rather than an instance label that means
|
||||
/// nothing outside the config file.
|
||||
#[test]
|
||||
fn transports_row_names_the_interface() {
|
||||
let data = json!({
|
||||
"transports": [{
|
||||
"transport_id": 1,
|
||||
"type": "ethernet",
|
||||
"state": "up",
|
||||
"mtu": 1497,
|
||||
"name": "lan",
|
||||
"interface": {
|
||||
"name": "br-lan",
|
||||
"presence": "present",
|
||||
"carrier": true,
|
||||
"policy": "required",
|
||||
"since_secs": 3600,
|
||||
"binds": 1,
|
||||
"failed_attempts": 0
|
||||
},
|
||||
"stats": {}
|
||||
}]
|
||||
});
|
||||
let mut app = app_with(Tab::Transports, data);
|
||||
let buf = testkit::render(80, 20, |frame, area| {
|
||||
super::transports::draw(frame, &mut app, area);
|
||||
});
|
||||
|
||||
assert!(testkit::contains_row(&buf, "br-lan"));
|
||||
// A bound interface reads as the transport state; only the exceptional
|
||||
// case displaces it, or the annotation stops being read.
|
||||
assert!(testkit::contains_row(&buf, "up"));
|
||||
assert!(!testkit::contains_row(&buf, "absent"));
|
||||
}
|
||||
|
||||
/// The narrow layout keeps the columns an operator is scanning for.
|
||||
///
|
||||
/// 80x24 is the OpenWrt serial console and the xterm/tmux default. The full
|
||||
/// layout's fixed columns sum to 96 against ~77 usable, and ratatui resolves
|
||||
/// an over-subscribed layout by shrinking *every* column — so the overflow
|
||||
/// does not clip the rightmost one, it clips all of them, and
|
||||
/// `mesh0 (optional)` became `mesh0 (optio`. The marker is the one thing on
|
||||
/// that row worth reading.
|
||||
///
|
||||
/// This is the gap that let the regression through: the test asserting the
|
||||
/// marker rendered at width 110, and the one rendering at 80 asserted only the
|
||||
/// netdev name.
|
||||
#[test]
|
||||
fn the_optional_marker_survives_an_80_column_terminal() {
|
||||
let data = json!({
|
||||
"transports": [{
|
||||
"transport_id": 2,
|
||||
"type": "ethernet",
|
||||
"state": "up",
|
||||
"mtu": 1499,
|
||||
"name": "mesh0",
|
||||
"interface": {
|
||||
"name": "fips-mesh0",
|
||||
"presence": "absent",
|
||||
"carrier": false,
|
||||
"policy": "optional",
|
||||
"since_secs": 252,
|
||||
"binds": 0,
|
||||
"failed_attempts": 0
|
||||
},
|
||||
"stats": {}
|
||||
}]
|
||||
});
|
||||
let mut app = app_with(Tab::Transports, data);
|
||||
let buf = testkit::render(80, 20, |frame, area| {
|
||||
super::transports::draw(frame, &mut app, area);
|
||||
});
|
||||
|
||||
assert!(
|
||||
testkit::contains_row(&buf, "mesh0 (optional)"),
|
||||
"the policy marker must not be the thing that clips at 80 columns"
|
||||
);
|
||||
assert!(testkit::contains_row(&buf, "fips-mesh0"));
|
||||
assert!(testkit::contains_row(&buf, "absent"));
|
||||
}
|
||||
|
||||
/// An absent interface displaces the transport state in the State column, and
|
||||
/// its severity follows the absence policy.
|
||||
///
|
||||
/// `state` reads `up` from the moment the transport starts, whether or not it
|
||||
/// is bound to anything, so the row would otherwise look identical to a
|
||||
/// working one — the "looks healthy, reaches nothing" failure dynamic binding
|
||||
/// exists to make visible.
|
||||
///
|
||||
/// Optional is a warning: a dock adapter that is not plugged in, or a radio
|
||||
/// this board never had, is the case `optional: true` was added to describe,
|
||||
/// and the daemon stays `Full` for it. Red there would train the operator to
|
||||
/// ignore red.
|
||||
#[test]
|
||||
fn an_absent_optional_interface_warns() {
|
||||
let data = json!({
|
||||
"transports": [{
|
||||
"transport_id": 2,
|
||||
"type": "ethernet",
|
||||
"state": "up",
|
||||
"mtu": 1499,
|
||||
"name": "mesh0",
|
||||
"interface": {
|
||||
"name": "fips-mesh0",
|
||||
"presence": "absent",
|
||||
"carrier": false,
|
||||
"policy": "optional",
|
||||
"since_secs": 252,
|
||||
"binds": 0,
|
||||
"failed_attempts": 0
|
||||
},
|
||||
"stats": {}
|
||||
}]
|
||||
});
|
||||
let mut app = app_with(Tab::Transports, data);
|
||||
let buf = testkit::render(110, 20, |frame, area| {
|
||||
super::transports::draw(frame, &mut app, area);
|
||||
});
|
||||
|
||||
assert!(testkit::contains_row(&buf, "fips-mesh0"));
|
||||
// Policy rides with the instance name, and only the exception prints.
|
||||
assert!(testkit::contains_row(&buf, "mesh0 (optional)"));
|
||||
// The State column carries presence, because `up` is what it would
|
||||
// otherwise say about an interface that has never existed.
|
||||
assert!(testkit::contains_row(&buf, "absent"));
|
||||
assert_eq!(
|
||||
testkit::fg_at(&buf, "fips-mesh0"),
|
||||
Some(ratatui::style::Color::Yellow),
|
||||
"an interface whose absence is normal is a warning, not an error"
|
||||
);
|
||||
}
|
||||
|
||||
/// A required interface being absent is an error.
|
||||
///
|
||||
/// Naming an interface without `optional: true` is a statement that you expect
|
||||
/// it, and the daemon reports `Degraded` while it is missing. The view has to
|
||||
/// agree, or the colour stops carrying the same meaning as the health state.
|
||||
#[test]
|
||||
fn an_absent_required_interface_errors() {
|
||||
let data = json!({
|
||||
"transports": [{
|
||||
"transport_id": 2,
|
||||
"type": "ethernet",
|
||||
"state": "up",
|
||||
"mtu": 1499,
|
||||
"name": "wan",
|
||||
"interface": {
|
||||
"name": "eth0",
|
||||
"presence": "absent",
|
||||
"carrier": false,
|
||||
"policy": "required",
|
||||
"since_secs": 30,
|
||||
"binds": 0,
|
||||
"failed_attempts": 0
|
||||
},
|
||||
"stats": {}
|
||||
}]
|
||||
});
|
||||
let mut app = app_with(Tab::Transports, data);
|
||||
let buf = testkit::render(110, 20, |frame, area| {
|
||||
super::transports::draw(frame, &mut app, area);
|
||||
});
|
||||
|
||||
// `required` is the default and is spelled by omission — a column of it
|
||||
// on nearly every row would be a column of noise.
|
||||
assert!(!testkit::contains_row(&buf, "required"));
|
||||
assert!(!testkit::contains_row(&buf, "(optional)"));
|
||||
assert!(testkit::contains_row(&buf, "wan"));
|
||||
assert_eq!(
|
||||
testkit::fg_at(&buf, "eth0"),
|
||||
Some(ratatui::style::Color::Red),
|
||||
"an interface the config says to expect is an error while it is gone"
|
||||
);
|
||||
}
|
||||
|
||||
/// The detail pane reports presence, carrier, policy and bind counts —
|
||||
/// everything `state` cannot say.
|
||||
#[test]
|
||||
fn transport_detail_reports_interface_presence() {
|
||||
let data = json!({
|
||||
"transports": [{
|
||||
"transport_id": 3,
|
||||
"type": "ethernet",
|
||||
"state": "up",
|
||||
"mtu": 1497,
|
||||
"name": "wan",
|
||||
"interface": {
|
||||
"name": "eth0",
|
||||
"presence": "present",
|
||||
"carrier": false,
|
||||
"policy": "required",
|
||||
"since_secs": 90,
|
||||
"binds": 3,
|
||||
"failed_attempts": 2
|
||||
},
|
||||
"stats": {}
|
||||
}]
|
||||
});
|
||||
let mut app = app_with(Tab::Transports, data);
|
||||
app.detail_view = Some(crate::app::DetailView { scroll: 0 });
|
||||
let buf = testkit::render(120, 30, |frame, area| {
|
||||
super::transports::draw(frame, &mut app, area);
|
||||
});
|
||||
|
||||
assert!(testkit::contains_row(&buf, "Interface"));
|
||||
assert!(testkit::contains_row(&buf, "eth0"));
|
||||
assert!(testkit::contains_row(&buf, "present"));
|
||||
// Carrier is reported rather than acted on, so "bound with no carrier" —
|
||||
// a bridge with nothing plugged into it — has to be legible as its own
|
||||
// state rather than inferred from silence.
|
||||
assert!(testkit::contains_row(&buf, "Carrier"));
|
||||
// A rebind count and a failed-bind count distinguish an interface that is
|
||||
// flapping from one that is there and refusing.
|
||||
assert!(testkit::contains_row(&buf, "rebound"));
|
||||
assert!(testkit::contains_row(&buf, "Failed binds"));
|
||||
}
|
||||
|
||||
/// Instance names and interfaces line up down the list.
|
||||
///
|
||||
/// Packed into one label they did not: `ethernet dongle2 en25` over
|
||||
/// `ethernet wifi en0` left the netdev names ragged, which is the column an
|
||||
/// operator scans to find the interface they are looking for. Separate facts,
|
||||
/// separate columns.
|
||||
#[test]
|
||||
fn transports_columns_align_across_mixed_types() {
|
||||
let data = json!({
|
||||
"transports": [
|
||||
{
|
||||
"transport_id": 1, "type": "udp", "state": "up", "mtu": 1472,
|
||||
"local_addr": "0.0.0.0:2121", "stats": {}
|
||||
},
|
||||
{
|
||||
"transport_id": 2, "type": "tcp", "state": "up", "mtu": 1400,
|
||||
"local_addr": "0.0.0.0:8443", "stats": {}
|
||||
},
|
||||
{
|
||||
"transport_id": 3, "type": "ethernet", "state": "up", "mtu": 1497,
|
||||
"name": "dongle2",
|
||||
"interface": {
|
||||
"name": "en25", "presence": "absent", "carrier": false,
|
||||
"policy": "required", "since_secs": 12, "binds": 0,
|
||||
"failed_attempts": 0
|
||||
},
|
||||
"stats": {}
|
||||
},
|
||||
{
|
||||
"transport_id": 4, "type": "ethernet", "state": "up", "mtu": 1497,
|
||||
"name": "wifi",
|
||||
"interface": {
|
||||
"name": "en0", "presence": "present", "carrier": true,
|
||||
"policy": "optional", "since_secs": 900, "binds": 1,
|
||||
"failed_attempts": 0
|
||||
},
|
||||
"stats": {}
|
||||
}
|
||||
]
|
||||
});
|
||||
let mut app = app_with(Tab::Transports, data);
|
||||
let buf = testkit::render(100, 20, |frame, area| {
|
||||
super::transports::draw(frame, &mut app, area);
|
||||
});
|
||||
|
||||
let col_of = |needle: &str| testkit::find(&buf, needle).map(|(x, _)| x);
|
||||
|
||||
// The instance column starts at one x for every row that has one.
|
||||
assert_eq!(col_of("dongle2"), col_of("wifi"));
|
||||
// ... and so does the interface column, across long and short names and
|
||||
// across transports that name a netdev and ones that name a socket.
|
||||
assert_eq!(col_of("en25"), col_of("en0"));
|
||||
assert_eq!(col_of("en25"), col_of("0.0.0.0:2121"));
|
||||
assert_eq!(col_of("0.0.0.0:2121"), col_of("0.0.0.0:8443"));
|
||||
|
||||
// The header sits over the column it names.
|
||||
assert_eq!(col_of("Bound to"), col_of("en25"));
|
||||
assert_eq!(col_of("Instance"), col_of("dongle2"));
|
||||
}
|
||||
|
||||
/// A link's remote address shares the `Bound to` column with its parent's
|
||||
/// interface.
|
||||
///
|
||||
/// Same question — what is this attached to — so the same column: a netdev for
|
||||
/// the transport, a remote endpoint for the link. Keeping a full MAC out of
|
||||
/// the first column is also what lets the three identifying columns sit
|
||||
/// against the left edge rather than being pushed right by the widest link
|
||||
/// row.
|
||||
#[test]
|
||||
fn a_link_puts_its_remote_address_in_the_bound_to_column() {
|
||||
let transports = json!({
|
||||
"transports": [{
|
||||
"transport_id": 3, "type": "ethernet", "state": "up", "mtu": 1497,
|
||||
"name": "wifi",
|
||||
"interface": {
|
||||
"name": "en0", "presence": "present", "carrier": true,
|
||||
"policy": "optional", "since_secs": 60, "binds": 1,
|
||||
"failed_attempts": 0
|
||||
},
|
||||
"stats": {}
|
||||
}]
|
||||
});
|
||||
let links = json!({
|
||||
"links": [{
|
||||
"link_id": 1, "transport_id": 3, "direction": "Outbound",
|
||||
"remote_addr": "aa:bb:cc:dd:ee:ff", "state": "connected"
|
||||
}]
|
||||
});
|
||||
|
||||
let mut app = app_with(Tab::Transports, transports);
|
||||
app.data.insert(Tab::Links, links);
|
||||
app.expanded_transports.insert(3);
|
||||
|
||||
let buf = testkit::render(110, 20, |frame, area| {
|
||||
super::transports::draw(frame, &mut app, area);
|
||||
});
|
||||
|
||||
let col_of = |needle: &str| testkit::find(&buf, needle).map(|(x, _)| x);
|
||||
|
||||
// A right-aligned State clips from the left on overflow, so `connected`
|
||||
// reading as `onnected` is a width bug rather than an honest truncation.
|
||||
assert!(testkit::contains_row(&buf, "connected"));
|
||||
|
||||
// The MAC is not truncated and sits under the same header as the netdev.
|
||||
assert!(testkit::contains_row(&buf, "aa:bb:cc:dd:ee:ff"));
|
||||
assert_eq!(col_of("aa:bb:cc:dd:ee:ff"), col_of("en0"));
|
||||
assert_eq!(col_of("Bound to"), col_of("en0"));
|
||||
|
||||
// The link keeps its direction and tree glyph in the first column, which
|
||||
// is therefore narrow enough to leave the identifying columns at the left.
|
||||
assert!(testkit::contains_row(&buf, "Out"));
|
||||
// Packed at the left rather than floating in the middle. With the first
|
||||
// column flexible, as it was, it absorbed every spare column of a wide
|
||||
// terminal and pushed these two past the halfway mark.
|
||||
assert!(col_of("Instance").unwrap() <= 20);
|
||||
assert!(
|
||||
col_of("Bound to").unwrap() < 45,
|
||||
"the identifying columns must stay against the left edge"
|
||||
);
|
||||
}
|
||||
|
||||
/// `State` is right-aligned, so the values share a right edge and the column
|
||||
/// reads as a status strip rather than as ragged text.
|
||||
#[test]
|
||||
fn the_state_column_is_right_aligned() {
|
||||
let data = json!({
|
||||
"transports": [
|
||||
{
|
||||
"transport_id": 1, "type": "udp", "state": "up", "mtu": 1472,
|
||||
"local_addr": "0.0.0.0:2121", "stats": {}
|
||||
},
|
||||
{
|
||||
"transport_id": 2, "type": "ethernet", "state": "up", "mtu": 1499,
|
||||
"name": "mesh0",
|
||||
"interface": {
|
||||
"name": "fips-mesh0", "presence": "absent", "carrier": false,
|
||||
"policy": "optional", "since_secs": 10, "binds": 0,
|
||||
"failed_attempts": 0
|
||||
},
|
||||
"stats": {}
|
||||
}
|
||||
]
|
||||
});
|
||||
let mut app = app_with(Tab::Transports, data);
|
||||
let buf = testkit::render(110, 20, |frame, area| {
|
||||
super::transports::draw(frame, &mut app, area);
|
||||
});
|
||||
|
||||
let end_of = |needle: &str| testkit::find(&buf, needle).map(|(x, _)| x + needle.len() as u16);
|
||||
|
||||
// Same right edge for a two-character value and a six-character one, and
|
||||
// the header shares it.
|
||||
assert_eq!(end_of("up"), end_of("absent"));
|
||||
assert_eq!(end_of("up"), end_of("State"));
|
||||
}
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
use ratatui::Frame;
|
||||
use ratatui::layout::{Constraint, Layout, Rect};
|
||||
use ratatui::layout::{Alignment, Constraint, Layout, Rect};
|
||||
use ratatui::style::{Color, Modifier, Style};
|
||||
use ratatui::text::{Line, Span};
|
||||
use ratatui::widgets::{
|
||||
@@ -23,6 +23,21 @@ enum TreeRow {
|
||||
},
|
||||
}
|
||||
|
||||
/// Below this width the table drops its byte counters and keeps its
|
||||
/// identifying columns.
|
||||
///
|
||||
/// The full layout's fixed columns sum to 90, plus six single-column gaps —
|
||||
/// 96 against the ~77 usable inside an 80-column terminal's border and
|
||||
/// scrollbar. Ratatui resolves an over-subscribed layout by shrinking every
|
||||
/// column, so the overflow does not clip the rightmost column, it clips *all*
|
||||
/// of them: `mesh0 (optional)` becomes `mesh0 (optio`, losing the one marker
|
||||
/// on the row worth reading.
|
||||
const NARROW_TABLE_WIDTH: u16 = 100;
|
||||
|
||||
/// Below this width the detail view stacks above/below the table instead of
|
||||
/// beside it.
|
||||
const SIDE_BY_SIDE_MIN_WIDTH: u16 = 110;
|
||||
|
||||
pub fn draw(frame: &mut Frame, app: &mut App, area: Rect) {
|
||||
let transports = get_transports(app);
|
||||
let links = get_links(app);
|
||||
@@ -33,8 +48,18 @@ pub fn draw(frame: &mut Frame, app: &mut App, area: Rect) {
|
||||
update_selected_tree_item(app, &tree_rows);
|
||||
|
||||
if app.detail_view.is_some() {
|
||||
let chunks = Layout::horizontal([Constraint::Percentage(40), Constraint::Percentage(60)])
|
||||
.split(area);
|
||||
// Side by side only where the table half can still show its
|
||||
// identifying columns. At 80 columns — the OpenWrt serial console, and
|
||||
// the xterm/tmux default — a 40% split leaves the table 32 columns for
|
||||
// a layout that needs 65 even in its narrow form, and ratatui resolves
|
||||
// that by giving the trailing columns everything and rendering
|
||||
// Transport, Instance, Bound-to and State at width zero. Stacking
|
||||
// keeps both panes readable instead of keeping both unreadable.
|
||||
let chunks = if area.width < SIDE_BY_SIDE_MIN_WIDTH {
|
||||
Layout::vertical([Constraint::Percentage(50), Constraint::Percentage(50)]).split(area)
|
||||
} else {
|
||||
Layout::horizontal([Constraint::Percentage(40), Constraint::Percentage(60)]).split(area)
|
||||
};
|
||||
|
||||
draw_table(frame, app, chunks[0], &transports, &links, &tree_rows);
|
||||
draw_detail(frame, app, chunks[1], &transports, &links, &tree_rows);
|
||||
@@ -113,6 +138,19 @@ fn update_selected_tree_item(app: &mut App, tree_rows: &[TreeRow]) {
|
||||
};
|
||||
}
|
||||
|
||||
/// Build a row, dropping the trailing byte counters in the narrow layout.
|
||||
///
|
||||
/// Both row shapes carry the same seven cells in the same order, so the
|
||||
/// narrow variant is the same list with its tail cut — keeping one place
|
||||
/// where the column count is decided, rather than two that must agree.
|
||||
fn table_row<'a>(cells: Vec<Cell<'a>>, narrow: bool) -> Row<'a> {
|
||||
let mut cells = cells;
|
||||
if narrow {
|
||||
cells.truncate(5);
|
||||
}
|
||||
Row::new(cells)
|
||||
}
|
||||
|
||||
fn draw_table(
|
||||
frame: &mut Frame,
|
||||
app: &mut App,
|
||||
@@ -121,14 +159,24 @@ fn draw_table(
|
||||
links: &[serde_json::Value],
|
||||
tree_rows: &[TreeRow],
|
||||
) {
|
||||
let header = Row::new(vec![
|
||||
// Tx/Rx are the first thing to go when width is short: they are the only
|
||||
// columns whose absence costs nothing an operator is scanning this table
|
||||
// to find, and the detail pane carries them in full.
|
||||
let narrow = area.width < NARROW_TABLE_WIDTH;
|
||||
|
||||
let mut header_cells = vec![
|
||||
Cell::from("Transport / Link"),
|
||||
Cell::from("State"),
|
||||
Cell::from("Instance"),
|
||||
Cell::from("Bound to"),
|
||||
Cell::from(Line::from("State").alignment(Alignment::Right)),
|
||||
Cell::from("Peer"),
|
||||
Cell::from("Tx"),
|
||||
Cell::from("Rx"),
|
||||
])
|
||||
.style(
|
||||
];
|
||||
if narrow {
|
||||
header_cells.truncate(5);
|
||||
}
|
||||
let header = Row::new(header_cells).style(
|
||||
Style::default()
|
||||
.fg(Color::Yellow)
|
||||
.add_modifier(Modifier::BOLD),
|
||||
@@ -153,29 +201,72 @@ fn draw_table(
|
||||
let typ = helpers::str_field(t, "type");
|
||||
let name = t.get("name").and_then(|v| v.as_str()).unwrap_or("");
|
||||
let addr = t.get("local_addr").and_then(|v| v.as_str()).unwrap_or("");
|
||||
let label = if !name.is_empty() {
|
||||
format!("{indicator}{typ} {name}")
|
||||
} else if typ == "tor" {
|
||||
// An interface-bound transport is identified by the netdev it
|
||||
// names, not by its instance label: "ethernet lan" tells an
|
||||
// operator nothing, "ethernet lan br-lan" tells them where to
|
||||
// look. The presence marker is what makes the row honest —
|
||||
// `state` reads `up` from the moment the transport starts,
|
||||
// whether or not it is bound to anything.
|
||||
let iface = t.get("interface");
|
||||
let iface_name = iface
|
||||
.and_then(|i| i.get("name"))
|
||||
.and_then(|v| v.as_str())
|
||||
.unwrap_or("");
|
||||
let presence = iface
|
||||
.and_then(|i| i.get("presence"))
|
||||
.and_then(|v| v.as_str())
|
||||
.unwrap_or("");
|
||||
let policy = iface
|
||||
.and_then(|i| i.get("policy"))
|
||||
.and_then(|v| v.as_str())
|
||||
.unwrap_or("");
|
||||
|
||||
// Three cells, not one packed string. The instance name and
|
||||
// the thing the transport is bound to are separate facts about
|
||||
// separate columns of a table, and running them together left
|
||||
// the netdev names ragged down the list — the column an
|
||||
// operator scans to find the interface they are looking for.
|
||||
let label = if typ == "tor" {
|
||||
let mode = t
|
||||
.get("tor_mode")
|
||||
.and_then(|v| v.as_str())
|
||||
.unwrap_or("socks5");
|
||||
let onion_hint = t
|
||||
.get("onion_address")
|
||||
format!("{indicator}tor({mode})")
|
||||
} else {
|
||||
format!("{indicator}{typ}")
|
||||
};
|
||||
|
||||
// What this transport is attached to: a netdev for the
|
||||
// interface-bound ones, the bound socket address for IP
|
||||
// transports, an onion for tor. Different answers, one
|
||||
// question, so one column.
|
||||
let bound_to = if !iface_name.is_empty() {
|
||||
iface_name.to_string()
|
||||
} else if typ == "tor" {
|
||||
t.get("onion_address")
|
||||
.and_then(|v| v.as_str())
|
||||
.map(|a| {
|
||||
let short = if a.len() > 16 { &a[..16] } else { a };
|
||||
format!(" {short}..")
|
||||
format!("{short}..")
|
||||
})
|
||||
.unwrap_or_default();
|
||||
format!("{indicator}tor({mode}){onion_hint}")
|
||||
.unwrap_or_default()
|
||||
} else if !addr.is_empty() {
|
||||
format!("{indicator}{typ} {addr}")
|
||||
addr.to_string()
|
||||
} else {
|
||||
format!("{indicator}{typ} #{transport_id}")
|
||||
format!("#{transport_id}")
|
||||
};
|
||||
|
||||
let state = helpers::str_field(t, "state");
|
||||
// The State column carries presence for an interface-bound
|
||||
// transport, not the lifecycle state. `up` is true from the
|
||||
// moment the transport starts and stays true while its
|
||||
// interface is missing, so it is precisely the wrong answer in
|
||||
// the one case an operator is scanning this column for. There
|
||||
// is no room to show both, and only one of them is news.
|
||||
let state = if presence.is_empty() || presence == "present" {
|
||||
helpers::str_field(t, "state")
|
||||
} else {
|
||||
presence
|
||||
};
|
||||
let tx = t
|
||||
.get("stats")
|
||||
.and_then(|s| s.get("packets_sent").or_else(|| s.get("frames_sent")))
|
||||
@@ -189,14 +280,45 @@ fn draw_table(
|
||||
.map(|n| n.to_string())
|
||||
.unwrap_or_else(|| "-".into());
|
||||
|
||||
Row::new(vec![
|
||||
Cell::from(label),
|
||||
Cell::from(state.to_string()),
|
||||
Cell::from(""),
|
||||
Cell::from(tx),
|
||||
Cell::from(rx),
|
||||
])
|
||||
.style(Style::default().fg(Color::White))
|
||||
// Colour follows bindability, not lifecycle: an absent
|
||||
// interface is the case the operator most needs to spot, and
|
||||
// it is precisely the one `state` cannot show. Severity then
|
||||
// follows the absence policy, because that is what the policy
|
||||
// means — a dock adapter that is not plugged in is a warning,
|
||||
// an interface the config says to expect is an error. Same
|
||||
// split the daemon makes between `Degraded` and silence.
|
||||
let row_style = match presence {
|
||||
"" | "present" => Style::default().fg(Color::White),
|
||||
"binding" => Style::default().fg(Color::Yellow),
|
||||
_ if policy == "optional" => Style::default().fg(Color::Yellow),
|
||||
_ => Style::default().fg(Color::Red),
|
||||
};
|
||||
|
||||
// Policy rides with the instance name rather than owning a
|
||||
// column: `required` is the default and appears on nearly
|
||||
// every row, so a column of it is a column of noise. Only the
|
||||
// exception is worth printing, and its absence then means the
|
||||
// rule.
|
||||
let instance = match (name.is_empty(), policy == "optional") {
|
||||
(true, true) => "(optional)".to_string(),
|
||||
(true, false) => String::new(),
|
||||
(false, true) => format!("{name} (optional)"),
|
||||
(false, false) => name.to_string(),
|
||||
};
|
||||
|
||||
table_row(
|
||||
vec![
|
||||
Cell::from(label),
|
||||
Cell::from(instance),
|
||||
Cell::from(bound_to),
|
||||
Cell::from(Line::from(state.to_string()).alignment(Alignment::Right)),
|
||||
Cell::from(""),
|
||||
Cell::from(tx),
|
||||
Cell::from(rx),
|
||||
],
|
||||
narrow,
|
||||
)
|
||||
.style(row_style)
|
||||
}
|
||||
TreeRow::Link { index, is_last } => {
|
||||
let link = &links[*index];
|
||||
@@ -214,28 +336,42 @@ fn draw_table(
|
||||
// Wide enough to render a full MAC (~17) or `hci0/MAC`
|
||||
// (~22) without chopping mid-octet; the link detail view
|
||||
// shows the untruncated address.
|
||||
// The remote address goes in `Bound to`, not in the label. A
|
||||
// link is bound to a remote endpoint exactly as a transport is
|
||||
// bound to a netdev or a socket — same question, same column —
|
||||
// and keeping a full MAC out of the first column is what lets
|
||||
// the three left columns sit against the left edge instead of
|
||||
// being pushed right by the widest link row.
|
||||
let addr = helpers::truncate_hex(helpers::str_field(link, "remote_addr"), 24);
|
||||
let label = format!(" {tree_char} {dir_short} {addr}");
|
||||
let label = format!(" {tree_char} {dir_short}");
|
||||
|
||||
let state = helpers::str_field(link, "state");
|
||||
let peer_name = lookup_peer_for_link(app, link)
|
||||
.map(|p| helpers::str_field(&p, "display_name").to_string())
|
||||
.unwrap_or_default();
|
||||
|
||||
Row::new(vec![
|
||||
Cell::from(Span::styled(
|
||||
label,
|
||||
Style::default().fg(if dir == "Outbound" {
|
||||
Color::Cyan
|
||||
} else {
|
||||
Color::Green
|
||||
}),
|
||||
)),
|
||||
Cell::from(state.to_string()),
|
||||
Cell::from(peer_name),
|
||||
Cell::from(""),
|
||||
Cell::from(""),
|
||||
])
|
||||
table_row(
|
||||
vec![
|
||||
Cell::from(Span::styled(
|
||||
label,
|
||||
Style::default().fg(if dir == "Outbound" {
|
||||
Color::Cyan
|
||||
} else {
|
||||
Color::Green
|
||||
}),
|
||||
)),
|
||||
// A link has no instance name of its own — it inherits its
|
||||
// parent transport's, shown one row up — and no absence
|
||||
// policy, which is a property of an interface.
|
||||
Cell::from(""),
|
||||
Cell::from(addr),
|
||||
Cell::from(Line::from(state.to_string()).alignment(Alignment::Right)),
|
||||
Cell::from(peer_name),
|
||||
Cell::from(""),
|
||||
Cell::from(""),
|
||||
],
|
||||
narrow,
|
||||
)
|
||||
}
|
||||
})
|
||||
.collect();
|
||||
@@ -251,15 +387,44 @@ fn draw_table(
|
||||
format!(" Transports ({transport_count}) ")
|
||||
};
|
||||
|
||||
let widths = [
|
||||
Constraint::Min(28),
|
||||
Constraint::Length(12),
|
||||
Constraint::Length(14),
|
||||
Constraint::Length(9),
|
||||
Constraint::Length(9),
|
||||
// Every identifying column is fixed-width and packed against the left
|
||||
// edge; `Peer` takes the slack. The first column used to be `Min`, which
|
||||
// meant it absorbed all spare width and shoved Instance and Bound-to into
|
||||
// the middle of the terminal, away from the names an operator is scanning.
|
||||
//
|
||||
// It is sized for the widest label that lives in it — a link's
|
||||
// ` └─ Out` — rather than for a full MAC, because the address moved to
|
||||
// `Bound to` where it belongs. `Bound to` is sized for a MAC (17), which
|
||||
// also covers every netdev name and socket address that shares it.
|
||||
// Narrow: the same columns, sized down to what still reads. Instance keeps
|
||||
// 18 because `mesh0 (optional)` is 16 and the marker is the point; `Bound
|
||||
// to` keeps 17 because that is a full MAC.
|
||||
let widths: &[Constraint] = if narrow {
|
||||
&[
|
||||
Constraint::Length(12), // Transport / Link
|
||||
Constraint::Length(18), // Instance, plus "(optional)"
|
||||
Constraint::Length(17), // Bound to: a full MAC
|
||||
Constraint::Length(9), // State, right-aligned
|
||||
Constraint::Min(6), // Peer — takes the slack
|
||||
]
|
||||
} else {
|
||||
&FULL_WIDTHS
|
||||
};
|
||||
|
||||
const FULL_WIDTHS: [Constraint; 7] = [
|
||||
Constraint::Length(18), // Transport / Link
|
||||
Constraint::Length(20), // Instance, plus "(optional)" where it applies
|
||||
Constraint::Length(18), // Bound to: netdev, socket addr, onion, MAC
|
||||
// Wide enough for the longest value that lands here — `connected` (9)
|
||||
// and `binding` — because a right-aligned cell clips from the LEFT,
|
||||
// so an overflow reads as `onnected` rather than as a truncation.
|
||||
Constraint::Length(10), // State, right-aligned
|
||||
Constraint::Min(10), // Peer — takes the slack
|
||||
Constraint::Length(7), // Tx
|
||||
Constraint::Length(7), // Rx
|
||||
];
|
||||
|
||||
let table = Table::new(rows, widths)
|
||||
let table = Table::new(rows, widths.to_vec())
|
||||
.header(header)
|
||||
.block(Block::default().borders(Borders::ALL).title(title))
|
||||
.row_highlight_style(
|
||||
@@ -341,6 +506,73 @@ fn draw_transport_detail(frame: &mut Frame, app: &App, area: Rect, t: &serde_jso
|
||||
lines.push(helpers::kv_line("Local Addr", addr));
|
||||
}
|
||||
|
||||
// Interface presence, for the transports that are bound to a netdev.
|
||||
//
|
||||
// `State` above answers a lifecycle question — was this transport started
|
||||
// — and reads `up` for an interface that has never existed. That gap is
|
||||
// the whole reason interface binding is observable at all: the original
|
||||
// OpenWrt bug was expensive because the 802.11s link formed regardless, so
|
||||
// nothing an operator could see said the node was deaf. This is where they
|
||||
// see it.
|
||||
if let Some(iface) = t.get("interface") {
|
||||
lines.push(Line::from(""));
|
||||
lines.push(helpers::section_header("Interface"));
|
||||
lines.push(helpers::kv_line(
|
||||
"Interface",
|
||||
helpers::str_field(iface, "name"),
|
||||
));
|
||||
|
||||
let presence = helpers::str_field(iface, "presence");
|
||||
let since = iface
|
||||
.get("since_secs")
|
||||
.and_then(|v| v.as_u64())
|
||||
.map(|secs| helpers::format_duration_ms(secs.saturating_mul(1000)))
|
||||
.unwrap_or_else(|| "-".into());
|
||||
lines.push(helpers::kv_line(
|
||||
"Presence",
|
||||
&format!("{presence} for {since}"),
|
||||
));
|
||||
|
||||
// Carrier is reported, never acted on: presence is IFF_UP, so a bound
|
||||
// interface with no carrier is normal (a bridge with nothing plugged
|
||||
// into it) rather than a fault. Saying so beats an operator inferring
|
||||
// it from silence.
|
||||
let carrier = iface
|
||||
.get("carrier")
|
||||
.and_then(|v| v.as_bool())
|
||||
.map(|c| if c { "yes" } else { "no" })
|
||||
.unwrap_or("-");
|
||||
lines.push(helpers::kv_line("Carrier", carrier));
|
||||
|
||||
// The list marks only the exception, `(optional)`, beside the
|
||||
// instance name. The detail pane has room to spell out both, as the
|
||||
// consequence rather than the config key: `optional` is a statement
|
||||
// about what absence *means*, and someone who has opened this pane
|
||||
// wants the meaning.
|
||||
let policy = helpers::str_field(iface, "policy");
|
||||
let absence = if policy == "optional" {
|
||||
"optional (absence is normal)"
|
||||
} else {
|
||||
"required (absence degrades the node)"
|
||||
};
|
||||
lines.push(helpers::kv_line("On absence", absence));
|
||||
|
||||
// Binds past the first are rebinds, and a climbing failed-attempt
|
||||
// count is an interface that is there and refusing — a different
|
||||
// problem from one that is missing, and invisible without this.
|
||||
let binds = iface.get("binds").and_then(|v| v.as_u64()).unwrap_or(0);
|
||||
if binds > 1 {
|
||||
lines.push(helpers::kv_line("Binds", &format!("{binds} (rebound)")));
|
||||
} else {
|
||||
lines.push(helpers::kv_line("Binds", &binds.to_string()));
|
||||
}
|
||||
if let Some(failed) = iface.get("failed_attempts").and_then(|v| v.as_u64())
|
||||
&& failed > 0
|
||||
{
|
||||
lines.push(helpers::kv_line("Failed binds", &failed.to_string()));
|
||||
}
|
||||
}
|
||||
|
||||
// Tor-specific info
|
||||
if let Some(mode) = t.get("tor_mode").and_then(|v| v.as_str()) {
|
||||
lines.push(helpers::kv_line("Tor Mode", mode));
|
||||
|
||||
+104
-60
@@ -1057,6 +1057,68 @@ impl Config {
|
||||
|
||||
/// Validate cross-field configuration invariants.
|
||||
pub fn validate(&self) -> Result<(), ConfigError> {
|
||||
self.validate_ethernet_interfaces()?;
|
||||
self.validate_rendezvous()
|
||||
}
|
||||
|
||||
/// Reject interface names no kernel could ever hand back.
|
||||
///
|
||||
/// The presence machine deliberately cannot tell a typo from an interface
|
||||
/// that has not been created yet — both are simply absent, and waiting is
|
||||
/// the right answer for the second. That is what makes this check worth
|
||||
/// having: a name that is *impossible* is the one case still separable
|
||||
/// from "not there yet", and without it a typo costs a permanently
|
||||
/// `Degraded` node whose only symptom is an interface that never arrives.
|
||||
///
|
||||
/// Syntax only. Whether a well-formed name exists is the binder's
|
||||
/// question, asked once a second, forever.
|
||||
fn validate_ethernet_interfaces(&self) -> Result<(), ConfigError> {
|
||||
// Kernel limit: `IFNAMSIZ` is 16 including the terminating NUL, on
|
||||
// both Linux and the BSDs.
|
||||
const MAX_INTERFACE_NAME: usize = 15;
|
||||
|
||||
let mut seen: std::collections::HashMap<&str, &str> = std::collections::HashMap::new();
|
||||
|
||||
for (name, cfg) in self.transports.ethernet.iter() {
|
||||
let label = name.unwrap_or("ethernet");
|
||||
let iface = cfg.interface.as_str();
|
||||
|
||||
if iface.is_empty() {
|
||||
return Err(ConfigError::Validation(format!(
|
||||
"transport `{label}` has an empty `interface`"
|
||||
)));
|
||||
}
|
||||
if iface.len() > MAX_INTERFACE_NAME {
|
||||
return Err(ConfigError::Validation(format!(
|
||||
"transport `{label}` interface `{iface}` is {} bytes; \
|
||||
the kernel limit is {MAX_INTERFACE_NAME}, so no such \
|
||||
interface can exist",
|
||||
iface.len()
|
||||
)));
|
||||
}
|
||||
if iface.contains('/') || iface.chars().any(char::is_whitespace) {
|
||||
return Err(ConfigError::Validation(format!(
|
||||
"transport `{label}` interface `{iface}` contains a \
|
||||
character no interface name may hold"
|
||||
)));
|
||||
}
|
||||
|
||||
// Two transports on one netdev means two sockets on the same
|
||||
// device at the same ethertype, each receiving every frame the
|
||||
// other does.
|
||||
if let Some(prior) = seen.insert(iface, label) {
|
||||
return Err(ConfigError::Validation(format!(
|
||||
"transports `{prior}` and `{label}` both bind interface \
|
||||
`{iface}`"
|
||||
)));
|
||||
}
|
||||
}
|
||||
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Cross-checks between transports, peers and the Nostr rendezvous.
|
||||
fn validate_rendezvous(&self) -> Result<(), ConfigError> {
|
||||
let nostr = &self.node.rendezvous.nostr;
|
||||
|
||||
let any_transport_advertises_on_nostr = self
|
||||
@@ -1407,83 +1469,65 @@ node:
|
||||
}
|
||||
|
||||
/// The fips.yaml shipped in the OpenWrt package must keep parsing as the
|
||||
/// config schema evolves. Both the 802.11s mesh backhaul entries
|
||||
/// config schema evolves, and must keep the absence policy it depends on.
|
||||
///
|
||||
/// The 802.11s mesh backhaul entries
|
||||
/// (docs/how-to/set-up-80211s-mesh-backhaul.md) and the open-access SSID
|
||||
/// entries (docs/how-to/set-up-open-access-ssid.md) ship commented out —
|
||||
/// one per radio, so dual-band routers can run either on both bands — so
|
||||
/// a stock install that never creates fips-mesh*/fips-ap* logs no
|
||||
/// per-boot bind warning; `fips-mesh-setup`/`fips-ap-setup` uncomment the
|
||||
/// matching block when they create the interface. Verify both states
|
||||
/// parse: as shipped (both inactive), and after the uncomment the helpers
|
||||
/// perform.
|
||||
/// entries (docs/how-to/set-up-open-access-ssid.md) ship **enabled** — one
|
||||
/// per radio, so dual-band routers can run either on both bands — and
|
||||
/// marked `optional: true`. They used to ship commented out, with
|
||||
/// `fips-mesh-setup`/`fips-ap-setup` uncommenting the matching block; that
|
||||
/// was a workaround for a transport whose missing interface was skipped
|
||||
/// for the life of the process. The daemon now waits for the interface and
|
||||
/// binds it when it appears, and `optional: true` is what keeps a stock
|
||||
/// install that never runs those helpers quiet and un-`Degraded` about a
|
||||
/// radio it was never going to have.
|
||||
#[test]
|
||||
fn shipped_openwrt_config_parses() {
|
||||
let yaml = include_str!("../../packaging/openwrt-ipk/files/etc/fips/fips.yaml");
|
||||
|
||||
// As shipped: parses, and the mesh/ap entries are commented out (a
|
||||
// running daemon binds no fips-mesh*/fips-ap* transport, no warning).
|
||||
let config: Config = serde_yaml::from_str(yaml).expect("shipped OpenWrt fips.yaml");
|
||||
for name in ["mesh0", "mesh1", "ap0", "ap1"] {
|
||||
assert!(
|
||||
!config
|
||||
.transports
|
||||
.ethernet
|
||||
.iter()
|
||||
.any(|(n, _)| n == Some(name)),
|
||||
"{name} must ship commented out, not active, in fips.yaml"
|
||||
);
|
||||
}
|
||||
|
||||
// What `fips-mesh-setup`/`fips-ap-setup` produce: uncomment each
|
||||
// block, which must still parse into a transport bound to the right
|
||||
// netdev.
|
||||
let uncommented =
|
||||
uncomment_transport_blocks(&uncomment_transport_blocks(yaml, "mesh"), "ap");
|
||||
let config: Config = serde_yaml::from_str(&uncommented)
|
||||
.expect("fips.yaml with mesh and ap transports uncommented");
|
||||
let eth = |name: &str| {
|
||||
config
|
||||
.transports
|
||||
.ethernet
|
||||
.iter()
|
||||
.find(|(n, _)| *n == Some(name))
|
||||
.map(|(_, cfg)| cfg.clone())
|
||||
};
|
||||
|
||||
// The interfaces the setup helpers create: present, bound to the right
|
||||
// netdev, and optional. A regression to `optional: false` here would
|
||||
// report every stock router `Degraded` for a mesh it never configured.
|
||||
for (name, interface) in [
|
||||
("mesh0", "fips-mesh0"),
|
||||
("mesh1", "fips-mesh1"),
|
||||
("ap0", "fips-ap0"),
|
||||
("ap1", "fips-ap1"),
|
||||
] {
|
||||
let cfg = eth(name).unwrap_or_else(|| panic!("{name} entry missing from fips.yaml"));
|
||||
assert_eq!(cfg.interface, interface, "{name} binds the wrong netdev");
|
||||
assert!(cfg.optional(), "{name} must ship optional");
|
||||
}
|
||||
|
||||
// `phy0-sta0` only exists while a radio is in station mode.
|
||||
assert!(
|
||||
eth("wwan").expect("wwan entry").optional(),
|
||||
"wwan must ship optional"
|
||||
);
|
||||
|
||||
// The wired ports exist on every supported board, so their absence is
|
||||
// a real fault and must stay loud. Marking these optional too would
|
||||
// make the whole ethernet block silent, which is the failure mode the
|
||||
// presence mechanism exists to stop hiding.
|
||||
for name in ["wan", "lan"] {
|
||||
assert!(
|
||||
config
|
||||
.transports
|
||||
.ethernet
|
||||
.iter()
|
||||
.any(|(n, eth)| n == Some(name) && eth.interface == interface),
|
||||
"{name} entry missing after uncommenting shipped fips.yaml"
|
||||
!eth(name).expect("wired entry").optional(),
|
||||
"{name} must stay required"
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
/// Mirror the setup helpers' block uncomment: strip the ` # ` prefix
|
||||
/// from each `# <prefix><N>:` header and its ` # ` continuation
|
||||
/// lines, leaving every other comment untouched.
|
||||
fn uncomment_transport_blocks(yaml: &str, prefix: &str) -> String {
|
||||
let header = format!(" # {prefix}");
|
||||
let mut out = String::new();
|
||||
let mut in_block = false;
|
||||
for line in yaml.lines() {
|
||||
let is_header = line
|
||||
.strip_prefix(&header)
|
||||
.and_then(|r| r.strip_suffix(':'))
|
||||
.is_some_and(|n| !n.is_empty() && n.bytes().all(|b| b.is_ascii_digit()));
|
||||
if is_header {
|
||||
in_block = true;
|
||||
out.push_str(&line.replacen(" # ", " ", 1));
|
||||
} else if in_block && line.starts_with(" # ") {
|
||||
out.push_str(&line.replacen(" # ", " ", 1));
|
||||
} else {
|
||||
in_block = false;
|
||||
out.push_str(line);
|
||||
}
|
||||
out.push('\n');
|
||||
}
|
||||
out
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_parse_yaml_with_hex() {
|
||||
let yaml = r#"
|
||||
|
||||
@@ -302,6 +302,25 @@ pub struct EthernetConfig {
|
||||
/// Announcement beacon interval in seconds. Default: 30.
|
||||
#[serde(default, skip_serializing_if = "Option::is_none")]
|
||||
pub beacon_interval_secs: Option<u64>,
|
||||
|
||||
/// Whether the absence of this interface is normal. Default: false.
|
||||
///
|
||||
/// Naming an interface in configuration is a statement that you expect it,
|
||||
/// so the default is to complain: while the interface is missing the node
|
||||
/// reports `Degraded`, the edge is logged (`info` at startup, `warn` on a
|
||||
/// runtime detach), and an absence that outlasts the bring-up window — ten
|
||||
/// seconds, past which it is no longer a race against a radio or a
|
||||
/// container — is reported once at `error`. Set `optional: true` for
|
||||
/// hardware that is legitimately not always there — a dock adapter, a
|
||||
/// radio that only some boards carry — and its absence becomes silent
|
||||
/// (`info` on the edge, no health impact, no error).
|
||||
///
|
||||
/// This describes **the interface's presence, not the transport's
|
||||
/// importance**. An optional interface that is present is used exactly as
|
||||
/// hard as any other; setting it does not deprioritize the transport, and
|
||||
/// no value of this field makes a missing interface fatal at startup.
|
||||
#[serde(default, skip_serializing_if = "Option::is_none")]
|
||||
pub optional: Option<bool>,
|
||||
}
|
||||
|
||||
impl EthernetConfig {
|
||||
@@ -340,6 +359,11 @@ impl EthernetConfig {
|
||||
self.accept_connections.unwrap_or(false)
|
||||
}
|
||||
|
||||
/// Whether absence of the interface is normal. Default: false.
|
||||
pub fn optional(&self) -> bool {
|
||||
self.optional.unwrap_or(false)
|
||||
}
|
||||
|
||||
/// Get the beacon interval, clamped to minimum. Default: 30s.
|
||||
pub fn beacon_interval_secs(&self) -> u64 {
|
||||
self.beacon_interval_secs
|
||||
@@ -1071,4 +1095,40 @@ mod tests {
|
||||
serde_yaml::from_str("interface: eth0\nbogus: true\n");
|
||||
assert!(bogus.is_err());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn ethernet_absence_is_an_error_unless_opted_out() {
|
||||
// Naming an interface is a statement that you expect it, so the
|
||||
// default has to be the loud one. A default of `true` here would make
|
||||
// every missing interface silent, which is the failure mode the
|
||||
// whole presence mechanism exists to stop hiding.
|
||||
let bare: EthernetConfig = serde_yaml::from_str("interface: eth0\n").unwrap();
|
||||
assert_eq!(bare.optional, None);
|
||||
assert!(!bare.optional(), "absence must default to an error");
|
||||
|
||||
let opted: EthernetConfig =
|
||||
serde_yaml::from_str("interface: enx00e04c680001\noptional: true\n").unwrap();
|
||||
assert!(opted.optional());
|
||||
|
||||
let explicit: EthernetConfig =
|
||||
serde_yaml::from_str("interface: eth0\noptional: false\n").unwrap();
|
||||
assert!(!explicit.optional());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn ethernet_optional_survives_a_round_trip() {
|
||||
// The packaged OpenWrt config ships `optional: true` on the mesh and
|
||||
// access blocks; a serializer that dropped it would silently turn
|
||||
// every stock router Degraded.
|
||||
let opted: EthernetConfig =
|
||||
serde_yaml::from_str("interface: fips-mesh0\noptional: true\n").unwrap();
|
||||
let round: EthernetConfig =
|
||||
serde_yaml::from_str(&serde_yaml::to_string(&opted).unwrap()).unwrap();
|
||||
assert!(round.optional());
|
||||
|
||||
// ... and the default stays absent from the output rather than being
|
||||
// written back as an explicit `false`.
|
||||
let bare: EthernetConfig = serde_yaml::from_str("interface: eth0\n").unwrap();
|
||||
assert!(!serde_yaml::to_string(&bare).unwrap().contains("optional"));
|
||||
}
|
||||
}
|
||||
|
||||
+90
-2
@@ -1374,8 +1374,17 @@ pub(crate) fn show_connections_from_handle(
|
||||
|
||||
/// `show_transports` — Transport instances.
|
||||
pub fn show_transports(node: &Node) -> Value {
|
||||
let transports: Vec<Value> = node
|
||||
.transport_ids()
|
||||
// Ascending id, which is creation order: UDP, then Ethernet, then TCP,
|
||||
// Tor, Nym, BLE, with each type's instances in config order. The map
|
||||
// behind `transport_ids` is a `HashMap`, so without this the array order
|
||||
// is whatever the hash seed produced — arbitrary, and different on every
|
||||
// daemon restart. Anything scripting against this output, and every view
|
||||
// rendering it, inherits that. Sorting by id groups the list by transport
|
||||
// type for free, because the ids were handed out that way.
|
||||
let mut ids: Vec<_> = node.transport_ids().copied().collect();
|
||||
ids.sort_by_key(|id| id.as_u32());
|
||||
let transports: Vec<Value> = ids
|
||||
.iter()
|
||||
.map(|id| {
|
||||
let handle = node.get_transport(id).unwrap();
|
||||
let mut t_json = json!({
|
||||
@@ -1403,6 +1412,21 @@ pub fn show_transports(node: &Node) -> Value {
|
||||
t_json["tor_monitoring"] = serde_json::to_value(&monitoring).unwrap_or_default();
|
||||
}
|
||||
|
||||
// Interface presence, for the transports that have an interface.
|
||||
// Absent from the payload entirely for the ones that do not, rather
|
||||
// than reported as a permanently-`present` interface named "".
|
||||
if let Some(p) = handle.interface_presence() {
|
||||
t_json["interface"] = json!({
|
||||
"name": handle.interface_name().unwrap_or_default(),
|
||||
"presence": p.presence,
|
||||
"carrier": p.carrier,
|
||||
"policy": p.policy,
|
||||
"since_secs": p.since_secs,
|
||||
"binds": p.binds,
|
||||
"failed_attempts": p.failed_attempts,
|
||||
});
|
||||
}
|
||||
|
||||
t_json["stats"] = handle.transport_stats();
|
||||
|
||||
t_json
|
||||
@@ -1447,6 +1471,18 @@ pub(crate) fn show_transports_from_handle(handle: &super::read_handle::ControlRe
|
||||
t_json["tor_monitoring"] = monitoring.clone();
|
||||
}
|
||||
|
||||
if let Some(iface) = &t.interface {
|
||||
t_json["interface"] = json!({
|
||||
"name": iface.name,
|
||||
"presence": iface.presence,
|
||||
"carrier": iface.carrier,
|
||||
"policy": iface.policy,
|
||||
"since_secs": iface.since_secs,
|
||||
"binds": iface.binds,
|
||||
"failed_attempts": iface.failed_attempts,
|
||||
});
|
||||
}
|
||||
|
||||
t_json["stats"] = t.stats.clone();
|
||||
|
||||
t_json
|
||||
@@ -2568,6 +2604,8 @@ mod tests {
|
||||
"idle_ms",
|
||||
"first_seen_secs_ago",
|
||||
"last_contact_secs_ago",
|
||||
// Interface presence: elapsed since the current phase began.
|
||||
"since_secs",
|
||||
];
|
||||
|
||||
/// Build a Node with a fixed identity, default config, and empty
|
||||
@@ -2699,6 +2737,56 @@ mod tests {
|
||||
|
||||
// ---- 19 handler snapshot tests --------------------------------------
|
||||
|
||||
/// The `interface` block, which `build_test_node` cannot produce: it
|
||||
/// keeps every transport list empty, so the nineteen snapshots above pin
|
||||
/// `show_transports` only in its empty form. The block is emitted by two
|
||||
/// hand-duplicated sites (the live handler and the read-handle variant)
|
||||
/// that agree today with nothing enforcing it, and the control-socket
|
||||
/// reference states the response schema is pinned by these snapshots.
|
||||
///
|
||||
/// A separate node rather than a richer `build_test_node`, so the other
|
||||
/// snapshots keep their empty-state determinism.
|
||||
#[cfg(any(target_os = "linux", target_os = "macos"))]
|
||||
#[tokio::test]
|
||||
async fn snapshot_show_transports_with_interface() {
|
||||
use crate::config::EthernetConfig;
|
||||
use crate::transport::ethernet::EthernetTransport;
|
||||
use crate::transport::{TransportHandle, TransportId};
|
||||
|
||||
let mut node = build_test_node();
|
||||
|
||||
// An interface no host has, so presence is deterministically absent
|
||||
// and carrier deterministically false on every machine this runs on.
|
||||
let config = EthernetConfig {
|
||||
interface: "fips-absent-x0".to_string(),
|
||||
ethertype: None,
|
||||
mtu: None,
|
||||
recv_buf_size: None,
|
||||
send_buf_size: None,
|
||||
listen: Some(true),
|
||||
announce: Some(false),
|
||||
auto_connect: None,
|
||||
accept_connections: None,
|
||||
beacon_interval_secs: None,
|
||||
optional: Some(false),
|
||||
};
|
||||
let (tx, _rx) = crate::transport::packet_channel(8);
|
||||
let mut eth = EthernetTransport::new(TransportId::new(1), Some("lab".into()), config, tx);
|
||||
|
||||
// Started, because "up with its interface absent" is the state an
|
||||
// operator actually meets — and the one whose shape is new here.
|
||||
eth.start_async()
|
||||
.await
|
||||
.expect("absence is not a start failure");
|
||||
|
||||
node.insert_transport_for_test(TransportId::new(1), TransportHandle::Ethernet(eth));
|
||||
|
||||
assert_snapshot(
|
||||
"show_transports_with_interface",
|
||||
&render(show_transports(&node)),
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn snapshot_show_status() {
|
||||
let node = build_test_node();
|
||||
|
||||
@@ -778,6 +778,21 @@ pub(crate) struct TransportRow {
|
||||
pub onion_address: Option<String>,
|
||||
pub tor_monitoring: Option<serde_json::Value>,
|
||||
pub stats: serde_json::Value,
|
||||
/// Interface presence for interface-bound transports; `None` for the rest.
|
||||
pub interface: Option<InterfaceRow>,
|
||||
}
|
||||
|
||||
/// Interface name, presence and policy for an interface-bound transport, as
|
||||
/// `show_transports` renders it.
|
||||
#[derive(Clone, PartialEq)]
|
||||
pub(crate) struct InterfaceRow {
|
||||
pub name: String,
|
||||
pub presence: &'static str,
|
||||
pub carrier: bool,
|
||||
pub policy: &'static str,
|
||||
pub since_secs: u64,
|
||||
pub binds: u64,
|
||||
pub failed_attempts: u32,
|
||||
}
|
||||
|
||||
/// MMP trend labels for a peer's link-layer block in `show_mmp` (each present
|
||||
|
||||
@@ -0,0 +1,36 @@
|
||||
{
|
||||
"data": {
|
||||
"transports": [
|
||||
{
|
||||
"interface": {
|
||||
"binds": 0,
|
||||
"carrier": false,
|
||||
"failed_attempts": 0,
|
||||
"name": "fips-absent-x0",
|
||||
"policy": "required",
|
||||
"presence": "absent",
|
||||
"since_secs": "<redacted>"
|
||||
},
|
||||
"mtu": 1496,
|
||||
"name": "lab",
|
||||
"state": "up",
|
||||
"stats": {
|
||||
"beacons_dropped": 0,
|
||||
"beacons_recv": 0,
|
||||
"beacons_sent": 0,
|
||||
"bytes_recv": 0,
|
||||
"bytes_sent": 0,
|
||||
"frames_recv": 0,
|
||||
"frames_sent": 0,
|
||||
"frames_too_long": 0,
|
||||
"frames_too_short": 0,
|
||||
"recv_errors": 0,
|
||||
"send_errors": 0
|
||||
},
|
||||
"transport_id": 1,
|
||||
"type": "ethernet"
|
||||
}
|
||||
]
|
||||
},
|
||||
"status": "ok"
|
||||
}
|
||||
@@ -195,6 +195,31 @@ impl Node {
|
||||
None => None,
|
||||
};
|
||||
if let Some(e) = send_err {
|
||||
// A transient refusal is not a failed handshake. The
|
||||
// interface under this transport is absent or
|
||||
// mid-rebind, and the binder is already working to
|
||||
// bring it back — so the half-built link is left
|
||||
// exactly as it is for the initiator's msg1 resend to
|
||||
// land on, rather than being torn down and rebuilt.
|
||||
//
|
||||
// Tearing down here charged a *local* interface flap
|
||||
// to the remote: the reject counter it recorded means
|
||||
// "the peer sent something invalid", which is a
|
||||
// different thing entirely and one an operator reads
|
||||
// as the peer's fault.
|
||||
//
|
||||
// Nothing leaks by staying. An initiator that never
|
||||
// resends leaves a stale connection, which
|
||||
// `check_timeouts` reaps at `handshake_timeout_secs`
|
||||
// exactly as it reaps every other abandoned handshake.
|
||||
if e.is_transient() {
|
||||
debug!(
|
||||
link_id = %link,
|
||||
error = %e,
|
||||
"Deferred msg2: the transport is between interfaces"
|
||||
);
|
||||
return;
|
||||
}
|
||||
// Restored pre-refactor msg2-send-failure warn!
|
||||
// (`handle_msg1` L665): the send error text is surfaced
|
||||
// at the executor point where the failure is now handled.
|
||||
|
||||
@@ -109,6 +109,16 @@ impl Node {
|
||||
}
|
||||
};
|
||||
|
||||
// Interface-presence receiver, or a dummy channel — same pattern and
|
||||
// same reason as the child-liveness receiver above.
|
||||
let (mut presence_rx, _presence_guard) = match self.transport_presence_rx.take() {
|
||||
Some(rx) => (rx, None),
|
||||
None => {
|
||||
let (tx, rx) = tokio::sync::mpsc::channel(1);
|
||||
(rx, Some(tx))
|
||||
}
|
||||
};
|
||||
|
||||
let tick_period = Duration::from_secs(self.config().node.tick_interval_secs);
|
||||
let mut tick = tokio::time::interval(tick_period);
|
||||
|
||||
@@ -346,6 +356,71 @@ impl Node {
|
||||
self.supervisor.state = ns;
|
||||
}
|
||||
}
|
||||
// A transport child exiting leaves the bound set, so
|
||||
// it can be the one that was holding the node's egress
|
||||
// MTU down. `is_bound()` is `is_operational()` plus the
|
||||
// presence refinement, and this moves the first half.
|
||||
self.refresh_tun_mss_ceiling();
|
||||
}
|
||||
}
|
||||
// Interface presence. An interface-bound transport's binder
|
||||
// reports attach and detach; the FSM folds it into health.
|
||||
// Unlike `ChildExited` this is reversible in both directions —
|
||||
// the interface coming back republishes `Running` — which is
|
||||
// the whole point of `Degraded` being a level rather than a
|
||||
// latch.
|
||||
maybe_presence = presence_rx.recv() => {
|
||||
if let Some(edge) = maybe_presence {
|
||||
// Health is policy-filtered; the MTU floor below is
|
||||
// not. An `optional` interface's absence is normal and
|
||||
// must not move the node off `Full`, but it changes
|
||||
// the bound set all the same.
|
||||
if edge.health_relevant {
|
||||
let child = crate::node::lifecycle::supervisor::Child::Transport(
|
||||
edge.transport_id,
|
||||
);
|
||||
let event = if edge.present {
|
||||
crate::node::lifecycle::supervisor::Event::ChildPresent { child }
|
||||
} else {
|
||||
crate::node::lifecycle::supervisor::Event::ChildAbsent { child }
|
||||
};
|
||||
let actions = self.supervisor.fsm.step(event);
|
||||
for action in actions {
|
||||
if let crate::node::lifecycle::supervisor::Action::PublishState(
|
||||
ns,
|
||||
) = action
|
||||
{
|
||||
self.supervisor.state = ns;
|
||||
}
|
||||
}
|
||||
}
|
||||
// The bound set just changed, so the node's egress MTU
|
||||
// floor may have. Both directions: an interface that
|
||||
// binds can be the narrow one, and one that detaches
|
||||
// can be the reason the clamp was tight.
|
||||
self.refresh_tun_mss_ceiling();
|
||||
|
||||
// A peer reachable only through an interface that has
|
||||
// gone is not reachable. Withdraw it now rather than
|
||||
// leaving the liveness reaper to notice up to
|
||||
// `link_dead_timeout_secs` later, during which this
|
||||
// node both drops transit traffic in silence and keeps
|
||||
// advertising reachability it does not have.
|
||||
//
|
||||
// Not policy-filtered: whether an interface's absence
|
||||
// is normal is a statement about node *health*, not
|
||||
// about whether the routes over it still work.
|
||||
if !edge.present {
|
||||
let reaped =
|
||||
self.reap_peers_on_transport(edge.transport_id).await;
|
||||
if reaped > 0 {
|
||||
info!(
|
||||
transport_id = %edge.transport_id,
|
||||
peers = reaped,
|
||||
"Withdrew peers whose interface went away"
|
||||
);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
Some(ipv6_packet) = tun_outbound_rx.recv() => {
|
||||
|
||||
@@ -601,10 +601,72 @@ impl Node {
|
||||
}
|
||||
}
|
||||
|
||||
/// Route a link-dead liveness reap through the peer machine + executor.
|
||||
/// The shell already decided (the tick sweep's `plan_heartbeats` batch
|
||||
/// emitted this `ReapPeer` in phase order), so the machine only CONSUMES the
|
||||
/// decision via [`PeerEvent::LinkDeadSuspected`]. The resulting executor arms
|
||||
/// Reap every active peer reachable only through `transport_id`.
|
||||
///
|
||||
/// Called on a transport's detach edge. Until this existed, losing an
|
||||
/// interface withdrew nothing: the peers stayed in the registry, the
|
||||
/// routes through them stayed selectable, and this node kept advertising
|
||||
/// reachability it no longer had — so transit traffic was dropped in
|
||||
/// silence, and other nodes kept routing toward us for those destinations,
|
||||
/// until the liveness reaper noticed up to `link_dead_timeout_secs` later.
|
||||
/// Measured on real hardware that was 27 seconds of routing through a link
|
||||
/// that had already gone, with four alternative peers available the whole
|
||||
/// time.
|
||||
///
|
||||
/// The detach edge is both earlier and more certain than inactivity, so it
|
||||
/// is the better trigger. This routes through the same
|
||||
/// [`Self::route_link_dead`] the liveness reaper uses rather than
|
||||
/// open-coding a second teardown — every consequence of losing a peer
|
||||
/// (sessions, path MTU, session indices, the link, the control machine,
|
||||
/// tree cleanup and re-announce, bloom withdrawal) already hangs off that
|
||||
/// one path, and a parallel one would drift from it.
|
||||
///
|
||||
/// Deliberately undamped. A flapping interface cannot drive a reap storm
|
||||
/// through here, because `ChurnGuard` stops publishing presence edges
|
||||
/// after three short-lived bindings and does not resume until one lasts —
|
||||
/// so the edges this reacts to are already rate-limited at the source, and
|
||||
/// a second damper here would only add a way for the two to disagree.
|
||||
///
|
||||
/// Returns how many peers were reaped.
|
||||
pub(in crate::node) async fn reap_peers_on_transport(
|
||||
&mut self,
|
||||
transport_id: TransportId,
|
||||
) -> usize {
|
||||
let doomed: Vec<NodeAddr> = self
|
||||
.peers
|
||||
.iter()
|
||||
.filter(|(_, peer)| peer.transport_id() == Some(transport_id))
|
||||
.map(|(node_addr, _)| *node_addr)
|
||||
.collect();
|
||||
|
||||
if doomed.is_empty() {
|
||||
return 0;
|
||||
}
|
||||
|
||||
let now_ms = std::time::SystemTime::now()
|
||||
.duration_since(std::time::UNIX_EPOCH)
|
||||
.map(|d| d.as_millis() as u64)
|
||||
.unwrap_or(0);
|
||||
|
||||
let reaped = doomed.len();
|
||||
for node_addr in doomed {
|
||||
debug!(
|
||||
peer = %self.peer_display_name(&node_addr),
|
||||
%transport_id,
|
||||
"Removing peer: its interface went away"
|
||||
);
|
||||
self.route_link_dead(node_addr, now_ms).await;
|
||||
}
|
||||
reaped
|
||||
}
|
||||
|
||||
/// Route a link-dead reap through the peer machine + executor. Two callers
|
||||
/// decide: the tick sweep's `plan_heartbeats` batch emits a `ReapPeer` for a
|
||||
/// peer that has gone quiet, and [`Self::reap_peers_on_transport`] withdraws
|
||||
/// a transport's peers when its interface goes away. Mirrors
|
||||
/// [`route_rekey_cadence`](Node::route_rekey_cadence): the shell has already
|
||||
/// decided by the time this runs, so the machine only CONSUMES the decision
|
||||
/// via [`PeerEvent::LinkDeadSuspected`]. The resulting executor arms
|
||||
/// (`InvalidateSendState` → `remove_active_peer`, `ReportLost` →
|
||||
/// `note_link_dead`) reproduce the pre-refactor inline reap body exactly, in
|
||||
/// that order.
|
||||
@@ -614,9 +676,10 @@ impl Node {
|
||||
/// no-op, so we return; if the machine is absent (which should be impossible)
|
||||
/// we fall back to the byte-identical inline body under a `debug_assert`.
|
||||
///
|
||||
/// `now_ms` is the sweep's hoisted wall-clock ms (the same value the old reap
|
||||
/// fed `note_link_dead`); it flows to the executor `ReportLost` arm via
|
||||
/// `ambient.now_ms`.
|
||||
/// `now_ms` is the caller's hoisted wall-clock ms (the same value the old
|
||||
/// reap fed `note_link_dead`); it flows to the executor `ReportLost` arm via
|
||||
/// `ambient.now_ms`. Both callers hoist it once per batch, so every peer
|
||||
/// removed in one pass carries the same instant.
|
||||
async fn route_link_dead(&mut self, node_addr: NodeAddr, now_ms: u64) {
|
||||
let link = match self.peers.get(&node_addr) {
|
||||
Some(peer) => peer.link_id(),
|
||||
|
||||
@@ -385,11 +385,26 @@ impl Node {
|
||||
);
|
||||
}
|
||||
Err(e) => {
|
||||
warn!(
|
||||
peer = %self.peer_display_name(node_addr),
|
||||
error = %e,
|
||||
"Failed to send rekey msg1"
|
||||
);
|
||||
// The teardown here is already benign — this returns before
|
||||
// `set_rekey_state`, so the cycle simply does not start and
|
||||
// is retried when rekey next comes due, with nothing torn
|
||||
// down and nothing charged to the peer. Only the severity
|
||||
// is wrong for a transport between interfaces, which is a
|
||||
// local and self-clearing condition the presence machine
|
||||
// has already reported.
|
||||
if e.is_transient() {
|
||||
debug!(
|
||||
peer = %self.peer_display_name(node_addr),
|
||||
error = %e,
|
||||
"Deferred rekey msg1: the transport is between interfaces"
|
||||
);
|
||||
} else {
|
||||
warn!(
|
||||
peer = %self.peer_display_name(node_addr),
|
||||
error = %e,
|
||||
"Failed to send rekey msg1"
|
||||
);
|
||||
}
|
||||
let _ = self.index_allocator.free(our_index);
|
||||
self.stats_mut()
|
||||
.record_reject(RejectReason::Handshake(HandshakeReject::BadState));
|
||||
|
||||
+129
-38
@@ -1000,15 +1000,37 @@ impl Node {
|
||||
let connected = self.peers.contains_key(&node_addr);
|
||||
|
||||
if connected {
|
||||
// Active peer: skip a candidate whose path is already the
|
||||
// current, still-fresh one (avoid churning a healthy link).
|
||||
let transport_name = transport.transport_type().name;
|
||||
let peer_addr_candidate =
|
||||
PeerAddress::new(transport_name, remote_addr.to_string());
|
||||
if self.active_peer_candidate_is_fresh_enough_to_skip(
|
||||
&node_addr,
|
||||
std::slice::from_ref(&peer_addr_candidate),
|
||||
) {
|
||||
// Active peer: skip every candidate while the link we
|
||||
// already hold is live — the current path *and* any
|
||||
// alternate one.
|
||||
//
|
||||
// Only the same-path case used to be skipped, which left
|
||||
// the stated intent ("avoid churning a healthy link")
|
||||
// covering exactly the case that could not churn anything.
|
||||
// A peer reachable twice — the ordinary result of two
|
||||
// machines sharing a LAN and a cable, since each beacons on
|
||||
// both — was therefore re-dialled on its alternate path
|
||||
// every discovery tick, forever. Each dial that completed
|
||||
// promoted and displaced the incumbent, so the peer's link
|
||||
// migrated back and forth on a fixed cadence, tearing down
|
||||
// and re-establishing its session each time. Measured on
|
||||
// real hardware: seventeen dials to one peer in fifteen
|
||||
// minutes, alternating wifi and cable, displacing a link
|
||||
// reporting `etx = 1.0` and `loss = 0.0`.
|
||||
//
|
||||
// When that peer is the parent — which the best path
|
||||
// usually is — every migration also switched parents,
|
||||
// invalidating the downstream coordinate cache and
|
||||
// re-announcing to every peer. The cost of the churn was
|
||||
// therefore mesh-wide while the benefit was nil: the link
|
||||
// being replaced was already perfect.
|
||||
//
|
||||
// Failover is unaffected. Liveness is the gate, so a peer
|
||||
// that stops answering goes stale within a heartbeat
|
||||
// interval and every path, alternate included, is dialled
|
||||
// again. What is given up is switching away from a link
|
||||
// that is working, which is not a thing worth doing.
|
||||
if self.active_peer_link_is_live(&node_addr) {
|
||||
continue;
|
||||
}
|
||||
if self.is_connecting_to_peer_on_path(
|
||||
@@ -1617,6 +1639,17 @@ impl Node {
|
||||
self.child_exit_tx = Some(child_exit_tx);
|
||||
self.child_exit_rx = Some(child_exit_rx);
|
||||
|
||||
// Interface-presence channel. Created before `create_transports` so
|
||||
// every interface-bound transport gets the sender at construction and
|
||||
// its very first bind attempt — the one `start_async` makes inline —
|
||||
// is already reportable. A boot race therefore reaches the FSM while
|
||||
// it is still `Starting`, and start-completion health resolves to
|
||||
// `Degraded` on the first publish rather than publishing `Full` and
|
||||
// correcting it a moment later.
|
||||
let (presence_tx, presence_rx) = tokio::sync::mpsc::channel(16);
|
||||
self.transport_presence_tx = Some(presence_tx);
|
||||
self.transport_presence_rx = Some(presence_rx);
|
||||
|
||||
// Initialize transports first (before TUN, before Nostr discovery).
|
||||
// Creation allocates each transport's id; the supervisor FSM authors
|
||||
// the start order over those ids.
|
||||
@@ -1878,12 +1911,20 @@ impl Node {
|
||||
info!(" address: {}", device.address());
|
||||
info!(" mtu: {}", mtu);
|
||||
|
||||
// Calculate max MSS for TCP clamping
|
||||
// Seed the shared MSS ceiling from whatever is bound
|
||||
// right now. Both TUN threads read it live from here
|
||||
// on, so a transport binding or unbinding later moves
|
||||
// the clamp instead of leaving it at this instant's
|
||||
// value — see `crate::upper::tun::MssCeiling`.
|
||||
self.refresh_tun_mss_ceiling();
|
||||
let max_mss = self.tun_mss_ceiling.clone();
|
||||
let effective_mtu = self.effective_ipv6_mtu();
|
||||
let max_mss = effective_mtu.saturating_sub(40).saturating_sub(20); // IPv6 + TCP headers
|
||||
|
||||
info!("effective MTU: {} bytes", effective_mtu);
|
||||
debug!(" max TCP MSS: {} bytes", max_mss);
|
||||
debug!(
|
||||
" max TCP MSS: {} bytes",
|
||||
max_mss.load(std::sync::atomic::Ordering::Relaxed)
|
||||
);
|
||||
|
||||
// On macOS and FreeBSD, create a shutdown pipe. Writing to it
|
||||
// unblocks the reader thread's select() loop without closing
|
||||
@@ -1907,8 +1948,8 @@ impl Node {
|
||||
// Create writer (dups the fd for independent write access).
|
||||
// Pass path_mtu_lookup so inbound SYN-ACK clamp can read
|
||||
// per-destination path MTU learned via discovery.
|
||||
let (writer, tun_tx) =
|
||||
device.create_writer(max_mss, self.path_mtu_lookup.clone())?;
|
||||
let (writer, tun_tx) = device
|
||||
.create_writer(max_mss.clone(), self.path_mtu_lookup.clone())?;
|
||||
|
||||
// Spawn writer thread. On exit it self-reports
|
||||
// `Child::Tun` (sync context → `blocking_send`); TUN
|
||||
@@ -1934,7 +1975,6 @@ impl Node {
|
||||
// self-reports `Child::Tun` on exit (sync context →
|
||||
// `blocking_send`). Exactly one cfg variant compiles,
|
||||
// so the single clone is moved into that closure.
|
||||
let transport_mtu = self.transport_mtu();
|
||||
let path_mtu_lookup = self.path_mtu_lookup.clone();
|
||||
let reader_child_tx = self.child_exit_tx.clone();
|
||||
#[cfg(any(target_os = "macos", target_os = "freebsd"))]
|
||||
@@ -1945,7 +1985,7 @@ impl Node {
|
||||
our_addr,
|
||||
reader_tun_tx,
|
||||
outbound_tx,
|
||||
transport_mtu,
|
||||
max_mss,
|
||||
path_mtu_lookup,
|
||||
shutdown_read_fd,
|
||||
);
|
||||
@@ -1961,7 +2001,7 @@ impl Node {
|
||||
our_addr,
|
||||
reader_tun_tx,
|
||||
outbound_tx,
|
||||
transport_mtu,
|
||||
max_mss,
|
||||
path_mtu_lookup,
|
||||
);
|
||||
if let Some(tx) = &reader_child_tx {
|
||||
@@ -2092,6 +2132,16 @@ impl Node {
|
||||
}
|
||||
};
|
||||
|
||||
// Drain any presence edges this child's start produced *before*
|
||||
// reporting the child itself. An interface-bound transport whose
|
||||
// interface is missing reports absence from inside `start_async`
|
||||
// and then reports `SubstrateUp` (absence is a state, not a start
|
||||
// failure), so ordering the drain first means start-completion
|
||||
// health already knows about the absence when `pending` empties.
|
||||
// Otherwise a boot race publishes `Full` and corrects itself a
|
||||
// moment later, and every consumer sees a spurious transition.
|
||||
let _ = self.drain_transport_presence();
|
||||
|
||||
let feedback_actions = self.supervisor.fsm.step(feedback);
|
||||
for action in &feedback_actions {
|
||||
if let Action::PublishState(ns) = action {
|
||||
@@ -2100,6 +2150,12 @@ impl Node {
|
||||
}
|
||||
}
|
||||
|
||||
// Late edges: a transport that bound after its `SubstrateUp` was
|
||||
// reported, or one that detached during a later child's bring-up.
|
||||
if let Some(ns) = self.drain_transport_presence() {
|
||||
start_outcome = Some(ns);
|
||||
}
|
||||
|
||||
// Seams that never triggered inside the loop: the "Transports
|
||||
// initialized" info! when there was no non-transport child, and the
|
||||
// peer-connect when there was no Tun/Dns child (today it still runs,
|
||||
@@ -2138,8 +2194,10 @@ impl Node {
|
||||
// children. Enumerate them for the operator, then proceed —
|
||||
// a degraded node serves traffic.
|
||||
warn!(
|
||||
degraded_children = ?self.supervisor.fsm.failed(),
|
||||
"Node started DEGRADED: one or more configured optional children failed to start"
|
||||
degraded_children = ?self.supervisor.fsm.degraded_children(),
|
||||
absent_interfaces = ?self.supervisor.fsm.absent(),
|
||||
"Node started DEGRADED: one or more configured optional children failed to \
|
||||
start, or a configured interface is absent"
|
||||
);
|
||||
}
|
||||
_ => {}
|
||||
@@ -2498,6 +2556,50 @@ impl Node {
|
||||
}
|
||||
}
|
||||
|
||||
/// Feed every queued interface-presence edge to the supervisor FSM,
|
||||
/// returning the last [`NodeState`] it asked to publish (if any).
|
||||
///
|
||||
/// Non-blocking: it drains what is already queued and returns. Used during
|
||||
/// bring-up, where the rx_loop's presence arm is not running yet — from
|
||||
/// then on that arm owns the same translation.
|
||||
pub(in crate::node) fn drain_transport_presence(&mut self) -> Option<NodeState> {
|
||||
let mut edges = Vec::new();
|
||||
if let Some(rx) = self.transport_presence_rx.as_mut() {
|
||||
while let Ok(edge) = rx.try_recv() {
|
||||
edges.push(edge);
|
||||
}
|
||||
}
|
||||
|
||||
let mut published = None;
|
||||
let saw_edge = !edges.is_empty();
|
||||
for edge in edges {
|
||||
// Health is policy-filtered; the MTU refresh below is not. See
|
||||
// `saw_edge`.
|
||||
if !edge.health_relevant {
|
||||
continue;
|
||||
}
|
||||
let child = Child::Transport(edge.transport_id);
|
||||
let event = if edge.present {
|
||||
Event::ChildPresent { child }
|
||||
} else {
|
||||
Event::ChildAbsent { child }
|
||||
};
|
||||
for action in self.supervisor.fsm.step(event) {
|
||||
if let Action::PublishState(ns) = action {
|
||||
published = Some(ns);
|
||||
}
|
||||
}
|
||||
}
|
||||
if saw_edge {
|
||||
// A bind or unbind changes which transports are bound, and so the
|
||||
// node's egress MTU floor. During bring-up this runs before the
|
||||
// TUN threads exist, which is exactly when it must: they read the
|
||||
// ceiling this leaves behind.
|
||||
self.refresh_tun_mss_ceiling();
|
||||
}
|
||||
published
|
||||
}
|
||||
|
||||
/// Reconstruct the supervised up-set from observed runtime presence, so the
|
||||
/// FSM authors the teardown order regardless of how the node reached
|
||||
/// `Running`. Worker pools are deliberately excluded: today's teardown never
|
||||
@@ -3425,14 +3527,13 @@ impl Node {
|
||||
candidates
|
||||
}
|
||||
|
||||
pub(in crate::node) fn active_peer_candidate_is_fresh_enough_to_skip(
|
||||
&self,
|
||||
peer_node_addr: &NodeAddr,
|
||||
candidates: &[PeerAddress],
|
||||
) -> bool {
|
||||
if !self.active_peer_matches_any_candidate(peer_node_addr, candidates) {
|
||||
return false;
|
||||
}
|
||||
/// Whether the link we already hold to this peer is answering.
|
||||
///
|
||||
/// The gate on dialling an active peer at all. Phrased as liveness rather
|
||||
/// than as a property of the candidate, because the candidate's path is
|
||||
/// not the question: a live link should not be replaced by *any* path,
|
||||
/// and a dead one should be replaced by whichever path answers.
|
||||
pub(in crate::node) fn active_peer_link_is_live(&self, peer_node_addr: &NodeAddr) -> bool {
|
||||
!self.active_peer_needs_same_path_refresh(peer_node_addr)
|
||||
}
|
||||
|
||||
@@ -3449,17 +3550,7 @@ impl Node {
|
||||
peer.idle_time(Self::now_ms()) > stale_after_ms
|
||||
}
|
||||
|
||||
fn active_peer_matches_any_candidate(
|
||||
&self,
|
||||
peer_node_addr: &NodeAddr,
|
||||
candidates: &[PeerAddress],
|
||||
) -> bool {
|
||||
candidates
|
||||
.iter()
|
||||
.any(|candidate| self.active_peer_matches_candidate(peer_node_addr, candidate))
|
||||
}
|
||||
|
||||
fn active_peer_matches_candidate(
|
||||
pub(in crate::node) fn active_peer_matches_candidate(
|
||||
&self,
|
||||
peer_node_addr: &NodeAddr,
|
||||
candidate: &PeerAddress,
|
||||
|
||||
@@ -67,6 +67,39 @@
|
||||
//! when a task/thread dies at runtime) is **deferred**: start-completion health
|
||||
//! resolution is start-framed, and liveness monitoring is a substantial unbuilt
|
||||
//! mechanism. This commit is start-time health only.
|
||||
//!
|
||||
//! ## Scope: interface presence, and `Degraded` as a level (this commit)
|
||||
//!
|
||||
//! Interface-bound transports are now *sometimes bound*: a transport whose
|
||||
//! interface is missing at start comes up [`Absent`] and binds later, and one
|
||||
//! whose interface goes away at runtime unbinds and rebinds when it returns.
|
||||
//! Two things follow for this machine.
|
||||
//!
|
||||
//! - **A second reason set.** [`Event::ChildAbsent`] / [`Event::ChildPresent`]
|
||||
//! move a child in and out of `absent`, which feeds `Degraded` exactly like
|
||||
//! `failed` does. It is kept separate because it is *reversible* and `failed`
|
||||
//! is not: a child that failed to start stays failed for the bring-up, while
|
||||
//! an absent interface is expected to come back.
|
||||
//! - **`Degraded` is a level, not a latch.** Nothing ever removed from `failed`,
|
||||
//! which was correct while no child could recover — a monotonic set accurately
|
||||
//! described a one-way door. Once recovery exists the assumption inverts: plug
|
||||
//! the WAN back in and the node would stay `Degraded` until the process
|
||||
//! restarted, and `Degraded` would come to mean "something broke at some point
|
||||
//! since boot" rather than "something is broken now".
|
||||
//! [`Self::classify_health`](SupervisorFsm::classify_health) is therefore
|
||||
//! recomputed on every transition **in both directions**.
|
||||
//!
|
||||
//! An absent transport still counts as *up*. It came up — `start_async` returns
|
||||
//! `Ok` with the transport absent — so it does not push a single-transport node
|
||||
//! into the fatal [`FailReason::NoTransports`], which would make a node that
|
||||
//! merely booted before its wifi exit instead of waiting. Absence degrades; it
|
||||
//! never kills.
|
||||
//!
|
||||
//! There is deliberately **no restart action**. Rebinding is owned by the
|
||||
//! transport's own binder task, which is where the file descriptor and the
|
||||
//! presence watcher live; the FSM is told what happened and republishes health.
|
||||
//!
|
||||
//! [`Absent`]: crate::transport::ethernet::Presence::Absent
|
||||
|
||||
use std::collections::HashSet;
|
||||
use std::sync::Arc;
|
||||
@@ -174,6 +207,22 @@ pub(crate) enum Event {
|
||||
/// The child whose task/thread exited.
|
||||
child: Child,
|
||||
},
|
||||
/// An interface-bound child lost its interface — it was never there at
|
||||
/// start, or it went away at runtime. The child stays *up* (the transport
|
||||
/// object survives detach, only its socket goes) but contributes
|
||||
/// `Degraded`. Valid while `Running`; while `Starting` the edge is recorded
|
||||
/// so start-completion health already reflects it.
|
||||
ChildAbsent {
|
||||
/// The child whose interface is absent.
|
||||
child: Child,
|
||||
},
|
||||
/// An interface-bound child's interface came back and it rebound. Clears
|
||||
/// the absence and republishes health, which is how `Degraded` becomes
|
||||
/// reversible.
|
||||
ChildPresent {
|
||||
/// The child whose interface is present again.
|
||||
child: Child,
|
||||
},
|
||||
}
|
||||
|
||||
/// A driver-scheduled timer the supervisor can arm. Only the
|
||||
@@ -320,7 +369,16 @@ pub(crate) struct SupervisorFsm {
|
||||
up: HashSet<Child>,
|
||||
/// Configured children that failed to start during the current bring-up.
|
||||
/// Feeds the `Degraded` health determination when `pending` empties.
|
||||
///
|
||||
/// One-way within a bring-up: a start failure is not recoverable, so
|
||||
/// nothing removes from this set until the next `Start`.
|
||||
failed: HashSet<Child>,
|
||||
/// Children that are up but whose network interface is currently absent.
|
||||
///
|
||||
/// Reversible, unlike [`Self::failed`] — that is the whole reason it is a
|
||||
/// second set rather than more entries in the first. Feeds `Degraded` the
|
||||
/// same way, and empties as interfaces come back.
|
||||
absent: HashSet<Child>,
|
||||
}
|
||||
|
||||
impl SupervisorFsm {
|
||||
@@ -330,6 +388,7 @@ impl SupervisorFsm {
|
||||
state: SupState::Created,
|
||||
up: HashSet::new(),
|
||||
failed: HashSet::new(),
|
||||
absent: HashSet::new(),
|
||||
}
|
||||
}
|
||||
|
||||
@@ -348,6 +407,7 @@ impl SupervisorFsm {
|
||||
},
|
||||
up: up.into_iter().collect(),
|
||||
failed: HashSet::new(),
|
||||
absent: HashSet::new(),
|
||||
}
|
||||
}
|
||||
|
||||
@@ -357,13 +417,26 @@ impl SupervisorFsm {
|
||||
&self.state
|
||||
}
|
||||
|
||||
/// The configured children that failed to start during bring-up. The driver
|
||||
/// reads this on the `Degraded` start outcome to enumerate the degraded
|
||||
/// children in an operator-visible `warn!`.
|
||||
/// The configured children that failed to start during bring-up. Kept for
|
||||
/// the tests that pin the failure-vs-absence split; the driver reports
|
||||
/// [`Self::degraded_children`], which is the union of the two.
|
||||
#[cfg(test)]
|
||||
pub(in crate::node) fn failed(&self) -> &HashSet<Child> {
|
||||
&self.failed
|
||||
}
|
||||
|
||||
/// Children whose interface is currently absent.
|
||||
pub(in crate::node) fn absent(&self) -> &HashSet<Child> {
|
||||
&self.absent
|
||||
}
|
||||
|
||||
/// Every child currently contributing `Degraded` — the ones that failed to
|
||||
/// start plus the ones whose interface is away. This is what an operator
|
||||
/// wants named when the node reports `Degraded`.
|
||||
pub(in crate::node) fn degraded_children(&self) -> HashSet<Child> {
|
||||
self.failed.union(&self.absent).copied().collect()
|
||||
}
|
||||
|
||||
/// Whether the machine is in the bounded-drain window. The driver uses this
|
||||
/// after the rx loop returns to decide between the drain-teardown path and
|
||||
/// the immediate-`stop()` fallback.
|
||||
@@ -398,6 +471,8 @@ impl SupervisorFsm {
|
||||
Event::DrainDeadlineElapsed => self.on_drain_deadline_elapsed(),
|
||||
Event::ChildStopped { child } => self.on_child_stopped(child),
|
||||
Event::ChildExited { child } => self.on_child_exited(child),
|
||||
Event::ChildAbsent { child } => self.on_child_absent(child),
|
||||
Event::ChildPresent { child } => self.on_child_present(child),
|
||||
}
|
||||
}
|
||||
|
||||
@@ -444,6 +519,7 @@ impl SupervisorFsm {
|
||||
|
||||
self.up.clear();
|
||||
self.failed.clear();
|
||||
self.absent.clear();
|
||||
|
||||
// A node with no children at all resolves health immediately. Zero
|
||||
// transports up → `Failed` (this is the behavioral
|
||||
@@ -497,19 +573,30 @@ impl SupervisorFsm {
|
||||
self.classify_health()
|
||||
}
|
||||
|
||||
/// Classify health from the current `up` / `failed` sets and set the
|
||||
/// resulting state, returning the [`NodeState`] the driver should publish.
|
||||
/// Shared by start-completion ([`Self::resolve_start_health`]) and runtime
|
||||
/// child-exit ([`Self::on_child_exited`]):
|
||||
/// Classify health from the current `up` / `failed` / `absent` sets and set
|
||||
/// the resulting state, returning the [`NodeState`] the driver should
|
||||
/// publish. Shared by start-completion ([`Self::resolve_start_health`]),
|
||||
/// runtime child-exit ([`Self::on_child_exited`]) and the presence edges
|
||||
/// ([`Self::on_child_absent`] / [`Self::on_child_present`]):
|
||||
///
|
||||
/// - zero transports up → [`SupState::Failed`] / [`NodeState::Failed`];
|
||||
/// - ≥1 transport up but some child in `failed` → [`Health::Degraded`] /
|
||||
/// [`NodeState::Degraded`];
|
||||
/// - everything up and nothing failed → [`Health::Full`] / [`NodeState::Running`].
|
||||
/// - ≥1 transport up but some child in `failed` or `absent` →
|
||||
/// [`Health::Degraded`] / [`NodeState::Degraded`];
|
||||
/// - everything up, nothing failed, nothing absent → [`Health::Full`] /
|
||||
/// [`NodeState::Running`].
|
||||
///
|
||||
/// Worker-pool failures are captured in `failed` like any other optional
|
||||
/// child, so they contribute `Degraded` (never `Failed`) — the inline crypto
|
||||
/// fallback keeps the node correct without the pools.
|
||||
///
|
||||
/// A transport whose interface is absent is still counted among
|
||||
/// `transports_up`: it came up, it is simply not bound. Excluding it would
|
||||
/// make a single-ethernet node that booted before its wifi resolve to the
|
||||
/// fatal [`FailReason::NoTransports`] and exit — which is the failure this
|
||||
/// whole mechanism exists to remove. Absence degrades; it never kills.
|
||||
///
|
||||
/// Recomputed in **both** directions. This function is the reason
|
||||
/// `Degraded` is a level rather than a latch.
|
||||
fn classify_health(&mut self) -> NodeState {
|
||||
let transports_up = self
|
||||
.up
|
||||
@@ -521,10 +608,10 @@ impl SupervisorFsm {
|
||||
reason: FailReason::NoTransports,
|
||||
};
|
||||
NodeState::Failed
|
||||
} else if !self.failed.is_empty() {
|
||||
} else if !self.failed.is_empty() || !self.absent.is_empty() {
|
||||
self.state = SupState::Running {
|
||||
health: Health::Degraded {
|
||||
reasons: self.failed.clone(),
|
||||
reasons: self.degraded_children(),
|
||||
},
|
||||
};
|
||||
NodeState::Degraded
|
||||
@@ -536,6 +623,53 @@ impl SupervisorFsm {
|
||||
}
|
||||
}
|
||||
|
||||
/// An interface-bound child lost its interface.
|
||||
///
|
||||
/// The child stays in the up-set: the transport object survives detach —
|
||||
/// config, id, statistics and neighbor buffer persist, only the descriptor
|
||||
/// and its loops go. Only the reason set changes.
|
||||
///
|
||||
/// While `Starting` the edge is recorded silently; start-completion health
|
||||
/// picks it up when `pending` empties, so a boot race resolves to
|
||||
/// `Degraded` on the first publish rather than publishing `Full` and
|
||||
/// immediately correcting it. Inert outside `Starting` / `Running`: a
|
||||
/// teardown in flight owns its own bookkeeping.
|
||||
fn on_child_absent(&mut self, child: Child) -> Vec<Action> {
|
||||
match self.state {
|
||||
SupState::Starting { .. } => {
|
||||
self.absent.insert(child);
|
||||
Vec::new()
|
||||
}
|
||||
SupState::Running { .. } => {
|
||||
if !self.absent.insert(child) {
|
||||
return Vec::new();
|
||||
}
|
||||
vec![Action::PublishState(self.classify_health())]
|
||||
}
|
||||
_ => Vec::new(),
|
||||
}
|
||||
}
|
||||
|
||||
/// An interface-bound child's interface came back.
|
||||
///
|
||||
/// Clears the absence and republishes. A child that is not currently
|
||||
/// recorded absent produces nothing — a duplicate edge is not an event.
|
||||
fn on_child_present(&mut self, child: Child) -> Vec<Action> {
|
||||
match self.state {
|
||||
SupState::Starting { .. } => {
|
||||
self.absent.remove(&child);
|
||||
Vec::new()
|
||||
}
|
||||
SupState::Running { .. } => {
|
||||
if !self.absent.remove(&child) {
|
||||
return Vec::new();
|
||||
}
|
||||
vec![Action::PublishState(self.classify_health())]
|
||||
}
|
||||
_ => Vec::new(),
|
||||
}
|
||||
}
|
||||
|
||||
fn on_stop(&mut self) -> Vec<Action> {
|
||||
if !matches!(self.state, SupState::Running { .. }) {
|
||||
return Vec::new();
|
||||
@@ -588,6 +722,7 @@ impl SupervisorFsm {
|
||||
if let SupState::Stopping { pending } = &mut self.state {
|
||||
pending.remove(&child);
|
||||
self.up.remove(&child);
|
||||
self.absent.remove(&child);
|
||||
if pending.is_empty() {
|
||||
self.state = SupState::Stopped;
|
||||
}
|
||||
@@ -617,6 +752,8 @@ impl SupervisorFsm {
|
||||
if !self.up.remove(&child) {
|
||||
return Vec::new();
|
||||
}
|
||||
// An exit supersedes an absence: the child is gone, not waiting.
|
||||
self.absent.remove(&child);
|
||||
self.failed.insert(child);
|
||||
vec![Action::PublishState(self.classify_health())]
|
||||
}
|
||||
@@ -1393,4 +1530,289 @@ mod tests {
|
||||
);
|
||||
assert!(s.failed().is_empty());
|
||||
}
|
||||
|
||||
// ── Interface presence ────────────────────────────────────────────────
|
||||
//
|
||||
// The absence set is separate from `failed` because it is reversible, and
|
||||
// reversibility is what turns `Degraded` from a latch into a level.
|
||||
|
||||
/// Helper: bring a full node up cleanly and leave it `Running{Full}`.
|
||||
fn running_node() -> SupervisorFsm {
|
||||
let mut s = SupervisorFsm::new();
|
||||
s.step(start_full());
|
||||
for child in [
|
||||
Child::Transport(tid(1)),
|
||||
Child::Transport(tid(2)),
|
||||
Child::EncryptWorkers,
|
||||
Child::DecryptWorkers,
|
||||
Child::Nostr,
|
||||
Child::Mdns,
|
||||
Child::Tun,
|
||||
Child::Dns,
|
||||
] {
|
||||
s.step(Event::SubstrateUp { child });
|
||||
}
|
||||
assert_eq!(
|
||||
s.state(),
|
||||
&SupState::Running {
|
||||
health: Health::Full
|
||||
}
|
||||
);
|
||||
s
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn an_absent_interface_degrades_a_running_node() {
|
||||
let mut s = running_node();
|
||||
assert_eq!(
|
||||
s.step(Event::ChildAbsent {
|
||||
child: Child::Transport(tid(1))
|
||||
}),
|
||||
vec![Action::PublishState(NodeState::Degraded)]
|
||||
);
|
||||
assert!(s.absent().contains(&Child::Transport(tid(1))));
|
||||
// Absence is not failure: the two sets stay distinct.
|
||||
assert!(s.failed().is_empty());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_returning_interface_clears_degraded() {
|
||||
// The regression this whole design turns on: plug the WAN back in and
|
||||
// the node must leave `Degraded`, not carry it until the process
|
||||
// restarts. `Degraded` is a level, not a latch.
|
||||
let mut s = running_node();
|
||||
s.step(Event::ChildAbsent {
|
||||
child: Child::Transport(tid(1)),
|
||||
});
|
||||
assert_eq!(
|
||||
s.step(Event::ChildPresent {
|
||||
child: Child::Transport(tid(1))
|
||||
}),
|
||||
vec![Action::PublishState(NodeState::Running)]
|
||||
);
|
||||
assert_eq!(
|
||||
s.state(),
|
||||
&SupState::Running {
|
||||
health: Health::Full
|
||||
}
|
||||
);
|
||||
assert!(s.absent().is_empty());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn health_reflects_the_last_interface_still_away() {
|
||||
// Two interfaces away, one returns: still Degraded. Only the empty
|
||||
// absence set publishes Full.
|
||||
let mut s = running_node();
|
||||
s.step(Event::ChildAbsent {
|
||||
child: Child::Transport(tid(1)),
|
||||
});
|
||||
s.step(Event::ChildAbsent {
|
||||
child: Child::Transport(tid(2)),
|
||||
});
|
||||
assert_eq!(
|
||||
s.step(Event::ChildPresent {
|
||||
child: Child::Transport(tid(1))
|
||||
}),
|
||||
vec![Action::PublishState(NodeState::Degraded)]
|
||||
);
|
||||
assert_eq!(
|
||||
s.step(Event::ChildPresent {
|
||||
child: Child::Transport(tid(2))
|
||||
}),
|
||||
vec![Action::PublishState(NodeState::Running)]
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_duplicate_presence_edge_is_not_an_event() {
|
||||
let mut s = running_node();
|
||||
s.step(Event::ChildAbsent {
|
||||
child: Child::Transport(tid(1)),
|
||||
});
|
||||
assert_eq!(
|
||||
s.step(Event::ChildAbsent {
|
||||
child: Child::Transport(tid(1))
|
||||
}),
|
||||
vec![],
|
||||
"re-reporting the same absence must not republish"
|
||||
);
|
||||
s.step(Event::ChildPresent {
|
||||
child: Child::Transport(tid(1)),
|
||||
});
|
||||
assert_eq!(
|
||||
s.step(Event::ChildPresent {
|
||||
child: Child::Transport(tid(1))
|
||||
}),
|
||||
vec![],
|
||||
"re-reporting the same return must not republish"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn absence_during_bringup_resolves_degraded_on_the_first_publish() {
|
||||
// The boot race. The transport reports absence from inside its own
|
||||
// start, then reports up. Start-completion health must already know,
|
||||
// so the node never publishes `Full` and corrects itself.
|
||||
let mut s = SupervisorFsm::new();
|
||||
s.step(start_full());
|
||||
s.step(Event::ChildAbsent {
|
||||
child: Child::Transport(tid(1)),
|
||||
});
|
||||
for child in [
|
||||
Child::Transport(tid(1)),
|
||||
Child::Transport(tid(2)),
|
||||
Child::EncryptWorkers,
|
||||
Child::DecryptWorkers,
|
||||
Child::Nostr,
|
||||
Child::Mdns,
|
||||
Child::Tun,
|
||||
] {
|
||||
assert_eq!(s.step(Event::SubstrateUp { child }), vec![]);
|
||||
}
|
||||
assert_eq!(
|
||||
s.step(Event::SubstrateUp { child: Child::Dns }),
|
||||
vec![Action::PublishState(NodeState::Degraded)]
|
||||
);
|
||||
let mut reasons = HashSet::new();
|
||||
reasons.insert(Child::Transport(tid(1)));
|
||||
assert_eq!(
|
||||
s.state(),
|
||||
&SupState::Running {
|
||||
health: Health::Degraded { reasons }
|
||||
}
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_lone_absent_transport_degrades_rather_than_fails() {
|
||||
// A single-ethernet node that boots before its wifi. The transport
|
||||
// came up — absence is a state, not a start failure — so it counts
|
||||
// among the transports up and the node serves. Resolving to `Failed`
|
||||
// here would make the daemon exit on the very race this mechanism
|
||||
// exists to survive.
|
||||
let mut s = SupervisorFsm::new();
|
||||
s.step(Event::Start {
|
||||
transports: vec![tid(1)],
|
||||
encrypt_workers: false,
|
||||
decrypt_workers: false,
|
||||
nostr: false,
|
||||
mdns: false,
|
||||
tun: false,
|
||||
dns: false,
|
||||
});
|
||||
s.step(Event::ChildAbsent {
|
||||
child: Child::Transport(tid(1)),
|
||||
});
|
||||
assert_eq!(
|
||||
s.step(Event::SubstrateUp {
|
||||
child: Child::Transport(tid(1))
|
||||
}),
|
||||
vec![Action::PublishState(NodeState::Degraded)]
|
||||
);
|
||||
assert!(
|
||||
!matches!(s.state(), SupState::Failed { .. }),
|
||||
"an absent interface must never be the fatal no-transports case"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn an_exit_supersedes_an_absence() {
|
||||
// A transport that is away and then exits is gone, not waiting: it
|
||||
// leaves the up-set and moves from `absent` to `failed`, so a later
|
||||
// spurious `ChildPresent` cannot resurrect it into `Full`.
|
||||
let mut s = running_node();
|
||||
s.step(Event::ChildAbsent {
|
||||
child: Child::Transport(tid(1)),
|
||||
});
|
||||
s.step(Event::ChildExited {
|
||||
child: Child::Transport(tid(1)),
|
||||
});
|
||||
assert!(s.absent().is_empty());
|
||||
assert!(s.failed().contains(&Child::Transport(tid(1))));
|
||||
assert_eq!(
|
||||
s.step(Event::ChildPresent {
|
||||
child: Child::Transport(tid(1))
|
||||
}),
|
||||
vec![]
|
||||
);
|
||||
assert!(matches!(
|
||||
s.state(),
|
||||
SupState::Running {
|
||||
health: Health::Degraded { .. }
|
||||
}
|
||||
));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn presence_edges_are_inert_outside_starting_and_running() {
|
||||
// Teardown and drain own their own bookkeeping; a late edge from a
|
||||
// binder that has not noticed the stop must not author actions.
|
||||
let mut created = SupervisorFsm::new();
|
||||
assert_eq!(
|
||||
created.step(Event::ChildAbsent {
|
||||
child: Child::Transport(tid(1))
|
||||
}),
|
||||
vec![]
|
||||
);
|
||||
|
||||
let mut draining = running_node();
|
||||
draining.step(Event::Drain { deadline_ms: 1 });
|
||||
assert_eq!(
|
||||
draining.step(Event::ChildAbsent {
|
||||
child: Child::Transport(tid(1))
|
||||
}),
|
||||
vec![]
|
||||
);
|
||||
assert!(
|
||||
draining.is_draining(),
|
||||
"a presence edge must not end a drain"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_new_start_clears_the_previous_absences() {
|
||||
let mut s = running_node();
|
||||
s.step(Event::ChildAbsent {
|
||||
child: Child::Transport(tid(1)),
|
||||
});
|
||||
s.step(Event::Stop);
|
||||
for child in [
|
||||
Child::Dns,
|
||||
Child::Nostr,
|
||||
Child::Mdns,
|
||||
Child::Transport(tid(1)),
|
||||
Child::Transport(tid(2)),
|
||||
Child::Tun,
|
||||
] {
|
||||
s.step(Event::ChildStopped { child });
|
||||
}
|
||||
s.step(start_full());
|
||||
assert!(s.absent().is_empty());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn degraded_children_is_the_union_of_both_reason_sets() {
|
||||
let mut s = SupervisorFsm::new();
|
||||
s.step(start_full());
|
||||
s.step(Event::ChildAbsent {
|
||||
child: Child::Transport(tid(1)),
|
||||
});
|
||||
s.step(Event::SubstrateFailed { child: Child::Mdns });
|
||||
for child in [
|
||||
Child::Transport(tid(1)),
|
||||
Child::Transport(tid(2)),
|
||||
Child::EncryptWorkers,
|
||||
Child::DecryptWorkers,
|
||||
Child::Nostr,
|
||||
Child::Tun,
|
||||
Child::Dns,
|
||||
] {
|
||||
s.step(Event::SubstrateUp { child });
|
||||
}
|
||||
let degraded = s.degraded_children();
|
||||
assert!(degraded.contains(&Child::Mdns));
|
||||
assert!(degraded.contains(&Child::Transport(tid(1))));
|
||||
assert_eq!(degraded.len(), 2);
|
||||
}
|
||||
}
|
||||
|
||||
+156
-4
@@ -154,6 +154,17 @@ pub enum NodeError {
|
||||
#[error("send failed to {node_addr}: {reason}")]
|
||||
SendFailed { node_addr: NodeAddr, reason: String },
|
||||
|
||||
/// A send refused by a condition that is expected to clear on its own.
|
||||
///
|
||||
/// Distinct from [`Self::SendFailed`] because the right response differs:
|
||||
/// the state built around the send — a half-finished handshake, a route,
|
||||
/// a queued packet — is worth keeping across a transient refusal and
|
||||
/// worth tearing down after a terminal one. Carries the transport's own
|
||||
/// classification ([`TransportError::is_transient`]) rather than a
|
||||
/// re-derivation of it.
|
||||
#[error("send to {node_addr} unavailable: {reason}")]
|
||||
SendUnavailable { node_addr: NodeAddr, reason: String },
|
||||
|
||||
#[error("mtu exceeded forwarding to {node_addr}: packet {packet_size} > mtu {mtu}")]
|
||||
MtuExceeded {
|
||||
node_addr: NodeAddr,
|
||||
@@ -186,6 +197,32 @@ pub enum NodeError {
|
||||
NoOperationalTransports,
|
||||
}
|
||||
|
||||
impl Node {
|
||||
/// Test-only: place a transport into the node's map directly.
|
||||
///
|
||||
/// The snapshot tests live in `crate::control` and so cannot reach the
|
||||
/// private `transports` field, but the interface-presence block they need
|
||||
/// to pin only exists on a real interface-bound transport. Mirrors
|
||||
/// `isolate_peer_acl_for_test`: a narrow hook, so the fixture stays honest
|
||||
/// rather than the snapshot being hand-authored JSON that nothing
|
||||
/// produces.
|
||||
#[cfg(all(test, any(target_os = "linux", target_os = "macos")))]
|
||||
pub(crate) fn insert_transport_for_test(&mut self, id: TransportId, handle: TransportHandle) {
|
||||
self.transports.insert(id, handle);
|
||||
}
|
||||
}
|
||||
|
||||
impl NodeError {
|
||||
/// Whether this failure is expected to clear on its own.
|
||||
///
|
||||
/// Mirrors [`TransportError::is_transient`] across the node boundary, so
|
||||
/// a caller holding a `NodeError` can ask the same question a caller
|
||||
/// holding a `TransportError` can, and get the same answer.
|
||||
pub fn is_transient(&self) -> bool {
|
||||
matches!(self, Self::SendUnavailable { .. })
|
||||
}
|
||||
}
|
||||
|
||||
/// Node operational state.
|
||||
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
|
||||
pub enum NodeState {
|
||||
@@ -394,6 +431,16 @@ pub struct Node {
|
||||
/// SYN/SYN-ACK clamp can use the smaller of the local-egress floor
|
||||
/// and the learned per-destination path MTU.
|
||||
path_mtu_lookup: crate::upper::tun::PathMtuLookup,
|
||||
/// Node-global TCP MSS ceiling, shared live with the TUN reader and writer
|
||||
/// threads and recomputed whenever the set of *bound* transports changes.
|
||||
///
|
||||
/// Sits beside `path_mtu_lookup` because it answers the other half of the
|
||||
/// same question at the same moment: that map supplies the per-destination
|
||||
/// ceiling, this supplies the local-egress one, and the clamp takes the
|
||||
/// smaller. Both have to be read live — a transport that binds after start
|
||||
/// can be the narrow one, and one that unbinds can be the reason the node
|
||||
/// was clamped at all.
|
||||
tun_mss_ceiling: crate::upper::tun::MssCeiling,
|
||||
/// Which transport last supplied a *link seed* into `path_mtu_lookup`,
|
||||
/// per destination.
|
||||
///
|
||||
@@ -438,6 +485,20 @@ pub struct Node {
|
||||
/// rx_loop select arm that feeds `Event::ChildExited` to the supervisor FSM.
|
||||
child_exit_rx: Option<tokio::sync::mpsc::Receiver<crate::node::lifecycle::supervisor::Child>>,
|
||||
|
||||
// === Interface Presence Channel ===
|
||||
/// Sender half of the interface-presence channel, cloned into every
|
||||
/// interface-bound transport so its binder task can report attach and
|
||||
/// detach. Held on `self` for the rx_loop's lifetime as the keep-alive
|
||||
/// sender, exactly like [`Self::child_exit_tx`].
|
||||
///
|
||||
/// Separate from the child-exit channel because presence is *reversible*:
|
||||
/// an exit is one-way, an interface comes back.
|
||||
transport_presence_tx: Option<crate::transport::PresenceTx>,
|
||||
/// Receiver half of the interface-presence channel, `take()`-en by the
|
||||
/// rx_loop select arm that feeds `Event::ChildAbsent` / `Event::ChildPresent`
|
||||
/// to the supervisor FSM.
|
||||
transport_presence_rx: Option<crate::transport::PresenceRx>,
|
||||
|
||||
// === Per-Peer Control Machines ===
|
||||
/// Per-peer lifecycle control FSMs, keyed by the stable `LinkId` that spans
|
||||
/// the handshake→active lifetime. Each machine owns its handshake crypto
|
||||
@@ -840,6 +901,8 @@ impl Node {
|
||||
packet_rx: None,
|
||||
child_exit_tx: None,
|
||||
child_exit_rx: None,
|
||||
transport_presence_tx: None,
|
||||
transport_presence_rx: None,
|
||||
peer_machines: HashMap::new(),
|
||||
peer_timers: HashMap::new(),
|
||||
peers: HashMap::new(),
|
||||
@@ -908,6 +971,12 @@ impl Node {
|
||||
peer_acl,
|
||||
host_map,
|
||||
path_mtu_lookup: Arc::new(std::sync::RwLock::new(HashMap::new())),
|
||||
// Seeded at the IPv6 minimum, which is what `transport_mtu()`
|
||||
// itself falls back to when nothing is bound. Refreshed before
|
||||
// the TUN threads start and on every change to the bound set.
|
||||
tun_mss_ceiling: Arc::new(std::sync::atomic::AtomicU16::new(
|
||||
crate::upper::icmp::mss_ceiling(crate::upper::tun::IPV6_MIN_MTU),
|
||||
)),
|
||||
path_mtu_seeded_by: Arc::new(std::sync::RwLock::new(HashMap::new())),
|
||||
#[cfg(unix)]
|
||||
decrypt_registered_sessions: std::collections::HashSet::new(),
|
||||
@@ -1012,6 +1081,8 @@ impl Node {
|
||||
packet_rx: None,
|
||||
child_exit_tx: None,
|
||||
child_exit_rx: None,
|
||||
transport_presence_tx: None,
|
||||
transport_presence_rx: None,
|
||||
peer_machines: HashMap::new(),
|
||||
peer_timers: HashMap::new(),
|
||||
peers: HashMap::new(),
|
||||
@@ -1077,6 +1148,12 @@ impl Node {
|
||||
peer_acl,
|
||||
host_map,
|
||||
path_mtu_lookup: Arc::new(std::sync::RwLock::new(HashMap::new())),
|
||||
// Seeded at the IPv6 minimum, which is what `transport_mtu()`
|
||||
// itself falls back to when nothing is bound. Refreshed before
|
||||
// the TUN threads start and on every change to the bound set.
|
||||
tun_mss_ceiling: Arc::new(std::sync::atomic::AtomicU16::new(
|
||||
crate::upper::icmp::mss_ceiling(crate::upper::tun::IPV6_MIN_MTU),
|
||||
)),
|
||||
path_mtu_seeded_by: Arc::new(std::sync::RwLock::new(HashMap::new())),
|
||||
#[cfg(unix)]
|
||||
decrypt_registered_sessions: std::collections::HashSet::new(),
|
||||
@@ -1133,7 +1210,13 @@ impl Node {
|
||||
.collect();
|
||||
for (name, eth_config) in eth_instances {
|
||||
let transport_id = self.allocate_transport_id();
|
||||
let eth = EthernetTransport::new(transport_id, name, eth_config, packet_tx.clone());
|
||||
let mut eth =
|
||||
EthernetTransport::new(transport_id, name, eth_config, packet_tx.clone());
|
||||
// The binder task reports attach and detach here, so node
|
||||
// health tracks the interface in both directions.
|
||||
if let Some(tx) = self.transport_presence_tx.clone() {
|
||||
eth.set_presence_tx(tx);
|
||||
}
|
||||
transports.push(TransportHandle::Ethernet(eth));
|
||||
}
|
||||
}
|
||||
@@ -1445,6 +1528,46 @@ impl Node {
|
||||
crate::upper::icmp::effective_ipv6_mtu(self.transport_mtu())
|
||||
}
|
||||
|
||||
/// The TCP MSS ceiling the TUN threads are currently clamping to.
|
||||
#[cfg(test)]
|
||||
pub(crate) fn tun_mss_ceiling(&self) -> u16 {
|
||||
self.tun_mss_ceiling
|
||||
.load(std::sync::atomic::Ordering::Relaxed)
|
||||
}
|
||||
|
||||
/// Recompute the shared TUN MSS ceiling from the currently bound
|
||||
/// transports, and log it if it moved.
|
||||
///
|
||||
/// Called wherever the bound set can change — the presence edges that
|
||||
/// bind and unbind an interface-bound transport, and a child exiting —
|
||||
/// so the clamp the TUN threads apply keeps agreeing with the
|
||||
/// `effective_ipv6_mtu` this node reports in `show_status`.
|
||||
///
|
||||
/// Moves in **both** directions, deliberately. A narrow interface
|
||||
/// appearing has to tighten the ceiling or the clamp is wrong for
|
||||
/// traffic that will egress over it; that same interface going away has
|
||||
/// to release it, or unplugging a low-MTU adapter leaves the node
|
||||
/// over-clamped until it restarts. It is the same argument that makes
|
||||
/// `Degraded` a level rather than a latch: nothing here is one-way once
|
||||
/// a transport can come back.
|
||||
///
|
||||
/// Existing flows are not re-clamped — MSS is negotiated per connection
|
||||
/// at SYN time, so a change applies to connections opened after it.
|
||||
pub(crate) fn refresh_tun_mss_ceiling(&self) {
|
||||
use std::sync::atomic::Ordering;
|
||||
|
||||
let ceiling = crate::upper::icmp::mss_ceiling(self.transport_mtu());
|
||||
let previous = self.tun_mss_ceiling.swap(ceiling, Ordering::Relaxed);
|
||||
if previous != ceiling {
|
||||
tracing::info!(
|
||||
previous_max_mss = previous,
|
||||
max_mss = ceiling,
|
||||
effective_ipv6_mtu = self.effective_ipv6_mtu(),
|
||||
"Node egress MTU changed; TCP MSS ceiling updated for new connections"
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
/// Get the transport MTU governing the global TUN-boundary MSS clamp.
|
||||
///
|
||||
/// Returns the **minimum** MTU across all operational transports, or
|
||||
@@ -1460,10 +1583,17 @@ impl Node {
|
||||
/// to vary across HashMap iteration order + async-startup race) makes
|
||||
/// the clamp deterministic across daemon restarts.
|
||||
pub fn transport_mtu(&self) -> u16 {
|
||||
// `is_bound`, not `is_operational`. An interface-bound transport is
|
||||
// "operational" from the moment it starts, whether or not its
|
||||
// interface exists — so filtering on that let a transport whose
|
||||
// interface has never appeared clamp the whole node's IPv6 MTU to a
|
||||
// number derived from hardware that is not present. Before dynamic
|
||||
// binding a transport that could not bind was never inserted here at
|
||||
// all, so the distinction did not exist to get wrong.
|
||||
let min_operational = self
|
||||
.transports
|
||||
.values()
|
||||
.filter(|h| h.is_operational())
|
||||
.filter(|h| h.is_bound())
|
||||
.map(|h| h.mtu())
|
||||
.min();
|
||||
if let Some(mtu) = min_operational {
|
||||
@@ -2374,8 +2504,13 @@ impl Node {
|
||||
.collect();
|
||||
|
||||
// --- transports (show_transports) ---
|
||||
let transport_rows: Vec<snap::TransportRow> = self
|
||||
.transport_ids()
|
||||
// Ascending id, matching `show_transports`; see the note there. The
|
||||
// off-loop renderer reads this table verbatim, so the two paths would
|
||||
// otherwise disagree about ordering as well as being arbitrary.
|
||||
let mut transport_ids: Vec<_> = self.transport_ids().copied().collect();
|
||||
transport_ids.sort_by_key(|id| id.as_u32());
|
||||
let transport_rows: Vec<snap::TransportRow> = transport_ids
|
||||
.iter()
|
||||
.map(|id| {
|
||||
let handle = self.get_transport(id).unwrap();
|
||||
snap::TransportRow {
|
||||
@@ -2391,6 +2526,15 @@ impl Node {
|
||||
.tor_monitoring()
|
||||
.map(|m| serde_json::to_value(&m).unwrap_or_default()),
|
||||
stats: handle.transport_stats(),
|
||||
interface: handle.interface_presence().map(|p| snap::InterfaceRow {
|
||||
name: handle.interface_name().unwrap_or_default().to_string(),
|
||||
presence: p.presence,
|
||||
carrier: p.carrier,
|
||||
policy: p.policy,
|
||||
since_secs: p.since_secs,
|
||||
binds: p.binds,
|
||||
failed_attempts: p.failed_attempts,
|
||||
}),
|
||||
}
|
||||
})
|
||||
.collect();
|
||||
@@ -3817,6 +3961,14 @@ impl Node {
|
||||
packet_size,
|
||||
mtu,
|
||||
},
|
||||
// Preserve the transport's own classification instead of
|
||||
// flattening every non-MTU failure into one string. A caller
|
||||
// that wants to keep its half-built state across an interface
|
||||
// flap can only do that if the distinction survives to it.
|
||||
other if other.is_transient() => NodeError::SendUnavailable {
|
||||
node_addr: *node_addr,
|
||||
reason: format!("transport send: {}", other),
|
||||
},
|
||||
other => NodeError::SendFailed {
|
||||
node_addr: *node_addr,
|
||||
reason: format!("transport send: {}", other),
|
||||
|
||||
@@ -0,0 +1,203 @@
|
||||
//! Per-peer connected UDP sockets against the in-line decrypt path.
|
||||
//!
|
||||
//! A `connect(2)`-ed UDP socket is pinned to one 5-tuple. When the peer
|
||||
//! moves, the address the socket was opened against is gone, but the
|
||||
//! socket is still installed and the send path prefers it over the
|
||||
//! wildcard listen socket. `ActivePeer::set_current_addr` returns
|
||||
//! whether the address actually changed precisely so the caller can
|
||||
//! drop the stale socket, and both post-decrypt paths have to act on
|
||||
//! that return: the decrypt-worker completion path
|
||||
//! (`process_authentic_fmp_plaintext`) and the in-line one
|
||||
//! (`handle_encrypted_frame`). These tests cover the in-line path,
|
||||
//! which is the one the worker path's own coverage does not reach.
|
||||
//!
|
||||
//! `bool` carries no `#[must_use]`, so discarding the return here is
|
||||
//! silent under `-D warnings`; the assertions below are what makes the
|
||||
//! difference between binding it and dropping it observable.
|
||||
|
||||
use super::*;
|
||||
use crate::noise::NoiseSession;
|
||||
use crate::proto::fmp::wire::{build_encrypted, build_established_header, prepend_inner_header};
|
||||
|
||||
/// The address `seed_completed_connection` promotes a peer on, and so
|
||||
/// the peer's `current_addr` before anything rotates it.
|
||||
const PROMOTED_ADDR: &str = "127.0.0.1:5000";
|
||||
|
||||
/// The address the peer is made to move to.
|
||||
const ROAMED_ADDR: &str = "127.0.0.1:5001";
|
||||
|
||||
/// Build a promoted peer and hand back the far side's Noise session.
|
||||
///
|
||||
/// [`seed_completed_connection`] runs every leg of the handshake and
|
||||
/// then drops the responder, so nothing outside it can produce a frame
|
||||
/// the node will actually authenticate. This is the same seeding with
|
||||
/// the responder's session kept, which is what lets these tests reach
|
||||
/// the post-decrypt side effects rather than stopping at the AEAD.
|
||||
///
|
||||
/// Returns the node, the peer's `NodeAddr`, the session index an
|
||||
/// inbound frame must name to be routed to that peer, and the session
|
||||
/// to encrypt those frames with.
|
||||
fn promoted_peer_with_the_far_side_session(
|
||||
transport_id: TransportId,
|
||||
) -> (Node, NodeAddr, SessionIndex, NoiseSession) {
|
||||
let mut node = make_node();
|
||||
let link_id = LinkId::new(1);
|
||||
|
||||
let peer_identity_full = Identity::generate();
|
||||
// from_pubkey_full, not from_pubkey: the ECDH needs the parity bit.
|
||||
let peer_identity = PeerIdentity::from_pubkey_full(peer_identity_full.pubkey_full());
|
||||
|
||||
let our_index = node.index_allocator.allocate().unwrap();
|
||||
node.seed_handshake_machine(
|
||||
HandshakeSeed::outbound(link_id, peer_identity, 1_000)
|
||||
.with_our_index(our_index)
|
||||
.with_their_index(SessionIndex::new(42))
|
||||
.with_transport_id(transport_id)
|
||||
.with_source_addr(TransportAddr::from_string(PROMOTED_ADDR)),
|
||||
)
|
||||
.unwrap();
|
||||
|
||||
let our_keypair = node.identity().keypair();
|
||||
let startup_epoch = node.startup_epoch();
|
||||
let msg1 = node
|
||||
.peer_machines
|
||||
.get_mut(&link_id)
|
||||
.unwrap()
|
||||
.start_handshake(our_keypair, startup_epoch, 1_000)
|
||||
.unwrap();
|
||||
|
||||
let mut responder = inbound_leg(LinkId::new(999), 1_000);
|
||||
let mut responder_epoch = [0u8; 8];
|
||||
rand::Rng::fill_bytes(&mut rand::rng(), &mut responder_epoch);
|
||||
let msg2 = responder
|
||||
.receive_handshake_init(
|
||||
peer_identity_full.keypair(),
|
||||
responder_epoch,
|
||||
&msg1,
|
||||
None,
|
||||
1_000,
|
||||
)
|
||||
.unwrap();
|
||||
|
||||
// XX, so msg2 does not finish the responder: it holds no session until it
|
||||
// has processed msg3. The line this came from ran IK, where the responder
|
||||
// was done at msg2 and the third leg did not exist.
|
||||
let (msg3, _negotiation) = node
|
||||
.peer_machines
|
||||
.get_mut(&link_id)
|
||||
.unwrap()
|
||||
.complete_handshake(&msg2, None, 1_000)
|
||||
.unwrap();
|
||||
|
||||
responder
|
||||
.complete_handshake_msg3(&msg3, 1_000)
|
||||
.expect("the responder must accept a genuine msg3");
|
||||
|
||||
let far_side_session = responder
|
||||
.take_session()
|
||||
.expect("the responder holds a session once it has processed msg3");
|
||||
|
||||
node.promote_connection(link_id, peer_identity, 2_000)
|
||||
.unwrap();
|
||||
let node_addr = *peer_identity.node_addr();
|
||||
let our_index = node
|
||||
.get_peer(&node_addr)
|
||||
.and_then(|p| p.our_index())
|
||||
.expect("a promoted peer carries the index it was allocated");
|
||||
|
||||
(node, node_addr, our_index, far_side_session)
|
||||
}
|
||||
|
||||
/// Encrypt one well-formed established frame from the far side.
|
||||
///
|
||||
/// The link message is a heartbeat (`0x51`), which the dispatcher
|
||||
/// handles as a no-op — these tests are about the side effects that run
|
||||
/// before the dispatch, so the message must not have any of its own.
|
||||
fn far_side_frame(session: &mut NoiseSession, receiver_idx: SessionIndex) -> Vec<u8> {
|
||||
let inner = prepend_inner_header(0, &[0x51]);
|
||||
let counter = session.current_send_counter();
|
||||
let header = build_established_header(receiver_idx, counter, 0, inner.len() as u16);
|
||||
let ciphertext = session.encrypt_with_aad(&inner, &header).unwrap();
|
||||
build_encrypted(&header, &ciphertext)
|
||||
}
|
||||
|
||||
/// **The defect.**
|
||||
///
|
||||
/// The in-line decrypt path called `set_current_addr` as a bare
|
||||
/// statement and dropped its return, so a peer could roam, have its
|
||||
/// `current_addr` updated, and keep a connected socket pinned to the
|
||||
/// 5-tuple it had just left. The send path prefers that socket while it
|
||||
/// is installed, so every frame after the move goes out to an address
|
||||
/// the peer is no longer at.
|
||||
#[cfg(any(target_os = "linux", target_os = "macos"))]
|
||||
#[tokio::test]
|
||||
async fn a_peer_that_roams_loses_the_connected_socket_pinned_to_the_address_it_left() {
|
||||
let transport_id = TransportId::new(1);
|
||||
let (mut node, node_addr, our_index, mut far_side) =
|
||||
promoted_peer_with_the_far_side_session(transport_id);
|
||||
|
||||
install_connected_udp(&mut node, &node_addr, transport_id);
|
||||
assert!(
|
||||
node.get_peer(&node_addr).unwrap().connected_udp().is_some(),
|
||||
"precondition: the peer holds a connected socket before it moves"
|
||||
);
|
||||
|
||||
let frame = far_side_frame(&mut far_side, our_index);
|
||||
node.handle_encrypted_frame(ReceivedPacket::new(
|
||||
transport_id,
|
||||
TransportAddr::from_string(ROAMED_ADDR),
|
||||
frame,
|
||||
))
|
||||
.await;
|
||||
|
||||
let peer = node
|
||||
.get_peer(&node_addr)
|
||||
.expect("the peer survives an authentic frame");
|
||||
assert_eq!(
|
||||
peer.current_addr(),
|
||||
Some(&TransportAddr::from_string(ROAMED_ADDR)),
|
||||
"precondition for the assertion below: the frame must have been \
|
||||
authenticated and the rotation recorded, or the test proves nothing"
|
||||
);
|
||||
assert!(
|
||||
peer.connected_udp().is_none(),
|
||||
"a socket pinned to the address the peer has left must not survive \
|
||||
the rotation"
|
||||
);
|
||||
}
|
||||
|
||||
/// **The healthy path.**
|
||||
///
|
||||
/// A frame from the address the peer is already on changes nothing, so
|
||||
/// the connected socket has to stay. A fix that cleared unconditionally
|
||||
/// would tear down and reopen the socket on every single frame.
|
||||
#[cfg(any(target_os = "linux", target_os = "macos"))]
|
||||
#[tokio::test]
|
||||
async fn a_frame_from_the_address_the_peer_is_already_on_keeps_the_connected_socket() {
|
||||
let transport_id = TransportId::new(1);
|
||||
let (mut node, node_addr, our_index, mut far_side) =
|
||||
promoted_peer_with_the_far_side_session(transport_id);
|
||||
|
||||
install_connected_udp(&mut node, &node_addr, transport_id);
|
||||
|
||||
let frame = far_side_frame(&mut far_side, our_index);
|
||||
node.handle_encrypted_frame(ReceivedPacket::new(
|
||||
transport_id,
|
||||
TransportAddr::from_string(PROMOTED_ADDR),
|
||||
frame,
|
||||
))
|
||||
.await;
|
||||
|
||||
let peer = node
|
||||
.get_peer(&node_addr)
|
||||
.expect("the peer survives an authentic frame");
|
||||
assert_eq!(
|
||||
peer.current_addr(),
|
||||
Some(&TransportAddr::from_string(PROMOTED_ADDR)),
|
||||
"the peer has not moved"
|
||||
);
|
||||
assert!(
|
||||
peer.connected_udp().is_some(),
|
||||
"a frame from the address already in use must leave the socket alone"
|
||||
);
|
||||
}
|
||||
@@ -375,3 +375,74 @@ async fn a_failing_peer_is_retried_after_the_gap_and_not_before() {
|
||||
|
||||
cleanup_nodes(&mut nodes).await;
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Reaping peers when their interface goes away
|
||||
//
|
||||
// The detach edge is both earlier and more certain than inactivity, so it is
|
||||
// the better trigger for withdrawing what the interface carried. These drive
|
||||
// the same real two-node peering the liveness tests use, because a peer only
|
||||
// reaches the established context the reap acts on by actually peering.
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/// A peer reachable only through an interface that has gone is withdrawn on
|
||||
/// the detach edge, without waiting out `link_dead_timeout_secs`.
|
||||
///
|
||||
/// Note what is *not* set here: the link-dead timeout keeps its default, and
|
||||
/// no time is advanced. The peer is live by every liveness measure and is
|
||||
/// still withdrawn, because the transport under it is gone — which is the
|
||||
/// whole distinction this adds.
|
||||
#[tokio::test]
|
||||
async fn a_detached_transport_withdraws_the_peers_that_needed_it() {
|
||||
let mut nodes = run_tree_test(2, &[(0, 1)], false).await;
|
||||
verify_tree_convergence(&nodes);
|
||||
|
||||
let addr_1 = *nodes[1].node.node_addr();
|
||||
let transport_id = nodes[0]
|
||||
.node
|
||||
.get_peer(&addr_1)
|
||||
.expect("peer present")
|
||||
.transport_id()
|
||||
.expect("an established peer names its transport");
|
||||
|
||||
let reaped = nodes[0].node.reap_peers_on_transport(transport_id).await;
|
||||
|
||||
assert_eq!(reaped, 1);
|
||||
assert!(
|
||||
nodes[0].node.get_peer(&addr_1).is_none(),
|
||||
"a peer must not outlive the interface it was reachable through"
|
||||
);
|
||||
|
||||
cleanup_nodes(&mut nodes).await;
|
||||
}
|
||||
|
||||
/// The reap is scoped to the transport that detached.
|
||||
///
|
||||
/// The failure this guards is the one that would make the feature worse than
|
||||
/// the defect: an interface going away must not withdraw the peers that were
|
||||
/// never reachable through it, which on a mesh router is most of them.
|
||||
#[tokio::test]
|
||||
async fn a_detached_transport_leaves_other_transports_peers_alone() {
|
||||
let mut nodes = run_tree_test(2, &[(0, 1)], false).await;
|
||||
verify_tree_convergence(&nodes);
|
||||
|
||||
let addr_1 = *nodes[1].node.node_addr();
|
||||
let peer_transport = nodes[0]
|
||||
.node
|
||||
.get_peer(&addr_1)
|
||||
.expect("peer present")
|
||||
.transport_id()
|
||||
.expect("an established peer names its transport");
|
||||
|
||||
// A transport this peer was never reachable through.
|
||||
let unrelated = TransportId::new(peer_transport.as_u32() + 100);
|
||||
let reaped = nodes[0].node.reap_peers_on_transport(unrelated).await;
|
||||
|
||||
assert_eq!(reaped, 0, "an unrelated transport withdraws nothing");
|
||||
assert!(
|
||||
nodes[0].node.get_peer(&addr_1).is_some(),
|
||||
"a peer on a healthy transport must survive another one detaching"
|
||||
);
|
||||
|
||||
cleanup_nodes(&mut nodes).await;
|
||||
}
|
||||
|
||||
@@ -11,6 +11,7 @@ mod ble;
|
||||
mod bloom;
|
||||
mod bloom_poison;
|
||||
mod bootstrap;
|
||||
mod connected_udp;
|
||||
mod control;
|
||||
mod decrypt_failure;
|
||||
mod disconnect;
|
||||
@@ -57,6 +58,39 @@ pub(super) fn make_node_with(config: Config) -> Node {
|
||||
Node::new(config).unwrap()
|
||||
}
|
||||
|
||||
/// Install a real `connect()`-ed UDP socket on a peer, the way the tick-driven
|
||||
/// activation in `dataplane::connected_udp` does.
|
||||
///
|
||||
/// The socket is opened against the loopback discard port: nothing is ever sent
|
||||
/// through it, and the callers only care whether the handle is still installed
|
||||
/// afterwards.
|
||||
#[cfg(any(target_os = "linux", target_os = "macos"))]
|
||||
pub(super) fn install_connected_udp(
|
||||
node: &mut Node,
|
||||
addr: &NodeAddr,
|
||||
transport_id: crate::transport::TransportId,
|
||||
) {
|
||||
let local: std::net::SocketAddr = "0.0.0.0:0".parse().unwrap();
|
||||
let peer_sa: std::net::SocketAddr = "127.0.0.1:9".parse().unwrap();
|
||||
|
||||
let owned = crate::transport::udp::open_connected_fd(local, peer_sa, 65_536, 65_536)
|
||||
.expect("open a connected UDP socket");
|
||||
let bound = crate::transport::udp::ConnectedPeerSocket::from_fd(owned, peer_sa, local);
|
||||
let socket = std::sync::Arc::new(bound);
|
||||
let (packet_tx, _packet_rx) = packet_channel(8);
|
||||
let drain = crate::transport::udp::PeerRecvDrain::spawn(
|
||||
socket.clone(),
|
||||
transport_id,
|
||||
peer_sa,
|
||||
packet_tx,
|
||||
)
|
||||
.expect("spawn the peer recv drain");
|
||||
|
||||
node.get_peer_mut(addr)
|
||||
.expect("peer present")
|
||||
.set_connected_udp(socket, drain);
|
||||
}
|
||||
|
||||
/// Build a test node with an explicit `max_peers` limit (replaces the removed
|
||||
/// `set_max_peers` setter; resource limits are immutable post-construction).
|
||||
pub(super) fn make_node_with_max_peers(max_peers: usize) -> Node {
|
||||
|
||||
@@ -31,35 +31,6 @@ fn identity_of(nodes: &[TestNode], j: usize) -> PeerIdentity {
|
||||
PeerIdentity::from_pubkey_full(nodes[j].node.identity().pubkey_full())
|
||||
}
|
||||
|
||||
/// Install a real `connect()`-ed UDP socket on a peer, the way the tick-driven
|
||||
/// activation does.
|
||||
///
|
||||
/// The socket is opened against a discard port on loopback: nothing is ever
|
||||
/// sent through it, and the test only cares whether the handle survives a
|
||||
/// medium change.
|
||||
#[cfg(any(target_os = "linux", target_os = "macos"))]
|
||||
fn install_connected_udp(node: &mut Node, addr: &NodeAddr, transport_id: TransportId) {
|
||||
let local: std::net::SocketAddr = "0.0.0.0:0".parse().unwrap();
|
||||
let peer_sa: std::net::SocketAddr = "127.0.0.1:9".parse().unwrap();
|
||||
|
||||
let owned = crate::transport::udp::open_connected_fd(local, peer_sa, 65_536, 65_536)
|
||||
.expect("open a connected UDP socket");
|
||||
let bound = crate::transport::udp::ConnectedPeerSocket::from_fd(owned, peer_sa, local);
|
||||
let socket = std::sync::Arc::new(bound);
|
||||
let (packet_tx, _packet_rx) = crate::transport::packet_channel(8);
|
||||
let drain = crate::transport::udp::PeerRecvDrain::spawn(
|
||||
socket.clone(),
|
||||
transport_id,
|
||||
peer_sa,
|
||||
packet_tx,
|
||||
)
|
||||
.expect("spawn the peer recv drain");
|
||||
|
||||
node.get_peer_mut(addr)
|
||||
.expect("peer present")
|
||||
.set_connected_udp(socket, drain);
|
||||
}
|
||||
|
||||
/// **The defect this feature exists for.**
|
||||
///
|
||||
/// Established UDP peers get a per-peer `connect()`-ed socket. `open_connected_fd`
|
||||
|
||||
+172
-16
@@ -1258,7 +1258,7 @@ async fn test_try_peer_addresses_skips_connecting_peer() {
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn active_peer_same_path_discovery_skips_fresh_peer() {
|
||||
fn a_peer_heard_from_within_the_heartbeat_interval_has_a_live_link() {
|
||||
let mut node = make_node();
|
||||
let peer_full = Identity::generate();
|
||||
let peer_identity = PeerIdentity::from_pubkey_full(peer_full.pubkey_full());
|
||||
@@ -1268,16 +1268,14 @@ fn active_peer_same_path_discovery_skips_fresh_peer() {
|
||||
let mut active_peer = ActivePeer::new(peer_identity, LinkId::new(7), Node::now_ms());
|
||||
active_peer.set_current_addr(transport_id, current_addr.clone());
|
||||
node.peers.insert(peer_node_addr, active_peer);
|
||||
let candidate = crate::config::PeerAddress::new("udp", "127.0.0.1:9");
|
||||
|
||||
assert!(node.active_peer_candidate_is_fresh_enough_to_skip(
|
||||
&peer_node_addr,
|
||||
std::slice::from_ref(&candidate),
|
||||
));
|
||||
// A link heard from just now is live, so discovery must not dial this
|
||||
// peer at all — on this path or on any other.
|
||||
assert!(node.active_peer_link_is_live(&peer_node_addr));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn active_peer_same_path_discovery_refreshes_stale_peer() {
|
||||
fn a_peer_quiet_past_the_heartbeat_interval_no_longer_has_a_live_link() {
|
||||
let mut node = make_node();
|
||||
let peer_full = Identity::generate();
|
||||
let peer_identity = PeerIdentity::from_pubkey_full(peer_full.pubkey_full());
|
||||
@@ -1294,12 +1292,62 @@ fn active_peer_same_path_discovery_refreshes_stale_peer() {
|
||||
let mut active_peer = ActivePeer::new(peer_identity, LinkId::new(7), stale_at);
|
||||
active_peer.set_current_addr(transport_id, current_addr.clone());
|
||||
node.peers.insert(peer_node_addr, active_peer);
|
||||
let candidate = crate::config::PeerAddress::new("udp", "127.0.0.1:9");
|
||||
|
||||
assert!(!node.active_peer_candidate_is_fresh_enough_to_skip(
|
||||
&peer_node_addr,
|
||||
std::slice::from_ref(&candidate),
|
||||
));
|
||||
// Gone quiet past the heartbeat interval: every path is dialable again,
|
||||
// which is what keeps failover working now that a live link is never
|
||||
// displaced.
|
||||
assert!(!node.active_peer_link_is_live(&peer_node_addr));
|
||||
}
|
||||
|
||||
/// A bootstrap-held peer is never its own configured candidate.
|
||||
///
|
||||
/// `adopt_established_traversal` refuses a peer that is already connected, so
|
||||
/// an adopted NAT-traversal transport is the way *on* to a traversed path and
|
||||
/// not the way off. Now that beacon discovery asks only whether the link it
|
||||
/// holds is answering, the configured-peer refresh is the one automatic
|
||||
/// off-ramp left, and it works only because a bootstrap-held peer is refused
|
||||
/// as a match for its own address: a configured `udp` address can be
|
||||
/// byte-identical to the traversal's remote address and the transport kinds
|
||||
/// match, so without this carve-out `has_alternative` would be false and the
|
||||
/// peer could never be moved off the traversed socket at all.
|
||||
#[test]
|
||||
fn a_bootstrap_held_peer_is_never_its_own_configured_candidate() {
|
||||
let mut node = make_node();
|
||||
let peer_full = Identity::generate();
|
||||
let peer_identity = PeerIdentity::from_pubkey_full(peer_full.pubkey_full());
|
||||
let peer_node_addr = *peer_identity.node_addr();
|
||||
let npub = peer_identity.npub();
|
||||
let transport_id = TransportId::new(1);
|
||||
let current_addr = TransportAddr::from_string("203.0.113.5:41234");
|
||||
|
||||
let (packet_tx, _packet_rx) = packet_channel(8);
|
||||
let udp = UdpTransport::new(
|
||||
transport_id,
|
||||
Some("main".to_string()),
|
||||
crate::config::UdpConfig {
|
||||
bind_addr: Some("127.0.0.1:0".to_string()),
|
||||
..Default::default()
|
||||
},
|
||||
packet_tx,
|
||||
);
|
||||
node.transports
|
||||
.insert(transport_id, TransportHandle::Udp(udp));
|
||||
|
||||
let mut active_peer = ActivePeer::new(peer_identity, LinkId::new(7), Node::now_ms());
|
||||
active_peer.set_current_addr(transport_id, current_addr);
|
||||
node.peers.insert(peer_node_addr, active_peer);
|
||||
node.supervisor
|
||||
.nostr_rendezvous
|
||||
.insert_bootstrap_transport(transport_id, npub);
|
||||
|
||||
assert!(
|
||||
!node.active_peer_matches_candidate(
|
||||
&peer_node_addr,
|
||||
&crate::config::PeerAddress::new("udp", "203.0.113.5:41234")
|
||||
),
|
||||
"a bootstrap-held peer must not count its own traversal address as its \
|
||||
current path, or the configured-peer refresh can never migrate it off"
|
||||
);
|
||||
}
|
||||
|
||||
/// An instance-qualified candidate is the peer's *current* path only when it
|
||||
@@ -1341,12 +1389,11 @@ async fn an_instance_qualified_candidate_matches_only_its_own_instance() {
|
||||
active_peer.set_current_addr(main_id, TransportAddr::from_string("127.0.0.1:9"));
|
||||
node.peers.insert(peer_node_addr, active_peer);
|
||||
|
||||
// Path matching, which still gates the *configured-peer* refresh even
|
||||
// though beacon discovery now gates on liveness alone.
|
||||
let matches = |transport: &str| {
|
||||
let candidate = crate::config::PeerAddress::new(transport, "127.0.0.1:9");
|
||||
node.active_peer_candidate_is_fresh_enough_to_skip(
|
||||
&peer_node_addr,
|
||||
std::slice::from_ref(&candidate),
|
||||
)
|
||||
node.active_peer_matches_candidate(&peer_node_addr, &candidate)
|
||||
};
|
||||
|
||||
assert!(
|
||||
@@ -1845,6 +1892,115 @@ async fn test_transport_mtu_returns_min_across_operational() {
|
||||
}
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn the_tun_mss_ceiling_follows_a_transport_arriving_and_leaving() {
|
||||
// The TUN reader and writer used to be handed a `u16` computed once at
|
||||
// spawn. Every other consumer of `transport_mtu()` reads it live, so once
|
||||
// a transport could bind minutes after start the daemon reported one
|
||||
// effective MTU in `show_status` and clamped to another.
|
||||
//
|
||||
// Both directions. A narrow transport arriving has to tighten the ceiling
|
||||
// or traffic egressing over it is clamped too loose; the same transport
|
||||
// leaving has to release it, or unplugging a low-MTU adapter leaves the
|
||||
// node over-clamped until it restarts.
|
||||
let mut node = make_node();
|
||||
let (packet_tx, packet_rx) = packet_channel(64);
|
||||
node.supervisor.packet_tx = Some(packet_tx);
|
||||
node.packet_rx = Some(packet_rx);
|
||||
|
||||
let wide = make_udp_transport_with_mtu(1, 1452).await;
|
||||
node.transports.insert(TransportId::new(1), wide);
|
||||
node.refresh_tun_mss_ceiling();
|
||||
let wide_ceiling = node.tun_mss_ceiling();
|
||||
assert_eq!(
|
||||
wide_ceiling,
|
||||
crate::upper::icmp::mss_ceiling(1452),
|
||||
"the seeded ceiling must match the only bound transport"
|
||||
);
|
||||
|
||||
// A narrower transport arrives after the TUN threads would already be
|
||||
// running. The shared ceiling has to tighten.
|
||||
let narrow = make_udp_transport_with_mtu(2, 1280).await;
|
||||
node.transports.insert(TransportId::new(2), narrow);
|
||||
node.refresh_tun_mss_ceiling();
|
||||
let narrow_ceiling = node.tun_mss_ceiling();
|
||||
assert_eq!(narrow_ceiling, crate::upper::icmp::mss_ceiling(1280));
|
||||
assert!(
|
||||
narrow_ceiling < wide_ceiling,
|
||||
"a narrower transport must tighten the clamp, not be ignored"
|
||||
);
|
||||
assert_eq!(
|
||||
narrow_ceiling,
|
||||
crate::upper::icmp::mss_ceiling(node.transport_mtu()),
|
||||
"the clamp and the reported MTU must not disagree"
|
||||
);
|
||||
|
||||
// ...and leaving has to release it again.
|
||||
if let Some(mut gone) = node.transports.remove(&TransportId::new(2)) {
|
||||
gone.stop().await.ok();
|
||||
}
|
||||
node.refresh_tun_mss_ceiling();
|
||||
assert_eq!(
|
||||
node.tun_mss_ceiling(),
|
||||
wide_ceiling,
|
||||
"the ceiling must rise again when the narrow transport goes away"
|
||||
);
|
||||
|
||||
for transport in node.transports.values_mut() {
|
||||
transport.stop().await.ok();
|
||||
}
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn a_presence_edge_refreshes_the_tun_mss_ceiling_without_being_asked() {
|
||||
// The one above pins the arithmetic; this pins the wiring. A presence
|
||||
// edge arriving has to refresh the ceiling on its own — if the refresh is
|
||||
// dropped from the edge handlers the value silently stops tracking, which
|
||||
// is the defect in its original form.
|
||||
let mut node = make_node();
|
||||
let (packet_tx, packet_rx) = packet_channel(64);
|
||||
node.supervisor.packet_tx = Some(packet_tx);
|
||||
node.packet_rx = Some(packet_rx);
|
||||
|
||||
let (presence_tx, presence_rx) = tokio::sync::mpsc::channel(4);
|
||||
node.transport_presence_tx = Some(presence_tx.clone());
|
||||
node.transport_presence_rx = Some(presence_rx);
|
||||
|
||||
// Nothing bound: the conservative seed.
|
||||
node.refresh_tun_mss_ceiling();
|
||||
let seeded = node.tun_mss_ceiling();
|
||||
assert_eq!(seeded, crate::upper::icmp::mss_ceiling(1280));
|
||||
|
||||
// A wide transport appears, and an edge announces it. No explicit
|
||||
// refresh call here — draining the edge is the whole trigger.
|
||||
let wide = make_udp_transport_with_mtu(1, 1452).await;
|
||||
node.transports.insert(TransportId::new(1), wide);
|
||||
presence_tx
|
||||
.send(crate::transport::TransportPresence {
|
||||
transport_id: TransportId::new(1),
|
||||
present: true,
|
||||
health_relevant: true,
|
||||
})
|
||||
.await
|
||||
.expect("presence edge queued");
|
||||
node.drain_transport_presence();
|
||||
|
||||
assert_eq!(
|
||||
node.tun_mss_ceiling(),
|
||||
crate::upper::icmp::mss_ceiling(1452),
|
||||
"draining a presence edge must refresh the ceiling on its own"
|
||||
);
|
||||
assert_ne!(
|
||||
node.tun_mss_ceiling(),
|
||||
seeded,
|
||||
"the ceiling stayed at its seed, so the edge did not refresh it"
|
||||
);
|
||||
|
||||
for transport in node.transports.values_mut() {
|
||||
transport.stop().await.ok();
|
||||
}
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn test_transport_mtu_fallback_when_no_operational_transports() {
|
||||
// No transports configured at all → falls back to 1280 (IPv6 minimum).
|
||||
|
||||
@@ -9,6 +9,120 @@ use crate::transport::TransportError;
|
||||
/// Broadcast MAC address.
|
||||
pub const ETHERNET_BROADCAST: [u8; 6] = [0xff; 6];
|
||||
|
||||
/// Whether the named interface exists and is administratively up.
|
||||
///
|
||||
/// Presence is `IFF_UP` — the interface exists and the operator has enabled
|
||||
/// it — and deliberately **not** `IFF_RUNNING`.
|
||||
///
|
||||
/// Carrier is a different question from bindability, and only the second one
|
||||
/// belongs in a bind gate. An `AF_PACKET` socket on a carrier-less bridge is
|
||||
/// perfectly valid and starts carrying traffic the instant a member port comes
|
||||
/// up, with no rebind: the socket outlives the carrier. Gating on `IFF_RUNNING`
|
||||
/// bought nothing and cost three things —
|
||||
///
|
||||
/// - `br-lan` on a router with nothing plugged into its LAN ports is `UP` with
|
||||
/// `NO-CARRIER`, so a perfectly healthy wifi-only router reported `Degraded`
|
||||
/// forever;
|
||||
/// - every carrier flap the socket would have survived became an unbind /
|
||||
/// rebind cycle, which is churn the presence machine then has to damp;
|
||||
/// - an 802.11s mesh interface that reports `RUNNING` only once it has peered
|
||||
/// cannot peer, because peering needs beacons, which need a bound socket,
|
||||
/// which the gate refuses. A deadlock reachable on shipped hardware.
|
||||
///
|
||||
/// The signal `IFF_RUNNING` does carry — "is anything plugged in" — is not
|
||||
/// lost; it is reported alongside presence by [`interface_carrier`] and
|
||||
/// surfaced in `show_transports`, where an operator can read it without it
|
||||
/// steering the daemon.
|
||||
///
|
||||
/// `getifaddrs` rather than an `SIOCGIFFLAGS` ioctl: it needs no socket, so
|
||||
/// the presence watcher can poll before any file descriptor exists, and it is
|
||||
/// spelled the same on Linux and the BSDs.
|
||||
#[cfg(unix)]
|
||||
pub fn interface_present(interface: &str) -> bool {
|
||||
interface_present_probe(interface).unwrap_or(false)
|
||||
}
|
||||
|
||||
/// [`interface_present`], keeping "the probe failed" distinct from "absent".
|
||||
///
|
||||
/// `None` means the kernel would not answer. A caller deciding whether to
|
||||
/// *bind* can treat that as absence and retry on the next tick, which is what
|
||||
/// [`interface_present`] does. A caller deciding whether to *unbind* must not:
|
||||
/// see [`interface_has_flags`].
|
||||
#[cfg(unix)]
|
||||
pub fn interface_present_probe(interface: &str) -> Option<bool> {
|
||||
interface_has_flags(interface, libc::IFF_UP as u32)
|
||||
}
|
||||
|
||||
/// The kernel's index for the named interface, or `None` if it does not exist.
|
||||
///
|
||||
/// A name is not a device, and neither is a name that is still there. Both
|
||||
/// backends bind by index — `AF_PACKET` stores `sll_ifindex`, and a BPF
|
||||
/// descriptor follows the device it was attached to — so an interface deleted
|
||||
/// and recreated under the same name leaves the socket attached to a device
|
||||
/// that no longer exists while the *name* resolves perfectly well. Comparing
|
||||
/// the live index against the one captured at bind is what tells those apart.
|
||||
#[cfg(unix)]
|
||||
pub fn interface_index(interface: &str) -> Option<u32> {
|
||||
let c_name = std::ffi::CString::new(interface).ok()?;
|
||||
// Cheaper than `getifaddrs`: one syscall, no allocation, no walk.
|
||||
match unsafe { libc::if_nametoindex(c_name.as_ptr()) } {
|
||||
0 => None,
|
||||
idx => Some(idx),
|
||||
}
|
||||
}
|
||||
|
||||
/// Whether the named interface currently has carrier (`IFF_RUNNING`).
|
||||
///
|
||||
/// Reported, never acted on — see [`interface_present`]. `false` for an
|
||||
/// interface that does not exist, which keeps "no carrier" and "no interface"
|
||||
/// from being told apart here; presence answers that.
|
||||
#[cfg(unix)]
|
||||
pub fn interface_carrier(interface: &str) -> bool {
|
||||
// Report-only, so a probe failure reads the same as no carrier.
|
||||
interface_has_flags(interface, (libc::IFF_UP | libc::IFF_RUNNING) as u32).unwrap_or(false)
|
||||
}
|
||||
|
||||
/// Whether the named interface exists and has every flag in `wanted` set, or
|
||||
/// `None` if the question could not be asked.
|
||||
///
|
||||
/// The `None` matters. `getifaddrs` is a netlink dump on Linux and it does
|
||||
/// fail for reasons that have nothing to do with the interface — `ENOBUFS`
|
||||
/// under memory pressure or a busy netlink socket, `EMFILE`/`ENFILE` under fd
|
||||
/// exhaustion, since it opens a socket of its own. Answering `false` there
|
||||
/// reports a present interface as gone, and a caller holding a live binding
|
||||
/// would tear a working socket down over a transient syscall failure. Callers
|
||||
/// that can tell the two apart should.
|
||||
#[cfg(unix)]
|
||||
fn interface_has_flags(interface: &str, wanted: u32) -> Option<bool> {
|
||||
let Ok(c_name) = std::ffi::CString::new(interface) else {
|
||||
// An interior NUL is not a probe failure — no such interface can
|
||||
// exist, and no retry will change that.
|
||||
return Some(false);
|
||||
};
|
||||
|
||||
let mut addrs: *mut libc::ifaddrs = std::ptr::null_mut();
|
||||
if unsafe { libc::getifaddrs(&mut addrs) } != 0 {
|
||||
return None;
|
||||
}
|
||||
|
||||
let mut matched = false;
|
||||
let mut cur = addrs;
|
||||
while !cur.is_null() {
|
||||
let entry = unsafe { &*cur };
|
||||
if !entry.ifa_name.is_null()
|
||||
&& unsafe { libc::strcmp(entry.ifa_name, c_name.as_ptr()) } == 0
|
||||
&& entry.ifa_flags & wanted == wanted
|
||||
{
|
||||
matched = true;
|
||||
break;
|
||||
}
|
||||
cur = entry.ifa_next;
|
||||
}
|
||||
|
||||
unsafe { libc::freeifaddrs(addrs) };
|
||||
Some(matched)
|
||||
}
|
||||
|
||||
// Platform-specific PacketSocket implementation.
|
||||
#[cfg(target_os = "linux")]
|
||||
#[path = "io_linux.rs"]
|
||||
@@ -193,6 +307,12 @@ mod async_impl {
|
||||
/// A received frame: (payload, source_mac).
|
||||
type Frame = (Vec<u8>, [u8; 6]);
|
||||
|
||||
/// Consecutive failed BPF reads before the reader thread gives up.
|
||||
///
|
||||
/// Mirrors the receive loop's own error threshold: the point is not to
|
||||
/// tolerate errors but to end the task so the binder can rebind.
|
||||
const READ_ERROR_EXIT_THRESHOLD: u32 = 5;
|
||||
|
||||
pub struct AsyncPacketSocket {
|
||||
inner: Arc<PacketSocket>,
|
||||
/// `None` once shutdown has taken the receiver, which is what makes
|
||||
@@ -219,6 +339,9 @@ mod async_impl {
|
||||
let mut parse_buf = vec![0u8; bpf_buflen];
|
||||
let mut parse_offset: usize = 0;
|
||||
let mut parse_len: usize = 0;
|
||||
// Consecutive failed reads, to bound a socket whose
|
||||
// interface went away underneath it.
|
||||
let mut read_errors: u32 = 0;
|
||||
let nfds = bpf_fd.max(shutdown_fd) + 1;
|
||||
|
||||
loop {
|
||||
@@ -286,11 +409,31 @@ mod async_impl {
|
||||
if err.raw_os_error() == Some(libc::EBADF) {
|
||||
break;
|
||||
}
|
||||
if err.kind() == std::io::ErrorKind::Interrupted {
|
||||
continue;
|
||||
}
|
||||
}
|
||||
// Anything else — `ENXIO` is the one that matters,
|
||||
// which is what BPF answers once the interface it
|
||||
// was attached to is torn away — used to loop here
|
||||
// forever. That mattered beyond the spin: the
|
||||
// binder's detach check asks whether this thread is
|
||||
// still running, so a thread that never returns
|
||||
// reports a dead socket as a live one, and the
|
||||
// transport sits `present` and deaf until the name
|
||||
// or index happens to change too. Give up after a
|
||||
// streak and let the return close the channel,
|
||||
// which fails `recv_from`, which ends the tokio
|
||||
// task the binder is actually watching.
|
||||
read_errors += 1;
|
||||
if read_errors >= READ_ERROR_EXIT_THRESHOLD {
|
||||
break;
|
||||
}
|
||||
parse_len = 0;
|
||||
parse_offset = 0;
|
||||
continue;
|
||||
}
|
||||
read_errors = 0;
|
||||
parse_len = ret as usize;
|
||||
parse_offset = 0;
|
||||
}
|
||||
|
||||
@@ -59,6 +59,14 @@ impl PacketSocket {
|
||||
if ret < 0 {
|
||||
let err = std::io::Error::last_os_error();
|
||||
unsafe { libc::close(fd) };
|
||||
// The interface can disappear between the index lookup and the
|
||||
// bind. That is absence arriving a few microseconds late, not a
|
||||
// configuration fault, so it reports as absence.
|
||||
if matches!(err.raw_os_error(), Some(libc::ENODEV) | Some(libc::ENXIO)) {
|
||||
return Err(TransportError::InterfaceUnavailable {
|
||||
interface: interface.to_string(),
|
||||
});
|
||||
}
|
||||
return Err(TransportError::StartFailed(format!(
|
||||
"bind(AF_PACKET, {}) failed: {}",
|
||||
interface, err
|
||||
@@ -231,11 +239,11 @@ fn get_if_index(_fd: RawFd, interface: &str) -> Result<i32, TransportError> {
|
||||
|
||||
let idx = unsafe { libc::if_nametoindex(c_name.as_ptr()) };
|
||||
if idx == 0 {
|
||||
return Err(TransportError::StartFailed(format!(
|
||||
"interface not found: {} ({})",
|
||||
interface,
|
||||
std::io::Error::last_os_error()
|
||||
)));
|
||||
// Absence, not a fault: the caller's presence watcher rebinds when the
|
||||
// interface shows up. See `TransportError::InterfaceUnavailable`.
|
||||
return Err(TransportError::InterfaceUnavailable {
|
||||
interface: interface.to_string(),
|
||||
});
|
||||
}
|
||||
Ok(idx as i32)
|
||||
}
|
||||
|
||||
@@ -393,10 +393,18 @@ fn bind_to_interface(fd: RawFd, interface: &str) -> Result<(), TransportError> {
|
||||
|
||||
let ret = unsafe { libc::ioctl(fd, BIOCSETIF, ifreq.as_ptr()) };
|
||||
if ret < 0 {
|
||||
let err = std::io::Error::last_os_error();
|
||||
// BIOCSETIF answers ENXIO for an interface that is not there. The
|
||||
// interface can also vanish between the index lookup and this ioctl,
|
||||
// so absence is reported as absence rather than as a bind fault.
|
||||
if matches!(err.raw_os_error(), Some(libc::ENXIO) | Some(libc::ENODEV)) {
|
||||
return Err(TransportError::InterfaceUnavailable {
|
||||
interface: interface.to_string(),
|
||||
});
|
||||
}
|
||||
return Err(TransportError::StartFailed(format!(
|
||||
"BIOCSETIF({}) failed: {}",
|
||||
interface,
|
||||
std::io::Error::last_os_error()
|
||||
interface, err
|
||||
)));
|
||||
}
|
||||
Ok(())
|
||||
@@ -498,11 +506,11 @@ fn get_if_index(interface: &str) -> Result<i32, TransportError> {
|
||||
|
||||
let idx = unsafe { libc::if_nametoindex(c_name.as_ptr()) };
|
||||
if idx == 0 {
|
||||
return Err(TransportError::StartFailed(format!(
|
||||
"interface not found: {} ({})",
|
||||
interface,
|
||||
std::io::Error::last_os_error()
|
||||
)));
|
||||
// Absence, not a fault — see the Linux twin and
|
||||
// `TransportError::InterfaceUnavailable`.
|
||||
return Err(TransportError::InterfaceUnavailable {
|
||||
interface: interface.to_string(),
|
||||
});
|
||||
}
|
||||
Ok(idx as i32)
|
||||
}
|
||||
|
||||
+2075
-186
File diff suppressed because it is too large
Load Diff
@@ -25,6 +25,23 @@ pub mod ethernet;
|
||||
#[cfg(unix)]
|
||||
pub(crate) mod watcher;
|
||||
|
||||
/// Presence lifecycle for a transport bound to a local resource that can
|
||||
/// disappear and come back: the phase machine, the absence policy, and the
|
||||
/// damping that keeps a flapping resource from flapping node health with it.
|
||||
///
|
||||
/// Transport-agnostic on purpose. Only the *probe* — "is my thing there, and
|
||||
/// is it still the same one?" — is specific to what is bound, and that stays
|
||||
/// with the transport that knows how to ask.
|
||||
///
|
||||
/// Crate-internal on purpose, for the same reason as `watcher` above: it is a
|
||||
/// mechanism the crate's own transports share, not a surface an embedder
|
||||
/// builds against. `ethernet` re-exports the two types it used to own, so the
|
||||
/// published path stays `transport::ethernet::{AbsencePolicy, Presence}`.
|
||||
/// Gated with the one transport that binds through it today; widen the gate
|
||||
/// when a second binder arrives.
|
||||
#[cfg(any(target_os = "linux", target_os = "macos"))]
|
||||
pub(crate) mod presence;
|
||||
|
||||
#[cfg(ble_available)]
|
||||
pub mod ble;
|
||||
|
||||
@@ -112,6 +129,62 @@ pub fn packet_channel(buffer: usize) -> (PacketTx, PacketRx) {
|
||||
tokio::sync::mpsc::channel(buffer)
|
||||
}
|
||||
|
||||
/// Operator-visible interface presence, rendered by `show_transports`.
|
||||
///
|
||||
/// Worth as much as the retry itself. The original boot-race bug was expensive
|
||||
/// precisely because the 802.11s peer link formed regardless of the daemon, so
|
||||
/// nothing an operator could see said the node was deaf.
|
||||
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
|
||||
pub struct InterfacePresence {
|
||||
/// `absent`, `binding`, or `present`.
|
||||
pub presence: &'static str,
|
||||
/// Whether the interface currently has carrier (`IFF_RUNNING`).
|
||||
///
|
||||
/// Reported, never acted on. Presence is `IFF_UP`, because binding does
|
||||
/// not need carrier and a socket outlives a carrier flap — but "is
|
||||
/// anything plugged in" is still what an operator wants to know when a
|
||||
/// bound transport is carrying nothing, so it is reported here instead of
|
||||
/// steering the daemon.
|
||||
pub carrier: bool,
|
||||
/// `required` or `optional`.
|
||||
pub policy: &'static str,
|
||||
/// How long the current phase has been held.
|
||||
pub since_secs: u64,
|
||||
/// Successful binds since the transport was created (`1` after a clean
|
||||
/// start; more means it has rebound).
|
||||
pub binds: u64,
|
||||
/// Failed bind attempts since the last successful bind.
|
||||
pub failed_attempts: u32,
|
||||
}
|
||||
|
||||
/// A presence edge published by an interface-bound transport.
|
||||
///
|
||||
/// Absence and return are the same transition seen from two sides, so one
|
||||
/// event type carries both: `present: false` on detach (including a start
|
||||
/// where the interface was never there), `present: true` on every successful
|
||||
/// bind after the first observation.
|
||||
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
|
||||
pub struct TransportPresence {
|
||||
/// The transport whose interface changed presence.
|
||||
pub transport_id: TransportId,
|
||||
/// Whether the interface is now bound.
|
||||
pub present: bool,
|
||||
/// Whether this edge should move node health.
|
||||
///
|
||||
/// `false` for an `optional` interface, whose absence is normal and must
|
||||
/// not take the node off `Full`. The edge is still published, because the
|
||||
/// bound set changed either way and the node's egress MTU floor is derived
|
||||
/// from it — health and MTU are two different questions riding one
|
||||
/// channel, and only the first one is policy-filtered.
|
||||
pub health_relevant: bool,
|
||||
}
|
||||
|
||||
/// Channel sender for transport presence edges.
|
||||
pub type PresenceTx = tokio::sync::mpsc::Sender<TransportPresence>;
|
||||
|
||||
/// Channel receiver for transport presence edges.
|
||||
pub type PresenceRx = tokio::sync::mpsc::Receiver<TransportPresence>;
|
||||
|
||||
// ============================================================================
|
||||
// Errors
|
||||
// ============================================================================
|
||||
@@ -128,6 +201,20 @@ pub enum TransportError {
|
||||
#[error("transport failed to start: {0}")]
|
||||
StartFailed(String),
|
||||
|
||||
/// The named network interface is not usable right now: it does not exist,
|
||||
/// or it exists but is administratively down (no `IFF_UP`).
|
||||
///
|
||||
/// Distinct from [`TransportError::StartFailed`] because absence is a
|
||||
/// *state*, not a fault. Interface-bound transports treat it as "not bound
|
||||
/// yet" and keep a presence watcher running; a `StartFailed` carrying the
|
||||
/// same text could not be told apart from a typo'd interface name or a
|
||||
/// missing capability.
|
||||
#[error("interface unavailable: {interface}")]
|
||||
InterfaceUnavailable {
|
||||
/// The configured interface name.
|
||||
interface: String,
|
||||
},
|
||||
|
||||
#[error("transport shutdown failed: {0}")]
|
||||
ShutdownFailed(String),
|
||||
|
||||
@@ -159,6 +246,49 @@ pub enum TransportError {
|
||||
Io(#[from] std::io::Error),
|
||||
}
|
||||
|
||||
impl TransportError {
|
||||
/// Whether this failure is expected to clear on its own.
|
||||
///
|
||||
/// The distinction callers need is not *what* went wrong but whether
|
||||
/// waiting fixes it. A transient failure means the operation was refused
|
||||
/// by a condition the daemon is already working to resolve, so the state
|
||||
/// built up around it — a half-finished handshake, a route, a queued
|
||||
/// packet — is worth keeping. A terminal one means the state is worth
|
||||
/// tearing down.
|
||||
///
|
||||
/// This lives here, on the error, rather than being re-derived at each
|
||||
/// call site: `InterfaceUnavailable` used to be flattened into a
|
||||
/// formatted string on its way out of the transport layer, so every
|
||||
/// caller downstream saw a generic send failure and could only treat a
|
||||
/// two-second interface flap exactly as it treated a permanent fault.
|
||||
///
|
||||
/// Deliberately narrow. [`Self::Timeout`] and [`Self::ConnectionRefused`]
|
||||
/// are *not* transient here: they describe a remote that did not answer,
|
||||
/// which is a statement about the peer rather than about this node's
|
||||
/// ability to transmit, and the existing retry paths for them already sit
|
||||
/// at a different layer.
|
||||
pub fn is_transient(&self) -> bool {
|
||||
match self {
|
||||
// The interface is absent or mid-rebind. The binder is polling for
|
||||
// it and will bind it the moment it returns.
|
||||
Self::InterfaceUnavailable { .. } => true,
|
||||
Self::NotStarted
|
||||
| Self::AlreadyStarted
|
||||
| Self::StartFailed(_)
|
||||
| Self::ShutdownFailed(_)
|
||||
| Self::LinkFailed(_)
|
||||
| Self::SendFailed(_)
|
||||
| Self::RecvFailed(_)
|
||||
| Self::InvalidAddress(_)
|
||||
| Self::MtuExceeded { .. }
|
||||
| Self::Timeout
|
||||
| Self::ConnectionRefused
|
||||
| Self::NotSupported(_)
|
||||
| Self::Io(_) => false,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// ============================================================================
|
||||
// Transport Type Metadata
|
||||
// ============================================================================
|
||||
@@ -874,6 +1004,52 @@ impl TransportHandle {
|
||||
}
|
||||
}
|
||||
|
||||
/// Interface presence for interface-bound transports: the phase label, the
|
||||
/// absence policy, how long the phase has been held, and the failed-bind
|
||||
/// count since the last successful bind.
|
||||
///
|
||||
/// `None` for transports that are not bound to a named interface — for
|
||||
/// those, presence is not a concept and an operator should not be shown an
|
||||
/// always-`present` column.
|
||||
pub fn interface_presence(&self) -> Option<InterfacePresence> {
|
||||
match self {
|
||||
#[cfg(any(target_os = "linux", target_os = "macos"))]
|
||||
TransportHandle::Ethernet(t) => {
|
||||
let state = t.presence_state();
|
||||
Some(InterfacePresence {
|
||||
presence: t.presence().as_str(),
|
||||
carrier: t.has_carrier(),
|
||||
policy: t.absence_policy().as_str(),
|
||||
since_secs: state.since().as_secs(),
|
||||
binds: state.binds(),
|
||||
failed_attempts: state.attempts(),
|
||||
})
|
||||
}
|
||||
_ => None,
|
||||
}
|
||||
}
|
||||
|
||||
/// Whether this transport can actually put a frame on the wire *now*.
|
||||
///
|
||||
/// [`Self::is_operational`] answers a different question: it means the
|
||||
/// transport was started, which for an interface-bound transport no longer
|
||||
/// implies a live socket — that is the whole point of presence. Callers
|
||||
/// that are choosing a transport to use, or deriving a value from one,
|
||||
/// want this; callers reasoning about lifecycle want `is_operational`.
|
||||
///
|
||||
/// `true` for every transport that is not interface-bound, so this is
|
||||
/// `is_operational` with the presence refinement applied where it exists.
|
||||
pub fn is_bound(&self) -> bool {
|
||||
if !self.is_operational() {
|
||||
return false;
|
||||
}
|
||||
match self {
|
||||
#[cfg(any(target_os = "linux", target_os = "macos"))]
|
||||
TransportHandle::Ethernet(t) => t.presence() == ethernet::Presence::Present,
|
||||
_ => true,
|
||||
}
|
||||
}
|
||||
|
||||
/// Get the interface name (Ethernet only, returns None for other transports).
|
||||
pub fn interface_name(&self) -> Option<&str> {
|
||||
match self {
|
||||
@@ -1571,4 +1747,39 @@ mod tests {
|
||||
assert_eq!(handle.link_mtu(&addr), expected_mtu);
|
||||
assert_eq!(handle.link_mtu(&addr), handle.mtu());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn only_an_absent_interface_classifies_as_transient() {
|
||||
// The whole point of the classification is that it is narrow. An
|
||||
// interface the binder is already polling for will come back; nothing
|
||||
// else on this list resolves itself by waiting, and treating one of
|
||||
// them as transient would mean holding state open for a fault that is
|
||||
// never going to clear.
|
||||
assert!(
|
||||
TransportError::InterfaceUnavailable {
|
||||
interface: "eth0".into()
|
||||
}
|
||||
.is_transient()
|
||||
);
|
||||
|
||||
for terminal in [
|
||||
TransportError::NotStarted,
|
||||
TransportError::AlreadyStarted,
|
||||
TransportError::StartFailed("no CAP_NET_RAW".into()),
|
||||
TransportError::SendFailed("ENOBUFS".into()),
|
||||
TransportError::MtuExceeded {
|
||||
packet_size: 2000,
|
||||
mtu: 1500,
|
||||
},
|
||||
// Deliberately terminal: both describe a remote that did not
|
||||
// answer, not this node's inability to transmit.
|
||||
TransportError::Timeout,
|
||||
TransportError::ConnectionRefused,
|
||||
] {
|
||||
assert!(
|
||||
!terminal.is_transient(),
|
||||
"{terminal:?} must not be classified transient"
|
||||
);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,851 @@
|
||||
//! Interface presence state for interface-bound transports.
|
||||
//!
|
||||
//! A transport bound to a network interface is a long-lived object that is
|
||||
//! *sometimes bound*. The interface it names may not exist when the daemon
|
||||
//! starts, may appear minutes later, may vanish and return mid-operation, and
|
||||
//! may never appear at all. This module holds the state that makes that
|
||||
//! observable and drives the rebind loop in [`super`].
|
||||
//!
|
||||
//! ```text
|
||||
//! Absent ──attach──> Binding ──ok──> Present
|
||||
//! ^ │ │
|
||||
//! └──── fail/backoff ─┘ │
|
||||
//! └──────────── detach ──────────────┘
|
||||
//! ```
|
||||
//!
|
||||
//! Two invariants do the work:
|
||||
//!
|
||||
//! - **The transport object survives detach.** Config, `TransportId`,
|
||||
//! statistics, and the neighbor buffer persist; only the file descriptor and
|
||||
//! its loops go. A transport is never destroyed because its interface went
|
||||
//! away.
|
||||
//! - **Start-time absence and runtime detach are the same transition.** A node
|
||||
//! that boots before wifi and a node whose wifi reloads at 03:00 take one
|
||||
//! code path.
|
||||
//!
|
||||
//! Presence tracks `IFF_UP` — the interface exists and the operator has
|
||||
//! enabled it — and deliberately not `IFF_RUNNING`: binding needs no carrier,
|
||||
//! and a socket outlives a carrier flap. Carrier is reported alongside it
|
||||
//! rather than steering it. See
|
||||
//! the ethernet transport's `interface_present` for why.
|
||||
|
||||
use std::sync::atomic::{AtomicU8, AtomicU32, AtomicU64, Ordering};
|
||||
use std::sync::{PoisonError, RwLock, RwLockReadGuard, RwLockWriteGuard};
|
||||
use std::time::{Duration, Instant};
|
||||
|
||||
/// Read a lock, ignoring poisoning.
|
||||
///
|
||||
/// Every value guarded in this module is plain data — an `Instant`, an
|
||||
/// `Option<[u8; 6]>` — that a panic mid-write cannot leave logically
|
||||
/// inconsistent, so poisoning carries no information worth propagating.
|
||||
/// Treating it as a failure is what would hurt: the callers here are on the
|
||||
/// presence path, and "assume the worst" there means a transport that reports
|
||||
/// itself bound while every send fails, or a binder that tears down and
|
||||
/// rebinds every second forever. A stuck state is a worse outcome than
|
||||
/// reading a byte written by a thread that later panicked.
|
||||
fn read<T>(lock: &RwLock<T>) -> RwLockReadGuard<'_, T> {
|
||||
lock.read().unwrap_or_else(PoisonError::into_inner)
|
||||
}
|
||||
|
||||
/// Write a lock, ignoring poisoning. See [`read`].
|
||||
fn write<T>(lock: &RwLock<T>) -> RwLockWriteGuard<'_, T> {
|
||||
lock.write().unwrap_or_else(PoisonError::into_inner)
|
||||
}
|
||||
|
||||
/// How absence of the configured interface is reported.
|
||||
///
|
||||
/// Describes *the interface's presence*, not the transport's importance: an
|
||||
/// optional interface that is present is used exactly as hard as any other.
|
||||
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
|
||||
pub enum AbsencePolicy {
|
||||
/// Naming an interface in configuration is a statement that you expect it,
|
||||
/// so the default is to complain: absence degrades node health from the
|
||||
/// first edge, and if it outlasts [`ABSENCE_ERROR_AFTER`] — the window in
|
||||
/// which it could still have been an ordinary bring-up race — it is
|
||||
/// reported once at `error`.
|
||||
Required,
|
||||
/// Absence is normal for this interface (a dock adapter, a radio that only
|
||||
/// exists on some hardware): no health impact, `info` on the edge.
|
||||
Optional,
|
||||
}
|
||||
|
||||
impl AbsencePolicy {
|
||||
/// `optional: true` in configuration selects [`AbsencePolicy::Optional`].
|
||||
pub fn from_optional(optional: bool) -> Self {
|
||||
if optional {
|
||||
Self::Optional
|
||||
} else {
|
||||
Self::Required
|
||||
}
|
||||
}
|
||||
|
||||
/// Whether absence should be hidden from node health.
|
||||
pub fn is_optional(self) -> bool {
|
||||
matches!(self, Self::Optional)
|
||||
}
|
||||
|
||||
/// Operator-facing label, used by `show_transports`.
|
||||
pub fn as_str(self) -> &'static str {
|
||||
match self {
|
||||
Self::Required => "required",
|
||||
Self::Optional => "optional",
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Where a transport sits in the presence cycle.
|
||||
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
|
||||
pub enum Presence {
|
||||
/// The interface is not there (or is there without carrier). No socket, no
|
||||
/// loops; the watcher is waiting.
|
||||
Absent,
|
||||
/// The interface appeared and a bind is in flight, or a non-absence bind
|
||||
/// failure is backing off.
|
||||
Binding,
|
||||
/// Bound, with a live socket and running loops.
|
||||
Present,
|
||||
}
|
||||
|
||||
impl Presence {
|
||||
fn from_u8(v: u8) -> Self {
|
||||
match v {
|
||||
1 => Self::Binding,
|
||||
2 => Self::Present,
|
||||
_ => Self::Absent,
|
||||
}
|
||||
}
|
||||
|
||||
fn as_u8(self) -> u8 {
|
||||
match self {
|
||||
Self::Absent => 0,
|
||||
Self::Binding => 1,
|
||||
Self::Present => 2,
|
||||
}
|
||||
}
|
||||
|
||||
/// Operator-facing label, used by `show_transports`.
|
||||
pub fn as_str(self) -> &'static str {
|
||||
match self {
|
||||
Self::Absent => "absent",
|
||||
Self::Binding => "binding",
|
||||
Self::Present => "present",
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
impl std::fmt::Display for Presence {
|
||||
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
|
||||
f.write_str(self.as_str())
|
||||
}
|
||||
}
|
||||
|
||||
/// Shared, lock-light presence state.
|
||||
///
|
||||
/// Written by the transport's binder task, read by `send`, by the control
|
||||
/// plane, and by tests. Held behind an `Arc` so the binder task can outlive
|
||||
/// any particular borrow of the transport.
|
||||
#[derive(Debug)]
|
||||
pub struct PresenceState {
|
||||
phase: AtomicU8,
|
||||
/// Successful binds since the transport was created. `1` after a clean
|
||||
/// start; every increment past that is a rebind.
|
||||
binds: AtomicU64,
|
||||
/// Failed bind attempts since the last successful bind. Reset on bind.
|
||||
attempts: AtomicU32,
|
||||
/// When the current phase was entered.
|
||||
since: RwLock<Instant>,
|
||||
/// MAC observed at the last successful bind. A name reappearing with a
|
||||
/// different MAC is different hardware, not the same device returning.
|
||||
last_mac: RwLock<Option<[u8; 6]>>,
|
||||
}
|
||||
|
||||
impl Default for PresenceState {
|
||||
fn default() -> Self {
|
||||
Self::new()
|
||||
}
|
||||
}
|
||||
|
||||
impl PresenceState {
|
||||
/// A fresh tracker in [`Presence::Absent`].
|
||||
pub fn new() -> Self {
|
||||
Self {
|
||||
phase: AtomicU8::new(Presence::Absent.as_u8()),
|
||||
binds: AtomicU64::new(0),
|
||||
attempts: AtomicU32::new(0),
|
||||
since: RwLock::new(Instant::now()),
|
||||
last_mac: RwLock::new(None),
|
||||
}
|
||||
}
|
||||
|
||||
/// Current phase.
|
||||
pub fn presence(&self) -> Presence {
|
||||
Presence::from_u8(self.phase.load(Ordering::Acquire))
|
||||
}
|
||||
|
||||
/// Whether the transport currently holds a bound socket.
|
||||
pub fn is_present(&self) -> bool {
|
||||
self.presence() == Presence::Present
|
||||
}
|
||||
|
||||
/// How long the current presence *episode* has been held.
|
||||
///
|
||||
/// Episode, not phase: `Binding` is part of the absence episode until it
|
||||
/// succeeds. An interface that has been gone for a week while a bind is
|
||||
/// retried and refused every second must report a week, not one second —
|
||||
/// otherwise [`ABSENCE_ERROR_AFTER`] is never reached and the
|
||||
/// operator-facing `since_secs` reads as a healthy young absence forever.
|
||||
pub fn since(&self) -> Duration {
|
||||
read(&self.since).elapsed()
|
||||
}
|
||||
|
||||
/// Restart the episode clock, for a transport that is about to start.
|
||||
///
|
||||
/// `new()` stamps the clock at construction, but construction and
|
||||
/// `start_async` need not be adjacent — config load and supervisor staging
|
||||
/// sit between them. Left alone, a transport staged for longer than
|
||||
/// [`ABSENCE_ERROR_AFTER`] logs the sustained-absence error on its very
|
||||
/// first binder tick, having given the interface no bring-up window at
|
||||
/// all. The window is supposed to absorb exactly that race.
|
||||
pub fn mark_starting(&self) {
|
||||
*write(&self.since) = Instant::now();
|
||||
}
|
||||
|
||||
/// Test-only: age the episode clock, so deadline behaviour can be tested
|
||||
/// without sleeping through it.
|
||||
///
|
||||
/// `ABSENCE_ERROR_AFTER` is ten seconds. A test that waited it out would
|
||||
/// be ten seconds of nothing, which is how deadline logic ends up
|
||||
/// untested.
|
||||
#[cfg(test)]
|
||||
pub(crate) fn backdate_for_test(&self, by: Duration) {
|
||||
let mut since = write(&self.since);
|
||||
*since = since.checked_sub(by).unwrap_or(*since);
|
||||
}
|
||||
|
||||
/// Successful binds since creation (`1` after a clean start).
|
||||
pub fn binds(&self) -> u64 {
|
||||
self.binds.load(Ordering::Relaxed)
|
||||
}
|
||||
|
||||
/// Failed bind attempts since the last successful bind.
|
||||
pub fn attempts(&self) -> u32 {
|
||||
self.attempts.load(Ordering::Relaxed)
|
||||
}
|
||||
|
||||
/// MAC observed at the last successful bind, if any.
|
||||
pub fn last_mac(&self) -> Option<[u8; 6]> {
|
||||
*read(&self.last_mac)
|
||||
}
|
||||
|
||||
/// Move to `phase`, returning `true` if this was an actual edge.
|
||||
///
|
||||
/// Edge-vs-level is what the logging policy keys on: logged once on
|
||||
/// entering absence and once on recovery, never per retry attempt.
|
||||
///
|
||||
/// The episode clock ([`Self::since`]) restarts only when *boundness*
|
||||
/// changes — Present↔not-Present. A failed bind walks
|
||||
/// `Absent → Binding → Absent`, and resetting the clock on those would
|
||||
/// hide a permanent absence behind a timer that never gets past one
|
||||
/// second.
|
||||
pub fn transition(&self, phase: Presence) -> bool {
|
||||
let prev = Presence::from_u8(self.phase.swap(phase.as_u8(), Ordering::AcqRel));
|
||||
if prev == phase {
|
||||
return false;
|
||||
}
|
||||
if (prev == Presence::Present) != (phase == Presence::Present) {
|
||||
*write(&self.since) = Instant::now();
|
||||
}
|
||||
true
|
||||
}
|
||||
|
||||
/// Record a successful bind at `mac`.
|
||||
///
|
||||
/// Returns `true` when the interface came back as *different hardware* —
|
||||
/// the name reappeared with a MAC other than the one last bound. The
|
||||
/// caller drops cached neighbor state rather than silently resuming onto
|
||||
/// a different adapter.
|
||||
pub fn record_bind(&self, mac: [u8; 6]) -> bool {
|
||||
let changed = match *read(&self.last_mac) {
|
||||
Some(prev) => prev != mac,
|
||||
None => false,
|
||||
};
|
||||
*write(&self.last_mac) = Some(mac);
|
||||
self.binds.fetch_add(1, Ordering::Relaxed);
|
||||
self.attempts.store(0, Ordering::Relaxed);
|
||||
self.transition(Presence::Present);
|
||||
changed
|
||||
}
|
||||
|
||||
/// Record a failed bind attempt, returning the new attempt count.
|
||||
pub fn record_attempt(&self) -> u32 {
|
||||
self.attempts.fetch_add(1, Ordering::Relaxed) + 1
|
||||
}
|
||||
}
|
||||
|
||||
/// Backoff for bind failures that are *not* absence — permission denied,
|
||||
/// buffer sizing, a BPF device shortage. Absence itself does not back off
|
||||
/// where an event source is available: there is nothing to poll.
|
||||
///
|
||||
/// 1 s doubling to a 30 s ceiling.
|
||||
pub fn bind_backoff(attempts: u32) -> Duration {
|
||||
const BASE_SECS: u64 = 1;
|
||||
const CEILING_SECS: u64 = 30;
|
||||
let shift = attempts.saturating_sub(1).min(5);
|
||||
Duration::from_secs((BASE_SECS << shift).min(CEILING_SECS))
|
||||
}
|
||||
|
||||
/// How long a *required* interface may be absent before it is an error.
|
||||
///
|
||||
/// Absence is a state the presence machine handles, so it is not an error for
|
||||
/// happening — a daemon that wins the race against its own radio, or a cable
|
||||
/// out for two seconds, is the ordinary case this mechanism exists to absorb,
|
||||
/// and calling that an error at t=0 and "recovered" at t=0.2 s is cry-wolf.
|
||||
/// Past this window it is no longer a race: something an operator has to fix
|
||||
/// is wrong, and the log should say so once.
|
||||
///
|
||||
/// One window for both shapes of absence. A node that boots before its wifi
|
||||
/// and a node whose wifi reloads at 03:00 take one code path everywhere else
|
||||
/// in this module; giving them different deadlines would reintroduce exactly
|
||||
/// the start-versus-runtime asymmetry the presence machine removed.
|
||||
///
|
||||
/// Tuned against the platforms this exists for: comfortably past a veth or a
|
||||
/// container coming up, short enough that a mesh radio which never appears is
|
||||
/// named while somebody is still watching the boot. Raising it hides a real
|
||||
/// fault for longer; lowering it starts reporting ordinary bring-up races.
|
||||
pub const ABSENCE_ERROR_AFTER: Duration = Duration::from_secs(10);
|
||||
|
||||
/// Minimum lifetime for a binding to count as a real recovery.
|
||||
///
|
||||
/// A socket that dies sooner than this never really came back.
|
||||
pub const MIN_STABLE_BINDING: Duration = Duration::from_secs(10);
|
||||
|
||||
/// Consecutive short-lived bindings before the binder stops treating a
|
||||
/// successful bind as a recovery.
|
||||
pub const CHURN_THRESHOLD: u32 = 3;
|
||||
|
||||
/// Damping for the rebind loop.
|
||||
///
|
||||
/// Backoff covers *failed* binds; this covers the opposite and nastier case —
|
||||
/// binds that keep **succeeding** into a socket that dies moments later. A
|
||||
/// receive loop that gives up on a persistent error while the interface stays
|
||||
/// `UP` produces exactly that: tear down, rebind, succeed, fail again, once
|
||||
/// per second, forever. Undamped it is an `error!`/`info!` pair and a
|
||||
/// `Degraded`→`Running` health flap every cycle, which defeats both the
|
||||
/// "log edges, not attempts" rule and the meaning of `Degraded`.
|
||||
///
|
||||
/// So: count consecutive bindings that die young, back off between them on
|
||||
/// the same 1 s → 30 s curve, and once the streak reaches
|
||||
/// [`CHURN_THRESHOLD`] stop announcing each bind as a recovery — hold the
|
||||
/// node at its degraded reading until a binding actually survives
|
||||
/// [`MIN_STABLE_BINDING`]. A binding that holds ends the streak.
|
||||
///
|
||||
/// Pure state, driven by an injected clock, so the policy is testable without
|
||||
/// a network interface.
|
||||
#[derive(Debug, Default)]
|
||||
pub struct ChurnGuard {
|
||||
/// Consecutive bindings that died younger than [`MIN_STABLE_BINDING`].
|
||||
streak: u32,
|
||||
/// When the current binding was established.
|
||||
bound_at: Option<Instant>,
|
||||
/// Whether the current binding was announced as a recovery.
|
||||
announced: bool,
|
||||
}
|
||||
|
||||
/// What the caller should do about a successful bind.
|
||||
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
|
||||
pub struct BindOutcome {
|
||||
/// Log the recovery and publish presence now. `false` while churning:
|
||||
/// the bind is held back until it proves it will last.
|
||||
pub announce: bool,
|
||||
}
|
||||
|
||||
/// What the caller should do about a detach.
|
||||
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
|
||||
pub struct DetachOutcome {
|
||||
/// Publish absence. `false` when the binding that just died was never
|
||||
/// announced, so there is nothing to retract.
|
||||
pub retract: bool,
|
||||
/// Log the detach edge. `false` once churning — the fact has been said.
|
||||
pub log_edge: bool,
|
||||
/// This detach is the one that crossed [`CHURN_THRESHOLD`]; say so once.
|
||||
pub entered_churn: bool,
|
||||
/// Wait this long before trying to bind again.
|
||||
pub backoff: Option<Duration>,
|
||||
}
|
||||
|
||||
impl ChurnGuard {
|
||||
/// A guard with no history.
|
||||
pub fn new() -> Self {
|
||||
Self::default()
|
||||
}
|
||||
|
||||
/// Consecutive short-lived bindings, for logging and tests.
|
||||
pub fn streak(&self) -> u32 {
|
||||
self.streak
|
||||
}
|
||||
|
||||
/// Record a successful bind.
|
||||
pub fn bound(&mut self, now: Instant) -> BindOutcome {
|
||||
self.bound_at = Some(now);
|
||||
let announce = self.streak < CHURN_THRESHOLD;
|
||||
if announce {
|
||||
self.announced = true;
|
||||
}
|
||||
BindOutcome { announce }
|
||||
}
|
||||
|
||||
/// Called on every tick while bound. Returns `true` exactly once, at the
|
||||
/// moment a held-back binding has proved stable and should be announced.
|
||||
pub fn stabilized(&mut self, now: Instant) -> bool {
|
||||
let Some(bound_at) = self.bound_at else {
|
||||
return false;
|
||||
};
|
||||
if now.duration_since(bound_at) < MIN_STABLE_BINDING {
|
||||
return false;
|
||||
}
|
||||
let newly_announced = !self.announced;
|
||||
self.streak = 0;
|
||||
self.announced = true;
|
||||
newly_announced
|
||||
}
|
||||
|
||||
/// Record a detach.
|
||||
pub fn detached(&mut self, now: Instant) -> DetachOutcome {
|
||||
let young = self
|
||||
.bound_at
|
||||
.is_some_and(|t| now.duration_since(t) < MIN_STABLE_BINDING);
|
||||
self.bound_at = None;
|
||||
|
||||
if young {
|
||||
self.streak += 1;
|
||||
} else {
|
||||
self.streak = 0;
|
||||
}
|
||||
|
||||
DetachOutcome {
|
||||
retract: std::mem::take(&mut self.announced),
|
||||
log_edge: self.streak < CHURN_THRESHOLD,
|
||||
entered_churn: self.streak == CHURN_THRESHOLD,
|
||||
backoff: (self.streak > 0).then(|| bind_backoff(self.streak)),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
use std::sync::Arc;
|
||||
|
||||
#[test]
|
||||
fn presence_starts_absent() {
|
||||
let p = PresenceState::new();
|
||||
assert_eq!(p.presence(), Presence::Absent);
|
||||
assert_eq!(p.binds(), 0);
|
||||
assert_eq!(p.attempts(), 0);
|
||||
assert!(p.last_mac().is_none());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn transition_reports_edges_only() {
|
||||
let p = PresenceState::new();
|
||||
assert!(p.transition(Presence::Binding));
|
||||
assert!(!p.transition(Presence::Binding));
|
||||
assert!(p.transition(Presence::Present));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn bind_records_mac_and_clears_attempts() {
|
||||
let p = PresenceState::new();
|
||||
p.record_attempt();
|
||||
p.record_attempt();
|
||||
assert_eq!(p.attempts(), 2);
|
||||
|
||||
let mac = [0x02, 0, 0, 0, 0, 1];
|
||||
assert!(!p.record_bind(mac), "first bind is not a hardware change");
|
||||
assert_eq!(p.attempts(), 0);
|
||||
assert_eq!(p.binds(), 1);
|
||||
assert_eq!(p.presence(), Presence::Present);
|
||||
assert_eq!(p.last_mac(), Some(mac));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn rebind_on_same_mac_is_not_a_hardware_change() {
|
||||
let p = PresenceState::new();
|
||||
let mac = [0x02, 0, 0, 0, 0, 1];
|
||||
p.record_bind(mac);
|
||||
p.transition(Presence::Absent);
|
||||
assert!(!p.record_bind(mac));
|
||||
assert_eq!(p.binds(), 2);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn rebind_on_different_mac_is_a_hardware_change() {
|
||||
let p = PresenceState::new();
|
||||
p.record_bind([0x02, 0, 0, 0, 0, 1]);
|
||||
p.transition(Presence::Absent);
|
||||
assert!(p.record_bind([0x02, 0, 0, 0, 0, 2]));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn backoff_climbs_to_a_ceiling() {
|
||||
assert_eq!(bind_backoff(0), Duration::from_secs(1));
|
||||
assert_eq!(bind_backoff(1), Duration::from_secs(1));
|
||||
assert_eq!(bind_backoff(2), Duration::from_secs(2));
|
||||
assert_eq!(bind_backoff(3), Duration::from_secs(4));
|
||||
assert_eq!(bind_backoff(6), Duration::from_secs(30));
|
||||
assert_eq!(bind_backoff(u32::MAX), Duration::from_secs(30));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn the_error_deadline_outlasts_an_ordinary_bring_up_race() {
|
||||
// The window has to clear the races it exists to absorb — a veth
|
||||
// arriving a fraction of a second late, a container starting — while
|
||||
// staying short enough that a radio which never appears is named
|
||||
// during the boot somebody is watching. It is also the one deadline:
|
||||
// start-time absence and a runtime detach share it.
|
||||
assert!(
|
||||
ABSENCE_ERROR_AFTER >= Duration::from_secs(5),
|
||||
"shorter than a bring-up race would report the ordinary case"
|
||||
);
|
||||
assert!(
|
||||
ABSENCE_ERROR_AFTER <= Duration::from_secs(60),
|
||||
"longer and a required interface that never appears goes unsaid \
|
||||
for the whole boot"
|
||||
);
|
||||
}
|
||||
|
||||
// ── The absence clock measures an episode, not a phase ────────────────
|
||||
|
||||
#[test]
|
||||
fn a_failed_bind_does_not_restart_the_absence_clock() {
|
||||
// An interface that is present but refuses to bind walks
|
||||
// Absent → Binding → Absent on every retry. If those edges reset the
|
||||
// clock, `since_secs` reads as a one-second-old absence forever and
|
||||
// the error deadline is never reached — so a permission error would
|
||||
// sit silently behind a healthy-looking counter.
|
||||
let p = PresenceState::new();
|
||||
std::thread::sleep(Duration::from_millis(30));
|
||||
let before = p.since();
|
||||
|
||||
assert!(p.transition(Presence::Binding));
|
||||
assert!(p.transition(Presence::Absent));
|
||||
assert!(p.transition(Presence::Binding));
|
||||
assert!(p.transition(Presence::Absent));
|
||||
|
||||
assert!(
|
||||
p.since() >= before,
|
||||
"the absence clock ran backwards across failed binds"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn the_clock_restarts_only_when_boundness_changes() {
|
||||
let p = PresenceState::new();
|
||||
std::thread::sleep(Duration::from_millis(30));
|
||||
|
||||
// Absent → Present restarts it: a new episode began.
|
||||
p.record_bind([0x02, 0, 0, 0, 0, 1]);
|
||||
assert!(p.since() < Duration::from_millis(30));
|
||||
|
||||
std::thread::sleep(Duration::from_millis(30));
|
||||
let bound_for = p.since();
|
||||
|
||||
// Present → Present is not an edge at all.
|
||||
assert!(!p.transition(Presence::Present));
|
||||
assert!(p.since() >= bound_for);
|
||||
|
||||
// Present → Absent restarts it: the episode ended.
|
||||
assert!(p.transition(Presence::Absent));
|
||||
assert!(p.since() < Duration::from_millis(30));
|
||||
}
|
||||
|
||||
// ── Poisoning must not be a stuck state ───────────────────────────────
|
||||
|
||||
#[test]
|
||||
fn a_poisoned_lock_still_reports_presence() {
|
||||
// `.ok()`-style handling would make a poisoned lock read as "no MAC,
|
||||
// no socket, tasks dead" — a transport reporting itself present while
|
||||
// every send fails, and a binder rebinding once a second forever.
|
||||
// Poisoning carries no information about plain data, so it is ignored.
|
||||
let p = Arc::new(PresenceState::new());
|
||||
p.record_bind([0x02, 0, 0, 0, 0, 7]);
|
||||
|
||||
let poisoner = Arc::clone(&p);
|
||||
let panicked = std::thread::spawn(move || {
|
||||
let _guard = poisoner.last_mac.write().unwrap();
|
||||
panic!("poison the lock while holding it");
|
||||
})
|
||||
.join();
|
||||
assert!(panicked.is_err(), "the helper thread was supposed to panic");
|
||||
assert!(
|
||||
p.last_mac.is_poisoned(),
|
||||
"the lock was supposed to be poisoned"
|
||||
);
|
||||
|
||||
assert_eq!(
|
||||
p.last_mac(),
|
||||
Some([0x02, 0, 0, 0, 0, 7]),
|
||||
"a poisoned lock must not erase the binding"
|
||||
);
|
||||
// And the clock still answers rather than collapsing to zero.
|
||||
let _ = p.since();
|
||||
}
|
||||
|
||||
// ── Rebind churn ──────────────────────────────────────────────────────
|
||||
|
||||
#[test]
|
||||
fn a_healthy_bind_and_detach_is_not_churn() {
|
||||
let mut g = ChurnGuard::new();
|
||||
let t0 = Instant::now();
|
||||
assert!(g.bound(t0).announce, "a first bind is a recovery");
|
||||
|
||||
// Held well past the stability floor, then lost.
|
||||
let out = g.detached(t0 + MIN_STABLE_BINDING + Duration::from_secs(60));
|
||||
assert!(out.retract, "an announced binding must be retracted");
|
||||
assert!(out.log_edge, "an isolated detach is worth a line");
|
||||
assert!(!out.entered_churn);
|
||||
assert_eq!(out.backoff, None, "one clean outage must not back off");
|
||||
assert_eq!(g.streak(), 0);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn an_unseeded_guard_retracts_nothing_and_never_repairs_itself() {
|
||||
// Why `binder_loop` seeds the guard when it inherits a binding from
|
||||
// `start_async`, rather than leaving it fresh.
|
||||
//
|
||||
// A guard that was never told about a bind believes it has announced
|
||||
// nothing, so it asks for no retraction — and health, which learned
|
||||
// `present: true` from the inline bind, would keep reading `Full` with
|
||||
// the interface gone. `stabilized` cannot rescue it either: with no
|
||||
// `bound_at` there is nothing for it to judge stable.
|
||||
let mut g = ChurnGuard::new();
|
||||
let t0 = Instant::now();
|
||||
|
||||
assert!(
|
||||
!g.stabilized(t0 + MIN_STABLE_BINDING + Duration::from_secs(60)),
|
||||
"a guard with no recorded bind has nothing to stabilize"
|
||||
);
|
||||
|
||||
let out = g.detached(t0 + Duration::from_secs(60));
|
||||
assert!(
|
||||
!out.retract,
|
||||
"an unseeded guard retracts nothing — which is exactly why the \
|
||||
binder must seed it from the inline bind"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_seeded_guard_retracts_the_edge_the_inline_bind_published() {
|
||||
// The fix, from the binder's angle: seeding with `bound` is what makes
|
||||
// the first detach after a clean start reach node health.
|
||||
let mut g = ChurnGuard::new();
|
||||
let t0 = Instant::now();
|
||||
g.bound(t0);
|
||||
|
||||
let out = g.detached(t0 + MIN_STABLE_BINDING + Duration::from_secs(60));
|
||||
assert!(
|
||||
out.retract,
|
||||
"the edge `start_async` published must be retracted on detach"
|
||||
);
|
||||
assert!(out.log_edge);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn short_lived_bindings_back_off() {
|
||||
// The failure this guards: a receive loop that gives up on a
|
||||
// persistent error while the interface stays UP. Bind succeeds, dies,
|
||||
// rebinds, dies — once per second, forever, undamped.
|
||||
let mut g = ChurnGuard::new();
|
||||
let mut t = Instant::now();
|
||||
|
||||
for expected in [1u64, 2, 4] {
|
||||
g.bound(t);
|
||||
t += Duration::from_secs(1);
|
||||
let out = g.detached(t);
|
||||
assert_eq!(
|
||||
out.backoff,
|
||||
Some(Duration::from_secs(expected)),
|
||||
"streak {} should back off {expected}s",
|
||||
g.streak()
|
||||
);
|
||||
}
|
||||
assert_eq!(g.streak(), 3);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn churn_stops_announcing_and_stops_logging() {
|
||||
let mut g = ChurnGuard::new();
|
||||
let mut t = Instant::now();
|
||||
|
||||
// The detaches below the threshold are still news and still logged.
|
||||
for _ in 0..CHURN_THRESHOLD - 1 {
|
||||
assert!(g.bound(t).announce);
|
||||
t += Duration::from_secs(1);
|
||||
let out = g.detached(t);
|
||||
assert!(out.log_edge, "the first few detaches are still news");
|
||||
assert!(!out.entered_churn);
|
||||
}
|
||||
|
||||
// The detach that crosses the threshold reports the churn instead of
|
||||
// the edge: one line saying "this keeps happening", not two saying
|
||||
// "it happened" and "it keeps happening".
|
||||
assert!(g.bound(t).announce);
|
||||
t += Duration::from_secs(1);
|
||||
let crossing = g.detached(t);
|
||||
assert!(crossing.entered_churn, "crossing must be announced once");
|
||||
assert!(!crossing.log_edge, "the churn line replaces the edge line");
|
||||
|
||||
// Past the threshold: bindings are no longer announced as recoveries,
|
||||
// so node health stays put instead of flapping every second, and the
|
||||
// edges stop being logged.
|
||||
for _ in 0..5 {
|
||||
assert!(!g.bound(t).announce, "a churning bind is not a recovery");
|
||||
t += Duration::from_secs(1);
|
||||
let out = g.detached(t);
|
||||
assert!(!out.log_edge, "churn must not log per cycle");
|
||||
assert!(!out.retract, "nothing was announced, so nothing to retract");
|
||||
assert!(
|
||||
!out.entered_churn,
|
||||
"the threshold is crossed once, not repeatedly"
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn backoff_during_churn_is_capped() {
|
||||
let mut g = ChurnGuard::new();
|
||||
let mut t = Instant::now();
|
||||
let mut last = None;
|
||||
for _ in 0..12 {
|
||||
g.bound(t);
|
||||
t += Duration::from_secs(1);
|
||||
last = g.detached(t).backoff;
|
||||
}
|
||||
assert_eq!(
|
||||
last,
|
||||
Some(Duration::from_secs(30)),
|
||||
"churn backoff must climb to the ceiling and stop"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_binding_that_lasts_ends_the_streak_and_announces_once() {
|
||||
let mut g = ChurnGuard::new();
|
||||
let mut t = Instant::now();
|
||||
|
||||
// Churn into the held-back state.
|
||||
for _ in 0..CHURN_THRESHOLD + 1 {
|
||||
g.bound(t);
|
||||
t += Duration::from_secs(1);
|
||||
g.detached(t);
|
||||
}
|
||||
assert!(!g.bound(t).announce);
|
||||
|
||||
// Not yet stable: still nothing to say.
|
||||
assert!(!g.stabilized(t + Duration::from_secs(1)));
|
||||
|
||||
// Survived the floor: announce exactly once, and the streak is over.
|
||||
let stable_at = t + MIN_STABLE_BINDING;
|
||||
assert!(
|
||||
g.stabilized(stable_at),
|
||||
"a binding that lasts is a recovery"
|
||||
);
|
||||
assert!(
|
||||
!g.stabilized(stable_at + Duration::from_secs(60)),
|
||||
"recovery is announced once, not on every tick"
|
||||
);
|
||||
assert_eq!(g.streak(), 0);
|
||||
|
||||
// And the next detach behaves like an ordinary one again.
|
||||
let out = g.detached(stable_at + Duration::from_secs(60));
|
||||
assert!(out.retract);
|
||||
assert!(out.log_edge);
|
||||
assert_eq!(out.backoff, None);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn stabilized_is_silent_for_an_ordinary_binding() {
|
||||
// A bind that was announced immediately must not be announced again
|
||||
// when it passes the stability floor.
|
||||
let mut g = ChurnGuard::new();
|
||||
let t = Instant::now();
|
||||
assert!(g.bound(t).announce);
|
||||
assert!(!g.stabilized(t + MIN_STABLE_BINDING + Duration::from_secs(1)));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn policy_labels() {
|
||||
assert_eq!(AbsencePolicy::from_optional(true), AbsencePolicy::Optional);
|
||||
assert_eq!(AbsencePolicy::from_optional(false), AbsencePolicy::Required);
|
||||
assert!(AbsencePolicy::Optional.is_optional());
|
||||
assert!(!AbsencePolicy::Required.is_optional());
|
||||
assert_eq!(AbsencePolicy::Required.as_str(), "required");
|
||||
// Both labels, not just one. `show_transports` renders this string and
|
||||
// fipstop's severity split keys on it, so a swapped pair would paint
|
||||
// every expected interface as the tolerated kind and vice versa —
|
||||
// while a test that checks only `Required` stays green through it.
|
||||
assert_eq!(AbsencePolicy::Optional.as_str(), "optional");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_first_bind_is_not_a_hardware_change() {
|
||||
// The boundary the flush hangs off. `record_bind` returns "different
|
||||
// hardware", and on the very first bind there is no previous MAC to
|
||||
// differ from — so it must answer false, or every clean start would
|
||||
// drop a neighbour cache it had just built and log a hardware swap
|
||||
// that never happened.
|
||||
let state = PresenceState::new();
|
||||
assert!(
|
||||
!state.record_bind([1, 2, 3, 4, 5, 6]),
|
||||
"the first bind has nothing to differ from"
|
||||
);
|
||||
assert_eq!(state.binds(), 1);
|
||||
assert_eq!(state.presence(), Presence::Present);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_rebind_on_new_hardware_reports_the_change_once() {
|
||||
// And it reports the change once, not on every subsequent bind: the
|
||||
// caller drops its cached neighbours on a `true`, so a sticky answer
|
||||
// would flush the cache on every rebind forever.
|
||||
let state = PresenceState::new();
|
||||
state.record_bind([1, 2, 3, 4, 5, 6]);
|
||||
|
||||
assert!(
|
||||
state.record_bind([0xaa, 0xbb, 0xcc, 0xdd, 0xee, 0xff]),
|
||||
"a name returning on a different MAC is different hardware"
|
||||
);
|
||||
assert!(
|
||||
!state.record_bind([0xaa, 0xbb, 0xcc, 0xdd, 0xee, 0xff]),
|
||||
"the same hardware rebinding is not a change"
|
||||
);
|
||||
assert_eq!(state.binds(), 3);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn every_presence_label_is_distinct_and_round_trips() {
|
||||
// The labels are the operator-facing vocabulary — `show_transports`
|
||||
// emits them and the fipstop State column renders them — and
|
||||
// `Presence::as_str` had no test at all, so `binding` in particular was
|
||||
// never observed by anything.
|
||||
use std::collections::HashSet;
|
||||
let all = [Presence::Absent, Presence::Binding, Presence::Present];
|
||||
let labels: HashSet<&str> = all.iter().map(|p| p.as_str()).collect();
|
||||
assert_eq!(labels.len(), 3, "each phase needs its own label");
|
||||
assert!(labels.contains("binding"));
|
||||
|
||||
for phase in all {
|
||||
assert_eq!(
|
||||
Presence::from_u8(phase.as_u8()),
|
||||
phase,
|
||||
"{phase} must survive the atomic round trip the state uses"
|
||||
);
|
||||
assert_eq!(phase.to_string(), phase.as_str(), "Display must agree");
|
||||
}
|
||||
|
||||
// Anything outside the enum reads as absent rather than panicking: the
|
||||
// byte comes out of an AtomicU8 that a torn write could leave at any
|
||||
// value, and the safe answer there is "not bound".
|
||||
assert_eq!(Presence::from_u8(99), Presence::Absent);
|
||||
}
|
||||
}
|
||||
+81
-42
@@ -80,6 +80,35 @@ impl PathMtuEntry {
|
||||
/// address).
|
||||
pub type PathMtuLookup = Arc<RwLock<HashMap<FipsAddress, PathMtuEntry>>>;
|
||||
|
||||
/// The node-global TCP MSS ceiling, shared live with the TUN reader and
|
||||
/// writer threads.
|
||||
///
|
||||
/// Shared rather than copied because the value it is derived from moves at
|
||||
/// runtime. `Node::transport_mtu()` is the minimum across *bound* transports,
|
||||
/// and since dynamic interface binding a transport can bind minutes after
|
||||
/// start or unbind mid-operation — so a narrow interface appearing must
|
||||
/// tighten the clamp, and its departure must release it. Every other consumer
|
||||
/// of `transport_mtu()` already reads it live (`show_status`, the snapshot,
|
||||
/// the session-layer fragmentation check); these two threads captured a `u16`
|
||||
/// at spawn and were the only place left where the daemon could report one
|
||||
/// effective MTU and clamp to another.
|
||||
///
|
||||
/// A relaxed load per packet, beside the `PathMtuLookup` read that already
|
||||
/// happens on the same packet — strictly the cheaper of the two. Ordering is
|
||||
/// irrelevant: this is a clamp, and a packet that reads the previous value
|
||||
/// during the store is clamped by the ceiling that was correct a microsecond
|
||||
/// earlier. The per-flow ceiling has always had that property.
|
||||
/// The IPv6 minimum link MTU (RFC 8200): every compliant path accepts a packet
|
||||
/// this large, so an MSS derived from it fits anywhere.
|
||||
///
|
||||
/// Two callers, and they must not disagree: the cold-flow fallback in
|
||||
/// [`per_flow_max_mss`], and the seed for the node's [`MssCeiling`] before any
|
||||
/// transport has bound — which is the same value `Node::transport_mtu()` falls
|
||||
/// back to when nothing is bound.
|
||||
pub const IPV6_MIN_MTU: u16 = 1280;
|
||||
|
||||
pub type MssCeiling = Arc<std::sync::atomic::AtomicU16>;
|
||||
|
||||
/// Compute the effective TCP MSS ceiling for a packet given its peer
|
||||
/// address bytes (a 16-byte IPv6 destination on outbound, source on
|
||||
/// inbound). Returns `min(global_max_mss, learned_path_max_mss)` when
|
||||
@@ -115,7 +144,6 @@ pub(crate) fn per_flow_max_mss(
|
||||
// RFC 8200 IPv6-minimum MTU (1280) → effective FIPS-encapsulated
|
||||
// payload (1203) → TCP segment after IPv6+TCP headers (1143).
|
||||
// Used as the conservative ceiling for empty-lookup destinations.
|
||||
const IPV6_MIN_MTU: u16 = 1280;
|
||||
let conservative_max_mss = mss_ceiling(IPV6_MIN_MTU);
|
||||
let empty_lookup_ceiling = std::cmp::min(global_max_mss, conservative_max_mss);
|
||||
|
||||
@@ -435,13 +463,14 @@ impl TunDevice {
|
||||
/// a channel sender for submitting packets to be written.
|
||||
///
|
||||
/// `max_mss` is the global TCP MSS ceiling derived from the local
|
||||
/// `transport_mtu()` floor. `path_mtu_lookup` is a read-only handle to
|
||||
/// the per-destination path MTU map populated by discovery; the writer
|
||||
/// reads it on each inbound SYN-ACK to compute a per-flow ceiling that
|
||||
/// honors learned narrow paths through the mesh.
|
||||
/// `transport_mtu()` floor, shared live so a transport that binds or
|
||||
/// unbinds after start moves it (see [`MssCeiling`]). `path_mtu_lookup`
|
||||
/// is a read-only handle to the per-destination path MTU map populated by
|
||||
/// discovery; the writer reads both on each inbound SYN-ACK to compute a
|
||||
/// per-flow ceiling that honors learned narrow paths through the mesh.
|
||||
pub fn create_writer(
|
||||
&self,
|
||||
max_mss: u16,
|
||||
max_mss: MssCeiling,
|
||||
path_mtu_lookup: PathMtuLookup,
|
||||
) -> Result<(TunWriter, TunTx), TunError> {
|
||||
let fd = self.device.as_raw_fd();
|
||||
@@ -517,7 +546,7 @@ pub struct TunWriter {
|
||||
file: File,
|
||||
rx: mpsc::Receiver<Vec<u8>>,
|
||||
name: String,
|
||||
max_mss: u16,
|
||||
max_mss: MssCeiling,
|
||||
path_mtu_lookup: PathMtuLookup,
|
||||
}
|
||||
|
||||
@@ -531,16 +560,23 @@ impl TunWriter {
|
||||
pub fn run(mut self) {
|
||||
use super::tcp_mss::clamp_tcp_mss;
|
||||
|
||||
debug!(name = %self.name, max_mss = self.max_mss, "TUN writer starting");
|
||||
debug!(
|
||||
name = %self.name,
|
||||
max_mss = self.max_mss.load(std::sync::atomic::Ordering::Relaxed),
|
||||
"TUN writer starting"
|
||||
);
|
||||
|
||||
for mut packet in self.rx {
|
||||
// Read per packet, not once: a transport binding or unbinding
|
||||
// moves the node's egress floor at runtime. See `MssCeiling`.
|
||||
let global_max_mss = self.max_mss.load(std::sync::atomic::Ordering::Relaxed);
|
||||
// Per-destination clamp: peer IPv6 source address (bytes 8..24)
|
||||
// identifies the flow's remote end. If discovery has learned a
|
||||
// smaller path MTU for that peer, tighten the ceiling.
|
||||
let effective_max_mss = if packet.len() >= 24 {
|
||||
per_flow_max_mss(&self.path_mtu_lookup, &packet[8..24], self.max_mss)
|
||||
per_flow_max_mss(&self.path_mtu_lookup, &packet[8..24], global_max_mss)
|
||||
} else {
|
||||
self.max_mss
|
||||
global_max_mss
|
||||
};
|
||||
// Clamp TCP MSS on inbound SYN-ACK packets
|
||||
if clamp_tcp_mss(&mut packet, effective_max_mss) {
|
||||
@@ -622,17 +658,17 @@ pub fn run_tun_reader(
|
||||
our_addr: FipsAddress,
|
||||
tun_tx: TunTx,
|
||||
outbound_tx: TunOutboundTx,
|
||||
transport_mtu: u16,
|
||||
max_mss: MssCeiling,
|
||||
path_mtu_lookup: PathMtuLookup,
|
||||
) {
|
||||
let (name, mut buf, max_mss) = tun_reader_setup(device.name(), mtu, transport_mtu);
|
||||
let (name, mut buf) = tun_reader_setup(device.name(), mtu, &max_mss);
|
||||
|
||||
loop {
|
||||
match device.read_packet(&mut buf) {
|
||||
Ok(n) if n > 0 => {
|
||||
if !handle_tun_packet(
|
||||
&mut buf[..n],
|
||||
max_mss,
|
||||
max_mss.load(std::sync::atomic::Ordering::Relaxed),
|
||||
&name,
|
||||
our_addr,
|
||||
&tun_tx,
|
||||
@@ -685,13 +721,13 @@ pub fn run_tun_reader(
|
||||
our_addr: FipsAddress,
|
||||
tun_tx: TunTx,
|
||||
outbound_tx: TunOutboundTx,
|
||||
transport_mtu: u16,
|
||||
max_mss: MssCeiling,
|
||||
path_mtu_lookup: PathMtuLookup,
|
||||
shutdown_fd: std::os::unix::io::RawFd,
|
||||
) {
|
||||
let _shutdown_fd = ShutdownFd(shutdown_fd);
|
||||
let tun_fd = device.device().as_raw_fd();
|
||||
let (name, mut buf, max_mss) = tun_reader_setup(device.name(), mtu, transport_mtu);
|
||||
let (name, mut buf) = tun_reader_setup(device.name(), mtu, &max_mss);
|
||||
|
||||
// Set TUN fd to non-blocking so we can use select + read without blocking
|
||||
// past the point where select returns readable.
|
||||
@@ -741,7 +777,7 @@ pub fn run_tun_reader(
|
||||
Ok(n) if n > 0 => {
|
||||
if !handle_tun_packet(
|
||||
&mut buf[..n],
|
||||
max_mss,
|
||||
max_mss.load(std::sync::atomic::Ordering::Relaxed),
|
||||
&name,
|
||||
our_addr,
|
||||
&tun_tx,
|
||||
@@ -774,30 +810,25 @@ pub fn run_tun_reader(
|
||||
// _shutdown_fd closes on drop
|
||||
}
|
||||
|
||||
/// Common setup for TUN reader: allocates buffer, computes max MSS.
|
||||
fn tun_reader_setup(device_name: &str, mtu: u16, transport_mtu: u16) -> (String, Vec<u8>, u16) {
|
||||
use super::icmp::effective_ipv6_mtu;
|
||||
|
||||
/// Common setup for TUN reader: allocates the buffer and names the device.
|
||||
///
|
||||
/// The MSS ceiling is deliberately *not* returned. It is read from the shared
|
||||
/// [`MssCeiling`] on every packet, because a transport binding or unbinding
|
||||
/// moves it after this function has run; returning it here is what let the
|
||||
/// reader clamp to a floor derived from the transports that happened to be
|
||||
/// bound at spawn.
|
||||
fn tun_reader_setup(device_name: &str, mtu: u16, max_mss: &MssCeiling) -> (String, Vec<u8>) {
|
||||
let name = device_name.to_string();
|
||||
let buf = vec![0u8; mtu as usize + 100];
|
||||
|
||||
const IPV6_HEADER: u16 = 40;
|
||||
const TCP_HEADER: u16 = 20;
|
||||
let effective_mtu = effective_ipv6_mtu(transport_mtu);
|
||||
let max_mss = effective_mtu
|
||||
.saturating_sub(IPV6_HEADER)
|
||||
.saturating_sub(TCP_HEADER);
|
||||
|
||||
debug!(
|
||||
name = %name,
|
||||
tun_mtu = mtu,
|
||||
transport_mtu = transport_mtu,
|
||||
effective_mtu = effective_mtu,
|
||||
max_mss = max_mss,
|
||||
max_mss = max_mss.load(std::sync::atomic::Ordering::Relaxed),
|
||||
"TUN reader starting"
|
||||
);
|
||||
|
||||
(name, buf, max_mss)
|
||||
(name, buf)
|
||||
}
|
||||
|
||||
/// Process a single TUN packet. Returns `false` if the reader should exit.
|
||||
@@ -1076,12 +1107,13 @@ mod windows_tun {
|
||||
/// packets independently. Returns the writer and a channel sender for
|
||||
/// submitting packets to be written.
|
||||
///
|
||||
/// `max_mss` is the global TCP MSS ceiling. `path_mtu_lookup` is a
|
||||
/// read-only handle to per-destination path MTU learned via
|
||||
/// discovery.
|
||||
/// `max_mss` is the global TCP MSS ceiling, shared live so a transport
|
||||
/// that binds or unbinds after start moves it (see [`MssCeiling`]).
|
||||
/// `path_mtu_lookup` is a read-only handle to per-destination path MTU
|
||||
/// learned via discovery.
|
||||
pub fn create_writer(
|
||||
&self,
|
||||
max_mss: u16,
|
||||
max_mss: MssCeiling,
|
||||
path_mtu_lookup: PathMtuLookup,
|
||||
) -> Result<(TunWriter, TunTx), TunError> {
|
||||
let (tx, rx) = mpsc::channel();
|
||||
@@ -1119,7 +1151,7 @@ mod windows_tun {
|
||||
session: Arc<wintun::Session>,
|
||||
rx: mpsc::Receiver<Vec<u8>>,
|
||||
name: String,
|
||||
max_mss: u16,
|
||||
max_mss: MssCeiling,
|
||||
path_mtu_lookup: PathMtuLookup,
|
||||
}
|
||||
|
||||
@@ -1132,14 +1164,21 @@ mod windows_tun {
|
||||
use super::per_flow_max_mss;
|
||||
use crate::upper::tcp_mss::clamp_tcp_mss;
|
||||
|
||||
debug!(name = %self.name, max_mss = self.max_mss, "TUN writer starting");
|
||||
debug!(
|
||||
name = %self.name,
|
||||
max_mss = self.max_mss.load(std::sync::atomic::Ordering::Relaxed),
|
||||
"TUN writer starting"
|
||||
);
|
||||
|
||||
for mut packet in self.rx {
|
||||
// Read per packet, not once: a transport binding or unbinding
|
||||
// moves the node's egress floor. See `MssCeiling`.
|
||||
let global_max_mss = self.max_mss.load(std::sync::atomic::Ordering::Relaxed);
|
||||
// Per-destination clamp (peer source IPv6 = bytes 8..24)
|
||||
let effective_max_mss = if packet.len() >= 24 {
|
||||
per_flow_max_mss(&self.path_mtu_lookup, &packet[8..24], self.max_mss)
|
||||
per_flow_max_mss(&self.path_mtu_lookup, &packet[8..24], global_max_mss)
|
||||
} else {
|
||||
self.max_mss
|
||||
global_max_mss
|
||||
};
|
||||
// Clamp TCP MSS on inbound SYN-ACK packets
|
||||
if clamp_tcp_mss(&mut packet, effective_max_mss) {
|
||||
@@ -1188,17 +1227,17 @@ mod windows_tun {
|
||||
our_addr: FipsAddress,
|
||||
tun_tx: TunTx,
|
||||
outbound_tx: TunOutboundTx,
|
||||
transport_mtu: u16,
|
||||
max_mss: MssCeiling,
|
||||
path_mtu_lookup: PathMtuLookup,
|
||||
) {
|
||||
let (name, mut buf, max_mss) = super::tun_reader_setup(device.name(), mtu, transport_mtu);
|
||||
let (name, mut buf) = super::tun_reader_setup(device.name(), mtu, &max_mss);
|
||||
|
||||
loop {
|
||||
match device.read_packet(&mut buf) {
|
||||
Ok(n) if n > 0 => {
|
||||
if !super::handle_tun_packet(
|
||||
&mut buf[..n],
|
||||
max_mss,
|
||||
max_mss.load(std::sync::atomic::Ordering::Relaxed),
|
||||
&name,
|
||||
our_addr,
|
||||
&tun_tx,
|
||||
|
||||
@@ -85,6 +85,15 @@ End-to-end exercise of the production `fips0` nftables baseline at
|
||||
`packaging/common/fips.nft`, covering the default-deny, conntrack and
|
||||
drop-in semantics.
|
||||
|
||||
### [iface-binding/](iface-binding/) -- Dynamic Interface Binding
|
||||
|
||||
Two nodes whose only transports are interface-bound, started before the
|
||||
interface they name exists. Asserts the boot race (the daemon comes up
|
||||
`Degraded` and binds when the interface appears, with no restart), the flap
|
||||
(down/up in both directions), destroy-and-recreate, that an `optional`
|
||||
interface's absence never moves node health, and that absence is logged once
|
||||
on the edge rather than once per retry.
|
||||
|
||||
### [acl-allowlist/](acl-allowlist/) -- Peer ACL Enforcement
|
||||
|
||||
Six nodes with per-node allowlist files mounted at the runtime ACL
|
||||
@@ -146,6 +155,23 @@ matrix; a divergence fails the run. GitHub
|
||||
runs the same check as its own `ci-parity` job. `--check-parity` runs it
|
||||
alone (see [check-ci-parity.sh](check-ci-parity.sh)).
|
||||
|
||||
Note that `ci-local.sh` covers the integration suites and the glibc unit
|
||||
tests. GitHub additionally runs the library tests on macOS, Windows and
|
||||
**musl** (built for the musl target and run natively), and a `--features
|
||||
profiling` pass; the musl
|
||||
leg exists because interface presence is built on `getifaddrs`/`ifa_flags`,
|
||||
which musl reimplements independently, and OpenWrt is a musl target. A local
|
||||
green run does not certify those four.
|
||||
|
||||
The Linux and musl legs also create an address-less dummy interface
|
||||
(`fips-probe0`) and pass its name to the tests as
|
||||
`FIPS_TEST_ADDRLESS_IFACE`. That is the one assumption the interface-binding
|
||||
mechanism rests on that no ordinary test can reach: loopback has addresses, so
|
||||
probing it asks whether `getifaddrs` works rather than whether it reports an
|
||||
interface that has none — which is exactly what `fips-mesh0` and `fips-ap0`
|
||||
are on OpenWrt. Set the variable by hand to run the assertion locally against
|
||||
an interface you have created; leave it unset and the assertion does not run.
|
||||
|
||||
### Per-run isolation and the `FIPS_CI_RUN_ID` override
|
||||
|
||||
Every invocation derives a **run id** and scopes all of its Docker
|
||||
|
||||
@@ -0,0 +1,127 @@
|
||||
# Ethernet rebind under active traffic
|
||||
#
|
||||
# The one case dynamic interface binding has no coverage for anywhere:
|
||||
# traffic crossing an Ethernet link while the interface underneath it goes
|
||||
# away and comes back.
|
||||
#
|
||||
# The other Ethernet scenarios cannot reach it. `ethernet-only` and
|
||||
# `ethernet-mesh` both run with `traffic.enabled: false`, so no datagram
|
||||
# crosses an Ethernet link in either — framing, the length field that trims
|
||||
# NIC minimum-frame padding, and AEAD over Ethernet are all control-plane
|
||||
# assumptions there. And `link_flaps` cannot produce a rebind whatever it is
|
||||
# pointed at: it simulates a down link with netem 100% loss, so the interface
|
||||
# stays IFF_UP and the presence machine never sees an edge.
|
||||
#
|
||||
# `node_churn` is what actually moves an interface. Stopping a container
|
||||
# destroys its network namespace, which deletes every veth in it — and
|
||||
# deleting one end of a veth deletes its peer — so a *surviving* node watches
|
||||
# its Ethernet interface disappear outright. On restart the harness recreates
|
||||
# the pair (`NodeChurnManager._start_node`), and the survivor watches it come
|
||||
# back. That is a real detach and a real rebind, driven from outside the
|
||||
# daemon, with iperf3 running across the mesh throughout.
|
||||
#
|
||||
# Topology: a 4-node ring, so removing any single node leaves the remaining
|
||||
# three connected in a line and `protect_connectivity` has something to
|
||||
# protect.
|
||||
#
|
||||
# n01 ---eth--- n02
|
||||
# | |
|
||||
# eth eth
|
||||
# | |
|
||||
# n04 ---eth--- n03
|
||||
|
||||
scenario:
|
||||
name: "ethernet-churn"
|
||||
seed: 42
|
||||
duration_secs: 240
|
||||
|
||||
topology:
|
||||
algorithm: explicit
|
||||
num_nodes: 4
|
||||
default_transport: ethernet
|
||||
params:
|
||||
adjacency:
|
||||
- [n01, n02]
|
||||
- [n02, n03]
|
||||
- [n03, n04]
|
||||
- [n04, n01]
|
||||
|
||||
# Mild, and deliberately so. The variable under test is the interface going
|
||||
# away, not the link being bad while it is there; heavy loss here would make a
|
||||
# traffic shortfall ambiguous between the two.
|
||||
netem:
|
||||
enabled: true
|
||||
default_policy:
|
||||
delay_ms: [1, 5]
|
||||
jitter_ms: [0, 1]
|
||||
loss_pct: [0, 0.5]
|
||||
|
||||
# Off on purpose. A netem-simulated down link would add outage that never
|
||||
# reaches the presence machine, which is the opposite of what this isolates.
|
||||
link_flaps:
|
||||
enabled: false
|
||||
|
||||
traffic:
|
||||
enabled: true
|
||||
max_concurrent: 2
|
||||
interval_secs: {min: 10, max: 20}
|
||||
duration_secs: {min: 15, max: 25}
|
||||
parallel_streams: 2
|
||||
|
||||
# The mechanism. One node down at a time, long enough to outlast the
|
||||
# ten-second bring-up window so the absence is a real one rather than a race,
|
||||
# and to give traffic time to run against the reduced mesh before it returns.
|
||||
node_churn:
|
||||
enabled: true
|
||||
interval_secs: {min: 45, max: 60}
|
||||
max_down_nodes: 1
|
||||
down_duration_secs: {min: 20, max: 35}
|
||||
protect_connectivity: true
|
||||
|
||||
assertions:
|
||||
# The mesh re-forms after each interface comes back, fully.
|
||||
#
|
||||
# Calibrated 2026-09-10 against sixteen runs, eight on master-line code and
|
||||
# eight on next-line code, on a harness that fails a veth restore it cannot
|
||||
# complete and waits for the tree to settle before the final snapshot. Twelve
|
||||
# ran beside the other five CI chaos scenarios and four CPU busy loops; seeds
|
||||
# 42 (twelve runs), 7 and 1234:
|
||||
#
|
||||
# nodes answering 4 in all sixteen
|
||||
# distinct roots 1 in all sixteen
|
||||
# nodes parented 3 in all sixteen
|
||||
# settle 11-22 s
|
||||
# traffic 4-11 sessions per run, 3.8-11.7 GB
|
||||
#
|
||||
# So the thresholds are full convergence: every node answers, one root, three
|
||||
# parented. A red here is a mesh that did not re-form.
|
||||
#
|
||||
# READ BEFORE RETUNING: the 2026-09-02 calibration recorded 2 roots and 2
|
||||
# parented in all four of its runs, and that was not the daemon. The veth
|
||||
# restore then failed silently whenever a stopped node's namespace outlived
|
||||
# the stop, and in every later run that was inspected, the node islanded at
|
||||
# the snapshot was one whose links the harness had left down. Before widening these, read the red run's runner.log for a
|
||||
# harness fault; a failed restore now aborts the run rather than reaching
|
||||
# this assertion.
|
||||
baseline:
|
||||
min_nodes_reporting: 4
|
||||
max_roots: 1
|
||||
min_nodes_parented: 3
|
||||
|
||||
# The point of the scenario. Traffic must actually have crossed Ethernet
|
||||
# links while interfaces were being taken away underneath it — a green tree
|
||||
# with zero bytes moved is the failure this catches, and is exactly what a
|
||||
# control-plane-only assertion would have called a pass.
|
||||
min_traffic:
|
||||
min_sessions_ok: 2
|
||||
|
||||
# Chaos Ethernet transports are `optional: true` (see config_gen), precisely
|
||||
# because a neighbour's interface disappearing is the scenario here rather
|
||||
# than a fault. So absence must stay silent: any ERROR means something other
|
||||
# than the churn went wrong.
|
||||
max_errors:
|
||||
max_total: 0
|
||||
|
||||
logging:
|
||||
rust_log: "info"
|
||||
output_dir: "./sim-results"
|
||||
@@ -19,6 +19,7 @@ from dataclasses import dataclass
|
||||
|
||||
from .control import snapshot_all_bloom
|
||||
from .scenario import (
|
||||
MinTrafficAssertion,
|
||||
BaselineAssertion,
|
||||
BloomSendRateAssertion,
|
||||
CongestionSignalsAssertion,
|
||||
@@ -442,3 +443,61 @@ def evaluate_min_parent_switches(
|
||||
f"node win the root election?"
|
||||
),
|
||||
)
|
||||
|
||||
|
||||
def _session_bytes(result: dict) -> int:
|
||||
"""Bytes actually received in one iperf3 session, or 0 if it failed.
|
||||
|
||||
iperf3 reports a failed run as a top-level ``error`` string with no
|
||||
``end`` block, and a killed-but-partial run still carries whatever it
|
||||
managed. Both are handled by reading the received total and treating
|
||||
anything missing as zero, so a session only counts when it moved bytes.
|
||||
"""
|
||||
if not isinstance(result, dict) or result.get("error"):
|
||||
return 0
|
||||
end = result.get("end")
|
||||
if not isinstance(end, dict):
|
||||
return 0
|
||||
summary = end.get("sum_received") or end.get("sum_sent")
|
||||
if not isinstance(summary, dict):
|
||||
return 0
|
||||
value = summary.get("bytes", 0)
|
||||
return value if isinstance(value, int) and value > 0 else 0
|
||||
|
||||
|
||||
def evaluate_min_traffic(
|
||||
cfg: MinTrafficAssertion,
|
||||
results: list[dict],
|
||||
) -> AssertionOutcome:
|
||||
"""Floor on iperf3 sessions that actually carried data.
|
||||
|
||||
Without this the traffic generator is decoration: the results were
|
||||
written to disk and never read, so a scenario whose every session
|
||||
failed still passed on a healthy control plane. A rebind under load is
|
||||
exactly the case a tree snapshot cannot see.
|
||||
"""
|
||||
per_session = [_session_bytes(r) for r in results]
|
||||
ok = [b for b in per_session if b > 0]
|
||||
total = sum(ok)
|
||||
|
||||
if len(ok) >= cfg.min_sessions_ok and total >= cfg.min_bytes_total:
|
||||
return AssertionOutcome(
|
||||
name="min_traffic",
|
||||
passed=True,
|
||||
detail=(
|
||||
f"PASS min_traffic: {len(ok)}/{len(results)} session(s) moved "
|
||||
f"data (need {cfg.min_sessions_ok}); {total} byte(s) total "
|
||||
f"(need {cfg.min_bytes_total})"
|
||||
),
|
||||
)
|
||||
|
||||
return AssertionOutcome(
|
||||
name="min_traffic",
|
||||
passed=False,
|
||||
detail=(
|
||||
f"FAIL min_traffic: {len(ok)}/{len(results)} session(s) moved data "
|
||||
f"(need {cfg.min_sessions_ok}); {total} byte(s) total (need "
|
||||
f"{cfg.min_bytes_total}). A green control plane with no traffic "
|
||||
f"means the data path did not survive what the scenario did to it."
|
||||
),
|
||||
)
|
||||
|
||||
@@ -64,7 +64,22 @@ def generate_peers_block(
|
||||
|
||||
|
||||
def _build_ethernet_config(iface: str) -> dict:
|
||||
"""Build an Ethernet transport config dict for a single interface."""
|
||||
"""Build an Ethernet transport config dict for a single interface.
|
||||
|
||||
``optional: True`` because in this harness a neighbour's interface
|
||||
disappearing is the scenario, not a fault. ``node_churn`` stops a
|
||||
container, which destroys its netns and with it both ends of every veth
|
||||
pair it held (see ``nodes.py``: the veths are recreated on restart), so a
|
||||
surviving node watches a *required* interface vanish for the 30-90s the
|
||||
neighbour is down -- once per churn event, on every neighbour. The daemon
|
||||
reports a required interface absent past its bring-up window at ERROR,
|
||||
which is correct for a deployment and wrong for a harness that tears the
|
||||
interface down on purpose; the mesh-wide zero-ERROR ceiling would fail on
|
||||
injected chaos rather than on a defect.
|
||||
|
||||
Absence behaviour itself is asserted in ``testing/iface-binding/``, which
|
||||
exists for it and drives both policies deliberately.
|
||||
"""
|
||||
return {
|
||||
"interface": iface,
|
||||
"listen": True,
|
||||
@@ -72,6 +87,7 @@ def _build_ethernet_config(iface: str) -> dict:
|
||||
"auto_connect": True,
|
||||
"accept_connections": True,
|
||||
"beacon_interval_secs": 10,
|
||||
"optional": True,
|
||||
}
|
||||
|
||||
|
||||
|
||||
+30
-10
@@ -370,22 +370,42 @@ class NetemManager:
|
||||
# Re-apply Ethernet veth netem
|
||||
veth_states = self.veth_states.get(container)
|
||||
if veth_states:
|
||||
for iface, state in veth_states.items():
|
||||
cmd = (
|
||||
f"tc qdisc del dev {iface} root 2>/dev/null || true && "
|
||||
f"tc qdisc add dev {iface} root netem {state.params.to_tc_args()}"
|
||||
)
|
||||
result = docker_exec_quiet(container, cmd, timeout=10)
|
||||
if result is not None:
|
||||
log.debug("Re-applied veth netem on %s:%s", container, iface)
|
||||
else:
|
||||
log.warning("Failed to re-apply veth netem on %s:%s", container, iface)
|
||||
for state in veth_states.values():
|
||||
self._apply_veth(state)
|
||||
log.info(
|
||||
"Re-applied veth netem on %s (%d Ethernet peers)",
|
||||
container,
|
||||
len(veth_states),
|
||||
)
|
||||
|
||||
# And on each running neighbour's end of those links. The restore
|
||||
# recreates the whole pair, so the survivor's end is a new interface
|
||||
# with no qdisc, and without this that direction of every restored
|
||||
# link ran unshaped for the rest of the run.
|
||||
for peer_id in sorted(self.topology.nodes[node_id].peers):
|
||||
if peer_id in self.down_nodes:
|
||||
continue
|
||||
if self.topology.transport_for_edge(node_id, peer_id) != "ethernet":
|
||||
continue
|
||||
peer_container = self.topology.container_name(peer_id)
|
||||
state = self.veth_states.get(peer_container, {}).get(
|
||||
veth_interface_name(peer_id, node_id)
|
||||
)
|
||||
if state is not None:
|
||||
self._apply_veth(state)
|
||||
|
||||
def _apply_veth(self, state: VethNetemState):
|
||||
"""Install a veth end's current netem parameters as its root qdisc."""
|
||||
cmd = (
|
||||
f"tc qdisc del dev {state.iface} root 2>/dev/null || true && "
|
||||
f"tc qdisc add dev {state.iface} root netem {state.params.to_tc_args()}"
|
||||
)
|
||||
result = docker_exec_quiet(state.container, cmd, timeout=10)
|
||||
if result is not None:
|
||||
log.debug("Re-applied veth netem on %s:%s", state.container, state.iface)
|
||||
else:
|
||||
log.warning("Failed to re-apply veth netem on %s:%s", state.container, state.iface)
|
||||
|
||||
def mutate(self):
|
||||
"""Randomly mutate netem params on a fraction of links."""
|
||||
if not self.config.mutation.policies:
|
||||
|
||||
@@ -50,7 +50,10 @@ class NodeManager:
|
||||
self.rng = rng
|
||||
self.netem_mgr = netem_mgr
|
||||
self.veth_mgr = veth_mgr
|
||||
self.down_nodes = down_nodes or set()
|
||||
# `is not None`, not `or`: the runner passes its shared set while it
|
||||
# is still empty, and an empty set is falsy, so `or` replaced it with
|
||||
# a private one and no other manager ever saw a node go down.
|
||||
self.down_nodes = down_nodes if down_nodes is not None else set()
|
||||
self.on_node_restart = on_node_restart
|
||||
self.node_states: dict[str, NodeState] = {
|
||||
nid: NodeState(node_id=nid) for nid in topology.nodes
|
||||
@@ -173,7 +176,11 @@ class NodeManager:
|
||||
# Re-create veth pairs (container restart destroys netns)
|
||||
if self.veth_mgr:
|
||||
time.sleep(1)
|
||||
self.veth_mgr.setup_node(node_id)
|
||||
# Churn's own record, not the shared set: netem's safety net also
|
||||
# adds to that set, on a docker hiccup or a crash churn did not
|
||||
# cause, and a pair deferred on that basis would never be rebuilt.
|
||||
stopped = {nid for nid, ns in self.node_states.items() if ns.is_down}
|
||||
self.veth_mgr.setup_node(node_id, stopped)
|
||||
|
||||
# Re-apply netem after a brief delay for the container to initialize
|
||||
if self.netem_mgr:
|
||||
|
||||
+100
-20
@@ -20,6 +20,7 @@ from .assertions import (
|
||||
evaluate_max_errors,
|
||||
evaluate_max_parent_switches,
|
||||
evaluate_min_parent_switches,
|
||||
evaluate_min_traffic,
|
||||
evaluate_tree_parents,
|
||||
)
|
||||
from .compose import generate_compose
|
||||
@@ -41,11 +42,21 @@ from .veth import VethManager
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
# The final snapshot waits for this many identical consecutive tree reads,
|
||||
# taken this far apart, so the tree must hold still for two intervals.
|
||||
SETTLE_READS = 3
|
||||
SETTLE_INTERVAL_SECS = 5
|
||||
SETTLE_TIMEOUT_SECS = 90
|
||||
|
||||
|
||||
class SimRunner:
|
||||
def __init__(self, scenario: Scenario):
|
||||
self.scenario = scenario
|
||||
# Setup draws only: the topology and the ephemeral node choice, made
|
||||
# once and in a fixed order. Everything drawn while the simulation
|
||||
# runs comes from `_stream` instead.
|
||||
self.rng = random.Random(scenario.seed)
|
||||
self._streams: dict[str, random.Random] = {}
|
||||
self.topology: SimTopology | None = None
|
||||
self.compose_file: str | None = None
|
||||
# Claimed in _setup; the compose file refers to it as external, so
|
||||
@@ -316,14 +327,16 @@ class SimRunner:
|
||||
if s.netem.enabled:
|
||||
bw = s.bandwidth if s.bandwidth.enabled else None
|
||||
ig = s.ingress if s.ingress.enabled else None
|
||||
self.netem_mgr = NetemManager(self.topology, s.netem, self.rng, bandwidth=bw, ingress=ig)
|
||||
self.netem_mgr = NetemManager(
|
||||
self.topology, s.netem, self._stream("netem"), bandwidth=bw, ingress=ig
|
||||
)
|
||||
self.netem_mgr.down_nodes = self._down_nodes
|
||||
log.info("Applying initial per-link netem...")
|
||||
self.netem_mgr.setup_initial()
|
||||
|
||||
if s.link_flaps.enabled:
|
||||
self.link_mgr = LinkManager(
|
||||
self.topology, s.link_flaps, self.rng, netem_mgr=self.netem_mgr
|
||||
self.topology, s.link_flaps, self._stream("flaps"), netem_mgr=self.netem_mgr
|
||||
)
|
||||
|
||||
if s.link_swap.enabled:
|
||||
@@ -332,7 +345,7 @@ class SimRunner:
|
||||
"link_swap requires netem.enabled (depends on per-link tc state)"
|
||||
)
|
||||
self.link_swap_mgr = LinkSwapManager(
|
||||
self.topology, s.link_swap, self.netem_mgr, self.rng,
|
||||
self.topology, s.link_swap, self.netem_mgr, self._stream("swap"),
|
||||
)
|
||||
|
||||
if s.assertions.bloom_send_rate is not None:
|
||||
@@ -342,12 +355,12 @@ class SimRunner:
|
||||
|
||||
if s.traffic.enabled:
|
||||
self.traffic_mgr = TrafficManager(
|
||||
self.topology, s.traffic, self.rng, down_nodes=self._down_nodes
|
||||
self.topology, s.traffic, self._stream("traffic"), down_nodes=self._down_nodes
|
||||
)
|
||||
|
||||
if s.node_churn.enabled:
|
||||
self.node_mgr = NodeManager(
|
||||
self.topology, s.node_churn, self.rng,
|
||||
self.topology, s.node_churn, self._stream("churn"),
|
||||
netem_mgr=self.netem_mgr, down_nodes=self._down_nodes,
|
||||
veth_mgr=self.veth_mgr,
|
||||
on_node_restart=self._handle_node_restart,
|
||||
@@ -355,7 +368,7 @@ class SimRunner:
|
||||
|
||||
if s.peer_churn.enabled:
|
||||
self.peer_churn_mgr = PeerChurnManager(
|
||||
self.topology, s.peer_churn, self.rng,
|
||||
self.topology, s.peer_churn, self._stream("peer-churn"),
|
||||
down_nodes=self._down_nodes,
|
||||
ephemeral_nodes=self._ephemeral_nodes,
|
||||
)
|
||||
@@ -410,11 +423,11 @@ class SimRunner:
|
||||
log.info("Simulation running for %ds...", duration)
|
||||
|
||||
# Schedule first events
|
||||
next_netem = self._schedule_next(start, s.netem.mutation.interval_secs) if self.netem_mgr else float("inf")
|
||||
next_flap = self._schedule_next(start, s.link_flaps.interval_secs) if self.link_mgr else float("inf")
|
||||
next_traffic = self._schedule_next(start, s.traffic.interval_secs) if self.traffic_mgr else float("inf")
|
||||
next_churn = self._schedule_next(start, s.node_churn.interval_secs) if self.node_mgr else float("inf")
|
||||
next_peer_churn = self._schedule_next(start, s.peer_churn.interval_secs) if self.peer_churn_mgr else float("inf")
|
||||
next_netem = self._schedule_next(start, s.netem.mutation.interval_secs, "netem") if self.netem_mgr else float("inf")
|
||||
next_flap = self._schedule_next(start, s.link_flaps.interval_secs, "flaps") if self.link_mgr else float("inf")
|
||||
next_traffic = self._schedule_next(start, s.traffic.interval_secs, "traffic") if self.traffic_mgr else float("inf")
|
||||
next_churn = self._schedule_next(start, s.node_churn.interval_secs, "churn") if self.node_mgr else float("inf")
|
||||
next_peer_churn = self._schedule_next(start, s.peer_churn.interval_secs, "peer-churn") if self.peer_churn_mgr else float("inf")
|
||||
|
||||
# Bloom-send-rate assertion: sample at window_secs before end.
|
||||
bloom_window_start_at = float("inf")
|
||||
@@ -451,34 +464,34 @@ class SimRunner:
|
||||
# Netem mutation
|
||||
if self.netem_mgr and now >= next_netem:
|
||||
self.netem_mgr.mutate()
|
||||
next_netem = self._schedule_next(now, s.netem.mutation.interval_secs)
|
||||
next_netem = self._schedule_next(now, s.netem.mutation.interval_secs, "netem")
|
||||
|
||||
# Link flaps
|
||||
if self.link_mgr:
|
||||
if now >= next_flap:
|
||||
self.link_mgr.maybe_flap()
|
||||
next_flap = self._schedule_next(now, s.link_flaps.interval_secs)
|
||||
next_flap = self._schedule_next(now, s.link_flaps.interval_secs, "flaps")
|
||||
self.link_mgr.restore_expired()
|
||||
|
||||
# Traffic generation
|
||||
if self.traffic_mgr:
|
||||
if now >= next_traffic:
|
||||
self.traffic_mgr.maybe_spawn()
|
||||
next_traffic = self._schedule_next(now, s.traffic.interval_secs)
|
||||
next_traffic = self._schedule_next(now, s.traffic.interval_secs, "traffic")
|
||||
self.traffic_mgr.cleanup_expired()
|
||||
|
||||
# Node churn
|
||||
if self.node_mgr:
|
||||
if now >= next_churn:
|
||||
self.node_mgr.maybe_kill()
|
||||
next_churn = self._schedule_next(now, s.node_churn.interval_secs)
|
||||
next_churn = self._schedule_next(now, s.node_churn.interval_secs, "churn")
|
||||
self.node_mgr.restore_expired()
|
||||
|
||||
# Peer churn (topology mutation)
|
||||
if self.peer_churn_mgr:
|
||||
if now >= next_peer_churn:
|
||||
self.peer_churn_mgr.maybe_churn()
|
||||
next_peer_churn = self._schedule_next(now, s.peer_churn.interval_secs)
|
||||
next_peer_churn = self._schedule_next(now, s.peer_churn.interval_secs, "peer-churn")
|
||||
|
||||
# Status line
|
||||
down_links = self.link_mgr.down_count if self.link_mgr else 0
|
||||
@@ -699,6 +712,7 @@ class SimRunner:
|
||||
self.node_mgr.restore_all()
|
||||
|
||||
# Collect iperf3 throughput results before containers stop
|
||||
iperf_results: list[dict] = []
|
||||
if self.traffic_mgr:
|
||||
iperf_results = self.traffic_mgr.collect_results()
|
||||
if iperf_results:
|
||||
@@ -707,7 +721,11 @@ class SimRunner:
|
||||
json.dump(iperf_results, f, indent=2)
|
||||
log.info("Saved %d iperf3 results to %s", len(iperf_results), iperf_path)
|
||||
|
||||
# Take final tree snapshot while nodes are still running
|
||||
# Take final tree snapshot while nodes are still running, once the
|
||||
# tree has stopped moving. A node restored a moment ago is its own
|
||||
# root until it re-parents, so a snapshot taken straight after the
|
||||
# restore reads a mesh still converging.
|
||||
self._settle_tree()
|
||||
self._take_snapshot("final")
|
||||
|
||||
# Collect logs before stopping containers
|
||||
@@ -758,6 +776,13 @@ class SimRunner:
|
||||
if err_cfg is not None:
|
||||
outcome = evaluate_max_errors(err_cfg, result.errors)
|
||||
self.assertion_outcomes.append(outcome)
|
||||
|
||||
# Traffic. Evaluated even when no session completed, because
|
||||
# "nothing ran" is the failure this exists to catch.
|
||||
traffic_cfg = self.scenario.assertions.min_traffic
|
||||
if traffic_cfg is not None:
|
||||
outcome = evaluate_min_traffic(traffic_cfg, iperf_results)
|
||||
self.assertion_outcomes.append(outcome)
|
||||
if outcome.passed:
|
||||
log.info("%s", outcome.detail)
|
||||
else:
|
||||
@@ -838,6 +863,46 @@ class SimRunner:
|
||||
|
||||
return result
|
||||
|
||||
def _settle_tree(self):
|
||||
"""Wait until consecutive tree reads agree, or the settle time runs out.
|
||||
|
||||
Compares each answering node's root and parent. A fixed delay would
|
||||
either waste time on a mesh that settled at once or cut off one that
|
||||
had not. Running out is logged and is not a failure in itself: the
|
||||
final snapshot is taken anyway, and the assertions judge what it
|
||||
shows.
|
||||
|
||||
A read that no node answered never counts toward agreement. It does
|
||||
not catch a node that stays its own root for longer than the reads
|
||||
span, which is a tree that is stable and wrong, and is left to the
|
||||
assertions.
|
||||
"""
|
||||
started = time.time()
|
||||
previous = None
|
||||
agreeing = 0
|
||||
while not self._interrupted:
|
||||
trees = snapshot_all_trees(self.topology)
|
||||
shape = {
|
||||
nid: (data.get("root"), data.get("parent"))
|
||||
for nid, data in trees.items()
|
||||
}
|
||||
if not shape:
|
||||
agreeing = 0
|
||||
else:
|
||||
agreeing = agreeing + 1 if shape == previous else 1
|
||||
previous = shape
|
||||
waited = time.time() - started
|
||||
if agreeing >= SETTLE_READS:
|
||||
log.info("Tree settled after %.0fs", waited)
|
||||
return
|
||||
if waited >= SETTLE_TIMEOUT_SECS:
|
||||
log.warning(
|
||||
"Tree still changing after %.0fs; taking the final snapshot anyway",
|
||||
waited,
|
||||
)
|
||||
return
|
||||
self._sleep(SETTLE_INTERVAL_SECS)
|
||||
|
||||
def _take_snapshot(self, label: str):
|
||||
"""Query all nodes via control socket and save tree/MMP/congestion snapshots."""
|
||||
if not self.topology:
|
||||
@@ -874,9 +939,24 @@ class SimRunner:
|
||||
len(self.topology.nodes),
|
||||
)
|
||||
|
||||
def _schedule_next(self, now: float, interval) -> float:
|
||||
"""Schedule the next event using a Range interval."""
|
||||
return now + self.rng.uniform(interval.min, interval.max)
|
||||
def _schedule_next(self, now: float, interval, kind: str) -> float:
|
||||
"""Schedule the next event of one kind using a Range interval."""
|
||||
return now + self._stream(f"{kind}-schedule").uniform(interval.min, interval.max)
|
||||
|
||||
def _stream(self, name: str) -> random.Random:
|
||||
"""Return the random stream for one consumer, derived from the seed.
|
||||
|
||||
One stream shared by every manager was drawn in wall-clock order, and
|
||||
a manager that returns early draws nothing, so host load changed
|
||||
which node the churn stopped next. Under heavy host load that walked
|
||||
the stops around the ring. A stream per consumer means one manager's
|
||||
draws no longer shift another's. It does not make a schedule a
|
||||
function of the seed alone: a manager's own draws can still depend on
|
||||
mesh state at the tick, such as which nodes are down.
|
||||
"""
|
||||
if name not in self._streams:
|
||||
self._streams[name] = random.Random(f"{self.scenario.seed}:{name}")
|
||||
return self._streams[name]
|
||||
|
||||
def _sleep(self, seconds: float):
|
||||
"""Sleep in small increments so SIGINT can break out."""
|
||||
|
||||
@@ -323,6 +323,27 @@ class MaxErrorsAssertion:
|
||||
max_total: int = 0
|
||||
|
||||
|
||||
@dataclass
|
||||
class MinTrafficAssertion:
|
||||
"""Floor on how much iperf3 traffic actually completed.
|
||||
|
||||
Traffic has always been generated and its results saved to
|
||||
``iperf3-results.json``, but nothing read them: a scenario could carry
|
||||
``traffic.enabled: true``, have every single session fail, and still
|
||||
exit 0 on a green control plane. That gap matters most for the
|
||||
scenarios where traffic is the point — a datagram crossing an Ethernet
|
||||
link while its interface rebinds underneath is not observable in the
|
||||
tree snapshot at all.
|
||||
|
||||
``min_sessions_ok`` counts sessions that finished with bytes actually
|
||||
received. ``min_bytes_total`` is the aggregate floor across them; 0
|
||||
disables it and leaves the session count as the only gate.
|
||||
"""
|
||||
|
||||
min_sessions_ok: int = 1
|
||||
min_bytes_total: int = 0
|
||||
|
||||
|
||||
@dataclass
|
||||
class AssertionsConfig:
|
||||
"""Optional post-run assertions evaluated against control-socket data."""
|
||||
@@ -334,6 +355,7 @@ class AssertionsConfig:
|
||||
congestion_signals: CongestionSignalsAssertion | None = None
|
||||
tree_parents: TreeParentsAssertion | None = None
|
||||
baseline: BaselineAssertion | None = None
|
||||
min_traffic: MinTrafficAssertion | None = None
|
||||
|
||||
|
||||
@dataclass
|
||||
@@ -414,6 +436,7 @@ _SECTION_KEYS = {
|
||||
"assertions": {
|
||||
"bloom_send_rate", "min_parent_switches", "max_parent_switches",
|
||||
"max_errors", "congestion_signals", "tree_parents", "baseline",
|
||||
"min_traffic",
|
||||
},
|
||||
"logging": {"rust_log", "output_dir"},
|
||||
}
|
||||
@@ -422,6 +445,7 @@ _ASSERTION_KEYS = {
|
||||
"min_parent_switches": {"min_total"},
|
||||
"max_parent_switches": {"max_total", "node"},
|
||||
"max_errors": {"max_total"},
|
||||
"min_traffic": {"min_sessions_ok", "min_bytes_total"},
|
||||
"congestion_signals": {
|
||||
"min_nodes_detected", "min_nodes_ce_forwarded", "min_nodes_ce_received",
|
||||
},
|
||||
@@ -739,6 +763,27 @@ def load_scenario(path: str) -> Scenario:
|
||||
f"got {err_total}"
|
||||
)
|
||||
s.assertions.max_errors = MaxErrorsAssertion(max_total=err_total)
|
||||
|
||||
if "min_traffic" in asrt:
|
||||
mt = asrt["min_traffic"]
|
||||
_reject_unknown(
|
||||
mt, _ASSERTION_KEYS["min_traffic"], "assertions.min_traffic",
|
||||
)
|
||||
sessions = mt.get("min_sessions_ok", 1)
|
||||
if not isinstance(sessions, int) or isinstance(sessions, bool) or sessions < 1:
|
||||
raise ValueError(
|
||||
"assertions.min_traffic: min_sessions_ok must be a positive "
|
||||
f"integer, got {sessions!r} — a floor of zero asserts nothing"
|
||||
)
|
||||
min_bytes = mt.get("min_bytes_total", 0)
|
||||
if not isinstance(min_bytes, int) or isinstance(min_bytes, bool) or min_bytes < 0:
|
||||
raise ValueError(
|
||||
"assertions.min_traffic: min_bytes_total must be a "
|
||||
f"non-negative integer, got {min_bytes!r}"
|
||||
)
|
||||
s.assertions.min_traffic = MinTrafficAssertion(
|
||||
min_sessions_ok=sessions, min_bytes_total=min_bytes,
|
||||
)
|
||||
else:
|
||||
# Default-on. See MaxErrorsAssertion for why this one assertion is
|
||||
# applied without being asked for: it is the floor on what a green
|
||||
@@ -835,14 +880,21 @@ def load_scenario(path: str) -> Scenario:
|
||||
+ ", ".join(sorted(_ASSERTION_KEYS["baseline"]))
|
||||
+ "; a block with none asserts nothing"
|
||||
)
|
||||
if vals.get("min_nodes_parented", 0) > 0 or vals.get("max_roots"):
|
||||
if vals.get("min_nodes_parented", 0) > 0:
|
||||
# Every root is its own parent, so a mesh with R roots has at most
|
||||
# n - R nodes parented. Checked against the roots ceiling rather
|
||||
# than against one root: a floor above n - max_roots reds a run
|
||||
# that sits exactly at the ceiling the file itself allows, so the
|
||||
# two thresholds contradict each other there.
|
||||
n = s.topology.num_nodes
|
||||
if vals.get("min_nodes_parented", 0) > n - 1:
|
||||
roots = vals.get("max_roots", 1)
|
||||
floor = vals["min_nodes_parented"]
|
||||
if floor > n - roots:
|
||||
raise ValueError(
|
||||
f"assertions.baseline.min_nodes_parented: "
|
||||
f"{vals['min_nodes_parented']} exceeds {n - 1}, the most a "
|
||||
f"{n}-node mesh can reach — the root is its own parent, so "
|
||||
f"this could never pass"
|
||||
f"assertions.baseline.min_nodes_parented: {floor} exceeds "
|
||||
f"{n - roots}, the most a {n}-node mesh with {roots} "
|
||||
f"root(s) can reach — each root is its own parent, so a "
|
||||
f"run at the max_roots ceiling could never pass"
|
||||
)
|
||||
s.assertions.baseline = BaselineAssertion(**vals)
|
||||
|
||||
|
||||
@@ -44,7 +44,10 @@ class TrafficManager:
|
||||
self.topology = topology
|
||||
self.config = config
|
||||
self.rng = rng
|
||||
self.down_nodes = down_nodes or set()
|
||||
# `is not None`, not `or`: the runner passes its shared set while it
|
||||
# is still empty, and an empty set is falsy, so `or` replaced it with
|
||||
# a private one and no other manager ever saw a node go down.
|
||||
self.down_nodes = down_nodes if down_nodes is not None else set()
|
||||
self.npub_cache = npub_cache or {}
|
||||
self.active_sessions: list[TrafficSession] = []
|
||||
self.completed_results: list[dict] = []
|
||||
|
||||
+177
-52
@@ -34,12 +34,36 @@ from __future__ import annotations
|
||||
|
||||
import logging
|
||||
import subprocess
|
||||
import time
|
||||
|
||||
from .docker_exec import docker_exec_quiet
|
||||
from .docker_exec import DockerExecError, docker_exec, docker_exec_quiet
|
||||
from .topology import SimTopology, veth_interface_name
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
# How long a freshly raised veth end may take to report operstate up. The
|
||||
# kernel publishes the carrier change through linkwatch, which is deferred, so
|
||||
# a read straight after `ip link set up` can still see the old state. Long
|
||||
# enough for at least two reads under host load, since one slow `docker exec`
|
||||
# must not abort a run whose link is up.
|
||||
OPERSTATE_WAIT_SECS = 15
|
||||
|
||||
# Timeout for each `docker exec` a restore makes. The run blocks on these, and
|
||||
# a timeout aborts it, so this errs long: a loaded host is the normal case.
|
||||
EXEC_TIMEOUT_SECS = 30
|
||||
|
||||
|
||||
class VethSetupError(RuntimeError):
|
||||
"""A veth pair could not be created, placed, renamed or raised.
|
||||
|
||||
Raised rather than logged. A pair that fails part way leaves a ring link
|
||||
down while the run carries on, and what reaches the verdict is then a tree
|
||||
that did not converge: a daemon failure in every respect a reader can see.
|
||||
Letting this propagate aborts the run, so the fault is reported as the
|
||||
harness's own.
|
||||
"""
|
||||
|
||||
|
||||
class VethManager:
|
||||
"""Manages veth pairs for Ethernet-transport edges."""
|
||||
|
||||
@@ -102,12 +126,13 @@ class VethManager:
|
||||
len(self._host_pairs),
|
||||
)
|
||||
|
||||
def setup_node(self, node_id: str):
|
||||
def setup_node(self, node_id: str, down_nodes: set[str] | None = None):
|
||||
"""Re-create veth endpoints for a single node after container restart.
|
||||
|
||||
When a container restarts (node churn), its network namespace is
|
||||
destroyed. We re-create the veth pairs for all Ethernet edges
|
||||
involving this node.
|
||||
involving this node. ``down_nodes`` names the neighbours churn has
|
||||
stopped, whose pairs are left for their own restart.
|
||||
"""
|
||||
image = self._get_image()
|
||||
for a, b in self.topology.ethernet_edges():
|
||||
@@ -117,7 +142,7 @@ class VethManager:
|
||||
host_a = self.topology.veth_host_name(a, b, "a")
|
||||
_run_host(["ip", "link", "delete", host_a], image, check=False)
|
||||
# Re-create
|
||||
self._create_veth_pair(a, b, image)
|
||||
self._create_veth_pair(a, b, image, down_nodes or set())
|
||||
|
||||
def teardown_all(self):
|
||||
"""Clean up all veth pairs."""
|
||||
@@ -126,18 +151,40 @@ class VethManager:
|
||||
_run_host(["ip", "link", "delete", host_a], image, check=False)
|
||||
self._host_pairs.clear()
|
||||
|
||||
def _create_veth_pair(self, node_a: str, node_b: str, image: str):
|
||||
"""Create a single veth pair between two containers."""
|
||||
def _create_veth_pair(
|
||||
self, node_a: str, node_b: str, image: str, down_nodes: set[str] | None = None
|
||||
):
|
||||
"""Create a single veth pair between two containers.
|
||||
|
||||
``down_nodes`` is None at first setup, when every container must be
|
||||
running. On a restore it names the nodes churn has stopped.
|
||||
"""
|
||||
container_a = self.topology.container_name(node_a)
|
||||
container_b = self.topology.container_name(node_b)
|
||||
|
||||
# Get container PIDs
|
||||
pid_a = _get_container_pid(container_a)
|
||||
pid_b = _get_container_pid(container_b)
|
||||
pid_a = _container_pid(container_a)
|
||||
pid_b = _container_pid(container_b)
|
||||
if pid_a is None or pid_b is None:
|
||||
log.warning(
|
||||
"Cannot create veth %s--%s: container PID not found", node_a, node_b
|
||||
)
|
||||
stopped = [n for n, pid in ((node_a, pid_a), (node_b, pid_b)) if pid is None]
|
||||
names = ", ".join(stopped)
|
||||
if down_nodes is None:
|
||||
raise VethSetupError(
|
||||
f"veth {node_a}--{node_b}: {names} not running at setup"
|
||||
)
|
||||
if all(n in down_nodes for n in stopped):
|
||||
# Stopped by churn: it has no namespace to join, and its own
|
||||
# restart recreates this pair.
|
||||
log.info("Veth %s--%s deferred: %s down", node_a, node_b, names)
|
||||
else:
|
||||
# Not raised: a container that exited on its own is a daemon
|
||||
# failure, and aborting here would report it as the harness's.
|
||||
# Nothing recreates this pair, so say so loudly.
|
||||
log.warning(
|
||||
"Veth %s--%s not recreated: %s not running, and churn did "
|
||||
"not stop all of them; the link stays absent for the rest "
|
||||
"of the run",
|
||||
node_a, node_b, names,
|
||||
)
|
||||
return
|
||||
|
||||
# Generate names
|
||||
@@ -146,68 +193,146 @@ class VethManager:
|
||||
final_a = veth_interface_name(node_a, node_b)
|
||||
final_b = veth_interface_name(node_b, node_a)
|
||||
|
||||
# Clear both final names and both temporary names out of the
|
||||
# containers first. A stopped node's network namespace can outlive
|
||||
# the stop by minutes, and while it does the survivor still holds its
|
||||
# old `ve-X-Y`, whose peer sits in that namespace; the rename below
|
||||
# then fails with "File exists". Deleting the survivor's end removes
|
||||
# its peer too. A temporary name is left behind only by a restore
|
||||
# that failed after the move, and it blocks the next move the same
|
||||
# way.
|
||||
_purge_links(container_a, [final_a, host_a])
|
||||
_purge_links(container_b, [final_b, host_b])
|
||||
|
||||
# Clean up a stale pair left by this scenario. The token makes the
|
||||
# name unique to this run, so a pair orphaned by an earlier run is
|
||||
# no longer reclaimed here — `ci-cleanup.sh` reaps those instead.
|
||||
_run_host(["ip", "link", "delete", host_a], image, check=False)
|
||||
|
||||
# Create veth pair on host
|
||||
ok = _run_host([
|
||||
"ip", "link", "add", host_a, "type", "veth", "peer", "name", host_b,
|
||||
], image)
|
||||
if not ok:
|
||||
log.warning("Failed to create veth pair %s/%s", host_a, host_b)
|
||||
return
|
||||
|
||||
# Move into container namespaces
|
||||
_run_host(["ip", "link", "set", host_a, "netns", str(pid_a)], image)
|
||||
_run_host(["ip", "link", "set", host_b, "netns", str(pid_b)], image)
|
||||
|
||||
# Rename and bring up inside containers
|
||||
docker_exec_quiet(
|
||||
container_a,
|
||||
f"ip link set {host_a} name {final_a} && ip link set {final_a} up",
|
||||
timeout=10,
|
||||
)
|
||||
docker_exec_quiet(
|
||||
container_b,
|
||||
f"ip link set {host_b} name {final_b} && ip link set {final_b} up",
|
||||
timeout=10,
|
||||
_require_host(
|
||||
["ip", "link", "add", host_a, "type", "veth", "peer", "name", host_b],
|
||||
image,
|
||||
)
|
||||
_require_host(["ip", "link", "set", host_a, "netns", str(pid_a)], image)
|
||||
_require_host(["ip", "link", "set", host_b, "netns", str(pid_b)], image)
|
||||
|
||||
# Query MAC addresses
|
||||
_raise_link(container_a, host_a, final_a)
|
||||
_raise_link(container_b, host_b, final_b)
|
||||
_await_up(container_a, final_a)
|
||||
_await_up(container_b, final_b)
|
||||
|
||||
# Read only after both ends are proven renamed and up. Before, a
|
||||
# failed rename left the old interface under the final name, and this
|
||||
# read its MAC and reported the restore as a success.
|
||||
mac_a = _get_mac_in_container(container_a, final_a)
|
||||
mac_b = _get_mac_in_container(container_b, final_b)
|
||||
|
||||
if mac_a:
|
||||
self.topology.nodes[node_a].ethernet_macs[node_b] = mac_a
|
||||
if mac_b:
|
||||
self.topology.nodes[node_b].ethernet_macs[node_a] = mac_b
|
||||
if not mac_a or not mac_b:
|
||||
raise VethSetupError(
|
||||
f"veth {node_a}--{node_b} is up but its MAC could not be read "
|
||||
f"({final_a}: {mac_a or '?'}, {final_b}: {mac_b or '?'})"
|
||||
)
|
||||
self.topology.nodes[node_a].ethernet_macs[node_b] = mac_a
|
||||
self.topology.nodes[node_b].ethernet_macs[node_a] = mac_b
|
||||
|
||||
self._host_pairs.append((node_a, node_b, host_a, host_b))
|
||||
|
||||
log.info(
|
||||
"Veth %s(%s) -- %s(%s) MAC: %s / %s",
|
||||
node_a, final_a, node_b, final_b,
|
||||
mac_a or "?", mac_b or "?",
|
||||
node_a, final_a, node_b, final_b, mac_a, mac_b,
|
||||
)
|
||||
|
||||
|
||||
def _get_container_pid(container: str) -> int | None:
|
||||
"""Get the PID of a running Docker container."""
|
||||
def _in_container(container: str, cmd: str, what: str, timeout: int = EXEC_TIMEOUT_SECS) -> str:
|
||||
"""Run a command inside a container, raising `VethSetupError` on failure."""
|
||||
try:
|
||||
return docker_exec(container, cmd, timeout=timeout)
|
||||
except (DockerExecError, subprocess.TimeoutExpired) as e:
|
||||
raise VethSetupError(f"{what} in {container} failed: {e}") from e
|
||||
|
||||
|
||||
def _purge_links(container: str, names: list[str]):
|
||||
"""Delete each named interface in a container if it exists, in one exec.
|
||||
|
||||
An absent name is the ordinary case and is skipped. A delete that fails
|
||||
raises, since the rename that follows would then fail on the name too,
|
||||
unless the name is gone by then: a lingering namespace can be reaped
|
||||
between the check and the delete, taking the survivor's end with it.
|
||||
"""
|
||||
script = "; ".join(
|
||||
f"if ip link show {name} >/dev/null 2>&1; then "
|
||||
f"ip link delete {name} || ! ip link show {name} >/dev/null 2>&1 || exit 1; fi"
|
||||
for name in names
|
||||
)
|
||||
_in_container(container, script, f"deleting stale {', '.join(names)}")
|
||||
|
||||
|
||||
def _require_host(cmd: list[str], image: str):
|
||||
"""Run a host-namespace ``ip`` command, raising `VethSetupError` on failure."""
|
||||
if not _run_host(cmd, image):
|
||||
raise VethSetupError(f"host command failed: {' '.join(cmd)}")
|
||||
|
||||
|
||||
def _raise_link(container: str, temp: str, final: str):
|
||||
"""Rename a moved veth end to its final name and set it up."""
|
||||
_in_container(
|
||||
container,
|
||||
f"ip link set {temp} name {final} && ip link set {final} up",
|
||||
f"renaming {temp} to {final}",
|
||||
)
|
||||
|
||||
|
||||
def _await_up(container: str, iface: str):
|
||||
"""Wait for an interface's operstate to read ``up``, else raise.
|
||||
|
||||
A veth end reports up only once both ends are up, so this proves the
|
||||
pair is joined as well as that this end was raised.
|
||||
"""
|
||||
deadline = time.monotonic() + OPERSTATE_WAIT_SECS
|
||||
state = None
|
||||
reads = 0
|
||||
while True:
|
||||
state = docker_exec_quiet(
|
||||
container, f"cat /sys/class/net/{iface}/operstate", timeout=EXEC_TIMEOUT_SECS
|
||||
)
|
||||
state = state.strip() if state is not None else None
|
||||
reads += 1
|
||||
if state == "up":
|
||||
return
|
||||
if reads >= 2 and time.monotonic() >= deadline:
|
||||
raise VethSetupError(
|
||||
f"{iface} in {container} reads operstate {state or '?'}, "
|
||||
f"not up, {OPERSTATE_WAIT_SECS}s after it was raised"
|
||||
)
|
||||
time.sleep(0.2)
|
||||
|
||||
|
||||
def _container_pid(container: str) -> int | None:
|
||||
"""Return a container's PID, or None when docker reports it not running.
|
||||
|
||||
Raises `VethSetupError` when docker cannot be asked. A timed-out or failed
|
||||
inspect of a running container used to read as "not running", so under
|
||||
host load a live link was skipped and never recreated.
|
||||
"""
|
||||
try:
|
||||
result = subprocess.run(
|
||||
["docker", "inspect", "-f", "{{.State.Pid}}", container],
|
||||
capture_output=True,
|
||||
text=True,
|
||||
timeout=10,
|
||||
timeout=EXEC_TIMEOUT_SECS,
|
||||
)
|
||||
if result.returncode == 0:
|
||||
pid = int(result.stdout.strip())
|
||||
return pid if pid > 0 else None
|
||||
except (subprocess.TimeoutExpired, ValueError):
|
||||
pass
|
||||
return None
|
||||
except subprocess.TimeoutExpired as e:
|
||||
raise VethSetupError(f"docker inspect {container} timed out") from e
|
||||
if result.returncode != 0:
|
||||
raise VethSetupError(
|
||||
f"docker inspect {container} failed: {result.stderr.strip()}"
|
||||
)
|
||||
try:
|
||||
pid = int(result.stdout.strip())
|
||||
except ValueError as e:
|
||||
raise VethSetupError(
|
||||
f"docker inspect {container} returned no PID: {result.stdout.strip()!r}"
|
||||
) from e
|
||||
return pid if pid > 0 else None
|
||||
|
||||
|
||||
def _get_mac_in_container(container: str, iface: str) -> str | None:
|
||||
@@ -215,7 +340,7 @@ def _get_mac_in_container(container: str, iface: str) -> str | None:
|
||||
result = docker_exec_quiet(
|
||||
container,
|
||||
f"cat /sys/class/net/{iface}/address",
|
||||
timeout=5,
|
||||
timeout=EXEC_TIMEOUT_SECS,
|
||||
)
|
||||
if result is not None:
|
||||
return result.strip()
|
||||
|
||||
+13
-7
@@ -204,13 +204,19 @@ VETH_NODE_ID='(0[0-9]|[1-9][0-9]+)'
|
||||
# from the simulation itself rather than repeated here, so widening it cannot
|
||||
# leave this matching the old width. Empty output means "reap nothing".
|
||||
#
|
||||
# Two producers, two shapes. The chaos simulation makes vh{token}{NN}{MM}{a,b};
|
||||
# the NAT lab (nat/scripts/setup-topology.sh) makes vn{a,b}{token}{0,1}, using
|
||||
# the RUN-wide suffix rather than any chaos scenario's. Widening this regex is
|
||||
# Three producers, three shapes. The chaos simulation makes
|
||||
# vh{token}{NN}{MM}{a,b}; the NAT lab (nat/scripts/setup-topology.sh) makes
|
||||
# vn{a,b}{token}{0,1}; the interface-binding suite
|
||||
# (iface-binding/test.sh) makes vhifb{token}{a,b,c,d}. The last two use the
|
||||
# RUN-wide suffix rather than any chaos scenario's. Widening this regex is
|
||||
# only half the fix: the token set is derived separately below, so a suffix
|
||||
# list carrying no NAT suffix leaves the NAT half matching nothing while
|
||||
# list carrying no run-wide suffix leaves those halves matching nothing while
|
||||
# looking correct. ci-local.sh's ci_teardown therefore appends the run-wide
|
||||
# suffix to --veth-suffixes.
|
||||
#
|
||||
# vhifb is listed before the bare vh shape in each alternation only for
|
||||
# readability; the shapes are anchored and cannot overlap, because a node id
|
||||
# is digits and `ifb` is not.
|
||||
veth_pattern() {
|
||||
# vh{token}{NN}{MM}{a,b}: the token is 4 hex or wholly absent — never a
|
||||
# part of one — and the two node ids follow. Anchored and shaped this
|
||||
@@ -219,7 +225,7 @@ veth_pattern() {
|
||||
# single run, no missing or empty suffix list may widen this back out to
|
||||
# every run.
|
||||
if [[ -z "$RUN_ID" && -z "$VETH_SUFFIXES" ]]; then
|
||||
printf '^vh([0-9a-f]{4})?%s%s[ab]$|^vn[ab]([0-9a-f]{4})?[01]$' \
|
||||
printf '^vhifb([0-9a-f]{4})?[a-d]$|^vh([0-9a-f]{4})?%s%s[ab]$|^vn[ab]([0-9a-f]{4})?[01]$' \
|
||||
"$VETH_NODE_ID" "$VETH_NODE_ID"
|
||||
return 0
|
||||
fi
|
||||
@@ -253,8 +259,8 @@ veth_pattern() {
|
||||
veth_warn "no interface tokens derived from: ${sfx[*]}"
|
||||
return 0
|
||||
fi
|
||||
printf '^vh(%s)%s%s[ab]$|^vn[ab](%s)[01]$' \
|
||||
"$alt" "$VETH_NODE_ID" "$VETH_NODE_ID" "$alt"
|
||||
printf '^vhifb(%s)[a-d]$|^vh(%s)%s%s[ab]$|^vn[ab](%s)[01]$' \
|
||||
"$alt" "$alt" "$VETH_NODE_ID" "$VETH_NODE_ID" "$alt"
|
||||
}
|
||||
|
||||
# ip(8) in a privileged --net=host container, matching how the simulation
|
||||
|
||||
+81
-4
@@ -30,7 +30,8 @@
|
||||
# firewall, nat-cone, nat-symmetric,
|
||||
# nat-lan, nostr-publish-consume, stun-faults,
|
||||
# chaos-churn-mixed-10, chaos-ethernet-mesh,
|
||||
# chaos-ethernet-only, chaos-tcp-mesh, chaos-congestion-stress,
|
||||
# chaos-ethernet-only, chaos-ethernet-churn, chaos-tcp-mesh,
|
||||
# chaos-congestion-stress,
|
||||
# sidecar, dns-resolver, deb-install, medium-change
|
||||
#
|
||||
# Opt-in (require --with-tor; depend on live Tor network):
|
||||
@@ -164,6 +165,7 @@ CHAOS_SUITES=(
|
||||
"churn-mixed-10 churn-mixed --nodes 10 --duration 120"
|
||||
"ethernet-mesh ethernet-mesh"
|
||||
"ethernet-only ethernet-only"
|
||||
"ethernet-churn ethernet-churn"
|
||||
"tcp-mesh tcp-mesh"
|
||||
"congestion-stress congestion-stress"
|
||||
)
|
||||
@@ -208,6 +210,7 @@ CHAOS_SUITES=(
|
||||
GATEWAY_SUITES=(gateway)
|
||||
SIDECAR_SUITES=(sidecar)
|
||||
FIREWALL_SUITES=(firewall)
|
||||
IFACE_BINDING_SUITES=(iface-binding)
|
||||
NAT_SUITES=(cone symmetric lan)
|
||||
NOSTR_RELAY_SUITES=(nostr-publish-consume)
|
||||
STUN_FAULTS_SUITES=(stun-faults)
|
||||
@@ -247,6 +250,9 @@ list_suites() {
|
||||
echo " Firewall baseline:"
|
||||
for s in "${FIREWALL_SUITES[@]}"; do echo " $s"; done
|
||||
echo ""
|
||||
echo " Dynamic interface binding:"
|
||||
for s in "${IFACE_BINDING_SUITES[@]}"; do echo " $s"; done
|
||||
echo ""
|
||||
echo " NAT scenarios:"
|
||||
for s in "${NAT_SUITES[@]}"; do echo " nat-$s"; done
|
||||
echo ""
|
||||
@@ -669,6 +675,47 @@ run_static() {
|
||||
record "static-$topology" $rc
|
||||
}
|
||||
|
||||
# Lines kept from each node's log when a red scenario's results are printed.
|
||||
CHAOS_DUMP_LINES=300
|
||||
|
||||
# Print a red chaos scenario's results into the run log.
|
||||
#
|
||||
# A CI worker that runs each job in a worktree it deletes afterwards, under a
|
||||
# private /tmp, loses the results directory when the run ends, and the log is
|
||||
# the only record that survives. Without this such a red cannot be traced node
|
||||
# by node. The runner's own log is not repeated here: it already reached the run
|
||||
# log as the scenario's output. Each node log is capped, so the total grows with
|
||||
# node count; the largest CI scenario has ten nodes.
|
||||
chaos_dump() {
|
||||
local name="$1" base="$2" dir f
|
||||
dir="$(find "$base" -mindepth 1 -maxdepth 1 -type d 2>/dev/null | sort | tail -n 1)"
|
||||
if [[ -z "$dir" ]]; then
|
||||
echo "[chaos/$name] no results directory under $base"
|
||||
return 0
|
||||
fi
|
||||
echo "===== chaos-$name results from $dir ====="
|
||||
for f in status.txt assertions.txt; do
|
||||
[[ -f "$dir/$f" ]] || continue
|
||||
echo "----- $f -----"
|
||||
cat "$dir/$f"
|
||||
done
|
||||
if [[ -f "$dir/tree-snapshot-final.json" ]]; then
|
||||
echo "----- final tree: node, root, parent -----"
|
||||
python3 -c '
|
||||
import json, sys
|
||||
for node, tree in sorted(json.load(open(sys.argv[1])).items()):
|
||||
print(node, tree.get("root"), tree.get("parent"))
|
||||
' "$dir/tree-snapshot-final.json" || true
|
||||
fi
|
||||
for f in "$dir"/fips-node-*.log; do
|
||||
[[ -f "$f" ]] || continue
|
||||
echo "----- $(basename "$f"), last $CHAOS_DUMP_LINES lines -----"
|
||||
tail -n "$CHAOS_DUMP_LINES" "$f"
|
||||
done
|
||||
echo "===== end of chaos-$name results ====="
|
||||
return 0
|
||||
}
|
||||
|
||||
# Run a chaos scenario
|
||||
run_chaos() {
|
||||
local name="$1"
|
||||
@@ -686,11 +733,17 @@ run_chaos() {
|
||||
suffix="$(ci_chaos_suffix "$name")"
|
||||
local -x FIPS_CI_NAME_SUFFIX="$suffix"
|
||||
|
||||
# The same scoping for the results, so a red can find the directory this
|
||||
# run wrote rather than guess among earlier runs' timestamps.
|
||||
local results="$SCRIPT_DIR/chaos/sim-results/ci$suffix"
|
||||
local -x FIPS_SIM_OUTPUT="$results"
|
||||
|
||||
info "[chaos/$name] Running simulation"
|
||||
if bash testing/chaos/scripts/chaos.sh "$@" 2>&1; then
|
||||
rc=0
|
||||
else
|
||||
rc=1
|
||||
chaos_dump "$name" "$results"
|
||||
fi
|
||||
|
||||
record "chaos-$name" $rc
|
||||
@@ -796,6 +849,22 @@ run_sidecar() {
|
||||
record "sidecar" $rc
|
||||
}
|
||||
|
||||
# Run the dynamic interface binding integration test.
|
||||
#
|
||||
# Creates a veth pair from the host namespace after the daemons are up, so it
|
||||
# needs the same privileged ip(8) helper the chaos harness uses. Scoped by
|
||||
# COMPOSE_PROJECT_NAME like every other suite; the veth names carry the run
|
||||
# suffix themselves (see testing/iface-binding/test.sh).
|
||||
run_iface_binding() {
|
||||
export COMPOSE_PROJECT_NAME="$(ci_project iface_binding)"
|
||||
info "[iface-binding] Running integration test"
|
||||
if bash testing/iface-binding/test.sh --skip-build 2>&1; then
|
||||
record "iface-binding" 0
|
||||
else
|
||||
record "iface-binding" 1
|
||||
fi
|
||||
}
|
||||
|
||||
# Run firewall baseline integration test
|
||||
run_firewall() {
|
||||
export COMPOSE_PROJECT_NAME="$(ci_project firewall)"
|
||||
@@ -1262,6 +1331,9 @@ run_integration() {
|
||||
# Firewall baseline
|
||||
run_firewall
|
||||
|
||||
# Dynamic interface binding
|
||||
run_iface_binding
|
||||
|
||||
# NAT scenarios (sequential — each owns its compose project)
|
||||
for scenario in "${NAT_SUITES[@]}"; do
|
||||
run_nat "$scenario"
|
||||
@@ -1334,9 +1406,12 @@ run_integration() {
|
||||
record "chaos-$scenario" 0
|
||||
else
|
||||
record "chaos-$scenario" 1
|
||||
# Show tail of failure log
|
||||
echo "--- chaos-$scenario output (last 20 lines) ---"
|
||||
tail -20 "$logfile" 2>/dev/null || true
|
||||
# The whole log, not its tail: the child printed the results
|
||||
# directory into it, and this is the only place that reaches
|
||||
# the run log. The sed keeps the last frame of each line the
|
||||
# progress display redraws with carriage returns.
|
||||
echo "--- chaos-$scenario output ---"
|
||||
sed 's/.*\r//' "$logfile" 2>/dev/null || true
|
||||
echo "---"
|
||||
fi
|
||||
rm -f "$logfile"
|
||||
@@ -1375,6 +1450,8 @@ run_suite() {
|
||||
run_gateway ;;
|
||||
firewall)
|
||||
run_firewall ;;
|
||||
iface-binding)
|
||||
run_iface_binding ;;
|
||||
nat-cone|nat-symmetric|nat-lan)
|
||||
run_nat "${suite#nat-}" ;;
|
||||
nostr-publish-consume)
|
||||
|
||||
@@ -0,0 +1,57 @@
|
||||
# Dynamic Interface Binding
|
||||
|
||||
Two FIPS daemons whose **only** transports are bound to network interfaces,
|
||||
exercised against a veth pair the harness creates, downs, deletes and recreates
|
||||
underneath them while they run.
|
||||
|
||||
```
|
||||
node-a node-b
|
||||
lab ve-lab0 required ── veth ── ve-lab0 required
|
||||
dock fips-dock0 optional fips-dock0 optional
|
||||
```
|
||||
|
||||
`ve-lab0` does not exist when the daemons start. `fips-dock0` never exists at
|
||||
all, on any host, ever — it is the negative control for `optional: true`.
|
||||
|
||||
## What it asserts
|
||||
|
||||
| | Behavior |
|
||||
| - | -------- |
|
||||
| (a) | A daemon whose only interface is missing **starts**, reports the transport `absent`, and reports `Degraded` — it does not exit on `NoTransports`, and it does not skip the transport for the life of the process |
|
||||
| (b) | The interface appears; both daemons bind it with no restart, `Degraded` clears, and they discover and peer over it |
|
||||
| (c) | The interface goes down and comes back; presence and health follow it in **both** directions, and the rebind is counted |
|
||||
| (d) | The interface is deleted outright and recreated; both daemons rebind and re-peer — the case the old ENXIO beacon-socket reopen half-covered |
|
||||
| (e) | An `optional` interface that never appears logs at `info` and never moves node health |
|
||||
| | Absence is logged **once on the edge**, not once per retry |
|
||||
|
||||
Health is asserted through `fipsctl show status` (`state`), presence through
|
||||
`fipsctl show transports` (the per-transport `interface` block: `presence`,
|
||||
`policy`, `binds`, `since_secs`).
|
||||
|
||||
## Running
|
||||
|
||||
```sh
|
||||
./test.sh # builds the image first
|
||||
./test.sh --skip-build # reuse an existing image
|
||||
./test.sh --keep-up # leave the containers running for inspection
|
||||
```
|
||||
|
||||
Via the local CI runner:
|
||||
|
||||
```sh
|
||||
./testing/ci-local.sh --only iface-binding
|
||||
```
|
||||
|
||||
## Notes
|
||||
|
||||
The containers run under `FIPS_TEST_MODE=default`, **not** `chaos`. The chaos
|
||||
entrypoint waits up to 30 s for every configured Ethernet interface before
|
||||
starting the daemon — which is exactly the workaround this mechanism retires.
|
||||
The daemon has to do its own waiting here or the suite proves nothing.
|
||||
|
||||
Every `ip link` operation on the host network stack runs inside a short-lived
|
||||
privileged container sharing the host network and PID namespaces, for the
|
||||
reason [chaos/sim/veth.py](../chaos/sim/veth.py) documents: on macOS the
|
||||
containers live in the Docker VM, so ip(8) run on the macOS host could never
|
||||
reach them, while on Linux the shared namespaces make it identical to running
|
||||
ip(8) directly.
|
||||
@@ -0,0 +1,80 @@
|
||||
networks:
|
||||
# Management bridge only. The FIPS transport under test is raw Ethernet on a
|
||||
# veth pair the harness creates *after* the daemons are already running —
|
||||
# that is the whole point of the suite — so no FIPS traffic crosses this
|
||||
# network. No subnet is requested, so two concurrent runs cannot collide on
|
||||
# one address range.
|
||||
#
|
||||
# The compose project name is still fixed, so two runs that do not set
|
||||
# COMPOSE_PROJECT_NAME share a project and the second `up` recreates the
|
||||
# first's containers. The local CI runner scopes it externally
|
||||
# (run_iface_binding in ci-local.sh); a bare hand run does not.
|
||||
ifb-net:
|
||||
driver: bridge
|
||||
labels:
|
||||
- "com.corganlabs.fips-ci=1"
|
||||
|
||||
x-fips-common: &fips-common
|
||||
build:
|
||||
# The harness scopes its build context per run and passes it here; the
|
||||
# shared directory is the hand-run default. Compose resolves a relative
|
||||
# value against THIS file's directory, so the harness must export an
|
||||
# absolute path.
|
||||
context: ${FIPS_BUILD_CONTEXT:-../docker}
|
||||
image: ${FIPS_TEST_IMAGE:-fips-test:latest}
|
||||
entrypoint: ["/usr/local/bin/entrypoint.sh"]
|
||||
cap_add:
|
||||
- NET_ADMIN
|
||||
- NET_RAW
|
||||
restart: "no"
|
||||
environment:
|
||||
# `default`, deliberately — NOT `chaos`. The chaos entrypoint waits up to
|
||||
# 30 s for every configured Ethernet interface to appear before it starts
|
||||
# the daemon, which is precisely the workaround this mechanism retires. The
|
||||
# daemon must do its own waiting here or the suite proves nothing.
|
||||
- FIPS_TEST_MODE=default
|
||||
- RUST_LOG=info,fips::transport::ethernet=debug,fips::node=debug
|
||||
networks:
|
||||
- ifb-net
|
||||
|
||||
services:
|
||||
node-a:
|
||||
<<: *fips-common
|
||||
container_name: fips-ifb-node-a${FIPS_CI_NAME_SUFFIX:-}
|
||||
hostname: host-a
|
||||
volumes:
|
||||
- ../docker/resolv.conf:/etc/resolv.conf:ro
|
||||
- ./generated-configs${FIPS_CI_NAME_SUFFIX:-}/node-a/fips.yaml:/etc/fips/fips.yaml:ro
|
||||
- ./generated-configs${FIPS_CI_NAME_SUFFIX:-}/node-a/fips.key:/etc/fips/fips.key:ro
|
||||
|
||||
node-b:
|
||||
<<: *fips-common
|
||||
container_name: fips-ifb-node-b${FIPS_CI_NAME_SUFFIX:-}
|
||||
hostname: host-b
|
||||
volumes:
|
||||
- ../docker/resolv.conf:/etc/resolv.conf:ro
|
||||
- ./generated-configs${FIPS_CI_NAME_SUFFIX:-}/node-b/fips.yaml:/etc/fips/fips.yaml:ro
|
||||
- ./generated-configs${FIPS_CI_NAME_SUFFIX:-}/node-b/fips.key:/etc/fips/fips.key:ro
|
||||
|
||||
# The clean-start case. Its interface exists before its daemon does, which is
|
||||
# the ordinary state of a booted router and the one ordering node-a and
|
||||
# node-b cannot produce: their interface is created after they are already
|
||||
# running, so they can only ever bind through the binder loop.
|
||||
#
|
||||
# The gate is what buys that ordering. The harness needs a running container
|
||||
# to have a netns to move a veth into, but the daemon must not start until
|
||||
# after the move — so the container comes up, parks on this file, and the
|
||||
# harness releases it once the interface is in place.
|
||||
node-c:
|
||||
<<: *fips-common
|
||||
container_name: fips-ifb-node-c${FIPS_CI_NAME_SUFFIX:-}
|
||||
hostname: host-c
|
||||
entrypoint: ["/bin/sh", "-c"]
|
||||
command:
|
||||
- |
|
||||
while [ ! -e /tmp/fips-go ]; do sleep 0.2; done
|
||||
exec /usr/local/bin/entrypoint.sh
|
||||
volumes:
|
||||
- ../docker/resolv.conf:/etc/resolv.conf:ro
|
||||
- ./generated-configs${FIPS_CI_NAME_SUFFIX:-}/node-c/fips.yaml:/etc/fips/fips.yaml:ro
|
||||
- ./generated-configs${FIPS_CI_NAME_SUFFIX:-}/node-c/fips.key:/etc/fips/fips.key:ro
|
||||
Executable
+136
@@ -0,0 +1,136 @@
|
||||
#!/bin/bash
|
||||
# Generate fixtures for the dynamic interface binding integration test.
|
||||
#
|
||||
# Two FIPS nodes, each with two Ethernet transports and nothing else:
|
||||
#
|
||||
# lab ve-lab0 required — does not exist when the daemon starts; the
|
||||
# harness creates the veth pair afterwards
|
||||
# dock fips-dock0 optional — never exists, on any host, ever
|
||||
#
|
||||
# Plus a third node whose single required interface exists *before* its daemon
|
||||
# starts — the one ordering the other two cannot produce, and the one that
|
||||
# `start_async`'s inline bind takes. See node-c in test.sh case (f).
|
||||
#
|
||||
# There is deliberately no UDP transport. A node whose only transports are
|
||||
# interface-bound is the case that used to be unrecoverable: every transport
|
||||
# skipped at start, nothing retried, and the node up and deaf.
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
|
||||
# Scoped by the per-run suffix because this directory is wiped and rewritten
|
||||
# below: two runs sharing one output directory would delete each other's
|
||||
# fixtures out from under running containers.
|
||||
GENERATED_DIR="$SCRIPT_DIR/generated-configs${FIPS_CI_NAME_SUFFIX:-}"
|
||||
|
||||
# Deterministic test identities (mirrors the firewall/acl-allowlist style).
|
||||
KEY_A="0102030405060708090a0b0c0d0e0f101112131415161718191a1b1c1d1e1f20"
|
||||
KEY_B="b102030405060708090a0b0c0d0e0f101112131415161718191a1b1c1d1e1fb0"
|
||||
KEY_C="c102030405060708090a0b0c0d0e0f101112131415161718191a1b1c1d1e1fc0"
|
||||
|
||||
write_file() {
|
||||
local path="$1"
|
||||
mkdir -p "$(dirname "$path")"
|
||||
cat > "$path"
|
||||
}
|
||||
|
||||
# Peers are found by beacon, not configured: a MAC address that does not exist
|
||||
# until the harness creates the veth pair cannot be written into a config file
|
||||
# ahead of time. Discovery over the late-bound interface is part of what the
|
||||
# suite asserts.
|
||||
write_node_config() {
|
||||
write_file "$GENERATED_DIR/$1/fips.yaml" <<EOF
|
||||
node:
|
||||
identity:
|
||||
persistent: true
|
||||
|
||||
# No TUN and no DNS: this suite is about transport binding, and every extra
|
||||
# child is another way for a failure to be misattributed.
|
||||
tun:
|
||||
enabled: false
|
||||
|
||||
dns:
|
||||
enabled: false
|
||||
|
||||
transports:
|
||||
ethernet:
|
||||
lab:
|
||||
interface: "$LAB_IFACE"
|
||||
listen: true
|
||||
announce: true
|
||||
auto_connect: true
|
||||
accept_connections: true
|
||||
# Fast beacons so peering after a late bind is observed in seconds
|
||||
# rather than in the 30 s production default.
|
||||
beacon_interval_secs: 2
|
||||
dock:
|
||||
interface: "$DOCK_IFACE"
|
||||
# Absence is normal for this one, so it must never move node health.
|
||||
optional: true
|
||||
listen: true
|
||||
announce: true
|
||||
auto_connect: true
|
||||
accept_connections: true
|
||||
beacon_interval_secs: 2
|
||||
|
||||
peers: []
|
||||
EOF
|
||||
}
|
||||
|
||||
# node-c: one required interface, present at daemon start.
|
||||
#
|
||||
# The other two nodes can only ever reach `Present` through the binder loop,
|
||||
# because their interface does not exist until the harness makes it. That left
|
||||
# the inline bind in `start_async` — the ordinary case on a booted router —
|
||||
# with no coverage at all, which is exactly where the churn guard went unseeded
|
||||
# and the first detach stopped reaching node health.
|
||||
write_boot_node_config() {
|
||||
write_file "$GENERATED_DIR/node-c/fips.yaml" <<EOF
|
||||
node:
|
||||
identity:
|
||||
persistent: true
|
||||
|
||||
tun:
|
||||
enabled: false
|
||||
|
||||
dns:
|
||||
enabled: false
|
||||
|
||||
transports:
|
||||
ethernet:
|
||||
boot:
|
||||
interface: "$BOOT_IFACE"
|
||||
listen: true
|
||||
announce: true
|
||||
auto_connect: true
|
||||
accept_connections: true
|
||||
beacon_interval_secs: 2
|
||||
|
||||
peers: []
|
||||
EOF
|
||||
}
|
||||
|
||||
LAB_IFACE="${LAB_IFACE:-ve-lab0}"
|
||||
DOCK_IFACE="${DOCK_IFACE:-fips-dock0}"
|
||||
BOOT_IFACE="${BOOT_IFACE:-ve-boot0}"
|
||||
|
||||
echo "Generating interface-binding fixtures..."
|
||||
rm -rf "$GENERATED_DIR"
|
||||
|
||||
write_node_config node-a
|
||||
write_file "$GENERATED_DIR/node-a/fips.key" <<EOF
|
||||
$KEY_A
|
||||
EOF
|
||||
|
||||
write_node_config node-b
|
||||
write_file "$GENERATED_DIR/node-b/fips.key" <<EOF
|
||||
$KEY_B
|
||||
EOF
|
||||
|
||||
write_boot_node_config
|
||||
write_file "$GENERATED_DIR/node-c/fips.key" <<EOF
|
||||
$KEY_C
|
||||
EOF
|
||||
|
||||
echo "Interface-binding fixtures written to $GENERATED_DIR"
|
||||
Executable
+632
@@ -0,0 +1,632 @@
|
||||
#!/bin/bash
|
||||
# Integration test for dynamic interface binding.
|
||||
#
|
||||
# Asserts the five behaviors the presence machine exists to provide, against
|
||||
# real daemons and a real veth pair:
|
||||
#
|
||||
# (a) boot race — a daemon whose interface does not exist yet starts,
|
||||
# reports the transport ABSENT, and reports Degraded
|
||||
# (b) late attach — the interface appears; the daemon binds it with no
|
||||
# restart, clears Degraded, and peers over it
|
||||
# (c) flap — the interface goes down and comes back; presence and
|
||||
# health follow it in BOTH directions
|
||||
# (d) destroy/recreate — the interface is deleted outright and recreated; the
|
||||
# daemon rebinds (the case the old ENXIO beacon hack
|
||||
# half-covered)
|
||||
# (e) optional — an interface that never appears logs at info and
|
||||
# never moves node health
|
||||
#
|
||||
# plus the log-hygiene property the design is explicit about: absence is logged
|
||||
# once on the edge, never once per retry.
|
||||
#
|
||||
# The daemons run under FIPS_TEST_MODE=default, NOT chaos: the chaos entrypoint
|
||||
# waits for Ethernet interfaces before starting the daemon, which is exactly
|
||||
# the workaround being retired. The daemon must wait for itself here.
|
||||
#
|
||||
# Usage: ./test.sh [--skip-build] [--keep-up]
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
TESTING_DIR="$(cd "$SCRIPT_DIR/.." && pwd)"
|
||||
COMPOSE_FILE="$SCRIPT_DIR/docker-compose.yml"
|
||||
|
||||
NODE_A="fips-ifb-node-a${FIPS_CI_NAME_SUFFIX:-}"
|
||||
NODE_B="fips-ifb-node-b${FIPS_CI_NAME_SUFFIX:-}"
|
||||
NODE_C="fips-ifb-node-c${FIPS_CI_NAME_SUFFIX:-}"
|
||||
|
||||
# The interface each node binds. Same name on both sides: they are in separate
|
||||
# network namespaces, and using one name keeps the fixtures identical.
|
||||
LAB_IFACE="ve-lab0"
|
||||
# The interface that never exists. `optional: true` in both configs.
|
||||
DOCK_IFACE="fips-dock0"
|
||||
# node-c's interface, present before its daemon starts. See case (f).
|
||||
BOOT_IFACE="ve-boot0"
|
||||
|
||||
# Host-side veth names, scoped per run: these live in the host (or Docker VM)
|
||||
# namespace for the moment between creation and the move into the containers,
|
||||
# where two concurrent runs would otherwise collide on one name.
|
||||
#
|
||||
# The scope comes from a four-hex hash of the run suffix, not from the suffix
|
||||
# itself. An interface name gets 15 characters and the suffix alone can spend
|
||||
# 24 of them (`-20260910t025205-2802512`), so interpolating it produced
|
||||
# `vhifb-20260910t025205-2802512b` and ip(8) refused the name outright. The
|
||||
# chaos simulation had already met this and answers it in `sim.naming`, which
|
||||
# the NAT topology script also calls; this uses the same token so the reaper
|
||||
# in `ci-cleanup.sh` can match these interfaces by the same rule.
|
||||
#
|
||||
# Empty for an empty suffix, so a bare run keeps short unscoped names.
|
||||
veth_token() {
|
||||
local suffix="${FIPS_CI_NAME_SUFFIX:-}"
|
||||
if [ -z "$suffix" ]; then
|
||||
echo ""
|
||||
return 0
|
||||
fi
|
||||
PYTHONPATH="$TESTING_DIR/chaos" python3 -m sim.naming "$suffix"
|
||||
}
|
||||
|
||||
VETH_TOKEN="$(veth_token)"
|
||||
HOST_VETH_A="vhifb${VETH_TOKEN}a"
|
||||
HOST_VETH_B="vhifb${VETH_TOKEN}b"
|
||||
HOST_VETH_C="vhifb${VETH_TOKEN}c"
|
||||
HOST_VETH_D="vhifb${VETH_TOKEN}d"
|
||||
|
||||
SKIP_BUILD=false
|
||||
KEEP_UP=false
|
||||
|
||||
while [ $# -gt 0 ]; do
|
||||
case "$1" in
|
||||
--skip-build) SKIP_BUILD=true; shift ;;
|
||||
--keep-up) KEEP_UP=true; shift ;;
|
||||
*) echo "Unknown option: $1" >&2; exit 1 ;;
|
||||
esac
|
||||
done
|
||||
|
||||
log() { echo "=== $*"; }
|
||||
pass() { echo "PASS: $*"; }
|
||||
fail() { echo "FAIL: $*" >&2; dump_diagnostics; exit 1; }
|
||||
|
||||
dump_diagnostics() {
|
||||
echo "--- node-a transports ---" >&2
|
||||
docker exec "$NODE_A" fipsctl show transports >&2 2>&1 || true
|
||||
echo "--- node-a status ---" >&2
|
||||
docker exec "$NODE_A" fipsctl show status >&2 2>&1 || true
|
||||
echo "--- node-a log (tail) ---" >&2
|
||||
docker logs --tail 80 "$NODE_A" >&2 2>&1 || true
|
||||
echo "--- node-b log (tail) ---" >&2
|
||||
docker logs --tail 80 "$NODE_B" >&2 2>&1 || true
|
||||
return 0
|
||||
}
|
||||
|
||||
cleanup() {
|
||||
# Remove the veth pair wherever it survived: inside a container if the move
|
||||
# succeeded, on the host if the run died between creation and the move.
|
||||
docker exec "$NODE_A" ip link del "$LAB_IFACE" >/dev/null 2>&1 || true
|
||||
ip_host "ip link del $HOST_VETH_A" >/dev/null 2>&1 || true
|
||||
docker exec "$NODE_C" ip link del "$BOOT_IFACE" >/dev/null 2>&1 || true
|
||||
ip_host "ip link del $HOST_VETH_C" >/dev/null 2>&1 || true
|
||||
if [ "$KEEP_UP" = false ]; then
|
||||
docker compose -f "$COMPOSE_FILE" down --volumes --remove-orphans >/dev/null 2>&1 || true
|
||||
fi
|
||||
return 0
|
||||
}
|
||||
|
||||
# ── Host-namespace ip(8) ─────────────────────────────────────────────────
|
||||
#
|
||||
# Every `ip link` operation that touches the host network stack runs inside a
|
||||
# short-lived privileged container sharing the host network and PID namespaces,
|
||||
# for the reason testing/chaos/sim/veth.py documents at length: on macOS the
|
||||
# containers live in the Docker VM, so running ip(8) on the macOS host could
|
||||
# never reach them, while on Linux the shared namespaces make it identical to
|
||||
# running ip(8) directly.
|
||||
ip_host() {
|
||||
docker run --rm --privileged --network host --pid host \
|
||||
--entrypoint /bin/sh "$IMAGE" -c "$1"
|
||||
}
|
||||
|
||||
# Docker's own view of a container, so the gate case can wait for a netns
|
||||
# without implying the daemon inside it has started.
|
||||
container_state() {
|
||||
docker inspect -f '{{.State.Status}}' "$1" 2>/dev/null || true
|
||||
}
|
||||
|
||||
container_pid() {
|
||||
docker inspect -f '{{.State.Pid}}' "$1"
|
||||
}
|
||||
|
||||
# Create the veth pair and move one end into each container.
|
||||
create_veth() {
|
||||
local pid_a pid_b
|
||||
pid_a="$(container_pid "$NODE_A")"
|
||||
pid_b="$(container_pid "$NODE_B")"
|
||||
|
||||
# One invocation, not three. The host-side names exist only between the
|
||||
# `add` and the two `netns` moves, and ci-cleanup.sh's host-veth sweep is
|
||||
# deliberately shaped to the chaos simulation's names and does not cover
|
||||
# these — so the window in which a hard kill could strand them is kept to
|
||||
# a single command, with the EXIT trap covering the rest.
|
||||
ip_host "set -e
|
||||
ip link add $HOST_VETH_A type veth peer name $HOST_VETH_B
|
||||
ip link set $HOST_VETH_A netns $pid_a name $LAB_IFACE
|
||||
ip link set $HOST_VETH_B netns $pid_b name $LAB_IFACE" >/dev/null
|
||||
|
||||
# A moved link arrives down. Presence is IFF_UP, so the daemon correctly
|
||||
# does not bind until this runs — which is also why (c) can flap it with
|
||||
# nothing but `ip link set down`.
|
||||
docker exec "$NODE_A" ip link set "$LAB_IFACE" up
|
||||
docker exec "$NODE_B" ip link set "$LAB_IFACE" up
|
||||
return 0
|
||||
}
|
||||
|
||||
# ── Daemon introspection ─────────────────────────────────────────────────
|
||||
|
||||
# Field of a named transport's `interface` block, or "" if the transport, the
|
||||
# block, or the daemon is not there.
|
||||
iface_field() {
|
||||
docker exec "$1" fipsctl show transports 2>/dev/null | python3 -c '
|
||||
import json, sys
|
||||
try:
|
||||
data = json.load(sys.stdin)
|
||||
except Exception:
|
||||
print(""); raise SystemExit
|
||||
for t in data.get("transports", []):
|
||||
if t.get("name") == sys.argv[1]:
|
||||
print(t.get("interface", {}).get(sys.argv[2], ""))
|
||||
break
|
||||
else:
|
||||
print("")
|
||||
' "$2" "$3"
|
||||
}
|
||||
|
||||
node_state() {
|
||||
docker exec "$1" fipsctl show status 2>/dev/null | python3 -c '
|
||||
import json, sys
|
||||
try:
|
||||
print(json.load(sys.stdin).get("state", ""))
|
||||
except Exception:
|
||||
print("")
|
||||
'
|
||||
}
|
||||
|
||||
peer_count() {
|
||||
docker exec "$1" fipsctl show peers 2>/dev/null | python3 -c '
|
||||
import json, sys
|
||||
try:
|
||||
data = json.load(sys.stdin)
|
||||
except Exception:
|
||||
print(0); raise SystemExit
|
||||
print(len(data.get("peers", [])))
|
||||
'
|
||||
}
|
||||
|
||||
# Count of a literal in a container log. Used for the log-hygiene assertion.
|
||||
log_count() {
|
||||
docker logs "$1" 2>&1 | grep -c -- "$2" || true
|
||||
}
|
||||
|
||||
# Poll `expr` until it prints `want`, up to `timeout` seconds.
|
||||
# Usage: wait_for <timeout> <want> <command...>
|
||||
wait_for() {
|
||||
local timeout="$1" want="$2"; shift 2
|
||||
local i got
|
||||
for i in $(seq 1 "$timeout"); do
|
||||
got="$("$@" || true)"
|
||||
if [ "$got" = "$want" ]; then
|
||||
return 0
|
||||
fi
|
||||
sleep 1
|
||||
done
|
||||
echo " (last value: '${got:-}', wanted '$want')" >&2
|
||||
return 1
|
||||
}
|
||||
|
||||
# Poll until the command prints a value no greater than `want`.
|
||||
wait_for_at_most() {
|
||||
local timeout="$1" want="$2"; shift 2
|
||||
local i got
|
||||
for i in $(seq 1 "$timeout"); do
|
||||
got="$("$@" || true)"
|
||||
if [ -n "$got" ] && [ "$got" -le "$want" ] 2>/dev/null; then
|
||||
return 0
|
||||
fi
|
||||
sleep 1
|
||||
done
|
||||
echo " (last value: '${got:-}', wanted <= '$want')" >&2
|
||||
return 1
|
||||
}
|
||||
|
||||
# Poll until the command prints a value that is at least `want`.
|
||||
wait_for_at_least() {
|
||||
local timeout="$1" want="$2"; shift 2
|
||||
local i got
|
||||
for i in $(seq 1 "$timeout"); do
|
||||
got="$("$@" || true)"
|
||||
if [ -n "$got" ] && [ "$got" -ge "$want" ] 2>/dev/null; then
|
||||
return 0
|
||||
fi
|
||||
sleep 1
|
||||
done
|
||||
echo " (last value: '${got:-}', wanted >= '$want')" >&2
|
||||
return 1
|
||||
}
|
||||
|
||||
# ── Run ──────────────────────────────────────────────────────────────────
|
||||
|
||||
IMAGE="${FIPS_TEST_IMAGE:-fips-test:latest}"
|
||||
|
||||
trap cleanup EXIT
|
||||
|
||||
if [ "$SKIP_BUILD" = false ]; then
|
||||
log "Building test image"
|
||||
bash "$TESTING_DIR/scripts/build.sh"
|
||||
fi
|
||||
|
||||
log "Generating fixtures"
|
||||
LAB_IFACE="$LAB_IFACE" DOCK_IFACE="$DOCK_IFACE" BOOT_IFACE="$BOOT_IFACE" \
|
||||
bash "$SCRIPT_DIR/generate-configs.sh"
|
||||
|
||||
log "Starting nodes with $LAB_IFACE absent"
|
||||
docker compose -f "$COMPOSE_FILE" up -d
|
||||
|
||||
# The daemon must reach a serving state without the interface. Waiting on the
|
||||
# control socket answering at all is the first half of assertion (a): a daemon
|
||||
# that exited on `NoTransports` never answers.
|
||||
if ! wait_for 40 "degraded" node_state "$NODE_A"; then
|
||||
fail "(a) node-a did not come up Degraded with $LAB_IFACE absent"
|
||||
fi
|
||||
pass "(a) daemon started and serves with its only interface absent"
|
||||
|
||||
# ── (a) presence and policy are visible to an operator ───────────────────
|
||||
|
||||
[ "$(iface_field "$NODE_A" lab presence)" = "absent" ] \
|
||||
|| fail "(a) lab transport is not reported ABSENT"
|
||||
[ "$(iface_field "$NODE_A" lab policy)" = "required" ] \
|
||||
|| fail "(a) lab transport is not reported required"
|
||||
[ "$(iface_field "$NODE_A" lab name)" = "$LAB_IFACE" ] \
|
||||
|| fail "(a) lab transport does not name $LAB_IFACE"
|
||||
[ "$(iface_field "$NODE_A" dock presence)" = "absent" ] \
|
||||
|| fail "(e) dock transport is not reported ABSENT"
|
||||
[ "$(iface_field "$NODE_A" dock policy)" = "optional" ] \
|
||||
|| fail "(e) dock transport is not reported optional"
|
||||
pass "(a) presence and policy are visible in show_transports"
|
||||
|
||||
# ── log hygiene, measured across the absence ─────────────────────────────
|
||||
#
|
||||
# The required interface's absence must be logged once, on the edge — not once
|
||||
# per retry. Measured over an interval long enough for many retries (the binder
|
||||
# polls every second). The 12 s also carries the next assertion past the 10 s
|
||||
# bring-up window, so both the edge rule and the deadline are covered by one
|
||||
# wait.
|
||||
|
||||
# Matched on the edge line specifically. The deadline error below is a
|
||||
# different line by design, and counting both here would read the deadline
|
||||
# firing during the sleep as a repeated edge.
|
||||
EDGE_LINE="Ethernet interface absent; waiting"
|
||||
|
||||
absent_before="$(log_count "$NODE_A" "$EDGE_LINE")"
|
||||
sleep 12
|
||||
absent_after="$(log_count "$NODE_A" "$EDGE_LINE")"
|
||||
if [ "$absent_after" -ne "$absent_before" ]; then
|
||||
fail "absence was logged $((absent_after - absent_before)) more times over 12s of retries; \
|
||||
the edge must be logged once, not per attempt"
|
||||
fi
|
||||
# Two edges total: one for the required lab interface, one for the optional
|
||||
# dock interface. More would mean the edge is not an edge.
|
||||
[ "$absent_before" -eq 2 ] \
|
||||
|| fail "expected exactly 2 absence edges at start, saw $absent_before"
|
||||
pass "absence is logged once on the edge, not per retry"
|
||||
|
||||
# ── a sustained absence errors exactly once ──────────────────────────────
|
||||
#
|
||||
# The edge is not an error: an interface missing for a moment at boot and bound
|
||||
# a moment later is the ordinary case this mechanism exists to absorb. Past the
|
||||
# 10 s bring-up window it is no longer a race, and a *required* interface says
|
||||
# so — once. The 12 s above put us on the far side of that window.
|
||||
#
|
||||
# Exactly one line, from lab. The optional dock interface has been absent just
|
||||
# as long and must be silent, which is what `optional` means; two here would
|
||||
# mean the policy is not being consulted.
|
||||
|
||||
startup_errors="$(log_count "$NODE_A" " ERROR ")"
|
||||
if [ "$startup_errors" -ne 1 ]; then
|
||||
docker logs "$NODE_A" 2>&1 | grep -- " ERROR " >&2 || true
|
||||
fail "expected exactly 1 ERROR line for the required interface past the \
|
||||
bring-up window, saw $startup_errors (an optional interface must contribute none)"
|
||||
fi
|
||||
[ "$(log_count "$NODE_A" "still missing past the bring-up window")" -eq 1 ] \
|
||||
|| fail "the ERROR line is not the sustained-absence report"
|
||||
pass "a required interface absent past the window errors, an optional one does not"
|
||||
|
||||
# ...and does not keep saying it. Duration is state, published as
|
||||
# interface.since_secs and as Degraded; re-announcing it on a timer is what
|
||||
# the old 1 m / 10 m / 1 h ladder did.
|
||||
sleep 12
|
||||
[ "$(log_count "$NODE_A" " ERROR ")" -eq 1 ] \
|
||||
|| fail "the sustained-absence error repeated; it must be said once per episode"
|
||||
pass "the sustained-absence error is said once, not on a schedule"
|
||||
|
||||
# ── (b) late attach ──────────────────────────────────────────────────────
|
||||
|
||||
log "Creating the veth pair"
|
||||
create_veth
|
||||
|
||||
if ! wait_for 30 "present" iface_field "$NODE_A" lab presence; then
|
||||
fail "(b) node-a did not bind $LAB_IFACE after it appeared"
|
||||
fi
|
||||
if ! wait_for 30 "present" iface_field "$NODE_B" lab presence; then
|
||||
fail "(b) node-b did not bind $LAB_IFACE after it appeared"
|
||||
fi
|
||||
pass "(b) both daemons bound the interface with no restart"
|
||||
|
||||
# Health must clear. This is `Degraded` behaving as a level rather than a
|
||||
# latch — the property that made the old monotonic failed-set wrong.
|
||||
if ! wait_for 20 "running" node_state "$NODE_A"; then
|
||||
fail "(b) node-a stayed Degraded after its interface returned"
|
||||
fi
|
||||
pass "(b) Degraded cleared when the interface came back"
|
||||
|
||||
# The optional interface is still absent and must still not matter.
|
||||
[ "$(iface_field "$NODE_A" dock presence)" = "absent" ] \
|
||||
|| fail "(e) dock unexpectedly bound"
|
||||
pass "(e) an absent optional interface does not degrade the node"
|
||||
|
||||
# Discovery and peering over the late-bound interface: the point of binding at
|
||||
# all. Without this the suite would prove the daemon can open a socket, not
|
||||
# that traffic flows over it.
|
||||
if ! wait_for_at_least 45 1 peer_count "$NODE_A"; then
|
||||
fail "(b) node-a found no peer over the late-bound interface"
|
||||
fi
|
||||
if ! wait_for_at_least 45 1 peer_count "$NODE_B"; then
|
||||
fail "(b) node-b found no peer over the late-bound interface"
|
||||
fi
|
||||
pass "(b) nodes discovered and peered over the late-bound interface"
|
||||
|
||||
# ── (c) flap ─────────────────────────────────────────────────────────────
|
||||
|
||||
errors_before_detach="$(log_count "$NODE_A" " ERROR ")"
|
||||
|
||||
log "Taking $LAB_IFACE down on node-a"
|
||||
docker exec "$NODE_A" ip link set "$LAB_IFACE" down
|
||||
|
||||
detach_start=$SECONDS
|
||||
if ! wait_for 20 "absent" iface_field "$NODE_A" lab presence; then
|
||||
fail "(c) node-a did not notice the interface going down"
|
||||
fi
|
||||
detach_elapsed=$(( SECONDS - detach_start ))
|
||||
if ! wait_for 20 "degraded" node_state "$NODE_A"; then
|
||||
fail "(c) node-a did not report Degraded while its interface was down"
|
||||
fi
|
||||
pass "(c) a link going down is observed as absence and degrades health"
|
||||
|
||||
# A detach is reported, at warn. The edge is not an error — a link coming and
|
||||
# going is the weather in a mesh daemon, and the error is the 10 s deadline's
|
||||
# to give, not the edge's.
|
||||
if ! wait_for_at_least 10 1 log_count "$NODE_A" "Ethernet interface detached"; then
|
||||
fail "(c) a runtime detach was not reported"
|
||||
fi
|
||||
|
||||
# Only meaningful while we are still inside the bring-up window. Detection is
|
||||
# sub-second over netlink, so this is the ordinary path; if the runner was slow
|
||||
# enough that the deadline could have fired, the check has nothing to say and
|
||||
# says so rather than failing on the harness's own latency.
|
||||
if [ "$detach_elapsed" -lt 8 ]; then
|
||||
detach_errors="$(log_count "$NODE_A" " ERROR ")"
|
||||
if [ "$detach_errors" -ne "$errors_before_detach" ]; then
|
||||
docker logs "$NODE_A" 2>&1 | grep -- " ERROR " >&2 || true
|
||||
fail "(c) the detach edge logged an ERROR after ${detach_elapsed}s; the \
|
||||
edge is a warn and only outlasting the window earns an error"
|
||||
fi
|
||||
pass "(c) a detach is reported without crying error"
|
||||
else
|
||||
echo " (skipped the edge-not-an-error check: detach took ${detach_elapsed}s,"
|
||||
echo " which is inside the deadline's reach)"
|
||||
fi
|
||||
|
||||
# The peers that interface carried must go with it, and go *now*.
|
||||
#
|
||||
# `link_dead_timeout_secs` is at its 30 s default here, so a withdrawal inside
|
||||
# 15 s can only have come from the detach edge and not from the liveness
|
||||
# reaper. That gap is the whole point: until the edge drove the teardown, this
|
||||
# node kept the peer, kept selecting routes through it, and kept advertising
|
||||
# reachability it no longer had — dropping transit traffic in silence for the
|
||||
# whole timeout, with alternative paths sitting unused.
|
||||
if ! wait_for_at_most 15 0 peer_count "$NODE_A"; then
|
||||
fail "(c) node-a kept a peer that was only reachable over the downed \
|
||||
interface; the detach edge did not withdraw it"
|
||||
fi
|
||||
pass "(c) the peers the interface carried were withdrawn on the detach edge"
|
||||
|
||||
log "Bringing $LAB_IFACE back up on node-a"
|
||||
docker exec "$NODE_A" ip link set "$LAB_IFACE" up
|
||||
|
||||
if ! wait_for 30 "present" iface_field "$NODE_A" lab presence; then
|
||||
fail "(c) node-a did not rebind after the interface came back"
|
||||
fi
|
||||
if ! wait_for 20 "running" node_state "$NODE_A"; then
|
||||
fail "(c) node-a stayed Degraded after the interface came back"
|
||||
fi
|
||||
if ! wait_for_at_least 45 2 iface_field "$NODE_A" lab binds; then
|
||||
fail "(c) the rebind was not counted"
|
||||
fi
|
||||
pass "(c) the interface flapped and the daemon followed it both ways"
|
||||
|
||||
# And the withdrawal is not a one-way door: the peer comes back over the
|
||||
# rebound interface on its own, by beacon, with no operator action.
|
||||
if ! wait_for_at_least 60 1 peer_count "$NODE_A"; then
|
||||
fail "(c) node-a did not re-peer after its interface came back"
|
||||
fi
|
||||
pass "(c) peering re-established over the rebound interface"
|
||||
|
||||
# ── (d) destroy and recreate ─────────────────────────────────────────────
|
||||
#
|
||||
# Deleting the netdev outright is the case the old ENXIO beacon-socket reopen
|
||||
# half-covered: the veth is gone, the socket underneath is stale, and the name
|
||||
# comes back a moment later. One presence machine now covers it.
|
||||
|
||||
log "Deleting and recreating the veth pair"
|
||||
docker exec "$NODE_A" ip link del "$LAB_IFACE"
|
||||
|
||||
if ! wait_for 20 "absent" iface_field "$NODE_A" lab presence; then
|
||||
fail "(d) node-a did not notice the interface being deleted"
|
||||
fi
|
||||
if ! wait_for 20 "absent" iface_field "$NODE_B" lab presence; then
|
||||
fail "(d) node-b did not notice its end of the pair disappearing"
|
||||
fi
|
||||
|
||||
create_veth
|
||||
|
||||
if ! wait_for 30 "present" iface_field "$NODE_A" lab presence; then
|
||||
fail "(d) node-a did not rebind the recreated interface"
|
||||
fi
|
||||
if ! wait_for 30 "present" iface_field "$NODE_B" lab presence; then
|
||||
fail "(d) node-b did not rebind the recreated interface"
|
||||
fi
|
||||
if ! wait_for 20 "running" node_state "$NODE_A"; then
|
||||
fail "(d) node-a stayed Degraded after the interface was recreated"
|
||||
fi
|
||||
pass "(d) a destroyed and recreated interface is rebound"
|
||||
|
||||
# Peering must re-establish over the new hardware. A recreated veth has a new
|
||||
# MAC, so this also exercises the "same name, different hardware" path that
|
||||
# drops cached neighbors instead of resuming onto them.
|
||||
if ! wait_for_at_least 60 1 peer_count "$NODE_A"; then
|
||||
fail "(d) node-a did not re-peer after the interface was recreated"
|
||||
fi
|
||||
pass "(d) peering re-established over the recreated interface"
|
||||
|
||||
# ── (f) an interface present before the daemon starts ────────────────────
|
||||
#
|
||||
# Everything above binds through the binder loop, because the interface does
|
||||
# not exist until the harness makes it. The ordinary case on a booted router is
|
||||
# the opposite one: the interface is already there and `start_async` binds it
|
||||
# inline, before the loop is running.
|
||||
#
|
||||
# That path published its presence edge outside the churn guard, so the guard
|
||||
# believed it had announced nothing and the *first* detach asked for no
|
||||
# retraction. The node kept reporting Running with its only required interface
|
||||
# gone, and stayed that way until a second detach happened to repair the guard.
|
||||
# Nothing in cases (a)-(e) can reach it.
|
||||
log "(f) starting node-c with its interface already present"
|
||||
|
||||
docker compose -f "$COMPOSE_FILE" up -d node-c
|
||||
|
||||
# The container parks on the gate, so this is the netns and not yet the daemon.
|
||||
if ! wait_for 30 "running" container_state "$NODE_C"; then
|
||||
fail "(f) node-c container did not start"
|
||||
fi
|
||||
|
||||
pid_c="$(container_pid "$NODE_C")"
|
||||
ip_host "set -e
|
||||
ip link add $HOST_VETH_C type veth peer name $HOST_VETH_D
|
||||
ip link set $HOST_VETH_C netns $pid_c name $BOOT_IFACE
|
||||
ip link set $HOST_VETH_D netns $pid_c name ${BOOT_IFACE}p" >/dev/null
|
||||
docker exec "$NODE_C" ip link set "$BOOT_IFACE" up
|
||||
docker exec "$NODE_C" ip link set "${BOOT_IFACE}p" up
|
||||
|
||||
# Release the gate. The daemon now starts with the interface already up.
|
||||
docker exec "$NODE_C" touch /tmp/fips-go
|
||||
|
||||
if ! wait_for 40 "running" node_state "$NODE_C"; then
|
||||
fail "(f) node-c did not come up Running with its interface present at start"
|
||||
fi
|
||||
[ "$(iface_field "$NODE_C" boot presence)" = "present" ] \
|
||||
|| fail "(f) node-c did not bind $BOOT_IFACE inline at start"
|
||||
pass "(f) an interface present at start is bound inline and reports Running"
|
||||
|
||||
# The assertion. One detach, on a binding this loop did not create.
|
||||
docker exec "$NODE_C" ip link set "$BOOT_IFACE" down
|
||||
|
||||
if ! wait_for 30 "absent" iface_field "$NODE_C" boot presence; then
|
||||
fail "(f) node-c did not notice $BOOT_IFACE going down"
|
||||
fi
|
||||
if ! wait_for 30 "degraded" node_state "$NODE_C"; then
|
||||
fail "(f) node-c stayed Running after its only required interface went \
|
||||
away — the start-time bind never reached node health"
|
||||
fi
|
||||
pass "(f) the first detach after a clean start degrades the node"
|
||||
|
||||
# And it is still a level, not a latch, on this path too.
|
||||
docker exec "$NODE_C" ip link set "$BOOT_IFACE" up
|
||||
if ! wait_for 30 "running" node_state "$NODE_C"; then
|
||||
fail "(f) node-c stayed Degraded after its interface returned"
|
||||
fi
|
||||
pass "(f) health clears again when the interface returns"
|
||||
|
||||
# ── (g) the link-event fast path is actually the one in use ──────────────
|
||||
#
|
||||
# The whole suite would pass with `open_link_socket()` hardcoded to Err: the
|
||||
# 1 s poll is a complete fallback and covers every `wait_for` window here, so
|
||||
# nothing else asserts that the netlink path exists, let alone that it is what
|
||||
# detected anything. The binder says which backing it got at startup, so ask
|
||||
# it directly rather than inferring from timing that the poll would also
|
||||
# satisfy.
|
||||
# `log_count`, not `grep -q`: under `set -o pipefail` a `grep -q` that exits on
|
||||
# its first match closes the pipe, `docker logs` takes SIGPIPE, and the
|
||||
# pipeline reports failure even though the line was found. `grep -c` reads the
|
||||
# stream to the end.
|
||||
if [ "$(log_count "$NODE_A" "event_driven=true")" -eq 0 ]; then
|
||||
docker logs "$NODE_A" 2>&1 | grep -i "binder started" >&2 || true
|
||||
fail "(g) the binder fell back to polling; the netlink link-event source \
|
||||
did not open, and every timing assertion in this suite would still pass"
|
||||
fi
|
||||
pass "(g) detection is driven by netlink events, not by the poll fallback"
|
||||
|
||||
# ── (h) churn damping engages on a genuinely flapping interface ──────────
|
||||
#
|
||||
# This is load-bearing twice over. It is what stops a flapping interface
|
||||
# logging a recovery per cycle, and — since the detach edge now withdraws the
|
||||
# peers that interface carried — it is also the only thing bounding how often
|
||||
# that withdrawal can fire. Nothing exercised it: every flap elsewhere in this
|
||||
# suite is a single down/up with long settles either side, which is precisely
|
||||
# the shape the damper ignores.
|
||||
#
|
||||
# Four bindings that each die well inside MIN_STABLE_BINDING (10 s). The
|
||||
# streak crosses CHURN_THRESHOLD (3) on the third, which is the edge that
|
||||
# announces itself.
|
||||
log "(h) flapping $LAB_IFACE to drive the churn guard"
|
||||
for _ in 1 2 3 4; do
|
||||
docker exec "$NODE_A" ip link set "$LAB_IFACE" down
|
||||
sleep 1
|
||||
docker exec "$NODE_A" ip link set "$LAB_IFACE" up
|
||||
sleep 2
|
||||
done
|
||||
|
||||
if ! wait_for_at_least 30 1 log_count "$NODE_A" "keeps dying immediately after binding"; then
|
||||
docker logs "$NODE_A" 2>&1 | grep -i "ethernet" | tail -20 >&2
|
||||
fail "(h) four short-lived bindings did not engage the churn guard"
|
||||
fi
|
||||
pass "(h) a flapping interface engages churn damping"
|
||||
|
||||
# Having engaged, the guard must hold health rather than announcing each bind.
|
||||
# The failure this catches is a damper that counts but does not damp.
|
||||
recoveries_during_churn="$(log_count "$NODE_A" "Ethernet interface recovered")"
|
||||
if [ "$recoveries_during_churn" -gt 6 ]; then
|
||||
fail "(h) node-a announced $recoveries_during_churn recoveries; the guard \
|
||||
counted the churn but kept announcing through it"
|
||||
fi
|
||||
pass "(h) churn suppressed the per-cycle recovery announcements"
|
||||
|
||||
# And it is not a latch: once a binding lasts, the interface is announced
|
||||
# again and the node returns to Running on its own.
|
||||
log "(h) letting $LAB_IFACE settle"
|
||||
docker exec "$NODE_A" ip link set "$LAB_IFACE" up >/dev/null 2>&1 || true
|
||||
if ! wait_for 60 "present" iface_field "$NODE_A" lab presence; then
|
||||
fail "(h) node-a did not rebind after the flapping stopped"
|
||||
fi
|
||||
if ! wait_for 60 "running" node_state "$NODE_A"; then
|
||||
fail "(h) node-a stayed Degraded after the flapping stopped"
|
||||
fi
|
||||
pass "(h) a settled interface is announced again after churn"
|
||||
|
||||
# ── final log hygiene ────────────────────────────────────────────────────
|
||||
#
|
||||
# Four outages happened above (start, down, delete, and node-b's end of the
|
||||
# delete). A generous ceiling still catches the failure mode that matters: a
|
||||
# retry loop logging per attempt would be in the hundreds by now.
|
||||
edges="$(log_count "$NODE_A" "Ethernet interface")"
|
||||
[ "$edges" -lt 40 ] \
|
||||
|| fail "node-a logged $edges interface lines; the edges are not edges"
|
||||
pass "log volume stayed proportional to edges, not to retries"
|
||||
|
||||
echo
|
||||
echo "ALL PASSED"
|
||||
Reference in New Issue
Block a user