mirror of
https://github.com/jmcorgan/fips.git
synced 2026-10-05 11:08:25 +00:00
Two daemons whose only transports are interface-bound, run against a veth pair the harness creates, downs, deletes and recreates underneath them. Asserts the boot race (a daemon whose only interface is missing starts, reports the transport absent and the node Degraded, rather than exiting on NoTransports or skipping the transport for the life of the process), the late attach and discovery over it, the flap in both directions, destroy-and-recreate, and that an optional interface which never appears never moves node health. Also the log policy, which is the half that is easy to regress silently: absence is logged once on the edge and not once per retry; a required interface still absent past the ten-second bring-up window errors exactly once, while the optional one — absent just as long — stays silent; and that error is not repeated on a schedule. The detach edge is checked not to error, guarded by how long detection actually took, so a slow runner skips the check rather than failing on the harness's own latency. The containers run FIPS_TEST_MODE=default, not chaos. The chaos entrypoint waits up to 30 s for every configured Ethernet interface before starting the daemon, which is precisely the workaround under test — the daemon has to do its own waiting here or the suite proves nothing. Host-namespace ip(8) runs in a short-lived privileged container sharing the host network and PID namespaces, for the reason chaos/sim/veth.py documents: on macOS the containers live in the Docker VM, so ip(8) run on the macOS host could never reach them. Chaos ethernet transports are marked optional: true. In that harness a neighbour's interface disappearing is the scenario, not a fault — node_churn stops a container, which destroys its netns and with it both ends of every veth it held, so a surviving node watches a required interface vanish for the 30-90 s the neighbour is down, once per churn event. Reporting that at error is right for a deployment and wrong for a harness that tears the interface down on purpose; the mesh-wide zero-ERROR ceiling would have failed on injected chaos rather than on a defect. test(iface-binding): cover an interface present before the daemon starts Every scenario in the suite created its interface after the daemons were already running — that ordering is the boot race the suite was written for. But it means both nodes could only ever reach Present through binder_loop, so the inline bind in start_async, which is the ordinary case on a booted router, had no end-to-end coverage at all. That is where the churn guard went unseeded and the first detach stopped reaching node health, and no existing case could reach it: they all detach from a binding the loop created, which seeds the guard as a side effect. Case (f) adds a third node whose single required interface exists before its daemon does. The gate is what buys that ordering — the harness needs a running container to have a netns to move a veth into, but the daemon must not start until after the move, so node-c comes up parked on a file and the harness releases it once the interface is in place. Then one detach, on a binding the loop did not create, and the node must degrade. Verified against the defect rather than only against the fix: with the guard seed reverted, cases (a) through (e) all still pass and (f) is the only failure. A regression test that has never been seen to fail is a claim, not a test. It also asserts the reverse edge, so Degraded stays a level rather than a latch on this path too. test(iface-binding): assert the fast path and the churn guard Three gaps, two of them in tests that existed and asserted nothing. **The netlink path was never asserted to be in use.** The 1 s poll is a complete fallback and covers every wait in the suite, so the whole thing passed with `open_link_socket()` hardcoded to Err — the fast path could have been dead for a release and no test would have said so. The binder reports which backing it got at startup, so case (g) asks it directly rather than inferring from timing the poll would also satisfy, and the unit test that used to write `let _ = w.is_event_driven();` now asserts it on Linux, where the source is an unprivileged `AF_NETLINK` socket and falling back is a real loss rather than a sandbox's prerogative. **Churn damping had no end-to-end coverage**, which now matters twice over: it bounds the recovery announcements, and since the detach edge withdraws peers it is also the only thing bounding how often that withdrawal fires. Every flap elsewhere in the suite is a single down/up with long settles either side — exactly the shape the damper ignores. Case (h) drives four bindings that each die inside `MIN_STABLE_BINDING`, asserts the guard engages, asserts it then *suppresses* rather than merely counting, and asserts it is not a latch. **`a_poisoned_binding_does_not_strand_the_transport` discarded its result.** `let _ = eth.binding.tasks_alive();` left the entire point unasserted: reading a poisoned lock as "alive" would have the binder believe a dead binding healthy and never rebind, and treating it as an error would strand the transport. `false` is what routes it back through detach and rebind, so say so. `a_stop_racing_a_bind_leaves_nothing_behind` now asserts the error *kind*. `bind_and_spawn` refuses at its presence probe long before the post-store shutdown check, so `is_err()` alone passed on absence and would still pass with that check deleted. The test keeps the coverage it genuinely has — stop raises the flag before teardown, teardown leaves no socket and no loops — and says plainly that the race it is named for needs a bind that succeeds, which needs privilege no unit test has. Both new cases were verified against the defect: with the netlink source forced to Err, (g) fails; with `CHURN_THRESHOLD` raised out of reach, (h) fails. Nothing else in the suite notices either. One case was attempted and removed rather than shipped: `"interface replaced"` cannot be produced deterministically, because the delete that changes an ifindex fires a netlink event the binder acts on within microseconds, so `gone` wins the race. It passed about one run in three. reference/notes.md records the measurement and the two approaches that could work. Also fixes a real bug in the harness: `grep -q` under `set -o pipefail` exits on its first match, `docker logs` takes SIGPIPE, and the pipeline reports failure even though the line was found. That cost two false failures before it was spotted; `log_count` reads the stream to the end. test(chaos): cover an Ethernet rebind under active traffic The one case dynamic interface binding had no coverage for anywhere: a datagram crossing an Ethernet link while the interface underneath it goes away and comes back. No existing scenario could reach it, for two separate reasons. `ethernet-only` and `ethernet-mesh` both run with `traffic.enabled: false`, so no datagram crosses an Ethernet link in any test — `ethernet-only`'s own comment says exactly that, and names framing, the length field that trims NIC minimum-frame padding, and AEAD over Ethernet as unexercised because of it. And `link_flaps` cannot produce a rebind whatever it is pointed at: it simulates a down link with netem 100% loss, so the interface stays IFF_UP and the presence machine never sees an edge. `ethernet-mesh` has had link flaps enabled all along without once exercising a rebind. `node_churn` is what actually moves an interface. Stopping a container destroys its network namespace, deleting every veth in it — and deleting one end of a veth deletes its peer — so a *surviving* node watches its Ethernet interface disappear outright, and watches it return when the harness recreates the pair on restart. That is a real detach and a real rebind, driven from outside the daemon. The new scenario is a 4-node Ethernet ring with traffic on and one node churned at a time, with link flaps deliberately off so the only outage is a genuine interface removal and a traffic shortfall cannot be ambiguous between the two. Measured across four runs: 206-388 MB moved over Ethernet links while interfaces were being taken away underneath. It also needed an assertion that did not exist. Traffic results have always been written to `iperf3-results.json` and never read, so a scenario carrying `traffic.enabled: true` could have every session fail and still exit 0 on a green control plane — and a rebind under load is precisely what a tree snapshot cannot see. `min_traffic` counts sessions that finished with bytes actually received, treating iperf3's top-level `error` and a missing `end` block as zero, so a session only counts when it moved data. The baseline is calibrated against four runs rather than assumed: `max_roots` starts at the observed maximum plus one, and the site records the sample, its size, and why four runs is thin. The first draft asserted a single root and failed every run — the harness restores stopped nodes immediately before the final snapshot, so a just-restarted node has not re-parented yet and is briefly its own root. That is the scenario working. Wired into both runners, since a chaos scenario on one side only makes "local green" and "GitHub green" stop meaning the same thing; check-ci-parity was confirmed to fail on a one-sided addition before this was committed. The iface-binding suite's entry in the GitHub workflow's integration matrix moves here from the commit that introduced the presence machine. That commit declared the suite on GitHub before testing/iface-binding/ existed and before testing/ci-local.sh knew about it, so testing/check-ci-parity.sh failed there and the three workflow steps named files that were not yet in the tree. Registering both runners in the commit that adds the suite settles both.
363 lines
14 KiB
Rust
363 lines
14 KiB
Rust
//! Link-event sources for the interface presence watcher.
|
|
//!
|
|
//! The presence machine works on a 1-second poll alone. This module removes
|
|
//! the latency: where the kernel offers a link-event source, the binder blocks
|
|
//! on it and reacts in sub-second time, and the poll stays underneath as a
|
|
//! backstop rather than as the mechanism.
|
|
//!
|
|
//! | Platform | Source |
|
|
//! | -------- | ------ |
|
|
//! | Linux | netlink `RTNLGRP_LINK` (`RTM_NEWLINK` / `RTM_DELLINK`) |
|
|
//! | macOS, FreeBSD | `PF_ROUTE` socket, `RTM_IFINFO` |
|
|
//! | Fallback | poll `getifaddrs` + flags, 1 s |
|
|
//!
|
|
//! The messages themselves are deliberately **not parsed**. A link event is a
|
|
//! hint to re-run the presence probe, which is cheap and authoritative;
|
|
//! decoding `nlmsghdr`/`ifinfomsg` payloads to reach the same answer would add
|
|
//! a parser whose bugs would be presence bugs. Any event on the socket wakes
|
|
//! the binder, which then asks
|
|
//! [`interface_present`](super::io::interface_present).
|
|
//!
|
|
//! Construction is best-effort. A kernel or sandbox that refuses the socket
|
|
//! yields a watcher that never fires, and the binder degrades to its poll.
|
|
|
|
use std::os::unix::io::{AsRawFd, RawFd};
|
|
use std::sync::atomic::{AtomicBool, AtomicU32, Ordering};
|
|
use std::time::Duration;
|
|
|
|
use tokio::io::unix::AsyncFd;
|
|
use tracing::{debug, warn};
|
|
|
|
/// Consecutive receive errors before the event source is abandoned for the
|
|
/// caller's poll.
|
|
const ERROR_GIVE_UP: u32 = 5;
|
|
|
|
/// An owned link-event socket. Closes its descriptor on drop.
|
|
struct LinkEventSocket {
|
|
fd: RawFd,
|
|
}
|
|
|
|
impl AsRawFd for LinkEventSocket {
|
|
fn as_raw_fd(&self) -> RawFd {
|
|
self.fd
|
|
}
|
|
}
|
|
|
|
impl Drop for LinkEventSocket {
|
|
fn drop(&mut self) {
|
|
unsafe { libc::close(self.fd) };
|
|
}
|
|
}
|
|
|
|
impl LinkEventSocket {
|
|
fn recv(&self, buf: &mut [u8]) -> std::io::Result<usize> {
|
|
let n = unsafe { libc::recv(self.fd, buf.as_mut_ptr() as *mut libc::c_void, buf.len(), 0) };
|
|
if n < 0 {
|
|
Err(std::io::Error::last_os_error())
|
|
} else {
|
|
Ok(n as usize)
|
|
}
|
|
}
|
|
}
|
|
|
|
/// Open the platform's link-event socket, non-blocking.
|
|
#[cfg(target_os = "linux")]
|
|
fn open_link_socket() -> std::io::Result<LinkEventSocket> {
|
|
// RTMGRP_LINK. Spelled as a literal because the constant's name and
|
|
// availability differ across libc versions; the value is ABI.
|
|
const RTMGRP_LINK: u32 = 1;
|
|
|
|
let fd = unsafe {
|
|
libc::socket(
|
|
libc::AF_NETLINK,
|
|
libc::SOCK_RAW | libc::SOCK_NONBLOCK | libc::SOCK_CLOEXEC,
|
|
libc::NETLINK_ROUTE,
|
|
)
|
|
};
|
|
if fd < 0 {
|
|
return Err(std::io::Error::last_os_error());
|
|
}
|
|
let socket = LinkEventSocket { fd };
|
|
|
|
let mut sa: libc::sockaddr_nl = unsafe { std::mem::zeroed() };
|
|
sa.nl_family = libc::AF_NETLINK as u16;
|
|
sa.nl_groups = RTMGRP_LINK;
|
|
let ret = unsafe {
|
|
libc::bind(
|
|
fd,
|
|
&sa as *const libc::sockaddr_nl as *const libc::sockaddr,
|
|
std::mem::size_of::<libc::sockaddr_nl>() as libc::socklen_t,
|
|
)
|
|
};
|
|
if ret < 0 {
|
|
return Err(std::io::Error::last_os_error());
|
|
}
|
|
Ok(socket)
|
|
}
|
|
|
|
/// Open the platform's link-event socket, non-blocking.
|
|
#[cfg(not(target_os = "linux"))]
|
|
fn open_link_socket() -> std::io::Result<LinkEventSocket> {
|
|
// PF_ROUTE delivers RTM_IFINFO (and the rest of the routing messages) to
|
|
// every reader; no bind and no group selection exist for it.
|
|
let fd = unsafe { libc::socket(libc::PF_ROUTE, libc::SOCK_RAW, libc::AF_UNSPEC) };
|
|
if fd < 0 {
|
|
return Err(std::io::Error::last_os_error());
|
|
}
|
|
let socket = LinkEventSocket { fd };
|
|
|
|
let flags = unsafe { libc::fcntl(fd, libc::F_GETFL) };
|
|
if flags < 0 {
|
|
return Err(std::io::Error::last_os_error());
|
|
}
|
|
if unsafe { libc::fcntl(fd, libc::F_SETFL, flags | libc::O_NONBLOCK) } < 0 {
|
|
return Err(std::io::Error::last_os_error());
|
|
}
|
|
Ok(socket)
|
|
}
|
|
|
|
/// A source of "something about the links changed" wake-ups.
|
|
pub(crate) struct LinkWatcher {
|
|
/// `None` when no event source could be opened — the caller's poll is then
|
|
/// the whole mechanism, which is exactly the documented fallback.
|
|
inner: Option<AsyncFd<LinkEventSocket>>,
|
|
/// Consecutive receive errors. Reset by any successful read.
|
|
errors: AtomicU32,
|
|
/// Set once the source has been abandoned for good.
|
|
///
|
|
/// Abandonment has to outlive the future that decided it. `changed()` is
|
|
/// called fresh on every pass of the caller's `select!` and dropped
|
|
/// whenever the poll ticker wins, so a `pending()` inside that future
|
|
/// parks nothing beyond the current pass — without this flag the next pass
|
|
/// re-reads the dead socket, re-counts the error, and re-logs the
|
|
/// give-up warning, once per wake-up, forever.
|
|
given_up: AtomicBool,
|
|
}
|
|
|
|
impl LinkWatcher {
|
|
/// Open the platform link-event source, falling back to nothing.
|
|
pub(crate) fn new() -> Self {
|
|
let inner = match open_link_socket() {
|
|
Ok(socket) => match AsyncFd::new(socket) {
|
|
Ok(afd) => Some(afd),
|
|
Err(e) => {
|
|
debug!(error = %e, "Link event socket not registrable; polling instead");
|
|
None
|
|
}
|
|
},
|
|
Err(e) => {
|
|
debug!(error = %e, "No link event source available; polling instead");
|
|
None
|
|
}
|
|
};
|
|
Self {
|
|
inner,
|
|
errors: AtomicU32::new(0),
|
|
given_up: AtomicBool::new(false),
|
|
}
|
|
}
|
|
|
|
/// Whether an event source is actually backing this watcher.
|
|
pub(crate) fn is_event_driven(&self) -> bool {
|
|
self.inner.is_some()
|
|
}
|
|
|
|
/// Resolve when the kernel reports a link change.
|
|
///
|
|
/// Never resolves when no event source is available, which makes it safe
|
|
/// to `select!` against the poll ticker: the ticker simply always wins.
|
|
pub(crate) async fn changed(&self) {
|
|
let Some(afd) = &self.inner else {
|
|
std::future::pending::<()>().await;
|
|
unreachable!("pending never resolves")
|
|
};
|
|
|
|
// Already abandoned on an earlier pass. Park without touching the
|
|
// socket, so giving up costs one syscall in total rather than one per
|
|
// caller wake-up for the life of the process.
|
|
if self.given_up.load(Ordering::Relaxed) {
|
|
std::future::pending::<()>().await;
|
|
unreachable!("pending never resolves")
|
|
}
|
|
|
|
loop {
|
|
let Ok(mut guard) = afd.readable().await else {
|
|
// The registration died. Stop firing rather than spinning; the
|
|
// caller's poll continues to cover presence.
|
|
std::future::pending::<()>().await;
|
|
unreachable!("pending never resolves")
|
|
};
|
|
|
|
// Drain to WouldBlock so a burst of link messages is one wake-up
|
|
// and the socket buffer does not fill behind us.
|
|
let mut buf = [0u8; 4096];
|
|
let mut saw_event = false;
|
|
let mut failure = None;
|
|
loop {
|
|
match guard.try_io(|inner| inner.get_ref().recv(&mut buf)) {
|
|
Ok(Ok(n)) if n > 0 => saw_event = true,
|
|
// A zero-length read. Readiness is *not* cleared by
|
|
// `try_io` here — it clears only on `WouldBlock` — so
|
|
// breaking out plainly would leave `readable()` instantly
|
|
// ready with nothing to read, and this loop would spin
|
|
// without ever returning `Pending`. That starves the
|
|
// caller's `select!` of its poll ticker entirely, which
|
|
// takes presence detection down with it. Clear it by hand
|
|
// and treat it as a fault, so the give-up path applies.
|
|
Ok(Ok(_)) => {
|
|
guard.clear_ready();
|
|
failure = Some(std::io::Error::from(std::io::ErrorKind::UnexpectedEof));
|
|
break;
|
|
}
|
|
// A genuine socket error. Distinct from WouldBlock, and
|
|
// the distinction is the whole point: `try_io` clears
|
|
// readiness only on WouldBlock, so breaking out of a real
|
|
// error leaves `readable()` instantly ready, `recv`
|
|
// failing again, and the loop spinning a core flat with
|
|
// nothing logged. Clear it by hand and back off.
|
|
Ok(Err(e)) => {
|
|
guard.clear_ready();
|
|
failure = Some(e);
|
|
break;
|
|
}
|
|
// WouldBlock — readiness is cleared, drain complete.
|
|
Err(_) => break,
|
|
}
|
|
}
|
|
|
|
if saw_event {
|
|
self.errors.store(0, Ordering::Relaxed);
|
|
return;
|
|
}
|
|
|
|
if let Some(e) = failure {
|
|
let errors = self.errors.fetch_add(1, Ordering::Relaxed) + 1;
|
|
if errors == 1 {
|
|
// ENOBUFS is the realistic one: a burst of link events
|
|
// overflowed the socket buffer, so the kernel dropped some.
|
|
// Losing events is survivable — the caller polls — but the
|
|
// spin is not, and neither is doing it silently.
|
|
warn!(error = %e, "Link event source read failed");
|
|
}
|
|
if errors >= ERROR_GIVE_UP {
|
|
self.given_up.store(true, Ordering::Relaxed);
|
|
warn!(
|
|
errors,
|
|
"Link event source is not recoverable; falling back to \
|
|
polling for interface presence"
|
|
);
|
|
std::future::pending::<()>().await;
|
|
unreachable!("pending never resolves")
|
|
}
|
|
tokio::time::sleep(Duration::from_millis(100) * errors).await;
|
|
}
|
|
}
|
|
}
|
|
}
|
|
|
|
#[cfg(test)]
|
|
mod tests {
|
|
use super::*;
|
|
|
|
/// The watcher must construct on any host, with or without a usable event
|
|
/// source, because the binder builds one unconditionally.
|
|
#[tokio::test]
|
|
async fn watcher_constructs_and_reports_its_backing() {
|
|
let w = LinkWatcher::new();
|
|
|
|
// On Linux the source is a plain `AF_NETLINK` socket in the
|
|
// `RTNLGRP_LINK` group, which needs no capability and no privilege —
|
|
// so on this platform "a sandbox might refuse it" is not a licence to
|
|
// accept either answer. Discarding the result, which this test used
|
|
// to do, meant nothing anywhere asserted that the event path exists:
|
|
// the 1 s poll is a complete fallback, so the entire suite passed with
|
|
// the source unavailable and no test could tell.
|
|
#[cfg(target_os = "linux")]
|
|
assert!(
|
|
w.is_event_driven(),
|
|
"the netlink link-event source must open on Linux; \
|
|
falling back to the poll here is a silent loss of the fast path"
|
|
);
|
|
|
|
// Elsewhere both answers are legitimate, so pin only that asking is
|
|
// safe and that a watcher with no source parks rather than fires.
|
|
#[cfg(not(target_os = "linux"))]
|
|
{
|
|
let backed = w.is_event_driven();
|
|
assert!(
|
|
backed
|
|
|| tokio::time::timeout(Duration::from_millis(50), w.changed())
|
|
.await
|
|
.is_err(),
|
|
"a watcher with no source must never resolve"
|
|
);
|
|
}
|
|
}
|
|
|
|
/// A descriptor whose `recv` always fails must not become a busy loop.
|
|
///
|
|
/// `try_io` clears readiness only on `WouldBlock`. Breaking out of a real
|
|
/// error left `readable()` instantly ready, `recv` failing again, and the
|
|
/// loop spinning a core flat with nothing logged — the realistic trigger
|
|
/// being `ENOBUFS` when a burst of link events overflows the socket
|
|
/// buffer. A pipe stands in for that here: `recv` on one answers
|
|
/// `ENOTSOCK`, every time, which is exactly the shape of a persistent
|
|
/// error.
|
|
#[tokio::test]
|
|
async fn a_persistently_failing_source_gives_up_instead_of_spinning() {
|
|
let mut fds = [0i32; 2];
|
|
assert_eq!(unsafe { libc::pipe(fds.as_mut_ptr()) }, 0, "pipe()");
|
|
let (read_fd, write_fd) = (fds[0], fds[1]);
|
|
|
|
// AsyncFd requires a non-blocking descriptor.
|
|
let flags = unsafe { libc::fcntl(read_fd, libc::F_GETFL) };
|
|
assert!(unsafe { libc::fcntl(read_fd, libc::F_SETFL, flags | libc::O_NONBLOCK) } >= 0);
|
|
|
|
let watcher = LinkWatcher {
|
|
inner: Some(AsyncFd::new(LinkEventSocket { fd: read_fd }).expect("register")),
|
|
errors: AtomicU32::new(0),
|
|
given_up: AtomicBool::new(false),
|
|
};
|
|
|
|
// Keep producing readiness edges. The error arm calls `clear_ready`,
|
|
// and a descriptor that was already readable before re-registration
|
|
// may never deliver another edge on its own — which would stall the
|
|
// loop at one error and hide whether the give-up path works. A steady
|
|
// trickle stands in for the burst of link events that provokes the
|
|
// real failure.
|
|
let writer = tokio::task::spawn_blocking(move || {
|
|
for _ in 0..200 {
|
|
if unsafe { libc::write(write_fd, b"x".as_ptr().cast(), 1) } < 0 {
|
|
break;
|
|
}
|
|
std::thread::sleep(Duration::from_millis(25));
|
|
}
|
|
unsafe { libc::close(write_fd) };
|
|
});
|
|
|
|
// Never resolves — there is no event to report — but it must reach the
|
|
// give-up state rather than burn until the timeout.
|
|
let fired = tokio::time::timeout(Duration::from_secs(5), watcher.changed()).await;
|
|
assert!(fired.is_err(), "a failing source must not report an event");
|
|
assert!(
|
|
watcher.errors.load(Ordering::Relaxed) >= ERROR_GIVE_UP,
|
|
"the error path must count, back off and stop, not spin silently"
|
|
);
|
|
|
|
writer.abort();
|
|
}
|
|
|
|
/// A watcher with no event source must never resolve, so a `select!`
|
|
/// against the poll ticker degrades cleanly instead of spinning.
|
|
#[tokio::test]
|
|
async fn a_sourceless_watcher_never_fires() {
|
|
let w = LinkWatcher {
|
|
inner: None,
|
|
errors: AtomicU32::new(0),
|
|
given_up: AtomicBool::new(false),
|
|
};
|
|
let fired = tokio::time::timeout(std::time::Duration::from_millis(50), w.changed()).await;
|
|
assert!(fired.is_err(), "sourceless watcher resolved");
|
|
}
|
|
}
|