docs(netmon): state the probe's measured syscall cost

The per-peer probe's cost was given as five syscalls read off the code
rather than measured, and the backstop timer's as a few syscalls per
period. Counted with strace on Linux, a probe that gets an answer makes
five (socket, bind, connect, getsockname, close) and one with no route
makes four. A build with debug assertions adds an fcntl before each close,
from the standard library's open-descriptor check, which is why a test
binary counts six.

At the 128-peer default a sample is 640 syscalls, and a reported change
under the default debounce costs two to nine samples, 1280 to 5760
syscalls. State those figures where the cost is described, and describe
the backstop as costing one sample per period.
This commit is contained in:
Johnathan Corgan
2026-09-14 16:09:10 +00:00
parent cca6ed2c62
commit 3ec75e8385
2 changed files with 33 additions and 19 deletions
+9 -4
View File
@@ -275,10 +275,15 @@ one would put a DNS lookup on the sample path; the address becomes numeric as
soon as an authenticated packet arrives from the peer). A node holding no peers
detects nothing, which is correct — it has nothing bound to the old path.
The cost is five non-blocking syscalls per peer per sample, read from the
probe's own code rather than measured: `socket(2)` and `bind(2)`, a `connect(2)`
that sends no packet, a `getsockname(2)`, and the `close(2)` the socket takes on
drop. Nothing goes on the wire and no name is resolved.
The cost is five non-blocking syscalls per peer per sample, counted with
`strace` on Linux: `socket(2)` and `bind(2)`, a `connect(2)` that sends no
packet, a `getsockname(2)`, and the `close(2)` the socket takes on drop. A peer
with no route costs four, because the lookup fails at `connect(2)`. Nothing goes
on the wire and no name is resolved. At the default `max_peers` of 128 that is
640 syscalls per sample. A detected change is resampled until it settles, so
with the default `debounce_ms` it costs between two samples and nine, which is
between 1280 and 5760 syscalls at 128 peers; the backstop timer also takes one
sample every `poll_interval_secs` whether or not anything moved.
`node.limits.max_peers` bounds the per-sample total only where it is set: at
`max_peers: 0`, which means unlimited, there is no bound and the cost tracks the
live peer count instead.
+24 -15
View File
@@ -51,9 +51,9 @@
//! One local source address per peer: for every peer whose transport address is
//! a numeric IP endpoint, the address the kernel would pick to reach *that
//! peer*. A connected-but-never-sending UDP socket makes the kernel run its
//! route lookup and bind the source address it would use; five syscalls, read
//! off [`NetFingerprint::sample`] rather than measured, no packets, no name
//! resolution, and it works identically on every platform std supports.
//! route lookup and bind the source address it would use; five syscalls per
//! peer (see [`NetFingerprint::sample`] for the measured cost), no packets, no
//! name resolution, and it works identically on every platform std supports.
//!
//! Keying on peers bounds the *reaction* — only the peers a change names are
//! acted on — and does not bound the *sampling*. One roaming peer still makes
@@ -222,18 +222,26 @@ struct PeerPath {
impl NetFingerprint {
/// Probe every target and record the local address the kernel picks.
///
/// Five non-blocking syscalls per target: `socket(2)` and `bind(2)` behind
/// `UdpSocket::bind`, a `connect(2)` that sends no packet, a
/// `getsockname(2)`, and the `close(2)` the socket takes on drop. No I/O
/// wait, no name resolution, and no allocation beyond the map.
/// Five non-blocking syscalls per target, counted with `strace -f` on
/// Linux: `socket(2)` and `bind(2)` behind `UdpSocket::bind`, a
/// `connect(2)` that sends no packet, a `getsockname(2)`, and the
/// `close(2)` the socket takes on drop. A target with no route costs four,
/// because `connect(2)` fails and `getsockname(2)` is never reached. A
/// build with debug assertions on adds a sixth to each, the `fcntl(2)` std
/// uses to check a descriptor is still open before closing it, so count
/// against a release build. No I/O wait, no name resolution, and no
/// allocation beyond the map.
///
/// The count matters because a debounced handover resamples: up to
/// `MAX_DEBOUNCE_ROUNDS` rounds plus the settled sample, times the peers
/// held. `node.limits.max_peers` bounds that only where it is set —
/// the value 0 means unlimited, and there the cost tracks the live peer
/// count instead. It runs inline in the detector's own task rather than
/// through `spawn_blocking`, which is what keeps it off every other task
/// regardless.
/// The count matters because a debounced handover resamples: the sample
/// that saw the move, then up to `MAX_DEBOUNCE_ROUNDS` more until two
/// consecutive samples agree. At 128 peers, the `node.limits.max_peers`
/// default, that is 640 syscalls per sample, and a reported change under a
/// non-zero debounce costs from 1280 (settled on the first resample) to
/// 5760 (still moving after every round). `node.limits.max_peers` bounds
/// that only where it is set — the value 0 means unlimited, and there the
/// cost tracks the live peer count instead. It runs inline in the
/// detector's own task rather than through `spawn_blocking`, which is what
/// keeps it off every other task regardless.
pub(in crate::node) fn sample(targets: &[ProbeTarget]) -> Self {
Self {
sources: targets
@@ -541,7 +549,8 @@ struct WakeSource {
/// either of which would otherwise leave the node noticing nothing at all.
/// Keeping the period the poller would have used makes an event-driven
/// backend a strict latency improvement rather than a replacement that can
/// regress, for the cost of a few syscalls per period.
/// regress, for the cost of one sample per period: five syscalls per probed
/// peer, or 640 at the default of 128 peers.
timer: tokio::time::Interval,
}