Files
fips/docs/reference
fr34akyandJohnathan Corgan 957e05b283 feat(netmon): fingerprint the path to each peer, not the host
The medium-change detector sampled two host-wide signals: the source
address the routing table would pick for an off-link destination, and the
set of up, non-loopback interface addresses. The second one was the
problem. It enumerated every address the host had, so a docker bridge
coming up, a VPN connecting or a container network appearing moved the
fingerprint with no peering affected at all — and the reaction to a moved
fingerprint is to drop every connected UDP socket and heartbeat every
peer. Self-healing, so it cost work rather than connectivity, but on any
host running containers it could fire repeatedly for nothing.

Ask the question per peer instead. For each peer whose transport address
is a numeric IP endpoint, `connect(2)` a UDP socket to it and read back
the local address — the same no-packets operation, aimed at the peers we
actually hold rather than at a documentation prefix. The fingerprint
becomes the set of local addresses the kernel would use to reach our
peers.

That is the quantity the reaction cares about. The stale `connect(2)`
this subsystem exists to repair pinned a local source address chosen for
one destination, so measuring the same thing for the same destinations
asks the kernel the question the bug is about rather than a proxy for it.
Three things follow:

- Interfaces the node does not peer over cannot move it, by construction
  rather than by a filter guessing which interface names are
  infrastructure. `docker compose up` moves nothing.
- On-link peers become visible. A peer on the same LAN is reached by its
  subnet route, and the old probe followed the *default* route by
  construction, so it looked straight past that path.
- A more specific route moving under one peer is representable at all,
  which no single host-wide sample could be.

Samples are compared over the *intersection* of their peer sets, never
the union, so peers joining and leaving cannot fire the fan-out on their
own. The sample is still adopted on the no-change path, or `last` would
freeze on the peer set the detector started with. A peer whose probe
stops answering is a move to "no route" and does count: that peer is
exactly the one now stranded.

**A peer's first sample is judged against its socket, not against
nothing.** The intersection rule skips a peer present in only one sample,
which is right for peer churn and wrong for the sample in which a peer
first appears, because that sample may already be the post-change one.
`last` gains a peer only at the first wake after it shows up in the
entity snapshot, so a medium change inside that window is consumed rather
than delayed: the peer's connected socket stays pinned to the path the
host has just left, and the peering black-holes until the liveness
timeout tears it down. A first-seen peer is therefore compared against
the source its own connected socket is bound to, where it has one.

One residual on that rule, stated exactly rather than understated: the
tick publishes the entity snapshot *before* it installs connected
sockets, so a socket installed on tick N is first visible on tick N+1,
and a peer whose path moves inside that window is first seen with `bound`
still `None` while genuinely holding a pinned socket. About one
`tick_interval_secs` per join. A hole, not a harmless skip.

**The probe binds the way the send path binds.** `open_connected_fd`
binds `local_addr` verbatim before connecting, so the socket keeps a
configured address whatever the route says, while the probe took the
kernel's choice. Under a non-wildcard `bind_addr` the two answered
different questions and every first-seen peer reported a phantom move.
The probe now binds what the transport binds, address only and port 0;
under the default wildcard bind nothing changes.

Operator-facing corrections in the same surface. `PeerSourceMove::before`
was recorded on every move and then wildcarded away by the only thing
that read it, so the log said where a peer moved to but not where from —
and the address it moved *from* is the one the stale `connect(2)` had
pinned. Both ends are rendered now. The peer id used a private four-byte
hex helper that duplicated `NodeAddr::short_hex` and dropped the `...`
suffix every other operator surface prints; the duplicate is gone. The
cost note said three syscalls per target: `UdpSocket::bind` is a
`socket(2)` and a `bind(2)`, and the socket takes a `close(2)` on drop,
so it is five. In the same place, "bounded by `node.limits.max_peers`"
does not hold when that value is 0, which the configuration defines as
unlimited.

This deletes `interface_addrs()`'s only call site, and with it the
`getifaddrs` walk and its `sockaddr` decoding. It therefore absorbs the
Android `getifaddrs` issue rather than leaving it to be fixed separately.

The peer table is reached through `entities_snapshot` rather than a new
sharing primitive, so the detector stays a detached task holding no node
state and taking no node lock. `PeerRow` gains a typed `probe_target`
rather than having the detector re-parse the display string next to it,
so a change to that string's rendering cannot silently leave the detector
with an empty table and no way to notice.

Peers that are not probeable IP destinations contribute nothing and need
no per-transport special-casing here: a MAC on Ethernet or BLE, a .onion
or Nym recipient behind a local proxy, a scoped IPv6 literal, and a peer
still carrying its configured hostname all arrive as `None`. The last is
deliberate — resolving one would put a DNS lookup with its timeouts on
the sample path — and the window is small, since the address is replaced
by the observed numeric source the first time an authenticated packet
arrives. A node with no peers detects nothing, which is right: nothing is
bound to the old path.

Also corrects two doc claims that did not match the code, both in the
text being rewritten: `transport::watcher` has no consumer besides this
module, so it is not "shared with the interface binder"; and a
connection-oriented transport does not "re-dial on send" in the case that
matters, because `send_async` only dials when the pool holds no
connection for the address and a connection stranded by a medium change
is still in the pool — it is evicted after a write to it fails.

Tests, each run against the defect it guards rather than only against the
fix:

- The churn rules are mutation-checked: iterating the union instead of
  the intersection fails four tests, and dropping the sample-adoption
  fails the one that pins a newly joined peer entering the comparison.
- The snapshot seam is table-driven over six address shapes and fails if
  the publish site stops populating `probe_target` — nothing renders that
  field, so nothing else would have caught it.
- A namespace test brings up a dummy interface with its own subnet and
  asserts the fingerprint does not move, then puts a more specific route
  to the peer out of that same interface and asserts it does, so the
  negative half cannot pass because sampling quietly stopped working. The
  same test now asserts that an unconstrained probe answers with the
  carrier and a constrained one with the other address, and that the two
  differ — the disagreement that would otherwise report a phantom move on
  every first-seen peer. Ignoring the constraint in `preferred_source`
  fails it.
- The two links carrying the pinned source were asserted by nothing.
  Substituting `None` where `probe_targets` reads the row, or where
  `sample` writes the fingerprint, left the whole suite green while
  silently restoring the bug the first-sight rule exists to fix. Both are
  asserted now, and both mutations fail.
- Every live-probe test passed if `preferred_source` returned `None` for
  everything: two all-`None` samples are self-consistent, the
  recorded-keys test never inspected a value, and the loopback assertion
  skipped through its `if let`. A probe to loopback must now answer with
  loopback, which holds on any host that can run the suite, including one
  started with `--network none`.
- `reports_are_spaced_out_under_clean_flapping` polled at one second
  against a one-second pacing floor, so the two were indistinguishable
  and deleting the pacing block still passed. The wake is 100ms now.
- The pinned-source publish test was Linux-gated though `open_connected_fd`
  and the field it asserts are available on macOS too. Widened, along
  with the two sibling tests on the same helper.
- The detector's netlink subscription asserts the route groups rather
  than logging a decline, so a wrong group mask reds instead of passing
  quietly.
2026-09-09 18:45:42 +00:00
..
2026-08-30 10:42:59 +00:00
2026-08-30 10:42:59 +00:00

Reference

Information-oriented technical descriptions for lookup on demand. Reference content describes what is: wire formats, configuration keys, command-line flags, control-socket commands, default values, file paths, exit codes. It is consulted, not read end-to-end.

Reference is austere by design: minimal narrative, no opinions, no guidance on when to use a feature. The "why" lives in design/; the "how do I accomplish X" lives in how-to/.

Available Reference

Document Scope
wire-formats.md All FMP and FSP message byte layouts, encapsulation walkthrough
configuration.md Full YAML configuration reference for the daemon and gateway
security.md nftables baseline, peer ACL, cryptographic primitives, rekey defaults, threat-resistance matrix
nostr-events.md Kind 37195 advert, Kind 21059 traversal signaling, Kind 10050 inbox relays
transports.md Per-transport statistics counter inventory
control-socket.md Line-delimited JSON control protocol for the daemon and gateway
native-api.md Native datagram API: the Rust surface, addressing and ports, errno table, ceilings, line protocol, command reference
cli-fips.md fips daemon CLI: options, exit codes, environment, files
cli-fipsctl.md fipsctl control-client: subcommands, options, exit codes
cli-fipstop.md fipstop live-status TUI: tabs, keybindings
cli-fips-gateway.md fips-gateway service CLI: options, exit codes, files