docs: document dynamic interface binding

The transport-layer design gains an Interface Presence section: the three
deployment scenarios that motivated it, the presence machine and its two
invariants, why presence is IFF_UP and not IFF_UP|IFF_RUNNING, interface
identity and the recreated-netdev case, the detection sources and their two
rate limits, the optional policy, health as a level rather than a latch, the
egress-MTU consequence, the logging rules, and what the mechanism retires.
It records why the error deadline is a single ten-second window rather than
a repeating severity ladder, so the ladder does not come back.

The control-socket and fipsctl references described show_transports without
its `interface` block, and the transport-layer state machine still implied
that Up meant bound. Both now say what Up and presence each describe and how
they differ, since "transport up, interface absent" is a normal state an
operator will meet and would otherwise read as a contradiction.

The configuration reference documents `optional` with the absence, log and
retry behaviour of each setting, and the ground-up tutorial's dongle example
gains optional: true, which is what that example is actually for.
This commit is contained in:
Arjen
2026-09-01 09:54:43 +01:00
parent d059ef1434
commit cc3c7c4588
6 changed files with 564 additions and 2 deletions
+170
View File
@@ -7,6 +7,176 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
## [Unreleased]
### Added
- Dynamic interface binding for the Ethernet transport. An interface-bound
transport is now a long-lived object that is *sometimes bound*: the interface
it names need not exist when the daemon starts, may appear minutes later, and
may vanish and return mid-operation. `start_async` returns `Ok` with the
transport **absent** rather than failing, and a per-transport binder task
binds when the interface appears, unbinds when it goes away, and rebinds when
it returns. Start-time absence and runtime detach are one code path.
Detection is event-driven where the kernel offers a source — netlink
`RTNLGRP_LINK` on Linux, `PF_ROUTE` on macOS and FreeBSD — with a 1 s
`getifaddrs` poll underneath as a backstop. Presence means
`IFF_UP` — the interface exists and is administratively up — and
deliberately not `IFF_RUNNING`. Binding needs no carrier and a socket
outlives a carrier flap, so a bridge with nothing plugged into it (`br-lan`
on a wifi-only router) is bound and healthy rather than permanently
`Degraded`, and starts carrying traffic the moment a port comes up. Whether
an interface has carrier is reported separately as `interface.carrier` in
`show_transports`, never acted on.
This closes the OpenWrt boot race (procd starts `fips` before wifi has
created `fips-mesh0` / `fips-ap0`; both transports were skipped for the life
of the process while the 802.11s peer link formed anyway, so the node looked
healthy and reached nothing), the intermittent-adapter case, and the
mid-operation `wifi reload` that destroyed and recreated an interface under a
live socket.
- `transports.ethernet.*.optional` (bool, default `false`). Naming an interface
in configuration is a statement that you expect it, so the default is to
complain: while a required interface is missing the node reports `Degraded`
and logs the edge, at a severity that follows how long the absence lasts
(see below). `optional: true` makes absence silent (`info` on the edge, no
health impact) for hardware that is legitimately not always there. It describes the interface's *presence*, not the transport's
importance — an optional interface that is present is used exactly as hard as
any other — and no value of it makes a missing interface fatal at startup.
- `fipsctl show transports` reports interface presence per transport under a new
`interface` block: `name`, `presence` (`absent` / `binding` / `present`),
`policy` (`required` / `optional`), `since_secs`, `binds` and
`failed_attempts`. The original boot-race bug was expensive precisely because
nothing an operator could see said the node was deaf.
- `testing/iface-binding/` integration suite (`ci-local.sh --only
iface-binding`, and a GitHub matrix leg): two daemons whose only transports
are interface-bound, run against a veth pair the harness creates, downs,
deletes and recreates underneath them. Asserts the boot race, the late
attach and peering over it, the flap in both directions,
destroy-and-recreate, that an `optional` interface never moves node health,
and that absence is logged once on the edge rather than once per retry.
### Changed
- `Degraded` is now a level rather than a latch. The supervisor's reason set
was monotonic, which was correct while no child could recover; with recovery
it would have meant "something broke at some point since boot" rather than
"something is broken now". Interface absence is tracked in its own reversible
set and node health is recomputed on every transition **in both directions**,
so plugging the WAN back in clears `Degraded` without a restart. A transport
whose interface is absent still counts as up, so a single-interface node that
boots before its wifi degrades rather than exiting on "no transports".
- The OpenWrt package ships the `mesh0`/`mesh1` and `ap0`/`ap1` Ethernet
transports **enabled** with `optional: true`, instead of commented out.
`fips-mesh-setup` and `fips-ap-setup` no longer comment-toggle blocks in
`fips.yaml`, and no longer tell the operator to restart the daemon after
creating an interface — the daemon binds it on its own. `phy0-sta0` (`wwan`)
is marked `optional: true` for the same reason: it only exists while a radio
is in station mode.
- The Ethernet receive loop backs off and exits on a dead socket instead of
spinning on `Err` with a `warn!` per iteration, and the ad-hoc ENXIO
socket-reopen in the beacon sender is gone. Both hand recovery to the
presence machine: one mechanism for every cause rather than one hack per
symptom. Beacons pause while an interface is absent.
- The absence edge is not itself an error, and there is exactly one deadline
after it. An interface missing when the daemon starts logs at `info` — that
is the boot race the mechanism exists to absorb, not a fault — and a runtime
detach at `warn`, because a link coming and going is ordinary weather for a
mesh daemon. Ten seconds is the whole grace: past it, absence is no longer a
race against a radio or a container, so a **required** interface still
missing is reported once at `error`. Start-time absence and a runtime detach
share that one deadline rather than getting one each. An `optional`
interface never reaches `error`. Node health does not wait for any of it,
publishing `Degraded` on the first edge either way.
- The TUN boundary's TCP MSS clamp now tracks the node's egress MTU at
runtime instead of freezing it at startup. `transport_mtu()` is the minimum
across *bound* transports, so a transport that binds minutes after start can
be the narrow one — but the TUN reader and writer were handed a `u16`
computed once when they spawned, while every other consumer
(`show_status`, the control-socket snapshot, the session-layer fragmentation
check) read it live. A node could therefore report one effective IPv6 MTU
and clamp to another. The ceiling is now shared with those threads and
recomputed whenever the bound set changes, in both directions: a narrow
interface appearing tightens it, and its departure releases it. MSS is
negotiated per connection, so a change applies to connections opened after
it; existing ones are not disturbed.
- Rebinds that keep succeeding into a socket that dies moments later are
damped: consecutive bindings shorter than ten seconds back off on the
1 s → 30 s curve, and past three of them the binder stops announcing each
bind as a recovery until one lasts. Undamped, a persistently broken socket
behind a healthy interface produced a log pair and a `Degraded`→`Running`
health flap every second.
- A bind failure that is **not** absence — no `CAP_NET_RAW`, no readable
`/dev/bpf*`, a buffer the kernel refused — fails the daemon's start as it
always has, rather than being waited out. It is a fault, not a state, and
will not resolve on its own; only a missing interface is retried at start.
A non-absence failure during a later rebind still backs off, since the node
is serving by then.
- The binder cannot outlive its transport, and a teardown that races a bind
cannot leave a live receive loop on a socket nothing owns. A shared stop flag
is raised before teardown and checked by the binder after it stores a
binding, so whichever order the two interleave exactly one of them cleans up;
`EthernetTransport` gained a `Drop` that raises the flag, aborts the binder
and releases the socket, for handles dropped without `stop_async`.
- Presence edges are published with `try_send` and retried on the next tick
rather than awaited. A bounded channel could previously park the binder
mid-publish — a health channel able to deadlock the machine whose health it
carries — freezing the interface in whatever state it held.
- `TransportHandle::is_bound()` joins `is_operational()`: the latter means the
transport was *started*, which for an interface-bound transport no longer
implies a live socket. `Node::transport_mtu` now filters on the former,
because an interface that has never existed was clamping the whole node's
IPv6 MTU to a number derived from absent hardware.
- An interface deleted and recreated under the same name is detected as a
detach. Both backends bind by device rather than by name, so the old socket
is attached to nothing while the name still resolves — and a stale
`AF_PACKET` socket never becomes readable, so nothing errors and nothing
exits. Detection previously rested entirely on the beacon sender failing,
which a node with `announce: false` does not have. The bound interface index
is now captured at bind and compared on every poll.
- The link-event watcher distinguishes a genuine receive error from
`WouldBlock`. `try_io` clears readiness only on the latter, so a persistent
error — `ENOBUFS` after a burst of link events overflows the socket buffer —
span a core flat with nothing logged. Errors are now counted, logged once,
backed off, and after five the source is abandoned for the presence poll.
- Presence probes are coalesced to at most ten a second. Linux netlink is
filtered to `RTNLGRP_LINK`, but `PF_ROUTE` has no group filter, so the macOS
source delivers every routing message on the host — route churn, ARP, DHCP
renewals, a VPN going up and down — and each would otherwise drive a full
`getifaddrs` walk.
- CI runs the library tests on musl (Alpine) as well as glibc. Presence is
built on `getifaddrs` and `ifa_flags`, musl reimplements both independently,
and the interfaces this feature exists for (`fips-mesh0`, `fips-ap0` on
OpenWrt) are unbridged with no IP address at all — the case where
implementations most plausibly differ. It was previously asserted on a libc
no test had ever run it against, on the target it was written for.
- Interface presence state ignores lock poisoning. Treating a poisoned lock as
a failure meant reading "no socket, tasks dead", which is the destructive
direction: a transport reporting itself present while every send fails, or a
binder tearing down and rebinding every second while teardown silently
declined to abort anything.
### Added (internal)
- `TransportError::InterfaceUnavailable { interface }`. A missing interface and
a typo'd interface name were previously the same flat
`StartFailed(String)`; nothing downstream could branch on absence.
### Fixed
#### Discovery
+307
View File
@@ -1010,6 +1010,313 @@ Transports begin in `Configured` state with all parameters set. `start()`
transitions through `Starting` to `Up` (operational). `stop()` moves to
`Down`. Transport failures move to `Failed`.
## Interface Presence
`Up` describes the *transport*, not the socket. An interface-bound transport
(today: Ethernet) carries a second, orthogonal state — whether it is bound
right now — and the two are independent: a transport is `Up` from the moment
it starts, whether or not the interface it names exists.
### The Gap This Closes
Three deployment scenarios exercise one missing mechanism:
- **Boot ordering.** On OpenWrt, procd starts `fips` before wifi has created
`fips-mesh0` / `fips-ap0`. Both transports were skipped and never retried,
while the 802.11s peer link formed anyway — that is mac80211, not the
daemon — so the node looked healthy and reached nothing. The failure was
expensive precisely because nothing an operator could see said the node was
deaf.
- **Intermittent hardware.** A USB ethernet adapter named in `fips.yaml` is
plugged in some days and not others. Its absence is normal and must be
silent; its arrival must bind without operator action.
- **Mid-operation restart.** `wifi reload` for a channel change destroys and
recreates the mesh interface within a couple of seconds. The socket dies,
the receive loop spun on `Err` with no backoff and no exit, and nothing
rebound.
These are not three features. They are one presence machine plus one policy
field. Before it existed, the first observation was final: an interface
missing at start was logged once and skipped for the life of the process, and
one that disappeared at runtime published a health change but was never
rebound.
### The Presence Machine
```text
Absent ──attach──> Binding ──ok──> Present
^ │ │
└──── fail/backoff ─┘ │
└──────────── detach ──────────────┘
```
`start_async` binds if it can and otherwise returns `Ok` with the transport
`Up` and `Absent`; a per-transport binder task then binds when the interface
appears, tears the socket down when it goes away, and rebinds when it
returns. Two invariants do the work:
- **The transport object survives detach.** Config, `TransportId`,
statistics and the neighbor buffer persist; only the file descriptor and
its loops go. A transport is never destroyed because its interface went
away.
- **Start-time absence and runtime detach are the same transition.** A node
that boots before its wifi and a node whose wifi reloads at 03:00 take one
code path. The old asymmetry — skip forever at start, publish health at
runtime — is gone.
`TransportError::InterfaceUnavailable` is what makes absence branchable. A
missing interface and a typo'd interface name were the same flat
`StartFailed(String)`, so nothing downstream could tell a state from a fault.
### What Counts as Present
Presence means `IFF_UP` — the interface exists and the operator has enabled
it — and deliberately **not** `IFF_RUNNING`.
Carrier and bindability are different questions, and only the second belongs
in a bind gate. An `AF_PACKET` socket on a carrier-less bridge is valid and
starts carrying traffic the instant a member port comes up, with no rebind:
the socket outlives the carrier. Gating on `IFF_RUNNING` bought nothing and
cost three things:
- `br-lan` on a router with nothing in its LAN ports is `UP` with
`NO-CARRIER`, so a healthy wifi-only router reported `Degraded` forever and
errored for a fault it did not have;
- every carrier flap the socket would have survived became an unbind/rebind
cycle — churn the presence machine then has to damp, a mechanism
compensating for a policy error;
- an 802.11s interface that reports `RUNNING` only once it has peered cannot
peer, because peering needs beacons, beacons need a bound socket, and the
gate refuses to bind. A deadlock reachable on the hardware this mechanism
was written for.
The signal `IFF_RUNNING` carries is not lost: `show_transports` reports
`interface.carrier` beside presence, so an operator can still tell a bound
transport carrying nothing from a working one. It is reported rather than
obeyed.
The probe is `getifaddrs` plus `ifa_flags` rather than an `SIOCGIFFLAGS`
ioctl: it needs no socket, so the watcher can probe before any file
descriptor exists, and it is spelled the same on Linux and the BSDs.
### Interface Identity
The configured name is the key, but a name is not a device. Both backends
bind by *device* — `AF_PACKET` stores `sll_ifindex`, a BPF descriptor follows
the interface it was attached to — so an interface deleted and recreated
under the same name leaves the socket attached to something that no longer
exists while the name resolves perfectly well.
Nothing else notices. A stale `AF_PACKET` socket never becomes readable, so
the receive loop neither errors nor exits, and send failures go to the caller
rather than to the binder. A listen-only node (`announce: false`, so no
beacon sender to fail) therefore sat `present` and deaf indefinitely after a
`wifi reload` — the original bug wearing a different hat. The bound index is
captured at bind and compared on every poll; a mismatch is a detach.
Hardware can also change underneath a name. If the name reappears with a MAC
other than the one last bound, that is a different device, so the cached
neighbor entries for that transport are dropped rather than resumed onto, and
the swap is logged at `warn`. Richer selectors (`match: { name | mac |
id_path }`) are deliberately deferred; the requirement here is only that FIPS
never silently resumes onto different hardware.
### Detection
| Platform | Source |
| -------- | ------ |
| Linux | netlink `RTNLGRP_LINK` (`RTM_NEWLINK` / `RTM_DELLINK`) |
| macOS, FreeBSD† | `PF_ROUTE` socket, `RTM_IFINFO` |
| Fallback | poll `getifaddrs` + flags, 1 s |
† Aspirational: the Ethernet transport is
`cfg(any(target_os = "linux", target_os = "macos"))`, so FreeBSD has no
interface-bound transport for a watcher to serve. The `PF_ROUTE` branch
compiles for the BSD family, but only macOS reaches it.
Where an event source exists, detection is sub-second. The poll stays
underneath as a backstop rather than as the mechanism, and must stay at ~1 s:
the probe is cheap, and letting the interval drift to tens of seconds
reintroduces exactly the latency the event source was added to remove.
Construction is best-effort — a kernel or sandbox that refuses the socket
yields a watcher that never fires, and the binder degrades to its poll.
Link-event payloads are **not parsed**. An event is a hint to re-run the
presence probe, which is cheap and authoritative; decoding
`nlmsghdr`/`ifinfomsg` to reach the same answer would add a parser whose bugs
would be presence bugs.
Two rate limits protect the binder from its own event source. Probes are
coalesced to ten a second, because `PF_ROUTE` has no group filter and
delivers every routing message on the host — route churn, ARP, DHCP renewals,
a VPN going up and down — each of which would otherwise drive a full
`getifaddrs` walk. And a persistently failing event source is counted, logged
once, backed off, and after five consecutive errors abandoned for the poll:
losing events is survivable because the poll is the backstop, but burning a
core on a socket that is readable-but-erroring is not.
Bind failures that are *not* absence back off 1 s → 30 s. Absence itself does
not back off; there is nothing to poll but the probe.
### Policy: `optional`
One field per transport, `transports.ethernet.*.optional`, default `false`:
| | absence | log | retries |
| --- | --- | --- | --- |
| `optional: false` (default) | node reports `Degraded` | `info` at boot / `warn` on a runtime detach, then `error` once if it lasts past 10 s | forever |
| `optional: true` | no health impact | `info`, and nothing after | forever |
Naming an interface in configuration is a statement that you expect it, so
the default is to complain; silence is opted into.
`optional` describes **the interface's presence, not the transport's
importance**. An optional interface that is present is used exactly as hard
as any other. No value of it makes a missing interface fatal at startup: the
only fatal case remains "no transports at all came up". If a deployment ever
needs absence to abort startup, that arrives as an explicit `on_absent: exit`
— never as a second meaning for `optional`.
A bind failure that is not absence — no `CAP_NET_RAW`, no readable
`/dev/bpf*`, a buffer the kernel refused — is a fault, not a state, and still
fails the daemon's start. Retrying those forever would convert a hard,
actionable deployment error into a daemon that retries a socket it can never
open behind a `Degraded` nobody is watching. Only absence is waited out at
start; a non-absence failure during a later *rebind* does back off, since by
then the node is serving and killing it would be the worse answer.
### Health
Node health is recomputed on every presence transition **in both
directions**, so a returning interface clears `Degraded` without a restart.
That makes `Degraded` a level rather than a latch: the supervisor's reason set
was monotonic, which was correct while no child could recover, but with
recovery it would have come to mean "something broke at some point since boot"
rather than "something is broken now". Absence lives in its own reversible
set, separate from the one-way `failed` set a start failure enters.
An absent transport still counts as *up*. It came up — `start_async` returned
`Ok` — so it does not push a single-transport node into the fatal
`NoTransports`, which would make a node that merely booted before its wifi
exit instead of waiting. Absence degrades; it never kills.
There is deliberately **no restart action in the supervisor FSM.** The
presence watcher and the rebind loop live inside the transport, next to the
file descriptor they manage, and once that exists a supervisor-authored retry
has nothing left to do — it would be a second mechanism racing the first for
the same socket. The supervisor learns about presence
(`Event::ChildAbsent` / `Event::ChildPresent`) and republishes health; it does
not drive rebinding.
Peer state needs no separate grace period. A send over an absent interface
returns `InterfaceUnavailable` and the peer entry survives untouched, so peers
are already held across a detach and resume when the interface returns; the
liveness reaper is the effective linger bound. A recreated mesh interface
comes back with the same MAC (it is derived from the phy) and Noise sessions
are keyed on the remote peer, so a local rebind is invisible to peers. A
dedicated linger timer would be a second answer to a question already
answered.
### Logging
Edges, never attempts. A loop that logs per attempt reproduces the hot log
spin this mechanism removed, at 1–30 s intervals forever on any router with
an unplugged WAN — and operators learn to filter it, which is how the next
real failure gets missed.
The edge itself is not an error. An interface missing when the daemon starts
and bound a moment later is the ordinary case the mechanism exists to absorb,
so it is `info`; calling it an error at t=0 and "recovered" at t=0.2 s is the
cry-wolf failure this rule exists to prevent. A runtime detach is `warn` — a
link coming and going is ordinary weather for a mesh daemon.
There is exactly one deadline. Ten seconds is the window in which absence
could still be a race — a radio, a container, a veth arriving late. Past it a
**required** interface is a fault an operator has to fix, and it is reported
once at `error`. Start-time absence and a runtime detach share that deadline
rather than getting one each, for the same reason they share a code path
everywhere else here. An `optional` interface never reaches `error`; that is
what `optional` means.
Once, not repeated. This was a 1 m / 10 m / 1 h ladder that re-announced the
same fact at rising severity and then went permanently quiet after an hour,
which got both halves wrong: it used the log as a store for something already
published continuously as state, and it stopped mentioning a fault that was
still live. Duration belongs in `interface.since_secs` and in how long
`Degraded` has been held, where a monitor can threshold it per deployment
instead of the daemon compiling one in.
Node health does not wait for the deadline. `Degraded` publishes on the first
edge, which is the signal an operator actually watches.
Successful rebinds are damped. Backoff covers failed binds; the opposite and
nastier case is binds that keep *succeeding* into a socket that dies moments
later, which a receive loop giving up on a persistent error while the
interface stays `UP` produces once per second, forever. The binder counts
consecutive bindings that die inside ten seconds, backs off on the same
1 s → 30 s curve, and past three of them stops announcing each bind as a
recovery — holding health where it is until a binding lasts.
### Egress MTU
`transport_mtu()` is the minimum across *bound* transports — `is_bound()`,
not `is_operational()`, because an interface-bound transport is operational
from the moment it starts whether or not it holds a socket. Filtering on the
weaker predicate let a transport whose interface had never appeared set the
whole node's IPv6 MTU from hardware that was not present.
Since a transport can now bind long after start, that minimum moves at
runtime, and every consumer has to read it live. `show_status`, the
control-socket snapshot and the session-layer fragmentation check always did.
The TUN reader and writer did not: they were handed a `u16` at spawn, so a
narrow interface binding later never tightened the TCP MSS clamp and the node
reported one effective MTU while clamping to another. The ceiling is now
shared with those threads — an atomic beside the per-destination
`path_mtu_lookup` they already read on the same packet — and recomputed on
every change to the bound set.
Both directions, for the same reason `Degraded` is a level rather than a
latch: a narrow interface arriving must tighten the clamp or traffic
egressing over it is clamped too loose, and that interface leaving must
release it or unplugging a low-MTU adapter leaves the node over-clamped until
it restarts. MSS is negotiated per connection at SYN time, so a change binds
connections opened after it and leaves established ones alone.
### Observability
`show_transports` carries an `interface` block per interface-bound transport:
netdev name, `presence` (`absent` / `binding` / `present`), `carrier`,
`policy` (`required` / `optional`), `since_secs`, `binds` and
`failed_attempts`. The two counters separate an interface that is flapping
from one that is there and refusing to bind, and `since_secs` measures the
absence *episode* rather than the phase — a bind that fails walks
`Absent → Binding → Absent`, and restarting the clock on those edges would
report a permanently unbindable interface as one second old forever.
`fipstop`'s transports view names the netdev and the absence policy in their
own columns and shows presence in the State column for these transports,
because `state` reads `up` from the moment the transport starts and is
therefore precisely the wrong answer in the one case someone is scanning that
column for. Both render sites sort by ascending transport id — creation
order, and so grouped by transport type — rather than by `HashMap` iteration
order, which was arbitrary and differed on every daemon restart.
### What This Retires
- The `hotplug.d/net` rule that restarted the daemon when the FIPS radio
interfaces appeared, and the `wifi down; wifi up; sleep` dance provisioning
performed to sequence around the race.
- The YAML comment-toggling in `fips-mesh-setup` / `fips-ap-setup` — the mesh
and AP blocks ship enabled with `optional: true` and simply wait.
- The ad-hoc ENXIO socket reopen in the beacon sender: beacons now pause while
absent because the task does not exist then, and recovery is the presence
machine's job. One mechanism for every cause rather than one hack per
symptom.
- The start-time versus runtime asymmetry in the supervisor.
It also covers the case none of those workarounds did: an interface that flaps
while the daemon is running.
## Implementation Status
| Transport | Status | Notes |
+1 -1
View File
@@ -52,7 +52,7 @@ prints the response's `data` object as pretty JSON.
| `show mmp` | `show_mmp` | MMP metrics summary: per-peer link-layer metrics and per-session session-layer metrics. |
| `show cache` | `show_cache` | Coordinate cache: TTL, fill ratio, per-destination coords and path MTU. |
| `show connections` | `show_connections` | Pending handshake connections: state, idle time, resend count. |
| `show transports` | `show_transports` | Transport instances: type, state, MTU, local address, per-transport stats. |
| `show transports` | `show_transports` | Transport instances: type, state, MTU, local address, per-transport stats, and — for interface-bound transports — interface presence (`absent` / `binding` / `present`), carrier, absence policy (`required` / `optional`), time in the current phase, and bind/failed-attempt counts. |
| `show routing` | `show_routing` | Routing summary: pending lookups, retry state, forwarding/discovery/error/congestion counters. |
| `show identity-cache` | `show_identity_cache` | Cached `(node_addr → npub)` entries with last-seen timestamps. |
| `show native-flows` | `show_native_flows` | Native datagram API: open and pending flows with their ports, queue depth and age, bound listeners with their backlog, and the `native` counters. |
+74
View File
@@ -566,6 +566,80 @@ root; on macOS it requires read/write access to a `/dev/bpf*` device.
| `auto_connect` | bool | `false` | Auto-connect to discovered peers |
| `accept_connections` | bool | `false` | Accept incoming connection attempts from discovered peers |
| `beacon_interval_secs` | u64 | `30` | Announcement beacon interval in seconds (minimum 10) |
| `optional` | bool | `false` | Whether absence of the interface is normal. See below |
**Dynamic binding.** The interface does not have to exist when the daemon
starts. A transport whose interface is missing comes up *absent*: it is not a
start failure, it is not skipped, and it binds on its own the moment the
interface appears — sub-second where the kernel offers link events (netlink on
Linux, `PF_ROUTE` on the BSDs), within a second otherwise. An interface that
goes away at runtime unbinds and rebinds by the same path, so a `wifi reload`
or an unplugged adapter needs no restart. Presence means `IFF_UP` — the
interface exists and is administratively up — and deliberately not
`IFF_RUNNING`: binding needs no carrier, and the socket keeps working across a
carrier flap without rebinding. A bridge with nothing plugged into it, such as
`br-lan` on a wifi-only router, is therefore bound and healthy rather than
permanently `Degraded`, and starts carrying traffic the moment a port comes up.
Whether an interface has carrier is reported separately, as `interface.carrier`
in `show_transports`.
`optional` selects how that absence is reported:
| | absence | log | retries |
| --- | --- | --- | --- |
| `optional: false` (default) | node reports `Degraded` | `info` at boot / `warn` on a runtime detach, then `error` once if it lasts past 10 s | forever |
| `optional: true` | no health impact | `info`, and nothing after | forever |
Naming an interface in configuration is a statement that you expect it, so the
default is to complain; silence is opted into. Set `optional: true` for
hardware that is legitimately not always there — a dock adapter, a radio only
some boards carry.
`optional` describes **the interface's presence, not the transport's
importance**. An optional interface that is present is used exactly as hard as
any other. No value of it makes a missing interface fatal at startup: the only
fatal case remains "no transports at all came up".
Either way the edge is logged once, on entering absence and on recovery —
never once per retry.
The edge itself is not an error. An interface missing when the daemon starts
and bound a moment later is the ordinary boot race this mechanism exists to
absorb, so it is `info`; a runtime detach is `warn`, because a link coming and
going is ordinary weather for a mesh daemon and a cable unplugged for two
seconds does not need a human.
Ten seconds is the whole grace. Past that it is no longer a race against a
radio or a container coming up, so a **required** interface still missing is
reported once at `error` and stays `Degraded` until it returns. Start-time
absence and a runtime detach share the one deadline — they are the same
transition throughout this mechanism. An `optional` interface never reaches
`error`; that is what `optional` means.
Said once, not repeated. How long the absence has lasted is a *state*, and it
is published as one: `interface.since_secs` in `show_transports`, and
`Degraded` for as long as it holds. Re-announcing it on a timer would put a
second, lossier copy of that in the log.
Node health does not wait for the ten seconds — `Degraded` is published on the
first edge, and that is the signal to watch.
If a binding keeps dying moments after it is established (a socket that errors
persistently while the interface stays up), the binder stops treating each
bind as a recovery: it backs off on the same 1 s → 30 s curve, holds node
health at its degraded reading, and stays quiet until a binding survives ten
seconds. Without that damping a broken socket produces a health flap and a log
pair every second, which is the same cry-wolf failure the edge-only logging
rule exists to prevent.
A bind failure that is **not** absence — no `CAP_NET_RAW`, no readable
`/dev/bpf*`, a buffer the kernel refused — is a fault, not a state, and fails
the daemon's start as it always has. Only a missing interface is waited out.
`fipsctl show transports` reports the current state per transport under
`interface`: `presence` (`absent` / `binding` / `present`), `carrier`,
`policy` (`required` / `optional`), `since_secs`, `binds`, and
`failed_attempts`.
**Named instances.** Multiple Ethernet interfaces can be configured by
using named sub-keys instead of flat parameters:
+1 -1
View File
@@ -124,7 +124,7 @@ table below lists every command currently registered.
| `show_mmp` | — | `peers[]` (link-layer per peer), `sessions[]` (session-layer per session). Each entry includes loss/RTT/ETX/goodput, smoothed values, trends. |
| `show_cache` | — | `count`, `max_entries`, `fill_ratio`, `default_ttl_ms`, `expired`, `avg_age_ms`, `entries[]` — per-destination coords, depth, age, last-used, optional `path_mtu`. |
| `show_connections` | — | `connections[]` — pending handshakes: `link_id`, `direction`, `handshake_state`, `started_at_ms`, `idle_ms`, `resend_count`, optional `expected_peer`. |
| `show_transports` | — | `transports[]` — `transport_id`, `type`, `state`, `mtu`, `name`, `local_addr`, optional `tor_mode`, `onion_address`, `tor_monitoring`, `stats`. |
| `show_transports` | — | `transports[]` — `transport_id`, `type`, `state`, `mtu`, `name`, `local_addr`, optional `tor_mode`, `onion_address`, `tor_monitoring`, `stats`, and `interface` for interface-bound transports. `interface` carries `name` (the configured netdev), `presence` (`absent` / `binding` / `present`), `carrier` (whether the link has `IFF_RUNNING` — reported only, never acted on: presence is `IFF_UP`, so a bound interface with no carrier is normal), `policy` (`required` / `optional`), `since_secs` (how long the current presence phase has been held), `binds` (successful binds since the transport was created — `1` after a clean start, more means it has rebound) and `failed_attempts` (failed binds since the last success). Absent entirely for transports that are not bound to a named interface, rather than reported as a permanently-`present` interface named `""`. Note that `state` describes the *transport* (`up` once started) and `interface.presence` describes the *socket*: an `up` transport whose interface is `absent` is started and waiting, which is a normal state and not a failure. |
| `show_routing` | — | `coord_cache_entries`, `identity_cache_entries`, `pending_lookups[]`, `pending_tun_destinations`, `pending_tun_packets`, `recent_requests`, `retries[]`, `forwarding`, `discovery` (request/response sub-counters; includes `req_deduplicated` — requests suppressed as recent duplicates — and `req_dedup_cache_full` — requests admitted because the dedup cache was full), `error_signals`, `congestion`. |
| `show_identity_cache` | — | `entries[]`, `count`, `max_entries`. Each entry: `node_addr`, `npub`, `display_name`, `ipv6_addr`, `last_seen_ms`, `age_ms`. |
| `show_native_flows` | — | `flows[]`, `listeners[]`, `stats` (the `native` counter family). Each flow: `flow_id`, `peer` (the peer's npub, which is its address; always present, because the flow carries the key its client named or its session authenticated), `peer_addr` (the 16-byte node address in hex — a truncated hash of the same key, kept because it is what `show_sessions` and `show_routing` key on), `local_port`, `remote_port`, `state` (`established` / `pending_accept`), `queued` (datagrams the node is holding for the flow), `age_ms` (time since the flow reached its current state: opened for a flow this node opened, accepted for one taken off a listener, announced for one still pending — accepting a pending flow restarts the clock). Each listener: `local_port`, `backlog`. |
+11
View File
@@ -223,6 +223,7 @@ is "all four flags on both ends."
> accept_connections: true
> dongle:
> interface: "enx00aabbccddee"
> optional: true
> announce: true
> # ...
> ```
@@ -231,6 +232,16 @@ is "all four flags on both ends."
> A single ground-up link only needs the flat form shown
> first; named instances become useful when the same node
> bridges multiple physical segments.
>
> `optional: true` on the dongle says its absence is normal —
> a USB adapter that is plugged in some days and not others.
> Without it, naming an interface is a statement that you
> expect it, and while it is missing the node reports
> `Degraded` and logs at `error`. Either way the interface
> does not have to exist when the daemon starts: a transport
> whose interface is missing waits and binds when it appears,
> and rebinds if it later goes away. Watch that with
> `fipsctl show transports`.
## Step 3: Grant the daemon permission to open raw sockets