An interface-bound transport is a long-lived object that is *sometimes bound*. The interface it names may not exist when the daemon starts, may appear minutes later, may vanish and return mid-operation, and may never appear at all. Until now the first observation was final: an interface missing at start was logged once and skipped for the life of the process, and one that disappeared at runtime published a health change but was never rebound. On OpenWrt that is a live bug. procd starts fips before wifi has created fips-mesh0 / fips-ap0, both transports are skipped and never retried, and the 802.11s peer link forms anyway (that is mac80211, not the daemon) — so the node looks healthy and reaches nothing. Give every interface-bound transport presence state and a binder task. start_async now returns Ok with the transport ABSENT rather than failing; the binder binds when the interface appears, tears the socket down when it goes away, and rebinds when it returns. Start-time absence and runtime detach are one transition and one code path. TransportError::InterfaceUnavailable makes absence branchable: a missing interface and a typo'd interface name were the same flat StartFailed (String), so nothing downstream could tell a state from a fault. A bind failure that is *not* absence — no CAP_NET_RAW, no readable /dev/bpf* — still fails the start, because retrying a socket that can never open behind a Degraded nobody is watching is worse than dying loudly at boot. **Presence is IFF_UP, not IFF_UP|IFF_RUNNING.** Carrier and bindability are different questions and only the second belongs in a bind gate. An AF_PACKET socket on a carrier-less bridge is valid and starts carrying traffic the instant a member port comes up, with no rebind. Gating on carrier would report a wifi-only router Degraded forever for an empty br-lan, turn every carrier flap the socket would have survived into an unbind/rebind cycle, and deadlock an 802.11s interface that reports RUNNING only once it has peered — peering needs beacons, beacons need a bound socket. Carrier is reported beside presence in show_transports instead. **A name is not a device.** Both backends bind by device: AF_PACKET stores sll_ifindex, a BPF descriptor follows the interface it was attached to. A netdev deleted and recreated under the same name leaves the socket attached to something gone while the name resolves perfectly well, and nothing notices — a stale AF_PACKET socket never becomes readable, so the receive loop neither errors nor exits, and send failures go to the caller rather than the binder. A listen-only node (announce: false, no beacon sender to fail) sat present and deaf indefinitely after a wifi reload. The bound index is captured at bind and compared on every poll. A name reappearing with a different MAC is different hardware: cached neighbors are dropped rather than resumed onto. **Detection is event-driven** where the kernel offers a source — netlink RTNLGRP_LINK on Linux, PF_ROUTE on macOS and FreeBSD — with a 1 s getifaddrs poll underneath as a backstop. Link-event payloads are not parsed: an event is a hint to re-run the probe, which is cheap and authoritative, and a parser's bugs would be presence bugs. Probes are coalesced to ten a second because PF_ROUTE has no group filter and delivers every routing message on the host; a persistently failing event source is logged once, backed off, and abandoned for the poll after five errors rather than spinning a core silently on ENOBUFS. **Degraded becomes a level, not a latch.** The supervisor's reason set was monotonic, which was correct while no child could recover; with recovery it would have come to mean "something broke at some point since boot" rather than "something is broken now". Absence lives in its own reversible set, health is recomputed on every transition in both directions, and an absent transport still counts as up so a single-interface node that boots before its wifi degrades rather than exiting on NoTransports. There is deliberately no restart action in the FSM: rebinding belongs next to the file descriptor, and a supervisor-authored retry would be a second mechanism racing the first for the same socket. **The node's egress MTU follows the bound set.** transport_mtu filters on is_bound(), not is_operational() — an interface-bound transport is operational from the moment it starts, so the weaker predicate let hardware that had never appeared clamp the whole node's IPv6 MTU. Because that minimum now moves at runtime, the TUN reader and writer read the TCP MSS ceiling from a shared atomic instead of a u16 captured at spawn; every other consumer (show_status, the snapshot, the session-layer fragmentation check) already read it live, so the clamp was the one place the daemon could report one effective MTU and enforce another. It moves in both directions: a narrow interface appearing tightens it, its departure releases it. MSS is negotiated per connection, so a change binds connections opened after it. **Logging is per edge, never per attempt**, with one deadline after it. The edge itself is not an error — info at boot, warn on a runtime detach, since an interface bound a moment later is the ordinary case this mechanism exists to absorb and crying error at t=0 then "recovered" at t=0.2s is the failure the rule exists to prevent. Ten seconds is the whole grace: past it, absence is no longer a race against a radio or a container, so a required interface still missing is reported once at error. Start-time absence and a runtime detach share that one deadline, as they share everything else here. Said once, not repeated: how long an absence has lasted is state, published as interface.since_secs and as Degraded, and a monitor can threshold it per deployment rather than the daemon compiling a schedule in. Successful rebinds are damped. Backoff covers failed binds; the nastier case is binds that keep succeeding into a socket that dies moments later, which a receive loop giving up on a persistent error while the interface stays UP produces once per second forever. Consecutive bindings dying inside ten seconds back off on the 1 s → 30 s curve, and past three the binder stops announcing each bind as a recovery until one lasts. The receive loop backs off and exits on a dead socket instead of spinning on Err with a warn per iteration, and the ad-hoc ENXIO socket reopen in the beacon sender is gone: both hand recovery to the presence machine, one mechanism for every cause rather than one hack per symptom. Beacons pause while absent because the task simply does not exist then. transports.ethernet.*.optional (default false) decides how loudly absence is reported. Naming an interface is a statement that you expect it, so the default complains: Degraded, and the error at the deadline. optional: true is silent — info on the edge, no health impact, no error — for hardware legitimately not always there. It describes the interface's presence, not the transport's importance, and no value of it makes a missing interface fatal at startup. show_transports grows an `interface` block — presence, carrier, policy, how long the phase has been held, bind count, failed attempts. The original boot race was expensive precisely because nothing an operator could see said the node was deaf. ci: exercise the presence probe against musl and an address-less interface Two legs, both about the assumption the presence probe rests on. OpenWrt — the platform dynamic interface binding exists for — is musl, and musl reimplements getifaddrs independently of glibc. `interface_present` reads ifa_flags out of it, and the interfaces this feature exists for (fips-mesh0, fips-ap0) are unbridged with no IP address at all, which is exactly where getifaddrs implementations differ. Every other leg is glibc, so without this one the probe was asserted on a libc no test had ever run it against, on the target it was written for. Building for the musl target rather than inside an Alpine container keeps the leg to a cross-build plus a run, without a second toolchain image to maintain. The address-less interface is created on both this leg and the glibc one. Pinning the contract on both libcs is what turns "glibc and musl agree about getifaddrs" from an assumption into a checked fact, and makes the glibc leg fail first if glibc is the one that changes. Loopback cannot stand in: it has 127.0.0.1, so probing it asks whether getifaddrs works rather than whether it reports this. Deliberately not tolerant of failure — a guard that quietly does not run is worse than no guard. The iface-binding suite is deliberately not added to the integration matrix here. Its files and its testing/ci-local.sh registration arrive in the next commit, and the GitHub workflow's matrix entry now goes with them, so both runners gain the suite together and testing/check-ci-parity.sh holds at every point in this history. docs: document dynamic interface binding The transport-layer design gains an Interface Presence section: the three deployment scenarios that motivated it, the presence machine and its two invariants, why presence is IFF_UP and not IFF_UP|IFF_RUNNING, interface identity and the recreated-netdev case, the detection sources and their two rate limits, the optional policy, health as a level rather than a latch, the egress-MTU consequence, the logging rules, and what the mechanism retires. It records why the error deadline is a single ten-second window rather than a repeating severity ladder, so the ladder does not come back. The control-socket and fipsctl references described show_transports without its `interface` block, and the transport-layer state machine still implied that Up meant bound. Both now say what Up and presence each describe and how they differ, since "transport up, interface absent" is a normal state an operator will meet and would otherwise read as a contradiction. The configuration reference documents `optional` with the absence, log and retry behaviour of each setting, and the ground-up tutorial's dongle example gains optional: true, which is what that example is actually for. fix(transport): pair every presence edge the binder publishes Six defects, all in the wiring around the presence machine rather than in the machine itself. The state machines were well tested; what went untested was how the binder feeds them, and every one of these lives there. **The first detach after a clean start never reached node health.** start_async binds inline and publishes that edge itself, before binder_loop exists. The loop then built a fresh ChurnGuard, which had therefore recorded no bind — and detached() takes `announced` to decide whether an edge is owed, so it asked for no retraction. A node that booted with its interface present and then lost it kept reporting Full for as long as the interface stayed away, while show_transports said absent and the log said detached. stabilized() could not repair it either: it early-returns while bound_at is None, which it is for a bind the loop did not perform, so the guard stayed unseeded until a second detach happened to fix it. The loop now seeds the guard from the binding it inherits. This is not the cable-unplug case. Presence is IFF_UP, so a carrier loss is correctly not a detach at all; it takes an admin down, a netdev delete, or a device removal — a wifi reload, a hostapd restart, a dongle pulled. **An optional interface published nothing, and the MTU floor rode the same channel.** Filtering the edge at the source conflated two questions that happen to share a transport: whether node health should move, and whether the bound set changed. Only the first is policy. The second determines the node's egress MTU floor, and refresh_tun_mss_ceiling has no other trigger — so an optional transport binding or unbinding at runtime left the TUN MSS clamp derived from a bound set that no longer existed, reporting one effective MTU and clamping to another. That is the defect the MssCeiling work was written to remove, reintroduced for exactly the transports the shipped OpenWrt config marks optional: five of seven. TransportPresence gains health_relevant; the edge always goes out and only health is filtered. **A permanent bind fault went unlogged if any absence race preceded it.** record_attempt ran ahead of the InterfaceUnavailable arm, so a probe that won a race the bind then lost burned the counter that gates the only error a real fault ever gets — emitted on attempts == 1. A flapping interface followed by CAP_NET_RAW being dropped or /dev/bpf* exhausting therefore reported nothing at all, for the life of the process, and the operator got report_sustained_absence's "still missing" instead: wrong, for an interface that is present. An absence race is not a bind attempt and no longer counts as one. **The watcher's give-up did not stick.** Abandoning the event source parked on pending() *inside the changed() future*, and that future is constructed fresh on every pass of the binder's select! and dropped whenever the poll ticker wins. So the next pass re-read the dead socket, re-counted the error and re-logged "not recoverable" — once per wake-up, forever, which at a 1 s tick across the shipped seven-transport config is seven warnings and seven failing syscalls a second on flash-backed logging. The 100 ms backoff lived in the dropped future too and never applied across passes. Abandonment is now state on the watcher. **A zero-length read livelocked the binder.** try_io clears readiness only on WouldBlock, which the sibling error arm handles by hand and this one did not — so breaking out left readable() instantly ready with nothing to read, and the loop never returned Pending. That starves the select! of its ticker entirely and takes presence detection down with it. Cleared and treated as a fault so the give-up path applies. **A failed presence probe read as an absent interface.** getifaddrs is a netlink dump and fails for reasons that have nothing to do with the interface: ENOBUFS under memory pressure, EMFILE or ENFILE under fd exhaustion, since it opens a socket of its own. Answering "not present" there tore down a working socket and degraded health over a transient syscall failure, undiagnosably, and under fd exhaustion the rebind could not have succeeded anyway. interface_has_flags now distinguishes the two; the detach gate holds its binding when the kernel will not answer, while the bind gate still treats it as absence and retries. **On macOS the reader thread could not die, so a dead socket read as live.** Any read() failure other than EBADF — ENXIO being the one that matters, which is what BPF answers once the interface it was attached to is torn away — reset the parse buffer and continued. The thread never returned, so the channel never closed, so recv_from never failed, so the tokio task never exited, so tasks_alive() reported a dead socket as a live one. Detach detection on macOS reduced to the name and the index, and an interface reset in place left the transport present and deaf. It now gives up after a streak, and the return closes the channel the binder is actually watching. Two smaller pairings while here: stop_async retracts the edge it would otherwise leave standing for a socket that is gone, and start_async hands a refused edge to the binder to retry rather than dropping the only edge either consumer will see until the interface next moves. The absence deadline is stamped at start rather than at construction, so a transport staged for longer than the bring-up window no longer reports sustained absence on its first tick having given the interface no window at all. feat(config): reject impossible ethernet interface names at load Config::validate never inspected transports.ethernet, so an empty interface name, a name past the kernel's 15-byte limit, and two transports naming the same netdev all loaded cleanly and failed only at runtime — the first two as a permanent absence indistinguishable from an interface that has not been created yet. That indistinguishability is deliberate and worth keeping: waiting is the right answer for an interface the operator has not made yet, and the daemon cannot know which of the two it is looking at. Which is exactly why the syntactic gate earns its place. A name that is *impossible* is the one case still separable from "not there yet", and without the check a typo costs a permanently Degraded node whose only symptom is an interface that never arrives — the failure mode the presence machine exists to make legible, reintroduced one level up. Syntax only. Whether a well-formed name exists stays the binder's question, asked once a second, forever. The duplicate check is a different fault: two transports on one netdev means two sockets on the same device at the same ethertype, each receiving every frame the other does. test(transport): close the vacuous and uncovered branches Six tests that asserted nothing, or asserted less than they claimed. `a_bind_fault_still_fails_the_start` returned early whenever the socket *could* be opened, so it was vacuous as root and on any developer machine with a group-readable /dev/bpf* — the fail-fast path it is named for went unchecked exactly where someone was most likely to run it. It now asserts in both halves, and the privileged half is worth more than the fix: a present, bindable interface binding inline is the ordinary case on a booted router, and no other unit test reaches it, because every other one here names an interface that does not exist. The bind-success path had no unit coverage at all. `an_interface_with_no_addresses_is_still_present` returned early when the fixture was absent — no fixture, pass. It still has to skip on a machine with no address-less interface, so the guard is the runner declaring that it has fixtures: CI now sets FIPS_TEST_REQUIRE_FIXTURES beside FIPS_TEST_ADDRLESS_IFACE, and the test fails rather than skips if the fixture step is ever removed or renamed. Its `let _ = interface_carrier(...)` is now asserted too: if presence and carrier ever collapsed into one read, a carrier-less bridge would report absent and the whole IFF_UP-not-IFF_RUNNING decision would be silently undone. `policy_labels` checked `Required.as_str()` and not `Optional.as_str()`, so a swapped pair would paint every expected interface as the tolerated kind and stay green. `Presence::as_str` had no test at all — `binding` was never observed by anything, anywhere. Three binder branches had no coverage: a transport restarting (the second `start_async` clearing the previous run's stop flag — only a second *stop* was tested, so a transport that could never restart passed everything), the episode clock being restamped at start rather than at construction, and the refused-edge retry actually delivering. The last one matters most: the existing test filled the channel and dropped the receiver, so a slot that captured an edge and never re-sent it would pass while health sat on a stale level forever. And the hardware-change boundary `record_bind` returns, which the neighbour flush hangs off: false on a first bind (or every clean start would drop a cache it had just built) and true once, not stickily, on a MAC change. Each new test was verified against the defect it guards — comment out the `shutdown.store(false)`, the `mark_starting()`, or the seeded `unpublished`, and the corresponding test goes red while the rest stay green. fix(test): correct the dummy-carrier assertion, and pin Darwin's presence probe The carrier assertion added ind6240698was wrong and would have failed both Linux legs. A `dummy` interface brought up reports `<BROADCAST,NOARP,UP, LOWER_UP>` — `IFF_RUNNING` is set, so it *has* carrier. Verified against the exact fixture CI builds, `addrgenmode none` and all, rather than against the comment: the original code discarded the result and its comment claimed "up but not running", which is what made asserting it look safe. So the fixture pins address-less *presence* and cannot demonstrate the presence-vs-carrier split at all — an interface up with no carrier is a bridge with nothing plugged in, which no fixture here creates. `carrier_is_reported_separately_from_presence` pins that split from the other side. The corrected assertion is Linux-only, because the expected answer is a property of the fixture device rather than of the code. That is also what lets the macOS fixture land. The Linux legs pin the address-less contract on glibc and musl, but the BSD-derived `getifaddrs` the macOS backend actually calls had no coverage — the test skipped itself silently on that runner, which is precisely the shape the previous commit was removing. `feth` is macOS's fake-Ethernet pseudo-interface and is created address-less; the step fails the leg rather than testing the wrong thing if the runner hands it an address anyway, mirroring why the Linux step needs `addrgenmode none`. Both branches of the fixture guard were exercised: unset skips and passes, and declared-but-missing fails loudly. The corrected test was run against a real Linux dummy inside a container, not reasoned about. fix(ci): put the macOS presence fixture on the macOS job The `feth` fixture step added in6b9faa0clanded on the Linux `test` job, not on `test-macos`, and has failed CI ever since: create: Host name lookup failure ifconfig: `--help' gives usage information. That is Linux net-tools `ifconfig`, which has no `create` subcommand. So the Linux job ran two address-less-fixture steps — its own correct `ip link add type dummy` one, then a macOS one that cannot work there — while `test-macos` had none at all, leaving the Darwin `getifaddrs` path exactly as uncovered as before. Cause was a pattern-anchored edit: `Install cargo-nextest` followed by `Run unit tests` appears in three jobs, and the insert hit the first match. Moving it then hit the *last* match, which is `test-windows`. It is now placed by job boundary rather than by pattern, and verified per job: `test` and `test-musl` carry the `ip link`/dummy fixture, `test-macos` carries `ifconfig`/`feth`, and `feth` appears exactly once in the file. The verification that missed this was counting steps in `test-macos` and reading 6 as confirmation. Six was the count *before* the insert; seven is what a successful insert looks like. Now asserted by job and by which tool each fixture step uses, so a step in the wrong place fails the check rather than matching a total. test(transport): cover the detach branch and the stop-race check Two of the three untested branches, by two different routes. **The detach decision is now a pure function.** `classify_detach(gone, replaced, dead)` replaces the inline three-way `if` in the binder loop, and all eight input combinations are asserted, plus the precedence between them and the `reason=` labels the integration suite and operators grep for. This does not make `Replaced` or `SocketDied` reachable from a test — both need a bind that succeeded and then a specific external event, and the integration suite cannot arrange the recreate deterministically either, for the reason recorded in reference/notes.md: the netlink event from a delete is acted on within microseconds, so `Gone` wins that race in practice. What it does is split the untested thing in two. The three inputs each already had tests (`interface_present`, `device_replaced`, `tasks_alive`); the branch between them did not, and that half is now total and exhaustive. Precedence is asserted rather than assumed: an interface that has gone away has also trivially been "replaced" and its socket is also dead, so the most specific true statement has to win, and callers pass `replaced` already masked by `!gone` — the function no longer depends on them having masked correctly. **The post-store shutdown check is now tested directly.** It needs a bind that *succeeds*, which is why `a_stop_racing_a_bind_leaves_nothing_behind` could never reach it: that test's interface does not exist, so `bind_and_spawn` refuses at the presence probe several steps earlier. Binding loopback as root reaches it, and the assertion is that a bind completing after a stop undoes its own socket, its own loops, and its own presence. Unprivileged runners skip it, but loudly: a runner declaring `FIPS_TEST_PRIVILEGED` and unable to open a raw socket fails instead of skipping, the same guard shape as `FIPS_TEST_REQUIRE_FIXTURES`. No CI leg sets that yet — unit tests run unprivileged — so today it exercises the branch only under a root container. Verified there, including the negative control: with the post-store check removed the test fails, and with it present it passes. test(transport): run the bind-success half in CI, and cover the replaced device Adds the privileged unit-test step the rest of this depends on, then uses it. **The privileged step.** Every unit-test leg runs unprivileged, so `PacketSocket::open` cannot succeed on any of them and everything past a successful bind runs nowhere in CI: the post-store shutdown check, the `Present` arm of the binder loop, `bind_now` itself. The step builds as the runner user and executes only the test binary under sudo — running `cargo` as root would use root's CARGO_HOME and discard the cache the job just restored. The binary-path extraction was verified locally before being written into the workflow. `FIPS_TEST_PRIVILEGED` is what makes it honest. Tests needing a raw socket skip quietly without it; with it set, a test that cannot open one fails and says so. A runner that stops granting the capability shows up as a red leg rather than as silence. **The replaced device.** `"interface replaced"` — a netdev recreated under the same name, the `wifi reload` case in #125 — could not be reached by the integration suite: with link events live the kernel's `RTM_DELLINK` is acted on within microseconds, so the binder observes "gone" first and takes the branch already covered. Attempting it there passed about one run in three. Rather than race the binder, the test asserts the *inputs*: after a real delete-and-recreate the name still resolves and the bound index no longer matches. Paired with the exhaustive classifier test, which pins that `(gone: false, replaced: true)` maps to `Replaced`, the path is covered without depending on scheduling. A watcher-disable hook was written for the racing approach and removed once this one made it unnecessary. **Two smaller gaps.** `report_sustained_absence` had no direct test — the integration suite infers it from counting ERROR lines, and ten seconds of real time is how a deadline ends up asserted by proxy. A test-only `backdate_for_test` ages the episode clock instead. And the probe floor was asserted only as a relation between two constants; it now has behaviour, via a `probe_delay` helper extracted from the loop. That extraction was not neutral, and its own test caught it: `checked_sub` yields `Some(0)` at exact equality where the loop used a strict `<`, so a zero-length sleep would have replaced no sleep at all. Harmless in effect, wrong in meaning, and fixed. Every new test was run against the defect it guards, in a root container: remove the post-store shutdown check, or make `device_replaced` always answer false, and the corresponding test fails. Mark publish_presence must_use. The defect this commit fixes was a discarded return value, so the fix is made self-guarding: a future call site that drops the pending edge instead of storing it now fails the lint rather than silently reintroducing the bug. Option is not must_use in std, unlike Result, so the attribute has to be explicit.
80 KiB
FIPS Configuration
FIPS uses YAML-based configuration with a cascading multi-file priority system. All parameters have sensible defaults; a node can run with no configuration file at all (it will generate an ephemeral identity and listen on default addresses).
Configuration Loading
Search Paths
When started without the -c flag, FIPS searches for fips.yaml in these
locations, lowest to highest priority:
| Priority | Path | Purpose |
|---|---|---|
| 1 (lowest) | /usr/local/etc/fips/fips.yaml (macOS, FreeBSD), /etc/fips/fips.yaml (other Unix) |
System-wide defaults |
| 2 | ~/.config/fips/fips.yaml |
User preferences |
| 3 | ~/.fips.yaml |
Legacy user config |
| 4 (highest) | ./fips.yaml |
Deployment-specific overrides |
All found files are loaded and merged in priority order. Values from higher
priority files override those from lower priority files. This allows a system
administrator to set site-wide defaults in the priority 1 path above,
/usr/local/etc/fips/fips.yaml on macOS and FreeBSD and
/etc/fips/fips.yaml on other Unix systems, while individual
deployments override specific values in ./fips.yaml.
On macOS and FreeBSD both directories are probed: /etc/fips first,
then /usr/local/etc/fips, so the packaged file wins over a leftover
/etc/fips copy from an earlier install.
CLI Option
fips -c /path/to/config.yaml
When -c is specified, only that file is loaded (search paths are skipped).
Partial Configuration
Every field has a built-in default. A configuration file only needs to specify values that differ from defaults. For example, a minimal config might contain only the identity and peer list, inheriting all other defaults.
YAML Structure
The configuration is organized into six top-level sections (gateway:
is Linux only):
node: # Node behavior, protocol parameters, and tuning
tun: # TUN virtual interface
dns: # DNS responder for .fips domain
transports: # Network transports (UDP, Ethernet, Bluetooth, Tor, ...)
peers: # Static peer list
gateway: # LAN gateway service (Linux only)
Control Socket (node.control.*)
| Parameter | Type | Default | Description |
|---|---|---|---|
node.control.enabled |
bool | true |
Enable the control socket |
node.control.socket_path |
string | (auto) | Unix: Socket file path. Resolution is shared by daemon and clients: /run/fips/control.sock when /run/fips exists; then /var/run/fips/control.sock on macOS/FreeBSD when its private directory exists; then $XDG_RUNTIME_DIR/fips/control.sock; finally /tmp/fips-control.sock. A privileged macOS daemon selects /var/run/fips/control.sock even when the private directory must be created after boot. Windows: TCP port number (default: 21210); the control socket listens on 127.0.0.1 at this port. |
The control socket provides access to node state and runtime management
via the fipsctl command-line tool. In addition to read-only status
queries, fipsctl connect and fipsctl disconnect enable runtime peer
management. See the fipsctl reference for the
command list.
On Unix, the control socket is a Unix domain socket with filesystem
permissions (mode 0770, group fips). On Windows, it is a TCP listener
on localhost. TCP does not provide filesystem-level ACLs, so any local
user can connect to the control port.
Security note (Windows): The TCP control socket on Windows is a known limitation. Any process running on the local machine can connect to the control port and issue commands, including
disconnect,connect, andinject-config. This is acceptable for single-user workstations but may be inappropriate for shared machines. Future improvements may include named pipe support (with Windows ACLs) or an authentication token mechanism. On shared Windows systems, consider using firewall rules to restrict access to the control port.
All tunable protocol parameters live under node.*, organized as sysctl-style
dotted paths. The top-level sections (tun, dns, transports, peers)
handle infrastructure concerns only.
Node Parameters (node.*)
Identity (node.identity.*)
| Parameter | Type | Default | Description |
|---|---|---|---|
node.identity.nsec |
string | (none) | Secret key in nsec (bech32) or hex format. If omitted, behavior depends on persistent. |
node.identity.persistent |
bool | false |
Persist identity across restarts via key file. |
Identity resolution follows a three-tier priority:
- Explicit
nsecin config — always used when present, regardless ofpersistent - Persistent key file — when
persistent: trueand nonsec, loads fromfips.keyadjacent to the config file; if no key file exists, generates a new keypair and saves it - Ephemeral — when
persistent: false(default) and nonsec, generates a fresh keypair on each start
Key files (fips.key with mode 0600, fips.pub with mode 0644) are written adjacent
to the highest-priority config file for operator visibility, even in ephemeral mode.
General
| Parameter | Type | Default | Description |
|---|---|---|---|
node.leaf_only |
bool | false |
Leaf-only mode: node does not forward traffic or participate in routing |
node.tick_interval_secs |
u64 | 1 |
Periodic maintenance tick interval (retry checks, timeout cleanup, tree refresh) |
node.base_rtt_ms |
u64 | 100 |
Initial RTT estimate for new links before measurements converge |
node.heartbeat_interval_secs |
u64 | 10 |
Heartbeat send interval per peer for liveness detection |
node.link_dead_timeout_secs |
u64 | 30 |
No-traffic timeout before a peer is declared dead and removed |
node.drain_timeout_secs |
u64 | 2 |
Upper bound in seconds on the Draining shutdown phase. On shutdown the node broadcasts Disconnect to its peers and then waits up to this long for the links to clear, exiting as soon as the last peer is gone. 0 skips the wait. The key is absent from a default config file rather than written with its default value, so an unset key and the 2-second default are the same thing |
node.log_level |
string | "info" |
Tracing filter default. Case-insensitive; one of trace, debug, info, warn, error. Overridden by the RUST_LOG environment variable when set |
Resource Limits (node.limits.*)
Controls capacity for connections, peers, and links.
| Parameter | Type | Default | Description |
|---|---|---|---|
node.limits.max_connections |
usize | 256 |
Max handshake-phase connections |
node.limits.max_peers |
usize | 128 |
Max authenticated peers |
node.limits.max_links |
usize | 256 |
Max active links |
node.limits.max_pending_inbound |
usize | 1000 |
Max pending inbound handshakes |
Rate Limiting (node.rate_limit.*)
Handshake rate limiting protects against DoS on the Noise IK handshake path.
| Parameter | Type | Default | Description |
|---|---|---|---|
node.rate_limit.handshake_burst |
u32 | 100 |
Token bucket burst capacity |
node.rate_limit.handshake_rate |
f64 | 10.0 |
Tokens per second refill rate |
node.rate_limit.handshake_timeout_secs |
u64 | 30 |
Stale handshake cleanup timeout |
node.rate_limit.handshake_resend_interval_ms |
u64 | 1000 |
Initial handshake message resend interval |
node.rate_limit.handshake_resend_backoff |
f64 | 2.0 |
Resend backoff multiplier (1s, 2s, 4s, 8s, 16s with defaults) |
node.rate_limit.handshake_max_resends |
u32 | 5 |
Max resends per handshake attempt |
node.rate_limit.established_handshake_burst |
u32 | derived | Burst capacity of the established-link bucket. Derived default is node.limits.max_peers (128) |
node.rate_limit.established_handshake_rate |
f64 | derived | Refill rate of that bucket. Derived default is (max_peers / max(node.rekey.after_secs, 1)) * (1 + handshake_max_resends), floored at 1.0/s — 6.4/s at shipped defaults |
node.rate_limit.session_setup_burst |
u32 | 64 |
Per-link-peer burst for inbound session-setup messages that would open a new session |
node.rate_limit.session_setup_rate |
f64 | 16.0 |
Per-link-peer refill rate for those messages, in tokens per second |
Msg1 whose source matches an established link (rekey and restart
maintenance traffic) draws on a second bucket rather than competing with
stranger admission. Both keys are optional; leaving them unset keeps the
derived sizing, which tracks max_peers and the rekey period
automatically instead of becoming a constant nobody revisits.
max_peers: 0 (unlimited) has no peer-count-derived size, so the
derivation falls back to handshake_burst / handshake_rate.
The node's total admitted msg1 rate is the sum of the two buckets: 228 burst and 16.4/s at shipped defaults, of which the established half is reachable only by a source that already matches a live link. Size against the sum when budgeting handshake crypto load for a host.
The session_setup_* pair is a separate limiter on the session layer, not
the link layer. It is keyed on the authenticated link peer a session datagram
arrived over, so each neighbour gets its own budget and a flood is
attributable. Setup messages naming a peer this node is already established
with (inbound rekey and restart traffic) draw on a second per-link bucket
derived from max_peers, node.rekey.after_secs and handshake_max_resends,
exactly as established_handshake_* is, so a stranger flood cannot suppress
rekey traffic sharing the link.
At the defaults one neighbour can force at most
session_setup_rate * handshake_timeout_secs half-open entries (480) and
session_setup_rate * (1 + handshake_max_resends) acks per second (96). The
node-wide ceiling is still that times the peer count, since the limiter bounds
each neighbour rather than the aggregate. A legitimate peer whose traffic
reaches this node over the same link as an attacker's shares that
attacker's stranger bucket, so establishment behind a flooded neighbour is
refused until the bucket refills; the initiator's own resend schedule (1s, 2s,
4s, 8s, 16s) covers a short drain.
Retry / Backoff (node.retry.*)
Connection retry with exponential backoff.
| Parameter | Type | Default | Description |
|---|---|---|---|
node.retry.max_retries |
u32 | 5 |
Max connection retry attempts |
node.retry.base_interval_secs |
u64 | 5 |
Base backoff interval |
node.retry.max_backoff_secs |
u64 | 300 |
Cap on exponential backoff (5 minutes) |
Auto-reconnect (triggered by MMP link-dead removal) uses the same backoff
parameters but bypasses max_retries, retrying indefinitely. See
peers[].auto_reconnect below.
Medium-Change Detection (node.netmon.*)
Detects that the host moved between transport media — WLAN to LAN, WLAN to 5G, an interface arriving or leaving — and rebinds the send path immediately.
| Parameter | Type | Default | Description |
|---|---|---|---|
node.netmon.enabled |
bool | true |
Whether medium-change detection runs |
node.netmon.poll_interval_secs |
u64 | 5 |
How often the path to each peer is sampled (backstop period where an event-driven backend exists). Must be at least 1 while enabled; 0 is refused at startup. |
node.netmon.debounce_ms |
u64 | 250 |
How long to wait for the picture to settle before acting (0 disables). Refused at startup when debounce_ms × 8 reaches node.link_dead_timeout_secs — see below. |
Both refusals stop the node rather than degrade it, so they are worth knowing before they are met.
A handover is ridden out for up to 8 settling rounds (MAX_DEBOUNCE_ROUNDS)
of debounce_ms each before a change is reported. If that worst case reaches
node.link_dead_timeout_secs, the liveness reaper tears the peering down before
the detector ever reports, so the machinery runs and cannot help — the node
refuses to start rather than run in that shape. At the shipped defaults the
margin is wide (8 × 250 ms = 2 s against 30 s), but the constraint couples two
keys in different blocks: shortening link_dead_timeout_secs for fast
failover can make an untouched debounce_ms illegal. The refusal names both
values and the multiplier.
Established UDP peers use a per-peer connect()-ed socket for the send fast
path. connect(2) makes the kernel resolve the route once and pin the local
source address to whichever interface carried it then; it never re-evaluates.
Without detection, a medium change therefore leaves every peer transmitting
from an abandoned address while the peer answers where it last heard the node —
the peering reports itself connected and carries nothing until
link_dead_timeout_secs tears it down, typically 60–90s per switch.
On a detected change the node drops the sockets of the peers the change names (the wildcard listen socket resolves a route per packet, so sends keep working, and a correctly bound connected socket is reinstalled on a later tick) and heartbeats those of them on a connectionless transport at once so the far side re-pins to the new source address. A peer reached over TCP, Tor, Nym or BLE is left to its periodic heartbeat, since sending to it here would block the node's receive loop on a stream the medium change has very likely just stranded. A peer the change does not name is left alone entirely. No peering is torn down: sessions, tree positions and routes survive the switch.
What counts as a change. For each peer whose transport address is a numeric
IP endpoint, the node asks the kernel which local address it would use to reach
that peer — a connect(2) on a UDP socket, which resolves the route and sends
nothing. A change is reported when a peer present in two consecutive samples is
now reached from a different local address, or has stopped being reachable at
all.
Because the question is asked per peer, an interface the node does not peer over
cannot trigger anything: a container bridge, a VPN, a veth pair or a tunnel
appearing is not the route to any peer, so it does not enter the sample. A peer
on the same LAN, reached by its subnet route rather than the default route, is
covered as well as one across the internet, and so is a more specific route
moving under a single peer.
The converse is the residual. The probe answers for the peer's current address, and that address is the source of the last authentic packet it sent, so a peer that roams between two of its own addresses which leave this host by different interfaces is indistinguishable from a local path move. The reaction is scoped to the peers named in the change, so such a peer moves nothing but its own send path — but it is the peer, not this host, that decided the fingerprint changed.
Peers appearing and leaving are ignored on their own — that is ordinary node
behaviour and says nothing about the medium. A peer seen for the first time is
the one exception, and it is not judged against history but against its own
send path: if its connect()-ed socket is pinned to a source the routing table
would no longer choose, it is reported. Without that, a medium change in the
window between a peer authenticating and the detector's next sample would be
the detector's first sight of that peer, and would be adopted silently while
the peer's socket stayed pinned to the path the host had just left. A peer
joining onto a path that has not moved has its socket pinned exactly where its
traffic goes, so it still reports nothing. A peer whose address is not a
probeable IP endpoint contributes nothing: a MAC on Ethernet or BLE, a .onion
or Nym recipient reached through a local proxy, an IPv6 literal with a scope
suffix, or a peer still carrying the hostname it was configured with (resolving
one would put a DNS lookup on the sample path; the address becomes numeric as
soon as an authenticated packet arrives from the peer). A node holding no peers
detects nothing, which is correct — it has nothing bound to the old path.
The cost is five non-blocking syscalls per peer per sample, read from the
probe's own code rather than measured: socket(2) and bind(2), a connect(2)
that sends no packet, a getsockname(2), and the close(2) the socket takes on
drop. Nothing goes on the wire and no name is resolved.
node.limits.max_peers bounds the per-sample total only where it is set: at
max_peers: 0, which means unlimited, there is no bound and the cost tracks the
live peer count instead.
Detection uses the best backend the platform has:
| Platform | Backend | Latency |
|---|---|---|
| Linux, Android | NETLINK_ROUTE multicast (as ip monitor) |
kernel event, milliseconds |
| macOS, FreeBSD | PF_ROUTE socket |
kernel event, milliseconds |
| Windows, iOS | timer | up to poll_interval_secs |
Where a backend exists, poll_interval_secs is only a backstop: a netlink
socket drops messages under memory pressure and the subscription can fail to
start in a restricted sandbox, so the timer keeps running underneath. A backend
that cannot start is logged once at warn and the node falls back to the timer.
None of this covers Bluetooth: a BLE adapter's state is not an IP attachment and is invisible to this detector. The connected-socket fast path is Linux and macOS only; elsewhere there are no pinned sockets to rebind, and the heartbeat alone carries the new address.
Cache Parameters (node.cache.*)
Controls caching of tree coordinates and identity mappings.
| Parameter | Type | Default | Description |
|---|---|---|---|
node.cache.coord_size |
usize | 50000 |
Max entries in coordinate cache |
node.cache.coord_ttl_secs |
u64 | 300 |
Coordinate cache entry TTL (5 minutes) |
node.cache.identity_size |
usize | 10000 |
Max entries in identity cache (LRU, no TTL) |
Mesh Lookup (node.lookup.*)
Controls bloom-guided mesh lookup (LookupRequest/LookupResponse): finding the current coordinates of a mesh address the node already knows.
Renamed in v0.5.0. These six keys were
node.discovery.*. See Deprecated keys for the full mapping. A deployednode.discovery:block still loads and still applies, with a one-time deprecation warning at startup.
| Parameter | Type | Default | Description |
|---|---|---|---|
node.lookup.ttl |
u8 | 64 |
Hop limit for LookupRequest forwarding |
node.lookup.attempt_timeouts_secs |
array<u64> | [1, 2, 4, 8] |
Per-attempt timeouts. Each entry is the deadline for one LookupRequest before sending the next attempt with a fresh request_id. Length determines total attempt count; default gives 4 attempts and a 15s total budget |
node.lookup.recent_expiry_secs |
u64 | 10 |
Dedup cache expiry for recent request IDs |
node.lookup.backoff_base_secs |
u64 | 0 |
Optional post-failure suppression base in seconds; doubles per consecutive failure. 0 disables (default); the per-attempt sequence is the only retry pacing |
node.lookup.backoff_max_secs |
u64 | 0 |
Cap on optional post-failure backoff |
node.lookup.forward_min_interval_secs |
u64 | 2 |
Transit-side rate limiting: minimum interval between forwarded lookups for the same target |
Peer Rendezvous (node.rendezvous.*)
How the node finds peers to connect to at all, over the Nostr overlay and on the local link. Distinct from mesh lookup above, which resolves coordinates for a mesh address that is already known.
Renamed in v0.5.0.
node.discovery.nostr.*is nownode.rendezvous.nostr.*, andnode.discovery.lan.*is nownode.rendezvous.lan.*. See Deprecated keys.
Nostr Rendezvous (node.rendezvous.nostr.*)
Optional Nostr-mediated overlay rendezvous. This layer publishes replaceable
endpoint adverts (fips-overlay-v1), consumes advert-derived endpoint
fallbacks for configured peers, and can optionally discover non-configured
peers (policy: open). udp:nat remains the trigger for NAT traversal
offer/answer + punch-through, after which the established UDP socket is handed
into the normal FIPS transport/session stack.
Inbox-relay discovery falls back to the local DM relay list if remote relay
metadata cannot be fetched.
The Nostr discovery runtime is compiled into every build of the crate; it
is enabled at runtime via node.rendezvous.nostr.enabled: true and stays
inert otherwise.
| Parameter | Type | Default | Description |
|---|---|---|---|
node.rendezvous.nostr.enabled |
bool | false |
Enable Nostr-mediated overlay discovery |
node.rendezvous.nostr.policy |
string | "configured_only" |
Advert discovery policy: disabled, configured_only, open |
node.rendezvous.nostr.open_discovery_max_pending |
usize | 64 |
Max open-discovery peers queued in outbound retry/connection state at once |
node.rendezvous.nostr.max_concurrent_incoming_offers |
usize | 16 |
Max concurrent inbound traversal offers processed at once (rate limit against offer spam) |
node.rendezvous.nostr.max_concurrent_offers_per_npub |
usize | 4 |
Max concurrent inbound traversal offers accepted from any one sender npub, so a single identity cannot hold the whole pool. Sits inside max_concurrent_incoming_offers, which stays the outer bound; a larger value is inert. Zero is rejected, since it refuses every inbound offer rather than disabling the limit |
node.rendezvous.nostr.advert_cache_max_entries |
usize | 2048 |
Max cached overlay adverts retained from relay traffic |
node.rendezvous.nostr.seen_sessions_max_entries |
usize | 2048 |
Max seen-session IDs retained for replay detection |
node.rendezvous.nostr.advertise |
bool | true |
Publish local endpoint adverts |
node.rendezvous.nostr.advert_relays |
list[string] | ["wss://relay.damus.io", "wss://nos.lol", "wss://offchain.pub"] |
Relays used for service adverts |
node.rendezvous.nostr.dm_relays |
list[string] | ["wss://relay.damus.io", "wss://nos.lol", "wss://offchain.pub"] |
Relays used for encrypted signaling events |
node.rendezvous.nostr.stun_servers |
list[string] | ["stun:stun.l.google.com:19302", "stun:stun.cloudflare.com:3478", "stun:global.stun.twilio.com:3478"] |
STUN servers used for local reflexive address discovery |
node.rendezvous.nostr.share_local_candidates |
bool | false |
Whether to advertise local (RFC 1918 / ULA) interface addresses as host candidates in the traversal offer. Off by default: in most deployments peers aren't on the same broadcast domain, and sharing private host candidates causes misleading punch successes when an asymmetric L3 path (VPN, Tailscale subnet route, overlapping address space) makes a peer's private IP one-way reachable. Enable only when peers are on the same physical LAN |
node.rendezvous.nostr.app |
string | "fips-overlay-v1" |
Traversal application namespace, published in the advert's protocol tag (the d tag itself is hardcoded to fips-overlay-v1) |
node.rendezvous.nostr.signal_ttl_secs |
u64 | 120 |
Signaling TTL in seconds |
node.rendezvous.nostr.attempt_timeout_secs |
u64 | 10 |
Overall traversal attempt timeout in seconds |
node.rendezvous.nostr.replay_window_secs |
u64 | 300 |
Replay tracking retention window in seconds |
node.rendezvous.nostr.punch_start_delay_ms |
u64 | 2000 |
Delay before punch traffic starts |
node.rendezvous.nostr.punch_interval_ms |
u64 | 200 |
Interval between punch packets |
node.rendezvous.nostr.punch_duration_ms |
u64 | 10000 |
How long to keep punching before failure |
node.rendezvous.nostr.advert_ttl_secs |
u64 | 3600 |
Advert TTL in seconds |
node.rendezvous.nostr.advert_refresh_secs |
u64 | 1800 |
How often adverts are refreshed in seconds |
node.rendezvous.nostr.startup_sweep_delay_secs |
u64 | 5 |
Settle delay after Nostr discovery starts before the one-shot startup advert sweep runs (only used under policy: open). Allows the relay subscription backlog to populate the in-memory advert cache before the sweep fires |
node.rendezvous.nostr.startup_sweep_max_age_secs |
u64 | 3600 |
Maximum advert age (now - created_at) considered by the one-shot startup sweep (only used under policy: open). Adverts older than this are skipped on startup; the per-tick sweep still considers them up to valid_until_ms |
node.rendezvous.nostr.failure_streak_threshold |
u32 | 5 |
Consecutive NAT-traversal failures against a peer before an extended cooldown is applied. At this threshold the daemon also actively re-fetches the peer's advert from advert_relays to evict cache entries for peers that have gone away |
node.rendezvous.nostr.extended_cooldown_secs |
u64 | 1800 |
Cooldown applied to a peer once failure_streak_threshold is hit. Suppresses both open-discovery sweep enqueues and per-attempt retry firings until elapsed (30 minutes default) |
node.rendezvous.nostr.warn_log_interval_secs |
u64 | 300 |
Minimum interval between NAT traversal failed WARN log lines for the same peer. Subsequent failures inside the window log at DEBUG to reduce log spam on public-test nodes with many cache-learned peers |
node.rendezvous.nostr.failure_state_max_entries |
usize | 4096 |
Maximum entries retained in the per-npub failure-state map. Bounds memory under high cache turnover; oldest entries (by last failure time) are evicted when the cap is exceeded |
node.rendezvous.nostr.protocol_mismatch_cooldown_secs |
u64 | 86400 |
Cooldown applied after observing a fatal protocol mismatch on a Nostr-adopted bootstrap transport (e.g. Unknown FMP version from a peer running a different FMP-protocol version). Independent of extended_cooldown_secs and much longer (24 hours default) because the mismatch is structural; re-traversing is wasted effort until one side upgrades |
If stun_servers is omitted, the built-in default list above is used. If it is
specified in YAML, the configured list fully overrides the defaults.
Initiators use only this local list for outbound STUN queries; peer-advertised
STUN values are published for diagnostics/interoperability but are not used as
arbitrary egress targets.
The built-in advert and DM relay defaults point at widely-operated public
relays (Damus, nos.lol, Primal) as best-effort endpoints; operators are
encouraged to override them with their own relay preferences for production
deployments.
Advert freshness is enforced semantically: events with expired NIP-40
expiration tags are dropped, and adverts are also bounded by a created-at
staleness window derived from advert_ttl_secs (with a grace multiplier).
The current in-tree STUN parser handles IPv4 and IPv6 mapped-address
attributes. Local traversal candidates include active non-loopback private
interface addresses (RFC1918 IPv4 and IPv6 ULA) plus probed local egress
addresses for the punch socket port.
During punching, compatible private-subnet candidates and reflexive candidates
are attempted in parallel; the first successful path wins.
LAN Rendezvous (node.rendezvous.lan.*)
Peer rendezvous on the local link via mDNS / DNS-SD (RFC 6762 / RFC
6763). When enabled, the node publishes a _fips._udp.local. service
advert carrying its npub (and optional scope) and concurrently
browses for the same service type to learn same-broadcast-domain peers.
The result is sub-second peer pairing with no Nostr-relay roundtrip,
STUN observation, or NAT traversal: the observed endpoint is by
construction routable from the consumer's LAN.
mDNS adverts are unauthenticated, so a LAN advert is treated only as a routing hint. Identity is still proven end-to-end by the Noise IK handshake the node initiates against the observed endpoint; a spoofed advert carrying another peer's npub fails the handshake and is dropped. LAN discovery requires an active UDP transport (peers dial the advertised UDP port to begin the handshake).
| Parameter | Type | Default | Description |
|---|---|---|---|
node.rendezvous.lan.enabled |
bool | false |
Master switch. Opt-in: enable for sub-second same-LAN pairing. Default-off avoids reintroducing a per-LAN identity broadcast on nodes that have deliberately disabled other discovery channels |
node.rendezvous.lan.service_type |
string | "_fips._udp.local." |
DNS-SD service type. Primarily an override for integration tests running multiple isolated services on one loopback interface; leave at the default in production |
node.rendezvous.lan.scope |
string | (none) | Optional application/network scope carried in a scope=<name> TXT entry. Browsers with a scope set only surface peers advertising the same scope, so nodes on the same physical LAN configured for different mesh networks do not cross-feed. Intentionally separate from node.rendezvous.nostr.app so relay-visible adverts can stay generic while LAN discovery is isolated per private network |
Spanning Tree (node.tree.*)
Controls tree construction and parent selection.
| Parameter | Type | Default | Description |
|---|---|---|---|
node.tree.announce_min_interval_ms |
u64 | 500 |
Per-peer TreeAnnounce rate limit |
node.tree.parent_hysteresis |
f64 | 0.2 |
Cost improvement fraction required for same-root parent switch (0.0–1.0) |
node.tree.hold_down_secs |
u64 | 30 |
Suppress non-mandatory re-evaluation after parent switch |
node.tree.reeval_interval_secs |
u64 | 60 |
Periodic cost-based parent re-evaluation interval (0 = disabled) |
node.tree.flap_threshold |
u32 | 4 |
Parent switches in window before dampening engages |
node.tree.flap_window_secs |
u64 | 60 |
Sliding window for counting parent switches |
node.tree.flap_dampening_secs |
u64 | 120 |
Extended hold-down duration when flap threshold exceeded |
Bloom Filter (node.bloom.*)
| Parameter | Type | Default | Description |
|---|---|---|---|
node.bloom.update_debounce_ms |
u64 | 500 |
Debounce interval for filter update propagation |
node.bloom.max_inbound_fpr |
f64 | 0.20 |
Antipoison cap: reject inbound FilterAnnounce frames whose advertised false-positive rate exceeds this value. Valid range (0.0, 1.0). The default 0.20 corresponds to fill 0.7248 at k=5 (≈2,114 entries on the 1 KB filter); a saturated/poisoned filter is still ~100% FPR and rejected |
Bloom filter size (1 KB), hash count (5), and size classes are protocol constants and not configurable.
ECN Signaling (node.ecn.*)
Controls hop-by-hop ECN (Explicit Congestion Notification) signaling. When enabled, transit nodes detect congestion on outgoing links (via MMP loss/ETX metrics or kernel buffer drops) and set the CE flag on forwarded FMP frames. Destination nodes mark ECN-capable IPv6 packets with CE before TUN delivery per RFC 3168, enabling end-host TCP congestion control to react.
| Parameter | Type | Default | Description |
|---|---|---|---|
node.ecn.enabled |
bool | true |
Enable ECN congestion signaling (CE flag relay and local congestion detection) |
node.ecn.loss_threshold |
f64 | 0.05 |
MMP loss rate threshold for CE marking (0.0–1.0). When the outgoing link's loss rate meets or exceeds this value, forwarded packets are CE-marked. |
node.ecn.etx_threshold |
f64 | 3.0 |
MMP ETX threshold for CE marking (≥1.0). When the outgoing link's ETX meets or exceeds this value, forwarded packets are CE-marked. |
Congestion detection triggers on any of: outgoing link loss ≥ loss_threshold,
outgoing link ETX ≥ etx_threshold, or kernel receive buffer drops detected on
any local transport. CE is relayed hop-by-hop: once set on any hop, the flag
stays set for all subsequent hops to the destination.
Rekey (node.rekey.*)
Controls periodic Noise rekey for forward secrecy. When enabled, both FMP (link-layer IK) and FSP (session-layer XK) sessions perform fresh Diffie-Hellman key exchanges after a time or message count threshold, whichever comes first. A 10-second drain window keeps the old session active for decryption during cutover.
| Parameter | Type | Default | Description |
|---|---|---|---|
node.rekey.enabled |
bool | true |
Initiate periodic Noise rekey on links and sessions. A peer-driven session rekey is still answered when this is off, so session keys can still rotate |
node.rekey.after_secs |
u64 | 120 |
Initiate rekey after this many seconds on a session |
node.rekey.after_messages |
u64 | 65536 |
Initiate rekey after this many messages sent on a session |
Session / Data Plane (node.session.*)
Controls end-to-end session behavior and packet queuing.
| Parameter | Type | Default | Description |
|---|---|---|---|
node.session.default_ttl |
u8 | 64 |
Default SessionDatagram TTL |
node.session.pending_packets_per_dest |
usize | 16 |
Queue depth per destination during session establishment |
node.session.pending_max_destinations |
usize | 256 |
Max destinations with pending packets |
node.session.idle_timeout_secs |
u64 | 90 |
Idle session timeout; established sessions with no application data for this duration are removed. MMP reports (SenderReport, ReceiverReport, PathMtuNotification) do not count as activity |
node.session.coords_warmup_packets |
u8 | 5 |
Number of initial data packets per session that include the CP flag for transit cache warmup; also the reset count on CoordsRequired/PathBroken receipt |
node.session.coords_response_interval_ms |
u64 | 2000 |
Minimum interval (ms) between standalone CoordsWarmup responses to CoordsRequired/PathBroken signals per destination |
The anti-replay window size (2048 packets) is a compile-time constant and not configurable.
Link-Layer MMP (node.mmp.*)
Metrics Measurement Protocol for per-peer link measurement. See ../design/fips-mesh-layer.md for behavioral details.
| Parameter | Type | Default | Description |
|---|---|---|---|
node.mmp.mode |
string | "full" |
Operating mode: full (sender + receiver reports), lightweight (receiver reports only), or minimal (spin bit + CE echo only, no reports) |
node.mmp.log_interval_secs |
u64 | 30 |
Periodic operator log interval for link metrics |
node.mmp.owd_window_size |
usize | 32 |
One-way delay trend ring buffer size |
Session-Layer MMP (node.session_mmp.*)
Metrics Measurement Protocol for end-to-end session measurement. Configured independently from link-layer MMP because session reports are routed through every transit link, consuming bandwidth proportional to path length.
| Parameter | Type | Default | Description |
|---|---|---|---|
node.session_mmp.mode |
string | "full" |
Operating mode: full, lightweight, or minimal |
node.session_mmp.log_interval_secs |
u64 | 30 |
Periodic operator log interval for session metrics |
node.session_mmp.owd_window_size |
usize | 32 |
One-way delay trend ring buffer size |
Internal Buffers (node.buffers.*)
Channel sizes affecting throughput and memory. Primarily useful for performance tuning under high load or on memory-constrained devices.
| Parameter | Type | Default | Description |
|---|---|---|---|
node.buffers.packet_channel |
usize | 1024 |
Transport to Node packet channel capacity |
node.buffers.tun_channel |
usize | 1024 |
TUN to Node outbound channel capacity |
node.buffers.dns_channel |
usize | 64 |
DNS to Node identity channel capacity |
Native Datagram API (node.native_api.*)
Experimental, off by default, and built on Linux, FreeBSD and macOS only. A client process connects to a Unix socket and asks either to open a flow to a remote pubkey or to hold a local port. Both answers carry a file descriptor: a flow's, which the client sends and receives datagrams on, or a listener's, which arriving flows are delivered on. There is no IPv6 emulation and no TUN device on this path.
The surface is not stable, is not a reliability layer, and is not the v2 external process API. No compatibility promise is made about it: the keys below, the line protocol behind them, and the Rust client that hides it may change or be withdrawn in any release.
The listener is not built on macOS or Windows, and this section is ignored
there. Two separate things bound that: Windows has no SCM_RIGHTS and so no
way to hand a file descriptor to another process at all, while macOS has
SCM_RIGHTS but does not implement SOCK_SEQPACKET for AF_UNIX.
| Parameter | Type | Default | Description |
|---|---|---|---|
node.native_api.enabled |
bool | false |
Enable the native API socket |
node.native_api.socket_path |
string | (auto) | Socket file path. Resolved the same way as the control socket, with the filename api.sock: /run/fips/api.sock when /run/fips exists; then /var/run/fips/api.sock on FreeBSD when its private directory exists; then $XDG_RUNTIME_DIR/fips/api.sock; finally /tmp/fips-api.sock |
node.native_api.pending_per_flow |
usize | 16 |
Datagrams held for one flow while it waits to be accepted, or while an established flow's client is slow to read. Refused above 64 at startup: the whole batch is written onto a socket pair the client cannot read yet. Refused below 1: a flow that can hold nothing loses its peer's opening datagram between the arrival being announced and the client taking the flow |
node.native_api.backlog |
usize | 16 |
Flows announced on one listener and not yet taken by its task. Refused below 1 at startup: a listener with no backlog admits no flow, so every arrival would be dropped |
node.native_api.max_flows |
usize | 256 |
Flows this node holds at once |
node.native_api.debug_commands |
bool | false |
Answer the inject, stats and arrive debug commands. Not a supported interface |
Security note: the socket is mode
0770, owned by groupfips, and that is the whole of the authorization model. Any user in thefipsgroup can impersonate this node on the mesh. A process that can open the socket can send datagrams under this node's identity to any peer it names, and can hold a port and receive mesh traffic addressed to this node on it. There is no per-client authentication, no capability check and no audit trail beyond the daemon's own logs. On a node with the native API enabled, treatfipsgroup membership exactly as you would treat the node's private key. This is why the API is disabled by default, and why enabling it is an explicit operator decision rather than something a package turns on. See security.md.
A client may hold ports 1024 through 65535. Ports 0 through 255 are reserved for protocol use and 256 through 1023 for FIPS standard services (the IPv6 shim among them), and the daemon refuses both ranges by name. A client that names no local port is given one from 49152 upward.
debug_commands is a separate gate on three commands that exist only so the
test harness can drive the receive and dispatch paths without a wire.
inject makes the daemon write bytes the client chose into one of that
client's own flows, and arrive makes it dispatch a datagram as though a peer
had sent it, reaching any listener this node holds. A node with the key off
refuses each by name, so a client can tell "this node will not do that" from
"this build has no such command". Leave it off outside the test harness.
fipsctl show native-flows reports the open and pending flows, the bound
listeners and the native counters; see the fipsctl
reference and
control-socket.md. A Rust program links
the crate and speaks the API through the fips::native::client module, which
hides the line protocol; see
../how-to/use-the-native-datagram-api.md.
TUN Interface (tun.*)
| Parameter | Type | Default | Description |
|---|---|---|---|
tun.enabled |
bool | false |
Enable TUN virtual interface |
tun.name |
string | "fips0" |
Interface name |
tun.mtu |
u16 | 1280 |
Interface MTU (IPv6 minimum) |
DNS Responder (dns.*)
Resolves <npub>.fips queries to FIPS IPv6 addresses. Resolution is pure
computation (npub to public key to address); resolved identities are registered
with the node for routing.
| Parameter | Type | Default | Description |
|---|---|---|---|
dns.enabled |
bool | true |
Enable DNS responder |
dns.bind_addr |
string | "::1" |
Bind address. Default is IPv6 loopback only; the shipped fips-dns-setup configures systemd-resolved to forward .fips queries to [::1]:5354. To expose the responder to mesh peers (or to the gateway over IPv4), override (e.g., "::" for all interfaces). |
dns.port |
u16 | 5354 |
Listen port |
dns.ttl |
u32 | 300 |
AAAA record TTL in seconds |
The dns.ttl value should not exceed node.cache.coord_ttl_secs to avoid
stale address mappings.
Host Mapping
The DNS resolver checks a host map before falling back to direct npub
resolution, enabling names like gateway.fips instead of npub1...fips.
The host map is populated from two sources:
- Peer aliases — the
aliasfield on configured peers inpeers:. - Hosts file —
/etc/fips/hosts, onehostname npub1...per line. Blank lines and#comments are allowed.
On conflict, hosts-file entries take precedence over peer aliases. The hosts file is auto-reloaded on modification (mtime change) without restarting the daemon. Hostnames are case-insensitive.
The installer ships /etc/fips/hosts pre-populated with the public test
mesh roster (test-us01 … test-uk01). Operator-style guide for
adding entries and the precedence rules:
../how-to/host-aliases.md.
Transports (transports.*)
UDP (transports.udp.*)
| Parameter | Type | Default | Description |
|---|---|---|---|
transports.udp.bind_addr |
string | "0.0.0.0:2121" |
UDP bind address and port. Ignored when outbound_only: true (kernel-assigned ephemeral port is used regardless). |
transports.udp.mtu |
u16 | 1280 |
Transport MTU |
transports.udp.recv_buf_size |
usize | 2097152 |
UDP socket receive buffer size in bytes (2 MB). Linux kernel doubles the requested value internally. Host net.core.rmem_max must be >= this value. |
transports.udp.send_buf_size |
usize | 2097152 |
UDP socket send buffer size in bytes (2 MB). Host net.core.wmem_max must be >= this value. |
transports.udp.advertise_on_nostr |
bool | false |
Include this UDP transport in Nostr endpoint adverts. Implicitly forced false when outbound_only: true. |
transports.udp.public |
bool | false |
If advertised: true publishes direct host:port; false publishes udp:nat rendezvous |
transports.udp.external_addr |
string | (none) | Explicit advertise-as override. Bare IP ("203.0.113.45" — bind port is appended) or full host:port. Takes precedence over the bound address and STUN autodiscovery. Useful when the public IP isn't on a local interface (cloud 1:1 NAT, EIP) or to skip STUN for a deterministic value. |
transports.udp.outbound_only |
bool | false |
Pure-client posture. When true, the transport binds to 0.0.0.0:0 (kernel-assigned ephemeral port) regardless of bind_addr, refuses inbound handshake msg1, and is never advertised on Nostr regardless of advertise_on_nostr. |
transports.udp.accept_connections |
bool | true |
Accept inbound handshake msg1 from new peers. Combine with outbound_only: false and accept_connections: false (plus auto_connect on peer entries) for a node that initiates outbound links but rejects fresh inbound handshakes. The handshake handler carves out msg1 from peers already established on this transport so rekey continues to work. |
Ethernet (transports.ethernet.*)
Ethernet transport sends raw frames over the platform's raw-frame
socket: AF_PACKET SOCK_DGRAM on Linux, BPF (/dev/bpf*) on macOS.
Linux and macOS only. On Linux it requires CAP_NET_RAW or running as
root; on macOS it requires read/write access to a /dev/bpf* device.
| Parameter | Type | Default | Description |
|---|---|---|---|
interface |
string | (required) | Network interface name (e.g., "eth0", "enp3s0") |
ethertype |
u16 | 0x2121 |
EtherType |
mtu |
u16 | (auto) | Override MTU. Default: interface MTU minus 3 (for frame type + length prefix) |
recv_buf_size |
usize | 2097152 |
Socket receive buffer size in bytes (2 MB) |
send_buf_size |
usize | 2097152 |
Socket send buffer size in bytes (2 MB) |
listen |
bool | true |
Listen for neighbor beacons from other nodes. Renamed from discovery in v0.5.0; the old key is still accepted as an alias, so a deployed config loads unchanged |
announce |
bool | false |
Broadcast announcement beacons on the LAN |
auto_connect |
bool | false |
Auto-connect to discovered peers |
accept_connections |
bool | false |
Accept incoming connection attempts from discovered peers |
beacon_interval_secs |
u64 | 30 |
Announcement beacon interval in seconds (minimum 10) |
optional |
bool | false |
Whether absence of the interface is normal. See below |
Dynamic binding. The interface does not have to exist when the daemon
starts. A transport whose interface is missing comes up absent: it is not a
start failure, it is not skipped, and it binds on its own the moment the
interface appears — sub-second where the kernel offers link events (netlink on
Linux, PF_ROUTE on the BSDs), within a second otherwise. An interface that
goes away at runtime unbinds and rebinds by the same path, so a wifi reload
or an unplugged adapter needs no restart. Presence means IFF_UP — the
interface exists and is administratively up — and deliberately not
IFF_RUNNING: binding needs no carrier, and the socket keeps working across a
carrier flap without rebinding. A bridge with nothing plugged into it, such as
br-lan on a wifi-only router, is therefore bound and healthy rather than
permanently Degraded, and starts carrying traffic the moment a port comes up.
Whether an interface has carrier is reported separately, as interface.carrier
in show_transports.
optional selects how that absence is reported:
| absence | log | retries | |
|---|---|---|---|
optional: false (default) |
node reports Degraded |
info at boot / warn on a runtime detach, then error once if it lasts past 10 s |
forever |
optional: true |
no health impact | info, and nothing after |
forever |
Naming an interface in configuration is a statement that you expect it, so the
default is to complain; silence is opted into. Set optional: true for
hardware that is legitimately not always there — a dock adapter, a radio only
some boards carry.
optional describes the interface's presence, not the transport's
importance. An optional interface that is present is used exactly as hard as
any other. No value of it makes a missing interface fatal at startup: the only
fatal case remains "no transports at all came up".
Either way the edge is logged once, on entering absence and on recovery — never once per retry.
The edge itself is not an error. An interface missing when the daemon starts
and bound a moment later is the ordinary boot race this mechanism exists to
absorb, so it is info; a runtime detach is warn, because a link coming and
going is ordinary weather for a mesh daemon and a cable unplugged for two
seconds does not need a human.
Ten seconds is the whole grace. Past that it is no longer a race against a
radio or a container coming up, so a required interface still missing is
reported once at error and stays Degraded until it returns. Start-time
absence and a runtime detach share the one deadline — they are the same
transition throughout this mechanism. An optional interface never reaches
error; that is what optional means.
Said once, not repeated. How long the absence has lasted is a state, and it
is published as one: interface.since_secs in show_transports, and
Degraded for as long as it holds. Re-announcing it on a timer would put a
second, lossier copy of that in the log.
Node health does not wait for the ten seconds — Degraded is published on the
first edge, and that is the signal to watch.
If a binding keeps dying moments after it is established (a socket that errors persistently while the interface stays up), the binder stops treating each bind as a recovery: it backs off on the same 1 s → 30 s curve, holds node health at its degraded reading, and stays quiet until a binding survives ten seconds. Without that damping a broken socket produces a health flap and a log pair every second, which is the same cry-wolf failure the edge-only logging rule exists to prevent.
A bind failure that is not absence — no CAP_NET_RAW, no readable
/dev/bpf*, a buffer the kernel refused — is a fault, not a state, and fails
the daemon's start as it always has. Only a missing interface is waited out.
fipsctl show transports reports the current state per transport under
interface: presence (absent / binding / present), carrier,
policy (required / optional), since_secs, binds, and
failed_attempts.
Named instances. Multiple Ethernet interfaces can be configured by using named sub-keys instead of flat parameters:
transports:
ethernet:
lan:
interface: "eth0"
listen: true
announce: true
backbone:
interface: "eth1"
announce: false
Each named instance operates independently with its own socket and
neighbor state. The instance name is used in log messages and the
name() method on the Transport trait.
TCP (transports.tcp.*)
TCP transport enables firewall traversal on networks that block UDP but allow TCP (e.g., port 443). Uses FMP header-based framing with zero overhead.
| Parameter | Type | Default | Description |
|---|---|---|---|
transports.tcp.bind_addr |
string | (none) | Listen address (e.g., "0.0.0.0:8443"). If omitted, outbound-only mode. |
transports.tcp.mtu |
u16 | 1400 |
Default MTU. Per-connection MTU derived from TCP_MAXSEG when available. |
transports.tcp.connect_timeout_ms |
u64 | 5000 |
Outbound connect timeout in milliseconds |
transports.tcp.nodelay |
bool | true |
TCP_NODELAY (disable Nagle for low latency) |
transports.tcp.keepalive_secs |
u64 | 30 |
TCP keepalive interval in seconds (0 = disabled) |
transports.tcp.recv_buf_size |
usize | 2097152 |
Socket receive buffer size in bytes (2 MB) |
transports.tcp.send_buf_size |
usize | 2097152 |
Socket send buffer size in bytes (2 MB) |
transports.tcp.max_inbound_connections |
usize | 256 |
Maximum simultaneous inbound connections |
transports.tcp.advertise_on_nostr |
bool | false |
Include this TCP transport in Nostr endpoint adverts |
transports.tcp.external_addr |
string | (none) | Explicit advertise-as override. Bare IP or full host:port. Required when bind_addr is wildcard (e.g. "0.0.0.0:443") and advertise_on_nostr: true, since TCP has no STUN equivalent for autodiscovery. Common on cloud 1:1 NAT / EIP setups where the public IP isn't bindable on the host. |
Named instances. Like other transports, multiple TCP instances can be configured with named sub-keys:
transports:
tcp:
public:
bind_addr: "0.0.0.0:443"
internal:
bind_addr: "10.0.0.1:8443"
max_inbound_connections: 64
Tor (transports.tor.*)
Tor transport routes FIPS traffic through the Tor network for anonymity.
Requires an external Tor daemon providing a SOCKS5 proxy. Three modes:
socks5 for outbound-only, control_port for outbound + monitoring,
directory for outbound + inbound via Tor-managed onion service.
| Parameter | Type | Default | Description |
|---|---|---|---|
transports.tor.mode |
string | "socks5" |
Tor access mode: socks5 (outbound only), control_port (outbound + monitoring), or directory (outbound + inbound onion service) |
transports.tor.socks5_addr |
string | "127.0.0.1:9050" |
SOCKS5 proxy address (host:port) |
transports.tor.connect_timeout_ms |
u64 | 120000 |
Connect timeout in milliseconds. Tor circuits take 10–60s. |
transports.tor.mtu |
u16 | 1400 |
Default MTU |
transports.tor.control_addr |
string | "/run/tor/control" |
Tor control port address: Unix socket path or host:port. Used in control_port mode; optional in directory mode for monitoring. |
transports.tor.control_auth |
string | "cookie" |
Control port authentication: "cookie", "cookie:/path/to/cookie", or "password:<secret>". |
transports.tor.cookie_path |
string | "/var/run/tor/control.authcookie" |
Path to Tor control cookie file. Used when control_auth is "cookie". |
transports.tor.max_inbound_connections |
usize | 64 |
Maximum inbound connections via onion service. |
transports.tor.directory_service.hostname_file |
string | "/var/lib/tor/fips_onion_service/hostname" |
Path to Tor-managed hostname file containing the .onion address. |
transports.tor.directory_service.bind_addr |
string | "127.0.0.1:8443" |
Local bind address for the listener that Tor forwards inbound connections to. Must match HiddenServicePort target in torrc. |
transports.tor.advertise_on_nostr |
bool | false |
Include this Tor transport in Nostr endpoint adverts. Requires node.rendezvous.nostr.enabled: true; setting it while Nostr rendezvous is disabled is a config-load error. advertised_port has no effect unless this is true. |
transports.tor.advertised_port |
u16 | 443 |
Public-facing onion port published in Nostr overlay adverts. Must match the virtual port in torrc's HiddenServicePort <port> 127.0.0.1:<bind_port> directive — that is the port other peers will use to reach this onion. |
Named instances. Like other transports, multiple Tor instances can be configured with named sub-keys for different SOCKS5 proxy endpoints.
Directory mode (recommended for production). Tor manages the onion
service via HiddenServiceDir in torrc. FIPS reads the .onion
address from the hostname file and binds a local TCP listener. This
enables Tor's Sandbox 1 (seccomp-bpf). If control_addr is also
set, the transport connects to the control port for daemon monitoring
(non-fatal on failure).
Control port mode. Connects to the Tor daemon's control port for
monitoring only (bootstrap status, circuit health, traffic stats).
No inbound connections. Both control_addr and control_auth are
required.
UDP + Tor Bridge Example
A node bridging clearnet (UDP) and anonymous (Tor) portions of the mesh:
node:
identity:
persistent: true
tun:
enabled: true
transports:
udp:
bind_addr: "0.0.0.0:2121"
mtu: 1472
tor:
socks5_addr: "127.0.0.1:9050"
peers:
- npub: "npub1abc..."
alias: "clearnet-peer"
addresses:
- transport: udp
addr: "203.0.113.5:2121"
- npub: "npub1def..."
alias: "anonymous-peer"
addresses:
- transport: tor
addr: "abc123...xyz.onion:2121"
Tor Directory Mode Example
A node accepting inbound connections via Tor-managed onion service (recommended for production — enables Sandbox 1):
node:
identity:
persistent: true
tun:
enabled: true
transports:
tor:
mode: "directory"
socks5_addr: "127.0.0.1:9050"
control_addr: "/run/tor/control" # optional, for monitoring
control_auth: "cookie"
directory_service:
hostname_file: "/var/lib/tor/fips/hostname"
bind_addr: "127.0.0.1:8444"
peers:
- npub: "npub1abc..."
alias: "tor-peer"
addresses:
- transport: tor
addr: "abcdef...xyz.onion:8443"
Requires a corresponding torrc:
HiddenServiceDir /var/lib/tor/fips
HiddenServicePort 8443 127.0.0.1:8444
Nym (transports.nym.*)
Nym transport routes FIPS traffic through the Nym mixnet for
metadata-resistant anonymity. Outbound-only: connections are made
through a nym-socks5-client SOCKS5 proxy that must be running
separately (e.g. as a service running alongside the fips daemon or as a
container). There is no inbound listener — a Nym-only node initiates
outbound links but is not reachable for unsolicited inbound handshakes.
| Parameter | Type | Default | Description |
|---|---|---|---|
transports.nym.socks5_addr |
string | "127.0.0.1:1080" |
nym-socks5-client SOCKS5 proxy address (host:port) |
transports.nym.connect_timeout_ms |
u64 | 300000 |
Outbound connect timeout in milliseconds. Mixnet SOCKS5 connections traverse 3 mix nodes with timing obfuscation and can take several minutes, so this is generous (300s). |
transports.nym.mtu |
u16 | 1400 |
Default MTU |
transports.nym.startup_timeout_secs |
u64 | 120 |
Seconds to wait for nym-socks5-client to become ready at startup before giving up |
Named instances. Like other transports, multiple Nym instances can be configured with named sub-keys for different SOCKS5 proxy endpoints.
BLE (transports.ble.*)
Bluetooth Low Energy transport using L2CAP Connection-Oriented Channels.
Compiled on glibc Linux and on Android. At build time, build.rs sets
bluer_available from the target triple (Linux and not musl) and sets
ble_available for that or Android; the BLE runtime is gated behind
#[cfg(ble_available)], with bluer_available gating only the BlueZ
backend inside it. There is no Cargo feature flag to toggle. On musl
Linux or any other platform, BLE config still parses but the transport
runtime is absent and config entries become no-ops. On glibc Linux the
transport communicates with BlueZ via D-Bus through the bluer crate;
on Android the radio is supplied by the embedding application.
| Parameter | Type | Default | Description |
|---|---|---|---|
transports.ble.adapter |
string | "hci0" |
HCI adapter name |
transports.ble.psm |
u16 | 0x0085 (133) |
L2CAP Protocol/Service Multiplexer |
transports.ble.mtu |
u16 | 2048 |
Default MTU. Actual MTU is negotiated per-link during L2CAP connection setup. |
transports.ble.max_connections |
usize | 7 |
Maximum concurrent BLE connections |
transports.ble.connect_timeout_ms |
u64 | 10000 |
Outbound connect timeout in milliseconds |
transports.ble.advertise |
bool | true |
Broadcast BLE beacon advertisements for peer discovery |
transports.ble.scan |
bool | true |
Listen for BLE beacon advertisements from other nodes |
transports.ble.auto_connect |
bool | false |
Automatically connect to discovered peers |
transports.ble.accept_connections |
bool | true |
Accept incoming L2CAP connections |
transports.ble.probe_cooldown_secs |
u64 | 30 |
Cooldown before re-probing the same BLE address |
Address format. BLE peer addresses use the form
"adapter/device_address" — for example, "hci0/AA:BB:CC:DD:EE:FF".
Advertising and scanning. When advertise is enabled, the transport
advertises the FIPS service UUID continuously so that nearby nodes can
discover and connect via L2CAP, plus the L2CAP PSM its listener actually
bound, as a service-data structure (see src/transport/ble/psm.rs for the
wire layout and why platforms with OS-assigned PSMs need it). The
advertisement carries no device name — alongside the PSM a name no longer
fits the 31-byte legacy PDU, so the node shows up in generic Bluetooth
scanners as an unnamed device with the FIPS UUID. When scan is enabled,
the transport continuously scans for other FIPS nodes' advertisements and
learns each peer's advertised PSM; a peer that advertises none is dialled
at the configured psm. Discovered
peers are probed immediately (L2CAP connect + pubkey exchange) with a
cooldown (probe_cooldown_secs) to prevent rapid re-probing of the same
address. If two nodes probe each other at the same time (cross-probe),
a deterministic tie-breaker based on NodeAddr comparison ensures only
one connection is established.
Connection pool. The max_connections parameter limits the number of
concurrent BLE connections. When the pool is full, the least-recently-used
connection is evicted to make room for new connections.
BLE Example
A node using BLE for local mesh discovery alongside UDP for internet peers:
node:
identity:
persistent: true
tun:
enabled: true
transports:
udp:
bind_addr: "0.0.0.0:2121"
ble:
adapter: "hci0"
advertise: true
scan: true
auto_connect: true
accept_connections: true
peers:
- npub: "npub1abc..."
alias: "internet-peer"
addresses:
- transport: udp
addr: "203.0.113.5:2121"
connect_policy: auto_connect
BLE peers on the local radio range are discovered automatically via beacons — no static peer entries needed. Internet peers still require explicit configuration.
Peers (peers[])
Static peer list. Each entry defines a peer to connect to.
| Parameter | Type | Default | Description |
|---|---|---|---|
peers[].npub |
string | (required) | Peer's Nostr public key (npub-encoded) |
peers[].alias |
string | (none) | Human-readable name for logging |
peers[].addresses |
list | [] |
Transport addresses for the peer. May be left empty (or omitted) when via_nostr: true, in which case the daemon resolves endpoints from the peer's Nostr advert at dial time. |
peers[].addresses[].transport |
string | (required) | Transport type: udp, tcp, ethernet, tor, nym, or ble. A udp entry may be qualified with a named instance as udp/<instance> (see below). |
peers[].addresses[].addr |
string | (required) | Transport address. UDP/TCP: "host:port" (IP or DNS hostname). Ethernet: "interface/mac" (e.g., "eth0/aa:bb:cc:dd:ee:ff"). BLE: "adapter/device_address" (e.g., "hci0/AA:BB:CC:DD:EE:FF"). Tor: ".onion:port" or "host:port" |
peers[].addresses[].priority |
u8 | 100 |
Address priority (lower = preferred) |
peers[].connect_policy |
string | "auto_connect" |
Connection policy: auto_connect, on_demand, or manual. Note: on_demand and manual are reserved for future use; the only policy currently honored at runtime is auto_connect. |
peers[].auto_reconnect |
bool | true |
Automatically reconnect after MMP link-dead removal (exponential backoff, unlimited retries) |
peers[].via_nostr |
bool | false |
Append Nostr advert-derived endpoints after static addresses for this peer |
Named UDP instances. Where several UDP transports are configured
under named sub-keys, a peer address can name the one it belongs to by
writing the transport field as udp/<instance>, for example
udp/aware. A bare udp matches any instance. The qualifier resolves
only for udp: writing it on any other transport type, or naming a UDP
instance that is not configured, fails config load with a validation
error rather than falling back to another instance.
Gateway (gateway.*)
The gateway.* block configures the optional fips-gateway
service, which lets unmodified LAN hosts reach mesh destinations
through DNS proxy + virtual-IP NAT (and, optionally, exposes
LAN-side services back into the mesh through inbound port forwards).
The gateway is a separate service from the FIPS daemon but reads the
same fips.yaml file. The block is read only when fips-gateway is
running; the fips daemon ignores it. Linux only — the field is
gated behind #[cfg(target_os = "linux")]. For setup, see
../how-to/deploy-gateway.md; for the
end-to-end design, see
../design/fips-gateway.md.
| Parameter | Type | Default | Description |
|---|---|---|---|
gateway.enabled |
bool | false |
Enable the gateway. Must be true for fips-gateway to start. |
gateway.pool |
string | (required) | Virtual IPv6 pool CIDR (e.g., "fd01::/112"). Must not overlap with the FIPS mesh address space (fd00::/8) or any address space already in use on the LAN. The /112 size yields 65 535 usable virtual IPs (address 0 in the pool is skipped), which is the gateway's hard cap regardless of CIDR width. |
gateway.lan_interface |
string | (required) | LAN-facing network interface name (e.g., "enp3s0"). Used for proxy-NDP entry installation so LAN clients can resolve the link-layer address of allocated virtual IPs. |
gateway.pool_grace_period |
u64 | 60 |
Seconds a virtual-IP allocation is retained after its last referencing session ends, before the address is returned to the free pool. Larger values reduce churn for short-lived flows; smaller values reclaim addresses faster. |
Gateway DNS (gateway.dns.*)
Settings for the gateway's DNS listener and its upstream link to the
FIPS daemon's .fips resolver. The gateway proxies .fips queries to
the daemon's resolver, which returns mesh addresses; the gateway then
allocates a virtual IP from the pool and rewrites the response.
Non-.fips queries are answered with REFUSED.
| Parameter | Type | Default | Description |
|---|---|---|---|
gateway.dns.listen |
string | "[::1]:5353" |
DNS listen address. The default binds IPv6 loopback on an unprivileged port, matching the canonical deployment where another resolver on the host (dnsmasq, systemd-resolved, BIND) holds port 53 and forwards .fips queries to the gateway over loopback. Bind on the LAN-side IP (e.g., "192.168.1.1:53") or wildcard ("[::]:53") only on hosts with no other resolver on 53 and where LAN clients query the gateway directly. See ../how-to/troubleshoot-gateway.md. |
gateway.dns.upstream |
string | "[::1]:5354" |
Upstream FIPS daemon resolver. Must match the daemon's dns.bind_addr and dns.port. Defaults match the daemon defaults (::1:5354). A v4 upstream ("127.0.0.1:5354") cannot reach a daemon bound on [::1]:5354 — Linux IPv6 sockets bound to explicit ::1 do not accept v4-mapped traffic. If you change the daemon's dns.bind_addr, update this field accordingly. |
gateway.dns.ttl |
u32 | 60 |
TTL in seconds on AAAA responses returned to LAN clients. Smaller values let the gateway recycle pool addresses faster; larger values reduce LAN-side query traffic. |
Conntrack (gateway.conntrack.*)
Linux conntrack timeout overrides for the gateway's NAT table. These adjust the kernel-default timeouts for NAT sessions installed by the gateway. All values are in seconds; omit any field to inherit the gateway's built-in default (which itself usually matches the kernel default for that protocol).
| Parameter | Type | Default | Description |
|---|---|---|---|
gateway.conntrack.tcp_established |
u64 | 432000 |
TCP established-state timeout (5 days). Long-lived TCP flows (SSH, persistent HTTP) keep their NAT mapping alive for at least this long without traffic. |
gateway.conntrack.udp_timeout |
u64 | 30 |
UDP unreplied timeout. Applied until reply traffic is observed in the reverse direction. |
gateway.conntrack.udp_assured |
u64 | 180 |
UDP assured (bidirectional) timeout. Applied once reply traffic has been observed. |
gateway.conntrack.icmp_timeout |
u64 | 30 |
ICMP echo / error timeout. |
Inbound Port Forwards (gateway.port_forwards[])
Optional list of inbound port-forward rules. Each rule maps a TCP or
UDP port on the gateway's fips0 mesh-side address to a host:port
on the LAN. Mesh peers connect to the gateway's mesh address on the
listen port; the gateway terminates the connection and forwards the
payload to the LAN target. This is the inverse of the outbound mode:
the LAN service is exposed to the mesh, not the other way around. See
../how-to/deploy-gateway.md for the
operator recipe.
| Parameter | Type | Default | Description |
|---|---|---|---|
gateway.port_forwards[].listen_port |
u16 | (required) | Port on fips0 that mesh peers connect to. Must be non-zero. The (listen_port, proto) pair must be unique across the list. |
gateway.port_forwards[].proto |
string | (required) | Transport protocol: tcp or udp. |
gateway.port_forwards[].target |
string | (required) | LAN destination as IPv6 [addr]:port (e.g., "[fd12:3456::10]:80"). IPv4 targets are rejected at config-load time. |
Gateway Example
A typical gateway with both outbound (LAN-to-mesh) and inbound (mesh-to-LAN) modes enabled:
gateway:
enabled: true
pool: "fd01::/112"
lan_interface: "enp3s0"
dns:
listen: "[::1]:5353"
upstream: "[::1]:5354"
ttl: 60
pool_grace_period: 60
conntrack:
tcp_established: 432000
udp_assured: 180
port_forwards:
- listen_port: 8080
proto: tcp
target: "[fd12:3456::10]:80"
- listen_port: 5353
proto: udp
target: "[fd12:3456::10]:53"
Minimal Example
A typical node configuration enabling TUN, DNS, and a single peer:
node:
identity:
nsec: "0102030405060708090a0b0c0d0e0f101112131415161718191a1b1c1d1e1f20"
tun:
enabled: true
name: fips0
mtu: 1280
dns:
enabled: true
bind_addr: "127.0.0.1"
port: 53
transports:
udp:
bind_addr: "0.0.0.0:2121"
mtu: 1472
peers:
- npub: "npub1tdwa4vjrjl33pcjdpf2t4p027nl86xrx24g4d3avg4vwvayr3g8qhd84le"
alias: "node-b"
addresses:
- transport: udp
addr: "172.20.0.11:2121"
connect_policy: auto_connect
Mixed UDP + Ethernet Example
A node bridging internet peers (UDP) and a local Ethernet segment with neighbor beacons:
node:
identity:
nsec: "0102030405060708090a0b0c0d0e0f101112131415161718191a1b1c1d1e1f20"
tun:
enabled: true
transports:
udp:
bind_addr: "0.0.0.0:2121"
mtu: 1472
ethernet:
interface: "eth0"
listen: true
announce: true
auto_connect: true
accept_connections: true
peers:
- npub: "npub1tdwa4vjrjl33pcjdpf2t4p027nl86xrx24g4d3avg4vwvayr3g8qhd84le"
alias: "internet-peer"
addresses:
- transport: udp
addr: "203.0.113.5:2121"
connect_policy: auto_connect
Ethernet peers on the local segment are discovered automatically via beacons — no static peer entries needed. Internet peers still require explicit configuration.
All node.* parameters use their defaults. To override specific values, add
only the relevant sections:
node:
identity:
nsec: "..."
limits:
max_peers: 64
retry:
max_retries: 10
max_backoff_secs: 600
cache:
coord_size: 100000
Deprecated Keys
Every key below still loads. Nothing has been removed, so a config file written against v0.4.x keeps working after an upgrade. The old spellings are scheduled for removal at the next wire-protocol cutover, so migrate when convenient rather than urgently.
node.discovery.* split into node.lookup.* and node.rendezvous.*
The single node.discovery table mixed two unrelated jobs: resolving
coordinates for a mesh address already known (lookup), and finding peers
to connect to in the first place (rendezvous). It is now two tables.
| Deprecated key | Replacement |
|---|---|
node.discovery.ttl |
node.lookup.ttl |
node.discovery.attempt_timeouts_secs |
node.lookup.attempt_timeouts_secs |
node.discovery.recent_expiry_secs |
node.lookup.recent_expiry_secs |
node.discovery.backoff_base_secs |
node.lookup.backoff_base_secs |
node.discovery.backoff_max_secs |
node.lookup.backoff_max_secs |
node.discovery.forward_min_interval_secs |
node.lookup.forward_min_interval_secs |
node.discovery.nostr.* (whole sub-table) |
node.rendezvous.nostr.* |
node.discovery.lan.* (whole sub-table) |
node.rendezvous.lan.* |
Behaviour of a deployed node.discovery: block: each config file is
folded as it is parsed, before the cross-file merge, and a single warning
is logged on the fips::config target naming the move. Only the keys
actually present in the old block are applied; the rest keep their
defaults. The compat block is never written back out, so anything that
re-serializes the configuration emits the new spelling only.
Mixing the two spellings inside one file is not a merge. The fold
runs after that file is parsed, so a value under node.discovery
overwrites whatever the corresponding node.lookup or node.rendezvous
key held in the same file. Use one spelling per file.
transports.ethernet.discovery renamed to transports.ethernet.listen
| Deprecated key | Replacement |
|---|---|
transports.ethernet.discovery |
transports.ethernet.listen |
This one is a plain alias rather than a compat fold, so both spellings
parse into the same field and no warning is logged. The name changed
because the key never controlled discovery in the node.discovery
sense: it decides whether the interface listens for neighbour beacons.
Complete Reference
The full YAML structure with all defaults:
node:
identity:
nsec: null # secret key in nsec or hex (null = depends on persistent)
persistent: false # true = load/save fips.key; false = ephemeral each start
leaf_only: false
tick_interval_secs: 1
base_rtt_ms: 100
heartbeat_interval_secs: 10
link_dead_timeout_secs: 30
# drain_timeout_secs: 2 # bounded Draining phase; absent = 2s
limits:
max_connections: 256
max_peers: 128
max_links: 256
max_pending_inbound: 1000
rate_limit:
handshake_burst: 100
handshake_rate: 10.0
handshake_timeout_secs: 30
handshake_resend_interval_ms: 1000
handshake_resend_backoff: 2.0
handshake_max_resends: 5
session_setup_burst: 64
session_setup_rate: 16.0
retry:
max_retries: 5
base_interval_secs: 5
max_backoff_secs: 300
netmon:
enabled: true
poll_interval_secs: 5
debounce_ms: 250
cache:
coord_size: 50000
coord_ttl_secs: 300
identity_size: 10000
lookup:
ttl: 64
attempt_timeouts_secs: [1, 2, 4, 8]
recent_expiry_secs: 10
backoff_base_secs: 0
backoff_max_secs: 0
forward_min_interval_secs: 2
rendezvous:
# nostr: # uncomment to enable Nostr rendezvous
# enabled: true # opt-in, default false
# policy: configured_only # disabled | configured_only | open
# lan: # uncomment to enable mDNS LAN rendezvous
# enabled: true # opt-in, default false
# scope: "my-mesh" # optional per-network scope filter
tree:
announce_min_interval_ms: 500
parent_hysteresis: 0.2 # cost improvement fraction for parent switch
hold_down_secs: 30 # suppress re-evaluation after switch
reeval_interval_secs: 60 # periodic cost-based re-evaluation (0 = disabled)
flap_threshold: 4 # parent switches before dampening
flap_window_secs: 60 # sliding window for flap detection
flap_dampening_secs: 120 # extended hold-down on flap
bloom:
update_debounce_ms: 500
max_inbound_fpr: 0.20 # antipoison cap on inbound FilterAnnounce FPR
session:
default_ttl: 64
pending_packets_per_dest: 16
pending_max_destinations: 256
idle_timeout_secs: 90
coords_warmup_packets: 5
coords_response_interval_ms: 2000
mmp:
mode: full # full | lightweight | minimal
log_interval_secs: 30
owd_window_size: 32
session_mmp:
mode: full # full | lightweight | minimal
log_interval_secs: 30
owd_window_size: 32
ecn:
enabled: true # ECN congestion signaling (CE flag relay)
loss_threshold: 0.05 # MMP loss rate threshold for CE marking (5%)
etx_threshold: 3.0 # MMP ETX threshold for CE marking
rekey:
enabled: true # periodic Noise rekey for forward secrecy
after_secs: 120 # rekey interval (seconds)
after_messages: 65536 # rekey after N messages sent
control:
enabled: true
socket_path: null # null = auto (platform runtime dir → XDG → /tmp)
# native_api: # uncomment to enable the experimental native datagram API
# enabled: true # opt-in, default false; not on Windows
# socket_path: /run/fips/api.sock # omit the key for the resolution above
# pending_per_flow: 16 # datagrams held for one flow; 1..=64
# backlog: 16 # flows announced on one listener, awaiting its task; at least 1
# max_flows: 256 # flows this node holds at once
# debug_commands: false # inject/stats/arrive; test harness only
buffers:
packet_channel: 1024
tun_channel: 1024
dns_channel: 64
tun:
enabled: false
name: "fips0"
mtu: 1280
dns:
enabled: true
bind_addr: "::1"
port: 5354
ttl: 300
transports:
udp:
bind_addr: "0.0.0.0:2121"
mtu: 1280
recv_buf_size: 2097152 # 2 MB (kernel doubles to 4 MB actual)
send_buf_size: 2097152 # 2 MB
# ethernet: # uncomment to enable (requires CAP_NET_RAW)
# interface: "eth0" # required: network interface name
# ethertype: 0x2121 # default EtherType
# mtu: null # null = interface MTU - 3 (typically 1497)
# recv_buf_size: 2097152 # 2 MB
# send_buf_size: 2097152 # 2 MB
# listen: true # listen for beacons
# announce: false # broadcast beacons
# auto_connect: false # connect to discovered peers
# accept_connections: false # accept inbound handshakes
# beacon_interval_secs: 30 # beacon interval (min 10)
# tcp: # uncomment to enable TCP transport
# bind_addr: "0.0.0.0:8443" # listen address (omit for outbound-only)
# mtu: 1400 # default MTU
# connect_timeout_ms: 5000 # outbound connect timeout
# nodelay: true # TCP_NODELAY
# keepalive_secs: 30 # keepalive interval (0 = disabled)
# recv_buf_size: 2097152 # 2 MB
# send_buf_size: 2097152 # 2 MB
# max_inbound_connections: 256 # resource protection limit
# tor: # uncomment to enable Tor transport
# mode: "socks5" # "socks5", "control_port", or "directory"
# socks5_addr: "127.0.0.1:9050" # SOCKS5 proxy address
# connect_timeout_ms: 120000 # connect timeout (120s for Tor circuits)
# mtu: 1400 # default MTU
# # monitoring (control_port mode, or optional in directory mode):
# # control_addr: "/run/tor/control" # Unix socket or host:port
# # control_auth: "cookie" # "cookie" or "password:<secret>"
# # cookie_path: "/var/run/tor/control.authcookie"
# # directory mode (inbound via Tor-managed onion service):
# # directory_service:
# # hostname_file: "/var/lib/tor/fips_onion_service/hostname"
# # bind_addr: "127.0.0.1:8443"
# # max_inbound_connections: 64
# # advertise_on_nostr: false # publish this onion in Nostr adverts
# # # (requires node.rendezvous.nostr.enabled)
# # advertised_port: 443 # public-facing onion port for Nostr adverts
# nym: # uncomment to enable Nym mixnet transport (outbound-only)
# socks5_addr: "127.0.0.1:1080" # nym-socks5-client SOCKS5 proxy address
# connect_timeout_ms: 300000 # connect timeout (300s for mixnet)
# mtu: 1400 # default MTU
# startup_timeout_secs: 120 # wait for nym-socks5-client to be ready
# ble: # uncomment to enable BLE transport (Linux only, requires BlueZ)
# adapter: "hci0" # HCI adapter name
# psm: 0x0085 # L2CAP PSM (133)
# mtu: 2048 # default MTU (negotiated per-link)
# max_connections: 7 # max concurrent BLE connections
# connect_timeout_ms: 10000 # outbound connect timeout
# advertise: true # broadcast BLE beacons
# scan: true # listen for BLE beacons
# auto_connect: false # connect to discovered peers
# accept_connections: true # accept incoming L2CAP connections
# probe_cooldown_secs: 30 # cooldown before re-probing same address
peers: # static peer list
# - npub: "npub1..."
# alias: "node-b"
# addresses:
# - transport: udp
# addr: "10.0.0.2:2121"
# priority: 100
# connect_policy: auto_connect
# auto_reconnect: true # reconnect after link-dead removal