Files
fips/docs/reference/configuration.md
T
ArjenandJohnathan Corgan abe3f0b40a feat(transport): bind and rebind network interfaces dynamically
An interface-bound transport is a long-lived object that is *sometimes
bound*. The interface it names may not exist when the daemon starts, may
appear minutes later, may vanish and return mid-operation, and may never
appear at all. Until now the first observation was final: an interface
missing at start was logged once and skipped for the life of the process,
and one that disappeared at runtime published a health change but was never
rebound.

On OpenWrt that is a live bug. procd starts fips before wifi has created
fips-mesh0 / fips-ap0, both transports are skipped and never retried, and
the 802.11s peer link forms anyway (that is mac80211, not the daemon) — so
the node looks healthy and reaches nothing.

Give every interface-bound transport presence state and a binder task.
start_async now returns Ok with the transport ABSENT rather than failing;
the binder binds when the interface appears, tears the socket down when it
goes away, and rebinds when it returns. Start-time absence and runtime
detach are one transition and one code path.

TransportError::InterfaceUnavailable makes absence branchable: a missing
interface and a typo'd interface name were the same flat StartFailed
(String), so nothing downstream could tell a state from a fault. A bind
failure that is *not* absence — no CAP_NET_RAW, no readable /dev/bpf* —
still fails the start, because retrying a socket that can never open behind
a Degraded nobody is watching is worse than dying loudly at boot.

**Presence is IFF_UP, not IFF_UP|IFF_RUNNING.** Carrier and bindability are
different questions and only the second belongs in a bind gate. An
AF_PACKET socket on a carrier-less bridge is valid and starts carrying
traffic the instant a member port comes up, with no rebind. Gating on
carrier would report a wifi-only router Degraded forever for an empty
br-lan, turn every carrier flap the socket would have survived into an
unbind/rebind cycle, and deadlock an 802.11s interface that reports RUNNING
only once it has peered — peering needs beacons, beacons need a bound
socket. Carrier is reported beside presence in show_transports instead.

**A name is not a device.** Both backends bind by device: AF_PACKET stores
sll_ifindex, a BPF descriptor follows the interface it was attached to. A
netdev deleted and recreated under the same name leaves the socket attached
to something gone while the name resolves perfectly well, and nothing
notices — a stale AF_PACKET socket never becomes readable, so the receive
loop neither errors nor exits, and send failures go to the caller rather
than the binder. A listen-only node (announce: false, no beacon sender to
fail) sat present and deaf indefinitely after a wifi reload. The bound index
is captured at bind and compared on every poll. A name reappearing with a
different MAC is different hardware: cached neighbors are dropped rather
than resumed onto.

**Detection is event-driven** where the kernel offers a source — netlink
RTNLGRP_LINK on Linux, PF_ROUTE on macOS and FreeBSD — with a 1 s
getifaddrs poll underneath as a backstop. Link-event payloads are not
parsed: an event is a hint to re-run the probe, which is cheap and
authoritative, and a parser's bugs would be presence bugs. Probes are
coalesced to ten a second because PF_ROUTE has no group filter and delivers
every routing message on the host; a persistently failing event source is
logged once, backed off, and abandoned for the poll after five errors
rather than spinning a core silently on ENOBUFS.

**Degraded becomes a level, not a latch.** The supervisor's reason set was
monotonic, which was correct while no child could recover; with recovery it
would have come to mean "something broke at some point since boot" rather
than "something is broken now". Absence lives in its own reversible set,
health is recomputed on every transition in both directions, and an absent
transport still counts as up so a single-interface node that boots before
its wifi degrades rather than exiting on NoTransports. There is deliberately
no restart action in the FSM: rebinding belongs next to the file descriptor,
and a supervisor-authored retry would be a second mechanism racing the first
for the same socket.

**The node's egress MTU follows the bound set.** transport_mtu filters on
is_bound(), not is_operational() — an interface-bound transport is
operational from the moment it starts, so the weaker predicate let hardware
that had never appeared clamp the whole node's IPv6 MTU. Because that
minimum now moves at runtime, the TUN reader and writer read the TCP MSS
ceiling from a shared atomic instead of a u16 captured at spawn; every other
consumer (show_status, the snapshot, the session-layer fragmentation check)
already read it live, so the clamp was the one place the daemon could report
one effective MTU and enforce another. It moves in both directions: a narrow
interface appearing tightens it, its departure releases it. MSS is
negotiated per connection, so a change binds connections opened after it.

**Logging is per edge, never per attempt**, with one deadline after it. The
edge itself is not an error — info at boot, warn on a runtime detach, since
an interface bound a moment later is the ordinary case this mechanism exists
to absorb and crying error at t=0 then "recovered" at t=0.2s is the failure
the rule exists to prevent. Ten seconds is the whole grace: past it, absence
is no longer a race against a radio or a container, so a required interface
still missing is reported once at error. Start-time absence and a runtime
detach share that one deadline, as they share everything else here. Said
once, not repeated: how long an absence has lasted is state, published as
interface.since_secs and as Degraded, and a monitor can threshold it per
deployment rather than the daemon compiling a schedule in.

Successful rebinds are damped. Backoff covers failed binds; the nastier case
is binds that keep succeeding into a socket that dies moments later, which a
receive loop giving up on a persistent error while the interface stays UP
produces once per second forever. Consecutive bindings dying inside ten
seconds back off on the 1 s → 30 s curve, and past three the binder stops
announcing each bind as a recovery until one lasts.

The receive loop backs off and exits on a dead socket instead of spinning on
Err with a warn per iteration, and the ad-hoc ENXIO socket reopen in the
beacon sender is gone: both hand recovery to the presence machine, one
mechanism for every cause rather than one hack per symptom. Beacons pause
while absent because the task simply does not exist then.

transports.ethernet.*.optional (default false) decides how loudly absence is
reported. Naming an interface is a statement that you expect it, so the
default complains: Degraded, and the error at the deadline. optional: true
is silent — info on the edge, no health impact, no error — for hardware
legitimately not always there. It describes the interface's presence, not
the transport's importance, and no value of it makes a missing interface
fatal at startup.

show_transports grows an `interface` block — presence, carrier, policy, how
long the phase has been held, bind count, failed attempts. The original boot
race was expensive precisely because nothing an operator could see said the
node was deaf.

ci: exercise the presence probe against musl and an address-less interface

Two legs, both about the assumption the presence probe rests on.

OpenWrt — the platform dynamic interface binding exists for — is musl, and
musl reimplements getifaddrs independently of glibc. `interface_present`
reads ifa_flags out of it, and the interfaces this feature exists for
(fips-mesh0, fips-ap0) are unbridged with no IP address at all, which is
exactly where getifaddrs implementations differ. Every other leg is glibc,
so without this one the probe was asserted on a libc no test had ever run it
against, on the target it was written for. Building for the musl target
rather than inside an Alpine container keeps the leg to a cross-build plus a
run, without a second toolchain image to maintain.

The address-less interface is created on both this leg and the glibc one.
Pinning the contract on both libcs is what turns "glibc and musl agree about
getifaddrs" from an assumption into a checked fact, and makes the glibc leg
fail first if glibc is the one that changes. Loopback cannot stand in: it
has 127.0.0.1, so probing it asks whether getifaddrs works rather than
whether it reports this. Deliberately not tolerant of failure — a guard that
quietly does not run is worse than no guard.

The iface-binding suite is deliberately not added to the integration
matrix here. Its files and its testing/ci-local.sh registration arrive
in the next commit, and the GitHub workflow's matrix entry now goes with
them, so both runners gain the suite together and
testing/check-ci-parity.sh holds at every point in this history.

docs: document dynamic interface binding

The transport-layer design gains an Interface Presence section: the three
deployment scenarios that motivated it, the presence machine and its two
invariants, why presence is IFF_UP and not IFF_UP|IFF_RUNNING, interface
identity and the recreated-netdev case, the detection sources and their two
rate limits, the optional policy, health as a level rather than a latch, the
egress-MTU consequence, the logging rules, and what the mechanism retires.
It records why the error deadline is a single ten-second window rather than
a repeating severity ladder, so the ladder does not come back.

The control-socket and fipsctl references described show_transports without
its `interface` block, and the transport-layer state machine still implied
that Up meant bound. Both now say what Up and presence each describe and how
they differ, since "transport up, interface absent" is a normal state an
operator will meet and would otherwise read as a contradiction.

The configuration reference documents `optional` with the absence, log and
retry behaviour of each setting, and the ground-up tutorial's dongle example
gains optional: true, which is what that example is actually for.

fix(transport): pair every presence edge the binder publishes

Six defects, all in the wiring around the presence machine rather than in
the machine itself. The state machines were well tested; what went untested
was how the binder feeds them, and every one of these lives there.

**The first detach after a clean start never reached node health.**
start_async binds inline and publishes that edge itself, before binder_loop
exists. The loop then built a fresh ChurnGuard, which had therefore recorded
no bind — and detached() takes `announced` to decide whether an edge is
owed, so it asked for no retraction. A node that booted with its interface
present and then lost it kept reporting Full for as long as the interface
stayed away, while show_transports said absent and the log said detached.
stabilized() could not repair it either: it early-returns while bound_at is
None, which it is for a bind the loop did not perform, so the guard stayed
unseeded until a second detach happened to fix it. The loop now seeds the
guard from the binding it inherits.

This is not the cable-unplug case. Presence is IFF_UP, so a carrier loss is
correctly not a detach at all; it takes an admin down, a netdev delete, or a
device removal — a wifi reload, a hostapd restart, a dongle pulled.

**An optional interface published nothing, and the MTU floor rode the same
channel.** Filtering the edge at the source conflated two questions that
happen to share a transport: whether node health should move, and whether
the bound set changed. Only the first is policy. The second determines the
node's egress MTU floor, and refresh_tun_mss_ceiling has no other trigger —
so an optional transport binding or unbinding at runtime left the TUN MSS
clamp derived from a bound set that no longer existed, reporting one
effective MTU and clamping to another. That is the defect the MssCeiling
work was written to remove, reintroduced for exactly the transports the
shipped OpenWrt config marks optional: five of seven. TransportPresence
gains health_relevant; the edge always goes out and only health is filtered.

**A permanent bind fault went unlogged if any absence race preceded it.**
record_attempt ran ahead of the InterfaceUnavailable arm, so a probe that
won a race the bind then lost burned the counter that gates the only error
a real fault ever gets — emitted on attempts == 1. A flapping interface
followed by CAP_NET_RAW being dropped or /dev/bpf* exhausting therefore
reported nothing at all, for the life of the process, and the operator got
report_sustained_absence's "still missing" instead: wrong, for an interface
that is present. An absence race is not a bind attempt and no longer counts
as one.

**The watcher's give-up did not stick.** Abandoning the event source parked
on pending() *inside the changed() future*, and that future is constructed
fresh on every pass of the binder's select! and dropped whenever the poll
ticker wins. So the next pass re-read the dead socket, re-counted the error
and re-logged "not recoverable" — once per wake-up, forever, which at a 1 s
tick across the shipped seven-transport config is seven warnings and seven
failing syscalls a second on flash-backed logging. The 100 ms backoff lived
in the dropped future too and never applied across passes. Abandonment is
now state on the watcher.

**A zero-length read livelocked the binder.** try_io clears readiness only
on WouldBlock, which the sibling error arm handles by hand and this one did
not — so breaking out left readable() instantly ready with nothing to read,
and the loop never returned Pending. That starves the select! of its ticker
entirely and takes presence detection down with it. Cleared and treated as a
fault so the give-up path applies.

**A failed presence probe read as an absent interface.** getifaddrs is a
netlink dump and fails for reasons that have nothing to do with the
interface: ENOBUFS under memory pressure, EMFILE or ENFILE under fd
exhaustion, since it opens a socket of its own. Answering "not present"
there tore down a working socket and degraded health over a transient
syscall failure, undiagnosably, and under fd exhaustion the rebind could not
have succeeded anyway. interface_has_flags now distinguishes the two; the
detach gate holds its binding when the kernel will not answer, while the
bind gate still treats it as absence and retries.

**On macOS the reader thread could not die, so a dead socket read as live.**
Any read() failure other than EBADF — ENXIO being the one that matters,
which is what BPF answers once the interface it was attached to is torn away
— reset the parse buffer and continued. The thread never returned, so the
channel never closed, so recv_from never failed, so the tokio task never
exited, so tasks_alive() reported a dead socket as a live one. Detach
detection on macOS reduced to the name and the index, and an interface reset
in place left the transport present and deaf. It now gives up after a
streak, and the return closes the channel the binder is actually watching.

Two smaller pairings while here: stop_async retracts the edge it would
otherwise leave standing for a socket that is gone, and start_async hands a
refused edge to the binder to retry rather than dropping the only edge
either consumer will see until the interface next moves. The absence
deadline is stamped at start rather than at construction, so a transport
staged for longer than the bring-up window no longer reports sustained
absence on its first tick having given the interface no window at all.

feat(config): reject impossible ethernet interface names at load

Config::validate never inspected transports.ethernet, so an empty interface
name, a name past the kernel's 15-byte limit, and two transports naming the
same netdev all loaded cleanly and failed only at runtime — the first two as
a permanent absence indistinguishable from an interface that has not been
created yet.

That indistinguishability is deliberate and worth keeping: waiting is the
right answer for an interface the operator has not made yet, and the daemon
cannot know which of the two it is looking at. Which is exactly why the
syntactic gate earns its place. A name that is *impossible* is the one case
still separable from "not there yet", and without the check a typo costs a
permanently Degraded node whose only symptom is an interface that never
arrives — the failure mode the presence machine exists to make legible,
reintroduced one level up.

Syntax only. Whether a well-formed name exists stays the binder's question,
asked once a second, forever.

The duplicate check is a different fault: two transports on one netdev means
two sockets on the same device at the same ethertype, each receiving every
frame the other does.

test(transport): close the vacuous and uncovered branches

Six tests that asserted nothing, or asserted less than they claimed.

`a_bind_fault_still_fails_the_start` returned early whenever the socket
*could* be opened, so it was vacuous as root and on any developer machine with
a group-readable /dev/bpf* — the fail-fast path it is named for went unchecked
exactly where someone was most likely to run it. It now asserts in both
halves, and the privileged half is worth more than the fix: a present,
bindable interface binding inline is the ordinary case on a booted router, and
no other unit test reaches it, because every other one here names an interface
that does not exist. The bind-success path had no unit coverage at all.

`an_interface_with_no_addresses_is_still_present` returned early when the
fixture was absent — no fixture, pass. It still has to skip on a machine with
no address-less interface, so the guard is the runner declaring that it has
fixtures: CI now sets FIPS_TEST_REQUIRE_FIXTURES beside
FIPS_TEST_ADDRLESS_IFACE, and the test fails rather than skips if the fixture
step is ever removed or renamed. Its `let _ = interface_carrier(...)` is now
asserted too: if presence and carrier ever collapsed into one read, a
carrier-less bridge would report absent and the whole IFF_UP-not-IFF_RUNNING
decision would be silently undone.

`policy_labels` checked `Required.as_str()` and not `Optional.as_str()`, so a
swapped pair would paint every expected interface as the tolerated kind and
stay green. `Presence::as_str` had no test at all — `binding` was never
observed by anything, anywhere.

Three binder branches had no coverage: a transport restarting (the second
`start_async` clearing the previous run's stop flag — only a second *stop* was
tested, so a transport that could never restart passed everything), the
episode clock being restamped at start rather than at construction, and the
refused-edge retry actually delivering. The last one matters most: the
existing test filled the channel and dropped the receiver, so a slot that
captured an edge and never re-sent it would pass while health sat on a stale
level forever.

And the hardware-change boundary `record_bind` returns, which the neighbour
flush hangs off: false on a first bind (or every clean start would drop a
cache it had just built) and true once, not stickily, on a MAC change.

Each new test was verified against the defect it guards — comment out the
`shutdown.store(false)`, the `mark_starting()`, or the seeded `unpublished`,
and the corresponding test goes red while the rest stay green.

fix(test): correct the dummy-carrier assertion, and pin Darwin's presence probe

The carrier assertion added in d6240698 was wrong and would have failed both
Linux legs. A `dummy` interface brought up reports `<BROADCAST,NOARP,UP,
LOWER_UP>` — `IFF_RUNNING` is set, so it *has* carrier. Verified against the
exact fixture CI builds, `addrgenmode none` and all, rather than against the
comment: the original code discarded the result and its comment claimed "up
but not running", which is what made asserting it look safe.

So the fixture pins address-less *presence* and cannot demonstrate the
presence-vs-carrier split at all — an interface up with no carrier is a bridge
with nothing plugged in, which no fixture here creates.
`carrier_is_reported_separately_from_presence` pins that split from the other
side. The corrected assertion is Linux-only, because the expected answer is a
property of the fixture device rather than of the code.

That is also what lets the macOS fixture land. The Linux legs pin the
address-less contract on glibc and musl, but the BSD-derived `getifaddrs` the
macOS backend actually calls had no coverage — the test skipped itself
silently on that runner, which is precisely the shape the previous commit was
removing. `feth` is macOS's fake-Ethernet pseudo-interface and is created
address-less; the step fails the leg rather than testing the wrong thing if the
runner hands it an address anyway, mirroring why the Linux step needs
`addrgenmode none`.

Both branches of the fixture guard were exercised: unset skips and passes, and
declared-but-missing fails loudly. The corrected test was run against a real
Linux dummy inside a container, not reasoned about.

fix(ci): put the macOS presence fixture on the macOS job

The `feth` fixture step added in 6b9faa0c landed on the Linux `test` job, not
on `test-macos`, and has failed CI ever since:

    create: Host name lookup failure
    ifconfig: `--help' gives usage information.

That is Linux net-tools `ifconfig`, which has no `create` subcommand. So the
Linux job ran two address-less-fixture steps — its own correct `ip link add
type dummy` one, then a macOS one that cannot work there — while `test-macos`
had none at all, leaving the Darwin `getifaddrs` path exactly as uncovered as
before.

Cause was a pattern-anchored edit: `Install cargo-nextest` followed by `Run
unit tests` appears in three jobs, and the insert hit the first match. Moving
it then hit the *last* match, which is `test-windows`. It is now placed by job
boundary rather than by pattern, and verified per job: `test` and `test-musl`
carry the `ip link`/dummy fixture, `test-macos` carries `ifconfig`/`feth`, and
`feth` appears exactly once in the file.

The verification that missed this was counting steps in `test-macos` and
reading 6 as confirmation. Six was the count *before* the insert; seven is what
a successful insert looks like. Now asserted by job and by which tool each
fixture step uses, so a step in the wrong place fails the check rather than
matching a total.

test(transport): cover the detach branch and the stop-race check

Two of the three untested branches, by two different routes.

**The detach decision is now a pure function.** `classify_detach(gone,
replaced, dead)` replaces the inline three-way `if` in the binder loop, and all
eight input combinations are asserted, plus the precedence between them and the
`reason=` labels the integration suite and operators grep for.

This does not make `Replaced` or `SocketDied` reachable from a test — both need
a bind that succeeded and then a specific external event, and the integration
suite cannot arrange the recreate deterministically either, for the reason
recorded in reference/notes.md: the netlink event from a delete is acted on
within microseconds, so `Gone` wins that race in practice. What it does is
split the untested thing in two. The three inputs each already had tests
(`interface_present`, `device_replaced`, `tasks_alive`); the branch between them
did not, and that half is now total and exhaustive.

Precedence is asserted rather than assumed: an interface that has gone away has
also trivially been "replaced" and its socket is also dead, so the most
specific true statement has to win, and callers pass `replaced` already masked
by `!gone` — the function no longer depends on them having masked correctly.

**The post-store shutdown check is now tested directly.** It needs a bind that
*succeeds*, which is why `a_stop_racing_a_bind_leaves_nothing_behind` could
never reach it: that test's interface does not exist, so `bind_and_spawn`
refuses at the presence probe several steps earlier. Binding loopback as root
reaches it, and the assertion is that a bind completing after a stop undoes its
own socket, its own loops, and its own presence.

Unprivileged runners skip it, but loudly: a runner declaring
`FIPS_TEST_PRIVILEGED` and unable to open a raw socket fails instead of
skipping, the same guard shape as `FIPS_TEST_REQUIRE_FIXTURES`. No CI leg sets
that yet — unit tests run unprivileged — so today it exercises the branch only
under a root container. Verified there, including the negative control: with
the post-store check removed the test fails, and with it present it passes.

test(transport): run the bind-success half in CI, and cover the replaced device

Adds the privileged unit-test step the rest of this depends on, then uses it.

**The privileged step.** Every unit-test leg runs unprivileged, so
`PacketSocket::open` cannot succeed on any of them and everything past a
successful bind runs nowhere in CI: the post-store shutdown check, the
`Present` arm of the binder loop, `bind_now` itself. The step builds as the
runner user and executes only the test binary under sudo — running `cargo` as
root would use root's CARGO_HOME and discard the cache the job just restored.
The binary-path extraction was verified locally before being written into the
workflow.

`FIPS_TEST_PRIVILEGED` is what makes it honest. Tests needing a raw socket skip
quietly without it; with it set, a test that cannot open one fails and says so.
A runner that stops granting the capability shows up as a red leg rather than
as silence.

**The replaced device.** `"interface replaced"` — a netdev recreated under the
same name, the `wifi reload` case in #125 — could not be reached by the
integration suite: with link events live the kernel's `RTM_DELLINK` is acted on
within microseconds, so the binder observes "gone" first and takes the branch
already covered. Attempting it there passed about one run in three.

Rather than race the binder, the test asserts the *inputs*: after a real
delete-and-recreate the name still resolves and the bound index no longer
matches. Paired with the exhaustive classifier test, which pins that
`(gone: false, replaced: true)` maps to `Replaced`, the path is covered without
depending on scheduling. A watcher-disable hook was written for the racing
approach and removed once this one made it unnecessary.

**Two smaller gaps.** `report_sustained_absence` had no direct test — the
integration suite infers it from counting ERROR lines, and ten seconds of real
time is how a deadline ends up asserted by proxy. A test-only `backdate_for_test`
ages the episode clock instead. And the probe floor was asserted only as a
relation between two constants; it now has behaviour, via a `probe_delay`
helper extracted from the loop.

That extraction was not neutral, and its own test caught it: `checked_sub`
yields `Some(0)` at exact equality where the loop used a strict `<`, so a
zero-length sleep would have replaced no sleep at all. Harmless in effect,
wrong in meaning, and fixed.

Every new test was run against the defect it guards, in a root container:
remove the post-store shutdown check, or make `device_replaced` always answer
false, and the corresponding test fails.
Mark publish_presence must_use. The defect this commit fixes was a discarded
return value, so the fix is made self-guarding: a future call site that drops
the pending edge instead of storing it now fails the lint rather than silently
reintroducing the bug. Option is not must_use in std, unlike Result, so the
attribute has to be explicit.
2026-09-10 19:18:09 +00:00

80 KiB
Raw Blame History

FIPS Configuration

FIPS uses YAML-based configuration with a cascading multi-file priority system. All parameters have sensible defaults; a node can run with no configuration file at all (it will generate an ephemeral identity and listen on default addresses).

Configuration Loading

Search Paths

When started without the -c flag, FIPS searches for fips.yaml in these locations, lowest to highest priority:

Priority Path Purpose
1 (lowest) /usr/local/etc/fips/fips.yaml (macOS, FreeBSD), /etc/fips/fips.yaml (other Unix) System-wide defaults
2 ~/.config/fips/fips.yaml User preferences
3 ~/.fips.yaml Legacy user config
4 (highest) ./fips.yaml Deployment-specific overrides

All found files are loaded and merged in priority order. Values from higher priority files override those from lower priority files. This allows a system administrator to set site-wide defaults in the priority 1 path above, /usr/local/etc/fips/fips.yaml on macOS and FreeBSD and /etc/fips/fips.yaml on other Unix systems, while individual deployments override specific values in ./fips.yaml.

On macOS and FreeBSD both directories are probed: /etc/fips first, then /usr/local/etc/fips, so the packaged file wins over a leftover /etc/fips copy from an earlier install.

CLI Option

fips -c /path/to/config.yaml

When -c is specified, only that file is loaded (search paths are skipped).

Partial Configuration

Every field has a built-in default. A configuration file only needs to specify values that differ from defaults. For example, a minimal config might contain only the identity and peer list, inheriting all other defaults.

YAML Structure

The configuration is organized into six top-level sections (gateway: is Linux only):

node:        # Node behavior, protocol parameters, and tuning
tun:         # TUN virtual interface
dns:         # DNS responder for .fips domain
transports:  # Network transports (UDP, Ethernet, Bluetooth, Tor, ...)
peers:       # Static peer list
gateway:     # LAN gateway service (Linux only)

Control Socket (node.control.*)

Parameter Type Default Description
node.control.enabled bool true Enable the control socket
node.control.socket_path string (auto) Unix: Socket file path. Resolution is shared by daemon and clients: /run/fips/control.sock when /run/fips exists; then /var/run/fips/control.sock on macOS/FreeBSD when its private directory exists; then $XDG_RUNTIME_DIR/fips/control.sock; finally /tmp/fips-control.sock. A privileged macOS daemon selects /var/run/fips/control.sock even when the private directory must be created after boot. Windows: TCP port number (default: 21210); the control socket listens on 127.0.0.1 at this port.

The control socket provides access to node state and runtime management via the fipsctl command-line tool. In addition to read-only status queries, fipsctl connect and fipsctl disconnect enable runtime peer management. See the fipsctl reference for the command list.

On Unix, the control socket is a Unix domain socket with filesystem permissions (mode 0770, group fips). On Windows, it is a TCP listener on localhost. TCP does not provide filesystem-level ACLs, so any local user can connect to the control port.

Security note (Windows): The TCP control socket on Windows is a known limitation. Any process running on the local machine can connect to the control port and issue commands, including disconnect, connect, and inject-config. This is acceptable for single-user workstations but may be inappropriate for shared machines. Future improvements may include named pipe support (with Windows ACLs) or an authentication token mechanism. On shared Windows systems, consider using firewall rules to restrict access to the control port.

All tunable protocol parameters live under node.*, organized as sysctl-style dotted paths. The top-level sections (tun, dns, transports, peers) handle infrastructure concerns only.

Node Parameters (node.*)

Identity (node.identity.*)

Parameter Type Default Description
node.identity.nsec string (none) Secret key in nsec (bech32) or hex format. If omitted, behavior depends on persistent.
node.identity.persistent bool false Persist identity across restarts via key file.

Identity resolution follows a three-tier priority:

  1. Explicit nsec in config — always used when present, regardless of persistent
  2. Persistent key file — when persistent: true and no nsec, loads from fips.key adjacent to the config file; if no key file exists, generates a new keypair and saves it
  3. Ephemeral — when persistent: false (default) and no nsec, generates a fresh keypair on each start

Key files (fips.key with mode 0600, fips.pub with mode 0644) are written adjacent to the highest-priority config file for operator visibility, even in ephemeral mode.

General

Parameter Type Default Description
node.leaf_only bool false Leaf-only mode: node does not forward traffic or participate in routing
node.tick_interval_secs u64 1 Periodic maintenance tick interval (retry checks, timeout cleanup, tree refresh)
node.base_rtt_ms u64 100 Initial RTT estimate for new links before measurements converge
node.heartbeat_interval_secs u64 10 Heartbeat send interval per peer for liveness detection
node.link_dead_timeout_secs u64 30 No-traffic timeout before a peer is declared dead and removed
node.drain_timeout_secs u64 2 Upper bound in seconds on the Draining shutdown phase. On shutdown the node broadcasts Disconnect to its peers and then waits up to this long for the links to clear, exiting as soon as the last peer is gone. 0 skips the wait. The key is absent from a default config file rather than written with its default value, so an unset key and the 2-second default are the same thing
node.log_level string "info" Tracing filter default. Case-insensitive; one of trace, debug, info, warn, error. Overridden by the RUST_LOG environment variable when set

Resource Limits (node.limits.*)

Controls capacity for connections, peers, and links.

Parameter Type Default Description
node.limits.max_connections usize 256 Max handshake-phase connections
node.limits.max_peers usize 128 Max authenticated peers
node.limits.max_links usize 256 Max active links
node.limits.max_pending_inbound usize 1000 Max pending inbound handshakes

Rate Limiting (node.rate_limit.*)

Handshake rate limiting protects against DoS on the Noise IK handshake path.

Parameter Type Default Description
node.rate_limit.handshake_burst u32 100 Token bucket burst capacity
node.rate_limit.handshake_rate f64 10.0 Tokens per second refill rate
node.rate_limit.handshake_timeout_secs u64 30 Stale handshake cleanup timeout
node.rate_limit.handshake_resend_interval_ms u64 1000 Initial handshake message resend interval
node.rate_limit.handshake_resend_backoff f64 2.0 Resend backoff multiplier (1s, 2s, 4s, 8s, 16s with defaults)
node.rate_limit.handshake_max_resends u32 5 Max resends per handshake attempt
node.rate_limit.established_handshake_burst u32 derived Burst capacity of the established-link bucket. Derived default is node.limits.max_peers (128)
node.rate_limit.established_handshake_rate f64 derived Refill rate of that bucket. Derived default is (max_peers / max(node.rekey.after_secs, 1)) * (1 + handshake_max_resends), floored at 1.0/s — 6.4/s at shipped defaults
node.rate_limit.session_setup_burst u32 64 Per-link-peer burst for inbound session-setup messages that would open a new session
node.rate_limit.session_setup_rate f64 16.0 Per-link-peer refill rate for those messages, in tokens per second

Msg1 whose source matches an established link (rekey and restart maintenance traffic) draws on a second bucket rather than competing with stranger admission. Both keys are optional; leaving them unset keeps the derived sizing, which tracks max_peers and the rekey period automatically instead of becoming a constant nobody revisits. max_peers: 0 (unlimited) has no peer-count-derived size, so the derivation falls back to handshake_burst / handshake_rate.

The node's total admitted msg1 rate is the sum of the two buckets: 228 burst and 16.4/s at shipped defaults, of which the established half is reachable only by a source that already matches a live link. Size against the sum when budgeting handshake crypto load for a host.

The session_setup_* pair is a separate limiter on the session layer, not the link layer. It is keyed on the authenticated link peer a session datagram arrived over, so each neighbour gets its own budget and a flood is attributable. Setup messages naming a peer this node is already established with (inbound rekey and restart traffic) draw on a second per-link bucket derived from max_peers, node.rekey.after_secs and handshake_max_resends, exactly as established_handshake_* is, so a stranger flood cannot suppress rekey traffic sharing the link.

At the defaults one neighbour can force at most session_setup_rate * handshake_timeout_secs half-open entries (480) and session_setup_rate * (1 + handshake_max_resends) acks per second (96). The node-wide ceiling is still that times the peer count, since the limiter bounds each neighbour rather than the aggregate. A legitimate peer whose traffic reaches this node over the same link as an attacker's shares that attacker's stranger bucket, so establishment behind a flooded neighbour is refused until the bucket refills; the initiator's own resend schedule (1s, 2s, 4s, 8s, 16s) covers a short drain.

Retry / Backoff (node.retry.*)

Connection retry with exponential backoff.

Parameter Type Default Description
node.retry.max_retries u32 5 Max connection retry attempts
node.retry.base_interval_secs u64 5 Base backoff interval
node.retry.max_backoff_secs u64 300 Cap on exponential backoff (5 minutes)

Auto-reconnect (triggered by MMP link-dead removal) uses the same backoff parameters but bypasses max_retries, retrying indefinitely. See peers[].auto_reconnect below.

Medium-Change Detection (node.netmon.*)

Detects that the host moved between transport media — WLAN to LAN, WLAN to 5G, an interface arriving or leaving — and rebinds the send path immediately.

Parameter Type Default Description
node.netmon.enabled bool true Whether medium-change detection runs
node.netmon.poll_interval_secs u64 5 How often the path to each peer is sampled (backstop period where an event-driven backend exists). Must be at least 1 while enabled; 0 is refused at startup.
node.netmon.debounce_ms u64 250 How long to wait for the picture to settle before acting (0 disables). Refused at startup when debounce_ms × 8 reaches node.link_dead_timeout_secs — see below.

Both refusals stop the node rather than degrade it, so they are worth knowing before they are met.

A handover is ridden out for up to 8 settling rounds (MAX_DEBOUNCE_ROUNDS) of debounce_ms each before a change is reported. If that worst case reaches node.link_dead_timeout_secs, the liveness reaper tears the peering down before the detector ever reports, so the machinery runs and cannot help — the node refuses to start rather than run in that shape. At the shipped defaults the margin is wide (8 × 250 ms = 2 s against 30 s), but the constraint couples two keys in different blocks: shortening link_dead_timeout_secs for fast failover can make an untouched debounce_ms illegal. The refusal names both values and the multiplier.

Established UDP peers use a per-peer connect()-ed socket for the send fast path. connect(2) makes the kernel resolve the route once and pin the local source address to whichever interface carried it then; it never re-evaluates. Without detection, a medium change therefore leaves every peer transmitting from an abandoned address while the peer answers where it last heard the node — the peering reports itself connected and carries nothing until link_dead_timeout_secs tears it down, typically 60–90s per switch.

On a detected change the node drops the sockets of the peers the change names (the wildcard listen socket resolves a route per packet, so sends keep working, and a correctly bound connected socket is reinstalled on a later tick) and heartbeats those of them on a connectionless transport at once so the far side re-pins to the new source address. A peer reached over TCP, Tor, Nym or BLE is left to its periodic heartbeat, since sending to it here would block the node's receive loop on a stream the medium change has very likely just stranded. A peer the change does not name is left alone entirely. No peering is torn down: sessions, tree positions and routes survive the switch.

What counts as a change. For each peer whose transport address is a numeric IP endpoint, the node asks the kernel which local address it would use to reach that peer — a connect(2) on a UDP socket, which resolves the route and sends nothing. A change is reported when a peer present in two consecutive samples is now reached from a different local address, or has stopped being reachable at all.

Because the question is asked per peer, an interface the node does not peer over cannot trigger anything: a container bridge, a VPN, a veth pair or a tunnel appearing is not the route to any peer, so it does not enter the sample. A peer on the same LAN, reached by its subnet route rather than the default route, is covered as well as one across the internet, and so is a more specific route moving under a single peer.

The converse is the residual. The probe answers for the peer's current address, and that address is the source of the last authentic packet it sent, so a peer that roams between two of its own addresses which leave this host by different interfaces is indistinguishable from a local path move. The reaction is scoped to the peers named in the change, so such a peer moves nothing but its own send path — but it is the peer, not this host, that decided the fingerprint changed.

Peers appearing and leaving are ignored on their own — that is ordinary node behaviour and says nothing about the medium. A peer seen for the first time is the one exception, and it is not judged against history but against its own send path: if its connect()-ed socket is pinned to a source the routing table would no longer choose, it is reported. Without that, a medium change in the window between a peer authenticating and the detector's next sample would be the detector's first sight of that peer, and would be adopted silently while the peer's socket stayed pinned to the path the host had just left. A peer joining onto a path that has not moved has its socket pinned exactly where its traffic goes, so it still reports nothing. A peer whose address is not a probeable IP endpoint contributes nothing: a MAC on Ethernet or BLE, a .onion or Nym recipient reached through a local proxy, an IPv6 literal with a scope suffix, or a peer still carrying the hostname it was configured with (resolving one would put a DNS lookup on the sample path; the address becomes numeric as soon as an authenticated packet arrives from the peer). A node holding no peers detects nothing, which is correct — it has nothing bound to the old path.

The cost is five non-blocking syscalls per peer per sample, read from the probe's own code rather than measured: socket(2) and bind(2), a connect(2) that sends no packet, a getsockname(2), and the close(2) the socket takes on drop. Nothing goes on the wire and no name is resolved. node.limits.max_peers bounds the per-sample total only where it is set: at max_peers: 0, which means unlimited, there is no bound and the cost tracks the live peer count instead.

Detection uses the best backend the platform has:

Platform Backend Latency
Linux, Android NETLINK_ROUTE multicast (as ip monitor) kernel event, milliseconds
macOS, FreeBSD PF_ROUTE socket kernel event, milliseconds
Windows, iOS timer up to poll_interval_secs

Where a backend exists, poll_interval_secs is only a backstop: a netlink socket drops messages under memory pressure and the subscription can fail to start in a restricted sandbox, so the timer keeps running underneath. A backend that cannot start is logged once at warn and the node falls back to the timer.

None of this covers Bluetooth: a BLE adapter's state is not an IP attachment and is invisible to this detector. The connected-socket fast path is Linux and macOS only; elsewhere there are no pinned sockets to rebind, and the heartbeat alone carries the new address.

Cache Parameters (node.cache.*)

Controls caching of tree coordinates and identity mappings.

Parameter Type Default Description
node.cache.coord_size usize 50000 Max entries in coordinate cache
node.cache.coord_ttl_secs u64 300 Coordinate cache entry TTL (5 minutes)
node.cache.identity_size usize 10000 Max entries in identity cache (LRU, no TTL)

Mesh Lookup (node.lookup.*)

Controls bloom-guided mesh lookup (LookupRequest/LookupResponse): finding the current coordinates of a mesh address the node already knows.

Renamed in v0.5.0. These six keys were node.discovery.*. See Deprecated keys for the full mapping. A deployed node.discovery: block still loads and still applies, with a one-time deprecation warning at startup.

Parameter Type Default Description
node.lookup.ttl u8 64 Hop limit for LookupRequest forwarding
node.lookup.attempt_timeouts_secs array<u64> [1, 2, 4, 8] Per-attempt timeouts. Each entry is the deadline for one LookupRequest before sending the next attempt with a fresh request_id. Length determines total attempt count; default gives 4 attempts and a 15s total budget
node.lookup.recent_expiry_secs u64 10 Dedup cache expiry for recent request IDs
node.lookup.backoff_base_secs u64 0 Optional post-failure suppression base in seconds; doubles per consecutive failure. 0 disables (default); the per-attempt sequence is the only retry pacing
node.lookup.backoff_max_secs u64 0 Cap on optional post-failure backoff
node.lookup.forward_min_interval_secs u64 2 Transit-side rate limiting: minimum interval between forwarded lookups for the same target

Peer Rendezvous (node.rendezvous.*)

How the node finds peers to connect to at all, over the Nostr overlay and on the local link. Distinct from mesh lookup above, which resolves coordinates for a mesh address that is already known.

Renamed in v0.5.0. node.discovery.nostr.* is now node.rendezvous.nostr.*, and node.discovery.lan.* is now node.rendezvous.lan.*. See Deprecated keys.

Nostr Rendezvous (node.rendezvous.nostr.*)

Optional Nostr-mediated overlay rendezvous. This layer publishes replaceable endpoint adverts (fips-overlay-v1), consumes advert-derived endpoint fallbacks for configured peers, and can optionally discover non-configured peers (policy: open). udp:nat remains the trigger for NAT traversal offer/answer + punch-through, after which the established UDP socket is handed into the normal FIPS transport/session stack. Inbox-relay discovery falls back to the local DM relay list if remote relay metadata cannot be fetched. The Nostr discovery runtime is compiled into every build of the crate; it is enabled at runtime via node.rendezvous.nostr.enabled: true and stays inert otherwise.

Parameter Type Default Description
node.rendezvous.nostr.enabled bool false Enable Nostr-mediated overlay discovery
node.rendezvous.nostr.policy string "configured_only" Advert discovery policy: disabled, configured_only, open
node.rendezvous.nostr.open_discovery_max_pending usize 64 Max open-discovery peers queued in outbound retry/connection state at once
node.rendezvous.nostr.max_concurrent_incoming_offers usize 16 Max concurrent inbound traversal offers processed at once (rate limit against offer spam)
node.rendezvous.nostr.max_concurrent_offers_per_npub usize 4 Max concurrent inbound traversal offers accepted from any one sender npub, so a single identity cannot hold the whole pool. Sits inside max_concurrent_incoming_offers, which stays the outer bound; a larger value is inert. Zero is rejected, since it refuses every inbound offer rather than disabling the limit
node.rendezvous.nostr.advert_cache_max_entries usize 2048 Max cached overlay adverts retained from relay traffic
node.rendezvous.nostr.seen_sessions_max_entries usize 2048 Max seen-session IDs retained for replay detection
node.rendezvous.nostr.advertise bool true Publish local endpoint adverts
node.rendezvous.nostr.advert_relays list[string] ["wss://relay.damus.io", "wss://nos.lol", "wss://offchain.pub"] Relays used for service adverts
node.rendezvous.nostr.dm_relays list[string] ["wss://relay.damus.io", "wss://nos.lol", "wss://offchain.pub"] Relays used for encrypted signaling events
node.rendezvous.nostr.stun_servers list[string] ["stun:stun.l.google.com:19302", "stun:stun.cloudflare.com:3478", "stun:global.stun.twilio.com:3478"] STUN servers used for local reflexive address discovery
node.rendezvous.nostr.share_local_candidates bool false Whether to advertise local (RFC 1918 / ULA) interface addresses as host candidates in the traversal offer. Off by default: in most deployments peers aren't on the same broadcast domain, and sharing private host candidates causes misleading punch successes when an asymmetric L3 path (VPN, Tailscale subnet route, overlapping address space) makes a peer's private IP one-way reachable. Enable only when peers are on the same physical LAN
node.rendezvous.nostr.app string "fips-overlay-v1" Traversal application namespace, published in the advert's protocol tag (the d tag itself is hardcoded to fips-overlay-v1)
node.rendezvous.nostr.signal_ttl_secs u64 120 Signaling TTL in seconds
node.rendezvous.nostr.attempt_timeout_secs u64 10 Overall traversal attempt timeout in seconds
node.rendezvous.nostr.replay_window_secs u64 300 Replay tracking retention window in seconds
node.rendezvous.nostr.punch_start_delay_ms u64 2000 Delay before punch traffic starts
node.rendezvous.nostr.punch_interval_ms u64 200 Interval between punch packets
node.rendezvous.nostr.punch_duration_ms u64 10000 How long to keep punching before failure
node.rendezvous.nostr.advert_ttl_secs u64 3600 Advert TTL in seconds
node.rendezvous.nostr.advert_refresh_secs u64 1800 How often adverts are refreshed in seconds
node.rendezvous.nostr.startup_sweep_delay_secs u64 5 Settle delay after Nostr discovery starts before the one-shot startup advert sweep runs (only used under policy: open). Allows the relay subscription backlog to populate the in-memory advert cache before the sweep fires
node.rendezvous.nostr.startup_sweep_max_age_secs u64 3600 Maximum advert age (now - created_at) considered by the one-shot startup sweep (only used under policy: open). Adverts older than this are skipped on startup; the per-tick sweep still considers them up to valid_until_ms
node.rendezvous.nostr.failure_streak_threshold u32 5 Consecutive NAT-traversal failures against a peer before an extended cooldown is applied. At this threshold the daemon also actively re-fetches the peer's advert from advert_relays to evict cache entries for peers that have gone away
node.rendezvous.nostr.extended_cooldown_secs u64 1800 Cooldown applied to a peer once failure_streak_threshold is hit. Suppresses both open-discovery sweep enqueues and per-attempt retry firings until elapsed (30 minutes default)
node.rendezvous.nostr.warn_log_interval_secs u64 300 Minimum interval between NAT traversal failed WARN log lines for the same peer. Subsequent failures inside the window log at DEBUG to reduce log spam on public-test nodes with many cache-learned peers
node.rendezvous.nostr.failure_state_max_entries usize 4096 Maximum entries retained in the per-npub failure-state map. Bounds memory under high cache turnover; oldest entries (by last failure time) are evicted when the cap is exceeded
node.rendezvous.nostr.protocol_mismatch_cooldown_secs u64 86400 Cooldown applied after observing a fatal protocol mismatch on a Nostr-adopted bootstrap transport (e.g. Unknown FMP version from a peer running a different FMP-protocol version). Independent of extended_cooldown_secs and much longer (24 hours default) because the mismatch is structural; re-traversing is wasted effort until one side upgrades

If stun_servers is omitted, the built-in default list above is used. If it is specified in YAML, the configured list fully overrides the defaults. Initiators use only this local list for outbound STUN queries; peer-advertised STUN values are published for diagnostics/interoperability but are not used as arbitrary egress targets. The built-in advert and DM relay defaults point at widely-operated public relays (Damus, nos.lol, Primal) as best-effort endpoints; operators are encouraged to override them with their own relay preferences for production deployments. Advert freshness is enforced semantically: events with expired NIP-40 expiration tags are dropped, and adverts are also bounded by a created-at staleness window derived from advert_ttl_secs (with a grace multiplier). The current in-tree STUN parser handles IPv4 and IPv6 mapped-address attributes. Local traversal candidates include active non-loopback private interface addresses (RFC1918 IPv4 and IPv6 ULA) plus probed local egress addresses for the punch socket port. During punching, compatible private-subnet candidates and reflexive candidates are attempted in parallel; the first successful path wins.

LAN Rendezvous (node.rendezvous.lan.*)

Peer rendezvous on the local link via mDNS / DNS-SD (RFC 6762 / RFC 6763). When enabled, the node publishes a _fips._udp.local. service advert carrying its npub (and optional scope) and concurrently browses for the same service type to learn same-broadcast-domain peers. The result is sub-second peer pairing with no Nostr-relay roundtrip, STUN observation, or NAT traversal: the observed endpoint is by construction routable from the consumer's LAN.

mDNS adverts are unauthenticated, so a LAN advert is treated only as a routing hint. Identity is still proven end-to-end by the Noise IK handshake the node initiates against the observed endpoint; a spoofed advert carrying another peer's npub fails the handshake and is dropped. LAN discovery requires an active UDP transport (peers dial the advertised UDP port to begin the handshake).

Parameter Type Default Description
node.rendezvous.lan.enabled bool false Master switch. Opt-in: enable for sub-second same-LAN pairing. Default-off avoids reintroducing a per-LAN identity broadcast on nodes that have deliberately disabled other discovery channels
node.rendezvous.lan.service_type string "_fips._udp.local." DNS-SD service type. Primarily an override for integration tests running multiple isolated services on one loopback interface; leave at the default in production
node.rendezvous.lan.scope string (none) Optional application/network scope carried in a scope=<name> TXT entry. Browsers with a scope set only surface peers advertising the same scope, so nodes on the same physical LAN configured for different mesh networks do not cross-feed. Intentionally separate from node.rendezvous.nostr.app so relay-visible adverts can stay generic while LAN discovery is isolated per private network

Spanning Tree (node.tree.*)

Controls tree construction and parent selection.

Parameter Type Default Description
node.tree.announce_min_interval_ms u64 500 Per-peer TreeAnnounce rate limit
node.tree.parent_hysteresis f64 0.2 Cost improvement fraction required for same-root parent switch (0.0–1.0)
node.tree.hold_down_secs u64 30 Suppress non-mandatory re-evaluation after parent switch
node.tree.reeval_interval_secs u64 60 Periodic cost-based parent re-evaluation interval (0 = disabled)
node.tree.flap_threshold u32 4 Parent switches in window before dampening engages
node.tree.flap_window_secs u64 60 Sliding window for counting parent switches
node.tree.flap_dampening_secs u64 120 Extended hold-down duration when flap threshold exceeded

Bloom Filter (node.bloom.*)

Parameter Type Default Description
node.bloom.update_debounce_ms u64 500 Debounce interval for filter update propagation
node.bloom.max_inbound_fpr f64 0.20 Antipoison cap: reject inbound FilterAnnounce frames whose advertised false-positive rate exceeds this value. Valid range (0.0, 1.0). The default 0.20 corresponds to fill 0.7248 at k=5 (≈2,114 entries on the 1 KB filter); a saturated/poisoned filter is still ~100% FPR and rejected

Bloom filter size (1 KB), hash count (5), and size classes are protocol constants and not configurable.

ECN Signaling (node.ecn.*)

Controls hop-by-hop ECN (Explicit Congestion Notification) signaling. When enabled, transit nodes detect congestion on outgoing links (via MMP loss/ETX metrics or kernel buffer drops) and set the CE flag on forwarded FMP frames. Destination nodes mark ECN-capable IPv6 packets with CE before TUN delivery per RFC 3168, enabling end-host TCP congestion control to react.

Parameter Type Default Description
node.ecn.enabled bool true Enable ECN congestion signaling (CE flag relay and local congestion detection)
node.ecn.loss_threshold f64 0.05 MMP loss rate threshold for CE marking (0.0–1.0). When the outgoing link's loss rate meets or exceeds this value, forwarded packets are CE-marked.
node.ecn.etx_threshold f64 3.0 MMP ETX threshold for CE marking (≥1.0). When the outgoing link's ETX meets or exceeds this value, forwarded packets are CE-marked.

Congestion detection triggers on any of: outgoing link loss ≥ loss_threshold, outgoing link ETX ≥ etx_threshold, or kernel receive buffer drops detected on any local transport. CE is relayed hop-by-hop: once set on any hop, the flag stays set for all subsequent hops to the destination.

Rekey (node.rekey.*)

Controls periodic Noise rekey for forward secrecy. When enabled, both FMP (link-layer IK) and FSP (session-layer XK) sessions perform fresh Diffie-Hellman key exchanges after a time or message count threshold, whichever comes first. A 10-second drain window keeps the old session active for decryption during cutover.

Parameter Type Default Description
node.rekey.enabled bool true Initiate periodic Noise rekey on links and sessions. A peer-driven session rekey is still answered when this is off, so session keys can still rotate
node.rekey.after_secs u64 120 Initiate rekey after this many seconds on a session
node.rekey.after_messages u64 65536 Initiate rekey after this many messages sent on a session

Session / Data Plane (node.session.*)

Controls end-to-end session behavior and packet queuing.

Parameter Type Default Description
node.session.default_ttl u8 64 Default SessionDatagram TTL
node.session.pending_packets_per_dest usize 16 Queue depth per destination during session establishment
node.session.pending_max_destinations usize 256 Max destinations with pending packets
node.session.idle_timeout_secs u64 90 Idle session timeout; established sessions with no application data for this duration are removed. MMP reports (SenderReport, ReceiverReport, PathMtuNotification) do not count as activity
node.session.coords_warmup_packets u8 5 Number of initial data packets per session that include the CP flag for transit cache warmup; also the reset count on CoordsRequired/PathBroken receipt
node.session.coords_response_interval_ms u64 2000 Minimum interval (ms) between standalone CoordsWarmup responses to CoordsRequired/PathBroken signals per destination

The anti-replay window size (2048 packets) is a compile-time constant and not configurable.

Metrics Measurement Protocol for per-peer link measurement. See ../design/fips-mesh-layer.md for behavioral details.

Parameter Type Default Description
node.mmp.mode string "full" Operating mode: full (sender + receiver reports), lightweight (receiver reports only), or minimal (spin bit + CE echo only, no reports)
node.mmp.log_interval_secs u64 30 Periodic operator log interval for link metrics
node.mmp.owd_window_size usize 32 One-way delay trend ring buffer size

Session-Layer MMP (node.session_mmp.*)

Metrics Measurement Protocol for end-to-end session measurement. Configured independently from link-layer MMP because session reports are routed through every transit link, consuming bandwidth proportional to path length.

Parameter Type Default Description
node.session_mmp.mode string "full" Operating mode: full, lightweight, or minimal
node.session_mmp.log_interval_secs u64 30 Periodic operator log interval for session metrics
node.session_mmp.owd_window_size usize 32 One-way delay trend ring buffer size

Internal Buffers (node.buffers.*)

Channel sizes affecting throughput and memory. Primarily useful for performance tuning under high load or on memory-constrained devices.

Parameter Type Default Description
node.buffers.packet_channel usize 1024 Transport to Node packet channel capacity
node.buffers.tun_channel usize 1024 TUN to Node outbound channel capacity
node.buffers.dns_channel usize 64 DNS to Node identity channel capacity

Native Datagram API (node.native_api.*)

Experimental, off by default, and built on Linux, FreeBSD and macOS only. A client process connects to a Unix socket and asks either to open a flow to a remote pubkey or to hold a local port. Both answers carry a file descriptor: a flow's, which the client sends and receives datagrams on, or a listener's, which arriving flows are delivered on. There is no IPv6 emulation and no TUN device on this path.

The surface is not stable, is not a reliability layer, and is not the v2 external process API. No compatibility promise is made about it: the keys below, the line protocol behind them, and the Rust client that hides it may change or be withdrawn in any release.

The listener is not built on macOS or Windows, and this section is ignored there. Two separate things bound that: Windows has no SCM_RIGHTS and so no way to hand a file descriptor to another process at all, while macOS has SCM_RIGHTS but does not implement SOCK_SEQPACKET for AF_UNIX.

Parameter Type Default Description
node.native_api.enabled bool false Enable the native API socket
node.native_api.socket_path string (auto) Socket file path. Resolved the same way as the control socket, with the filename api.sock: /run/fips/api.sock when /run/fips exists; then /var/run/fips/api.sock on FreeBSD when its private directory exists; then $XDG_RUNTIME_DIR/fips/api.sock; finally /tmp/fips-api.sock
node.native_api.pending_per_flow usize 16 Datagrams held for one flow while it waits to be accepted, or while an established flow's client is slow to read. Refused above 64 at startup: the whole batch is written onto a socket pair the client cannot read yet. Refused below 1: a flow that can hold nothing loses its peer's opening datagram between the arrival being announced and the client taking the flow
node.native_api.backlog usize 16 Flows announced on one listener and not yet taken by its task. Refused below 1 at startup: a listener with no backlog admits no flow, so every arrival would be dropped
node.native_api.max_flows usize 256 Flows this node holds at once
node.native_api.debug_commands bool false Answer the inject, stats and arrive debug commands. Not a supported interface

Security note: the socket is mode 0770, owned by group fips, and that is the whole of the authorization model. Any user in the fips group can impersonate this node on the mesh. A process that can open the socket can send datagrams under this node's identity to any peer it names, and can hold a port and receive mesh traffic addressed to this node on it. There is no per-client authentication, no capability check and no audit trail beyond the daemon's own logs. On a node with the native API enabled, treat fips group membership exactly as you would treat the node's private key. This is why the API is disabled by default, and why enabling it is an explicit operator decision rather than something a package turns on. See security.md.

A client may hold ports 1024 through 65535. Ports 0 through 255 are reserved for protocol use and 256 through 1023 for FIPS standard services (the IPv6 shim among them), and the daemon refuses both ranges by name. A client that names no local port is given one from 49152 upward.

debug_commands is a separate gate on three commands that exist only so the test harness can drive the receive and dispatch paths without a wire. inject makes the daemon write bytes the client chose into one of that client's own flows, and arrive makes it dispatch a datagram as though a peer had sent it, reaching any listener this node holds. A node with the key off refuses each by name, so a client can tell "this node will not do that" from "this build has no such command". Leave it off outside the test harness.

fipsctl show native-flows reports the open and pending flows, the bound listeners and the native counters; see the fipsctl reference and control-socket.md. A Rust program links the crate and speaks the API through the fips::native::client module, which hides the line protocol; see ../how-to/use-the-native-datagram-api.md.

TUN Interface (tun.*)

Parameter Type Default Description
tun.enabled bool false Enable TUN virtual interface
tun.name string "fips0" Interface name
tun.mtu u16 1280 Interface MTU (IPv6 minimum)

DNS Responder (dns.*)

Resolves <npub>.fips queries to FIPS IPv6 addresses. Resolution is pure computation (npub to public key to address); resolved identities are registered with the node for routing.

Parameter Type Default Description
dns.enabled bool true Enable DNS responder
dns.bind_addr string "::1" Bind address. Default is IPv6 loopback only; the shipped fips-dns-setup configures systemd-resolved to forward .fips queries to [::1]:5354. To expose the responder to mesh peers (or to the gateway over IPv4), override (e.g., "::" for all interfaces).
dns.port u16 5354 Listen port
dns.ttl u32 300 AAAA record TTL in seconds

The dns.ttl value should not exceed node.cache.coord_ttl_secs to avoid stale address mappings.

Host Mapping

The DNS resolver checks a host map before falling back to direct npub resolution, enabling names like gateway.fips instead of npub1...fips. The host map is populated from two sources:

  1. Peer aliases — the alias field on configured peers in peers:.
  2. Hosts file — /etc/fips/hosts, one hostname npub1... per line. Blank lines and # comments are allowed.

On conflict, hosts-file entries take precedence over peer aliases. The hosts file is auto-reloaded on modification (mtime change) without restarting the daemon. Hostnames are case-insensitive.

The installer ships /etc/fips/hosts pre-populated with the public test mesh roster (test-us01 … test-uk01). Operator-style guide for adding entries and the precedence rules: ../how-to/host-aliases.md.

Transports (transports.*)

UDP (transports.udp.*)

Parameter Type Default Description
transports.udp.bind_addr string "0.0.0.0:2121" UDP bind address and port. Ignored when outbound_only: true (kernel-assigned ephemeral port is used regardless).
transports.udp.mtu u16 1280 Transport MTU
transports.udp.recv_buf_size usize 2097152 UDP socket receive buffer size in bytes (2 MB). Linux kernel doubles the requested value internally. Host net.core.rmem_max must be >= this value.
transports.udp.send_buf_size usize 2097152 UDP socket send buffer size in bytes (2 MB). Host net.core.wmem_max must be >= this value.
transports.udp.advertise_on_nostr bool false Include this UDP transport in Nostr endpoint adverts. Implicitly forced false when outbound_only: true.
transports.udp.public bool false If advertised: true publishes direct host:port; false publishes udp:nat rendezvous
transports.udp.external_addr string (none) Explicit advertise-as override. Bare IP ("203.0.113.45" — bind port is appended) or full host:port. Takes precedence over the bound address and STUN autodiscovery. Useful when the public IP isn't on a local interface (cloud 1:1 NAT, EIP) or to skip STUN for a deterministic value.
transports.udp.outbound_only bool false Pure-client posture. When true, the transport binds to 0.0.0.0:0 (kernel-assigned ephemeral port) regardless of bind_addr, refuses inbound handshake msg1, and is never advertised on Nostr regardless of advertise_on_nostr.
transports.udp.accept_connections bool true Accept inbound handshake msg1 from new peers. Combine with outbound_only: false and accept_connections: false (plus auto_connect on peer entries) for a node that initiates outbound links but rejects fresh inbound handshakes. The handshake handler carves out msg1 from peers already established on this transport so rekey continues to work.

Ethernet (transports.ethernet.*)

Ethernet transport sends raw frames over the platform's raw-frame socket: AF_PACKET SOCK_DGRAM on Linux, BPF (/dev/bpf*) on macOS. Linux and macOS only. On Linux it requires CAP_NET_RAW or running as root; on macOS it requires read/write access to a /dev/bpf* device.

Parameter Type Default Description
interface string (required) Network interface name (e.g., "eth0", "enp3s0")
ethertype u16 0x2121 EtherType
mtu u16 (auto) Override MTU. Default: interface MTU minus 3 (for frame type + length prefix)
recv_buf_size usize 2097152 Socket receive buffer size in bytes (2 MB)
send_buf_size usize 2097152 Socket send buffer size in bytes (2 MB)
listen bool true Listen for neighbor beacons from other nodes. Renamed from discovery in v0.5.0; the old key is still accepted as an alias, so a deployed config loads unchanged
announce bool false Broadcast announcement beacons on the LAN
auto_connect bool false Auto-connect to discovered peers
accept_connections bool false Accept incoming connection attempts from discovered peers
beacon_interval_secs u64 30 Announcement beacon interval in seconds (minimum 10)
optional bool false Whether absence of the interface is normal. See below

Dynamic binding. The interface does not have to exist when the daemon starts. A transport whose interface is missing comes up absent: it is not a start failure, it is not skipped, and it binds on its own the moment the interface appears — sub-second where the kernel offers link events (netlink on Linux, PF_ROUTE on the BSDs), within a second otherwise. An interface that goes away at runtime unbinds and rebinds by the same path, so a wifi reload or an unplugged adapter needs no restart. Presence means IFF_UP — the interface exists and is administratively up — and deliberately not IFF_RUNNING: binding needs no carrier, and the socket keeps working across a carrier flap without rebinding. A bridge with nothing plugged into it, such as br-lan on a wifi-only router, is therefore bound and healthy rather than permanently Degraded, and starts carrying traffic the moment a port comes up. Whether an interface has carrier is reported separately, as interface.carrier in show_transports.

optional selects how that absence is reported:

absence log retries
optional: false (default) node reports Degraded info at boot / warn on a runtime detach, then error once if it lasts past 10 s forever
optional: true no health impact info, and nothing after forever

Naming an interface in configuration is a statement that you expect it, so the default is to complain; silence is opted into. Set optional: true for hardware that is legitimately not always there — a dock adapter, a radio only some boards carry.

optional describes the interface's presence, not the transport's importance. An optional interface that is present is used exactly as hard as any other. No value of it makes a missing interface fatal at startup: the only fatal case remains "no transports at all came up".

Either way the edge is logged once, on entering absence and on recovery — never once per retry.

The edge itself is not an error. An interface missing when the daemon starts and bound a moment later is the ordinary boot race this mechanism exists to absorb, so it is info; a runtime detach is warn, because a link coming and going is ordinary weather for a mesh daemon and a cable unplugged for two seconds does not need a human.

Ten seconds is the whole grace. Past that it is no longer a race against a radio or a container coming up, so a required interface still missing is reported once at error and stays Degraded until it returns. Start-time absence and a runtime detach share the one deadline — they are the same transition throughout this mechanism. An optional interface never reaches error; that is what optional means.

Said once, not repeated. How long the absence has lasted is a state, and it is published as one: interface.since_secs in show_transports, and Degraded for as long as it holds. Re-announcing it on a timer would put a second, lossier copy of that in the log.

Node health does not wait for the ten seconds — Degraded is published on the first edge, and that is the signal to watch.

If a binding keeps dying moments after it is established (a socket that errors persistently while the interface stays up), the binder stops treating each bind as a recovery: it backs off on the same 1 s → 30 s curve, holds node health at its degraded reading, and stays quiet until a binding survives ten seconds. Without that damping a broken socket produces a health flap and a log pair every second, which is the same cry-wolf failure the edge-only logging rule exists to prevent.

A bind failure that is not absence — no CAP_NET_RAW, no readable /dev/bpf*, a buffer the kernel refused — is a fault, not a state, and fails the daemon's start as it always has. Only a missing interface is waited out.

fipsctl show transports reports the current state per transport under interface: presence (absent / binding / present), carrier, policy (required / optional), since_secs, binds, and failed_attempts.

Named instances. Multiple Ethernet interfaces can be configured by using named sub-keys instead of flat parameters:

transports:
  ethernet:
    lan:
      interface: "eth0"
      listen: true
      announce: true
    backbone:
      interface: "eth1"
      announce: false

Each named instance operates independently with its own socket and neighbor state. The instance name is used in log messages and the name() method on the Transport trait.

TCP (transports.tcp.*)

TCP transport enables firewall traversal on networks that block UDP but allow TCP (e.g., port 443). Uses FMP header-based framing with zero overhead.

Parameter Type Default Description
transports.tcp.bind_addr string (none) Listen address (e.g., "0.0.0.0:8443"). If omitted, outbound-only mode.
transports.tcp.mtu u16 1400 Default MTU. Per-connection MTU derived from TCP_MAXSEG when available.
transports.tcp.connect_timeout_ms u64 5000 Outbound connect timeout in milliseconds
transports.tcp.nodelay bool true TCP_NODELAY (disable Nagle for low latency)
transports.tcp.keepalive_secs u64 30 TCP keepalive interval in seconds (0 = disabled)
transports.tcp.recv_buf_size usize 2097152 Socket receive buffer size in bytes (2 MB)
transports.tcp.send_buf_size usize 2097152 Socket send buffer size in bytes (2 MB)
transports.tcp.max_inbound_connections usize 256 Maximum simultaneous inbound connections
transports.tcp.advertise_on_nostr bool false Include this TCP transport in Nostr endpoint adverts
transports.tcp.external_addr string (none) Explicit advertise-as override. Bare IP or full host:port. Required when bind_addr is wildcard (e.g. "0.0.0.0:443") and advertise_on_nostr: true, since TCP has no STUN equivalent for autodiscovery. Common on cloud 1:1 NAT / EIP setups where the public IP isn't bindable on the host.

Named instances. Like other transports, multiple TCP instances can be configured with named sub-keys:

transports:
  tcp:
    public:
      bind_addr: "0.0.0.0:443"
    internal:
      bind_addr: "10.0.0.1:8443"
      max_inbound_connections: 64

Tor (transports.tor.*)

Tor transport routes FIPS traffic through the Tor network for anonymity. Requires an external Tor daemon providing a SOCKS5 proxy. Three modes: socks5 for outbound-only, control_port for outbound + monitoring, directory for outbound + inbound via Tor-managed onion service.

Parameter Type Default Description
transports.tor.mode string "socks5" Tor access mode: socks5 (outbound only), control_port (outbound + monitoring), or directory (outbound + inbound onion service)
transports.tor.socks5_addr string "127.0.0.1:9050" SOCKS5 proxy address (host:port)
transports.tor.connect_timeout_ms u64 120000 Connect timeout in milliseconds. Tor circuits take 10–60s.
transports.tor.mtu u16 1400 Default MTU
transports.tor.control_addr string "/run/tor/control" Tor control port address: Unix socket path or host:port. Used in control_port mode; optional in directory mode for monitoring.
transports.tor.control_auth string "cookie" Control port authentication: "cookie", "cookie:/path/to/cookie", or "password:<secret>".
transports.tor.cookie_path string "/var/run/tor/control.authcookie" Path to Tor control cookie file. Used when control_auth is "cookie".
transports.tor.max_inbound_connections usize 64 Maximum inbound connections via onion service.
transports.tor.directory_service.hostname_file string "/var/lib/tor/fips_onion_service/hostname" Path to Tor-managed hostname file containing the .onion address.
transports.tor.directory_service.bind_addr string "127.0.0.1:8443" Local bind address for the listener that Tor forwards inbound connections to. Must match HiddenServicePort target in torrc.
transports.tor.advertise_on_nostr bool false Include this Tor transport in Nostr endpoint adverts. Requires node.rendezvous.nostr.enabled: true; setting it while Nostr rendezvous is disabled is a config-load error. advertised_port has no effect unless this is true.
transports.tor.advertised_port u16 443 Public-facing onion port published in Nostr overlay adverts. Must match the virtual port in torrc's HiddenServicePort <port> 127.0.0.1:<bind_port> directive — that is the port other peers will use to reach this onion.

Named instances. Like other transports, multiple Tor instances can be configured with named sub-keys for different SOCKS5 proxy endpoints.

Directory mode (recommended for production). Tor manages the onion service via HiddenServiceDir in torrc. FIPS reads the .onion address from the hostname file and binds a local TCP listener. This enables Tor's Sandbox 1 (seccomp-bpf). If control_addr is also set, the transport connects to the control port for daemon monitoring (non-fatal on failure).

Control port mode. Connects to the Tor daemon's control port for monitoring only (bootstrap status, circuit health, traffic stats). No inbound connections. Both control_addr and control_auth are required.

UDP + Tor Bridge Example

A node bridging clearnet (UDP) and anonymous (Tor) portions of the mesh:

node:
  identity:
    persistent: true

tun:
  enabled: true

transports:
  udp:
    bind_addr: "0.0.0.0:2121"
    mtu: 1472
  tor:
    socks5_addr: "127.0.0.1:9050"

peers:
  - npub: "npub1abc..."
    alias: "clearnet-peer"
    addresses:
      - transport: udp
        addr: "203.0.113.5:2121"
  - npub: "npub1def..."
    alias: "anonymous-peer"
    addresses:
      - transport: tor
        addr: "abc123...xyz.onion:2121"

Tor Directory Mode Example

A node accepting inbound connections via Tor-managed onion service (recommended for production — enables Sandbox 1):

node:
  identity:
    persistent: true

tun:
  enabled: true

transports:
  tor:
    mode: "directory"
    socks5_addr: "127.0.0.1:9050"
    control_addr: "/run/tor/control"    # optional, for monitoring
    control_auth: "cookie"
    directory_service:
      hostname_file: "/var/lib/tor/fips/hostname"
      bind_addr: "127.0.0.1:8444"

peers:
  - npub: "npub1abc..."
    alias: "tor-peer"
    addresses:
      - transport: tor
        addr: "abcdef...xyz.onion:8443"

Requires a corresponding torrc:

HiddenServiceDir /var/lib/tor/fips
HiddenServicePort 8443 127.0.0.1:8444

Nym (transports.nym.*)

Nym transport routes FIPS traffic through the Nym mixnet for metadata-resistant anonymity. Outbound-only: connections are made through a nym-socks5-client SOCKS5 proxy that must be running separately (e.g. as a service running alongside the fips daemon or as a container). There is no inbound listener — a Nym-only node initiates outbound links but is not reachable for unsolicited inbound handshakes.

Parameter Type Default Description
transports.nym.socks5_addr string "127.0.0.1:1080" nym-socks5-client SOCKS5 proxy address (host:port)
transports.nym.connect_timeout_ms u64 300000 Outbound connect timeout in milliseconds. Mixnet SOCKS5 connections traverse 3 mix nodes with timing obfuscation and can take several minutes, so this is generous (300s).
transports.nym.mtu u16 1400 Default MTU
transports.nym.startup_timeout_secs u64 120 Seconds to wait for nym-socks5-client to become ready at startup before giving up

Named instances. Like other transports, multiple Nym instances can be configured with named sub-keys for different SOCKS5 proxy endpoints.

BLE (transports.ble.*)

Bluetooth Low Energy transport using L2CAP Connection-Oriented Channels. Compiled on glibc Linux and on Android. At build time, build.rs sets bluer_available from the target triple (Linux and not musl) and sets ble_available for that or Android; the BLE runtime is gated behind #[cfg(ble_available)], with bluer_available gating only the BlueZ backend inside it. There is no Cargo feature flag to toggle. On musl Linux or any other platform, BLE config still parses but the transport runtime is absent and config entries become no-ops. On glibc Linux the transport communicates with BlueZ via D-Bus through the bluer crate; on Android the radio is supplied by the embedding application.

Parameter Type Default Description
transports.ble.adapter string "hci0" HCI adapter name
transports.ble.psm u16 0x0085 (133) L2CAP Protocol/Service Multiplexer
transports.ble.mtu u16 2048 Default MTU. Actual MTU is negotiated per-link during L2CAP connection setup.
transports.ble.max_connections usize 7 Maximum concurrent BLE connections
transports.ble.connect_timeout_ms u64 10000 Outbound connect timeout in milliseconds
transports.ble.advertise bool true Broadcast BLE beacon advertisements for peer discovery
transports.ble.scan bool true Listen for BLE beacon advertisements from other nodes
transports.ble.auto_connect bool false Automatically connect to discovered peers
transports.ble.accept_connections bool true Accept incoming L2CAP connections
transports.ble.probe_cooldown_secs u64 30 Cooldown before re-probing the same BLE address

Address format. BLE peer addresses use the form "adapter/device_address" — for example, "hci0/AA:BB:CC:DD:EE:FF".

Advertising and scanning. When advertise is enabled, the transport advertises the FIPS service UUID continuously so that nearby nodes can discover and connect via L2CAP, plus the L2CAP PSM its listener actually bound, as a service-data structure (see src/transport/ble/psm.rs for the wire layout and why platforms with OS-assigned PSMs need it). The advertisement carries no device name — alongside the PSM a name no longer fits the 31-byte legacy PDU, so the node shows up in generic Bluetooth scanners as an unnamed device with the FIPS UUID. When scan is enabled, the transport continuously scans for other FIPS nodes' advertisements and learns each peer's advertised PSM; a peer that advertises none is dialled at the configured psm. Discovered peers are probed immediately (L2CAP connect + pubkey exchange) with a cooldown (probe_cooldown_secs) to prevent rapid re-probing of the same address. If two nodes probe each other at the same time (cross-probe), a deterministic tie-breaker based on NodeAddr comparison ensures only one connection is established.

Connection pool. The max_connections parameter limits the number of concurrent BLE connections. When the pool is full, the least-recently-used connection is evicted to make room for new connections.

BLE Example

A node using BLE for local mesh discovery alongside UDP for internet peers:

node:
  identity:
    persistent: true

tun:
  enabled: true

transports:
  udp:
    bind_addr: "0.0.0.0:2121"
  ble:
    adapter: "hci0"
    advertise: true
    scan: true
    auto_connect: true
    accept_connections: true

peers:
  - npub: "npub1abc..."
    alias: "internet-peer"
    addresses:
      - transport: udp
        addr: "203.0.113.5:2121"
    connect_policy: auto_connect

BLE peers on the local radio range are discovered automatically via beacons — no static peer entries needed. Internet peers still require explicit configuration.

Peers (peers[])

Static peer list. Each entry defines a peer to connect to.

Parameter Type Default Description
peers[].npub string (required) Peer's Nostr public key (npub-encoded)
peers[].alias string (none) Human-readable name for logging
peers[].addresses list [] Transport addresses for the peer. May be left empty (or omitted) when via_nostr: true, in which case the daemon resolves endpoints from the peer's Nostr advert at dial time.
peers[].addresses[].transport string (required) Transport type: udp, tcp, ethernet, tor, nym, or ble. A udp entry may be qualified with a named instance as udp/<instance> (see below).
peers[].addresses[].addr string (required) Transport address. UDP/TCP: "host:port" (IP or DNS hostname). Ethernet: "interface/mac" (e.g., "eth0/aa:bb:cc:dd:ee:ff"). BLE: "adapter/device_address" (e.g., "hci0/AA:BB:CC:DD:EE:FF"). Tor: ".onion:port" or "host:port"
peers[].addresses[].priority u8 100 Address priority (lower = preferred)
peers[].connect_policy string "auto_connect" Connection policy: auto_connect, on_demand, or manual. Note: on_demand and manual are reserved for future use; the only policy currently honored at runtime is auto_connect.
peers[].auto_reconnect bool true Automatically reconnect after MMP link-dead removal (exponential backoff, unlimited retries)
peers[].via_nostr bool false Append Nostr advert-derived endpoints after static addresses for this peer

Named UDP instances. Where several UDP transports are configured under named sub-keys, a peer address can name the one it belongs to by writing the transport field as udp/<instance>, for example udp/aware. A bare udp matches any instance. The qualifier resolves only for udp: writing it on any other transport type, or naming a UDP instance that is not configured, fails config load with a validation error rather than falling back to another instance.

Gateway (gateway.*)

The gateway.* block configures the optional fips-gateway service, which lets unmodified LAN hosts reach mesh destinations through DNS proxy + virtual-IP NAT (and, optionally, exposes LAN-side services back into the mesh through inbound port forwards). The gateway is a separate service from the FIPS daemon but reads the same fips.yaml file. The block is read only when fips-gateway is running; the fips daemon ignores it. Linux only — the field is gated behind #[cfg(target_os = "linux")]. For setup, see ../how-to/deploy-gateway.md; for the end-to-end design, see ../design/fips-gateway.md.

Parameter Type Default Description
gateway.enabled bool false Enable the gateway. Must be true for fips-gateway to start.
gateway.pool string (required) Virtual IPv6 pool CIDR (e.g., "fd01::/112"). Must not overlap with the FIPS mesh address space (fd00::/8) or any address space already in use on the LAN. The /112 size yields 65 535 usable virtual IPs (address 0 in the pool is skipped), which is the gateway's hard cap regardless of CIDR width.
gateway.lan_interface string (required) LAN-facing network interface name (e.g., "enp3s0"). Used for proxy-NDP entry installation so LAN clients can resolve the link-layer address of allocated virtual IPs.
gateway.pool_grace_period u64 60 Seconds a virtual-IP allocation is retained after its last referencing session ends, before the address is returned to the free pool. Larger values reduce churn for short-lived flows; smaller values reclaim addresses faster.

Gateway DNS (gateway.dns.*)

Settings for the gateway's DNS listener and its upstream link to the FIPS daemon's .fips resolver. The gateway proxies .fips queries to the daemon's resolver, which returns mesh addresses; the gateway then allocates a virtual IP from the pool and rewrites the response. Non-.fips queries are answered with REFUSED.

Parameter Type Default Description
gateway.dns.listen string "[::1]:5353" DNS listen address. The default binds IPv6 loopback on an unprivileged port, matching the canonical deployment where another resolver on the host (dnsmasq, systemd-resolved, BIND) holds port 53 and forwards .fips queries to the gateway over loopback. Bind on the LAN-side IP (e.g., "192.168.1.1:53") or wildcard ("[::]:53") only on hosts with no other resolver on 53 and where LAN clients query the gateway directly. See ../how-to/troubleshoot-gateway.md.
gateway.dns.upstream string "[::1]:5354" Upstream FIPS daemon resolver. Must match the daemon's dns.bind_addr and dns.port. Defaults match the daemon defaults (::1:5354). A v4 upstream ("127.0.0.1:5354") cannot reach a daemon bound on [::1]:5354 — Linux IPv6 sockets bound to explicit ::1 do not accept v4-mapped traffic. If you change the daemon's dns.bind_addr, update this field accordingly.
gateway.dns.ttl u32 60 TTL in seconds on AAAA responses returned to LAN clients. Smaller values let the gateway recycle pool addresses faster; larger values reduce LAN-side query traffic.

Conntrack (gateway.conntrack.*)

Linux conntrack timeout overrides for the gateway's NAT table. These adjust the kernel-default timeouts for NAT sessions installed by the gateway. All values are in seconds; omit any field to inherit the gateway's built-in default (which itself usually matches the kernel default for that protocol).

Parameter Type Default Description
gateway.conntrack.tcp_established u64 432000 TCP established-state timeout (5 days). Long-lived TCP flows (SSH, persistent HTTP) keep their NAT mapping alive for at least this long without traffic.
gateway.conntrack.udp_timeout u64 30 UDP unreplied timeout. Applied until reply traffic is observed in the reverse direction.
gateway.conntrack.udp_assured u64 180 UDP assured (bidirectional) timeout. Applied once reply traffic has been observed.
gateway.conntrack.icmp_timeout u64 30 ICMP echo / error timeout.

Inbound Port Forwards (gateway.port_forwards[])

Optional list of inbound port-forward rules. Each rule maps a TCP or UDP port on the gateway's fips0 mesh-side address to a host:port on the LAN. Mesh peers connect to the gateway's mesh address on the listen port; the gateway terminates the connection and forwards the payload to the LAN target. This is the inverse of the outbound mode: the LAN service is exposed to the mesh, not the other way around. See ../how-to/deploy-gateway.md for the operator recipe.

Parameter Type Default Description
gateway.port_forwards[].listen_port u16 (required) Port on fips0 that mesh peers connect to. Must be non-zero. The (listen_port, proto) pair must be unique across the list.
gateway.port_forwards[].proto string (required) Transport protocol: tcp or udp.
gateway.port_forwards[].target string (required) LAN destination as IPv6 [addr]:port (e.g., "[fd12:3456::10]:80"). IPv4 targets are rejected at config-load time.

Gateway Example

A typical gateway with both outbound (LAN-to-mesh) and inbound (mesh-to-LAN) modes enabled:

gateway:
  enabled: true
  pool: "fd01::/112"
  lan_interface: "enp3s0"
  dns:
    listen: "[::1]:5353"
    upstream: "[::1]:5354"
    ttl: 60
  pool_grace_period: 60
  conntrack:
    tcp_established: 432000
    udp_assured: 180
  port_forwards:
    - listen_port: 8080
      proto: tcp
      target: "[fd12:3456::10]:80"
    - listen_port: 5353
      proto: udp
      target: "[fd12:3456::10]:53"

Minimal Example

A typical node configuration enabling TUN, DNS, and a single peer:

node:
  identity:
    nsec: "0102030405060708090a0b0c0d0e0f101112131415161718191a1b1c1d1e1f20"

tun:
  enabled: true
  name: fips0
  mtu: 1280

dns:
  enabled: true
  bind_addr: "127.0.0.1"
  port: 53

transports:
  udp:
    bind_addr: "0.0.0.0:2121"
    mtu: 1472

peers:
  - npub: "npub1tdwa4vjrjl33pcjdpf2t4p027nl86xrx24g4d3avg4vwvayr3g8qhd84le"
    alias: "node-b"
    addresses:
      - transport: udp
        addr: "172.20.0.11:2121"
    connect_policy: auto_connect

Mixed UDP + Ethernet Example

A node bridging internet peers (UDP) and a local Ethernet segment with neighbor beacons:

node:
  identity:
    nsec: "0102030405060708090a0b0c0d0e0f101112131415161718191a1b1c1d1e1f20"

tun:
  enabled: true

transports:
  udp:
    bind_addr: "0.0.0.0:2121"
    mtu: 1472
  ethernet:
    interface: "eth0"
    listen: true
    announce: true
    auto_connect: true
    accept_connections: true

peers:
  - npub: "npub1tdwa4vjrjl33pcjdpf2t4p027nl86xrx24g4d3avg4vwvayr3g8qhd84le"
    alias: "internet-peer"
    addresses:
      - transport: udp
        addr: "203.0.113.5:2121"
    connect_policy: auto_connect

Ethernet peers on the local segment are discovered automatically via beacons — no static peer entries needed. Internet peers still require explicit configuration.

All node.* parameters use their defaults. To override specific values, add only the relevant sections:

node:
  identity:
    nsec: "..."
  limits:
    max_peers: 64
  retry:
    max_retries: 10
    max_backoff_secs: 600
  cache:
    coord_size: 100000

Deprecated Keys

Every key below still loads. Nothing has been removed, so a config file written against v0.4.x keeps working after an upgrade. The old spellings are scheduled for removal at the next wire-protocol cutover, so migrate when convenient rather than urgently.

node.discovery.* split into node.lookup.* and node.rendezvous.*

The single node.discovery table mixed two unrelated jobs: resolving coordinates for a mesh address already known (lookup), and finding peers to connect to in the first place (rendezvous). It is now two tables.

Deprecated key Replacement
node.discovery.ttl node.lookup.ttl
node.discovery.attempt_timeouts_secs node.lookup.attempt_timeouts_secs
node.discovery.recent_expiry_secs node.lookup.recent_expiry_secs
node.discovery.backoff_base_secs node.lookup.backoff_base_secs
node.discovery.backoff_max_secs node.lookup.backoff_max_secs
node.discovery.forward_min_interval_secs node.lookup.forward_min_interval_secs
node.discovery.nostr.* (whole sub-table) node.rendezvous.nostr.*
node.discovery.lan.* (whole sub-table) node.rendezvous.lan.*

Behaviour of a deployed node.discovery: block: each config file is folded as it is parsed, before the cross-file merge, and a single warning is logged on the fips::config target naming the move. Only the keys actually present in the old block are applied; the rest keep their defaults. The compat block is never written back out, so anything that re-serializes the configuration emits the new spelling only.

Mixing the two spellings inside one file is not a merge. The fold runs after that file is parsed, so a value under node.discovery overwrites whatever the corresponding node.lookup or node.rendezvous key held in the same file. Use one spelling per file.

transports.ethernet.discovery renamed to transports.ethernet.listen

Deprecated key Replacement
transports.ethernet.discovery transports.ethernet.listen

This one is a plain alias rather than a compat fold, so both spellings parse into the same field and no warning is logged. The name changed because the key never controlled discovery in the node.discovery sense: it decides whether the interface listens for neighbour beacons.

Complete Reference

The full YAML structure with all defaults:

node:
  identity:
    nsec: null                       # secret key in nsec or hex (null = depends on persistent)
    persistent: false                # true = load/save fips.key; false = ephemeral each start
  leaf_only: false
  tick_interval_secs: 1
  base_rtt_ms: 100
  heartbeat_interval_secs: 10
  link_dead_timeout_secs: 30
  # drain_timeout_secs: 2            # bounded Draining phase; absent = 2s
  limits:
    max_connections: 256
    max_peers: 128
    max_links: 256
    max_pending_inbound: 1000
  rate_limit:
    handshake_burst: 100
    handshake_rate: 10.0
    handshake_timeout_secs: 30
    handshake_resend_interval_ms: 1000
    handshake_resend_backoff: 2.0
    handshake_max_resends: 5
    session_setup_burst: 64
    session_setup_rate: 16.0
  retry:
    max_retries: 5
    base_interval_secs: 5
    max_backoff_secs: 300
  netmon:
    enabled: true
    poll_interval_secs: 5
    debounce_ms: 250
  cache:
    coord_size: 50000
    coord_ttl_secs: 300
    identity_size: 10000
  lookup:
    ttl: 64
    attempt_timeouts_secs: [1, 2, 4, 8]
    recent_expiry_secs: 10
    backoff_base_secs: 0
    backoff_max_secs: 0
    forward_min_interval_secs: 2
  rendezvous:
    # nostr:                           # uncomment to enable Nostr rendezvous
    #   enabled: true                  # opt-in, default false
    #   policy: configured_only        # disabled | configured_only | open
    # lan:                             # uncomment to enable mDNS LAN rendezvous
    #   enabled: true                  # opt-in, default false
    #   scope: "my-mesh"               # optional per-network scope filter
  tree:
    announce_min_interval_ms: 500
    parent_hysteresis: 0.2              # cost improvement fraction for parent switch
    hold_down_secs: 30                  # suppress re-evaluation after switch
    reeval_interval_secs: 60            # periodic cost-based re-evaluation (0 = disabled)
    flap_threshold: 4                    # parent switches before dampening
    flap_window_secs: 60                 # sliding window for flap detection
    flap_dampening_secs: 120             # extended hold-down on flap
  bloom:
    update_debounce_ms: 500
    max_inbound_fpr: 0.20            # antipoison cap on inbound FilterAnnounce FPR
  session:
    default_ttl: 64
    pending_packets_per_dest: 16
    pending_max_destinations: 256
    idle_timeout_secs: 90
    coords_warmup_packets: 5
    coords_response_interval_ms: 2000
  mmp:
    mode: full                       # full | lightweight | minimal
    log_interval_secs: 30
    owd_window_size: 32
  session_mmp:
    mode: full                       # full | lightweight | minimal
    log_interval_secs: 30
    owd_window_size: 32
  ecn:
    enabled: true                    # ECN congestion signaling (CE flag relay)
    loss_threshold: 0.05             # MMP loss rate threshold for CE marking (5%)
    etx_threshold: 3.0               # MMP ETX threshold for CE marking
  rekey:
    enabled: true                    # periodic Noise rekey for forward secrecy
    after_secs: 120                  # rekey interval (seconds)
    after_messages: 65536            # rekey after N messages sent
  control:
    enabled: true
    socket_path: null                # null = auto (platform runtime dir → XDG → /tmp)
  # native_api:                      # uncomment to enable the experimental native datagram API
  #   enabled: true                  # opt-in, default false; not on Windows
  #   socket_path: /run/fips/api.sock  # omit the key for the resolution above
  #   pending_per_flow: 16           # datagrams held for one flow; 1..=64
  #   backlog: 16                    # flows announced on one listener, awaiting its task; at least 1
  #   max_flows: 256                 # flows this node holds at once
  #   debug_commands: false          # inject/stats/arrive; test harness only
  buffers:
    packet_channel: 1024
    tun_channel: 1024
    dns_channel: 64

tun:
  enabled: false
  name: "fips0"
  mtu: 1280

dns:
  enabled: true
  bind_addr: "::1"
  port: 5354
  ttl: 300

transports:
  udp:
    bind_addr: "0.0.0.0:2121"
    mtu: 1280
    recv_buf_size: 2097152           # 2 MB (kernel doubles to 4 MB actual)
    send_buf_size: 2097152           # 2 MB
  # ethernet:                        # uncomment to enable (requires CAP_NET_RAW)
  #   interface: "eth0"              # required: network interface name
  #   ethertype: 0x2121              # default EtherType
  #   mtu: null                      # null = interface MTU - 3 (typically 1497)
  #   recv_buf_size: 2097152         # 2 MB
  #   send_buf_size: 2097152         # 2 MB
  #   listen: true                   # listen for beacons
  #   announce: false                # broadcast beacons
  #   auto_connect: false            # connect to discovered peers
  #   accept_connections: false      # accept inbound handshakes
  #   beacon_interval_secs: 30       # beacon interval (min 10)
  # tcp:                             # uncomment to enable TCP transport
  #   bind_addr: "0.0.0.0:8443"     # listen address (omit for outbound-only)
  #   mtu: 1400                      # default MTU
  #   connect_timeout_ms: 5000       # outbound connect timeout
  #   nodelay: true                  # TCP_NODELAY
  #   keepalive_secs: 30             # keepalive interval (0 = disabled)
  #   recv_buf_size: 2097152         # 2 MB
  #   send_buf_size: 2097152         # 2 MB
  #   max_inbound_connections: 256   # resource protection limit
  # tor:                             # uncomment to enable Tor transport
  #   mode: "socks5"                 # "socks5", "control_port", or "directory"
  #   socks5_addr: "127.0.0.1:9050" # SOCKS5 proxy address
  #   connect_timeout_ms: 120000    # connect timeout (120s for Tor circuits)
  #   mtu: 1400                     # default MTU
  #   # monitoring (control_port mode, or optional in directory mode):
  #   # control_addr: "/run/tor/control"   # Unix socket or host:port
  #   # control_auth: "cookie"             # "cookie" or "password:<secret>"
  #   # cookie_path: "/var/run/tor/control.authcookie"
  #   # directory mode (inbound via Tor-managed onion service):
  #   # directory_service:
  #   #   hostname_file: "/var/lib/tor/fips_onion_service/hostname"
  #   #   bind_addr: "127.0.0.1:8443"
  #   # max_inbound_connections: 64
  #   # advertise_on_nostr: false      # publish this onion in Nostr adverts
  #   #                                # (requires node.rendezvous.nostr.enabled)
  #   # advertised_port: 443           # public-facing onion port for Nostr adverts
  # nym:                              # uncomment to enable Nym mixnet transport (outbound-only)
  #   socks5_addr: "127.0.0.1:1080" # nym-socks5-client SOCKS5 proxy address
  #   connect_timeout_ms: 300000    # connect timeout (300s for mixnet)
  #   mtu: 1400                     # default MTU
  #   startup_timeout_secs: 120     # wait for nym-socks5-client to be ready
  # ble:                              # uncomment to enable BLE transport (Linux only, requires BlueZ)
  #   adapter: "hci0"                 # HCI adapter name
  #   psm: 0x0085                     # L2CAP PSM (133)
  #   mtu: 2048                       # default MTU (negotiated per-link)
  #   max_connections: 7              # max concurrent BLE connections
  #   connect_timeout_ms: 10000       # outbound connect timeout
  #   advertise: true                 # broadcast BLE beacons
  #   scan: true                      # listen for BLE beacons
  #   auto_connect: false             # connect to discovered peers
  #   accept_connections: true         # accept incoming L2CAP connections
  #   probe_cooldown_secs: 30         # cooldown before re-probing same address

peers:                               # static peer list
  # - npub: "npub1..."
  #   alias: "node-b"
  #   addresses:
  #     - transport: udp
  #       addr: "10.0.0.2:2121"
  #       priority: 100
  #   connect_policy: auto_connect
  #   auto_reconnect: true           # reconnect after link-dead removal