macOS does not implement SOCK_SEQPACKET for AF_UNIX, so the listener was gated to Linux and FreeBSD and a Mac got no API at all. It now uses SOCK_DGRAM there, which macOS does implement and which keeps the message boundaries the API's contract with its clients rests on. Both kernels were measured rather than reasoned about, and the Linux answer alone refuted the replacement the source had proposed. On Linux 6.8 a connected SOCK_DGRAM pair reports a closed peer not at all: revents stays empty and recv returns EAGAIN, which is exactly what an idle socket with a live peer does. SOCK_SEQPACKET on the same kernel sets POLLHUP and returns a zero-byte read, which is what the receive path keyed on. Darwin does report the close, with ECONNRESET, errno 54, and does not set POLLHUP. So the receive path treats ECONNRESET as end of file alongside the existing POLLHUP rule. One rule accepting either signal is correct on both kernels, where a rule split by platform would be silently wrong on whichever one it guessed at. EAGAIN is deliberately not in that company: it means the socket is empty and the peer alive, so it stays an error and the caller waits again. The measurements are asserted rather than only written down, so a kernel that gains or loses the signal reds a test and reopens the question instead of leaving a stale comment behind. Each carries its result in the assertion message, since a passing test prints nothing and a negative result would otherwise be as uninformative as no result: the close probe reports the poll return, the whole revents bitmask broken out by flag, the recv result and the errno, which is enough to write the real rule without another round trip. A portable test walks datagram sizes upward, because Darwin bounds a unix-domain datagram with the net.local.dgram.maxdgram sysctl, whose default is small and which is a system tunable rather than something this process controls, while the API advertises 1362 bytes to its clients. Every test recv in the seqpacket suite is bounded in time. This is not tidying. Simulating the Darwin configuration on a kernel that does not report the close, the end-of-file test hung for over ten minutes rather than failing, and a hang is not a red: it would have wedged the macOS runner with no diagnostic instead of naming the assertion. The same simulation now fails by name in five seconds. Three things in the native tree compiled on one platform only, all of them in code that had never been built for Darwin before. suseconds_t is i64 on Linux and i32 there, so the timeval microseconds field is a cast, matching the tv_sec line above it; it cannot truncate, because subsec_micros is below 1_000_000 by construction, and the cast is the only form that compiles on both, since From does not exist for the narrower width and try_from is a clippy error on the wider one. MSG_CMSG_CLOEXEC does not exist on Apple, so the recvmsg flags are chosen per platform and each received descriptor is marked close-on-exec with fcntl where there is no flag to pass; a failure to set it is reported rather than ignored, since the descriptor is live either way and the caller must not be told the receive was clean. socketpair takes SOCK_CLOEXEC in its type argument on Linux and FreeBSD and rejects it on macOS, so Darwin sets FD_CLOEXEC with a second fcntl. Both windows between a call and its fcntl are stated in the code rather than closed, since the daemon spawns no child on this path, and a test asserts both halves of a pair are close-on-exec on every platform, because the failure is a silent descriptor leak into a child and nothing else would report it. Every libc item the native tree uses was then checked against the crate's own Apple definitions rather than from memory, and those two constants are the only ones absent. The close difference turned out to be unhandled in six further places, and the whole native API agrees on it now. Darwin reports a closed AF_UNIX SOCK_DGRAM peer as ECONNRESET, and a later send on the disconnected survivor as EDESTADDRREQ, where Linux SOCK_SEQPACKET gives EPIPE on a write and a zero-byte read plus POLLHUP on a read. Each site below promised one of those spellings and saw another. - The client's recv and send passed ECONNRESET through, so a closed daemon half surfaced as errno 54 against the EPIPE the documentation promises. The translation is in one function in seqpacket rather than at each call site, since only this one condition has two spellings. - accept propagated the same errno instead of its documented EPIPE. The listener's read reports a closed peer as the empty chunk both of its callers already read as the far end going away, which leaves accept's contract true on both platforms without either caller knowing which it is on. - why() classified a failed hand-off by BrokenPipe alone, so every ordinary macOS listener close was counted under the counter an operator reads to find a client that stopped reading. It recognises all three errnos now, with a test over each. - A full client buffer ended a flow's only writer. On Linux that never arrives, because the send reports EAGAIN and waits for the client to drain; Darwin has no sender-side queue to wait on and reports ENOBUFS on the send itself. Returning left the registration, the port and the reader alive while every later inbound datagram was counted as a full queue for the rest of the flow's life, and a client that resumed reading never recovered. The datagram is dropped instead, which is what a datagram API does when the far end cannot take it. - The flow pair was never sized, and the two kernels charge a queued message to different ends: Linux to the sender's SO_SNDBUF, BSD to the receiver's so_rcv. Sizing only the sender, as the listener pair does, left the flow pair bounded on Darwin by a system default small enough that a batch held for an arriving client could not fit, and the whole flow was destroyed before its client ever saw it. Both halves are sized now. - peer_hung_up polled with an empty events field, on the rule that POLLHUP is reported whether or not it is requested. That holds on Linux, where it was measured, and not on Darwin, where a poll requesting nothing registers no filter. Nothing observable depended on it, because ECONNRESET arrives first and both callers act on it earlier. The cost was elsewhere: three assertions written as tripwires for a change in Darwin's behaviour could not fail there, which is a guard that executes and proves nothing. Requesting POLLIN fixes the function and the guards together. One difference is not an errno at all, and reading the kernel source rather than a manual page is what found it. Darwin's unp_disconnect sets SS_CANTRCVMORE and runs soisdisconnected on both ends for SOCK_STREAM. For SOCK_DGRAM it removes the reflink, clears SS_ISCONNECTED and stops: no sorwakeup, no socantrcvmore, no soisdisconnected. The closing peer deposits ECONNRESET in the survivor's so_error and wakes no knote. The registration is edge-triggered and was made while the socket was healthy, so nothing re-evaluates it, and recv awaited readiness before its syscall, which left the ECONNRESET arm sitting behind an await that never returns. A client closing its descriptor left the daemon's reader parked for ever, and the flow's port and registry entry held for the node's lifetime. recv reads before it waits now, because the latched error is visible to a syscall and only to a syscall, so the attempt that precedes the wait is what sees a close that has already happened. A close can also land while the task is parked, which no first attempt can catch, so on Darwin the wait is bounded and the syscall retried; the error is latched until a read consumes it, so the bound sets how long a dead flow holds its port rather than deciding whether the close is seen at all. On Linux this is one extra recv returning EAGAIN before the wait and changes nothing else, and everywhere else the readiness is authoritative and the wait stays unbounded. The three tests this predicted are the three that had failed: end of file on a closed client half, a listener's port unbound on close, and one flow's port freed while its connection stays open. One test asserted a delivery detail rather than the rule it exists to guard. a_descriptor_lands_on_the_last_complete_line_of_the_read_that_carried_it asserted that a plain write and the sendmsg following it arrive in one recvmsg. Linux coalesces them, so the read returns both lines and the descriptor together; Darwin stops a stream read at the ancillary boundary, so the plain line arrives by itself and the descriptor-bearing line comes on the next read. The rule the module rests on is unaffected, and Darwin satisfies it more easily than Linux, because the read it arrives on holds nothing later. The test fills until both lines are queued and asserts the rule instead of the number of reads it took. The client compiled in /run/fips/api.sock on every platform, and macOS has no /run for that path to be in. The daemon never had this problem: it resolves its socket at startup by looking for a directory, and its macOS branch lands on /var/run/fips. The constant is conditional the same way now, so a client that is told nothing looks where a packaged daemon on its own platform actually is. The reference documentation described that branch as FreeBSD-only and describes both. Windows stays excluded and cannot be included: it has no SCM_RIGHTS, so there is no way to pass a descriptor to another process at all, which is the whole mechanism rather than a detail of it. The platform statements in the source and in the shipped documentation all named Linux and FreeBSD and name macOS now, including the configuration reference, the security reference, the how-to and the walkthrough. The how-to also states how far the testing goes, because the person who would meet the gap first is the one enabling the API on a Mac. The end-to-end suite drives a client container against a node container over a shared volume, which is a Linux arrangement, so the socket lifecycle, the descriptor hand-off across a process boundary and the reclaiming of a port when a client exits are covered on macOS by unit tests rather than by anything that runs a daemon and a client as two real processes. That is a gap in testing and not a known defect, and it is a coverage gap rather than a discharged risk. The same place names the socket-type difference, since a reader who knows the descriptor is SOCK_DGRAM there can make sense of a close arriving as a different errno than the Linux documentation elsewhere describes. The changelog entry for the API is revised rather than followed by a second one: it now names the socket type each platform uses and the two end-of-file signals the receive path accepts. The entry describes what the release ships rather than the order the commits landed in.
16 KiB
Security Reference
Consolidated security reference covering the nftables baseline, peer ACL file format, cryptographic primitives, rekey defaults, replay window, filesystem permissions, threat-resistance matrix, and default network exposures per transport. For the threat-model design and rationale, see ../design/fips-security.md. For the operator activation steps and drop-in recipes, see ../how-to/enable-mesh-firewall.md.
nftables Baseline
The shipped baseline is /etc/fips/fips.nft. It defines a single
nftables table inet fips with one chain hooked at input, structured
as follows:
| Step | Rule | Effect |
|---|---|---|
| 1 | iifname != "fips0" return |
Match only traffic arriving on fips0; everything else short-circuits. |
| 2 | ct state established,related accept |
Allow conntrack replies and related ICMPv6 errors. |
| 3 | icmpv6 type echo-request accept |
Allow IPv6 echo (ping6 reachability). |
| 4 | include "/etc/fips/fips.d/*.nft" |
Splice in operator drop-ins (empty matches nothing). |
| 5 | counter drop |
Default-deny everything else; counter increments on every drop. |
Outbound from fips0 is unrestricted. The baseline is a documented
dpkg conffile — operator edits to /etc/fips/fips.nft are preserved
across upgrades.
The systemd unit is fips-firewall.service (oneshot). It is not
enabled by default; activation is an explicit operator gesture
documented in
../how-to/enable-mesh-firewall.md.
Drop-In File Format
Operator extensions live under /etc/fips/fips.d/ with the .nft
suffix. Each file is included inline into the inbound chain at the
marked point and may contain any nftables rule lines valid in that
context.
Naming convention: <purpose>-from-<source>.nft keeps drop-ins easy
to scan. Examples shipped in the design discussion:
ssh-from-bastion.nft— accept TCP/22 from a single mesh-node addresshttp-from-cluster.nft— accept TCP/80 from a/64mesh-address prefixdns-public.nft— accept UDP/53 and TCP/53 from any mesh nodegit-from-trusted.nft— accept TCP/9418 from a set of mesh-node addresses
After editing, reload via
sudo systemctl reload-or-restart fips-firewall.service (or
equivalently sudo nft -f /etc/fips/fips.nft since the file is
idempotent).
Cryptographic Primitives
| Component | Choice | Where Used |
|---|---|---|
| Curve | secp256k1 | FMP IK, FSP XK, Schnorr signatures |
| Diffie-Hellman | ECDH on secp256k1 (x-only normalized) | Noise IK, Noise XK |
| AEAD | ChaCha20-Poly1305 | FMP link encryption, FSP session encryption |
| Hash | SHA-256 | NodeAddr derivation, Noise key schedule |
| Key derivation | HKDF-SHA256 | Noise key schedule |
| Signatures | secp256k1 Schnorr | TreeAnnounce, LookupResponse proof, Nostr adverts |
| Noise pattern (link) | Noise_IK_secp256k1_ChaChaPoly_SHA256, with the deviation below |
FMP link layer (IK with epoch payload) |
| Noise pattern (session) | Noise_XK_secp256k1_ChaChaPoly_SHA256, with the deviation below |
FSP session layer (XK with epoch payload) |
These choices align with the Nostr cryptographic stack (secp256k1 + ChaCha20-Poly1305 + SHA-256) and the NIP-44 encrypted messaging standard.
Deviation: Empty Associated Data in the Handshake AEAD
Both Noise patterns above deviate from the standard construction in one
respect. The handshake AEAD uses an empty associated-data field where
standard Noise EncryptAndHash uses the handshake hash h.
The choice was deliberate. Using secp256k1 rather than 25519 already put the construction outside standard Noise, so no standard-Noise peer could be confused with it, and the transcript hash bought no distinguishing value.
That argument is about domain separation, and on those grounds it holds. It
does not cover transcript binding, which is the property actually absent.
Domain separation and DH binding survive through the chaining key ck, which
mix_key chains from ck = h, seeded from the protocol name in
SymmetricState::initialize (src/noise/handshake.rs). The handshake hash
h is maintained at every step and is never fed to the AEAD, so it binds
nothing.
Rekey Defaults
Both link-layer and session-layer Noise sessions rekey under one of
two triggers, configurable under node.rekey.*:
| Parameter | Default | Description |
|---|---|---|
enabled |
true |
Master switch. |
after_secs |
120 |
Time-based rekey threshold. |
after_messages |
65536 |
Message-count rekey threshold. |
In addition to the configurable triggers, the daemon retains the old
session keys for a fixed 10-second drain window after each
cutover (compile-time constant DRAIN_WINDOW_SECS in
src/node/handlers/rekey.rs). Rekey rotates the Noise key schedule
and the session indices; old session keys are kept in
previous_session for the drain window so in-flight packets
encrypted under the old keys still decrypt.
Replay Window
Both layers use explicit per-packet counters with a sliding bitmap
window for replay protection. The bitmap is 2048 entries at both
layers — large enough to accommodate UDP reordering and packet loss
without false-positive replay rejection. Counters older than the
window are rejected. The same ReplayWindow and
decrypt_with_replay_check() implementation is used at both the FMP
and FSP layers.
Peer ACL
Mesh-level ACL files at /etc/fips/peers.allow and
/etc/fips/peers.deny give the operator allowlist/blocklist control
over which npubs may complete the FMP Noise IK link handshake.
File format:
- One entry per line. An entry is either a bech32
npub1..., an alias defined in/etc/fips/hosts, or the literalALLwildcard (case-insensitive). - Lines beginning with
#are comments. - Blank lines are ignored.
Evaluation order (first match wins, default-allow on no match):
peers.allow— if the peer matches an entry here (orALLis inpeers.allow), the handshake is admitted, regardless of anypeers.denyentry.peers.deny— if the peer matches an entry here (orALLis inpeers.deny), the handshake is refused.- Otherwise the peer is admitted.
peers.allow is not an exclusive gate on its own: an unlisted
peer falls through to step 3 and is admitted unless it appears in
peers.deny. To turn peers.allow into a strict allowlist, place
ALL in peers.deny so every unlisted peer is rejected at step 2.
The ALL wildcard makes the operator's posture explicit:
ALLinpeers.allowadmits every peer (same effect as the default-allow behavior, but documented in the file).ALLinpeers.denyblocks every peer except those listed inpeers.allow— the "allowlist-strict" posture.
In practice this collapses to a few common postures:
- Default-allow with denylist: leave
peers.allowempty; populatepeers.deny. All npubs may peer except those listed. - Allowlist-strict: populate
peers.allowand putALLinpeers.deny. Only the listed npubs may peer; everyone else is rejected at step 2.
A populated peers.allow with an empty peers.deny is not a
strict allowlist — it is equivalent to default-allow plus an
explicit "always-admit" set. The strict variant requires ALL
in peers.deny.
Aliases are resolved through /etc/fips/hosts at file-load
time. If peers.allow lists core-vm and /etc/fips/hosts
maps core-vm to a specific npub, that npub is admitted. If
core-vm is later remapped to a different npub, the ACL
re-resolves on the next mtime change. Operators should be aware
that ACL semantics follow the hosts-file aliasing, not just
the literal npubs visible in the file.
Both files are reloaded automatically when their mtime changes — no daemon restart or signal is needed. ACL evaluation runs after msg1 decryption but before any further peer-state mutation; rate-limited msg1s never reach the ACL.
Filesystem Permissions
| Path | Owner | Mode | Purpose |
|---|---|---|---|
/etc/fips/fips.key |
root:root | 0600 |
Persistent identity private key (sensitive). |
/etc/fips/fips.pub |
root:root | 0644 |
Public key (npub). |
/etc/fips/fips.yaml |
root:root | 0644 |
Daemon configuration (dpkg conffile). |
/etc/fips/fips.nft |
root:root | 0644 |
nftables baseline (dpkg conffile). |
/etc/fips/fips.d/ |
root:root | 0755 |
Operator drop-in directory. |
/etc/fips/hosts |
root:root | 0644 |
Optional hostname → npub map (dpkg conffile). |
/etc/fips/peers.allow |
root:root | 0644 |
Optional peer allowlist. |
/etc/fips/peers.deny |
root:root | 0644 |
Optional peer denylist. |
/run/fips/control.sock |
root:fips | 0770 |
Control socket (members of fips group can use fipsctl). |
/run/fips/api.sock |
root:fips | 0770 |
Native datagram API socket, when node.native_api.enabled is set (experimental; absent otherwise). |
/run/fips/ |
root:fips | 0750 |
Socket parent directory. |
Adding a user to the fips group grants fipsctl access without
requiring root. The daemon chowns the control socket and its parent
directory at bind time, and does the same for the native API socket when
that is enabled.
Native Datagram API
Experimental. Disabled by default (node.native_api.enabled, default
false), and built on Linux, FreeBSD and macOS only. It is not a stable API
surface, not a reliability layer, and not the v2 external process API. No
compatibility promise is made about it.
Any user in the fips group can impersonate the node on the mesh. The
API socket is created at mode 0770 owned by group fips, and that is the
entire authorization model. A process that can open it can:
- send datagrams under this node's identity to any peer it names, which peers authenticate as coming from this node;
- hold any port from 1024 upward and receive mesh traffic addressed to this node on it, including traffic another local program expected;
- do both without authenticating, without a capability check, and without any record beyond the daemon's own logs.
Group membership is therefore equivalent to possession of the node's
identity for the purpose of sending on the mesh. On a node with the native
API enabled, treat membership of the fips group exactly as you would treat
/etc/fips/fips.key. Grant it to the accounts that are trusted to speak as
the node and to no others, and review it before enabling the API on a shared
machine.
The file descriptor carries the grant, not the connection. A setup call
hands the client a socket descriptor and the connection it was made on is then
closed; the flow or the held port lives until that descriptor is closed. A
descriptor is an ordinary kernel object, so it survives fork, survives
exec unless the client asked for it close-on-exec when it received it, and
can be handed to another process over SCM_RIGHTS. A process holding one can
send as this node on that flow, or receive on that port, without ever opening
the API socket and without being in the fips group.
Nothing revokes a descriptor already handed out. Restarting the daemon closes
its own halves and ends every flow and listener at once, and that is the only
revocation there is.
Two consequences follow for fipsctl access. First, the fips group is
already the control-socket group, so enabling the native API silently
upgrades every existing fipsctl user from "can read node state and manage
peers" to "can send as the node". Second, an operator who wants the two
audiences separated must not enable the API on a node whose fips group has
been handed out for monitoring.
node.native_api.debug_commands (default false) is a second, independent
gate. It admits three commands (inject, stats, arrive) that exist for
the test harness: arrive makes the daemon dispatch a datagram as though a
peer had sent it, reaching any listener on this node under any peer identity
the caller names. Leave it off outside a test harness; a packaged node does
not enable it.
The socket is local only. It is not reachable over the network, and nothing about it changes the mesh's own authentication: a peer still verifies the node's signature, which is precisely why a local caller that can send through this socket is indistinguishable from the node itself.
See configuration.md for the key list and ../how-to/use-the-native-datagram-api.md for the client.
Threat-Resistance Matrix
The link layer's threat-resistance matrix is consolidated here from the FMP design document:
| Threat | Mitigation |
|---|---|
| Connection exhaustion | Token-bucket rate limit + connection count limit |
| CPU exhaustion (msg1 flood) | Rate limit before crypto operations |
| Replay attacks | Counter-based nonces with sliding window (2048 entries) |
| State confusion | Strict handshake state machine validation |
| Spoofed encrypted packets | Index lookup + AEAD verification |
| Spoofed msg2 | Index lookup + Noise ephemeral key binding |
| Address spoofing | Cryptographic authority, not address-based |
| Session correlation | Index rotation on rekey |
Inbound exposure on fips0 |
Default-deny nftables baseline (operator opt-in) |
| Sybil identities | Discretionary peering + handshake rate limiting + optional peer ACL |
| Eclipse attack | Diverse peering across independent operators and transports |
| Unauthorized peer admission | Optional peers.allow allowlist consulted before handshake |
| Local impersonation via the native datagram API | API disabled by default; when enabled, fips group membership is the only gate and must be treated as key access |
See ../design/fips-mesh-layer.md for the unauthenticated-attack-surface analysis (only handshake msg1 is reachable by unauthenticated parties), and ../design/fips-mesh-operation.md for the metadata-privacy model and the rejection of onion routing.
Default Network Exposures by Transport
| Transport | Default Inbound | Default Bind | Opt-in |
|---|---|---|---|
| UDP | None until bind_addr set |
0.0.0.0:2121 typical |
Operator sets transports.udp.bind_addr |
| TCP | None until bind_addr set |
None — outbound-only without bind | Operator sets transports.tcp.bind_addr |
| Ethernet | Listens on configured interface (raw AF_PACKET) |
EtherType 0x2121 on selected interface | Per-flag listen, announce, auto_connect, accept_connections |
| Tor | None until directory_service configured |
127.0.0.1:8443 (loopback only) |
Operator sets transports.tor.directory_service and configures HiddenServiceDir in torrc |
| BLE | Off by default | n/a | Operator enables transports.ble.* |
| Nostr discovery | Off by default | n/a (relay client, not a listener) | Operator sets node.discovery.nostr.enabled: true |
The mesh-layer fips0 interface is reachable from any mesh node that
can route to you, not only direct peers — your direct peers forward
traffic from any reachable mesh node onto your fips0. The
default-deny nftables baseline (operator opt-in) is the recommended
way to restrict inbound traffic on fips0. See
../how-to/enable-mesh-firewall.md.
See also
- ../design/fips-security.md — threat
model and design rationale for the
fips0baseline - ../design/fips-mesh-layer.md — FMP link encryption, replay protection, rate limiting
- ../design/fips-session-layer.md — FSP end-to-end encryption, Noise XK, replay window
- ../how-to/enable-mesh-firewall.md — operator activation and drop-in recipes
- configuration.md — full
node.rekey.*,node.rate_limit.*parameter tables - ../how-to/use-the-native-datagram-api.md — enabling the experimental native datagram API, and what group membership grants once it is on