Johnathan Corgan 1ec62e5119 Build the native datagram API on macOS, over SOCK_DGRAM
macOS does not implement SOCK_SEQPACKET for AF_UNIX, so the listener was
gated to Linux and FreeBSD and a Mac got no API at all. It now uses
SOCK_DGRAM there, which macOS does implement and which keeps the message
boundaries the API's contract with its clients rests on.

Both kernels were measured rather than reasoned about, and the Linux answer
alone refuted the replacement the source had proposed. On Linux 6.8 a
connected SOCK_DGRAM pair reports a closed peer not at all: revents stays
empty and recv returns EAGAIN, which is exactly what an idle socket with a
live peer does. SOCK_SEQPACKET on the same kernel sets POLLHUP and returns a
zero-byte read, which is what the receive path keyed on. Darwin does report
the close, with ECONNRESET, errno 54, and does not set POLLHUP. So the
receive path treats ECONNRESET as end of file alongside the existing POLLHUP
rule. One rule accepting either signal is correct on both kernels, where a
rule split by platform would be silently wrong on whichever one it guessed
at. EAGAIN is deliberately not in that company: it means the socket is empty
and the peer alive, so it stays an error and the caller waits again.

The measurements are asserted rather than only written down, so a kernel that
gains or loses the signal reds a test and reopens the question instead of
leaving a stale comment behind. Each carries its result in the assertion
message, since a passing test prints nothing and a negative result would
otherwise be as uninformative as no result: the close probe reports the poll
return, the whole revents bitmask broken out by flag, the recv result and the
errno, which is enough to write the real rule without another round trip. A
portable test walks datagram sizes upward, because Darwin bounds a
unix-domain datagram with the net.local.dgram.maxdgram sysctl, whose default
is small and which is a system tunable rather than something this process
controls, while the API advertises 1362 bytes to its clients.

Every test recv in the seqpacket suite is bounded in time. This is not
tidying. Simulating the Darwin configuration on a kernel that does not report
the close, the end-of-file test hung for over ten minutes rather than
failing, and a hang is not a red: it would have wedged the macOS runner with
no diagnostic instead of naming the assertion. The same simulation now fails
by name in five seconds.

Three things in the native tree compiled on one platform only, all of them in
code that had never been built for Darwin before. suseconds_t is i64 on Linux
and i32 there, so the timeval microseconds field is a cast, matching the
tv_sec line above it; it cannot truncate, because subsec_micros is below
1_000_000 by construction, and the cast is the only form that compiles on
both, since From does not exist for the narrower width and try_from is a
clippy error on the wider one. MSG_CMSG_CLOEXEC does not exist on Apple, so
the recvmsg flags are chosen per platform and each received descriptor is
marked close-on-exec with fcntl where there is no flag to pass; a failure to
set it is reported rather than ignored, since the descriptor is live either
way and the caller must not be told the receive was clean. socketpair takes
SOCK_CLOEXEC in its type argument on Linux and FreeBSD and rejects it on
macOS, so Darwin sets FD_CLOEXEC with a second fcntl. Both windows between a
call and its fcntl are stated in the code rather than closed, since the
daemon spawns no child on this path, and a test asserts both halves of a pair
are close-on-exec on every platform, because the failure is a silent
descriptor leak into a child and nothing else would report it. Every libc
item the native tree uses was then checked against the crate's own Apple
definitions rather than from memory, and those two constants are the only
ones absent.

The close difference turned out to be unhandled in six further places, and
the whole native API agrees on it now. Darwin reports a closed AF_UNIX
SOCK_DGRAM peer as ECONNRESET, and a later send on the disconnected survivor
as EDESTADDRREQ, where Linux SOCK_SEQPACKET gives EPIPE on a write and a
zero-byte read plus POLLHUP on a read. Each site below promised one of those
spellings and saw another.

- The client's recv and send passed ECONNRESET through, so a closed daemon
  half surfaced as errno 54 against the EPIPE the documentation promises.
  The translation is in one function in seqpacket rather than at each call
  site, since only this one condition has two spellings.
- accept propagated the same errno instead of its documented EPIPE. The
  listener's read reports a closed peer as the empty chunk both of its
  callers already read as the far end going away, which leaves accept's
  contract true on both platforms without either caller knowing which it is
  on.
- why() classified a failed hand-off by BrokenPipe alone, so every ordinary
  macOS listener close was counted under the counter an operator reads to
  find a client that stopped reading. It recognises all three errnos now,
  with a test over each.
- A full client buffer ended a flow's only writer. On Linux that never
  arrives, because the send reports EAGAIN and waits for the client to drain;
  Darwin has no sender-side queue to wait on and reports ENOBUFS on the send
  itself. Returning left the registration, the port and the reader alive
  while every later inbound datagram was counted as a full queue for the rest
  of the flow's life, and a client that resumed reading never recovered. The
  datagram is dropped instead, which is what a datagram API does when the far
  end cannot take it.
- The flow pair was never sized, and the two kernels charge a queued message
  to different ends: Linux to the sender's SO_SNDBUF, BSD to the receiver's
  so_rcv. Sizing only the sender, as the listener pair does, left the flow
  pair bounded on Darwin by a system default small enough that a batch held
  for an arriving client could not fit, and the whole flow was destroyed
  before its client ever saw it. Both halves are sized now.
- peer_hung_up polled with an empty events field, on the rule that POLLHUP is
  reported whether or not it is requested. That holds on Linux, where it was
  measured, and not on Darwin, where a poll requesting nothing registers no
  filter. Nothing observable depended on it, because ECONNRESET arrives first
  and both callers act on it earlier. The cost was elsewhere: three
  assertions written as tripwires for a change in Darwin's behaviour could
  not fail there, which is a guard that executes and proves nothing.
  Requesting POLLIN fixes the function and the guards together.

One difference is not an errno at all, and reading the kernel source rather
than a manual page is what found it. Darwin's unp_disconnect sets
SS_CANTRCVMORE and runs soisdisconnected on both ends for SOCK_STREAM. For
SOCK_DGRAM it removes the reflink, clears SS_ISCONNECTED and stops: no
sorwakeup, no socantrcvmore, no soisdisconnected. The closing peer deposits
ECONNRESET in the survivor's so_error and wakes no knote. The registration is
edge-triggered and was made while the socket was healthy, so nothing
re-evaluates it, and recv awaited readiness before its syscall, which left
the ECONNRESET arm sitting behind an await that never returns. A client
closing its descriptor left the daemon's reader parked for ever, and the
flow's port and registry entry held for the node's lifetime. recv reads
before it waits now, because the latched error is visible to a syscall and
only to a syscall, so the attempt that precedes the wait is what sees a close
that has already happened. A close can also land while the task is parked,
which no first attempt can catch, so on Darwin the wait is bounded and the
syscall retried; the error is latched until a read consumes it, so the bound
sets how long a dead flow holds its port rather than deciding whether the
close is seen at all. On Linux this is one extra recv returning EAGAIN before
the wait and changes nothing else, and everywhere else the readiness is
authoritative and the wait stays unbounded. The three tests this predicted
are the three that had failed: end of file on a closed client half, a
listener's port unbound on close, and one flow's port freed while its
connection stays open.

One test asserted a delivery detail rather than the rule it exists to guard.
a_descriptor_lands_on_the_last_complete_line_of_the_read_that_carried_it
asserted that a plain write and the sendmsg following it arrive in one
recvmsg. Linux coalesces them, so the read returns both lines and the
descriptor together; Darwin stops a stream read at the ancillary boundary, so
the plain line arrives by itself and the descriptor-bearing line comes on the
next read. The rule the module rests on is unaffected, and Darwin satisfies
it more easily than Linux, because the read it arrives on holds nothing
later. The test fills until both lines are queued and asserts the rule
instead of the number of reads it took.

The client compiled in /run/fips/api.sock on every platform, and macOS has no
/run for that path to be in. The daemon never had this problem: it resolves
its socket at startup by looking for a directory, and its macOS branch lands
on /var/run/fips. The constant is conditional the same way now, so a client
that is told nothing looks where a packaged daemon on its own platform
actually is. The reference documentation described that branch as
FreeBSD-only and describes both.

Windows stays excluded and cannot be included: it has no SCM_RIGHTS, so there
is no way to pass a descriptor to another process at all, which is the whole
mechanism rather than a detail of it.

The platform statements in the source and in the shipped documentation all
named Linux and FreeBSD and name macOS now, including the configuration
reference, the security reference, the how-to and the walkthrough. The how-to
also states how far the testing goes, because the person who would meet the
gap first is the one enabling the API on a Mac. The end-to-end suite drives a
client container against a node container over a shared volume, which is a
Linux arrangement, so the socket lifecycle, the descriptor hand-off across a
process boundary and the reclaiming of a port when a client exits are covered
on macOS by unit tests rather than by anything that runs a daemon and a
client as two real processes. That is a gap in testing and not a known
defect, and it is a coverage gap rather than a discharged risk. The same
place names the socket-type difference, since a reader who knows the
descriptor is SOCK_DGRAM there can make sense of a close arriving as a
different errno than the Linux documentation elsewhere describes.

The changelog entry for the API is revised rather than followed by a second
one: it now names the socket type each platform uses and the two
end-of-file signals the receive path accepts. The entry describes what the
release ships rather than the order the commits landed in.
2026-08-21 05:48:39 +00:00
2026-02-22 20:52:55 +00:00

FIPS: Free Internetworking Peering System

banner License: MIT Rust Status

A self-organizing encrypted mesh network built on Nostr identities, capable of operating over arbitrary transports without central infrastructure.

FIPS is under active development. The protocol and APIs are not yet stable. See Status & roadmap below.

What FIPS does

A machine running FIPS becomes a node in the mesh with a self-generated cryptographic identity (a Nostr keypair). There are two equally-supported deployment modes.

As an overlay on top of existing IP networks, FIPS lets your node reach any other FIPS node wherever it sits — behind a NAT, on a different ISP, on a phone over cellular, on a laptop with only Bluetooth in range, or behind a Tor onion. The mesh forwards IPv6 traffic transparently and end-to-end encrypted, with no central VPN concentrator or coordinating server.

Ground up over raw Ethernet, WiFi, or Bluetooth, FIPS provides a complete permissionless network without any pre-existing IP infrastructure, ISP, or DNS. Any node that joins the link gets routable IPv6 addresses, peer discovery, and a path to every other node automatically.

Either way, existing networking software runs over it unchanged — SSH, HTTP servers, file transfer, anything IPv6-native works the same way it would on a local network.

Features

  • Self-organizing mesh routing. Spanning-tree coordinates with bloom-filter-guided discovery; no global routing tables, no flooding.
  • Multi-transport. UDP, TCP, Ethernet, Tor, Nym, and Bluetooth (BLE L2CAP) ship today; transports compose on a single mesh and a node may run several at once.
  • Two-layer encryption. Noise IK between peers (hop-by-hop) and Noise XK between mesh endpoints (independent end-to-end), with periodic rekey for forward secrecy.
  • Nostr-native identity. secp256k1 / schnorr keypairs as node addresses; self-generated, no registration, no central authority.
  • IPv6 adapter. A TUN interface maps each remote npub to an fd00::/8 address, so unmodified IPv6 software reaches mesh peers as <npub>.fips. Built-in .fips DNS resolver, with optional static name mapping via /etc/fips/hosts.
  • Nostr-mediated discovery and NAT traversal. Peers publish endpoint adverts on public Nostr relays, exchange candidates via NIP-59 gift-wrapped offers and answers, and establish direct paths through NATs using STUN-assisted hole punching. On the local network, mDNS LAN discovery finds peers directly without relays.
  • LAN gateway. Optional fips-gateway service folds an entire unmodified LAN into the mesh: outbound (LAN clients reach mesh destinations through a DNS-allocated virtual IPv6 pool and nftables NAT) and inbound (LAN-side services exposed to the mesh through 1:1 port forwards).
  • Per-link metrics. RTT, loss, jitter, and goodput on every hop, plus mesh-size estimation, via the Metrics Measurement Protocol.
  • ECN congestion signaling. Hop-by-hop CE-flag relay with RFC 3168 IPv6 marking and transport kernel-drop detection.
  • Mesh-interface security baseline. Optional default-deny nftables policy for fips0 shipped as a packaged conffile (/etc/fips/fips.nft) with an operator drop-in directory (/etc/fips/fips.d/) and a disabled-by-default fips-firewall.service. The baseline polices only the mesh interface, leaving Docker, Tor, and the host firewall untouched.
  • Operator visibility. fipsctl CLI for control and inspection with time-series stats history queryable for any metric, fipstop TUI for live status with inline sparkline dashboards, and a JSON-line control socket on each binary for direct programmatic access.
  • Reproducible builds with toolchain pinning and SOURCE_DATE_EPOCH.

Quick start

The shortest path on Debian / Ubuntu:

git clone https://github.com/jmcorgan/fips.git
cd fips
cargo install cargo-deb
cargo deb
sudo dpkg -i target/debian/fips_*.deb
sudo systemctl start fips

This installs the daemon, CLI tools (fipsctl, fipstop), the optional fips-gateway service, systemd units, and a default /etc/fips/fips.yaml you can edit before starting.

For macOS, Windows, OpenWrt, the systemd tarball, a Nix flake, or a from-source build, see docs/getting-started.md for the full multi-platform installation guide.

To join a live mesh and reach your first peer, follow the new-user tutorial progression starting at docs/tutorials/join-the-test-mesh.md.

Building from source

cargo build --release

Requires Rust 1.94.1+ (edition 2024). Linux, macOS, FreeBSD, and Windows run as standalone daemons; Android is supported as an embedded library (the host app owns the TUN, e.g. a VpnService). Transport availability varies by platform.

Transport Linux macOS FreeBSD Windows Android OpenWrt
UDP ✅ ✅ ✅ ✅ ✅ ✅
TCP ✅ ✅ ✅ ✅ ✅ ✅
Ethernet ✅ ✅ ❌ ❌ ❌ ✅
Tor ✅ ✅ ✅ ✅ ❌ ✅
Nym ✅ ✅ ✅ ✅ ❌ ❌
BLE ✅ ❌ ❌ ❌ ❌ ❌

On Linux, a source build requires libclang — the LAN gateway's nftables bindings are generated by bindgen at build time, which needs libclang.so on the build host. Install it before building (sudo apt install libclang-dev on Debian / Ubuntu); without it the build fails inside the rustables crate with an "Unable to find libclang" error. This is a build-time prerequisite only — it is not a runtime dependency, and the pre-built .deb artifacts do not need it.

BLE is optional and, on Linux, requires BlueZ and libdbus (sudo apt install bluez libdbus-1-dev on Debian / Ubuntu). It is gated on a build-script probe — install the dependencies first and the cargo build line above picks it up. The OpenWrt ipk omits BLE because libdbus is not available on the target.

Nym (mixnet) transport builds on all desktop platforms. The OpenWrt ❌ is provisional, pending verification of nym-socks5-client availability on the target; it will flip to ✅ only if confirmed buildable there.

Alternatively, the repo ships a Nix flake: nix develop drops you into a shell with the pinned toolchain and every build prerequisite (libclang, dbus, pkg-config) already provided, and nix build .#fips builds all four binaries with no host setup. See the Nix / NixOS section of packaging/README.md.

Documentation

docs/ is organised by reader purpose:

  • Tutorials — hand-held walk-throughs from a fresh install through to a participating mesh node, plus advanced deployments (gateway on OpenWrt, hosting services, ground-up two-device mesh).
  • How-to guides — operator recipes for specific tasks: firewall activation, Nostr discovery, Tor onion service, Bluetooth peering, LAN gateway deployment and troubleshooting, MTU diagnostics, host aliases, persistent identity, unprivileged-user setup, UDP buffer tuning.
  • Reference — fips.yaml configuration, wire formats, control-socket protocol, CLI references for each binary, security posture matrix, Nostr events catalog, transport statistics inventory.
  • Design — protocol-level architecture and layer specifications. Start with fips-concepts.md for the framing, then fips-architecture.md for the protocol stack.

If you want to contribute, see CONTRIBUTING.md and testing/README.md.

Examples

  • examples/sidecar-nostr-relay/ — Run a strfry Nostr relay reachable exclusively over the FIPS mesh. The relay container shares the FIPS sidecar's network namespace and is isolated from the host network.
  • examples/sidecar-nostr-mixnet-relay/ — Single-container demo of FIPS peering through a mixnet (implemented with Nym): the FIPS daemon, the mixnet proxy, and a strfry Nostr relay all in one isolated container, with the direct route to the peer firewalled off so traffic provably crosses the mixnet.
  • examples/k8s-sidecar/ — Run FIPS as a Kubernetes Pod sidecar. The sidecar creates fips0 in the Pod's shared network namespace so every other container in the Pod gets mesh access without modification.
  • examples/wireguard-sidecar-macos/ — Reach the FIPS mesh from a macOS host through a local Docker container over a WireGuard tunnel. Only traffic destined for fd00::/8 transits the sidecar; regular internet traffic continues to use the host network.

Project structure

src/          Rust source: library + fips, fipsctl, fipstop, fips-gateway binaries
docs/         Documentation: tutorials, how-to, reference, design
packaging/    Debian, macOS .pkg, Windows ZIP, OpenWrt ipk, AUR, systemd tarball
examples/     Deployment examples (Nostr relay, K8s sidecar, macOS WireGuard)
testing/      Docker-based integration test harnesses + chaos simulation

Status & roadmap

FIPS is at v0.5.0-dev on the master branch. v0.4.1 has shipped; this development line continues the testing-and-polishing track toward v0.5.0. The core protocol works end-to-end over UDP, TCP, Ethernet, Tor, Nym, and Bluetooth on a global, public test mesh of thousands of nodes. v0.4.0 added the Nym mixnet transport and mDNS LAN discovery alongside the existing Nostr-mediated peer discovery, UDP NAT traversal, peer ACL, and packaging hardening. New wire-format work continues to be staged on the next branch for the subsequent release line.

What works today

  • Spanning-tree construction with greedy coordinate routing.
  • Bloom-filter-guided destination discovery (no flooding, single-path with retry).
  • Two-layer Noise encryption (IK at the link, XK at the session) with periodic hitless rekey for forward secrecy at both layers.
  • Persistent or ephemeral node identity with key-file management.
  • IPv6 TUN adapter with built-in .fips DNS resolver and multi-backend auto-configuration (systemd dns-delegate, systemd-resolved, dnsmasq, NetworkManager).
  • Static hostname mapping (/etc/fips/hosts) with auto-reload.
  • Per-link metrics (RTT, loss, jitter, goodput) and mesh size estimation.
  • ECN congestion signaling (hop-by-hop CE relay, IPv6 CE marking, kernel-drop detection).
  • UDP, TCP, Ethernet, Tor, Nym (mixnet), and BLE transports (BLE via L2CAP CoC with per-link MTU negotiation).
  • Nostr-mediated overlay endpoint discovery and UDP hole punching for NAT traversal, plus mDNS LAN discovery for local peers.
  • LAN gateway (fips-gateway) with both outbound (LAN-to-mesh) and inbound (mesh-to-LAN port-forwarding) modes.
  • Peer ACL: per-npub allow / deny admission control at the link layer; opt-in mesh-firewall baseline at fips0 ingress.
  • Runtime inspection and peer management via fipsctl and fipstop.
  • Reproducible builds with toolchain pinning and SOURCE_DATE_EPOCH.
  • Linux (Debian, systemd tarball, OpenWrt, AUR), macOS (.pkg), FreeBSD (.pkg), and Windows (ZIP, service) packaging.
  • Docker-based integration and chaos testing.

Near-term priorities

  • Native API for FIPS-aware applications (npub:port addressing without the IPv6-shim path).
  • Security audit of the cryptographic protocols.

Longer-term

  • Mobile platform support.
  • Bandwidth-aware routing and QoS.
  • Protocol stability and a versioned wire format.
  • Published crate.

License

MIT — see LICENSE.

S
Description
The Free Internetworking Peering System
Readme MIT
66 MiB
Languages
Rust 82%
Shell 14.1%
Python 3.3%
PowerShell 0.2%
Nix 0.1%
Other 0.1%