ArjenandJohnathan Corgan 95866b7c7c test(iface-binding): cover the presence machine end to end
Two daemons whose only transports are interface-bound, run against a veth
pair the harness creates, downs, deletes and recreates underneath them.

Asserts the boot race (a daemon whose only interface is missing starts,
reports the transport absent and the node Degraded, rather than exiting on
NoTransports or skipping the transport for the life of the process), the
late attach and discovery over it, the flap in both directions,
destroy-and-recreate, and that an optional interface which never appears
never moves node health.

Also the log policy, which is the half that is easy to regress silently:
absence is logged once on the edge and not once per retry; a required
interface still absent past the ten-second bring-up window errors exactly
once, while the optional one — absent just as long — stays silent; and that
error is not repeated on a schedule. The detach edge is checked not to
error, guarded by how long detection actually took, so a slow runner skips
the check rather than failing on the harness's own latency.

The containers run FIPS_TEST_MODE=default, not chaos. The chaos entrypoint
waits up to 30 s for every configured Ethernet interface before starting the
daemon, which is precisely the workaround under test — the daemon has to do
its own waiting here or the suite proves nothing.

Host-namespace ip(8) runs in a short-lived privileged container sharing the
host network and PID namespaces, for the reason chaos/sim/veth.py documents:
on macOS the containers live in the Docker VM, so ip(8) run on the macOS
host could never reach them.

Chaos ethernet transports are marked optional: true. In that harness a
neighbour's interface disappearing is the scenario, not a fault — node_churn
stops a container, which destroys its netns and with it both ends of every
veth it held, so a surviving node watches a required interface vanish for
the 30-90 s the neighbour is down, once per churn event. Reporting that at
error is right for a deployment and wrong for a harness that tears the
interface down on purpose; the mesh-wide zero-ERROR ceiling would have
failed on injected chaos rather than on a defect.

test(iface-binding): cover an interface present before the daemon starts

Every scenario in the suite created its interface after the daemons were
already running — that ordering is the boot race the suite was written for.
But it means both nodes could only ever reach Present through binder_loop,
so the inline bind in start_async, which is the ordinary case on a booted
router, had no end-to-end coverage at all. That is where the churn guard
went unseeded and the first detach stopped reaching node health, and no
existing case could reach it: they all detach from a binding the loop
created, which seeds the guard as a side effect.

Case (f) adds a third node whose single required interface exists before its
daemon does. The gate is what buys that ordering — the harness needs a
running container to have a netns to move a veth into, but the daemon must
not start until after the move, so node-c comes up parked on a file and the
harness releases it once the interface is in place. Then one detach, on a
binding the loop did not create, and the node must degrade.

Verified against the defect rather than only against the fix: with the guard
seed reverted, cases (a) through (e) all still pass and (f) is the only
failure. A regression test that has never been seen to fail is a claim, not
a test.

It also asserts the reverse edge, so Degraded stays a level rather than a
latch on this path too.

test(iface-binding): assert the fast path and the churn guard

Three gaps, two of them in tests that existed and asserted nothing.

**The netlink path was never asserted to be in use.** The 1 s poll is a
complete fallback and covers every wait in the suite, so the whole thing
passed with `open_link_socket()` hardcoded to Err — the fast path could have
been dead for a release and no test would have said so. The binder reports
which backing it got at startup, so case (g) asks it directly rather than
inferring from timing the poll would also satisfy, and the unit test that used
to write `let _ = w.is_event_driven();` now asserts it on Linux, where the
source is an unprivileged `AF_NETLINK` socket and falling back is a real loss
rather than a sandbox's prerogative.

**Churn damping had no end-to-end coverage**, which now matters twice over: it
bounds the recovery announcements, and since the detach edge withdraws peers
it is also the only thing bounding how often that withdrawal fires. Every flap
elsewhere in the suite is a single down/up with long settles either side —
exactly the shape the damper ignores. Case (h) drives four bindings that each
die inside `MIN_STABLE_BINDING`, asserts the guard engages, asserts it then
*suppresses* rather than merely counting, and asserts it is not a latch.

**`a_poisoned_binding_does_not_strand_the_transport` discarded its result.**
`let _ = eth.binding.tasks_alive();` left the entire point unasserted: reading
a poisoned lock as "alive" would have the binder believe a dead binding
healthy and never rebind, and treating it as an error would strand the
transport. `false` is what routes it back through detach and rebind, so say so.

`a_stop_racing_a_bind_leaves_nothing_behind` now asserts the error *kind*.
`bind_and_spawn` refuses at its presence probe long before the post-store
shutdown check, so `is_err()` alone passed on absence and would still pass
with that check deleted. The test keeps the coverage it genuinely has — stop
raises the flag before teardown, teardown leaves no socket and no loops — and
says plainly that the race it is named for needs a bind that succeeds, which
needs privilege no unit test has.

Both new cases were verified against the defect: with the netlink source
forced to Err, (g) fails; with `CHURN_THRESHOLD` raised out of reach, (h)
fails. Nothing else in the suite notices either.

One case was attempted and removed rather than shipped: `"interface
replaced"` cannot be produced deterministically, because the delete that
changes an ifindex fires a netlink event the binder acts on within
microseconds, so `gone` wins the race. It passed about one run in three.
reference/notes.md records the measurement and the two approaches that could
work.

Also fixes a real bug in the harness: `grep -q` under `set -o pipefail` exits
on its first match, `docker logs` takes SIGPIPE, and the pipeline reports
failure even though the line was found. That cost two false failures before it
was spotted; `log_count` reads the stream to the end.

test(chaos): cover an Ethernet rebind under active traffic

The one case dynamic interface binding had no coverage for anywhere: a
datagram crossing an Ethernet link while the interface underneath it goes away
and comes back.

No existing scenario could reach it, for two separate reasons.

`ethernet-only` and `ethernet-mesh` both run with `traffic.enabled: false`, so
no datagram crosses an Ethernet link in any test — `ethernet-only`'s own
comment says exactly that, and names framing, the length field that trims NIC
minimum-frame padding, and AEAD over Ethernet as unexercised because of it.

And `link_flaps` cannot produce a rebind whatever it is pointed at: it
simulates a down link with netem 100% loss, so the interface stays IFF_UP and
the presence machine never sees an edge. `ethernet-mesh` has had link flaps
enabled all along without once exercising a rebind.

`node_churn` is what actually moves an interface. Stopping a container
destroys its network namespace, deleting every veth in it — and deleting one
end of a veth deletes its peer — so a *surviving* node watches its Ethernet
interface disappear outright, and watches it return when the harness recreates
the pair on restart. That is a real detach and a real rebind, driven from
outside the daemon.

The new scenario is a 4-node Ethernet ring with traffic on and one node
churned at a time, with link flaps deliberately off so the only outage is a
genuine interface removal and a traffic shortfall cannot be ambiguous between
the two. Measured across four runs: 206-388 MB moved over Ethernet links while
interfaces were being taken away underneath.

It also needed an assertion that did not exist. Traffic results have always
been written to `iperf3-results.json` and never read, so a scenario carrying
`traffic.enabled: true` could have every session fail and still exit 0 on a
green control plane — and a rebind under load is precisely what a tree
snapshot cannot see. `min_traffic` counts sessions that finished with bytes
actually received, treating iperf3's top-level `error` and a missing `end`
block as zero, so a session only counts when it moved data.

The baseline is calibrated against four runs rather than assumed: `max_roots`
starts at the observed maximum plus one, and the site records the sample, its
size, and why four runs is thin. The first draft asserted a single root and
failed every run — the harness restores stopped nodes immediately before the
final snapshot, so a just-restarted node has not re-parented yet and is
briefly its own root. That is the scenario working.

Wired into both runners, since a chaos scenario on one side only makes "local
green" and "GitHub green" stop meaning the same thing; check-ci-parity was
confirmed to fail on a one-sided addition before this was committed.

The iface-binding suite's entry in the GitHub workflow's integration matrix
moves here from the commit that introduced the presence machine. That commit
declared the suite on GitHub before testing/iface-binding/ existed and before
testing/ci-local.sh knew about it, so testing/check-ci-parity.sh failed there
and the three workflow steps named files that were not yet in the tree.
Registering both runners in the commit that adds the suite settles both.
2026-09-10 19:18:09 +00:00
2026-09-06 20:20:56 +00:00
2026-02-22 20:52:55 +00:00
2026-09-06 20:37:39 +00:00

FIPS: Free Internetworking Peering System

banner License: MIT Rust Status

A self-organizing encrypted mesh network built on Nostr identities, capable of operating over arbitrary transports without central infrastructure.

FIPS is under active development. The protocol and APIs are not yet stable. See Status & roadmap below.

What FIPS does

A machine running FIPS becomes a node in the mesh with a self-generated cryptographic identity, tunneling existing IPv6 traffic over the mesh or bypassing IP altogether and letting natively written applications communicate directly with each other. In either case all traffic between nodes is end-to-end encrypted and authenticated.

The mesh is self-organizing and permissionless. Any node can join and reach any other node without a central address registry, routing configuration, or coordination server. Peering between nodes can be manually configured or use auto-discovery.

There are two equally-supported deployment modes.

As an overlay on top of existing IP networks, FIPS lets your node reach any other FIPS node wherever it sits: behind a NAT, on a different ISP, on a phone over cellular, on a laptop with only Bluetooth in range, or behind a Tor onion.

Ground up over raw Ethernet, WiFi, or Bluetooth, FIPS provides a complete permissionless network without any pre-existing IP infrastructure, ISP, or DNS. Any node that joins the link gets routable IPv6 addresses, peer discovery, and a path to every other node automatically. Support exists in OpenWrt for turning a router radio into a backhaul link and for creating an open access SSID so a phone or laptop can join without any configuration.

Either way, existing networking software runs over it unchanged — SSH, HTTP servers, file transfer, anything IPv6-native works the same way it would on a local network. Applications written to the FIPS native API skip that layer entirely and address each other by public key, with no IPv6 emulation.

Features

The mesh

  • Self-organizing mesh routing. Spanning-tree coordinates with bloom-filter-guided discovery; no global routing tables, no flooding.
  • Multi-transport. UDP, TCP, Ethernet, Tor, Nym, and Bluetooth (BLE L2CAP) ship today; transports compose on a single mesh and a node may run several at once.
  • Self-assigned cryptographic identity. secp256k1 / schnorr keypairs as node addresses; no registration, no central authority.
  • Two-layer encryption. Noise IK between peers (hop-by-hop) and Noise XK between mesh endpoints (independent end-to-end), with periodic rekey for forward secrecy.
  • (Optional) Nostr-mediated discovery and NAT traversal. Peers may publish endpoint adverts on public Nostr relays, exchange peering candidates, and establish direct paths through NATs using STUN-assisted hole punching. On the local network, mDNS LAN discovery finds peers directly without relays.

Getting traffic onto it

  • IPv6 adapter. A TUN interface maps each remote npub to an fd00::/8 address, so unmodified IPv6 software reaches mesh peers as <npub>.fips. Built-in .fips DNS resolver, with optional static name mapping via /etc/fips/hosts.
  • Native datagram API. A local program moves bytes between two public keys over the mesh, addressing a peer as npub:port with no IPv6 emulation and no TUN device in the path. connect and bind take a key and a port, and from there it is ordinary socket calls.
  • LAN gateway. Optional fips-gateway service folds an entire unmodified LAN into the mesh: outbound (LAN clients reach mesh destinations through a DNS-allocated virtual IPv6 pool and nftables NAT) and inbound (LAN-side services exposed to the mesh through 1:1 port forwards).
  • OpenWrt support. FIPS ships as an OpenWrt package. Routers run 802.11s between themselves as a bare L2 link, with FIPS supplying the encryption, authentication and routing over it. A second helper brings up an open !FIPS SSID, the same on every router, which a FIPS client joins over WiFi without configuration.

Running a node

  • Operator visibility. fipsctl CLI for control and inspection with time-series stats history queryable for any metric, fipstop TUI for live status with inline sparkline dashboards, and a JSON-line control socket on each binary for direct programmatic access.
  • Per-link metrics. RTT, loss, jitter, and goodput on every hop, plus mesh-size estimation, via the Metrics Measurement Protocol.
  • ECN congestion signaling. Hop-by-hop CE-flag relay with RFC 3168 IPv6 marking and transport kernel-drop detection.
  • Mesh-interface security baseline. Optional default-deny nftables policy for fips0 shipped as a packaged conffile (/etc/fips/fips.nft) with an operator drop-in directory (/etc/fips/fips.d/) and a disabled-by-default fips-firewall.service. The baseline polices only the mesh interface, leaving Docker, Tor, and the host firewall untouched.
  • Reproducible builds with toolchain pinning and SOURCE_DATE_EPOCH.

Quick start

Start from a released package. Every packaged platform in the table below gets an installer built and published per release, with checksums, on the releases page. Building from source produces the same artifacts and the same post-install state, so it is the path to take when you want to modify FIPS rather than run it.

On Debian or Ubuntu, download fips_<version>_amd64.deb (or _arm64.deb) and install it:

sudo dpkg -i fips_<version>_amd64.deb
sudo systemctl start fips fips-dns

This installs the daemon, CLI tools (fipsctl, fipstop), the fips-dns service that wires .fips name resolution into the host resolver, the optional fips-gateway service, systemd units, and a default /etc/fips/fips.yaml you can edit before starting. The package enables fips and fips-dns but starts neither, which is why the second command is there.

For macOS, Windows, FreeBSD, OpenWrt, the systemd tarball or a Nix flake, see docs/getting-started.md for the full multi-platform installation guide.

To join a live mesh and reach your first peer, follow the new-user tutorial progression starting at docs/tutorials/join-the-test-mesh.md.

Building from source

To build the Debian package yourself rather than downloading it:

git clone https://github.com/jmcorgan/fips.git
cd fips
cargo install cargo-deb
cargo deb
sudo dpkg -i target/debian/fips_*.deb

For the binaries alone, without an installer:

cargo build --release

Requires Rust 1.94.1+ (edition 2024). Linux, macOS, FreeBSD, and Windows run as standalone daemons. FreeBSD is packaged for x86_64 only; no aarch64 FreeBSD artifact is built or tested. Android is supported as an embedded crate rather than as a standalone daemon: a compile-gated library surface where the host app owns the TUN (a VpnService, for example) and reaches the built-in resolver through Node::dns_local_addr(). There is no Android daemon artifact and no host-app guide. Transport and feature availability varies by platform.

Feature Debian/Ubuntu Arch NixOS macOS OpenWrt FreeBSD Android Windows
UDP ✅ ✅ ✅ ✅ ✅ ✅ ✅ ✅
TCP ✅ ✅ ✅ ✅ ✅ ✅ ✅ ✅
Tor ✅ ✅ ✅ ✅ ✅ ✅ ❌ ✅
Nym ✅ ✅ ✅ ✅ ❌ ✅ ❌ ✅
Ethernet ✅ ✅ ✅ ✅ ✅ ❌ ❌ ❌
BLE ✅ ✅ ✅ ❌ ❌ ❌ ✅ ❌
Native API ✅ ✅ ✅ ✅ ✅ ✅ ❌ ❌
Package format .deb AUR flake .pkg .ipk / .apk .pkg ❌ ZIP

A column records what builds and runs in a packaged daemon, FreeBSD on x86_64 only. Native API is the native datagram API, which is off by default; Windows cannot carry it, because it has no SCM_RIGHTS with which to pass a descriptor. Package format names the artifact you install, and a ❌ there means the platform ships none. Windows is the odd one: its ZIP is an archive you unpack yourself rather than a package an installer consumes, and there is no MSI.

Five of these columns are Linux: Debian/Ubuntu, Arch, NixOS, OpenWrt and Android. Linux is not one target. Debian, Ubuntu, Arch and NixOS are the same glibc build, and what differs is the packaging: Debian and Ubuntu take the same .deb, Arch takes fips from the AUR, and NixOS uses the Nix flake described below. Only the .deb is exercised by an install test, by the deb-install suite across debian12, debian13, ubuntu22, ubuntu24 and ubuntu26; neither the AUR package nor the flake is. That suite runs on every push and pull request, on x86_64, against a .deb built by the same pinned container as the released one. It does not run at a tag, and the arm64 package is install-tested by nothing: no workflow installs a published artifact, so the released packages are checked by hand. OpenWrt is a musl target rather than glibc, and it takes an .ipk on 24.x and earlier or an .apk on 25 and later; both carry the fips-mesh-setup and fips-ap-setup helpers.

Android records what compiles for aarch64-linux-android under the CI cross-check and nothing more: no transport in that column is exercised on a device or an emulator, so read it as "compiles", not "verified here". Being an embedded crate rather than a daemon platform, it has nothing to install, which is what its ❌ package format records. The BLE cell is narrower still: the transport compiles, but the radio behind it is supplied by the embedding application rather than by FIPS, and no part of that path is device-tested.

On Linux, a source build requires libclang — the LAN gateway's nftables bindings are generated by bindgen at build time, which needs libclang.so on the build host. Install it before building (sudo apt install libclang-dev on Debian / Ubuntu); without it the build fails inside the rustables crate with an "Unable to find libclang" error. This is a build-time prerequisite only — it is not a runtime dependency, and the pre-built .deb artifacts do not need it.

BLE compiles on every glibc Linux target and on Android, and is excluded on musl. On glibc Linux, libdbus is a hard build prerequisite (sudo apt install libdbus-1-dev pkg-config on Debian / Ubuntu) — without it the build fails inside libdbus-sys rather than skipping BLE. The BlueZ daemon itself is a runtime dependency, not a build one. The OpenWrt ipk is a musl target, so it omits BLE.

Nym (mixnet) transport builds on all desktop platforms. The OpenWrt ❌ is provisional, pending verification of nym-socks5-client availability on the target; it will flip to ✅ only if confirmed buildable there.

Alternatively, the repo ships a Nix flake: nix develop drops you into a shell with the pinned toolchain and every build prerequisite (libclang, dbus, pkg-config) already provided, and nix build .#fips builds all four binaries with no host setup. See the Nix / NixOS section of packaging/README.md.

Documentation

docs/ is organised by reader purpose:

  • Tutorials — hand-held walk-throughs from a fresh install through to a participating mesh node, plus advanced deployments (gateway on OpenWrt, hosting services, ground-up two-device mesh).
  • How-to guides — operator recipes for specific tasks: firewall activation, Nostr discovery, Tor onion service, Bluetooth peering, 802.11s mesh backhaul and the open access SSID on OpenWrt, LAN gateway deployment and troubleshooting, MTU diagnostics, host aliases, persistent identity, unprivileged-user setup, UDP buffer tuning.
  • Reference — fips.yaml configuration, wire formats, control-socket protocol, CLI references for each binary, security posture matrix, Nostr events catalog, transport statistics inventory.
  • Design — protocol-level architecture and layer specifications. Start with fips-concepts.md for the framing, then fips-architecture.md for the protocol stack.
  • Release notes — per-version notes, including v0.5.1.

If you want to contribute, see CONTRIBUTING.md and testing/README.md.

Examples

  • examples/sidecar-nostr-relay/ — Run a strfry Nostr relay reachable exclusively over the FIPS mesh. The relay container shares the FIPS sidecar's network namespace and is isolated from the host network.
  • examples/sidecar-nostr-mixnet-relay/ — Single-container demo of FIPS peering through a mixnet (implemented with Nym): the FIPS daemon, the mixnet proxy, and a strfry Nostr relay all in one isolated container, with the direct route to the peer firewalled off so traffic provably crosses the mixnet.
  • examples/k8s-sidecar/ — Run FIPS as a Kubernetes Pod sidecar. The sidecar creates fips0 in the Pod's shared network namespace so every other container in the Pod gets mesh access without modification.
  • examples/wireguard-sidecar-macos/ — Reach the FIPS mesh from a macOS host through a local Docker container over a WireGuard tunnel. Only traffic destined for fd00::/8 transits the sidecar; regular internet traffic continues to use the host network.

Project structure

src/          Rust source: library + fips, fipsctl, fipstop, fips-gateway binaries
docs/         Documentation: tutorials, how-to, reference, design
packaging/    Debian, AUR, systemd tarball, OpenWrt ipk/apk,
              macOS .pkg, FreeBSD .pkg, Windows ZIP
examples/     Deployment examples (Nostr relay, K8s sidecar, macOS WireGuard)
testing/      Docker-based integration test harnesses + chaos simulation

Status & roadmap

FIPS is at v0.6.0-dev on the master branch. v0.5.1 is the current release, a maintenance release on the v0.5.x line that makes the Linux packages install and run on Debian 12 and Ubuntu 22.04, where every artifact from v0.3.0 through v0.5.0 installed and then could not start. v0.5.0 was the last feature release; this development line continues the testing-and-polishing track toward v0.6.0. The core protocol works end-to-end over UDP, TCP, Ethernet, Tor, Nym, and Bluetooth on a global, public test mesh of thousands of nodes. v0.5.0 added FreeBSD as a packaged platform, OpenWrt setup helpers for an 802.11s mesh backhaul and an open client SSID, an Android embedding interface, a native datagram API addressed by public key, and published node health with a bounded shutdown drain. New wire-format work continues to be staged on the next branch for the subsequent release line.

What works today

  • Spanning-tree construction with greedy coordinate routing.
  • Bloom-filter-guided destination discovery (no flooding, single-path with retry).
  • Two-layer Noise encryption (IK at the link, XK at the session) with periodic hitless rekey for forward secrecy at both layers.
  • Persistent or ephemeral node identity with key-file management.
  • IPv6 TUN adapter with built-in .fips DNS resolver and multi-backend auto-configuration (systemd dns-delegate, systemd-resolved, dnsmasq, NetworkManager).
  • Native datagram API for FIPS-aware applications (npub:port addressing without the IPv6-shim path): off by default, with a surface that may still change.
  • Static hostname mapping (/etc/fips/hosts) with auto-reload.
  • Per-link metrics (RTT, loss, jitter, goodput) and mesh size estimation.
  • ECN congestion signaling (hop-by-hop CE relay, IPv6 CE marking, kernel-drop detection).
  • UDP, TCP, Ethernet, Tor, Nym (mixnet), and BLE transports (BLE via L2CAP CoC with per-link MTU negotiation).
  • Nostr-mediated overlay endpoint discovery and UDP hole punching for NAT traversal, plus mDNS LAN discovery for local peers.
  • LAN gateway (fips-gateway) with both outbound (LAN-to-mesh) and inbound (mesh-to-LAN port-forwarding) modes.
  • Peer ACL: per-npub allow / deny admission control at the link layer; opt-in mesh-firewall baseline at fips0 ingress.
  • Runtime inspection and peer management via fipsctl (including fipsctl probe for reachability diagnosis and fipsctl address for mesh-address derivation) and fipstop.
  • Reproducible builds with toolchain pinning and SOURCE_DATE_EPOCH.
  • Node lifecycle and health reporting (Starting, Running, Degraded, Failed, Draining) with a fatal start when no transport comes up and a bounded shutdown drain window.
  • OpenWrt setup helpers for an 802.11s mesh between routers (fips-mesh-setup) and for the open !FIPS client SSID (fips-ap-setup).
  • Linux (Debian, systemd tarball, OpenWrt .ipk and .apk, AUR), macOS (.pkg), FreeBSD (.pkg, x86_64 only), and Windows (ZIP, service) packaging.
  • Docker-based integration and chaos testing.

Near-term priorities

  • Security audit of the cryptographic protocols.

Longer-term

  • Packaged mobile applications: an Android host app, and iOS. The Android embedding interface ships today (see Building from source); what is absent is a packaged app on either platform.
  • Bandwidth-aware routing and QoS.
  • Protocol stability and a versioned wire format.
  • Published crate.

License

MIT — see LICENSE.

S
Description
The Free Internetworking Peering System
Readme MIT
64 MiB
Languages
Rust 81.7%
Shell 14.3%
Python 3.4%
PowerShell 0.2%
Nix 0.1%
Other 0.1%