mirror of
https://github.com/jmcorgan/fips.git
synced 2026-10-05 11:08:25 +00:00
An interface-bound transport is a long-lived object that is *sometimes bound*. The interface it names may not exist when the daemon starts, may appear minutes later, may vanish and return mid-operation, and may never appear at all. Until now the first observation was final: an interface missing at start was logged once and skipped for the life of the process, and one that disappeared at runtime published a health change but was never rebound. On OpenWrt that is a live bug. procd starts fips before wifi has created fips-mesh0 / fips-ap0, both transports are skipped and never retried, and the 802.11s peer link forms anyway (that is mac80211, not the daemon) — so the node looks healthy and reaches nothing. Give every interface-bound transport presence state and a binder task. start_async now returns Ok with the transport ABSENT rather than failing; the binder binds when the interface appears, tears the socket down when it goes away, and rebinds when it returns. Start-time absence and runtime detach are one transition and one code path. TransportError::InterfaceUnavailable makes absence branchable: a missing interface and a typo'd interface name were the same flat StartFailed (String), so nothing downstream could tell a state from a fault. A bind failure that is *not* absence — no CAP_NET_RAW, no readable /dev/bpf* — still fails the start, because retrying a socket that can never open behind a Degraded nobody is watching is worse than dying loudly at boot. **Presence is IFF_UP, not IFF_UP|IFF_RUNNING.** Carrier and bindability are different questions and only the second belongs in a bind gate. An AF_PACKET socket on a carrier-less bridge is valid and starts carrying traffic the instant a member port comes up, with no rebind. Gating on carrier would report a wifi-only router Degraded forever for an empty br-lan, turn every carrier flap the socket would have survived into an unbind/rebind cycle, and deadlock an 802.11s interface that reports RUNNING only once it has peered — peering needs beacons, beacons need a bound socket. Carrier is reported beside presence in show_transports instead. **A name is not a device.** Both backends bind by device: AF_PACKET stores sll_ifindex, a BPF descriptor follows the interface it was attached to. A netdev deleted and recreated under the same name leaves the socket attached to something gone while the name resolves perfectly well, and nothing notices — a stale AF_PACKET socket never becomes readable, so the receive loop neither errors nor exits, and send failures go to the caller rather than the binder. A listen-only node (announce: false, no beacon sender to fail) sat present and deaf indefinitely after a wifi reload. The bound index is captured at bind and compared on every poll. A name reappearing with a different MAC is different hardware: cached neighbors are dropped rather than resumed onto. **Detection is event-driven** where the kernel offers a source — netlink RTNLGRP_LINK on Linux, PF_ROUTE on macOS and FreeBSD — with a 1 s getifaddrs poll underneath as a backstop. Link-event payloads are not parsed: an event is a hint to re-run the probe, which is cheap and authoritative, and a parser's bugs would be presence bugs. Probes are coalesced to ten a second because PF_ROUTE has no group filter and delivers every routing message on the host; a persistently failing event source is logged once, backed off, and abandoned for the poll after five errors rather than spinning a core silently on ENOBUFS. **Degraded becomes a level, not a latch.** The supervisor's reason set was monotonic, which was correct while no child could recover; with recovery it would have come to mean "something broke at some point since boot" rather than "something is broken now". Absence lives in its own reversible set, health is recomputed on every transition in both directions, and an absent transport still counts as up so a single-interface node that boots before its wifi degrades rather than exiting on NoTransports. There is deliberately no restart action in the FSM: rebinding belongs next to the file descriptor, and a supervisor-authored retry would be a second mechanism racing the first for the same socket. **The node's egress MTU follows the bound set.** transport_mtu filters on is_bound(), not is_operational() — an interface-bound transport is operational from the moment it starts, so the weaker predicate let hardware that had never appeared clamp the whole node's IPv6 MTU. Because that minimum now moves at runtime, the TUN reader and writer read the TCP MSS ceiling from a shared atomic instead of a u16 captured at spawn; every other consumer (show_status, the snapshot, the session-layer fragmentation check) already read it live, so the clamp was the one place the daemon could report one effective MTU and enforce another. It moves in both directions: a narrow interface appearing tightens it, its departure releases it. MSS is negotiated per connection, so a change binds connections opened after it. **Logging is per edge, never per attempt**, with one deadline after it. The edge itself is not an error — info at boot, warn on a runtime detach, since an interface bound a moment later is the ordinary case this mechanism exists to absorb and crying error at t=0 then "recovered" at t=0.2s is the failure the rule exists to prevent. Ten seconds is the whole grace: past it, absence is no longer a race against a radio or a container, so a required interface still missing is reported once at error. Start-time absence and a runtime detach share that one deadline, as they share everything else here. Said once, not repeated: how long an absence has lasted is state, published as interface.since_secs and as Degraded, and a monitor can threshold it per deployment rather than the daemon compiling a schedule in. Successful rebinds are damped. Backoff covers failed binds; the nastier case is binds that keep succeeding into a socket that dies moments later, which a receive loop giving up on a persistent error while the interface stays UP produces once per second forever. Consecutive bindings dying inside ten seconds back off on the 1 s → 30 s curve, and past three the binder stops announcing each bind as a recovery until one lasts. The receive loop backs off and exits on a dead socket instead of spinning on Err with a warn per iteration, and the ad-hoc ENXIO socket reopen in the beacon sender is gone: both hand recovery to the presence machine, one mechanism for every cause rather than one hack per symptom. Beacons pause while absent because the task simply does not exist then. transports.ethernet.*.optional (default false) decides how loudly absence is reported. Naming an interface is a statement that you expect it, so the default complains: Degraded, and the error at the deadline. optional: true is silent — info on the edge, no health impact, no error — for hardware legitimately not always there. It describes the interface's presence, not the transport's importance, and no value of it makes a missing interface fatal at startup. show_transports grows an `interface` block — presence, carrier, policy, how long the phase has been held, bind count, failed attempts. The original boot race was expensive precisely because nothing an operator could see said the node was deaf. ci: exercise the presence probe against musl and an address-less interface Two legs, both about the assumption the presence probe rests on. OpenWrt — the platform dynamic interface binding exists for — is musl, and musl reimplements getifaddrs independently of glibc. `interface_present` reads ifa_flags out of it, and the interfaces this feature exists for (fips-mesh0, fips-ap0) are unbridged with no IP address at all, which is exactly where getifaddrs implementations differ. Every other leg is glibc, so without this one the probe was asserted on a libc no test had ever run it against, on the target it was written for. Building for the musl target rather than inside an Alpine container keeps the leg to a cross-build plus a run, without a second toolchain image to maintain. The address-less interface is created on both this leg and the glibc one. Pinning the contract on both libcs is what turns "glibc and musl agree about getifaddrs" from an assumption into a checked fact, and makes the glibc leg fail first if glibc is the one that changes. Loopback cannot stand in: it has 127.0.0.1, so probing it asks whether getifaddrs works rather than whether it reports this. Deliberately not tolerant of failure — a guard that quietly does not run is worse than no guard. The iface-binding suite is deliberately not added to the integration matrix here. Its files and its testing/ci-local.sh registration arrive in the next commit, and the GitHub workflow's matrix entry now goes with them, so both runners gain the suite together and testing/check-ci-parity.sh holds at every point in this history. docs: document dynamic interface binding The transport-layer design gains an Interface Presence section: the three deployment scenarios that motivated it, the presence machine and its two invariants, why presence is IFF_UP and not IFF_UP|IFF_RUNNING, interface identity and the recreated-netdev case, the detection sources and their two rate limits, the optional policy, health as a level rather than a latch, the egress-MTU consequence, the logging rules, and what the mechanism retires. It records why the error deadline is a single ten-second window rather than a repeating severity ladder, so the ladder does not come back. The control-socket and fipsctl references described show_transports without its `interface` block, and the transport-layer state machine still implied that Up meant bound. Both now say what Up and presence each describe and how they differ, since "transport up, interface absent" is a normal state an operator will meet and would otherwise read as a contradiction. The configuration reference documents `optional` with the absence, log and retry behaviour of each setting, and the ground-up tutorial's dongle example gains optional: true, which is what that example is actually for. fix(transport): pair every presence edge the binder publishes Six defects, all in the wiring around the presence machine rather than in the machine itself. The state machines were well tested; what went untested was how the binder feeds them, and every one of these lives there. **The first detach after a clean start never reached node health.** start_async binds inline and publishes that edge itself, before binder_loop exists. The loop then built a fresh ChurnGuard, which had therefore recorded no bind — and detached() takes `announced` to decide whether an edge is owed, so it asked for no retraction. A node that booted with its interface present and then lost it kept reporting Full for as long as the interface stayed away, while show_transports said absent and the log said detached. stabilized() could not repair it either: it early-returns while bound_at is None, which it is for a bind the loop did not perform, so the guard stayed unseeded until a second detach happened to fix it. The loop now seeds the guard from the binding it inherits. This is not the cable-unplug case. Presence is IFF_UP, so a carrier loss is correctly not a detach at all; it takes an admin down, a netdev delete, or a device removal — a wifi reload, a hostapd restart, a dongle pulled. **An optional interface published nothing, and the MTU floor rode the same channel.** Filtering the edge at the source conflated two questions that happen to share a transport: whether node health should move, and whether the bound set changed. Only the first is policy. The second determines the node's egress MTU floor, and refresh_tun_mss_ceiling has no other trigger — so an optional transport binding or unbinding at runtime left the TUN MSS clamp derived from a bound set that no longer existed, reporting one effective MTU and clamping to another. That is the defect the MssCeiling work was written to remove, reintroduced for exactly the transports the shipped OpenWrt config marks optional: five of seven. TransportPresence gains health_relevant; the edge always goes out and only health is filtered. **A permanent bind fault went unlogged if any absence race preceded it.** record_attempt ran ahead of the InterfaceUnavailable arm, so a probe that won a race the bind then lost burned the counter that gates the only error a real fault ever gets — emitted on attempts == 1. A flapping interface followed by CAP_NET_RAW being dropped or /dev/bpf* exhausting therefore reported nothing at all, for the life of the process, and the operator got report_sustained_absence's "still missing" instead: wrong, for an interface that is present. An absence race is not a bind attempt and no longer counts as one. **The watcher's give-up did not stick.** Abandoning the event source parked on pending() *inside the changed() future*, and that future is constructed fresh on every pass of the binder's select! and dropped whenever the poll ticker wins. So the next pass re-read the dead socket, re-counted the error and re-logged "not recoverable" — once per wake-up, forever, which at a 1 s tick across the shipped seven-transport config is seven warnings and seven failing syscalls a second on flash-backed logging. The 100 ms backoff lived in the dropped future too and never applied across passes. Abandonment is now state on the watcher. **A zero-length read livelocked the binder.** try_io clears readiness only on WouldBlock, which the sibling error arm handles by hand and this one did not — so breaking out left readable() instantly ready with nothing to read, and the loop never returned Pending. That starves the select! of its ticker entirely and takes presence detection down with it. Cleared and treated as a fault so the give-up path applies. **A failed presence probe read as an absent interface.** getifaddrs is a netlink dump and fails for reasons that have nothing to do with the interface: ENOBUFS under memory pressure, EMFILE or ENFILE under fd exhaustion, since it opens a socket of its own. Answering "not present" there tore down a working socket and degraded health over a transient syscall failure, undiagnosably, and under fd exhaustion the rebind could not have succeeded anyway. interface_has_flags now distinguishes the two; the detach gate holds its binding when the kernel will not answer, while the bind gate still treats it as absence and retries. **On macOS the reader thread could not die, so a dead socket read as live.** Any read() failure other than EBADF — ENXIO being the one that matters, which is what BPF answers once the interface it was attached to is torn away — reset the parse buffer and continued. The thread never returned, so the channel never closed, so recv_from never failed, so the tokio task never exited, so tasks_alive() reported a dead socket as a live one. Detach detection on macOS reduced to the name and the index, and an interface reset in place left the transport present and deaf. It now gives up after a streak, and the return closes the channel the binder is actually watching. Two smaller pairings while here: stop_async retracts the edge it would otherwise leave standing for a socket that is gone, and start_async hands a refused edge to the binder to retry rather than dropping the only edge either consumer will see until the interface next moves. The absence deadline is stamped at start rather than at construction, so a transport staged for longer than the bring-up window no longer reports sustained absence on its first tick having given the interface no window at all. feat(config): reject impossible ethernet interface names at load Config::validate never inspected transports.ethernet, so an empty interface name, a name past the kernel's 15-byte limit, and two transports naming the same netdev all loaded cleanly and failed only at runtime — the first two as a permanent absence indistinguishable from an interface that has not been created yet. That indistinguishability is deliberate and worth keeping: waiting is the right answer for an interface the operator has not made yet, and the daemon cannot know which of the two it is looking at. Which is exactly why the syntactic gate earns its place. A name that is *impossible* is the one case still separable from "not there yet", and without the check a typo costs a permanently Degraded node whose only symptom is an interface that never arrives — the failure mode the presence machine exists to make legible, reintroduced one level up. Syntax only. Whether a well-formed name exists stays the binder's question, asked once a second, forever. The duplicate check is a different fault: two transports on one netdev means two sockets on the same device at the same ethertype, each receiving every frame the other does. test(transport): close the vacuous and uncovered branches Six tests that asserted nothing, or asserted less than they claimed. `a_bind_fault_still_fails_the_start` returned early whenever the socket *could* be opened, so it was vacuous as root and on any developer machine with a group-readable /dev/bpf* — the fail-fast path it is named for went unchecked exactly where someone was most likely to run it. It now asserts in both halves, and the privileged half is worth more than the fix: a present, bindable interface binding inline is the ordinary case on a booted router, and no other unit test reaches it, because every other one here names an interface that does not exist. The bind-success path had no unit coverage at all. `an_interface_with_no_addresses_is_still_present` returned early when the fixture was absent — no fixture, pass. It still has to skip on a machine with no address-less interface, so the guard is the runner declaring that it has fixtures: CI now sets FIPS_TEST_REQUIRE_FIXTURES beside FIPS_TEST_ADDRLESS_IFACE, and the test fails rather than skips if the fixture step is ever removed or renamed. Its `let _ = interface_carrier(...)` is now asserted too: if presence and carrier ever collapsed into one read, a carrier-less bridge would report absent and the whole IFF_UP-not-IFF_RUNNING decision would be silently undone. `policy_labels` checked `Required.as_str()` and not `Optional.as_str()`, so a swapped pair would paint every expected interface as the tolerated kind and stay green. `Presence::as_str` had no test at all — `binding` was never observed by anything, anywhere. Three binder branches had no coverage: a transport restarting (the second `start_async` clearing the previous run's stop flag — only a second *stop* was tested, so a transport that could never restart passed everything), the episode clock being restamped at start rather than at construction, and the refused-edge retry actually delivering. The last one matters most: the existing test filled the channel and dropped the receiver, so a slot that captured an edge and never re-sent it would pass while health sat on a stale level forever. And the hardware-change boundary `record_bind` returns, which the neighbour flush hangs off: false on a first bind (or every clean start would drop a cache it had just built) and true once, not stickily, on a MAC change. Each new test was verified against the defect it guards — comment out the `shutdown.store(false)`, the `mark_starting()`, or the seeded `unpublished`, and the corresponding test goes red while the rest stay green. fix(test): correct the dummy-carrier assertion, and pin Darwin's presence probe The carrier assertion added ind6240698was wrong and would have failed both Linux legs. A `dummy` interface brought up reports `<BROADCAST,NOARP,UP, LOWER_UP>` — `IFF_RUNNING` is set, so it *has* carrier. Verified against the exact fixture CI builds, `addrgenmode none` and all, rather than against the comment: the original code discarded the result and its comment claimed "up but not running", which is what made asserting it look safe. So the fixture pins address-less *presence* and cannot demonstrate the presence-vs-carrier split at all — an interface up with no carrier is a bridge with nothing plugged in, which no fixture here creates. `carrier_is_reported_separately_from_presence` pins that split from the other side. The corrected assertion is Linux-only, because the expected answer is a property of the fixture device rather than of the code. That is also what lets the macOS fixture land. The Linux legs pin the address-less contract on glibc and musl, but the BSD-derived `getifaddrs` the macOS backend actually calls had no coverage — the test skipped itself silently on that runner, which is precisely the shape the previous commit was removing. `feth` is macOS's fake-Ethernet pseudo-interface and is created address-less; the step fails the leg rather than testing the wrong thing if the runner hands it an address anyway, mirroring why the Linux step needs `addrgenmode none`. Both branches of the fixture guard were exercised: unset skips and passes, and declared-but-missing fails loudly. The corrected test was run against a real Linux dummy inside a container, not reasoned about. fix(ci): put the macOS presence fixture on the macOS job The `feth` fixture step added in6b9faa0clanded on the Linux `test` job, not on `test-macos`, and has failed CI ever since: create: Host name lookup failure ifconfig: `--help' gives usage information. That is Linux net-tools `ifconfig`, which has no `create` subcommand. So the Linux job ran two address-less-fixture steps — its own correct `ip link add type dummy` one, then a macOS one that cannot work there — while `test-macos` had none at all, leaving the Darwin `getifaddrs` path exactly as uncovered as before. Cause was a pattern-anchored edit: `Install cargo-nextest` followed by `Run unit tests` appears in three jobs, and the insert hit the first match. Moving it then hit the *last* match, which is `test-windows`. It is now placed by job boundary rather than by pattern, and verified per job: `test` and `test-musl` carry the `ip link`/dummy fixture, `test-macos` carries `ifconfig`/`feth`, and `feth` appears exactly once in the file. The verification that missed this was counting steps in `test-macos` and reading 6 as confirmation. Six was the count *before* the insert; seven is what a successful insert looks like. Now asserted by job and by which tool each fixture step uses, so a step in the wrong place fails the check rather than matching a total. test(transport): cover the detach branch and the stop-race check Two of the three untested branches, by two different routes. **The detach decision is now a pure function.** `classify_detach(gone, replaced, dead)` replaces the inline three-way `if` in the binder loop, and all eight input combinations are asserted, plus the precedence between them and the `reason=` labels the integration suite and operators grep for. This does not make `Replaced` or `SocketDied` reachable from a test — both need a bind that succeeded and then a specific external event, and the integration suite cannot arrange the recreate deterministically either, for the reason recorded in reference/notes.md: the netlink event from a delete is acted on within microseconds, so `Gone` wins that race in practice. What it does is split the untested thing in two. The three inputs each already had tests (`interface_present`, `device_replaced`, `tasks_alive`); the branch between them did not, and that half is now total and exhaustive. Precedence is asserted rather than assumed: an interface that has gone away has also trivially been "replaced" and its socket is also dead, so the most specific true statement has to win, and callers pass `replaced` already masked by `!gone` — the function no longer depends on them having masked correctly. **The post-store shutdown check is now tested directly.** It needs a bind that *succeeds*, which is why `a_stop_racing_a_bind_leaves_nothing_behind` could never reach it: that test's interface does not exist, so `bind_and_spawn` refuses at the presence probe several steps earlier. Binding loopback as root reaches it, and the assertion is that a bind completing after a stop undoes its own socket, its own loops, and its own presence. Unprivileged runners skip it, but loudly: a runner declaring `FIPS_TEST_PRIVILEGED` and unable to open a raw socket fails instead of skipping, the same guard shape as `FIPS_TEST_REQUIRE_FIXTURES`. No CI leg sets that yet — unit tests run unprivileged — so today it exercises the branch only under a root container. Verified there, including the negative control: with the post-store check removed the test fails, and with it present it passes. test(transport): run the bind-success half in CI, and cover the replaced device Adds the privileged unit-test step the rest of this depends on, then uses it. **The privileged step.** Every unit-test leg runs unprivileged, so `PacketSocket::open` cannot succeed on any of them and everything past a successful bind runs nowhere in CI: the post-store shutdown check, the `Present` arm of the binder loop, `bind_now` itself. The step builds as the runner user and executes only the test binary under sudo — running `cargo` as root would use root's CARGO_HOME and discard the cache the job just restored. The binary-path extraction was verified locally before being written into the workflow. `FIPS_TEST_PRIVILEGED` is what makes it honest. Tests needing a raw socket skip quietly without it; with it set, a test that cannot open one fails and says so. A runner that stops granting the capability shows up as a red leg rather than as silence. **The replaced device.** `"interface replaced"` — a netdev recreated under the same name, the `wifi reload` case in #125 — could not be reached by the integration suite: with link events live the kernel's `RTM_DELLINK` is acted on within microseconds, so the binder observes "gone" first and takes the branch already covered. Attempting it there passed about one run in three. Rather than race the binder, the test asserts the *inputs*: after a real delete-and-recreate the name still resolves and the bound index no longer matches. Paired with the exhaustive classifier test, which pins that `(gone: false, replaced: true)` maps to `Replaced`, the path is covered without depending on scheduling. A watcher-disable hook was written for the racing approach and removed once this one made it unnecessary. **Two smaller gaps.** `report_sustained_absence` had no direct test — the integration suite infers it from counting ERROR lines, and ten seconds of real time is how a deadline ends up asserted by proxy. A test-only `backdate_for_test` ages the episode clock instead. And the probe floor was asserted only as a relation between two constants; it now has behaviour, via a `probe_delay` helper extracted from the loop. That extraction was not neutral, and its own test caught it: `checked_sub` yields `Some(0)` at exact equality where the loop used a strict `<`, so a zero-length sleep would have replaced no sleep at all. Harmless in effect, wrong in meaning, and fixed. Every new test was run against the defect it guards, in a root container: remove the post-store shutdown check, or make `device_replaced` always answer false, and the corresponding test fails. Mark publish_presence must_use. The defect this commit fixes was a discarded return value, so the fix is made self-guarding: a future call site that drops the pending edge instead of storing it now fails the lint rather than silently reintroducing the bug. Option is not must_use in std, unlike Result, so the attribute has to be explicit.
530 lines
21 KiB
Markdown
530 lines
21 KiB
Markdown
# Build a Mesh from the Ground Up
|
|
|
|
The earlier tutorials in this progression rode existing IP — your
|
|
daemon reached `test-us01` over the public internet through your
|
|
ISP, your ISP's upstream, and however many hops separate you from
|
|
the test node. That is the *overlay* deployment mode of FIPS:
|
|
useful, but not the new ground.
|
|
|
|
This tutorial is about the other mode. Two devices, a wire (or a
|
|
radio link) between them, no IP between them, and FIPS daemons on
|
|
each end. The two daemons discover each other over the raw link,
|
|
peer over Noise, and bring up an end-to-end mesh with addressing,
|
|
naming, and reachability — all from layer 2 up. There is no DHCP,
|
|
no router, no upstream. The mesh is the network.
|
|
|
|
This is the deployment mode FIPS was designed for. Overlay mode
|
|
exists because riding existing IP is a useful convenience; the
|
|
ground-up mode is what FIPS uniquely enables.
|
|
|
|
> **The two modes are not exclusive.** A node can carry overlay
|
|
> peers and ground-up peers at the same time — different transports
|
|
> on the same daemon. If you have already worked through
|
|
> [join-the-test-mesh](join-the-test-mesh.md), the static peer to
|
|
> `test-us01` you configured there can stay in place; the Ethernet
|
|
> peer you add in this tutorial sits alongside it. Traffic flows
|
|
> through whichever path is shortest by mesh metric, and a node on
|
|
> one side can reach a node on the other through your machine
|
|
> acting as a bridge between the two.
|
|
|
|
## What you'll build
|
|
|
|
```text
|
|
┌──────────────────────┐ raw Ethernet frames ┌──────────────────────┐
|
|
│ node A │ ─────────────────────── │ node B │
|
|
│ npub1aaa… │ EtherType 0x2121 │ npub1bbb… │
|
|
│ fips0 fd97:..:A │ no IP between them │ fips0 fd97:..:B │
|
|
└──────────────────────┘ └──────────────────────┘
|
|
│ │
|
|
│ a single Ethernet cable │
|
|
│ (or both NICs on the same │
|
|
│ unmanaged switch — no DHCP, │
|
|
│ no router, no IP at all) │
|
|
└─────────────────────────────────────────────────┘
|
|
```
|
|
|
|
Two machines, each running `fips`, joined by a physical Ethernet
|
|
link. After the worked example:
|
|
|
|
- The two daemons have discovered each other via L2 beacons on
|
|
the link, peered over Noise IK, and brought up an FMP link.
|
|
- Each `fips0` adapter has a routable mesh address; each can
|
|
ping the other by `<npub>.fips`.
|
|
- Nothing between the two machines speaks IP. The link carries
|
|
raw FIPS frames at EtherType `0x2121`.
|
|
|
|
The whole exercise should take about twenty minutes if you have
|
|
the hardware ready.
|
|
|
|
## Why ground-up
|
|
|
|
Most networking tutorials assume IP is already there: an address
|
|
arrived from DHCP, a default gateway routes you onward, DNS
|
|
resolves names. FIPS does not need any of that. Two devices and
|
|
a way to deliver bytes between them at layer 2 is enough — FIPS
|
|
supplies the rest:
|
|
|
|
- **Identity**: each daemon has an npub (the same kind you saw
|
|
in the overlay tutorials). Nothing in the ground-up case
|
|
depends on a network identity from a router; the npub is the
|
|
identity.
|
|
- **Addressing**: the `fips0` adapter takes an `fd97:...` ULA
|
|
derived from the npub. No DHCP. No SLAAC. The address is
|
|
cryptographically tied to the identity.
|
|
- **Neighbor detection**: each daemon broadcasts a small beacon on the
|
|
link advertising its npub; the other daemon's listener picks
|
|
it up and dials in over the same link.
|
|
- **Routing**: the FIPS mesh layer builds its own spanning tree
|
|
across whatever links it has. Add a third node (peered to
|
|
either A or B) and traffic reaches it transparently.
|
|
|
|
The point is not that ground-up replaces overlay. It's that
|
|
overlay is one of two modes the same daemon supports, and
|
|
ground-up is what unlocks the use cases overlay cannot —
|
|
ad-hoc local meshes, partitioned networks, situations where
|
|
no IP infrastructure exists or can be relied on.
|
|
|
|
## Prerequisites
|
|
|
|
Two devices (call them **node A** and **node B**) and a way to
|
|
join them at layer 2:
|
|
|
|
- Ethernet (the worked example): a direct cable between two
|
|
modern NICs (auto-MDI/MDIX handles crossover for you), or
|
|
both machines on a small unmanaged switch with no DHCP
|
|
server. USB-Ethernet dongles work; a typical "USB-to-RJ45"
|
|
adapter is fine on either end. The link does **not** need
|
|
to be the machine's primary network interface — a second
|
|
NIC dedicated to the mesh is the cleanest setup.
|
|
- WiFi (a one-line variation, covered later): both machines
|
|
associated to a common AP that has client (station)
|
|
isolation **off**.
|
|
- Bluetooth LE (a separate worked example via a how-to,
|
|
covered later): two BLE-capable Linux hosts within roughly
|
|
10 metres line of sight.
|
|
|
|
On both nodes:
|
|
|
|
- `fips` installed and running, per [getting-started](../getting-started.md).
|
|
- A persistent identity from
|
|
[persistent-identity](persistent-identity.md). Ephemeral
|
|
identities work, but on each restart the npub regenerates
|
|
and you'll have to re-check `fipsctl show peers` to see the
|
|
new identity. Persistent makes the lesson stick.
|
|
- The daemon running with `CAP_NET_RAW` (the shipped systemd
|
|
unit runs as root and gets this for free; running
|
|
interactively from a user account requires `setcap` —
|
|
noted at the relevant step below).
|
|
|
|
You do **not** need:
|
|
|
|
- An IP address on the chosen interface. The Ethernet
|
|
transport opens a raw socket directly; the kernel does not
|
|
need to assign an IP to the NIC.
|
|
- A default route. The mesh routes itself.
|
|
- DNS resolution between the machines via any external
|
|
service. The local `.fips` resolver supplies names from
|
|
the npubs the daemons exchange.
|
|
|
|
## Step 1: Identify the link interface on each node
|
|
|
|
On each node, list the network interfaces and pick the one that
|
|
sits on the link between the two machines. If it's a dedicated
|
|
NIC for the mesh, that NIC has no other purpose; if it's a
|
|
USB-Ethernet dongle, plug it in first so the kernel names it.
|
|
|
|
```sh
|
|
ip link show
|
|
```
|
|
|
|
Pick out the interface name. Common forms:
|
|
|
|
- `enp3s0`, `eno1` — built-in NICs under predictable naming.
|
|
- `eth0` — older or container-style naming.
|
|
- `enxAABBCCDDEEFF` — USB-Ethernet dongles often appear under
|
|
this MAC-derived form.
|
|
|
|
Bring the interface up if it isn't:
|
|
|
|
```sh
|
|
sudo ip link set dev <interface> up
|
|
```
|
|
|
|
Confirm:
|
|
|
|
```sh
|
|
ip -br link show <interface>
|
|
```
|
|
|
|
You want `UP` and `LOWER_UP` in the flags. The interface does
|
|
not need an IP address — `LOWER_UP` indicates the NIC sees
|
|
carrier (cable plugged into something at the other end), and
|
|
that is all the Ethernet transport needs.
|
|
|
|
For the rest of the tutorial we'll write the chosen interface
|
|
as `<eth>`. Substitute the actual name on each node when you
|
|
run the commands. Note that node A and node B may have
|
|
different interface names — that is normal.
|
|
|
|
> **No IP needed.** If your chosen interface has an address
|
|
> from a previous DHCP lease, leave it alone or remove it with
|
|
> `sudo ip addr flush dev <eth>` — the FIPS Ethernet transport
|
|
> uses raw `AF_PACKET` sockets that bypass the IP stack
|
|
> entirely. The interface needs to be `up` and `LOWER_UP`,
|
|
> nothing more.
|
|
|
|
## Step 2: Configure the Ethernet transport on each node
|
|
|
|
Edit `/etc/fips/fips.yaml` on **both** nodes. Under
|
|
`transports:`, add an `ethernet:` block. The key settings are
|
|
the four neighbor flags — both nodes must opt in to all four.
|
|
`listen` defaults on; the other three default to off:
|
|
|
|
```yaml
|
|
transports:
|
|
ethernet:
|
|
interface: "<eth>" # the name from Step 1
|
|
announce: true # broadcast our beacon on the link
|
|
listen: true # listen for beacons (default; shown for clarity)
|
|
auto_connect: true # dial peers we discover
|
|
accept_connections: true # accept dial-ins from peers we discover
|
|
```
|
|
|
|
Each flag does one thing:
|
|
|
|
- `announce: true` — emit a small beacon every
|
|
`beacon_interval_secs` (default 30s) carrying our npub.
|
|
- `listen: true` — listen for incoming beacons; populate a
|
|
candidate-peer list keyed by source MAC and observed npub.
|
|
- `auto_connect: true` — when we see a beacon from an npub
|
|
we have not yet peered with, initiate the outbound Noise
|
|
handshake.
|
|
- `accept_connections: true` — when a remote npub initiates
|
|
the handshake on this transport, complete it.
|
|
|
|
If only one node sets `announce`, the other won't see it; if
|
|
only one side sets `auto_connect` or `accept_connections`, the
|
|
roles are asymmetric and the link won't establish unless both
|
|
are configured. The cleanest pattern for a ground-up tutorial
|
|
is "all four flags on both ends."
|
|
|
|
> **Multiple Ethernet links.** If a node has more than one
|
|
> physical interface that participates in the mesh, configure
|
|
> each one as a *named instance* under `ethernet:`:
|
|
>
|
|
> ```yaml
|
|
> transports:
|
|
> ethernet:
|
|
> lan:
|
|
> interface: "eth0"
|
|
> announce: true
|
|
> listen: true
|
|
> auto_connect: true
|
|
> accept_connections: true
|
|
> dongle:
|
|
> interface: "enx00aabbccddee"
|
|
> optional: true
|
|
> announce: true
|
|
> # ...
|
|
> ```
|
|
>
|
|
> Each named instance runs its own socket and neighbor state.
|
|
> A single ground-up link only needs the flat form shown
|
|
> first; named instances become useful when the same node
|
|
> bridges multiple physical segments.
|
|
>
|
|
> `optional: true` on the dongle says its absence is normal —
|
|
> a USB adapter that is plugged in some days and not others.
|
|
> Without it, naming an interface is a statement that you
|
|
> expect it, and while it is missing the node reports
|
|
> `Degraded` and logs at `error`. Either way the interface
|
|
> does not have to exist when the daemon starts: a transport
|
|
> whose interface is missing waits and binds when it appears,
|
|
> and rebinds if it later goes away. Watch that with
|
|
> `fipsctl show transports`.
|
|
|
|
## Step 3: Grant the daemon permission to open raw sockets
|
|
|
|
The Ethernet transport opens an `AF_PACKET` `SOCK_DGRAM` socket
|
|
bound to the chosen interface. That requires `CAP_NET_RAW`.
|
|
|
|
If you installed FIPS via the Debian package and run via the
|
|
shipped systemd unit, the daemon runs as root and has
|
|
`CAP_NET_RAW` already — there is nothing to do here. Skip to
|
|
Step 4.
|
|
|
|
If you are running the daemon interactively as your user (a
|
|
from-source / development setup), grant the capability once on
|
|
the binary:
|
|
|
|
```sh
|
|
sudo setcap CAP_NET_RAW,CAP_NET_ADMIN+ep "$(which fips)"
|
|
```
|
|
|
|
`CAP_NET_ADMIN` is what the daemon needs for the `fips0` TUN
|
|
adapter regardless; `CAP_NET_RAW` is the ground-up addition.
|
|
The `setcap` invocation only needs to be repeated when the
|
|
binary is replaced.
|
|
|
|
## Step 4: Restart the daemon on each node
|
|
|
|
```sh
|
|
sudo systemctl restart fips
|
|
```
|
|
|
|
Or, if running interactively, restart your `fips` invocation
|
|
in whichever way you started it.
|
|
|
|
Watch the startup logs for the Ethernet transport coming up:
|
|
|
|
```sh
|
|
sudo journalctl -u fips -f --since="1 minute ago"
|
|
```
|
|
|
|
Look for landmarks like:
|
|
|
|
- A line indicating the Ethernet transport opened the chosen
|
|
interface and started its receive loop.
|
|
- Periodic outbound beacon messages (one per
|
|
`beacon_interval_secs` window).
|
|
- After the second beacon round on the *other* node, an
|
|
inbound beacon parsed and a candidate-peer entry created.
|
|
- Once each side dials, a Noise handshake completion log
|
|
message naming the remote npub.
|
|
|
|
Beacon interval defaults to 30s, so the first peering can take
|
|
up to a minute (one beacon window per side, plus handshake).
|
|
Lower the interval for the tutorial if you want faster
|
|
feedback:
|
|
|
|
```yaml
|
|
transports:
|
|
ethernet:
|
|
# ...
|
|
beacon_interval_secs: 10 # minimum allowed
|
|
```
|
|
|
|
## Step 5: Verify the link
|
|
|
|
On either node:
|
|
|
|
```sh
|
|
sudo fipsctl show peers
|
|
```
|
|
|
|
Expect one entry whose `npub` matches the **other** node and
|
|
whose `transport_type` reads `ethernet`. Your
|
|
existing overlay peers (if any from earlier tutorials) appear
|
|
alongside it. Each peer has its own row, and the link status
|
|
columns show whether the Noise session is up.
|
|
|
|
```sh
|
|
sudo fipsctl show transports
|
|
```
|
|
|
|
Confirms that the Ethernet transport is running and shows the
|
|
beacon counters incrementing. Both `beacons_sent` and
|
|
`beacons_recv` should be non-zero if the link is healthy.
|
|
|
|
## Step 6: Reach the other node by name
|
|
|
|
On node A, ping node B by `.fips` name. Get node B's npub
|
|
from its `fipsctl show status` output (it's the persistent
|
|
identity you established earlier), then:
|
|
|
|
```sh
|
|
ping6 npub1bbb…long-string….fips
|
|
```
|
|
|
|
Expect ICMPv6 echo replies. The path is:
|
|
|
|
1. The local `.fips` resolver translates the npub-form name
|
|
into an `fd97:...` mesh address (cryptographically derived
|
|
from the npub on both ends — the resolver does the
|
|
computation locally, with no network round trip).
|
|
2. The kernel routes the packet via `fips0`.
|
|
3. The FIPS daemon accepts it from the TUN, looks up the
|
|
mesh route, and hands it to the FMP link to node B.
|
|
4. The Ethernet transport on node A frames the FMP packet as
|
|
a raw EtherType `0x2121` Ethernet frame addressed to node
|
|
B's MAC, learned from B's beacons.
|
|
5. Node B's daemon receives the frame, peels off the
|
|
Ethernet/FIPS framing, and the packet emerges on node B's
|
|
`fips0`.
|
|
6. The kernel on node B sees an inbound ICMPv6 echo and
|
|
replies, and the same path runs in reverse.
|
|
|
|
If you have a hosts file with shortnames configured (see
|
|
[host-aliases](../how-to/host-aliases.md)), substitute the
|
|
shortname for the full npub form.
|
|
|
|
## Step 7: Try a forward composition
|
|
|
|
If node A also has the `test-us01` overlay peer from
|
|
[join-the-test-mesh](join-the-test-mesh.md), node B can
|
|
reach `test-us01` *through* node A — even though node B has
|
|
no direct internet path of its own:
|
|
|
|
On node B:
|
|
|
|
```sh
|
|
ping6 npub1qmc3cvfz0yu2hx96nq3gp55zdan2qclealn7xshgr448d3nh6lks7zel98.fips
|
|
```
|
|
|
|
The packet leaves B's `fips0`, traverses the Ethernet link to
|
|
A, gets forwarded by A across the overlay UDP transport to
|
|
`test-us01`, and the reply comes back the same way.
|
|
|
|
This is the composition the chapter intro flagged: the two
|
|
deployment modes coexist on a single daemon. Node A is
|
|
participating in the test mesh via the internet *and* in your
|
|
local Ethernet mesh. From node B's perspective, the test mesh
|
|
is reachable. From `test-us01`'s perspective, B is reachable.
|
|
The mesh handles the rest.
|
|
|
|
## Variations
|
|
|
|
### WiFi (AP mode), same shape as Ethernet
|
|
|
|
Replace `<eth>` with the WiFi interface name (typically
|
|
`wlan0` or `wlp3s0`) on each node. The WiFi NIC is presented
|
|
as an Ethernet-class interface to the kernel by the
|
|
`mac80211` abstraction; the FIPS Ethernet transport opens
|
|
the same `AF_PACKET` socket on it. No FIPS-side configuration
|
|
change beyond the interface name.
|
|
|
|
What you do need on the AP side:
|
|
|
|
- Both nodes associated to the same SSID.
|
|
- **Client (station) isolation must be OFF** on the AP.
|
|
Most consumer routers ship with it off; many guest
|
|
networks and "secure" enterprise APs ship with it on.
|
|
When client isolation is on, the AP refuses to forward
|
|
station-to-station frames — the broadcast beacons never
|
|
arrive at the other node, and neighbor detection fails silently.
|
|
If beacons aren't crossing, this is the first thing to
|
|
check.
|
|
|
|
There is no FIPS-specific configuration for WiFi versus
|
|
Ethernet on the daemon side; the choice is purely the
|
|
adapter name.
|
|
|
|
### Bluetooth LE (experimental but works)
|
|
|
|
BLE is a separate transport (`transports.ble.*`) with its own
|
|
neighbor-detection model — L2CAP advertisements rather than raw L2
|
|
broadcasts. The shape of the tutorial is the same (advertise +
|
|
scan + auto-connect + accept), but the prerequisites are
|
|
different: BlueZ, `bluetoothd`, an HCI adapter, and the
|
|
`bluetooth` group or capability set.
|
|
|
|
The full operator recipe is in
|
|
[../how-to/set-up-bluetooth-peer.md](../how-to/set-up-bluetooth-peer.md).
|
|
Mark this transport as experimental: it works in most
|
|
configurations but the BLE stack has more variability than
|
|
Ethernet — adapter quirks, BlueZ version differences, and the
|
|
shorter range all matter.
|
|
|
|
The BLE transport is **Linux-only** at present; macOS and
|
|
Windows builds skip it.
|
|
|
|
## What you've learned
|
|
|
|
- **Ground-up is the new ground.** FIPS does not need any IP
|
|
infrastructure between two devices to mesh them. A wire (or
|
|
a radio link), `CAP_NET_RAW`, and a few config flags on each
|
|
end are sufficient. The mesh supplies its own identity,
|
|
addressing, discovery, and routing.
|
|
- **Neighbor detection is a four-flag opt-in.** `announce`, `listen`,
|
|
`auto_connect`, and `accept_connections` each control one
|
|
thing; both ends must agree before a link will form.
|
|
- **The two modes coexist.** Overlay peers and ground-up peers
|
|
ride the same daemon — same FMP link layer, same FSP session
|
|
layer, same `fips0` adapter. A node can be a bridge between
|
|
the two without any extra plumbing.
|
|
- **No IP on the link.** The Ethernet transport bypasses the
|
|
kernel IP stack via `AF_PACKET`. Whether the interface has
|
|
an IP address is irrelevant; whether it has carrier is what
|
|
matters.
|
|
- **Names work the same way.** `<npub>.fips` resolves locally
|
|
via the cryptographically-derived ULA. The resolver does
|
|
not care whether the destination is reached over Ethernet,
|
|
UDP overlay, or some hop chain combining both.
|
|
|
|
## Troubleshooting
|
|
|
|
- **No beacons received.** On either node, `sudo fipsctl show
|
|
transports` should show `beacons_recv` incrementing
|
|
every `beacon_interval_secs` once the other node is also
|
|
running. If it stays at zero:
|
|
- Confirm the chosen interface is `LOWER_UP` (carrier
|
|
present).
|
|
- Confirm the other node is announcing (its `beacons_sent`
|
|
should be non-zero).
|
|
- On WiFi: confirm AP client isolation is off.
|
|
- On a switch: confirm the switch is unmanaged or that
|
|
EtherType `0x2121` is not being filtered. Most consumer
|
|
switches forward all EtherTypes; managed switches
|
|
sometimes don't.
|
|
- **Beacons received but no peer entry.** The handshake is
|
|
failing. Tail logs (`journalctl -u fips`) for Noise
|
|
handshake errors. Common causes: peer ACL active and not
|
|
including the remote npub (out of scope for this tutorial,
|
|
but check `/etc/fips/peers.allow` if you have set one);
|
|
daemon's clock drift large enough to fail freshness
|
|
checks (rare).
|
|
- **Daemon won't start with the Ethernet transport.** Likely
|
|
a permissions error. Check `journalctl -u fips` for an
|
|
`EPERM` or "operation not permitted" message; if running
|
|
interactively, confirm the binary has `CAP_NET_RAW`
|
|
(`getcap "$(which fips)"`).
|
|
- **Beacons in both directions, peers entries on both sides,
|
|
but ping6 times out.** The handshake completed but the FSP
|
|
session is not flowing data. Check `fipsctl show peers`'s
|
|
link status columns — if the FMP link is healthy but FSP
|
|
is not, the mesh-layer side is fine and the issue is one
|
|
layer up. The
|
|
[reach-mesh-services § Troubleshooting](reach-mesh-services.md#troubleshooting)
|
|
section covers symptoms at this level.
|
|
- **`AF_PACKET` socket bind fails on a kernel-protected
|
|
interface.** Some hardened kernels (`grsec`, certain
|
|
containers, certain VMs) restrict raw-socket access even
|
|
with `CAP_NET_RAW`. The daemon log will name the failing
|
|
syscall. The fix is host-side: relax the restriction or
|
|
pick a different interface.
|
|
|
|
## What's next
|
|
|
|
You now have the second deployment mode of FIPS in your
|
|
hands. From here:
|
|
|
|
- **Add a third node.** Bring up a third machine on the same
|
|
Ethernet segment, configure it identically, and watch all
|
|
three nodes form a mesh. The FIPS spanning tree picks a
|
|
root and routing converges within a few beacon intervals.
|
|
- **Mix transports.** Add an overlay peer (per
|
|
[join-the-test-mesh](join-the-test-mesh.md)) to one of
|
|
your ground-up nodes; the local mesh now reaches the test
|
|
mesh through that node, and vice versa.
|
|
- **Host services.** Anything you do on `fips0` with overlay
|
|
peers — bind an HTTP server (per
|
|
[host-a-service](host-a-service.md)), reach a service via
|
|
the daemon's IPv6 adapter (per
|
|
[reach-mesh-services](reach-mesh-services.md)) — works
|
|
identically on a ground-up mesh. The data plane is the
|
|
same.
|
|
|
|
For more depth on the link-layer machinery:
|
|
|
|
- [../reference/transports.md § Ethernet](../reference/transports.md)
|
|
— full Ethernet transport reference (counter inventory,
|
|
per-instance configuration, MTU model).
|
|
- [../reference/configuration.md § Ethernet](../reference/configuration.md#ethernet-transportsethernet)
|
|
— every configuration key and its default.
|
|
- [../how-to/set-up-bluetooth-peer.md](../how-to/set-up-bluetooth-peer.md)
|
|
— operator recipe for the BLE variant.
|
|
- [../design/fips-transport-layer.md](../design/fips-transport-layer.md)
|
|
— the design doc that describes the per-link MTU model and
|
|
why each transport is treated as link-layer rather than
|
|
network-layer.
|