Merge branch 'master' into next

Carries the experimental native datagram API up as its two commits: the API
itself, where a client opens a flow to a peer's public key on a chosen port
and reads datagrams off a descriptor the daemon hands it, and the macOS port
that builds the listener over SOCK_DGRAM because Darwin has no
SOCK_SEQPACKET for AF_UNIX. It is off by default and is not a stable
interface. It changes no wire format on either line: every FSP data packet
has carried a port pair inside its AEAD envelope since v0.2.0.

No conflicts, and no adaptation. `git diff --numstat` for what this merge
changes on next matches the master side file for file and count for count,
so the merge took master's diff verbatim rather than resolving anything.

The instrumentation step count needed no change this time: neither commit
touches `src/instr/recorder.rs`, so the pinned figure the previous merge
corrected to 28 still describes the tick.
This commit is contained in:
Johnathan Corgan
2026-08-21 05:55:35 +00:00
63 changed files with 14494 additions and 359 deletions
+57 -1
View File
@@ -221,7 +221,11 @@ jobs:
${{ runner.os }}-cargo-
- name: Build
run: cargo build --release
# --bins --examples rather than the bare default: the native datagram
# API's echo server is a cargo example, and the integration image needs
# it. Naming --bins keeps the daemon and its tools in the build, which
# --examples alone would drop. Mirrors testing/ci-local.sh.
run: cargo build --release --bins --examples
- name: SHA-256 hashes (Linux)
if: runner.os == 'Linux'
@@ -236,6 +240,15 @@ jobs:
shell: pwsh
run: Get-FileHash target\release\fips.exe, target\release\fipsctl.exe, target\release\fipstop.exe -Algorithm SHA256
# Cargo puts an example under target/release/examples. Staging them
# beside the bins keeps the artifact's common root at target/release, so
# every existing consumer still finds its file at _bin/<name>.
- name: Stage the native API examples beside the release binaries
if: matrix.os == 'ubuntu-latest'
run: |
cp target/release/examples/native-echo target/release/native-echo
cp target/release/examples/native-surface target/release/native-surface
# Upload the Linux binary so integration jobs can use it without rebuilding
- name: Upload Linux binary
if: matrix.os == 'ubuntu-latest'
@@ -247,6 +260,8 @@ jobs:
target/release/fipsctl
target/release/fipstop
target/release/fips-gateway
target/release/native-echo
target/release/native-surface
retention-days: 1
# ─────────────────────────────────────────────────────────────────────────────
@@ -521,6 +536,14 @@ jobs:
# loopback-delivered query was once misattributed to the mesh
# interface and dropped. Single matrix entry runs all 13
# scenarios sequentially; ~7-12 min warm, ~12-15 min cold.
# Native datagram API: a client process opening a pubkey-to-pubkey
# flow over the daemon's Unix socket. One single-node leg covering
# the socket, its access mode and the command surface, plus a
# two-node pair that sends a real datagram end to end. Fast: no
# per-distro images and no TUN. ~2-3 min.
- suite: native-api
type: native-api
- suite: dns-resolver
type: dns-resolver
@@ -544,6 +567,13 @@ jobs:
cp _bin/fipsctl testing/docker/fipsctl
[ -f _bin/fipstop ] && cp _bin/fipstop testing/docker/fipstop || true
[ -f _bin/fips-gateway ] && cp _bin/fips-gateway testing/docker/fips-gateway || true
# Not optional: the Dockerfile COPYs both native API examples
# unconditionally, and a missing source there fails the shared image
# build for every leg, not just native-api. Fail here instead, where
# the cause is legible.
chmod +x _bin/native-echo _bin/native-surface
cp _bin/native-echo testing/docker/native-echo
cp _bin/native-surface testing/docker/native-surface
docker build -t fips-test:latest testing/docker
docker build -t fips-test-app:latest -f testing/docker/Dockerfile.app testing/docker
@@ -734,6 +764,32 @@ jobs:
docker rm -f "$c" >/dev/null 2>&1 || true
done
# ── Native datagram API ─────────────────────────────────────────────
# Reads FIPS_TEST_IMAGE rather than defaulting to a name, so it runs
# against the image this workflow built. The two-node check creates and
# removes its own docker network.
- name: Run native-api test
if: matrix.type == 'native-api'
timeout-minutes: 15
env:
FIPS_TEST_IMAGE: fips-test:latest
run: bash testing/native-api/test.sh
- name: Collect logs on failure (native-api)
if: matrix.type == 'native-api' && failure()
run: |
docker ps -a --filter "name=fips-native" --format '{{.Names}}' | while read -r c; do
echo "--- ${c} ---"
docker logs "$c" 2>&1 | tail -100 || true
done
- name: Stop containers (native-api)
if: matrix.type == 'native-api' && always()
run: |
docker ps -a --filter "name=fips-native" --format '{{.Names}}' | while read -r c; do
docker rm -f "$c" >/dev/null 2>&1 || true
done
# ── DNS resolver multi-backend integration ──────────────────────────
# The dns-resolver harness builds its own fips binary from source in a
# Debian 12 builder image (shared cache layout with deb-install). Runs
+34
View File
@@ -339,6 +339,40 @@ with v0.4.x or earlier peers.
terminal leaves behind. `--json` emits exactly one document at the end, so a
script parsing the report does not have to skip past progress output.
- An **experimental** native datagram API addressed by public key, off by
default and not a stable interface. A client process opens a flow to a
peer's public key on a chosen port and sends and receives datagrams on a
file descriptor the daemon hands it: no IPv6 emulation, no TUN device and no
DNS, a datagram travelling from key to key. **The wire needs no change and
gets none.** Every FSP data packet has carried a port pair inside its AEAD
envelope since v0.2.0 and port 256 is simply the IPv6 shim, so what was
missing was a way for a program to ask for a port of its own and be handed
the traffic. The x-only public key is the address and an npub is that key
written in bech32, so converting between them is a local encoding rather
than a lookup or a name service; the 16-byte node address that travels on
the wire is a truncated hash of the key, does not invert, and appears
nowhere a client can see. A listener is a descriptor: the daemon writes one
message per arrival to it, carrying the new flow's descriptor and the peer's
address, so poll, select and epoll work on a listener and accepting is a
`recvmsg`. There is no accept command and no reject command, and refusing a
flow is closing the descriptor you were handed. The Rust surface mirrors
`std::net`, with `FipsStream::connect`, `FipsListener::bind`, `incoming`,
`accept`, `io::Result` and an errno mapping rather than a bespoke error
type, plus `set_nonblocking`, `AsFd` and four deadline methods under the
names and signatures `std::net` uses for the same jobs. One rule has no
counterpart in Berkeley sockets and a client author must know it: the v1
wire carries no half-close, so nothing peer-driven ever closes a flow, and a
server written to read until the flow ends waits for a signal that cannot
arrive. The listener uses `SOCK_SEQPACKET` on Linux and FreeBSD and
`SOCK_DGRAM` on macOS, which does not implement `SOCK_SEQPACKET` for
`AF_UNIX`; both keep the message boundaries the API's contract with its
clients rests on. The two kernels signal a closed peer differently and were
measured rather than reasoned about, so the receive path treats a Darwin
`ECONNRESET` as end of file alongside the `POLLHUP` and zero-byte read that
Linux gives. `EAGAIN` is deliberately not in that company: it means the
socket is empty and the peer alive, so it stays an error and the caller
waits again.
### Changed
- `node.rekey.enabled` now means "initiate rekeys" and nothing else. The
+1
View File
@@ -33,6 +33,7 @@ documents cover specific subsystems in detail.
| [fips-mesh-layer.md](fips-mesh-layer.md) | FIPS Mesh Protocol (FMP): peer authentication, link encryption, forwarding |
| [fips-session-layer.md](fips-session-layer.md) | FIPS Session Protocol (FSP): end-to-end encryption, sessions |
| [fips-ipv6-adapter.md](fips-ipv6-adapter.md) | IPv6 adaptation: TUN interface, DNS, MTU enforcement |
| [fips-native-api.md](fips-native-api.md) | Native datagram API: pubkey-addressed flows over FSP, and what it is instead of the TUN path |
### Cross-Cutting
@@ -0,0 +1,119 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 950 722" font-family="monospace" font-size="15">
<style>
rect.layer { rx: 5; }
rect.app { fill: #1a3a2a; stroke: #40a060; stroke-width: 2; }
rect.fmt { fill: #1a3a2a; stroke: #40a060; stroke-width: 2; }
rect.iface { fill: #16323a; stroke: #40a0b0; stroke-width: 2; }
rect.fsp { fill: #2a2040; stroke: #8060c0; stroke-width: 2; }
rect.fmp { fill: #2e3a5e; stroke: #5080c0; stroke-width: 2; }
rect.xport { fill: #2a1a1a; stroke: #c06040; stroke-width: 2; }
rect.trad { fill: #1a1d24; stroke: #565f6e; stroke-width: 2; }
rect.none { fill: #16181e; stroke: #3d434f; stroke-width: 2; stroke-dasharray: 7 5; }
text { fill: #e0e0e0; }
text.name { font-size: 18px; font-weight: bold; }
text.sub { font-size: 15px; font-weight: bold; }
text.desc { font-size: 13px; fill: #a0a0b0; }
text.dim { font-size: 13px; fill: #6a7180; }
text.tiny { font-size: 11.5px; fill: #6a7180; }
text.api { font-size: 14px; fill: #c8c8d4; }
text.endpt { font-size: 16px; font-weight: bold; }
text.tunnel { font-size: 12.5px; fill: #40a0b0; font-style: italic; }
text.head { font-size: 13px; fill: #808895; font-style: italic; }
text.caption { font-size: 15px; fill: #808090; font-style: italic; }
line.drop { stroke: #565f6e; stroke-width: 2; }
path.fork { stroke: #565f6e; stroke-width: 2; fill: none; }
path.tunnel { stroke: #40a0b0; stroke-width: 2.5; fill: none; }
</style>
<defs>
<marker id="a" markerWidth="9" markerHeight="9" refX="7" refY="4.5" orient="auto">
<path d="M0,1 L7,4.5 L0,8 Z" fill="#565f6e"/>
</marker>
<marker id="t" markerWidth="10" markerHeight="10" refX="8" refY="5" orient="auto">
<path d="M0,1 L8,5 L0,9 Z" fill="#40a0b0"/>
</marker>
</defs>
<rect x="10" y="10" width="930" height="702" fill="#0d1117" rx="10"/>
<!-- One application, two ways down -->
<rect x="50" y="36" width="850" height="58" class="layer app"/>
<text x="475" y="62" text-anchor="middle" class="name">Application</text>
<text x="475" y="82" text-anchor="middle" class="desc">unmodified on the left, written against the key on the right</text>
<!-- What an endpoint looks like on each side -->
<text x="230" y="118" text-anchor="middle" class="api">socket API</text>
<text x="230" y="140" text-anchor="middle" class="endpt" fill="#40a0b0">https://&lt;npub&gt;.fips:443</text>
<text x="230" y="158" text-anchor="middle" class="tiny">a name, resolved to an address</text>
<text x="720" y="118" text-anchor="middle" class="api">native datagram API</text>
<text x="720" y="140" text-anchor="middle" class="endpt" fill="#c8c8d4">fips://&lt;npub&gt;:443</text>
<text x="720" y="158" text-anchor="middle" class="tiny">the key itself, no lookup</text>
<line x1="230" y1="94" x2="230" y2="104" class="drop"/>
<line x1="720" y1="94" x2="720" y2="104" class="drop"/>
<line x1="230" y1="164" x2="230" y2="180" class="drop" marker-end="url(#a)"/>
<line x1="720" y1="164" x2="720" y2="180" class="drop" marker-end="url(#a)"/>
<!-- Row 1 — format -->
<rect x="50" y="180" width="360" height="72" class="layer trad"/>
<text x="68" y="208" class="name">HTTP</text>
<text x="68" y="230" class="desc">a format every client already speaks</text>
<rect x="540" y="180" width="360" height="72" class="layer fmt"/>
<text x="558" y="208" class="name">your own format</text>
<text x="558" y="230" class="desc">you design it; 1 send = 1 datagram</text>
<!-- Row 2 — secrecy. FSP is where the tunnel re-enters. -->
<rect x="50" y="258" width="360" height="72" class="layer trad"/>
<text x="68" y="286" class="name">TLS</text>
<text x="68" y="308" class="desc">trust anchored in a CA certificate</text>
<rect x="540" y="258" width="360" height="72" class="layer fsp"/>
<text x="558" y="286" class="name">FSP</text>
<text x="558" y="308" class="desc">end-to-end AEAD; trust in the key itself</text>
<!-- Row 3 — reliability, the row with no counterpart -->
<rect x="50" y="336" width="360" height="72" class="layer trad"/>
<text x="68" y="364" class="name">TCP</text>
<text x="68" y="386" class="desc">reliability, ordering, flow control</text>
<rect x="540" y="336" width="360" height="72" class="layer none"/>
<text x="558" y="362" class="name" fill="#8a919e">ROD</text>
<text x="558" y="380" class="dim">Reliable Object Delivery — not in v1</text>
<text x="558" y="397" class="dim">a v2 capability, may come forward</text>
<!-- Row 4 — routing -->
<rect x="50" y="414" width="360" height="72" class="layer trad"/>
<text x="68" y="442" class="name">IPv6</text>
<text x="68" y="464" class="desc">routing by address prefix</text>
<rect x="540" y="414" width="360" height="72" class="layer fmp"/>
<text x="558" y="442" class="name">FMP</text>
<text x="558" y="464" class="desc">spanning tree, bloom, per-hop auth</text>
<!-- Row 5 — where the IP stack lands: a wire, or the adapter -->
<path d="M230,486 L230,498 L130,498 L130,508" class="fork" marker-end="url(#a)"/>
<path d="M230,486 L230,498 L315,498 L315,508" class="fork" marker-end="url(#a)"/>
<rect x="50" y="508" width="160" height="66" class="layer none"/>
<text x="130" y="534" text-anchor="middle" class="sub" fill="#8a919e">eth0</text>
<text x="130" y="556" text-anchor="middle" class="tiny">the ordinary internet</text>
<rect x="220" y="508" width="190" height="66" class="layer iface"/>
<text x="315" y="532" text-anchor="middle" class="sub">fips0</text>
<text x="315" y="554" text-anchor="middle" class="tiny" fill="#8fb8c2">IPv6 adapter: TUN + DNS</text>
<rect x="540" y="508" width="360" height="66" class="layer xport"/>
<text x="558" y="534" class="name">Transport</text>
<text x="558" y="556" class="desc">UDP, Ethernet, WiFi, BLE, Tor, serial</text>
<!-- The tunnel: the adapter's output is FSP's input, back up the stack -->
<path d="M410,541 L470,541 L470,294 L540,294" class="tunnel" marker-end="url(#t)"/>
<text x="462" y="420" text-anchor="middle" class="tunnel" transform="rotate(-90 462 420)">tunnelled into FSP</text>
<text x="475" y="606" text-anchor="middle" class="head">Both endpoints name the same node and port, but not the same kind of port: on the left it is a TCP port inside the</text>
<text x="475" y="624" text-anchor="middle" class="head">tunnel, on the right an FSP port, whose tier rules are expected to change. The adapter is not a bottom layer — an</text>
<text x="475" y="642" text-anchor="middle" class="head">IPv6 packet reaching fips0 becomes an FSP payload, so the whole left stack runs inside the right one, TCP included.</text>
<text x="475" y="678" text-anchor="middle" class="caption">Two ways into one mesh: tunnel the IP stack, or address the key directly</text>
</svg>

After

Width:  |  Height:  |  Size: 6.8 KiB

+7
View File
@@ -83,6 +83,13 @@ addresses with header compression so unmodified IP applications can
use the network transparently, while the native datagram API
addresses destinations directly by npub.
The native datagram API is **experimental** and off by default. It is
not a stable API surface and carries no compatibility promise. See
[../how-to/use-the-native-datagram-api.md](../how-to/use-the-native-datagram-api.md)
for enabling it and writing against it, and
[../reference/security.md](../reference/security.md#native-datagram-api)
for what enabling it grants to the `fips` group.
![Node Architecture](diagrams/fips-node-architecture.svg)
The mesh routes application traffic across heterogeneous transports
+11
View File
@@ -24,6 +24,17 @@ a native FIPS datagram service, or through an IPv6 adaptation layer
that presents each node as an IPv6 endpoint for compatibility with
existing IP-based applications.
![Two ways into the mesh](diagrams/fips-native-api-stack-comparison.svg)
Both columns are the same mesh. On the left an unmodified program keeps
the stack it already has, and reaches the mesh through `fips0`, a
virtual network interface that carries its IPv6 packets. On the right a
FIPS-aware program names the far node by its key and skips the IP layers
altogether. The two protocols in the middle are the mesh's own: **FSP**
encrypts end to end between the two nodes, and **FMP** authenticates
each hop and decides where a packet goes next.
[fips-architecture.md](fips-architecture.md) takes them in order.
## Why FIPS?
**Self-sovereign identity**: FIPS nodes generate their own addresses,
+9
View File
@@ -18,6 +18,15 @@ connects to the kernel's IPv6 stack.
Applications that are FIPS-aware can bypass the adapter entirely and use the
native FIPS datagram API, addressing destinations directly by npub.
![Where the adapter sits](diagrams/fips-native-api-stack-comparison.svg)
The adapter is the `fips0` box, and the arrow leaving it is the point: an IPv6
packet arriving there does not go out to a wire. It becomes an FSP payload, so
the whole IP stack above runs inside the mesh's own stack — TCP included, which
is what lets an unmodified `ssh` or `curl` work across the mesh. The right-hand
column is the same mesh reached without the adapter, and
[fips-native-api.md](fips-native-api.md) covers that path.
## DNS Integration
### The Problem
+147
View File
@@ -0,0 +1,147 @@
# Native Datagram API
The native datagram API lets a local program move bytes between two public keys
over FSP, with no IPv6 emulation and no TUN device in the path. A program calls
`connect` for a flow to a public key and a port, or `bind` for a port to receive
flows on, and from then on uses ordinary socket calls.
This document explains what the interface is for and where its edges are. For
the surface itself — every type, method, errno and command — see
[../reference/native-api.md](../reference/native-api.md). For the steps to enable
it and write a program, see
[../how-to/use-the-native-datagram-api.md](../how-to/use-the-native-datagram-api.md).
## Where it sits
![Stack comparison](diagrams/fips-native-api-stack-comparison.svg)
The two endpoints at the top are the same node reached two ways. **The
`fips://` form is illustrative**: no code in this repository parses it, nothing
registers the scheme, and the API takes a key and a port as separate arguments
rather than a URL. It is drawn because it is the shape an address takes on that
side, against a `.fips` name the adapter's DNS really does resolve.
Read row by row, the native path replaces three layers and declines to replace a
fourth. FSP takes TLS's place and anchors trust in the key rather than in a
certificate authority. FMP takes IPv6's place and routes by spanning tree and
bloom filter rather than by address prefix, with the address derived from the
key. The transport layer takes the medium's place and can be several media at
once.
**There is nothing where TCP was**, and on the native path that is the single
most consequential row today. No acknowledgement, no retransmission, no ordering
and no flow control: a program that needs any of them builds it into its own
payload.
That row is marked **ROD — Reliable Object Delivery**, which is where the
capability is expected to land. ROD is a v2 capability and is not in v1; it may
be pulled forward. Until it is, treat the row as empty and design around it,
because a program written against a reliability layer that is not there yet
fails in the ways this document's "not a reliability layer" section describes.
**The two paths are not alternatives at the bottom.** They converge. An
unmodified IPv6 program does not stop at a wire: its packets reach `fips0`, and
the adapter hands each one to FSP as a payload. That is the arrow running up the
middle of the diagram, and it is why the left stack is drawn ending at an
interface rather than at Ethernet.
So the whole left column runs *inside* the right one. **TCP included** — which
is the practical answer to the empty row above it. A program that needs a
reliable ordered stream over the mesh already has one: run it over `fips0` and
let TCP do what TCP does, inside FSP's encryption. What the native API offers
instead is the same mesh with four layers of machinery removed, for a program
willing to do without them.
The bottom of the diagram is not always the bottom of the stack either. When
FIPS overlays an existing network its transport is UDP, which still rides IP and
Ethernet beneath; when the mesh *is* the network, a transport sits on a link
directly.
## What it is instead of
The fastest way to place the interface is by contrast with the TUN device, which
is the other way a program gets FIPS traffic.
| | TUN interface | Native datagram API |
| --- | ------------- | ------------------- |
| Addressing | IPv6 address | public key, written as an npub |
| Name resolution | DNS over the mesh | none: the program supplies the key |
| Kernel object | TUN device, routes | a `FipsStream` per peer |
| Encapsulation | IPv6 emulated over FSP | FSP port pair, no IP layer |
| Program sees | an IP network | a `FipsStream` |
| Privilege | `CAP_NET_ADMIN` to create the device | membership of group `fips` |
| Demultiplexing | by address and port | by flow, one stream each |
**The IPv6 emulation is not removed by this interface.** It continues to run
beside it on FSP port 256, which is why that port and the tier around it are
refused to a program. What the native API removes is a program's *dependence* on
it: a program that wants to move bytes between two known public keys no longer
has to acquire an IPv6 address, resolve a name, and hand its payload to a
protocol stack that will encapsulate it again.
Both paths reach the same place. A native datagram and an emulated IPv6 packet
are both FSP payloads with a port pair, carried in the same encrypted session to
the same peer. The difference is entirely on the local side of the daemon.
## Status
**The wire is connected**: a datagram sent on a flow leaves the node over FSP,
and one arriving on a held port reaches its flow.
**The interface around it is experimental.** It is not versioned, it has no
compatibility promise, and three of its five commands exist only to let the
daemon's own checks drive the receive path without a peer. It is Linux and
FreeBSD only, and it is off by default.
## What this is not
**Not a stable interface.** It is an experiment on the v1 wire. Names, fields,
reply shapes and the command set may change without a deprecation cycle.
**Not the v2 process API.** The v2 external process API is a separate and later
design, which retires ports entirely in favour of a listener, connection and
stream model. Nothing here governs it and nothing there governs this. The one
thing this interface takes from that work is the FSP port tiers, because port
256 already carries the IPv6 shim on the deployed wire and a new service must not
collide with it.
**Not a reliability layer.** There is no acknowledgement, no retransmission, no
ordering guarantee and no flow control between the two ends. A datagram is
carried or it is dropped. Some drops are counted inside the daemon and none are
reported to a program for real traffic. A program that needs delivery guarantees
builds them itself, on top, in the payload — or runs over `fips0` and lets TCP
provide them.
Reliable Object Delivery (ROD) is the v2 capability intended to fill this gap,
and it may be pulled forward into v1. **Nothing here anticipates it**: no field,
reply shape or command on this surface is reserved for it, and a program written
today should assume it does not exist.
**Not an authorization boundary.** The socket's group ownership is the whole of
the access control. Any process that can open it can send as this node's identity
and can receive mesh traffic on any port it can claim, and there is no per-program
separation beyond the port registry. Because the descriptor carries the flow, a
process handed one over `SCM_RIGHTS` can send as this node on that flow without
ever opening the socket. See
[../reference/security.md](../reference/security.md#native-datagram-api).
**Not multi-tenant.** `max_flows` is node-wide with no per-program share, so one
program can exhaust it, and every other program then sees `EMFILE` on `connect`
and silent drops on its listeners.
**Not a connection in the TCP sense.** A successful `connect` is a local
registration and contacts no peer. There is no handshake, no keepalive and no
notification that a peer went away. A flow ends when its descriptor closes, and
in no other way. In particular **a peer cannot end your flow: it has no close to
send.** That single fact shapes every program written against this interface,
and the consequences are drawn out in
[../how-to/use-the-native-datagram-api.md](../how-to/use-the-native-datagram-api.md#four-things-that-will-bite-you).
## See also
- [fips-session-layer.md](fips-session-layer.md) — FSP, which carries the
datagrams and owns the port pair
- [fips-ipv6-adapter.md](fips-ipv6-adapter.md) — the other consumer of FSP, and
what this interface is an alternative to
- [../reference/native-api.md](../reference/native-api.md) — the surface, the
line protocol and the command reference
+3
View File
@@ -28,3 +28,6 @@ X" to "X is done".
| [set-up-80211s-mesh-backhaul.md](set-up-80211s-mesh-backhaul.md) | Link OpenWrt FIPS routers over an open 802.11s radio backhaul (FIPS provides encryption, authentication, and routing) |
| [set-up-open-access-ssid.md](set-up-open-access-ssid.md) | Broadcast the open `!FIPS` access SSID so phones and laptops roam onto the mesh (one ESS: save once, roam every FIPS router) |
| [diagnose-mtu-issues.md](diagnose-mtu-issues.md) | Triage MTU-shaped failures and rule out their imposters (bufferbloat, transport saturation) |
| [use-the-native-datagram-api.md](use-the-native-datagram-api.md) | Enable the experimental native datagram API and write a program that sends and receives datagrams by pubkey and port (no IPv6 emulation, no TUN). Read the `fips` group warning first |
| [write-a-native-api-client.md](write-a-native-api-client.md) | Speak the native datagram API's line protocol directly from C, Python or Go, where there is no client library |
| [serve-many-peers-on-one-thread.md](serve-many-peers-on-one-thread.md) | Handle every native API flow from one `poll` loop instead of a thread per peer |
@@ -0,0 +1,196 @@
# Serve Many Peers on One Thread
**Goal:** handle every native datagram API flow from a single `poll` loop,
instead of dedicating a thread to each peer.
The straightforward listening program spawns a thread per flow. That is fine for
a handful of peers and wrong at the node's ceiling of 256, where it costs 256
threads mostly parked in `recv`.
Every object on this surface is a descriptor, so there is nothing to integrate:
both types implement `AsFd` and `AsRawFd` and go straight into a `poll`,
`select` or `epoll` set. **A listener is readable exactly when `accept` would not
block**, which is the property the whole shape rests on, and the crate asserts it
as a test rather than claiming it.
Read [use-the-native-datagram-api.md](use-the-native-datagram-api.md) first if
you have not opened a flow before.
## The whole program
One dependency beyond the crate, `libc`, for `poll` itself:
```toml
[dependencies]
fips = { git = "https://github.com/jmcorgan/fips" }
libc = "0.2"
```
```rust
//! Serve many flows from one poll loop, with no thread per peer.
//!
//! ```text
//! eventloop /run/fips/api.sock 4242
//! ```
use fips::native::client::{FipsListener, FipsStream};
use std::env;
use std::error::Error;
use std::os::fd::{AsRawFd, RawFd};
use std::path::Path;
use std::process::ExitCode;
use std::time::{Duration, Instant};
/// How long a flow may go without a datagram before this program closes it.
///
/// Nothing else will end one. The wire carries no far-end close, so a peer that
/// has stopped sending is indistinguishable from one that is thinking, and a
/// reactor with no deadline holds every flow it ever accepted until it exits.
const IDLE: Duration = Duration::from_secs(30);
/// Hold the port named on the command line and serve every flow from one loop.
fn run() -> Result<(), Box<dyn Error>> {
let mut args = env::args().skip(1);
let (Some(socket), Some(port)) = (args.next(), args.next()) else {
return Err("usage: eventloop <socket-path> <local-port>".into());
};
let port: u16 = port.parse()?;
let listener = FipsListener::bind_at(Path::new(&socket), port)?;
println!("holding {}", listener.local_addr());
// Each flow with the time its last datagram arrived, which is what the
// deadline is measured against.
let mut flows: Vec<(FipsStream, Instant)> = Vec::new();
loop {
let mut fds = vec![watch(listener.as_raw_fd())];
fds.extend(flows.iter().map(|(flow, _)| watch(flow.as_raw_fd())));
let count = fds.len() as libc::nfds_t;
// The timeout is what makes the deadline reachable: with no events at
// all the loop must still wake to notice a flow that has gone quiet.
let timeout = IDLE.as_millis() as libc::c_int;
// SAFETY: `fds` is a live slice of `pollfd` for the whole call, and
// `count` is its length.
if unsafe { libc::poll(fds.as_mut_ptr(), count, timeout) } < 0 {
return Err(std::io::Error::last_os_error().into());
}
// The established flows first, and backwards: removing a closed one
// must not renumber one not yet examined, and accepting below would
// otherwise push a flow this pass has no `revents` for. The listener
// occupies slot 0, hence the offset.
let now = Instant::now();
for index in (0..flows.len()).rev() {
if ready(&fds[index + 1]) {
if echo(&flows[index].0) {
flows[index].1 = now;
} else {
flows.remove(index); // dropping it releases the flow
}
} else if now.duration_since(flows[index].1) >= IDLE {
println!("closing an idle flow from {}", flows[index].0.peer_addr());
flows.remove(index);
}
}
// Once per readiness rather than in a loop: the descriptor is blocking,
// so a second `accept` with nothing queued would stall the whole loop.
if ready(&fds[0]) {
let (flow, peer) = listener.accept()?;
println!("flow from {peer}");
flows.push((flow, now));
}
}
}
/// A `pollfd` asking for readability on `fd`.
///
/// `POLLHUP` needs no asking for: it is reported in `revents` whether or not
/// it was requested, which is what lets one mask serve both cases.
fn watch(fd: RawFd) -> libc::pollfd {
libc::pollfd {
fd,
events: libc::POLLIN,
revents: 0,
}
}
/// Whether this descriptor has something to read or has hung up.
fn ready(poll: &libc::pollfd) -> bool {
poll.revents & (libc::POLLIN | libc::POLLHUP) != 0
}
/// Return one datagram, reporting whether the flow is still usable.
fn echo(flow: &FipsStream) -> bool {
let mut buf = vec![0u8; flow.max_payload()];
match flow.recv(&mut buf) {
// `Ok(0)` is an empty datagram and not a close, so it is echoed like
// any other. `EPIPE` is the daemon gone: no peer can close a flow.
Ok(len) => flow.send(&buf[..len]).is_ok(),
Err(_) => false,
}
}
/// Report a failure on stderr and exit non-zero.
fn main() -> ExitCode {
match run() {
Ok(()) => ExitCode::SUCCESS,
Err(error) => {
eprintln!("eventloop: {error}");
ExitCode::FAILURE
}
}
}
```
## The four rules
**Register the listener for readability only.** There is nothing else to ask it
for, and `POLLHUP` arrives in `revents` whether or not it was requested.
**Accept once per readiness, not in a loop.** The descriptor is blocking, so a
second `accept` with nothing queued stalls the whole loop. Looping until
`WouldBlock` is correct only after `listener.set_nonblocking(true)`, which is
also what an edge-triggered `epoll` requires.
**Give every flow a deadline, and the poll a timeout that makes the deadline
reachable.** This is the most important of the four. Nothing will tell a reactor
that a peer is finished, so a flow that goes quiet stays in the poll set for ever
unless the program removes it — and with no events at all the loop must still
wake in order to notice. **A reactor with a deadline but no poll timeout has a
deadline it can never reach.**
**A flow from `accept` arrives blocking, whatever the listener was set to.**
They are separate sockets and the daemon hands over a fresh one. The program
above deliberately leaves them blocking and makes exactly one `recv` per
readiness, which is safe on a blocking descriptor and is why it needs no flags at
all. If you want them otherwise, call `set_nonblocking` on the flow.
## Where this reaches past the client module
This program needs `libc` and an `unsafe` block, and it is the only one of the
API's example programs that needs anything.
That is a narrower gap than it once was. Readiness and a bounded wait were both
missing from the surface; `set_nonblocking` closed the first and
`set_read_timeout` the second, and a program wanting an option on a flow now has
a method for it. What is left is a different kind of thing: **an option on a flow
is something a surface can supply, and a reactor's own polling mechanism is
not.** No surface that stops at the descriptor can supply `poll`.
`AsRawFd` rather than `AsFd` here is deliberate: `libc::poll` takes a raw
descriptor. Prefer `AsFd` anywhere you **hold** a registration, because its
borrow cannot outlive the stream; this loop rebuilds its `pollfd` set from live
references on every pass, so it holds nothing across an iteration.
## See also
- [use-the-native-datagram-api.md](use-the-native-datagram-api.md) — opening
flows and receiving them, and the traps that apply to any program here
- [write-a-native-api-client.md](write-a-native-api-client.md) — the same loop
in a language with no client library
- [../reference/native-api.md](../reference/native-api.md#fipslistener) — what
`accept`, `incoming` and `set_nonblocking` guarantee
- [../reference/native-api.md](../reference/native-api.md#what-a-daemon-restart-costs)
— what a reactor sees when the daemon goes away
+314
View File
@@ -0,0 +1,314 @@
# Use the Native Datagram API
**Experimental.** The native datagram API lets a local program send
and receive datagrams addressed by pubkey and FSP port, with no IPv6
emulation and no TUN device in the path. A client connects to a Unix
socket, opens a flow to a remote pubkey, and is handed a file
descriptor it reads and writes datagrams on.
It is not a stable API surface, not a reliability layer, and not the
v2 external process API. No compatibility promise is made: the socket
protocol, the Rust client, and the configuration keys may change or be
withdrawn in any release. It is built on Linux, FreeBSD and macOS only.
**The macOS build is exercised less than the others, and you should know
by how much.** The socket type differs there: macOS has no
`SOCK_SEQPACKET` for `AF_UNIX`, so a flow descriptor is a `SOCK_DGRAM`
socket, which keeps message boundaries just the same but reports a
closed peer differently. The unit tests run on macOS in the project's
own automation and pass in full. The end-to-end suite does not: it
drives a client container against a node container over a shared
volume, and that arrangement is Linux only. So on macOS the socket
lifecycle, the descriptor hand-off across a process boundary, and the
reclaiming of a port when a client exits are covered by unit tests
rather than by anything that runs the daemon and a client as two real
processes.
That is a gap in testing, not a known defect. It is written here
because you are the one who would meet it first.
Before enabling it, read the security posture below. It is short and
it is the whole of the access control.
## Before you start: what enabling this grants
The API socket is mode `0770` owned by group `fips`, the same group as
the control socket, and that is the entire authorization model.
**Any user in the `fips` group can impersonate the node on the mesh.**
A process that can open the socket can send datagrams under this
node's identity to any peer it names, and can hold a port and receive
mesh traffic addressed to this node. Peers authenticate those
datagrams as coming from this node, because they did.
On a node with the API enabled, treat `fips` group membership exactly
as you would treat `/etc/fips/fips.key`. If the group has been handed
out so that people can run `fipsctl`, enabling the API upgrades every
one of those accounts from "can read node state" to "can speak as the
node". See
[../reference/security.md](../reference/security.md#native-datagram-api).
The API is disabled by default, and nothing in the packaging turns it
on.
## Step 1: Enable it on the daemon
Add to `/etc/fips/fips.yaml` (or a drop-in under `/etc/fips/fips.d/`):
```yaml
node:
native_api:
enabled: true
```
Every other key has a working default; see
[../reference/configuration.md](../reference/configuration.md#native-datagram-api-nodenative_api)
for the full list, including the flow ceiling and the per-flow queue
depth.
Leave `debug_commands` alone — it is off by default and belongs to the
test harness.
Restart the daemon and confirm the socket came up:
```sh
sudo systemctl restart fips
sudo journalctl -u fips | grep 'Native API socket listening'
```
The log line carries the path the daemon resolved, normally
`/run/fips/api.sock`.
## Step 2: Link the crate
The client is a module of the `fips` crate itself, so a program links
the crate and uses `fips::native::client`:
```toml
[dependencies]
fips = { git = "https://github.com/jmcorgan/fips" }
```
A checkout on the same machine can use a path dependency instead:
```toml
[dependencies]
fips = { path = "../fips" }
```
The client is blocking and std-only. It brings no async runtime, so a
plain `fn main` is enough, and a program using it does not need
`serde_json`, `libc`, or any knowledge of the socket's line protocol.
## Step 3: Open a flow and exchange datagrams
The surface is shaped like `std::net`: `FipsStream::connect` opens a
flow the way `TcpStream::connect` opens a connection, and the stream it
returns carries the datagrams.
An address is a public key and a port. The npub is that key written
down, and converting between the two is bech32 and nothing else: no
lookup, no resolution, no name service. So `("npub1...", 4600)` and
`"npub1...:4600"` name the same address, and one parameter takes either,
exactly as `ToSocketAddrs` does:
```rust
use fips::native::client::FipsStream;
use std::time::Duration;
fn main() -> Result<(), Box<dyn std::error::Error>> {
let flow = FipsStream::connect(("npub1...", 4600))?;
flow.send(b"hello")?;
// Nothing tells you a peer is never going to answer, so bound the
// wait yourself. Without this the recv below blocks for ever
// against a peer that is offline or listening elsewhere.
flow.set_read_timeout(Some(Duration::from_secs(5)))?;
// Size the buffer at the flow's own limit so no datagram it can
// carry is truncated on the way in.
let mut buf = vec![0u8; flow.max_payload()];
let len = flow.recv(&mut buf)?;
println!("{len} bytes back from {}", flow.peer_addr());
Ok(())
}
```
**`connect` contacts no peer.** It is a local registration at the
daemon, and nothing about it proves the peer exists, is reachable, or is
listening. The Berkeley shape invites the opposite reading, which is why
it is said here as well as in the API documentation.
`FipsStream::max_payload` is the daemon's answer for that flow: the
transport MTU less the FIPS encapsulation and the four-byte port header.
`FipsStream::send` refuses anything larger locally rather than letting it
be dropped further along.
`connect` uses the packaged socket path, `/run/fips/api.sock`, and an
ephemeral local port from 49152 upward. Use `connect_from(port, addr)`
when the local port matters, such as when the far end has been told it in
advance, and `connect_at(path, port, addr)` when the daemon's socket is
somewhere else.
**Bounding a wait.** `set_read_timeout` and `set_write_timeout` bound one
`recv` or one `send`, as their `TcpStream` counterparts do, and
`read_timeout` and `write_timeout` read them back. An expiring deadline
reports `WouldBlock`. `None` clears a deadline; a zero duration is
refused with `EINVAL`, because the kernel reads a zero timeout as "wait
for ever" and a caller passing zero means the opposite.
## Step 4: Receive flows from peers
`FipsListener::bind` holds a port, and `accept` blocks until a flow
arrives on it. `incoming()` is the same thing as an iterator, as it is on
`TcpListener`:
```rust
use fips::native::client::FipsListener;
fn main() -> Result<(), Box<dyn std::error::Error>> {
let listener = FipsListener::bind(4600)?;
println!("holding {}", listener.local_addr());
for arrival in listener.incoming() {
let flow = arrival?;
std::thread::spawn(move || {
let mut buf = vec![0u8; flow.max_payload()];
if let Ok(len) = flow.recv(&mut buf) {
let _ = flow.send(&buf[..len]);
}
});
}
Ok(())
}
```
`bind(0)` asks the daemon to pick the port, and `local_addr()` reports
the one it actually held: `getsockname` after `bind(2)` with port 0. Use
`bind_at(path, port)` when the daemon's socket is somewhere else. A
program wanting two independent accept loops binds two listeners.
The thread above serves **one datagram and then drops the flow**. That is
deliberate, and the fourth note below says why.
An accepted flow's `peer_addr()` names the far end by npub and port,
because the key is the address and it is the one that peer's session
authenticated. The 16-byte node address that travels on the wire is a
truncated hash of that key; it does not invert, and it appears nowhere on
this surface.
Both types implement `AsFd` and `AsRawFd`, which is the point of the
listener being a descriptor: it is readable exactly when `accept` would
not block, so a program with its own `poll`, `select` or `epoll` loop adds
it to that loop rather than dedicating a thread to blocking in `accept`.
**Prefer `AsFd`**: its borrow cannot outlive the stream, so a reactor
cannot hold a registration for a descriptor that has since been closed and
its number reissued to the next `connect`.
`set_nonblocking(true)` on either type turns a blocking call into
`WouldBlock`, which is the other way to drive a reactor. **A flow from
`accept` is blocking however its listener was set**: they are separate
sockets, so set it on the flow if the flow is what you poll.
## Four things that will bite you
**Setup leaves you nothing to keep alive.** `connect` and `bind` each
open a connection to the daemon socket, send one command, take the
descriptor off the reply and close that connection before returning. What
you hold afterwards is that descriptor and plain copies of what the reply
said, so a `FipsStream` and a `FipsListener` are `Send` and `'static`,
borrow nothing, and outlive nothing. A flow lasts exactly as long as its
own descriptor.
**Dropping a stream is how you close it.** There is no close command.
The daemon watches the descriptor and releases the flow and its port
when it goes away. Dropping a listener unbinds its port the same way, and
leaves the flows already accepted from it untouched. A program that parks
streams in a `Vec` and never removes them holds ports and flow slots
exactly as if it had leaked descriptors.
**Nothing peer-driven ever ends a flow, so your program has to.** The v1
wire carries no half-close. Nothing closes the daemon's half of a live
accepted flow, so a loop written as "echo until the flow closes", or one
that breaks on `POLLHUP`, waits for a signal that cannot arrive. The
mistake compiles, reads naturally, and passes every test that does not
involve a real daemon. What it costs is one blocked thread and one held
flow per peer, until the process dies; the node reaches its ceiling of 256
flows one silent peer at a time, and after that every `connect` on that
node, from any program, returns `EMFILE`.
Two shapes are correct, and a program that receives at all needs both.
**One exchange per flow** decides how many datagrams a flow carries, so
its end is decided rather than waited for. **A deadline you impose
yourself** — `set_read_timeout`, or a poll timeout in a reactor — bounds
a wait on a peer that may never speak again. The first decides when you
have said enough; the second decides when you have waited long enough.
A longer conversation needs an end-of-conversation marker in the payload,
because the protocol will not supply one.
Two signals are **not** a close. `Ok(0)` from `recv` is an empty datagram
and only that, so do not write `if n == 0 { break }` out of TCP habit.
`EPIPE` is real, and means the daemon went away — never that a peer
finished.
**Ports below 1024 are refused.** 0 through 255 are reserved for
protocol use and 256 through 1023 for FIPS standard services, the IPv6
shim among them. A client may hold 1024 through 65535. The refusal
applies to the remote port too, so a peer listening below 1024 is
unreachable from here.
## Step 5: Run the worked example
`examples/native-echo.rs` in the source tree is an echo server built
on nothing but this client. It holds a port and returns each datagram
to whoever sent it:
```sh
cargo run --example native-echo -- /run/fips/api.sock 4600
```
It prints `native-echo: holding port 4600` once the port is held, then
a line per datagram returned. It serves one datagram per flow by
design, for the reason the third note above gives.
## Step 6: Inspect what the node is holding
```sh
fipsctl show native-flows
```
This reports every flow the node holds — established and pending
accept — with its ports, its queue depth and its age, every bound
listener with its backlog, and the `native` counter family. The
counters are also in `fipsctl stats metrics` under `native`, where the
`drop_*` fields separate a datagram refused for having no listening
port from one dropped because a client was not reading fast enough.
For the response shape, see
[../reference/control-socket.md](../reference/control-socket.md#read-only-queries).
Reach for this when datagrams go missing. **Four places lose data with
nothing reported to your program**: a full per-flow queue, a listener that
does not accept fast enough, an outbound datagram sent before a session
exists, and an outbound datagram after the transport MTU has fallen.
[../reference/native-api.md](../reference/native-api.md#where-data-disappears)
describes each and what bounds it.
## See also
- [../reference/native-api.md](../reference/native-api.md)
— the whole surface: every type and method, the errno table, the
ceilings, the line protocol and the command reference
- [write-a-native-api-client.md](write-a-native-api-client.md)
— speaking the line protocol directly from another language
- [serve-many-peers-on-one-thread.md](serve-many-peers-on-one-thread.md)
— one `poll` loop instead of the thread per flow this guide spawns
- [../design/fips-native-api.md](../design/fips-native-api.md)
— why this exists beside the TUN path, and what it is not
- [../reference/configuration.md](../reference/configuration.md#native-datagram-api-nodenative_api)
— every `node.native_api.*` key and its default
- [../reference/security.md](../reference/security.md#native-datagram-api)
— what `fips` group membership grants once the API is on
- [../reference/control-socket.md](../reference/control-socket.md)
— the `show_native_flows` response shape
- [../reference/cli-fipsctl.md](../reference/cli-fipsctl.md)
— `fipsctl show native-flows`
+188
View File
@@ -0,0 +1,188 @@
# Write a Native API Client in Another Language
**Goal:** speak the native datagram API's line protocol directly, from C,
Python, Go or anything else, without the Rust client module.
Everything here is something the shipped Rust library already does. It is
written out so an author working where there is no such library knows what they
are reproducing. **Each of these was a real defect before it was a rule, and
each fails intermittently rather than outright.**
If you are writing Rust, you do not need this guide. Use
`fips::native::client` and see
[use-the-native-datagram-api.md](use-the-native-datagram-api.md).
For the protocol itself — the framing, the reply shapes, the commands and their
refusals — see
[../reference/native-api.md](../reference/native-api.md#the-line-protocol).
## Step 1: Read the setup connection with recvmsg, never with a buffered reader
**Every read on the setup connection is a `recvmsg` with an ancillary buffer.**
A plain `read` consumes a descriptor-bearing message's bytes with no control
buffer, and the kernel then closes the descriptor rather than queueing it. The
reply looks perfectly correct and the flow is silently gone.
This applies to both setup commands: a `listen` reply carries a descriptor as
much as a `connect` reply does. In any language it means no buffered reader, no
`BufReader`, no `readline`, and no library that wraps the socket in a stream
abstraction. Keep the line buffering in your own code, over `recvmsg`.
Read a listener's own descriptor the same way, for the ancillary data and the
close-on-exec flag. One message there is one arrival carrying exactly its own
descriptor, so there is nothing to associate.
## Step 2: Attach a descriptor to the last complete line of its read
**A descriptor belongs to the last complete line of the read that carried it,
never to the next line the reader assembles.** A `recvmsg` returning ancillary
data ends exactly at the end of the `sendmsg` that carried it, but it may begin
with any amount of data written before it.
A client that sends one command per connection reads one line and cannot hit
this. A client that pipelines two setup commands on one connection can: the
first reply and the second, descriptor-bearing reply arrive as one read, and a
reader that attached the descriptor to the first would hand the flow to the
wrong caller.
Two corollaries:
- A read that carries a descriptor and completes no line must be **reported**
rather than held. Holding it means guessing which later line it belongs to.
- A descriptor that arrives with a line you are going to discard must still be
**closed**, or the flow leaks.
Neither can happen while the daemon writes exactly one whole line per `sendmsg`
and treats a short write as an error, which it does. That is an invariant of
two programs, though, not of the socket type.
## Step 3: Treat a zero-byte read as end of file only when POLLHUP is set
**An empty datagram and a closed peer both produce a zero-byte read, and
`MSG_EOR` does not tell them apart.** On Linux 6.8, `recvmsg` on an `AF_UNIX`
`SOCK_SEQPACKET` socket returns `msg_flags == 0` for a normal message, an empty
message and end of file alike, so the flag carries no information.
`POLLHUP` does discriminate. After a zero-byte read, a queued empty datagram
leaves the socket with no events pending, while a closed peer leaves `POLLHUP`
set and latched. Poll with an events mask of **zero**, because `POLLHUP` is
reported in `revents` whether or not it was requested. The poll costs nothing:
it runs only on the zero-byte path and does not block.
Both directions of the mistake are real. Reading an empty datagram as a close
lets a peer tear down a live flow by sending nothing, and presents as a
spurious disconnect. Reading a close as an empty datagram leaves the caller
spinning on a dead flow.
Note what a close here means: the daemon went away, never a peer finishing.
## Step 4: Send with MSG_NOSIGNAL
A datagram written to a flow whose daemon half has gone, or a command written
to a daemon that has exited, raises `SIGPIPE`, whose default disposition kills
the process.
Rust ignores the signal at startup, and CPython sets it to `SIG_IGN`, so a
program in either language sees `EPIPE`. **A C or C++ client that has not
changed the disposition simply dies.** The daemon and the shipped client pass
the flag on every send.
## Step 5: Set a deadline on the setup socket
Set `SO_RCVTIMEO` on the setup connection and rewrite the resulting would-block
into `ETIMEDOUT`. The shipped client uses five seconds.
Without it, a daemon that accepted your connection and then stopped answering
blocks the setup call forever. This is also what keeps `ETIMEDOUT` to exactly
one producer on the surface, which is what lets a caller read it.
## Step 6: Keep descriptor hygiene
Five rules. Each one leaks a flow or loses one when broken.
**Request close-on-exec** with `MSG_CMSG_CLOEXEC` on the `recvmsg`, rather than
setting it afterwards. Without it the descriptor survives an `exec` into a
child, the child's reference holds the flow open after this process closes its
own, and the flow keeps its slot against the node's ceiling until the child
exits.
**Walk the whole control buffer**, not only the first header. Close extra
descriptors rather than dropping them on the floor.
**Check for truncation after taking the descriptors, not before.** A
`MSG_CTRUNC` test that returns early leaks whatever did arrive.
**Lift the descriptor out of an arrival you cannot parse** before discarding
the message. Refusing a flow is closing its descriptor; discarding the message
without taking it leaks the flow instead.
**Bound the partial line.** A daemon that stopped sending newlines would
otherwise grow your buffer without end. The shipped client caps it at 64 KiB,
well above any reply.
## Step 7: Read the errno name, never the message
The refusal's `data.errno` is the contract. The `message` is for an operator
reading a log.
A client that matched on English would break on a wording change. The shipped
client discards the message entirely so that no caller can come to depend on
it, and the daemon's own match over its error types is exhaustive precisely so
a new refusal cannot reach a client without a code.
A reply carrying no `errno` at all should be read as `ECONNREFUSED`, which
covers a daemon older than the field. The errno table is in
[../reference/native-api.md](../reference/native-api.md#the-errno-table).
## Step 8: Size the receive buffer, and add no framing
Size every receive buffer at the flow's `max_payload`, read from the reply and
never computed.
`SOCK_SEQPACKET` truncates a longer datagram, discards the remainder and
reports success. It is detectable: `recvmsg` sets `MSG_TRUNC` in `msg_flags`
when it dropped part of a message. A plain `recv` discards `msg_flags` and so
sees none of it, which is where the belief that truncation is silent comes
from. Test `MSG_TRUNC` as well, and a stale or misread `max_payload` is caught
rather than quietly corrupting a payload.
Do not add framing. There is no header and no length prefix in either
direction. One send is one datagram.
## Step 9: Decide when a flow is over, because nothing else will
Everything above is the library's job. **The termination condition is not**, in
any language.
The v1 wire carries no half-close. Nothing peer-driven closes the daemon's half
of a live flow, so a loop that reads until the flow ends does not terminate.
Your program decides when a flow is over, or nothing does.
The two shapes that work are a bounded exchange, where the program serves a
known number of datagrams per flow and then drops it, and an idle deadline,
where the program sets a read timeout and treats its expiry as the end. A
server that reads "until the flow closes" holds a thread per peer forever and
holds every flow against the node's `max_flows` ceiling.
## Verify it
The repository's test harness drives the line protocol from Python and is the
closest thing to a second implementation:
- `testing/native-api/client.py` — a thin RPC client that runs a script of
steps over one connection and checks the replies. It implements every rule
above.
- `testing/native-api/control.py` — reads `show_native_flows` back over the
control socket while a flow is open.
Check what the node actually holds with `fipsctl show native-flows`, and read
the per-cause drop counters with `fipsctl stats metrics` under `native`.
## See also
- [use-the-native-datagram-api.md](use-the-native-datagram-api.md) — the Rust
path, where none of this is your problem
- [../reference/native-api.md](../reference/native-api.md) — the surface, the
line protocol, the command reference and the errno table
- [../reference/control-socket.md](../reference/control-socket.md) — the same
line framing, for the control socket
+1
View File
@@ -19,6 +19,7 @@ guidance on when to use a feature. The "why" lives in design/; the
| [nostr-events.md](nostr-events.md) | Kind 37195 advert, Kind 21059 traversal signaling, Kind 10050 inbox relays |
| [transports.md](transports.md) | Per-transport statistics counter inventory |
| [control-socket.md](control-socket.md) | Line-delimited JSON control protocol for the daemon and gateway |
| [native-api.md](native-api.md) | Native datagram API: the Rust surface, addressing and ports, errno table, ceilings, line protocol, command reference |
| [cli-fips.md](cli-fips.md) | `fips` daemon CLI: options, exit codes, environment, files |
| [cli-fipsctl.md](cli-fipsctl.md) | `fipsctl` control-client: subcommands, options, exit codes |
| [cli-fipstop.md](cli-fipstop.md) | `fipstop` live-status TUI: tabs, keybindings |
+2 -1
View File
@@ -55,6 +55,7 @@ prints the response's `data` object as pretty JSON.
| `show transports` | `show_transports` | Transport instances: type, state, MTU, local address, per-transport stats. |
| `show routing` | `show_routing` | Routing summary: pending lookups, retry state, forwarding/discovery/error/congestion counters. |
| `show identity-cache` | `show_identity_cache` | Cached `(node_addr → npub)` entries with last-seen timestamps. |
| `show native-flows` | `show_native_flows` | Native datagram API: open and pending flows with their ports, queue depth and age, bound listeners with their backlog, and the `native` counters. |
### `acl <what>`
@@ -69,7 +70,7 @@ Time-series metrics from the in-process history rings.
| Subcommand | Control-socket command | Description |
| ---------- | ---------------------- | ----------- |
| `stats list` | `show_stats_list` | Enumerate available metrics, their units, and the per-ring retention windows. |
| `stats metrics` | `show_metrics` | Dump current counter values for every protocol metric family (`forwarding`, `discovery`, `tree`, `bloom`, `congestion`, `errors`). |
| `stats metrics` | `show_metrics` | Dump current counter values for every protocol metric family (`forwarding`, `discovery`, `tree`, `bloom`, `congestion`, `errors`, `native`). |
| `stats peers` | `show_stats_peers` | List peers tracked in stats history (active or recently active). |
| `stats history <metric> [options]` | `show_stats_history` | Fetch a time-series window for one metric. |
+68
View File
@@ -410,6 +410,67 @@ tuning under high load or on memory-constrained devices.
| `node.buffers.tun_channel` | usize | `1024` | TUN to Node outbound channel capacity |
| `node.buffers.dns_channel` | usize | `64` | DNS to Node identity channel capacity |
### Native Datagram API (`node.native_api.*`)
**Experimental, off by default, and built on Linux, FreeBSD and macOS only.** A
client process connects to a Unix socket and asks either to open a flow to a
remote pubkey or to hold a local port. Both answers carry a file descriptor:
a flow's, which the client sends and receives datagrams on, or a listener's,
which arriving flows are delivered on. There is no IPv6 emulation and no TUN
device on this path.
The surface is not stable, is not a reliability layer, and is not the v2
external process API. No compatibility promise is made about it: the keys
below, the line protocol behind them, and the Rust client that hides it may
change or be withdrawn in any release.
The listener is not built on macOS or Windows, and this section is ignored
there. Two separate things bound that: Windows has no `SCM_RIGHTS` and so no
way to hand a file descriptor to another process at all, while macOS has
`SCM_RIGHTS` but does not implement `SOCK_SEQPACKET` for `AF_UNIX`.
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `node.native_api.enabled` | bool | `false` | Enable the native API socket |
| `node.native_api.socket_path` | string | *(auto)* | Socket file path. Resolved the same way as the control socket, with the filename `api.sock`: `/run/fips/api.sock` when `/run/fips` exists; then `/var/run/fips/api.sock` on FreeBSD when its private directory exists; then `$XDG_RUNTIME_DIR/fips/api.sock`; finally `/tmp/fips-api.sock` |
| `node.native_api.pending_per_flow` | usize | `16` | Datagrams held for one flow while it waits to be accepted, or while an established flow's client is slow to read. Refused above 64 at startup: the whole batch is written onto a socket pair the client cannot read yet. Refused below 1: a flow that can hold nothing loses its peer's opening datagram between the arrival being announced and the client taking the flow |
| `node.native_api.backlog` | usize | `16` | Flows announced on one listener and not yet taken by its task. Refused below 1 at startup: a listener with no backlog admits no flow, so every arrival would be dropped |
| `node.native_api.max_flows` | usize | `256` | Flows this node holds at once |
| `node.native_api.debug_commands` | bool | `false` | Answer the `inject`, `stats` and `arrive` debug commands. Not a supported interface |
> **Security note:** the socket is mode `0770`, owned by group `fips`, and
> that is the whole of the authorization model. **Any user in the `fips` group
> can impersonate this node on the mesh.** A process that can open the socket
> can send datagrams under this node's identity to any peer it names, and can
> hold a port and receive mesh traffic addressed to this node on it. There is
> no per-client authentication, no capability check and no audit trail beyond
> the daemon's own logs. On a node with the native API enabled, treat `fips`
> group membership exactly as you would treat the node's private key. This is
> why the API is disabled by default, and why enabling it is an explicit
> operator decision rather than something a package turns on. See
> [security.md](security.md#native-datagram-api).
A client may hold ports 1024 through 65535. Ports 0 through 255 are reserved
for protocol use and 256 through 1023 for FIPS standard services (the IPv6
shim among them), and the daemon refuses both ranges by name. A client that
names no local port is given one from 49152 upward.
`debug_commands` is a separate gate on three commands that exist only so the
test harness can drive the receive and dispatch paths without a wire.
`inject` makes the daemon write bytes the client chose into one of that
client's own flows, and `arrive` makes it dispatch a datagram as though a peer
had sent it, reaching any listener this node holds. A node with the key off
refuses each by name, so a client can tell "this node will not do that" from
"this build has no such command". Leave it off outside the test harness.
`fipsctl show native-flows` reports the open and pending flows, the bound
listeners and the `native` counters; see the [`fipsctl`
reference](cli-fipsctl.md) and
[control-socket.md](control-socket.md#read-only-queries). A Rust program links
the crate and speaks the API through the `fips::native::client` module, which
hides the line protocol; see
[../how-to/use-the-native-datagram-api.md](../how-to/use-the-native-datagram-api.md).
## TUN Interface (`tun.*`)
| Parameter | Type | Default | Description |
@@ -1019,6 +1080,13 @@ node:
control:
enabled: true
socket_path: null # null = auto (platform runtime dir → XDG → /tmp)
# native_api: # uncomment to enable the experimental native datagram API
# enabled: true # opt-in, default false; not on Windows
# socket_path: /run/fips/api.sock # omit the key for the resolution above
# pending_per_flow: 16 # datagrams held for one flow; 1..=64
# backlog: 16 # flows announced on one listener, awaiting its task; at least 1
# max_flows: 256 # flows this node holds at once
# debug_commands: false # inject/stats/arrive; test harness only
buffers:
packet_channel: 1024
tun_channel: 1024
+2 -1
View File
@@ -125,9 +125,10 @@ table below lists every command currently registered.
| `show_transports` | — | `transports[]` — `transport_id`, `type`, `state`, `mtu`, `name`, `local_addr`, optional `tor_mode`, `onion_address`, `tor_monitoring`, `stats`. |
| `show_routing` | — | `coord_cache_entries`, `identity_cache_entries`, `pending_lookups[]`, `pending_tun_destinations`, `pending_tun_packets`, `recent_requests`, `retries[]`, `forwarding`, `discovery` (request/response sub-counters; includes `req_deduplicated` — requests suppressed as recent duplicates — and `req_dedup_cache_full` — requests admitted because the dedup cache was full), `error_signals`, `congestion`. |
| `show_identity_cache` | — | `entries[]`, `count`, `max_entries`. Each entry: `node_addr`, `npub`, `display_name`, `ipv6_addr`, `last_seen_ms`, `age_ms`. |
| `show_native_flows` | — | `flows[]`, `listeners[]`, `stats` (the `native` counter family). Each flow: `flow_id`, `peer` (the peer's npub, which is its address; always present, because the flow carries the key its client named or its session authenticated), `peer_addr` (the 16-byte node address in hex — a truncated hash of the same key, kept because it is what `show_sessions` and `show_routing` key on), `local_port`, `remote_port`, `state` (`established` / `pending_accept`), `queued` (datagrams the node is holding for the flow), `age_ms` (time since the flow reached its current state: opened for a flow this node opened, accepted for one taken off a listener, announced for one still pending — accepting a pending flow restarts the clock). Each listener: `local_port`, `backlog`. |
| `show_listening_sockets` | — | `fips0_addr`, `firewall_active` (bool — `inet fips` table loaded), `sockets[]`. Each entry: `proto` (`tcp` / `udp`), `local_addr` (`::` or the node's fd00::/8 address), `port`, `pid` (nullable), `process` (nullable), `wildcard_bind` (bool — `local_addr == ::`), `filter` (`accept` / `drop` / `unknown` / `no_firewall`). Linux-only; returns an empty `sockets[]` on other platforms. |
| `show_stats_list` | — | `metrics[]` (each with `name`, `unit`, `scope`), `fast_ring_seconds`, `slow_ring_minutes`, `peer_retention_seconds`. |
| `show_metrics` | — | Flat snapshot of every counter family in the metrics registry: `forwarding`, `discovery`, `tree`, `bloom`, `congestion`, `errors`. Each value is that family's counter snapshot object. Counter-only — gauges/histograms that need the live node are excluded. Served off the main loop. Silent-rejection sites classify their reason as a typed `RejectReason` and increment the matching per-family counter exposed here — see [Rejection reasons](#rejection-reasons). |
| `show_metrics` | — | Flat snapshot of every counter family in the metrics registry: `forwarding`, `discovery`, `tree`, `bloom`, `congestion`, `errors`, `native`. Each value is that family's counter snapshot object. Counter-only — gauges/histograms that need the live node are excluded. Served off the main loop. Silent-rejection sites classify their reason as a typed `RejectReason` and increment the matching per-family counter exposed here — see [Rejection reasons](#rejection-reasons). |
| `show_stats_history` | `metric` (req), `peer` (req for per-peer metrics), `window` (`<N>s` / `<N>m` / `<N>h`, default `10m`), `granularity` (`1s` / `1m`, default `1s`) | A single `Series`: `metric`, `unit`, `granularity_seconds`, `values[]`. |
| `show_stats_all_history` | `peer` (optional npub), `window`, `granularity` | `granularity_seconds`, `window_seconds`, `peer`, `series[]` (one per metric). |
| `show_stats_peers` | — | `peers[]`, `count`. Each entry: `npub`, `node_addr`, `display_name`, `is_active`, `first_seen_secs_ago`, `last_contact_secs_ago`. |
+741
View File
@@ -0,0 +1,741 @@
# Native Datagram API
The native datagram API lets a program send and receive datagrams over the
mesh addressed by public key and port, with no IPv6 emulation and no TUN
device. A program opens a flow to a peer, receives a socket descriptor, and
uses ordinary socket calls on it.
This document describes the surface. For the steps to enable it and write a
first program, see
[../how-to/use-the-native-datagram-api.md](../how-to/use-the-native-datagram-api.md).
**Status: experimental.** The Rust surface and the line protocol may both
change. The API is off by default and is Unix only.
## Enabling it
Every bound is a key under `node.native_api`. The defaults are the values
compiled into `src/config/node.rs`.
| Key | Default | What it bounds |
| --- | ------- | -------------- |
| `enabled` | `false` | Whether the socket is bound at all. |
| `socket_path` | resolved | See [Socket path](#socket-path). |
| `pending_per_flow` | `16` | Datagrams held for one flow, whether it awaits hand-off or its program is slow to read. Ceiling 64, floor 1, both applied when the configuration loads. |
| `backlog` | `16` | Flows announced on one listener and not yet taken. Floor 1. |
| `max_flows` | `256` | Flows this node holds at once, across every program. Established and pending flows both count. |
| `debug_commands` | `false` | Whether the three debug commands are answered. |
There is no cap on flows per program, and no notion of a program to hang one
on. One program can reach `max_flows` by itself, and the ceiling is shared
with every other program on that node.
See [configuration.md](configuration.md#native-datagram-api-nodenative_api)
for the full YAML reference, and [security.md](security.md#native-datagram-api)
for what `fips` group membership grants once the API is on.
## Socket path
`node.native_api.socket_path` has a default resolved at startup rather than a
fixed string. The resolver takes the first of these that applies and appends
`api.sock`:
1. `/run/fips/api.sock`, when `/run/fips` exists as a directory. This is the
packaged Linux convention, and the constant the shipped Rust client compiles
in on Linux.
2. `/var/run/fips/api.sock` on FreeBSD and macOS, whose service scripts and
LaunchDaemons create that directory. Linux skips this step, and macOS has no
`/run` for step 1 to find, so this is where a packaged macOS daemon lands and
is the constant the client compiles in there.
3. `$XDG_RUNTIME_DIR/fips/api.sock`, when that variable names an existing
directory. A development run usually gets this one.
4. `/tmp/fips-api.sock`, the last resort.
Selection is by existence of the directory, not by whether the client can
write to it. A client should therefore take the path as an argument with these
as its defaults, rather than compile one in.
## Addressing
An address is an x-only secp256k1 public key and a port. Nothing else
identifies an end. There is no node identifier, no address family, and no
scope.
The npub is the key written down. Conversion between the two is bech32 and
nothing else: no lookup, no resolution, no name service. The conversion is
local, needs no daemon, and fails identically whether or not one runs.
The 16-byte node address that the wire carries never appears on this surface.
It is a truncated hash of the public key, it does not invert, and no call here
accepts or returns one.
### Ports
Ports are FSP ports, carried inside the encrypted FSP envelope. They are
tiered.
| Range | Use | What a program may do |
| ----- | --- | --------------------- |
| 0-255 | Protocol use | Refused. |
| 256-1023 | FIPS standard services | Refused. Port 256 is the IPv6 shim. |
| 1024-49151 | Application, well known | Name it explicitly. |
| 49152-65535 | Application, ephemeral | Name it explicitly, or ask for 0 and be given one from this range. |
Nothing reserves the upper range against explicit use. A service that wants a
fixed port above 49151 may hold one, and ephemeral allocation will not take it
away, because the sweep skips a port already held.
**The refusal applies to the remote port as well as the local one.** A program
cannot send to a peer's port 256 and inject into its IPv6 plane. By the same
rule, a peer listening below 1024 is unreachable from here: the `connect` is
refused with `EADDRNOTAVAIL` before any node state is touched.
**Port 0 means "any port"**, for both calls, and it is the only spelling of
that. This matches `bind(2)` with port 0. `local_addr()` afterwards reports the
port actually held, which is `getsockname`.
A local port has exactly one owner, a listener or one connected flow, never
both. A connected flow owns its local port and frees it when the flow closes.
A flow accepted from a listener shares the listener's port and does not own it,
so closing an accepted flow frees no port; the port returns when the listener
drops. Ephemeral allocation sweeps forward from the last claim and wraps once,
so a port is not immediately reused after release.
## The Rust surface
The crate module is `fips::native::client`. It is gated to Unix targets.
```rust
use fips::native::client::{FipsAddr, FipsListener, FipsStream, ToFipsAddr};
```
`SOCKET` is a `&str` constant holding `/run/fips/api.sock`, the packaged path.
`XOnlyPublicKey` is re-exported from `secp256k1` so a caller needs no direct
dependency on that crate.
### The Berkeley mapping
The descriptor is a real socket. Everything a program does with a flow after
setup is the operating system, not this API.
| Berkeley | Here | Notes |
| -------- | ---- | ----- |
| `socket()` | none | There is no unbound object to make. `connect` and `bind` each return one already bound. |
| `bind()` | `FipsListener::bind` | One setup call. Port 0 asks the daemon to choose. |
| `listen()` | none | Folded into `bind`. The backlog is node configuration, not a caller's argument. |
| `connect()` | `FipsStream::connect` | One setup call. A local registration that contacts no peer. |
| `accept()` | `FipsListener::accept` | Blocks until a flow arrives. One `recvmsg`, no round trip to the daemon. |
| `send()` | `FipsStream::send` | One datagram, whole or not at all. One local size check, then one syscall. |
| `recv()` | `FipsStream::recv` | One datagram. `Ok(0)` is an empty one and only that. |
| `poll()`, `select()`, `epoll` | the same calls | Both types are `AsFd` and `AsRawFd`. |
| `fcntl(O_NONBLOCK)` | `set_nonblocking()` | On both types. Read-modify-write, so a flag the caller set survives. |
| `close()` | drop | Dropping releases the flow or unbinds the port. There is no method. |
| `getsockname()` | `local_addr()` | A field read of what setup reported. It cannot fail. |
| `getpeername()` | `peer_addr()` | The same, for the far end. |
| `SO_RCVTIMEO` | `set_read_timeout()` | On a flow, with `set_write_timeout` and both getters. |
| `SO_RCVBUF`, others | via `AsFd` | `setsockopt` on the descriptor. The surface exposes no option beyond the deadlines. |
| `inet_pton()` | npub decode | bech32, locally, with no lookup. |
| `inet_ntop()` | npub encode | Its exact inverse. |
Four differences from Berkeley remain. There is no `socket()` and no
`listen()`, because a descriptor cannot exist before the daemon has agreed to
make one. A successful `connect` proves nothing about the peer. `Ok(0)` from
`recv` is an empty datagram rather than end of file. And there is a payload
ceiling, reported per flow, which TCP has no equivalent of.
### FipsAddr
One end of a flow. The fields are private, so an address in hand always names
a valid key.
`Copy`, `Debug`, `Clone`, `PartialEq`, `Eq`. **It derives neither `Hash` nor
`Ord`**, so it cannot key a `HashMap` or a `BTreeMap`. Key on
`addr.key().serialize()` or on `addr.to_string()` instead.
`std::net::SocketAddr` derives both, so the omission is a surprise rather than
a convention.
| Call | Returns, and how it fails |
| ---- | ------------------------- |
| `new(key, port)` | The address. Infallible: both arguments are already well typed. Port 0 and the reserved tiers are representable and are refused later, by the daemon, at setup. |
| `key()` | The public key, by value. Infallible. |
| `port()` | The port. Infallible. |
| `Display` | Writes `npub1...:4242`. Pure bech32, no daemon, no lookup. |
| `FromStr` | Reads that back. `EINVAL` for no colon, a port that is not a `u16`, or a head that is not an npub. An nsec is refused despite being bech32 of the right length. |
`Display` and `FromStr` are exact inverses. Parsing splits on the last colon
and is purely local, so it fails identically with no daemon running. The colon
is unambiguous because the bech32 character set does not contain one.
### ToFipsAddr
The mirror of `std::net::ToSocketAddrs`. One generic parameter takes every
spelling of one address. It is taken by `&self`, so a caller never gives up
ownership, and **no implementation opens a socket or contacts the daemon**. The
only error any of them produces is `EINVAL`.
| Implementation | Notes |
| -------------- | ----- |
| `FipsAddr` | Returns a copy. Structurally infallible. |
| `(XOnlyPublicKey, u16)` | The binary address, already parsed. Structurally infallible. |
| `([u8; 32], u16)` | 32 raw bytes. The one binary form that can be rejected: `EINVAL` if the bytes are not a valid x-only public key, which `[0u8; 32]` is not. |
| `(&str, u16)` | An npub and a port. `EINVAL` if the npub does not decode. The port is taken as given and range-checked later. |
| `(String, u16)` | The same, for a string from `argv` or a config file. |
| `str` | The whole address as `"npub1...:4242"`. Delegates to `FromStr`. |
| `String` | The same. |
| `&T where T: ToFipsAddr` | A blanket implementation over references, so `&addr` and `&&str` work at all. |
There is **no associated iterator type**, unlike `ToSocketAddrs`. An address
here resolves to exactly one endpoint.
The concrete implementations are on the unsized `str` rather than on `&str`.
That is what lets the blanket reference implementation cover `&str` and
`&&str` alike. It makes no difference at a call site.
### FipsStream
One datagram flow, and the descriptor it rides on. A flow is an exact match of
both ends and both ports. **The descriptor is the flow**: it lives while a
process holds that descriptor and ends when the last one closes it.
`Send + Sync + 'static`, with no `Arc` and no borrow. There is **no `Clone` and
no `try_clone`**. Because `send` and `recv` both take `&self`, a shared borrow
across two scoped threads is a full-duplex reader and writer. A program wanting
an owned writer handle reaches for `Arc<FipsStream>`.
#### Constructors
| Call | What it does |
| ---- | ------------ |
| `connect(addr)` | A flow to `addr` from an ephemeral local port, at the packaged socket path. |
| `connect_from(local, addr)` | The same from a named local port. `connect_from(0, addr)` is exactly `connect(addr)`. |
| `connect_at(sock, local, addr)` | The full form, naming the daemon's socket. What a development node and a test harness need. |
All three resolve the address first, so a bad address is `EINVAL` with no
socket opened. Then, in this order: whatever `UnixStream::connect` gives for
the socket path (`NotFound` when no daemon runs or the API is disabled,
`PermissionDenied` on socket permissions, `ECONNREFUSED` when the path exists
but nothing is accepting); `EPIPE` if the daemon vanishes mid-write;
`ETIMEDOUT` if no answer arrives inside the five-second setup deadline; then
the daemon's own refusals; then `InvalidData` for a reply this client cannot
parse.
A stream does not outlive its setup connection, because the connection is not
an object at all. It is a local binding inside the constructor, dropped before
the call returns. What comes back holds an owned descriptor and four `Copy`
scalars.
#### Methods
**`send(&self, buf: &[u8]) -> io::Result<()>`** sends one datagram. There is no
byte count, because a `SOCK_SEQPACKET` message is delivered whole or not at
all. `buf.len() > max_payload()` is `EMSGSIZE` before any syscall, so a caller
learns which datagram was too large rather than finding a gap at the far end.
`EPIPE` means the daemon has gone. Takes `&self`, so it can be called
concurrently with `recv` from another thread.
The library sends with `MSG_NOSIGNAL` rather than writing to the descriptor. A
Rust binary ignores `SIGPIPE` at startup anyway, but a C program that has
loaded this crate would otherwise be killed by a write to a departed daemon
rather than told about it. `EINTR` is retried internally.
`Ok(())` means only that the local socket pair accepted the bytes. See
[Where data disappears](#where-data-disappears).
**`recv(&self, buf: &mut [u8]) -> io::Result<usize>`** receives one datagram
and returns its length. It **blocks by default**: the descriptor arrives
blocking and no timeout is set on it. `set_nonblocking(true)` makes it return
`WouldBlock` instead. `set_read_timeout` bounds a single wait without going
non-blocking. A datagram longer than `buf` is truncated and the remainder
discarded, which is `SOCK_SEQPACKET` behaviour and is not reported; size `buf`
at `max_payload()` and it cannot happen. `EINTR` is retried.
**`Ok(0)` is an empty datagram and only that.** An empty datagram and a closed
peer both produce a zero-byte read, and `MSG_EOR` does not tell them apart. On
the zero-byte path only, the library polls the descriptor with an events mask
of zero and reads `POLLHUP` from `revents`, returning `EPIPE` for a genuine
close and `Ok(0)` otherwise. Without that step a peer could tear down a live
flow by sending nothing. The `EPIPE` that `recv` returns means the daemon went
away, never that a peer finished.
**`peer_addr()`** and **`local_addr()`** return `FipsAddr`, not
`io::Result<FipsAddr>`, unlike their `TcpStream` counterparts. These are field
reads of what the setup reply already carried. `local_addr()` carries the
node's own public key, which on an accepted stream comes from the arrival
message rather than from the listener, so a node holding more than one identity
still answers correctly.
**`max_payload()`** returns the largest datagram this flow will carry.
Infallible. **It is a snapshot** taken when the flow opened and does not track
a later MTU change.
**`set_read_timeout(Option<Duration>)`** and **`set_write_timeout`** bound one
`recv` or one `send`, as their `TcpStream` counterparts do. `None` clears.
**A zero duration is `EINVAL`**, because the kernel reads a zero timeout as
"wait for ever" and a caller passing zero means the opposite. An expiring
deadline reports `WouldBlock`, the same answer a non-blocking descriptor gives.
**`read_timeout()`** and **`write_timeout()`** read them back, `None` when
unset. The two directions are separate options, and setting one leaves the
other alone.
**`set_nonblocking(bool)`** puts the flow in or out of non-blocking mode. In
that mode `recv` returns `WouldBlock` rather than waiting, and so does `send`
when the daemon is not draining the flow fast enough. **A flow from `accept` is
blocking however its listener was set**: they are separate sockets and the
daemon hands over a fresh one.
**`AsFd`** and **`AsRawFd`** both return the flow's descriptor. Prefer `AsFd`:
its borrow cannot outlive the stream, so a reactor cannot hold a registration
for a descriptor that has since been closed and its number reissued. **The
stream keeps ownership** either way. Do not close the descriptor and do not
wrap it in anything that takes ownership, because the stream's own drop is what
releases the flow. This is the route to `SO_RCVBUF` and every other socket
option the surface does not expose.
A descriptor that survives an `exec` into a child holds the flow open after
this process closes its own copy, and the flow keeps its slot against the
node's ceiling until the child exits. The library requests close-on-exec when
it takes the descriptor, so this happens only if a program deliberately clears
the flag.
### FipsListener
A held local port, and the descriptor arriving flows are delivered on. **The
listener is a descriptor**, so it joins an existing `poll`, `select` or `epoll`
loop with no new mechanism, and `accept` is one `recvmsg` on it.
`Send + Sync + 'static`, no `Clone`, no `try_clone`, no explicit `Drop`.
**`bind(port)`** holds `port`, or an ephemeral one when `port` is 0, at the
packaged path. **`bind_at(sock, port)`** is the same, naming the daemon's
socket. Failures are the socket-path and setup-deadline rows above, plus
`EADDRINUSE` when a listener or a flow already holds the port, and
`EADDRNOTAVAIL` for a reserved tier or an exhausted ephemeral range.
**`accept() -> io::Result<(FipsStream, FipsAddr)>`** blocks until a flow
arrives. One `recvmsg`, one arrival: a `SOCK_SEQPACKET` message carries exactly
its own descriptor. Takes `&self`, so several threads may accept on one
listener concurrently. It does not map a would-block to `ETIMEDOUT`, so a
caller that put the listener in non-blocking mode sees a raw `WouldBlock`.
Failures: `EPIPE` when the message carried no descriptor, meaning the daemon
closed its half; `InvalidData` for an arrival that is not JSON or is missing a
field; and `io::Error::other` if the control buffer overflowed and a descriptor
was lost.
Whatever the peer sent before `accept` returned is already queued on the
returned stream. **Refusing a flow is dropping the stream**, because there is
no other way to refuse one. An unparseable arrival is therefore reported only
after the descriptor it carried has been taken into ownership, so a parse
failure refuses the flow rather than leaking it.
**`incoming()`** returns an `Incoming<'_>`, which borrows the listener for the
iterator's lifetime, so the listener cannot be moved or dropped mid-iteration.
**`local_addr()`** returns the port actually held with the node's own key, not
an `io::Result`.
**`set_nonblocking(bool)`** puts the listener in or out of non-blocking mode.
In that mode `accept` returns `WouldBlock` when no flow has arrived, and so
does every `incoming` item, which makes that iterator spin unless the caller
waits on the descriptor between items. It does not reach the flows the listener
yields: each arrives blocking.
**`AsFd`** and **`AsRawFd`** return the listener's descriptor. **`poll` on it
reports readable exactly when `accept` would not block.**
**There is no `set_read_timeout` here**, following `TcpListener`, which has
none either. A bounded accept is `set_nonblocking` plus a wait of the caller's
own on the descriptor.
**Dropping** closes the descriptor and unbinds the port. Flows already accepted
from it are untouched; flows still pending on it go with it.
### Incoming
The iterator `incoming()` returns. Its item is `io::Result<FipsStream>`: it
calls `accept` and discards the peer address, which is recoverable as
`stream.peer_addr()`.
**It never returns `None`.** A listener has no last flow, and a failed accept
is yielded as an `Err` item rather than ending the iteration, so a `for` loop
over it never falls through. A reader who writes `let flow = flow?;` inside the
loop exits it on the first transient error, which is a different shape from
`TcpListener` habits.
## Errors
There is no bespoke error type. Every call returns `io::Result`, so the surface
matches `std::net` and a future C binding can return the number directly.
**Match on `err.raw_os_error()` against `libc` constants, never on
`err.kind()`.** `EMFILE` and `EMSGSIZE` both carry
`ErrorKind::Uncategorized`, which is `#[non_exhaustive]` and cannot be named in
a match arm, so `kind()` cannot distinguish the node's flow ceiling or an
oversize datagram from anything else.
### The errno table
Compare against the `libc` constant, never against a literal. The client maps
each name onto `libc::<NAME>` for the platform it was built for, so the number
differs between Linux, FreeBSD and macOS.
| errno | `kind()` | What causes it |
| ----- | -------- | -------------- |
| `EADDRINUSE` | `AddrInUse` | The local port is held, or a flow already exists between those two ends on those two ports. |
| `EADDRNOTAVAIL` | `AddrNotAvailable` | A port in a reserved tier, local or remote, or the ephemeral range exhausted. Reserved rather than `EACCES`, because no program however privileged may hold port 256. |
| `EMFILE` | `Uncategorized` | The node is at `max_flows`, or the socket pair could not be made. Unmatchable by kind. |
| `EINVAL` | `InvalidInput` | An address that does not resolve locally, or a malformed command. |
| `EMSGSIZE` | `Uncategorized` | A `send` above `max_payload()`. Raised locally, before the syscall. Unmatchable by kind. |
| `EPIPE` | `BrokenPipe` | The daemon went away: it exited, the node shut down, or its own read failed. Never a peer finishing. |
| `ETIMEDOUT` | `TimedOut` | The five-second setup deadline, and nothing else on this surface. |
| `ECONNREFUSED` | `ConnectionRefused` | The node is shutting down, a debug command is disabled, nothing is accepting on the socket path, or the reply carried an errno name this client has no row for. |
The last row is the catch-all, so `ECONNREFUSED` is the one code that does not
narrow the cause much. **The daemon sends errno names rather than numbers**,
because a number belongs to the platform the program was built for and the
daemon is not it. An unknown name is read as `ECONNREFUSED` for the same reason
a missing one is. The daemon also sends a human-readable message, and the
client discards it, so no caller can come to depend on prose.
### The errors that carry no errno
Three shapes have no `raw_os_error()` at all.
**`ErrorKind::InvalidData`** and **`io::Error::other`** both mean this daemon is
not speaking the protocol: a reply or an arrival that is not JSON, an unknown
or missing status, a missing or malformed field, a port outside `u16`, a reply
carrying no descriptor, a descriptor arriving on a read that completed no line,
more than 64 KiB with no newline, or a truncated control message. None is worth
retrying, and none is a condition a correct daemon produces.
**`ErrorKind::NotFound`** on a setup call means no daemon is running, or the
native API is disabled, which is the default.
### What is worth retrying
`EMFILE` and `EADDRNOTAVAIL` from an exhausted ephemeral range are load
conditions and may clear. A retry with a backoff is reasonable, and a program
that retries without one contributes to the exhaustion. `ETIMEDOUT` and `EPIPE`
mean the daemon is unhealthy or gone, so the useful retry is the whole setup
sequence and not the one call. `EADDRINUSE`, `EADDRNOTAVAIL` from a reserved
tier, `EINVAL` and `EMSGSIZE` are decisions about the arguments and produce the
same answer every time.
## Where data disappears
The API is datagram-unreliable. Nothing on this surface confirms delivery.
There is no acknowledgement, no retransmission, no ordering guarantee and no
flow control between the two ends. A program that needs confirmation gets it
from the peer, in the payload.
Four places lose data with nothing reported to the client.
**A full per-flow queue.** Inbound datagrams beyond `pending_per_flow` are
dropped with a trace log and no client-visible signal. A program that stops
reading a flow loses datagrams and is never told.
**A listener that does not accept fast enough.** Whole arriving flows are
discarded, counted inside the daemon as `ListenerNotReading`. The bound on how
many flows a stalled listener holds is not the backlog: the send buffer on a
listener's pair is sized generously, because the approximation must err toward
accepting an arrival a program would have read. A program that binds a listener
and stops reading it accumulates flows on the order of `max_flows` rather than
of `backlog`.
**An outbound datagram sent before a session exists.** The first `send` on a
new flow almost always takes this path, because `connect` contacts no peer and
leaves no FSP session behind it. The node holds the datagram and starts a
handshake. What holds it is bounded twice: at
`node.session.pending_packets_per_dest` datagrams for one destination, past
which a further datagram evicts the oldest one held, and at
`node.session.pending_max_destinations` destinations, past which a new
destination's datagram is dropped outright. The defaults are 16 and 256.
Neither eviction reaches the caller.
**An outbound datagram after the MTU has fallen.** `max_payload()` is a
snapshot taken at setup. The daemon re-checks each outbound datagram against
the node's current limit and drops it silently if the transport MTU has since
fallen.
### The drop causes, and what they mean
An inbound datagram can be refused for eight reasons, which render as seven
texts: a pending flow's full queue and an established flow's are distinct to a
counter and alike to a client. The texts are what `DropReason::as_str` produces;
the counter names are what `fipsctl stats metrics` reports under `native`.
| Text | Counter | Condition |
| ---- | ------- | --------- |
| `no listener or flow on that port` | `drop_no_port` | Nothing holds the destination port. |
| `listener backlog full` | `drop_backlog_full` | A listener holds the port but will not hold another pending flow. |
| `node flow ceiling reached` | `drop_too_many_flows` | The node is at `max_flows`. |
| `queue full` | `drop_pending_queue_full`, `drop_flow_queue_full` | A flow's queue is full, whether it is pending or established with a client that is not reading. Two counters, one text. |
| `arrival queue full` | `drop_arrival_queue_full` | The daemon's own queue to a listener's task is full. |
| `listener not reading arrivals` | `drop_listener_not_reading` | A listener's client is not reading its descriptor, so the arrival could not be written to it. |
| `listener closed` | `drop_listener_gone` | A listener's client closed its descriptor between the arrival being taken off the queue and being written. |
**None of these reaches a client for real traffic.** There is no drop event and
no reply reports one. The texts are visible only through the debug `arrive`
command, which reports `dropped: <text>` as its outcome; the counters are
readable at any time through the control socket.
`drop_oversize` is a ninth counter and is not in this table, because it is not a
dispatch refusal: it counts an **outbound** datagram the daemon discarded when
the transport MTU had fallen below it, which is the fourth case above.
### No framing, in either direction
There is no header and no length prefix. One send is one datagram and one
receive is one datagram. A program that adds its own length prefix on top of a
message-boundary-preserving transport is paying for something it already has.
### Seeing what a node holds
Neither type reports the node's state, and the client module exposes no
statistics. `fipsctl show native-flows` is the only way to see what a node
holds, and `fipsctl stats metrics` under `native` carries the per-cause drop
counters that no program is told about. See
[cli-fipsctl.md](cli-fipsctl.md) and
[control-socket.md](control-socket.md).
## What a daemon restart costs
**Every flow and every listener ends when the daemon does**, and descriptors do
not survive it. The daemon's halves close with the process, so a program reading
one gets `EPIPE` and a program writing one gets the same. There is no
reconnection and no resumption: a program that must survive a restart re-runs
its setup calls and gets new descriptors.
**Datagrams already sent and not yet forwarded are lost.** At an orderly flow
close there is nothing in flight, because a flow's reader forwards what the
client wrote and only then notices the flow is over. At daemon exit every reader
stops at once with no such notice. The window is small and nothing bounds it.
A program that needs to know its last datagram reached a peer needs an
acknowledgement from that peer. Neither this API nor the wire beneath it has one
to offer.
## The line protocol
The Rust client hides all of this. It is documented because a client in
another language has to implement it. For the obligations such a client
carries, see
[../how-to/write-a-native-api-client.md](../how-to/write-a-native-api-client.md).
### Framing
The encoding is the control socket's, so a client that speaks one speaks both.
One JSON object per line, terminated by `\n`.
A request is:
```json
{"command": "connect", "params": {"peer": "npub1...", "remote_port": 4242}}
```
Every command requires `params`. A request without the key is refused with
`command '<name>' requires params`.
A command line is capped at 8192 bytes. A longer one ends the connection with
`native API command too large`. **That is the only condition that ends the
connection instead of producing a reply.** The cap is applied whether or not
the chunk in hand holds the newline, so a well formed 9000-byte line ends the
connection exactly as a runaway one with no newline does. The connection is
dropped with nothing written back, so a client sees end of file and never a
message.
Two kinds of line come back, and only two. A success:
```json
{"status": "ok", "data": {}}
```
And a refusal, which carries the code a client acts on and prose it must not
match against:
```json
{"status": "error", "data": {"errno": "EADDRNOTAVAIL"},
"message": "port 256 is reserved for FIPS standard services"}
```
**There is no third kind.** No event is pushed on this connection and nothing
arrives on it unsolicited, so a reader that sends a command and takes the next
complete line as its answer is correct. A refusal is a normal reply line and
never drops the connection. A reply carrying no `errno` at all is read as
`ECONNREFUSED`, which covers a daemon older than the field.
### The arrival message
**An arriving flow is not a line.** It is one `SOCK_SEQPACKET` message on the
listener's own descriptor, with **no trailing newline**: the message boundary
is the framing, and a newline would offer a client a second framing to rely on.
Its seven fields are `flow_id`, `peer` and `node` as npubs, `local_port`,
`remote_port`, `max_payload` and `held`. The `node` field is carried so an
accepted flow can answer `local_addr` without consulting the listener that
produced it. `flow_id` is the same identifier a `connect` reply carries and the
name the debug commands take.
**`held` is the one field a client cannot infer.** It counts the datagrams the
daemon has already written onto the flow's descriptor before this message,
because the hand-off writes every held datagram first and the arrival last. It
says exactly how many `recv` calls a client may make on a newly accepted flow
without blocking. The Rust client reads neither `held` nor `flow_id`, because
neither names anything a caller can name. A client in another language may
ignore both on the same reasoning; it cannot ignore that they are there.
### Passing the descriptor
The daemon builds an `AF_UNIX` `SOCK_SEQPACKET` socket pair with
`SOCK_CLOEXEC`, keeps one half, and sends the other over `SCM_RIGHTS` in the
ancillary data of the same `sendmsg` that carries the reply line. It makes its
own half non-blocking and leaves the client's half blocking. `SOCK_SEQPACKET`
is what preserves message boundaries in both directions, which is why the
payload needs no framing.
A refused `connect` leaves the port free: the socket pair is built before the
port is claimed, so a failure to build it needs no rollback.
### The setup call
One `AF_UNIX` `SOCK_STREAM` connection to the socket path, one command line
written, one reply line read. Replies come back in command order, and several
commands are permitted on one connection, so a client may pipeline. The shipped
Rust client does not: it opens a connection per setup call and drops it before
returning.
**The connection owns nothing.** Closing it releases no flow and no listener,
and a descriptor kept across the close keeps working. What owns the flow is the
descriptor.
## Command reference
Five commands. Two are the interface; three are debug scaffolding, off by
default.
| Command | Carries FD | What it does |
| ------- | ---------- | ------------ |
| `connect` | yes | Open a flow to a peer named by npub. Returns the flow's descriptor. |
| `listen` | yes | Hold a local port. Returns the listener's descriptor. |
| `stats` | no | Debug. Report what the daemon received on a flow. |
| `inject` | no | Debug. Write bytes into a flow from the daemon's side. |
| `arrive` | no | Debug. Dispatch a datagram as though a peer had sent it. |
**There is no close command, no accept command, no reject command, and no
command to enumerate flows.** Closing a descriptor does the first three, and
`fipsctl show native-flows` does the fourth from the control socket.
### connect
| Parameter | Type | Meaning |
| --------- | ---- | ------- |
| `peer` | string | The far end, as an npub. |
| `remote_port` | u16 | The far end's port. Required. |
| `local_port` | u16 | Optional. Absent, `null` and `0` all mean an ephemeral port. |
Reply data: `flow_id`, `local_port`, `remote_port`, `peer` (the daemon's own
re-encode of the key, not an echo of what the caller wrote), `node` (this
node's own npub), and `max_payload`. Carries the flow's descriptor.
Refusals, in the order they are decided, which is what lets a client read an
`EADDRNOTAVAIL` as a tier refusal rather than an exhaustion:
1. `EADDRNOTAVAIL`: the port is in a reserved tier, either port. Decided before
any node state is touched.
2. `EINVAL`: the npub does not decode.
3. `EMFILE`: the socket pair could not be built, or the node holds its maximum
flows.
4. `EADDRINUSE`: the named local port is held, or a flow to that peer between
those two ports already exists.
5. `EADDRNOTAVAIL`: no ephemeral port is free.
6. `ECONNREFUSED`: the node is shutting down.
### listen
One parameter, `local_port` (u16), where absent, `null` and `0` all mean an
ephemeral port. Reply data: `local_port`, `node`, `backlog`. **Carries the
listener's descriptor**, and a reply that carried none is a daemon that is not
this one.
`backlog` reports the depth the daemon will hold for this listener, so an
operator can size a client's reader against it. Nothing on the Rust surface
takes a depth as a parameter or exposes the reported one.
Refusals: `EADDRNOTAVAIL` for a reserved tier or an exhausted ephemeral range;
`EADDRINUSE` when the port is held, whether by another listener or by a
connected flow; `EMFILE` when the socket pair could not be built;
`ECONNREFUSED` when the node is shutting down.
A flow announced on a listener but never taken is discarded after five seconds.
This is not an accept timeout and a client cannot reach it: the window it
bounds is a hand-off between two tasks inside the daemon, not a client's round
trip.
### The debug commands
`stats`, `inject` and `arrive` exist so the daemon's own checks can drive the
receive and dispatch paths without a peer. They are answered only where
`node.native_api.debug_commands` is set, which is off by default and which no
packaged node sets. A node with the key unset refuses all three by name, with
`ECONNREFUSED` and a message naming the key that would admit them, so a client
can tell "this node will not" from "this build cannot".
**Do not build a client on them.**
`stats` takes `flow_id` and reports `flow_id`, `local_port`, `rx_datagrams`,
`rx_bytes` and `closed`. The counters are the daemon's view of what the client
wrote into the descriptor, independent of whether any of it then reached a
peer.
`inject` takes `flow_id`, `data` (a hex string) and an optional `repeat`
(default 1, maximum 64). It writes `repeat` separate datagrams of those bytes
onto the flow's descriptor from the daemon's side. The flow table it names into
is the node's, not the connection's, so a caller that can reach this command
can write into any flow on the node.
`arrive` takes `peer` (npub), `src_port`, `dst_port` and `data` (hex), and
drives the same delivery decision the real receive path uses. It replies `ok`
whenever the node answers at all:
```json
{"status": "ok", "data": {"outcome": "announced", "flow_id": 9}}
```
`outcome` is `delivered`, `announced`, `held`, or `dropped: <cause>`.
`flow_id` is non-null only for `announced` and `held`. The peer's key is decoded
from the npub the caller named, which makes it client-asserted rather than
authenticated. Port tiers are not checked here.
## Compiled-in bounds
Three bounds are not configurable: a command line of 8192 bytes, a per-arrival
send-buffer allowance of 4096 bytes on a listener's pair, and a `repeat` of 64
on the debug `inject`. The daemon reads each descriptor with a 65535-byte
buffer, which is above anything the wire carries.
`pending_per_flow` has a compiled ceiling of 64, checked when the configuration
loads rather than at the first arrival: the whole held batch is written onto a
socket pair whose client half has not been sent yet, so a larger value could
leave a listener's task with a write it cannot complete. Both it and `backlog`
have a floor of 1. A `backlog` of zero admits no flow at all. A
`pending_per_flow` of zero is worse, because the arrival is announced and the
datagram that caused it is then refused, so a peer's opening message vanishes
with no refusal a client or an operator can see.
## See also
- [../how-to/use-the-native-datagram-api.md](../how-to/use-the-native-datagram-api.md)
— enable the API and write a first program
- [../how-to/write-a-native-api-client.md](../how-to/write-a-native-api-client.md)
— the obligations a client in another language carries
- [../how-to/serve-many-peers-on-one-thread.md](../how-to/serve-many-peers-on-one-thread.md)
— one `poll` loop instead of a thread per peer
- [../design/fips-native-api.md](../design/fips-native-api.md)
— what this interface is for, what it is instead of, and what it is not
- [configuration.md](configuration.md#native-datagram-api-nodenative_api)
— every `node.native_api.*` key and its default
- [security.md](security.md#native-datagram-api)
— what `fips` group membership grants once the API is on
- [control-socket.md](control-socket.md)
— the `show_native_flows` response shape
- [cli-fipsctl.md](cli-fipsctl.md)
— `fipsctl show native-flows`
+69 -2
View File
@@ -189,11 +189,74 @@ mutation; rate-limited msg1s never reach the ACL.
| `/etc/fips/peers.allow` | root:root | `0644` | Optional peer allowlist. |
| `/etc/fips/peers.deny` | root:root | `0644` | Optional peer denylist. |
| `/run/fips/control.sock` | root:fips | `0770` | Control socket (members of `fips` group can use `fipsctl`). |
| `/run/fips/` | root:fips | `0750` | Control socket parent directory. |
| `/run/fips/api.sock` | root:fips | `0770` | Native datagram API socket, when `node.native_api.enabled` is set (experimental; absent otherwise). |
| `/run/fips/` | root:fips | `0750` | Socket parent directory. |
Adding a user to the `fips` group grants `fipsctl` access without
requiring root. The daemon `chown`s the control socket and its parent
directory at bind time.
directory at bind time, and does the same for the native API socket when
that is enabled.
## Native Datagram API
**Experimental. Disabled by default** (`node.native_api.enabled`, default
`false`), and built on Linux, FreeBSD and macOS only. It is not a stable API
surface, not a reliability layer, and not the v2 external process API. No
compatibility promise is made about it.
**Any user in the `fips` group can impersonate the node on the mesh.** The
API socket is created at mode `0770` owned by group `fips`, and that is the
entire authorization model. A process that can open it can:
- send datagrams under this node's identity to any peer it names, which
peers authenticate as coming from this node;
- hold any port from 1024 upward and receive mesh traffic addressed to this
node on it, including traffic another local program expected;
- do both without authenticating, without a capability check, and without
any record beyond the daemon's own logs.
Group membership is therefore equivalent to possession of the node's
identity for the purpose of sending on the mesh. **On a node with the native
API enabled, treat membership of the `fips` group exactly as you would treat
`/etc/fips/fips.key`.** Grant it to the accounts that are trusted to speak as
the node and to no others, and review it before enabling the API on a shared
machine.
**The file descriptor carries the grant, not the connection.** A setup call
hands the client a socket descriptor and the connection it was made on is then
closed; the flow or the held port lives until that descriptor is closed. A
descriptor is an ordinary kernel object, so it survives `fork`, survives
`exec` unless the client asked for it close-on-exec when it received it, and
can be handed to another process over `SCM_RIGHTS`. A process holding one can
send as this node on that flow, or receive on that port, without ever opening
the API socket and without being in the `fips` group.
Nothing revokes a descriptor already handed out. Restarting the daemon closes
its own halves and ends every flow and listener at once, and that is the only
revocation there is.
Two consequences follow for `fipsctl` access. First, the `fips` group is
already the control-socket group, so enabling the native API silently
upgrades every existing `fipsctl` user from "can read node state and manage
peers" to "can send as the node". Second, an operator who wants the two
audiences separated must not enable the API on a node whose `fips` group has
been handed out for monitoring.
`node.native_api.debug_commands` (default `false`) is a second, independent
gate. It admits three commands (`inject`, `stats`, `arrive`) that exist for
the test harness: `arrive` makes the daemon dispatch a datagram as though a
peer had sent it, reaching any listener on this node under any peer identity
the caller names. Leave it off outside a test harness; a packaged node does
not enable it.
The socket is local only. It is not reachable over the network, and nothing
about it changes the mesh's own authentication: a peer still verifies the
node's signature, which is precisely why a local caller that can send through
this socket is indistinguishable from the node itself.
See [configuration.md](configuration.md#native-datagram-api-nodenative_api)
for the key list and
[../how-to/use-the-native-datagram-api.md](../how-to/use-the-native-datagram-api.md)
for the client.
## Threat-Resistance Matrix
@@ -214,6 +277,7 @@ the FMP design document:
| Sybil identities | Discretionary peering + handshake rate limiting + optional peer ACL |
| Eclipse attack | Diverse peering across independent operators and transports |
| Unauthorized peer admission | Optional `peers.allow` allowlist consulted before handshake |
| Local impersonation via the native datagram API | API disabled by default; when enabled, `fips` group membership is the only gate and must be treated as key access |
See [../design/fips-mesh-layer.md](../design/fips-mesh-layer.md) for
the unauthenticated-attack-surface analysis (only handshake msg1 is
@@ -251,3 +315,6 @@ way to restrict inbound traffic on `fips0`. See
— operator activation and drop-in recipes
- [configuration.md](configuration.md) — full `node.rekey.*`,
`node.rate_limit.*` parameter tables
- [../how-to/use-the-native-datagram-api.md](../how-to/use-the-native-datagram-api.md)
— enabling the experimental native datagram API, and what group
membership grants once it is on
+8 -2
View File
@@ -38,13 +38,19 @@ policy, and an understanding of both deployment modes — overlay
on top of existing IP, and ground-up where the mesh is the
network.
There is also a side trip you can take any time after tutorial 1:
There are also two side trips you can take:
- [ipv6-adapter-walkthrough.md](ipv6-adapter-walkthrough.md) —
trace one `ssh` from DNS query through session setup to the
far-side TUN, using `fipstop` and `fipsctl` to watch each step.
Optional, but if you like seeing how the pieces fit together,
this is the doc that shows you.
this is the doc that shows you. Take it any time after tutorial 1.
- [native-api-walkthrough.md](native-api-walkthrough.md) — write a
program against the experimental native datagram API, addressing a
peer by public key and port with no IPv6 emulation and no TUN. Runs
two throwaway nodes on one machine, so it needs no mesh and no root,
and you can take it without doing the tutorials first.
## Advanced
+343
View File
@@ -0,0 +1,343 @@
# Native Datagram API Walkthrough
A side trip. You will run two FIPS nodes on one machine, write a listening
program and a connecting program against the native datagram API, and watch one
datagram cross between them. Then you will look at the flow from the outside
with `fipsctl` while it is still open.
Nothing here touches the public mesh, and nothing needs root. Both nodes run
with no TUN device and no DNS, peered directly over loopback UDP, so the whole
session lives in one scratch directory you delete at the end.
**This is not part of the numbered progression.** Take it any time. It assumes
you can build the daemon from source and can read Rust; it does not assume you
have worked through the tutorials.
**The API is experimental.** Names, fields and the command set may change
without a deprecation cycle. It is Linux, FreeBSD and macOS only.
## What you will end up with
- Two nodes, each with its own identity, control socket and API socket.
- A listening program that holds port 4600 and echoes one datagram per flow.
- A connecting program that opens a flow to the other node's public key and
gets its datagram back.
- A reading of `fipsctl show native-flows` taken while the flow is open.
## Step 1: Build the daemon and its tools
From a checkout of the FIPS source:
```sh
cargo build --release --bins
```
That gives you `target/release/fips` and `target/release/fipsctl`. Put them on
your path for the rest of this walkthrough:
```sh
export PATH="$PWD/target/release:$PATH"
```
## Step 2: Make two identities
`keygen -s` prints a keypair to stdout and writes nothing:
```sh
fipsctl keygen -s
```
```text
nsec1...
npub1...
```
Run it twice and keep both pairs. Call them A and B. You need each node's
`nsec` for its own config, and each node's `npub` for the *other* node's peer
entry.
```sh
mkdir -p ~/napi-lab/a ~/napi-lab/b
cd ~/napi-lab
```
## Step 3: Write the two configs
Node A, at `~/napi-lab/a/fips.yaml`. Substitute A's `nsec` and B's `npub`:
```yaml
node:
identity:
nsec: "<A's nsec>"
control:
socket_path: "/home/YOU/napi-lab/a/control.sock"
native_api:
enabled: true
socket_path: "/home/YOU/napi-lab/a/api.sock"
tun:
enabled: false
dns:
enabled: false
transports:
udp:
bind_addr: "127.0.0.1:2121"
mtu: 1472
peers:
- npub: "<B's npub>"
alias: "node-b"
addresses:
- transport: udp
addr: "127.0.0.1:2122"
```
Node B, at `~/napi-lab/b/fips.yaml`, is the mirror image: B's `nsec`, A's
`npub`, its own sockets under `b/`, `bind_addr` on `2122`, and its peer address
pointing at `2121`.
> **Use absolute paths.** The daemon does not resolve a socket path relative to
> the config file. Putting an `nsec` in a config is fine for a throwaway lab
> node like this one; for anything you keep, use
> [../how-to/persistent-identity.md](../how-to/persistent-identity.md) instead.
Disabling TUN and DNS is what lets both nodes run as your own user. A node with
a TUN device needs `CAP_NET_ADMIN`, and this walkthrough does not need one:
the native API is the path that does not go through the IPv6 adapter.
## Step 4: Start both nodes
In two terminals:
```sh
fips --config ~/napi-lab/a/fips.yaml
```
```sh
fips --config ~/napi-lab/b/fips.yaml
```
Each should log that it bound its API socket:
```text
Native API socket listening on /home/YOU/napi-lab/a/api.sock
```
In a third terminal, confirm the two found each other:
```sh
fipsctl -s ~/napi-lab/a/control.sock show peers
```
Wait for B to appear with a session. The link forms over loopback UDP and
usually takes a second or two. **Wait for it before going on**: a `connect` on
a flow contacts no peer, so it will succeed whether or not the link is up, and
the datagram would simply be held and then dropped.
## Step 5: Write the listening program
Make a crate next to the lab directory:
```sh
cargo new --bin napi-listen
cd napi-listen
```
Point it at your FIPS checkout in `Cargo.toml`:
```toml
[dependencies]
fips = { path = "/path/to/your/fips/checkout" }
```
`src/main.rs`:
```rust
//! Hold a port and echo one datagram per flow.
use fips::native::client::{FipsListener, FipsStream};
use std::env;
use std::error::Error;
use std::path::Path;
use std::thread;
/// Return one datagram to where it came from, then release the flow.
fn serve(flow: FipsStream) {
// Sized at the flow's own limit, so no datagram it can carry is
// truncated on the way in and echoed short.
let mut buf = vec![0u8; flow.max_payload()];
match flow.recv(&mut buf) {
Ok(len) => {
let _ = flow.send(&buf[..len]);
println!("returned {len} bytes to {}", flow.peer_addr());
}
Err(error) => eprintln!("receiving: {error}"),
}
// Returning drops the flow, which closes its descriptor. That is what
// releases the flow at the daemon; there is no close call to make.
}
fn main() -> Result<(), Box<dyn Error>> {
let socket = env::args().nth(1).ok_or("usage: napi-listen <api-socket>")?;
let listener = FipsListener::bind_at(Path::new(&socket), 4600)?;
println!("holding {}", listener.local_addr());
for arrival in listener.incoming() {
// Detached rather than joined: the accept loop must not wait on one
// peer, and the thread owns everything it touches.
match arrival {
Ok(flow) => drop(thread::spawn(move || serve(flow))),
Err(error) => eprintln!("accepting: {error}"),
}
}
Ok(())
}
```
Run it against node B:
```sh
cargo run -- ~/napi-lab/b/api.sock
```
```text
holding npub1...:4600
```
**Note what `serve` does not do.** It does not loop reading until the flow
closes. The v1 wire carries no half-close, so nothing peer-driven would ever
end that loop; it would hold a thread and a flow slot per peer until the
process died. One exchange per flow is the program's own decision, and making
it is mandatory. See
[../how-to/use-the-native-datagram-api.md](../how-to/use-the-native-datagram-api.md#four-things-that-will-bite-you).
**Note also what `incoming()` does not do.** It never returns `None`, and a
failed accept arrives as an `Err` item rather than ending the iteration. Writing
`let flow = arrival?;` here would exit the loop on the first transient error,
which is a different shape from `TcpListener` habits.
## Step 6: Write the connecting program
```sh
cd ..
cargo new --bin napi-connect
cd napi-connect
```
Same dependency line. `src/main.rs`:
```rust
//! Open a flow to a peer, exchange one datagram, and exit.
use fips::native::client::FipsStream;
use std::env;
use std::error::Error;
use std::io;
use std::path::Path;
use std::time::Duration;
/// How long to wait for the peer's answer before giving up on it.
const REPLY: Duration = Duration::from_secs(10);
fn main() -> Result<(), Box<dyn Error>> {
let mut args = env::args().skip(1);
let (Some(socket), Some(peer)) = (args.next(), args.next()) else {
return Err("usage: napi-connect <api-socket> <peer-npub>".into());
};
// One setup call, and it contacts no peer: the daemon registers the flow
// locally and hands back the descriptor it rides on. Success here says
// nothing about the peer existing, being reachable, or listening.
let flow = FipsStream::connect_at(Path::new(&socket), 0, (peer, 4600))?;
println!("{} -> {}", flow.local_addr(), flow.peer_addr());
// Before the first recv and not after it, because the peer may never
// answer at all and the deadline is what makes that a failure rather
// than a hang.
flow.set_read_timeout(Some(REPLY))?;
flow.send(b"hello")?;
let mut buf = vec![0u8; flow.max_payload()];
match flow.recv(&mut buf) {
Ok(len) => println!("{}", String::from_utf8_lossy(&buf[..len])),
Err(error) if error.kind() == io::ErrorKind::WouldBlock => {
return Err(format!("no answer from {} in {REPLY:?}", flow.peer_addr()).into());
}
Err(error) => return Err(error.into()),
}
Ok(())
}
```
Run it against node A, naming node B's npub:
```sh
cargo run -- ~/napi-lab/a/api.sock <B's npub>
```
```text
npub1...:49152 -> npub1...:4600
hello
```
The listener's terminal reports the other half:
```text
returned 5 bytes to npub1...:49152
```
That datagram went from your connecting program, into node A over a Unix
socket, across loopback UDP inside an encrypted FSP session, into node B, and
out to your listening program on another Unix socket. No IPv6 address and no
TUN device was involved anywhere in it.
## Step 7: Watch a flow from the outside
The exchange above is over in milliseconds. To look at a live flow, make the
connector hold one open: add a `std::thread::sleep(Duration::from_secs(60));`
before the final `Ok(())` and run it again.
While it sleeps:
```sh
fipsctl -s ~/napi-lab/b/control.sock show native-flows
```
You get every flow node B holds, with its ports, its queue depth and its age,
plus every bound listener and its backlog. The counters are in:
```sh
fipsctl -s ~/napi-lab/b/control.sock stats metrics
```
under `native`, where the `drop_*` fields separate a datagram refused for
having no listening port from one dropped because a client was not reading fast
enough. Those counters are the only way to see a drop: **nothing on the API
surface reports one to your program.**
## Step 8: Clean up
Stop both daemons with Ctrl-C, then:
```sh
rm -rf ~/napi-lab
```
The identities were only ever in those config files, so removing the directory
removes them. Nothing was written outside it and nothing was published to any
relay.
## Where to go next
- [../how-to/use-the-native-datagram-api.md](../how-to/use-the-native-datagram-api.md)
— the same ground as a recipe, including enabling the API on a real node and
the security posture that grants
- [../reference/native-api.md](../reference/native-api.md)
— every type and method, the errno table, the ceilings, and what happens to
data that disappears
- [../how-to/write-a-native-api-client.md](../how-to/write-a-native-api-client.md)
— doing all of this from C, Python or Go, where there is no client library
and the obligations become yours
+126
View File
@@ -0,0 +1,126 @@
//! An echo server on the native datagram API, and the reference client for it.
//!
//! Run it beside a daemon whose native API socket is the first argument; it
//! holds the port given as the second and returns every datagram sent to that
//! port to whoever sent it.
//!
//! ```text
//! native-echo /run/fips/api.sock 4600
//! ```
//!
//! **One exchange, one flow.** A served flow takes a single datagram, sends it
//! back and closes. A datagram API carries no far-end close, so a server that
//! kept reading would hold the flow until its own process went away and would
//! be waiting for a signal that never comes; the accept loop takes the next
//! arrival instead.
//!
//! **It imports `fips::native::client` and nothing else from the crate.** That
//! is the point of the example as much as the echoing is: needing `serde_json`,
//! `libc`, or any knowledge of the line protocol here would mean the client
//! module had failed to hide something, and the fix would belong there.
//!
//! **Platform.** The client module is built on Linux and FreeBSD only, because
//! macOS has no `AF_UNIX` `SOCK_SEQPACKET` and Windows no `SCM_RIGHTS`. `main`
//! is gated to match rather than the file being Linux-only by accident:
//! `cargo clippy --all-targets` and `cargo nextest run` both compile example
//! targets, and both run on macOS, so an ungated file would break those runs
//! there instead of this program refusing to start.
#[cfg(any(target_os = "linux", target_os = "freebsd"))]
mod echo {
use fips::native::client::{FipsListener, FipsStream};
use std::env;
use std::io::{self, Write};
use std::path::Path;
use std::process::ExitCode;
use std::thread;
/// Hold the port named on the command line and echo what arrives on it.
///
/// Returns only on failure: a healthy server has no reason to stop, and the
/// caller kills it.
pub fn run() -> ExitCode {
let mut args = env::args().skip(1);
let (Some(socket), Some(port)) = (args.next(), args.next()) else {
eprintln!("usage: native-echo <socket-path> <local-port>");
return ExitCode::FAILURE;
};
let port: u16 = match port.parse() {
Ok(port) => port,
Err(error) => {
eprintln!("native-echo: {port:?} is not a port: {error}");
return ExitCode::FAILURE;
}
};
// "cannot hold" rather than "holding": a caller waits for the success
// line below by substring, and a failure line containing it would
// satisfy that wait and hide the reason.
let listener = match FipsListener::bind_at(Path::new(&socket), port) {
Ok(listener) => listener,
Err(error) => {
eprintln!("native-echo: cannot hold port {port} on {socket}: {error}");
return ExitCode::FAILURE;
}
};
// The port the daemon actually held, which is what a caller asking for
// an ephemeral one needs. A caller waits on this line before it sends
// anything, so it is flushed rather than left to the line buffer.
println!("native-echo: holding port {}", listener.local_addr().port());
let _ = io::stdout().flush();
for arrival in listener.incoming() {
match arrival {
// Detached rather than joined: the accept loop must not wait on
// one peer, and the thread owns everything it touches.
Ok(flow) => drop(thread::spawn(move || serve(flow))),
Err(error) => {
eprintln!("native-echo: accepting on port {port}: {error}");
return ExitCode::FAILURE;
}
}
}
// `incoming` has no end: a listener has no last flow. Reaching here at
// all would mean the iterator broke its contract.
ExitCode::FAILURE
}
/// Return one datagram to where it came from, then release the flow.
///
/// Dropping the flow closes its descriptor, which is what tells the daemon
/// the flow is finished; there is no close command to send.
fn serve(flow: FipsStream) {
// Sized at the flow's own limit, so no datagram the flow can carry is
// truncated on the way in and echoed short.
let mut buf = vec![0u8; flow.max_payload()];
let len = match flow.recv(&mut buf) {
Ok(len) => len,
Err(error) => {
eprintln!("native-echo: receiving from {}: {error}", flow.peer_addr());
return;
}
};
match flow.send(&buf[..len]) {
Ok(()) => println!("native-echo: returned {len} bytes to {}", flow.peer_addr()),
Err(error) => eprintln!("native-echo: returning to {}: {error}", flow.peer_addr()),
}
let _ = io::stdout().flush();
}
}
/// Serve until killed.
#[cfg(any(target_os = "linux", target_os = "freebsd"))]
fn main() -> std::process::ExitCode {
echo::run()
}
/// Refuse cleanly where the native API client is not built.
#[cfg(not(any(target_os = "linux", target_os = "freebsd")))]
fn main() -> std::process::ExitCode {
eprintln!(
"native-echo needs the native datagram API client, which is built on \
Linux and FreeBSD only"
);
std::process::ExitCode::FAILURE
}
+661
View File
@@ -0,0 +1,661 @@
//! Every public item of the native datagram API client, asserted against a
//! real daemon.
//!
//! ```text
//! native-surface walk /run/fips/api.sock npub1…
//! ```
//!
//! **This is an assertion harness, not a program shape to copy.** It opens
//! flows nobody answers, asks for deadlines only so it can watch them expire,
//! and reaches for `poll(2)` on a descriptor the surface hands out. A program
//! that wanted to do something useful with this API would look like
//! `native-echo`, which is the example to read first.
//!
//! **It reaches past `fips::native::client` for `libc`, and that is a
//! departure.** `native-echo` states as a design property that needing `libc`
//! would mean the client module had failed to hide something. Here the
//! descriptor is the thing under test: `AsRawFd` and `AsFd` exist so a caller
//! can put a flow or a listener in its own event loop, `poll(2)` has no `std`
//! spelling, and asserting that the number really is a pollable descriptor is
//! the whole point of those items. Every `libc` use is inside the platform-gated
//! module below, because `libc` is a `cfg(unix)` dependency and the Windows leg
//! of CI compiles examples.
//!
//! **Every assertion goes through [`surface::step`], and there is no other way
//! to record one.** The count in the terminal line is that recorder's counter
//! rather than a number written into the format string, so a caller comparing
//! it against what it expected catches a block that stopped running as well as
//! a block that failed.
//!
//! **Platform.** The client module is built on Linux and FreeBSD only, so
//! `main` is gated to match, for the reason `native-echo` gives: example
//! targets are compiled by `cargo clippy --all-targets` on every platform in
//! the build matrix, and an ungated file would break those runs rather than
//! this program refusing to start.
#[cfg(any(target_os = "linux", target_os = "freebsd"))]
mod surface {
use fips::native::client::{FipsAddr, FipsListener, FipsStream, SOCKET, XOnlyPublicKey};
use std::env;
use std::fmt;
use std::io::{self, Write};
use std::os::fd::{AsFd, AsRawFd, RawFd};
use std::path::Path;
use std::process::ExitCode;
use std::str::FromStr;
use std::sync::Mutex;
use std::sync::atomic::{AtomicUsize, Ordering};
use std::thread;
use std::time::{Duration, Instant};
/// How long the whole run may take before the watchdog calls it wedged.
///
/// Several of the defects these assertions exist to catch fail by blocking
/// for ever rather than by returning something wrong, and a container that
/// never exits reaches no accounting in the driver that started it.
const WATCHDOG: Duration = Duration::from_secs(30);
/// The read deadline the expiry assertion arms.
///
/// Long enough that scheduling noise cannot make the wait look absent, short
/// enough that waiting it out twice costs nothing.
const DEADLINE: Duration = Duration::from_millis(200);
/// The largest datagram a flow carries on the harness node.
///
/// A literal, because the number is the assertion: it is what the daemon
/// computes from `testing/native-api/node.yaml`'s 1472-byte UDP MTU, and the
/// harness's line-protocol checks assert the same 1362 from the other side.
/// A walk run against a node configured differently is expected to fail
/// here, which is why the port band and the socket path are checked too.
const PAYLOAD: usize = 1362;
/// The lowest port the daemon's ephemeral allocator ever hands out.
///
/// A floor rather than a value: the allocator is a forward-only cursor
/// shared with every check that ran before this one, so the exact port
/// depends on run order and only the range is a property of the surface.
const EPHEMERAL: u16 = 49152;
/// The port the walk asks a listener to hold by name.
///
/// From 4800-4809, a band no other check in `testing/native-api/` uses.
const HELD: u16 = 4800;
/// The local port the walk asks a flow to be opened from by name.
const NAMED: u16 = 4801;
/// How many assertions have held so far.
static PASSED: AtomicUsize = AtomicUsize::new(0);
/// The assertion currently running, so a wedge can say which one wedged.
static IN_FLIGHT: Mutex<&'static str> = Mutex::new("start-up");
/// Run one assertion, counting it when it holds and ending the run when not.
///
/// The only way to record an assertion, which is what makes [`PASSED`] a
/// measurement of what ran rather than a number kept by hand.
///
/// It exits rather than returning an error because the walk is a sequence:
/// a flow that could not be opened has nothing to assert about, and a run
/// that carried on would bury the first failure under the noise of every
/// assertion downstream of it.
pub fn step<T>(name: &'static str, body: impl FnOnce() -> Result<T, String>) -> T {
// The guard is dropped at the end of this statement rather than held
// across `body`, or the watchdog could not read the name it needs.
*IN_FLIGHT
.lock()
.unwrap_or_else(|poison| poison.into_inner()) = name;
match body() {
Ok(value) => {
PASSED.fetch_add(1, Ordering::SeqCst);
value
}
Err(detail) => {
eprintln!("native-surface: FAILED at {name}: {detail}");
let _ = io::stderr().flush();
std::process::exit(1);
}
}
}
/// Arm a thread that ends the run, by name, if it stops making progress.
///
/// A deadline that never reached the descriptor and a non-blocking mode that
/// was never set both fail as a `recv` that never returns. Without this the
/// container would run until something outside it lost patience, and the
/// evidence of which assertion was in flight would be gone.
fn watchdog() {
drop(thread::spawn(|| {
thread::sleep(WATCHDOG);
let name = *IN_FLIGHT
.lock()
.unwrap_or_else(|poison| poison.into_inner());
eprintln!(
"native-surface: FAILED at {name}: watchdog after {}s",
WATCHDOG.as_secs()
);
let _ = io::stderr().flush();
std::process::exit(1);
}));
}
/// Compare what a call reported against what the surface promises.
fn same<T: PartialEq + fmt::Debug>(what: &str, got: T, want: T) -> Result<(), String> {
if got == want {
return Ok(());
}
Err(format!("{what} is {got:?}, expected {want:?}"))
}
/// Assert a call was refused with a particular errno.
fn errno<T: fmt::Debug>(what: &str, got: io::Result<T>, want: i32) -> Result<(), String> {
match got {
Ok(value) => Err(format!(
"{what} succeeded with {value:?}, expected errno {want}"
)),
Err(error) if error.raw_os_error() == Some(want) => Ok(()),
Err(error) => Err(format!("{what} failed with {error}, expected errno {want}")),
}
}
/// Assert a call refused rather than waiting.
///
/// By kind rather than by errno: `WouldBlock` is what the surface documents,
/// and it is the one answer both an expired deadline and a non-blocking
/// descriptor give, which is why the two are told apart here by how long the
/// call took rather than by what it returned.
fn blocked<T: fmt::Debug>(what: &str, got: io::Result<T>) -> Result<(), String> {
match got {
Ok(value) => Err(format!(
"{what} succeeded with {value:?}, expected it to refuse to wait"
)),
Err(error) if error.kind() == io::ErrorKind::WouldBlock => Ok(()),
Err(error) => Err(format!(
"{what} failed with {error}, expected it to refuse to wait"
)),
}
}
/// Whether `poll(2)` says a descriptor has something to read, right now.
///
/// The one place this file reaches past the client module, and the reason it
/// is allowed to: what `AsRawFd` and `AsFd` promise is that the number names
/// a socket an event loop can wait on, and nothing in `std` asks that
/// question of a bare descriptor.
fn readable(fd: RawFd) -> Result<bool, String> {
let mut waiting = libc::pollfd {
fd,
events: libc::POLLIN,
revents: 0,
};
// SAFETY: the pointer and count describe one `pollfd` this frame owns,
// and the descriptor belongs to a stream or listener still alive here.
let rc = unsafe { libc::poll(std::ptr::from_mut(&mut waiting), 1, 0) };
if rc < 0 {
return Err(format!(
"poll on descriptor {fd}: {}",
io::Error::last_os_error()
));
}
// Without this the mechanism failing and the assertion holding are the
// same value: the only poll assertion here asserts a NEGATIVE, and a
// descriptor poll(2) rejects outright comes back rc=1 with POLLNVAL and
// no POLLIN, which would read as a quiet "nothing to read" pass.
let broken = waiting.revents & (libc::POLLNVAL | libc::POLLERR);
if broken != 0 {
return Err(format!(
"poll rejected descriptor {fd}, revents {broken:#x}, so it names no open socket"
));
}
Ok(waiting.revents & libc::POLLIN != 0)
}
/// Assert what a freshly opened flow says about the address it was given.
///
/// Through the accessors and `Display` rather than by comparing the whole
/// address, because those are themselves items under test.
fn opened(flow: &FipsStream, key: XOnlyPublicKey, port: u16) -> Result<(), String> {
let want = FipsAddr::new(key, port);
same("the peer key", flow.peer_addr().key(), key)?;
same("the peer port", flow.peer_addr().port(), port)?;
same(
"the peer address written out",
flow.peer_addr().to_string(),
want.to_string(),
)
}
/// Assert the whole surface against the daemon on `sock`.
///
/// `peer` is an npub nothing answers on, which is what most of these
/// assertions need: a native `connect` is a local registration, so a flow to
/// a peer that does not exist is a real flow with a real descriptor, and
/// nothing arriving on it is what makes a deadline observable.
fn walk(sock: &Path, peer: &str, key: XOnlyPublicKey) {
// ── Setup entry points ────────────────────────────────────────────
let held = step("bind_at holds the port it was told to hold", || {
let listener =
FipsListener::bind_at(sock, HELD).map_err(|e| format!("bind_at({HELD}): {e}"))?;
same("the port held", listener.local_addr().port(), HELD)?;
Ok(listener)
});
let ephemeral = step(
"bind with no socket path resolves SOCKET and takes an ephemeral port",
|| {
let listener = FipsListener::bind(0).map_err(|e| format!("bind(0): {e}"))?;
let port = listener.local_addr().port();
if port < EPHEMERAL {
return Err(format!(
"the port held is {port}, expected one from {EPHEMERAL} up"
));
}
Ok(listener)
},
);
let flow = step(
"connect_at opens a flow to the peer and port it was given",
|| {
let port = 4809;
let flow = FipsStream::connect_at(sock, 0, (key, port))
.map_err(|e| format!("connect_at(_, 0, (key, {port})): {e}"))?;
opened(&flow, key, port)?;
let local = flow.local_addr().port();
if local < EPHEMERAL {
return Err(format!(
"the flow's local port is {local}, expected one from {EPHEMERAL} up"
));
}
Ok(flow)
},
);
step(
"the node's key is the same on a listener and a flow, and is not the peer's",
|| {
same(
"the node key a flow reports",
flow.local_addr().key(),
ephemeral.local_addr().key(),
)?;
if flow.peer_addr().key() == flow.local_addr().key() {
return Err("a flow's peer key and node key are the same value".to_string());
}
Ok(())
},
);
// ── Addressing: one connect per ToFipsAddr impl ───────────────────
//
// Eight impls, eight calls, and the mapping is the audit: `grep -n
// 'impl.*ToFipsAddr for' src/native/client/mod.rs` returns eight lines.
// A ninth would have to appear here and in the harness's expected count
// before either could go green again, which is the intended friction.
let addr = FipsAddr::new(key, 4807);
step("connect takes an address by value", || {
let flow = FipsStream::connect(addr).map_err(|e| format!("connect(FipsAddr): {e}"))?;
// connect delegates to connect_at with a local port of 0, so a
// defect that passed some other port instead shows up here as a
// local port below the ephemeral floor.
let local = flow.local_addr().port();
if local < EPHEMERAL {
return Err(format!(
"the flow's local port is {local}, expected one from {EPHEMERAL} up"
));
}
opened(&flow, key, addr.port())
});
step("connect takes a key and a port as a tuple", || {
let port = 4802;
let flow = FipsStream::connect((key, port))
.map_err(|e| format!("connect((XOnlyPublicKey, {port})): {e}"))?;
opened(&flow, key, port)
});
step(
"connect takes a serialized key and a port as a tuple",
|| {
let port = 4803;
let flow = FipsStream::connect((key.serialize(), port))
.map_err(|e| format!("connect(([u8; 32], {port})): {e}"))?;
opened(&flow, key, port)
},
);
step("connect takes an npub slice and a port as a tuple", || {
let port = 4804;
let flow = FipsStream::connect((peer, port))
.map_err(|e| format!("connect((&str, {port})): {e}"))?;
opened(&flow, key, port)
});
step("connect takes an owned npub and a port as a tuple", || {
let port = 4805;
let flow = FipsStream::connect((peer.to_string(), port))
.map_err(|e| format!("connect((String, {port})): {e}"))?;
opened(&flow, key, port)
});
step(
"connect takes a whole address as one slice, parsed by FromStr",
|| {
// The only route to `impl ToFipsAddr for str`: a `str` is
// unsized, so the argument is a `&str` and the blanket impl for
// `&T` is what dispatches to it.
let text = format!("{peer}:4806");
let want =
FipsAddr::from_str(&text).map_err(|e| format!("parsing {text:?}: {e}"))?;
let flow = FipsStream::connect(text.as_str())
.map_err(|e| format!("connect({text:?} as &str): {e}"))?;
opened(&flow, want.key(), want.port())
},
);
step("connect takes a whole address as one owned string", || {
let text = format!("{peer}:4808");
let want = FipsAddr::from_str(&text).map_err(|e| format!("parsing {text:?}: {e}"))?;
let flow =
FipsStream::connect(text).map_err(|e| format!("connect(a String address): {e}"))?;
opened(&flow, want.key(), want.port())
});
// The borrow is the assertion, not an accident: `&addr` is the only
// thing that reaches the blanket `impl ToFipsAddr for &T`, and passing
// `addr` by value as clippy suggests would exercise `impl for FipsAddr`
// a second time and leave the blanket impl untested with this `T`.
#[allow(clippy::needless_borrows_for_generic_args)]
step(
"connect_from names the local port, taking the address by reference",
|| {
// The blanket impl again, with a different `T`, which is what
// makes this a separate exercise rather than a repeat.
let flow = FipsStream::connect_from(NAMED, &addr)
.map_err(|e| format!("connect_from({NAMED}, &FipsAddr): {e}"))?;
same("the local port asked for", flow.local_addr().port(), NAMED)?;
opened(&flow, key, addr.port())
},
);
// ── Deadlines ─────────────────────────────────────────────────────
step(
"read_timeout reads back the deadline set_read_timeout set",
|| {
flow.set_read_timeout(Some(DEADLINE))
.map_err(|e| format!("set_read_timeout(Some({DEADLINE:?})): {e}"))?;
let got = flow
.read_timeout()
.map_err(|e| format!("read_timeout: {e}"))?;
same("the read deadline", got, Some(DEADLINE))
},
);
step("recv gives up once the read deadline expires", || {
let mut buf = [0u8; 64];
let started = Instant::now();
let outcome = flow.recv(&mut buf);
let waited = started.elapsed();
blocked("recv on a flow no peer answers", outcome)?;
// Both bounds matter: too soon means the deadline never reached the
// descriptor and the answer came from somewhere else, and too late
// means it reached a different option than the one that was set.
if waited < DEADLINE - Duration::from_millis(50) {
return Err(format!(
"recv gave up after {waited:?}, too soon to have waited the {DEADLINE:?} deadline"
));
}
if waited > Duration::from_secs(5) {
return Err(format!(
"recv waited {waited:?}, far past the {DEADLINE:?} deadline"
));
}
Ok(())
});
step("setting the read deadline to None clears it", || {
flow.set_read_timeout(None)
.map_err(|e| format!("set_read_timeout(None): {e}"))?;
let got = flow
.read_timeout()
.map_err(|e| format!("read_timeout: {e}"))?;
same("the read deadline", got, None)
});
step("set_read_timeout refuses a zero duration", || {
errno(
"set_read_timeout(Some(0))",
flow.set_read_timeout(Some(Duration::ZERO)),
libc::EINVAL,
)
});
step(
"write_timeout reads back the deadline set_write_timeout set",
|| {
flow.set_write_timeout(Some(DEADLINE))
.map_err(|e| format!("set_write_timeout(Some({DEADLINE:?})): {e}"))?;
let got = flow
.write_timeout()
.map_err(|e| format!("write_timeout: {e}"))?;
same("the write deadline", got, Some(DEADLINE))
},
);
step("setting the write deadline to None clears it", || {
flow.set_write_timeout(None)
.map_err(|e| format!("set_write_timeout(None): {e}"))?;
let got = flow
.write_timeout()
.map_err(|e| format!("write_timeout: {e}"))?;
same("the write deadline", got, None)
});
step("set_write_timeout refuses a zero duration", || {
errno(
"set_write_timeout(Some(0))",
flow.set_write_timeout(Some(Duration::ZERO)),
libc::EINVAL,
)
});
// ── Non-blocking ──────────────────────────────────────────────────
step("a non-blocking flow refuses to wait in recv", || {
flow.set_nonblocking(true)
.map_err(|e| format!("set_nonblocking(true) on a flow: {e}"))?;
let mut buf = [0u8; 64];
let started = Instant::now();
let outcome = flow.recv(&mut buf);
let waited = started.elapsed();
blocked("recv on a non-blocking flow", outcome)?;
// The read deadline was cleared two assertions ago, so a
// set_nonblocking that did nothing would park here for ever and the
// watchdog would name this step. This bound therefore carries only
// the narrow shape where set_nonblocking armed a short deadline
// instead of the descriptor's mode. It is deliberately loose: it is
// still an order of magnitude under the cleared-deadline case, and
// tightening it buys no discrimination while inviting a flake when
// the runner is loaded.
if waited > Duration::from_millis(250) {
return Err(format!(
"recv on a non-blocking flow took {waited:?}, which is a wait rather than a refusal"
));
}
Ok(())
});
step("a non-blocking listener refuses to wait in accept", || {
held.set_nonblocking(true)
.map_err(|e| format!("set_nonblocking(true) on a listener: {e}"))?;
blocked("accept on a non-blocking listener", held.accept())
});
step(
"a failed accept is an incoming item rather than the end of the iteration",
|| match held.incoming().next() {
None => Err("incoming ended, and a listener has no last flow".to_string()),
Some(item) => blocked("the first incoming item", item),
},
);
// ── Descriptors ───────────────────────────────────────────────────
step(
"the flow's borrowed descriptor is the number its raw one gives",
|| {
same(
"the flow's descriptor",
flow.as_fd().as_raw_fd(),
flow.as_raw_fd(),
)
},
);
step(
"the listener's borrowed descriptor is the number its raw one gives",
|| {
same(
"the listener's descriptor",
held.as_fd().as_raw_fd(),
held.as_raw_fd(),
)
},
);
step("a flow and a listener hold different descriptors", || {
let (one, other) = (flow.as_raw_fd(), held.as_raw_fd());
if one == other {
return Err(format!("both report descriptor {one}"));
}
Ok(())
});
step(
"the flow's descriptor can be duplicated, so it names an open file",
|| {
flow.as_fd()
.try_clone_to_owned()
.map(drop)
.map_err(|e| format!("duplicating the flow's descriptor: {e}"))
},
);
step(
"the listener's descriptor can be duplicated, so it names an open file",
|| {
held.as_fd()
.try_clone_to_owned()
.map(drop)
.map_err(|e| format!("duplicating the listener's descriptor: {e}"))
},
);
step(
"poll reports neither the flow nor the listener readable while nothing has arrived",
|| {
if readable(flow.as_raw_fd())? {
return Err("the flow is readable and no peer has sent anything".to_string());
}
if readable(held.as_raw_fd())? {
return Err("the listener is readable and no flow has arrived".to_string());
}
Ok(())
},
);
// ── Limits ────────────────────────────────────────────────────────
step(
"max_payload is what the daemon computed for this transport",
|| {
same(
"the largest datagram this flow carries",
flow.max_payload(),
PAYLOAD,
)
},
);
step(
"a datagram of exactly max_payload bytes is accepted",
|| {
let datagram = vec![0x5a; flow.max_payload()];
flow.send(&datagram)
.map_err(|e| format!("send of {} bytes: {e}", datagram.len()))
},
);
step(
"a datagram one byte past max_payload is refused with EMSGSIZE",
|| {
let datagram = vec![0x5a; flow.max_payload() + 1];
errno(
"send of one byte past the limit",
flow.send(&datagram),
libc::EMSGSIZE,
)
},
);
}
/// Walk the surface, and print how many assertions held.
pub fn run() -> ExitCode {
let mut args = env::args().skip(1);
let (Some(mode), Some(sock), Some(peer)) = (args.next(), args.next(), args.next()) else {
eprintln!("usage: native-surface walk <socket-path> <peer-npub>");
return ExitCode::FAILURE;
};
if mode != "walk" {
eprintln!("native-surface: {mode:?} is not a mode; the modes are: walk");
return ExitCode::FAILURE;
}
// The walk is given a path because `connect_at` and `bind_at` are items
// in their own right and need one. The forms that take no path resolve
// SOCKET, which is the same daemon only when the caller mounted it
// there; a mismatch would fail those calls with ENOENT and read as a
// defect in the surface rather than in the invocation.
if sock != SOCKET {
eprintln!(
"native-surface: the no-path calls resolve {SOCKET}, so the walk has to be given \
that path, not {sock}"
);
return ExitCode::FAILURE;
}
let node = match FipsAddr::from_str(&format!("{peer}:0")) {
Ok(node) => node,
Err(error) => {
eprintln!("native-surface: {peer:?} is not an npub: {error}");
return ExitCode::FAILURE;
}
};
let key: XOnlyPublicKey = node.key();
watchdog();
walk(Path::new(&sock), &peer, key);
// The count is the recorder's, not a literal: a block that stopped
// running still reaches this line, and only the number betrays it.
println!(
"native-surface: walk complete, {} assertions passed",
PASSED.load(Ordering::SeqCst)
);
let _ = io::stdout().flush();
ExitCode::SUCCESS
}
}
/// Walk the surface once and report.
#[cfg(any(target_os = "linux", target_os = "freebsd"))]
fn main() -> std::process::ExitCode {
surface::run()
}
/// Refuse cleanly where the native API client is not built.
#[cfg(not(any(target_os = "linux", target_os = "freebsd")))]
fn main() -> std::process::ExitCode {
eprintln!(
"native-surface needs the native datagram API client, which is built on \
Linux and FreeBSD only"
);
std::process::ExitCode::FAILURE
}
+13
View File
@@ -177,6 +177,8 @@ enum ShowCommands {
Routing,
/// Identity cache entries (known node pubkeys)
IdentityCache,
/// Native datagram API flows and listeners
NativeFlows,
}
#[derive(Subcommand, Debug)]
@@ -200,6 +202,7 @@ impl ShowCommands {
ShowCommands::Transports => "show_transports",
ShowCommands::Routing => "show_routing",
ShowCommands::IdentityCache => "show_identity_cache",
ShowCommands::NativeFlows => "show_native_flows",
}
}
}
@@ -1481,6 +1484,16 @@ mod tests {
assert_eq!(AclCommands::Show.command_name(), "show_acl");
}
#[test]
fn test_cli_parses_show_native_flows_to_its_control_command() {
let cli = Cli::try_parse_from(["fipsctl", "show", "native-flows"]).unwrap();
let Commands::Show { what } = cli.command else {
panic!("expected a show subcommand");
};
assert_eq!(what.command_name(), "show_native_flows");
}
#[test]
fn test_cli_parses_acl_show() {
let cli = Cli::try_parse_from(["fipsctl", "acl", "show"]).unwrap();
+95 -2
View File
@@ -38,8 +38,8 @@ use zeroize::{Zeroize, Zeroizing};
pub use gateway::{ConntrackConfig, GatewayConfig, GatewayDnsConfig, PortForward, Proto};
pub use node::{
BloomConfig, BuffersConfig, CacheConfig, ControlConfig, LimitsConfig, LookupConfig, MmpConfig,
NodeConfig, NostrRendezvousConfig, NostrRendezvousPolicy, RateLimitConfig, RekeyConfig,
RendezvousConfig, RetryConfig, SessionConfig, SessionMmpConfig, TreeConfig,
NativeApiConfig, NodeConfig, NostrRendezvousConfig, NostrRendezvousPolicy, RateLimitConfig,
RekeyConfig, RendezvousConfig, RetryConfig, SessionConfig, SessionMmpConfig, TreeConfig,
};
pub use peer::{ConnectPolicy, PeerAddress, PeerConfig, TransportSpec};
pub use transport::{
@@ -1114,6 +1114,37 @@ impl Config {
}
}
let native = &self.node.native_api;
// Both floors refuse a node that would start, answer every setup call
// and then drop every datagram a peer sent. A zero `backlog` makes the
// registry refuse every arrival; a zero `pending_per_flow` makes it
// announce the arrival and then refuse the datagram that caused it, so
// a peer's opening message is lost with no refusal anywhere.
if native.backlog < 1 {
return Err(ConfigError::Validation(
"node.native_api.backlog is 0 but must be at least 1: a listener with no \
backlog admits no flow, so every arrival would be dropped"
.to_string(),
));
}
if native.pending_per_flow < 1 {
return Err(ConfigError::Validation(
"node.native_api.pending_per_flow is 0 but must be at least 1: a flow that \
can hold nothing loses its peer's opening datagram between the arrival \
being announced and the client taking the flow"
.to_string(),
));
}
if native.pending_per_flow > NativeApiConfig::MAX_PENDING_PER_FLOW {
return Err(ConfigError::Validation(format!(
"node.native_api.pending_per_flow is {} but must not exceed {}: the whole batch \
is written onto a socket pair the client cannot read yet, and a larger one \
would not fit the send buffer",
native.pending_per_flow,
NativeApiConfig::MAX_PENDING_PER_FLOW
)));
}
// Reject loopback UDP bind combined with non-loopback peer addresses.
// Linux pins the source IP to a loopback-bound socket, so packets
// sent from such a socket to external peers are dropped at the
@@ -2619,6 +2650,68 @@ node:
.expect("positive established-bucket values must validate");
}
#[test]
fn test_validate_pending_per_flow_above_its_ceiling_rejected() {
let mut config = Config::default();
config.node.native_api.pending_per_flow = NativeApiConfig::MAX_PENDING_PER_FLOW + 1;
let err = config.validate().expect_err("validation should fail");
let msg = err.to_string();
assert!(msg.contains("pending_per_flow"), "got: {msg}");
}
#[test]
fn test_validate_pending_per_flow_at_its_ceiling_accepted() {
// The boundary is accepted, or the check would be refusing a value the
// send buffer takes and the ceiling would be a different number than
// the one it is documented as.
let mut config = Config::default();
config.node.native_api.pending_per_flow = NativeApiConfig::MAX_PENDING_PER_FLOW;
config
.validate()
.expect("the documented ceiling itself must validate");
}
#[test]
fn test_validate_native_api_backlog_zero_rejected() {
// Zero would start a node whose listeners admit no flow at all: the
// registry compares the pending depth against this number before it
// announces anything, so every arrival is dropped.
let mut config = Config::default();
config.node.native_api.backlog = 0;
let err = config.validate().expect_err("validation should fail");
let msg = err.to_string();
assert!(msg.contains("node.native_api.backlog"), "got: {msg}");
}
#[test]
fn test_validate_native_api_pending_per_flow_zero_rejected() {
// Zero is the worse of the two: the arrival is announced and the
// datagram that caused it is then refused, so a peer's opening message
// is lost with no refusal a client or an operator can see.
let mut config = Config::default();
config.node.native_api.pending_per_flow = 0;
let err = config.validate().expect_err("validation should fail");
let msg = err.to_string();
assert!(msg.contains("pending_per_flow"), "got: {msg}");
}
#[test]
fn test_validate_native_api_floors_accept_one() {
// The boundary itself validates, or the floors would be refusing a
// working node rather than a broken one.
let mut config = Config::default();
config.node.native_api.backlog = 1;
config.node.native_api.pending_per_flow = 1;
config
.validate()
.expect("a depth of one is a working node, not a refused one");
}
#[test]
fn test_validate_rekey_after_messages_zero_rejected() {
let mut config = Config::default();
+159
View File
@@ -938,6 +938,142 @@ impl ControlConfig {
}
}
/// Native datagram API socket (`node.native_api.*`).
///
/// **Experimental, and built on Linux, FreeBSD and macOS only.** The API hands a
/// client a file descriptor over `SCM_RIGHTS`, which Windows has no equivalent
/// of, and does it over an `AF_UNIX` `SOCK_SEQPACKET` socket, which macOS does
/// not implement. No listener is built on either, and this section is ignored
/// there.
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
pub struct NativeApiConfig {
/// Enable the native API socket (`node.native_api.enabled`).
///
/// Disabled by default. Any process that can open the socket can send as
/// this node's identity, and can receive mesh traffic on a port it chooses,
/// so enabling it is an explicit operator decision rather than a default.
#[serde(default = "NativeApiConfig::default_enabled")]
pub enabled: bool,
/// Unix socket path (`node.native_api.socket_path`).
#[serde(default = "NativeApiConfig::default_socket_path")]
pub socket_path: String,
/// Datagrams held for one flow (`node.native_api.pending_per_flow`).
///
/// Applies while a flow waits to be accepted and while an established
/// flow's client is slow to read. Mirrors
/// [`SessionConfig::pending_packets_per_dest`], which bounds the same shape
/// of problem on the session layer.
///
/// Bounded above by [`NativeApiConfig::MAX_PENDING_PER_FLOW`] at config
/// load. The whole batch is written onto a socket pair no process can read
/// yet, so a value large enough to exceed the send buffer would leave the
/// listener's task with a write it cannot complete.
///
/// Bounded below by 1 at the same place. Zero announces an arrival and then
/// refuses the datagram that caused it, losing a peer's opening message
/// with no refusal a client or an operator can see.
#[serde(default = "NativeApiConfig::default_pending_per_flow")]
pub pending_per_flow: usize,
/// Flows awaiting accept on one listener (`node.native_api.backlog`).
///
/// A client that announces interest and never answers cannot make the node
/// hold more than this, whatever a peer does.
///
/// Bounded below by 1 at config load. Zero would admit no flow at all: the
/// registry compares a listener's pending depth against this before it
/// announces anything, so every arrival would be dropped.
#[serde(default = "NativeApiConfig::default_backlog")]
pub backlog: usize,
/// Flows this node holds at once (`node.native_api.max_flows`).
#[serde(default = "NativeApiConfig::default_max_flows")]
pub max_flows: usize,
/// Answer the debug commands (`node.native_api.debug_commands`).
///
/// **Off by default, and not a supported interface.** The three commands
/// it admits (`inject`, `stats`, `arrive`) exist so the test harness can
/// drive the receive and dispatch paths without a wire. `inject` makes
/// the daemon write bytes the client chose into one of that client's own
/// flows, and `arrive` makes it dispatch a datagram as though a peer had
/// sent it, which reaches any listener this node holds. None of the three
/// belongs in a packaged node, so this key is what the test harness turns
/// on and nothing else does.
#[serde(default = "NativeApiConfig::default_debug_commands")]
pub debug_commands: bool,
}
impl Default for NativeApiConfig {
fn default() -> Self {
Self {
enabled: Self::default_enabled(),
socket_path: Self::default_socket_path(),
pending_per_flow: Self::default_pending_per_flow(),
backlog: Self::default_backlog(),
max_flows: Self::default_max_flows(),
debug_commands: Self::default_debug_commands(),
}
}
}
impl NativeApiConfig {
/// Largest `pending_per_flow` a node will start with.
///
/// The held batch is at most this many datagrams of at most `max_payload`
/// bytes each, written without waiting onto a socket pair whose other half
/// is still on its way to the client. At 64 and a 1362-byte payload that is
/// about 87 KB, which an ordinary `AF_UNIX` send buffer takes. The bound is
/// checked at config load so a value that would wedge a listener's task is
/// refused at startup rather than at the first arrival.
pub const MAX_PENDING_PER_FLOW: usize = 64;
fn default_enabled() -> bool {
false
}
fn default_pending_per_flow() -> usize {
16
}
fn default_backlog() -> usize {
16
}
fn default_max_flows() -> usize {
256
}
fn default_debug_commands() -> bool {
false
}
/// Default native API socket path, resolved beside the control socket.
///
/// On Windows the path is empty: the API is not built there, so no value
/// would be meaningful.
fn default_socket_path() -> String {
#[cfg(unix)]
{
super::resolve_default_socket("api.sock")
}
#[cfg(windows)]
{
String::new()
}
}
/// Whether this section carries nothing but its defaults.
///
/// Drives `skip_serializing_if` so a config file that never named the
/// section does not gain one when the config is serialized back out.
fn is_default(&self) -> bool {
*self == Self::default()
}
}
/// Internal buffers (`node.buffers.*`).
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct BuffersConfig {
@@ -1172,6 +1308,11 @@ pub struct NodeConfig {
#[serde(default)]
pub control: ControlConfig,
/// Native datagram API (`node.native_api.*`). Experimental; the listener
/// is built on Linux, FreeBSD and macOS only.
#[serde(default, skip_serializing_if = "NativeApiConfig::is_default")]
pub native_api: NativeApiConfig,
/// Metrics Measurement Protocol — link layer (`node.mmp.*`).
#[serde(default)]
pub mmp: MmpConfig,
@@ -1217,6 +1358,7 @@ impl Default for NodeConfig {
session: SessionConfig::default(),
buffers: BuffersConfig::default(),
control: ControlConfig::default(),
native_api: NativeApiConfig::default(),
mmp: MmpConfig::default(),
session_mmp: SessionMmpConfig::default(),
ecn: EcnConfig::default(),
@@ -1420,6 +1562,23 @@ owd_window_size: 48
}
}
#[test]
fn test_native_api_is_off_and_undebuggable_by_default() {
// Both gates default closed, and neither has any other guard in the
// library: the harness case that leaves `debug_commands` out of its
// YAML proves the serde path, not the value it lands on.
let config = NativeApiConfig::default();
assert!(!config.enabled);
assert!(!config.debug_commands);
// Enabling the API must not drag the debug commands in with it, which
// is the shape a real operator config takes.
let yaml = "enabled: true\n";
let parsed: NativeApiConfig = serde_yaml::from_str(yaml).unwrap();
assert!(parsed.enabled);
assert!(!parsed.debug_commands);
}
#[cfg(windows)]
#[test]
fn test_default_socket_path_windows() {
+2 -236
View File
@@ -132,61 +132,6 @@ mod unix_impl {
use std::path::{Path, PathBuf};
use tokio::net::UnixListener;
/// Ensure the socket's parent exists and report whether this call created
/// the leaf directory.
///
/// `create_dir` gives us an atomic ownership decision: an `AlreadyExists`
/// result means another actor owns the existing directory, while success
/// means it is safe for this bind to apply FIPS ownership and mode. Missing
/// ancestors are created recursively, but only the requested leaf is later
/// treated as the socket's private directory.
fn ensure_socket_parent(parent: &Path) -> Result<bool, std::io::Error> {
if parent.as_os_str().is_empty() {
return Ok(false);
}
match std::fs::create_dir(parent) {
Ok(()) => Ok(true),
Err(error) if error.kind() == std::io::ErrorKind::AlreadyExists => {
if parent.is_dir() {
Ok(false)
} else {
Err(error)
}
}
Err(error) if error.kind() == std::io::ErrorKind::NotFound => {
let ancestor = parent.parent().ok_or(error)?;
ensure_socket_parent(ancestor)?;
ensure_socket_parent(parent)
}
Err(error) => Err(error),
}
}
/// Apply access policy to a newly bound control socket.
///
/// The socket is always group-owned. `managed_parent` is either a private
/// directory this bind created or a canonical FIPS runtime directory. A
/// shared or operator-owned existing parent is omitted so it retains its
/// ownership and mode.
fn set_control_socket_access(
socket_path: &Path,
managed_parent: Option<&Path>,
mut chown_to_fips_group: impl FnMut(&Path),
) -> Result<(), std::io::Error> {
use std::os::unix::fs::PermissionsExt;
std::fs::set_permissions(socket_path, std::fs::Permissions::from_mode(0o770))?;
chown_to_fips_group(socket_path);
if let Some(parent) = managed_parent {
std::fs::set_permissions(parent, std::fs::Permissions::from_mode(0o750))?;
chown_to_fips_group(parent);
}
Ok(())
}
/// Control socket listener (Unix domain socket).
///
/// Manages the Unix domain socket lifecycle: bind, accept, cleanup.
@@ -202,36 +147,7 @@ mod unix_impl {
/// and binds the Unix listener.
pub fn bind(config: &ControlConfig) -> Result<Self, std::io::Error> {
let socket_path = PathBuf::from(&config.socket_path);
// Creation is useful for diagnostics, but ownership is keyed to
// directory identity as well: systemd pre-creates /run/fips on
// every Linux service start and initially owns it as root:root.
let managed_parent = match socket_path.parent() {
Some(parent) => {
let created = ensure_socket_parent(parent)?;
if created {
debug!(path = %parent.display(), "Created private control socket directory");
}
(created || crate::config::is_managed_socket_parent(parent))
.then(|| parent.to_owned())
}
None => None,
};
// Remove stale socket if it exists
if socket_path.exists() {
Self::remove_stale_socket(&socket_path)?;
}
let listener = UnixListener::bind(&socket_path)?;
// Make the socket and its managed private directory group-accessible
// so fips group members can use fipsctl/fipstop.
set_control_socket_access(
&socket_path,
managed_parent.as_deref(),
Self::chown_to_fips_group,
)?;
let listener = crate::utils::sockbind::bind(&socket_path, "control")?;
info!(path = %socket_path.display(), "Control socket listening");
@@ -241,60 +157,6 @@ mod unix_impl {
})
}
/// Remove a stale socket file.
///
/// If the file exists but no one is listening, remove it so we can
/// bind. This handles unclean daemon exits.
fn remove_stale_socket(path: &Path) -> Result<(), std::io::Error> {
// Try connecting to see if someone is listening
match std::os::unix::net::UnixStream::connect(path) {
Ok(_) => {
// Someone is listening — don't remove it
Err(std::io::Error::new(
std::io::ErrorKind::AddrInUse,
format!("control socket already in use: {}", path.display()),
))
}
Err(_) => {
// No one listening — remove the stale socket
debug!(path = %path.display(), "Removing stale control socket");
std::fs::remove_file(path)?;
Ok(())
}
}
}
/// Set group ownership of a path to the 'fips' group (best-effort).
fn chown_to_fips_group(path: &Path) {
use std::ffi::CString;
use std::os::unix::ffi::OsStrExt;
// Look up the 'fips' group
let group_name = CString::new("fips").unwrap();
let grp = unsafe { libc::getgrnam(group_name.as_ptr()) };
if grp.is_null() {
debug!(
"'fips' group not found, skipping chown for {}",
path.display()
);
return;
}
let gid = unsafe { (*grp).gr_gid };
let c_path = match CString::new(path.as_os_str().as_bytes()) {
Ok(p) => p,
Err(_) => return,
};
let ret = unsafe { libc::chown(c_path.as_ptr(), u32::MAX, gid) };
if ret != 0 {
warn!(
path = %path.display(),
error = %std::io::Error::last_os_error(),
"Failed to chown control socket to 'fips' group"
);
}
}
/// Run the accept loop, forwarding requests to the main event loop via mpsc.
///
/// Each accepted connection is handled in a spawned task:
@@ -334,17 +196,7 @@ mod unix_impl {
/// Clean up the socket file.
fn cleanup(&self) {
if self.socket_path.exists() {
if let Err(e) = std::fs::remove_file(&self.socket_path) {
warn!(
path = %self.socket_path.display(),
error = %e,
"Failed to remove control socket"
);
} else {
debug!(path = %self.socket_path.display(), "Control socket removed");
}
}
crate::utils::sockbind::cleanup(&self.socket_path, "control");
}
}
@@ -353,92 +205,6 @@ mod unix_impl {
self.cleanup();
}
}
#[cfg(test)]
mod tests {
use super::{ensure_socket_parent, set_control_socket_access};
use std::os::unix::fs::PermissionsExt;
#[test]
fn parent_setup_distinguishes_existing_and_created_directories() {
let temp = tempfile::tempdir().unwrap();
let existing = temp.path().join("existing");
std::fs::create_dir(&existing).unwrap();
assert!(!ensure_socket_parent(&existing).unwrap());
let nested = temp.path().join("missing").join("fips");
assert!(ensure_socket_parent(&nested).unwrap());
assert!(nested.is_dir());
assert!(!ensure_socket_parent(&nested).unwrap());
}
#[test]
fn access_setup_leaves_an_existing_shared_parent_unchanged() {
let temp = tempfile::tempdir().unwrap();
let parent = temp.path().join("shared");
std::fs::create_dir(&parent).unwrap();
std::fs::set_permissions(&parent, std::fs::Permissions::from_mode(0o711)).unwrap();
let socket = parent.join("control.sock");
std::fs::File::create(&socket).unwrap();
let mut chowned = Vec::new();
set_control_socket_access(&socket, None, |path| chowned.push(path.to_path_buf()))
.unwrap();
assert_eq!(chowned, vec![socket.clone()]);
assert_eq!(
std::fs::metadata(&parent).unwrap().permissions().mode() & 0o777,
0o711
);
assert_eq!(
std::fs::metadata(&socket).unwrap().permissions().mode() & 0o777,
0o770
);
}
#[test]
fn access_setup_secures_a_new_private_parent() {
let temp = tempfile::tempdir().unwrap();
let parent = temp.path().join("fips");
std::fs::create_dir(&parent).unwrap();
let socket = parent.join("control.sock");
std::fs::File::create(&socket).unwrap();
let mut chowned = Vec::new();
set_control_socket_access(&socket, Some(&parent), |path| {
chowned.push(path.to_path_buf())
})
.unwrap();
assert_eq!(chowned, vec![socket, parent.clone()]);
assert_eq!(
std::fs::metadata(&parent).unwrap().permissions().mode() & 0o777,
0o750
);
}
#[test]
fn access_setup_secures_an_existing_managed_parent() {
let temp = tempfile::tempdir().unwrap();
let parent = temp.path().join("managed");
std::fs::create_dir(&parent).unwrap();
std::fs::set_permissions(&parent, std::fs::Permissions::from_mode(0o700)).unwrap();
let socket = parent.join("control.sock");
std::fs::File::create(&socket).unwrap();
let mut chowned = Vec::new();
set_control_socket_access(&socket, Some(&parent), |path| {
chowned.push(path.to_path_buf())
})
.unwrap();
assert_eq!(chowned, vec![socket, parent.clone()]);
assert_eq!(
std::fs::metadata(&parent).unwrap().permissions().mode() & 0o777,
0o750
);
}
}
}
// ============================================================================
+289 -1
View File
@@ -1634,6 +1634,124 @@ pub(crate) fn show_identity_cache_from_handle(
})
}
/// How `show_native_flows` names a flow's lifecycle state.
///
/// Two states and not three: the registry either holds a flow a client has
/// taken, or one it announced and is still waiting to be answered about. A
/// rejected or expired flow is gone from the registry entirely and has nothing
/// to report.
fn native_flow_state(established: bool) -> &'static str {
if established {
"established"
} else {
"pending_accept"
}
}
/// `show_native_flows` — Native datagram API flows and listeners.
///
/// Two representations of one peer, on purpose. `peer` is the npub, which is
/// the address a client names and the only form the native API itself reports.
/// `peer_addr` is the 16-byte node address, kept because this is an operator
/// surface and it is what `show_sessions` and `show_routing` key on. Both are
/// always present: the flow carries its peer's key, so nothing is resolved
/// here and nothing can be missing.
pub fn show_native_flows(node: &Node) -> Value {
let now = now_ms();
let flows: Vec<Value> = node
.native()
.flows()
.into_iter()
.map(|view| {
json!({
"flow_id": view.flow,
"peer": crate::identity::encode_npub(&view.pubkey),
"peer_addr": hex::encode(view.key.peer.as_bytes()),
"local_port": view.key.local,
"remote_port": view.key.remote,
"state": native_flow_state(view.established),
"queued": view.queued,
"age_ms": now.saturating_sub(view.at),
})
})
.collect();
let listeners: Vec<Value> = node
.native()
.listeners()
.into_iter()
.map(|view| {
json!({
"local_port": view.port,
"backlog": view.backlog,
})
})
.collect();
let native_stats = node.metrics().native.snapshot();
json!({
"flows": flows,
"listeners": listeners,
"stats": serde_json::to_value(&native_stats).unwrap_or_default(),
})
}
/// Off-loop variant of [`show_native_flows`]: renders from the tick-published
/// [`NativeSnapshot`](super::snapshot::NativeSnapshot) plus the `native`
/// counter family from the `MetricsRegistry`.
/// `age_ms` is derived at render time from the captured `since_ms`, exactly as
/// [`show_native_flows`] computed it. Output is byte-identical to
/// [`show_native_flows`].
///
/// A flow's `queued` depth is as of the last publish rather than as of the
/// read, which is the same point-in-time property every other snapshot cell
/// has; the underlying channel is drained by the client's own task and has no
/// value a control task could read anyway.
pub(crate) fn show_native_flows_from_handle(
handle: &super::read_handle::ControlReadHandle,
) -> Value {
let native = handle.native();
let now = now_ms();
let flows: Vec<Value> = native
.flows
.iter()
.map(|row| {
json!({
"flow_id": row.flow,
"peer": crate::identity::encode_npub(&row.peer_key),
"peer_addr": hex::encode(row.peer.as_bytes()),
"local_port": row.local_port,
"remote_port": row.remote_port,
"state": native_flow_state(row.established),
"queued": row.queued,
"age_ms": now.saturating_sub(row.since_ms),
})
})
.collect();
let listeners: Vec<Value> = native
.listeners
.iter()
.map(|row| {
json!({
"local_port": row.local_port,
"backlog": row.backlog,
})
})
.collect();
let native_stats = handle.metrics().native.snapshot();
json!({
"flows": flows,
"listeners": listeners,
"stats": serde_json::to_value(&native_stats).unwrap_or_default(),
})
}
/// `show_stats_list` — Enumerate available history metrics and their units.
pub fn show_stats_list() -> Value {
let metrics: Vec<Value> = ALL_METRICS
@@ -2353,6 +2471,7 @@ pub(crate) fn show_metrics_from_handle(handle: &super::read_handle::ControlReadH
"bloom": m.bloom.snapshot(),
"congestion": m.congestion.snapshot(),
"errors": m.errors.snapshot(),
"native": m.native.snapshot(),
})
}
@@ -2578,7 +2697,7 @@ mod tests {
}
}
// ---- 18 handler snapshot tests --------------------------------------
// ---- 19 handler snapshot tests --------------------------------------
#[test]
fn snapshot_show_status() {
@@ -2658,6 +2777,12 @@ mod tests {
assert_snapshot("show_identity_cache", &render(show_identity_cache(&node)));
}
#[test]
fn snapshot_show_native_flows() {
let node = build_test_node();
assert_snapshot("show_native_flows", &render(show_native_flows(&node)));
}
#[test]
fn snapshot_show_stats_list() {
// Static — no Node needed.
@@ -2769,6 +2894,7 @@ mod tests {
("show_connections", None),
("show_transports", None),
("show_mmp", None),
("show_native_flows", None),
];
for (cmd, params) in read_queries {
let req = Request {
@@ -2939,6 +3065,7 @@ mod tests {
("bloom", "accepted"),
("congestion", "ce_forwarded"),
("errors", "coords_required"),
("native", "flows_opened"),
];
assert_eq!(
obj.len(),
@@ -3305,6 +3432,167 @@ mod tests {
);
}
// ---- native datagram API coverage ------------------------------------
/// Freshness + fidelity: after a `record_stats_history()` tick (the native
/// publisher site) the off-loop `show_native_flows` render equals its
/// on-loop oracle byte-for-byte, and the query is served off-loop.
#[test]
fn native_snapshot_matches_on_loop_after_tick() {
use super::super::protocol::Request;
use super::super::read_handle::snapshot_dispatch;
let mut node = build_test_node();
// Put something in the registry first. Asserting the seed cell is empty
// against an empty registry would observe the absence of the very thing
// it checks and could not fail; with a listener already bound, an empty
// cell says the handle reads a published snapshot rather than the live
// registry.
let (arrivals, _arrivals_rx) = tokio::sync::mpsc::channel(8);
node.native_registry_for_test()
.listen(Some(4242), arrivals)
.expect("port 4242 is free on a fresh node");
let handle = node.control_read_handle();
assert!(
handle.native().flows.is_empty() && handle.native().listeners.is_empty(),
"the seed native snapshot is empty until the first tick publishes"
);
// Advance one tick (the publisher site).
node.record_stats_history();
let handle = node.control_read_handle();
assert_eq!(
handle.native().listeners.len(),
1,
"the tick publishes the listener the registry already held"
);
let req = Request {
command: "show_native_flows".to_string(),
params: None,
};
let resp =
snapshot_dispatch(&req, &handle).expect("show_native_flows must be served off-loop");
assert_eq!(
resp.status, "ok",
"show_native_flows off-loop response not ok"
);
assert_eq!(
render(show_native_flows(&node)),
render(show_native_flows_from_handle(&handle)),
"off-loop show_native_flows must match on-loop output"
);
}
/// The publisher carries every field the oracle emits, for flows it can only
/// get wrong once there are some to get wrong. The registry is populated
/// directly (the rx_loop's own handler is what does this in the daemon) and
/// then published: an empty-registry parity test passes just as happily
/// against a publisher that drops every per-flow field.
///
/// Both flow kinds appear, because they are rendered by different arms: a
/// flow the client opened, and one pending accept whose peer the node has
/// never had in any cache.
#[test]
fn native_snapshot_carries_a_populated_registry_faithfully() {
use crate::native::registry::Delivery;
let mut node = build_test_node();
// A flow the client opened, naming its peer by key.
let known = Identity::from_secret_bytes(&[0x11; 32]).expect("valid secret key");
let known_addr = *known.node_addr();
// And one a listener accepted, whose peer the node has never had in
// any cache: the key still reaches the report, because the flow
// carries it rather than the report resolving it.
let stranger = Identity::from_secret_bytes(&[0x5C; 32]).expect("valid secret key");
let unknown_addr = *stranger.node_addr();
let (sink, _sink_rx) = tokio::sync::mpsc::channel(8);
node.native_registry_for_test()
.connect(known.pubkey(), 5000, Some(6000), sink, 1_000)
.expect("a fresh node holds no flows");
let (arrivals, _arrivals_rx) = tokio::sync::mpsc::channel(8);
node.native_registry_for_test()
.listen(Some(4242), arrivals)
.expect("port 4242 is free on a fresh node");
let registry = node.native_registry_for_test();
let announced = match registry.deliver(unknown_addr, stranger.pubkey(), 5001, 4242, 2_000) {
Delivery::Arrived(_, arrival) => arrival.flow,
other => panic!("expected an arrival, got {other:?}"),
};
assert!(registry.hold(announced, b"held".to_vec()));
node.record_stats_history();
let handle = node.control_read_handle();
let on_loop = show_native_flows(&node);
let flows = on_loop["flows"].as_array().expect("flows is an array");
assert_eq!(flows.len(), 2, "both flows are reported");
assert_eq!(flows[0]["flow_id"], json!(1));
assert_eq!(
flows[0]["peer_addr"],
json!(hex::encode(known_addr.as_bytes()))
);
assert_eq!(flows[0]["peer"], json!(known.npub()));
assert_eq!(flows[0]["local_port"], json!(6000));
assert_eq!(flows[0]["remote_port"], json!(5000));
assert_eq!(flows[0]["state"], json!("established"));
assert_eq!(flows[0]["queued"], json!(0));
assert_eq!(flows[1]["flow_id"], json!(announced));
assert_eq!(
flows[1]["peer_addr"],
json!(hex::encode(unknown_addr.as_bytes()))
);
assert_eq!(
flows[1]["peer"],
json!(stranger.npub()),
"a flow learned from the wire reports its peer's npub too: the key \
was captured where the session authenticated it, so nothing has to \
invert the node address"
);
assert_eq!(flows[1]["local_port"], json!(4242));
assert_eq!(flows[1]["remote_port"], json!(5001));
assert_eq!(flows[1]["state"], json!("pending_accept"));
assert_eq!(flows[1]["queued"], json!(1), "the held datagram is queued");
let listeners = on_loop["listeners"]
.as_array()
.expect("listeners is an array");
assert_eq!(listeners.len(), 1, "the bound listener is reported");
assert_eq!(listeners[0]["local_port"], json!(4242));
assert_eq!(listeners[0]["backlog"], json!(1));
// `age_ms` is derived from `since_ms` and is redacted by `render`, so the
// parity assertion below cannot see the publisher dropping the flow
// timestamp. Pin the captured absolute times against the ones the
// registry was given.
assert_eq!(
handle.native().flows[0].since_ms,
1_000,
"the publisher carries the connect time the registry recorded"
);
assert_eq!(
handle.native().flows[1].since_ms,
2_000,
"the publisher carries the announce time the registry recorded"
);
// And the published snapshot renders the same thing.
assert_eq!(
render(on_loop),
render(show_native_flows_from_handle(&handle)),
"off-loop show_native_flows must match on-loop output for live flows"
);
}
/// Structural sharing: a republish in which only
/// one row changed re-allocates only that one `Arc<Row>` — every unchanged
/// row is reused by pointer (`Arc::ptr_eq`). Exercises
+27 -8
View File
@@ -13,8 +13,11 @@
//! - `entities` — `ArcSwap<EntitySnapshot>`: peers / sessions / links /
//! connections / transports, published from the tick with `Vec<Arc<Row>>`
//! structural sharing.
//! - `native` — `ArcSwap<NativeSnapshot>`: native datagram API flows and
//! listeners, published from the tick because the registry never leaves the
//! rx_loop.
//!
//! Publisher placement: all three snapshot cells are published from the
//! Publisher placement: all four snapshot cells are published from the
//! periodic tick, which runs as one arm of the rx_loop's `select!`. Publishing
//! therefore costs the rx_loop; what the handle removes is the read-side round
//! trip out to the rx_loop and back, not the cost of publishing. The
@@ -38,7 +41,7 @@ use crate::node::context::NodeContext;
use crate::node::metrics::MetricsRegistry;
use super::protocol::{Request, Response};
use super::snapshot::{EntitySnapshot, RoutingSnapshot, StatsSnapshot};
use super::snapshot::{EntitySnapshot, NativeSnapshot, RoutingSnapshot, StatsSnapshot};
/// Cloneable read-only view of node state for off-loop control serving.
///
@@ -62,6 +65,9 @@ pub(crate) struct ControlReadHandle {
/// connections / transports + mmp), published from the tick with
/// `Vec<Arc<Row>>` structural sharing.
entities: Arc<ArcSwap<EntitySnapshot>>,
/// Native datagram API read view (flows / listeners), published
/// from the tick because the registry is the rx_loop's alone.
native: Arc<ArcSwap<NativeSnapshot>>,
}
impl ControlReadHandle {
@@ -75,6 +81,7 @@ impl ControlReadHandle {
stats: Arc<ArcSwap<StatsSnapshot>>,
routing: Arc<ArcSwap<RoutingSnapshot>>,
entities: Arc<ArcSwap<EntitySnapshot>>,
native: Arc<ArcSwap<NativeSnapshot>>,
) -> Self {
Self {
context,
@@ -82,6 +89,7 @@ impl ControlReadHandle {
stats,
routing,
entities,
native,
}
}
@@ -112,6 +120,12 @@ impl ControlReadHandle {
pub(crate) fn entities(&self) -> arc_swap::Guard<Arc<EntitySnapshot>> {
self.entities.load()
}
/// Load the latest published native-API snapshot (freshest
/// available by construction; no staleness gate).
pub(crate) fn native(&self) -> arc_swap::Guard<Arc<NativeSnapshot>> {
self.native.load()
}
}
/// Attempt to serve a request entirely from the read handle, off the rx_loop.
@@ -120,12 +134,13 @@ impl ControlReadHandle {
/// snapshot cells, or `None` when it must take the mpsc → rx_loop path.
///
/// The queries served here read any of the cells the handle bundles —
/// `context`, `metrics`, `stats`, `routing`, `entities` — plus host-OS facts
/// gathered in [`super::listening`] and [`super::firewall_state`], so they
/// render in the control task without touching `Node`. Taking a parameter is
/// not what decides it: `show_stats_history` is parameterized and is served
/// here. What falls back is a query needing live `Node` state the snapshot does
/// not carry, and every mutation.
/// `context`, `metrics`, `stats`, `routing`, `entities`, `native` — plus
/// host-OS facts gathered in [`super::listening`] and
/// [`super::firewall_state`], so they render in the control task without
/// touching `Node`. Taking a parameter is not what decides it:
/// `show_stats_history` is parameterized and is served here. What falls back is
/// a query needing live `Node` state the snapshot does not carry, and every
/// mutation.
///
/// **It now also carries mutating commands**, namely the `profile_tick_*`
/// family under the `profiling` feature. They are served here rather than on
@@ -219,6 +234,10 @@ pub(crate) fn snapshot_dispatch(request: &Request, handle: &ControlReadHandle) -
"show_connections" => Some(Response::ok(queries::show_connections_from_handle(handle))),
"show_transports" => Some(Response::ok(queries::show_transports_from_handle(handle))),
"show_mmp" => Some(Response::ok(queries::show_mmp_from_handle(handle))),
// Served from the tick-published `NativeSnapshot`. Peer npubs are
// resolved at publish time, and the `native` counter family comes from
// the `MetricsRegistry`, so this renders entirely off-loop.
"show_native_flows" => Some(Response::ok(queries::show_native_flows_from_handle(handle))),
_ => None,
}
}
+71
View File
@@ -25,6 +25,7 @@ use crate::node::NodeState;
use crate::node::acl::PeerAclStatus;
use crate::node::stats_history::StatsHistory;
use crate::upper::tun::TunState;
use secp256k1::XOnlyPublicKey;
/// Read-only snapshot of the stats-history rings plus the scalar gauges and
/// counts `show_status` reports. Published from the tick.
@@ -395,6 +396,76 @@ pub(crate) struct IdentityRow {
pub last_seen_ms: u64,
}
// =====================================================================
// NativeSnapshot (native datagram API read view)
// =====================================================================
/// Read-only snapshot of the native datagram API registry that
/// `show_native_flows` renders. Published via `ArcSwap`.
///
/// **Publisher placement.** The registry lives inside `Node` and is touched
/// only by the rx_loop, which is precisely why the native receive path takes no
/// lock; a control task cannot read it at all, and putting a lock on it to allow
/// that would give back the property the design was built for. So the
/// projection is published from the tick beside the other snapshot cells.
///
/// **Cost.** A field read per flow. The peer's key is captured where its
/// session authenticated it and rides on the registry entry, so publishing a
/// row resolves nothing and consults no cache.
///
/// Time-relative fields (`age_ms`) are derived at render time from the captured
/// absolute timestamps, so a rendered age stays fresh relative to the read.
#[derive(Clone, Default)]
pub(crate) struct NativeSnapshot {
/// One row per flow, established and pending, ordered by identifier.
pub flows: Vec<NativeFlowRow>,
/// One row per listener, ordered by port.
pub listeners: Vec<NativeListenerRow>,
}
impl NativeSnapshot {
/// Build an empty snapshot for seeding the `ArcSwap` cell at construction,
/// before the first tick has published real state.
pub(crate) fn empty() -> Self {
Self::default()
}
}
/// One native API flow in `show_native_flows`.
#[derive(Clone)]
pub(crate) struct NativeFlowRow {
/// The identifier the client names the flow by.
pub flow: u64,
/// The far end, by node address. Kept because this is an operator surface
/// and the node address is what `show_sessions` and `show_routing` key on,
/// which is what lets a reader correlate the three.
pub peer: NodeAddr,
/// The far end's address, as the x-only public key. Always present: a
/// connected flow decoded it from the npub its client named, and an
/// accepted one took it from the session that authenticated the peer.
pub peer_key: XOnlyPublicKey,
/// This node's port.
pub local_port: u16,
/// The far end's port.
pub remote_port: u16,
/// Whether a client has taken the flow, as opposed to still awaiting accept.
pub established: bool,
/// Datagrams the node is holding for it.
pub queued: usize,
/// Absolute open / accept / announce time (Unix ms); `age_ms` derived at
/// render time.
pub since_ms: u64,
}
/// One native API listener in `show_native_flows`.
#[derive(Clone)]
pub(crate) struct NativeListenerRow {
/// The local port it holds.
pub local_port: u16,
/// Flows announced on it and not yet answered.
pub backlog: usize,
}
// =====================================================================
// EntitySnapshot (per-entity table read views)
// =====================================================================
@@ -0,0 +1,26 @@
{
"data": {
"flows": [],
"listeners": [],
"stats": {
"drop_arrival_queue_full": 0,
"drop_backlog_full": 0,
"drop_flow_queue_full": 0,
"drop_listener_gone": 0,
"drop_listener_not_reading": 0,
"drop_no_port": 0,
"drop_oversize": 0,
"drop_pending_queue_full": 0,
"drop_too_many_flows": 0,
"flows_accepted": 0,
"flows_closed": 0,
"flows_expired": 0,
"flows_opened": 0,
"received_bytes": 0,
"received_datagrams": 0,
"sent_bytes": 0,
"sent_datagrams": 0
}
},
"status": "ok"
}
+9 -83
View File
@@ -7,7 +7,7 @@
use crate::control::protocol::{Request, Response};
use crate::gateway::pool::{MappingInfo, MappingState, PoolStatus};
use std::path::{Path, PathBuf};
use std::path::PathBuf;
use std::time::Instant;
use tokio::io::{AsyncBufReadExt, AsyncWriteExt, BufReader};
use tokio::net::UnixListener;
@@ -54,35 +54,16 @@ pub struct GatewayControlSocket {
}
impl GatewayControlSocket {
/// Bind the gateway control socket.
/// Bind the gateway control socket under the shared FIPS access policy.
///
/// Creates parent directories if needed, removes stale socket files,
/// and sets `root:fips 0770` permissions.
/// The policy creates the parent directory, removes a stale socket file,
/// and applies mode `0770` plus the `fips` group. It also applies `0750` to
/// a parent it manages, which `/run/fips` is; systemd already creates that
/// directory at `0750` for the daemon this service requires, so under the
/// packaged deployment it is set to the value it already holds.
pub fn bind() -> Result<Self, std::io::Error> {
let socket_path = PathBuf::from(GATEWAY_SOCKET_PATH);
// Create parent directory if it doesn't exist
if let Some(parent) = socket_path.parent()
&& !parent.exists()
{
std::fs::create_dir_all(parent)?;
debug!(path = %parent.display(), "Created gateway control socket directory");
}
// Remove stale socket if it exists
if socket_path.exists() {
Self::remove_stale_socket(&socket_path)?;
}
let listener = UnixListener::bind(&socket_path)?;
// Set permissions to 0770 and chown to fips group
use std::os::unix::fs::PermissionsExt;
std::fs::set_permissions(&socket_path, std::fs::Permissions::from_mode(0o770))?;
Self::chown_to_fips_group(&socket_path);
if let Some(parent) = socket_path.parent() {
Self::chown_to_fips_group(parent);
}
let listener = crate::utils::sockbind::bind(&socket_path, "gateway control")?;
info!(path = %socket_path.display(), "Gateway control socket listening");
@@ -92,51 +73,6 @@ impl GatewayControlSocket {
})
}
/// Remove a stale socket file from a previous unclean exit.
fn remove_stale_socket(path: &Path) -> Result<(), std::io::Error> {
match std::os::unix::net::UnixStream::connect(path) {
Ok(_) => Err(std::io::Error::new(
std::io::ErrorKind::AddrInUse,
format!("gateway control socket already in use: {}", path.display()),
)),
Err(_) => {
debug!(path = %path.display(), "Removing stale gateway control socket");
std::fs::remove_file(path)?;
Ok(())
}
}
}
/// Set group ownership to the `fips` group (best-effort).
fn chown_to_fips_group(path: &Path) {
use std::ffi::CString;
use std::os::unix::ffi::OsStrExt;
let group_name = CString::new("fips").unwrap();
let grp = unsafe { libc::getgrnam(group_name.as_ptr()) };
if grp.is_null() {
debug!(
"'fips' group not found, skipping chown for {}",
path.display()
);
return;
}
let gid = unsafe { (*grp).gr_gid };
let c_path = match CString::new(path.as_os_str().as_bytes()) {
Ok(p) => p,
Err(_) => return,
};
let ret = unsafe { libc::chown(c_path.as_ptr(), u32::MAX, gid) };
if ret != 0 {
warn!(
path = %path.display(),
error = %std::io::Error::last_os_error(),
"Failed to chown gateway control socket to 'fips' group"
);
}
}
/// Run the accept loop, reading the latest snapshot from the watch channel.
pub async fn accept_loop(self, snapshot_rx: watch::Receiver<Option<GatewaySnapshot>>) {
loop {
@@ -218,17 +154,7 @@ impl GatewayControlSocket {
/// Clean up the socket file.
fn cleanup(&self) {
if self.socket_path.exists() {
if let Err(e) = std::fs::remove_file(&self.socket_path) {
warn!(
path = %self.socket_path.display(),
error = %e,
"Failed to remove gateway control socket"
);
} else {
debug!(path = %self.socket_path.display(), "Gateway control socket removed");
}
}
crate::utils::sockbind::cleanup(&self.socket_path, "gateway control");
}
}
+1
View File
@@ -21,6 +21,7 @@ pub mod identity;
#[macro_use]
pub(crate) mod instr;
pub mod mdns;
pub mod native;
pub mod node;
pub mod noise;
pub mod nostr;
+408
View File
@@ -0,0 +1,408 @@
//! The native API's line protocol, as a client sees it.
//!
//! Pure: no sockets, no descriptors and no node state, so every rule here is
//! testable without a running daemon. It is the client-side mirror of
//! [`crate::native::protocol`], which does the same job for the daemon: one
//! module that knows the encoding, and a shell that knows only I/O.
//!
//! The encoding is the control socket's, one JSON object per line. A client
//! sends `{"command": "...", "params": {...}}` and reads back one
//! [`Response`](crate::control::protocol::Response). **Replies only, in command
//! order**: nothing unsolicited arrives on this socket, so a line is always the
//! answer to the command just written.
//!
//! An arriving flow is not a line. It is one `SOCK_SEQPACKET` message on the
//! listener's own descriptor, carrying the same fields a `connect` reply
//! carries, which is why [`opened`] reads both.
//!
//! **Errors come back as `io::Error` carrying the errno the contract names**,
//! read from `data.errno`. The daemon's `message` is for an operator reading a
//! log and is deliberately dropped: a client that matched on it would be
//! relying on prose, and `io::Error::from_raw_os_error` is the only
//! constructor that leaves `raw_os_error` readable, which is what a C binding
//! would return from the corresponding call.
use crate::identity::{decode_npub, encode_npub};
use secp256k1::XOnlyPublicKey;
use serde_json::{Value, json};
use std::io;
/// An npub in text form, as an x-only public key.
///
/// The client-side half of `fips_pton`: bech32 in one direction, and nothing
/// else. No daemon, no socket, no lookup. A string that is not an npub is
/// `EINVAL`, which is what `inet_pton`'s caller gets for a malformed address.
pub fn pton(text: &str) -> io::Result<XOnlyPublicKey> {
decode_npub(text).map_err(|_| io::Error::from_raw_os_error(libc::EINVAL))
}
/// An x-only public key as the npub that names it.
///
/// The client-side half of `fips_ntop`, and the exact inverse of [`pton`].
pub fn ntop(key: &XOnlyPublicKey) -> String {
encode_npub(key)
}
/// The `connect` command line.
///
/// `local` of 0 is the request for an ephemeral port, which is `bind(2)` with
/// port 0. It is spelled out rather than omitted: the daemon reads an absent
/// field the same way, but one spelling means one rule for `connect` and
/// `listen` both.
pub fn connect(peer: &XOnlyPublicKey, remote: u16, local: u16) -> Vec<u8> {
request(
"connect",
json!({"peer": ntop(peer), "remote_port": remote, "local_port": local}),
)
}
/// The `listen` command line, where 0 asks the daemon to pick the port.
pub fn listen(local: u16) -> Vec<u8> {
request("listen", json!({"local_port": local}))
}
/// Encode one command as the newline-terminated line the daemon reads.
fn request(command: &str, params: Value) -> Vec<u8> {
let mut line = serde_json::to_vec(&json!({"command": command, "params": params}))
.expect("a JSON object built here always serializes");
line.push(b'\n');
line
}
/// The `data` of one reply, or the error the daemon reported.
///
/// A refusal becomes an `io::Error` here rather than at the call site, so every
/// setup path reports the same way and none of them has to know that a refusal
/// is spelled as a successful read of an error line.
pub fn reply(line: &[u8]) -> io::Result<Value> {
let value: Value = serde_json::from_slice(line).map_err(|error| {
io::Error::new(
io::ErrorKind::InvalidData,
format!("the daemon sent a line that is not JSON: {error}"),
)
})?;
match value.get("status").and_then(Value::as_str) {
Some("ok") => Ok(value.get("data").cloned().unwrap_or(Value::Null)),
// A reply carrying no `errno` is `ECONNREFUSED`, which covers a daemon
// older than the code that added the field and any refusal raised
// outside the registry.
Some("error") => Err(io::Error::from_raw_os_error(code(
value
.get("data")
.and_then(|data| data.get("errno"))
.and_then(Value::as_str)
.unwrap_or("ECONNREFUSED"),
))),
Some(other) => Err(io::Error::new(
io::ErrorKind::InvalidData,
format!("the daemon reported an unknown status '{other}'"),
)),
None => Err(missing("status")),
}
}
/// The errno a name from the daemon stands for on this platform.
///
/// The daemon sends names rather than numbers because the number belongs to the
/// platform the client was built for and the daemon is not it. An unknown name
/// is `ECONNREFUSED` for the same reason a missing one is: it is the daemon
/// refusing for a reason this client has no row for.
fn code(name: &str) -> i32 {
match name {
"EADDRINUSE" => libc::EADDRINUSE,
"EADDRNOTAVAIL" => libc::EADDRNOTAVAIL,
"EMFILE" => libc::EMFILE,
"EINVAL" => libc::EINVAL,
"EMSGSIZE" => libc::EMSGSIZE,
"EPIPE" => libc::EPIPE,
"ETIMEDOUT" => libc::ETIMEDOUT,
_ => libc::ECONNREFUSED,
}
}
/// What the daemon said about a flow it opened.
///
/// One shape for two producers. A `connect` reply and a listener's arrival
/// message carry the same fields, because the daemon re-encodes the peer's key
/// itself in both rather than echoing what a client wrote, so one reader serves
/// both and an accepted flow reports its peer exactly as a connected one does.
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct Opened {
/// The far end's public key. The address, not a name for one.
pub peer: XOnlyPublicKey,
/// This node's own public key, carried so an accepted flow can answer
/// `getsockname` without consulting the listener that produced it.
pub node: XOnlyPublicKey,
/// The local port the flow holds.
pub local: u16,
/// The far end's port.
pub remote: u16,
/// The largest datagram this flow carries.
pub max: usize,
}
/// Read a flow description out of a `connect` reply's data or an arrival
/// message.
///
/// The arrival's `flow_id` and `held` are not read. Neither addresses anything
/// a client can name: the flow is named by its descriptor, and the held
/// datagrams are already on that descriptor by the time the message carrying it
/// can be read.
pub fn opened(data: &Value) -> io::Result<Opened> {
let max = data
.get("max_payload")
.and_then(Value::as_u64)
.ok_or_else(|| missing("max_payload"))?;
Ok(Opened {
peer: key(data, "peer")?,
node: key(data, "node")?,
local: port(data, "local_port")?,
remote: port(data, "remote_port")?,
max: max as usize,
})
}
/// Read a flow description out of one arrival message.
///
/// The message is a whole JSON object with **no trailing newline**: it is one
/// `SOCK_SEQPACKET` message, so the boundary is the framing and a newline would
/// offer a client a second framing to rely on.
pub fn arrival(message: &[u8]) -> io::Result<Opened> {
let value: Value = serde_json::from_slice(message).map_err(|error| {
io::Error::new(
io::ErrorKind::InvalidData,
format!("the daemon sent an arrival that is not JSON: {error}"),
)
})?;
opened(&value)
}
/// What the daemon said about a port it bound.
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct Bound {
/// This node's own public key, which is half of the listener's address.
pub node: XOnlyPublicKey,
/// The port actually held, which is what a client that asked for 0 needs.
pub local: u16,
}
/// Read a listener description out of a `listen` reply's data.
///
/// `backlog` is not read. It reports the depth the daemon holds so an operator
/// can size a client's reader, and nothing on this surface takes a depth.
pub fn bound(data: &Value) -> io::Result<Bound> {
Ok(Bound {
node: key(data, "node")?,
local: port(data, "local_port")?,
})
}
/// Read one npub field as the key it encodes.
fn key(data: &Value, field: &'static str) -> io::Result<XOnlyPublicKey> {
let text = data
.get(field)
.and_then(Value::as_str)
.ok_or_else(|| missing(field))?;
pton(text).map_err(|_| {
io::Error::new(
io::ErrorKind::InvalidData,
format!("the daemon's '{field}' is not an npub"),
)
})
}
/// Read one port field, refusing a number that is not one.
fn port(data: &Value, field: &'static str) -> io::Result<u16> {
let value = data
.get(field)
.and_then(Value::as_u64)
.ok_or_else(|| missing(field))?;
u16::try_from(value).map_err(|_| {
io::Error::new(
io::ErrorKind::InvalidData,
format!("the daemon's '{field}' is outside the range a port has"),
)
})
}
/// The error for a field this client needs and the daemon did not send.
fn missing(field: &str) -> io::Error {
io::Error::new(
io::ErrorKind::InvalidData,
format!("the daemon's line has no usable '{field}'"),
)
}
#[cfg(test)]
mod tests {
use super::*;
/// An npub the daemon could have written. Nothing here reaches a peer.
const PEER: &str = "npub1sjlh2c3x9w7kjsqg2ay080n2lff2uvt325vpan33ke34rn8l5jcqawh57m";
/// A second one, standing in for the node's own identity.
const NODE: &str = "npub1n9lpnv0592cc2ps6nm0ca3qls642vx7yjsv35rkxqzj2vgds52sqgpverl";
/// The `data` of a connect reply, and of an arrival message but for two
/// fields this client does not read.
fn flow() -> Value {
json!({
"flow_id": 3, "local_port": 49152, "remote_port": 4242,
"peer": PEER, "node": NODE, "max_payload": 1362,
})
}
#[test]
fn an_npub_survives_a_round_trip_through_the_two_pure_conversions() {
let key = pton(PEER).unwrap();
assert_eq!(ntop(&key), PEER);
}
#[test]
fn a_string_that_is_not_an_npub_is_refused_as_einval() {
let error = pton("not-an-npub").unwrap_err();
assert_eq!(error.raw_os_error(), Some(libc::EINVAL));
// An nsec is bech32 and the right length, so only the prefix separates
// it from an address. Taking it would name a key nobody may address.
assert!(pton("nsec1vl029mgpspedva04g90vltkh6fvh240zqtv9k0t9af8935ke9laqsnlfe5").is_err());
}
#[test]
fn a_connect_command_names_the_peer_by_the_daemons_own_spelling() {
let key = pton(PEER).unwrap();
let line = connect(&key, 4242, 0);
assert_eq!(line.last(), Some(&b'\n'));
let value: Value = serde_json::from_slice(line.trim_ascii_end()).unwrap();
assert_eq!(value["command"], "connect");
assert_eq!(value["params"]["peer"], PEER);
assert_eq!(value["params"]["remote_port"], 4242);
// Spelled rather than omitted: 0 is the ephemeral request, and one
// spelling covers connect and listen both.
assert_eq!(value["params"]["local_port"], 0);
}
#[test]
fn a_listen_command_carries_the_port_the_caller_asked_for() {
let value: Value = serde_json::from_slice(listen(4242).trim_ascii_end()).unwrap();
assert_eq!(value["command"], "listen");
assert_eq!(value["params"]["local_port"], 4242);
}
#[test]
fn an_ok_reply_yields_its_data() {
let line = br#"{"status":"ok","data":{"local_port":4242}}"#;
assert_eq!(reply(line).unwrap()["local_port"], 4242);
}
#[test]
fn a_refusal_becomes_the_errno_the_daemon_named() {
let line = br#"{"status":"error","message":"port 4242 is already in use","data":{"errno":"EADDRINUSE"}}"#;
let error = reply(line).unwrap_err();
assert_eq!(error.raw_os_error(), Some(libc::EADDRINUSE));
}
#[test]
fn a_refusal_carrying_no_errno_falls_back_to_econnrefused() {
// A daemon older than the field, or a refusal raised outside the
// registry. Guessing a more specific code would tell a caller something
// the daemon did not say.
let line = br#"{"status":"error","message":"no such flow: 9"}"#;
assert_eq!(
reply(line).unwrap_err().raw_os_error(),
Some(libc::ECONNREFUSED)
);
let unknown = br#"{"status":"error","data":{"errno":"EWHATEVER"}}"#;
assert_eq!(
reply(unknown).unwrap_err().raw_os_error(),
Some(libc::ECONNREFUSED)
);
}
#[test]
fn every_errno_the_daemon_can_name_maps_to_this_platforms_number() {
// The contract is the number a C binding would return, so a name this
// client did not translate would silently become ECONNREFUSED and
// collapse a row of the error table.
for (name, want) in [
("EADDRINUSE", libc::EADDRINUSE),
("EADDRNOTAVAIL", libc::EADDRNOTAVAIL),
("EMFILE", libc::EMFILE),
("EINVAL", libc::EINVAL),
("EMSGSIZE", libc::EMSGSIZE),
("EPIPE", libc::EPIPE),
("ETIMEDOUT", libc::ETIMEDOUT),
] {
assert_eq!(code(name), want, "{name}");
}
}
#[test]
fn a_line_that_is_not_json_is_reported_as_such() {
let error = reply(b"not json at all").unwrap_err();
assert_eq!(error.kind(), io::ErrorKind::InvalidData);
}
#[test]
fn a_reply_with_no_status_is_refused_rather_than_guessed_at() {
let error = reply(br#"{"data":{}}"#).unwrap_err();
assert!(error.to_string().contains("status"), "{error}");
}
#[test]
fn one_reader_describes_a_connect_reply_and_an_arrival_alike() {
let want = Opened {
peer: pton(PEER).unwrap(),
node: pton(NODE).unwrap(),
local: 49152,
remote: 4242,
max: 1362,
};
assert_eq!(opened(&flow()).unwrap(), want);
// The arrival message is the same object with two fields this client
// does not read. It must not need a second reader.
let mut arrival = flow();
arrival["held"] = json!(2);
assert_eq!(opened(&arrival).unwrap(), want);
}
#[test]
fn a_flow_whose_peer_is_not_an_npub_is_refused_rather_than_carried() {
// The peer is the address. A client that accepted a hex node address
// here would report something that cannot be sent to.
let mut data = flow();
data["peer"] = json!("aabbccddeeff00112233445566778899");
let error = opened(&data).unwrap_err();
assert!(error.to_string().contains("not an npub"), "{error}");
}
#[test]
fn a_reply_missing_the_payload_limit_names_the_field() {
let mut data = flow();
data.as_object_mut().unwrap().remove("max_payload");
let error = opened(&data).unwrap_err();
assert!(error.to_string().contains("max_payload"), "{error}");
}
#[test]
fn a_port_above_the_range_a_port_has_is_refused() {
let mut data = flow();
data["local_port"] = json!(70000);
let error = opened(&data).unwrap_err();
assert!(error.to_string().contains("local_port"), "{error}");
}
#[test]
fn a_listen_reply_reports_the_port_actually_held_and_the_nodes_own_key() {
let data = json!({"local_port": 49152, "node": NODE, "backlog": 16});
assert_eq!(
bound(&data).unwrap(),
Bound {
node: pton(NODE).unwrap(),
local: 49152,
}
);
}
}
File diff suppressed because it is too large Load Diff
+321
View File
@@ -0,0 +1,321 @@
//! What a connected `AF_UNIX` `SOCK_DGRAM` pair does at end of file.
//!
//! This module is tests only. It exists to answer, by measurement on each
//! kernel rather than from the manual pages, the one question a macOS port of
//! this API turns on.
//!
//! [`super::seqpacket`] uses `SOCK_SEQPACKET`, which macOS does not implement
//! for `AF_UNIX`. The candidate replacement there is `SOCK_DGRAM`, which macOS
//! does implement and which also keeps message boundaries. What is not
//! transferable is the rule that tells a close apart from an empty datagram:
//! both produce a zero-byte read, and `seqpacket` resolves them with a latched
//! `POLLHUP` measured on Linux 6.8. Datagram poll semantics differ between
//! kernels, so that rule has to be re-established on Darwin before anything is
//! built on it.
//!
//! **The tests below assert the properties an implementation would need.** A
//! failure here is the measurement coming back negative, not a regression: it
//! says this kernel cannot support the `seqpacket` close rule on `SOCK_DGRAM`
//! and that a macOS port needs a different close signal. The same code runs on
//! every unix so the platforms can be compared without the test itself being a
//! variable.
//!
//! `SOCK_CLOEXEC` is deliberately not passed in the type argument, though
//! `super::seqpacket::pair` does pass it. Linux and FreeBSD accept it there and
//! macOS does not, and that difference belongs to the port rather than to this
//! measurement.
use std::io;
use std::os::fd::{AsRawFd, FromRawFd, OwnedFd, RawFd};
/// Create a connected `AF_UNIX` `SOCK_DGRAM` pair.
fn dgram_pair() -> io::Result<(OwnedFd, OwnedFd)> {
let mut fds = [0 as libc::c_int; 2];
// SAFETY: `fds` is a two-element array of the type socketpair writes, and
// the call either fills both entries or reports failure.
let rc = unsafe { libc::socketpair(libc::AF_UNIX, libc::SOCK_DGRAM, 0, fds.as_mut_ptr()) };
if rc != 0 {
return Err(io::Error::last_os_error());
}
// SAFETY: socketpair reported success, so both entries are open descriptors
// this frame now owns.
Ok(unsafe { (OwnedFd::from_raw_fd(fds[0]), OwnedFd::from_raw_fd(fds[1])) })
}
/// One non-blocking receive, returning the byte count or the errno.
fn recv(fd: RawFd, buf: &mut [u8]) -> io::Result<usize> {
// SAFETY: the descriptor is open and the pointer and length describe `buf`.
let n = unsafe { libc::recv(fd, buf.as_mut_ptr().cast(), buf.len(), libc::MSG_DONTWAIT) };
if n < 0 {
return Err(io::Error::last_os_error());
}
Ok(n as usize)
}
/// Send one datagram, returning the byte count or the errno.
fn send(fd: RawFd, buf: &[u8]) -> io::Result<usize> {
// SAFETY: the descriptor is open and the pointer and length describe `buf`.
let n = unsafe { libc::send(fd, buf.as_ptr().cast(), buf.len(), 0) };
if n < 0 {
return Err(io::Error::last_os_error());
}
Ok(n as usize)
}
/// Whether `POLLHUP` is set, by the same rule `seqpacket::peer_hung_up` uses.
///
/// `POLLIN` is requested rather than nothing. An empty `events` registers no
/// filter on Darwin, so a poll asking for nothing reports nothing there and
/// every assertion built on this helper would pass whatever the kernel did.
/// That is the shape of a guard that executes and cannot fail, so it is worth
/// more than a comment: with `POLLIN` requested, a Darwin that began reporting
/// `POLLHUP` would red the tests below instead of slipping past them.
fn hung_up(fd: RawFd) -> bool {
let mut poll = libc::pollfd {
fd,
events: libc::POLLIN,
revents: 0,
};
// SAFETY: `poll` points at one live pollfd and the call cannot block.
let rc = unsafe { libc::poll(&mut poll, 1, 0) };
rc > 0 && (poll.revents & libc::POLLHUP) != 0
}
#[test]
fn a_connected_dgram_pair_keeps_message_boundaries() {
let (a, b) = dgram_pair().expect("AF_UNIX SOCK_DGRAM socketpair");
assert_eq!(send(a.as_raw_fd(), &[1, 2, 3]).unwrap(), 3);
assert_eq!(send(a.as_raw_fd(), &[4, 5]).unwrap(), 2);
// Two sends must read back as two messages of their own lengths. A stream
// socket would hand back all five bytes in one read, which is the failure
// this discriminates.
let mut buf = [0u8; 64];
assert_eq!(recv(b.as_raw_fd(), &mut buf).unwrap(), 3);
assert_eq!(&buf[..3], &[1, 2, 3]);
assert_eq!(recv(b.as_raw_fd(), &mut buf).unwrap(), 2);
assert_eq!(&buf[..2], &[4, 5]);
}
#[test]
fn an_empty_datagram_reads_as_zero_bytes_and_is_not_a_hangup() {
let (a, b) = dgram_pair().expect("AF_UNIX SOCK_DGRAM socketpair");
assert_eq!(send(a.as_raw_fd(), &[]).unwrap(), 0);
let mut buf = [0u8; 64];
assert_eq!(
recv(b.as_raw_fd(), &mut buf).unwrap(),
0,
"an empty datagram must be delivered as a zero-byte message"
);
assert!(
!hung_up(b.as_raw_fd()),
"an empty datagram must not look like a closed peer: if this fails, a \
client could tear down its own flow by sending nothing"
);
}
/// Linux 6.8, measured 2026-08-20 with this test and cross-checked with a C
/// probe over both socket types: a connected `AF_UNIX` `SOCK_DGRAM` pair gives
/// **no close signal at all**. The peer closing leaves `revents` empty and
/// leaves `recv` returning `EAGAIN`, which is what an idle socket with a live
/// peer also does. `SOCK_SEQPACKET` on the same kernel sets `POLLHUP` and
/// returns a zero-byte read.
///
/// This is asserted rather than merely written down so that a kernel which
/// starts reporting the close reds this test and reopens the question.
#[cfg(target_os = "linux")]
#[test]
fn linux_gives_a_dgram_pair_no_close_signal_at_all() {
let (a, b) = dgram_pair().expect("AF_UNIX SOCK_DGRAM socketpair");
drop(a);
let mut buf = [0u8; 64];
let read = recv(b.as_raw_fd(), &mut buf);
let err = read.as_ref().err().map(|e| e.raw_os_error());
assert_eq!(
err,
Some(Some(libc::EAGAIN)),
"expected the closed peer to be indistinguishable from an idle socket, got {read:?}"
);
assert!(
!hung_up(b.as_raw_fd()),
"POLLHUP is now set on a closed SOCK_DGRAM peer: this kernel has gained \
the close signal Linux 6.8 did not have, and the macOS port's design \
question should be reopened"
);
}
/// The open question, and the only thing a Mac is needed for.
///
/// BSD kernels differ from Linux on datagram close reporting, so Darwin may
/// return `ECONNRESET`, or set `POLLHUP`, where Linux reports nothing. Either
/// would give the receive path something to key on.
///
/// **A failure here is the measurement coming back negative, not a
/// regression.** It says Darwin behaves as Linux does, that a `SOCK_DGRAM`
/// descriptor carries no close signal, and that a macOS port must take the
/// close from the client's line-protocol connection instead of from the flow
/// descriptor.
#[cfg(target_os = "macos")]
#[test]
fn darwin_reports_a_closed_dgram_peer_somehow() {
let (a, b) = dgram_pair().expect("AF_UNIX SOCK_DGRAM socketpair");
drop(a);
let mut buf = [0u8; 64];
let read = recv(b.as_raw_fd(), &mut buf);
let hup = hung_up(b.as_raw_fd());
let signalled = hup
|| matches!(&read, Ok(0))
|| read.as_ref().err().is_some_and(|e| {
e.raw_os_error() != Some(libc::EAGAIN) && e.raw_os_error() != Some(libc::EWOULDBLOCK)
});
assert!(
signalled,
"Darwin reports nothing when a connected SOCK_DGRAM peer closes: \
POLLHUP unset and recv gave {read:?}, which is what an idle socket \
gives. The flow descriptor cannot carry the close on this platform."
);
}
/// Which signal Darwin gives: `ECONNRESET`, and not `POLLHUP`.
///
/// Measured 2026-08-20 on `macos-latest`, run 32353220389, and identical across
/// all three of nextest's attempts, so it is the kernel's behaviour and not a
/// race. The exact reading was `poll` returning 0 with an empty `revents`, and
/// `recv` returning errno 54, `ECONNRESET`.
///
/// This is the opposite of `SOCK_SEQPACKET` on Linux, which sets `POLLHUP` and
/// returns a zero-byte read, and it is why
/// [`super::seqpacket::recv_once`](super::seqpacket) treats `ECONNRESET` as end
/// of file alongside the `POLLHUP` rule rather than choosing between them by
/// platform: one rule that accepts either signal is correct on both kernels.
///
/// Asserted rather than only written down, so that a Darwin release which moves
/// to `POLLHUP`, or stops reporting the close at all, reds this test instead of
/// silently changing what the receive path depends on.
#[cfg(target_os = "macos")]
#[test]
fn darwin_signals_a_closed_dgram_peer_with_econnreset() {
let (a, b) = dgram_pair().expect("AF_UNIX SOCK_DGRAM socketpair");
drop(a);
let mut poll = libc::pollfd {
fd: b.as_raw_fd(),
events: libc::POLLIN,
revents: 0,
};
// SAFETY: `poll` points at one live pollfd and the call cannot block.
let rc = unsafe { libc::poll(&mut poll, 1, 0) };
let mut buf = [0u8; 64];
let read = recv(b.as_raw_fd(), &mut buf);
let errno = read.as_ref().err().and_then(|e| e.raw_os_error());
assert_eq!(
errno,
Some(libc::ECONNRESET),
"Darwin no longer reports a closed SOCK_DGRAM peer as ECONNRESET. \
Measured: poll rc={rc}, revents=0x{:04x}, recv={read:?}. The receive \
path treats ECONNRESET as end of file and would now hang instead.",
poll.revents,
);
assert!(
(poll.revents & libc::POLLHUP) == 0,
"Darwin has gained POLLHUP on a closed SOCK_DGRAM peer, revents=0x{:04x}. \
Nothing breaks, since the receive path accepts either signal, but the \
record here is now wrong and the SOCK_SEQPACKET comparison it rests on \
should be re-read.",
poll.revents,
);
}
/// Whether a `SOCK_DGRAM` pair can carry the largest payload the API offers.
///
/// Darwin bounds a unix-domain datagram with the `net.local.dgram.maxdgram`
/// sysctl, whose default is small, and it is a system tunable rather than
/// something this process can rely on. Linux has no equivalent ceiling on an
/// `AF_UNIX` datagram beyond the socket buffer. The API advertises a payload
/// limit of 1362 bytes to its clients, so a kernel that refuses a datagram that
/// size would break the contract the client was told.
///
/// Written to answer on failure as well as on success: the assertion message
/// carries the largest size that did cross, so a negative result names the
/// actual ceiling rather than only saying the hoped-for one was not reached.
#[test]
fn a_dgram_pair_carries_the_largest_payload_the_api_advertises() {
/// The `max_payload` the API reports to a client, from `super::mod`'s
/// wire-derived limit. Duplicated rather than imported because the constant
/// lives inside a module gated to the platforms that have the listener, and
/// this test runs where that module does not exist.
const ADVERTISED_PAYLOAD: usize = 1362;
let (a, b) = dgram_pair().expect("AF_UNIX SOCK_DGRAM socketpair");
// Walk up rather than testing one size, so a failure reports the ceiling.
let mut largest = 0usize;
let mut buf = vec![0u8; ADVERTISED_PAYLOAD * 4];
for size in [64, 256, 1024, ADVERTISED_PAYLOAD, ADVERTISED_PAYLOAD * 2] {
let payload = vec![0xA5u8; size];
if send(a.as_raw_fd(), &payload).is_err() {
break;
}
match recv(b.as_raw_fd(), &mut buf) {
Ok(n) if n == size => largest = size,
_ => break,
}
}
assert!(
largest >= ADVERTISED_PAYLOAD,
"a SOCK_DGRAM pair carried at most {largest} bytes, below the \
{ADVERTISED_PAYLOAD} the API advertises to clients. On Darwin this is \
the net.local.dgram.maxdgram ceiling and the port has to raise it, or \
lower what it advertises, rather than let a client send what it was \
told it could."
);
}
#[test]
fn a_datagram_queued_before_the_close_is_still_readable() {
let (a, b) = dgram_pair().expect("AF_UNIX SOCK_DGRAM socketpair");
assert_eq!(send(a.as_raw_fd(), &[7, 7, 7]).unwrap(), 3);
drop(a);
// Data sent before the close must survive it. A kernel that discards the
// queue on close would lose a client's last datagram.
let mut buf = [0u8; 64];
assert_eq!(
recv(b.as_raw_fd(), &mut buf).unwrap(),
3,
"a datagram queued before the peer closed must still be delivered"
);
assert_eq!(&buf[..3], &[7, 7, 7]);
}
#[test]
fn an_empty_datagram_queued_before_the_close_is_not_read_as_the_close() {
let (a, b) = dgram_pair().expect("AF_UNIX SOCK_DGRAM socketpair");
assert_eq!(send(a.as_raw_fd(), &[]).unwrap(), 0);
drop(a);
// The ordering case the close rule is weakest against. The peer has closed,
// so POLLHUP is set, and the queued empty datagram also reads as zero
// bytes, so `zero && hung_up` cannot tell them apart. Whatever the kernel
// does here, the implementation has to handle it; this test records which
// kernel loses the datagram.
let mut buf = [0u8; 64];
let read = recv(b.as_raw_fd(), &mut buf);
let hup = hung_up(b.as_raw_fd());
assert!(
matches!(read, Ok(0)),
"expected the queued empty datagram to read as zero bytes, got {read:?}"
);
assert!(
!hup,
"POLLHUP is set while an empty datagram is still queued, so a zero-byte \
read plus POLLHUP cannot mean end of file on this kernel: the queued \
datagram would be swallowed and reported as a close"
);
}
+347
View File
@@ -0,0 +1,347 @@
//! Write a reply on the client connection, optionally carrying a descriptor.
//!
//! Passing a file descriptor between processes is `sendmsg` with an `SCM_RIGHTS`
//! control message, which neither std nor tokio exposes. The descriptor travels
//! in the ancillary data of the same `sendmsg` that carries the reply line, so a
//! client reads one message and gets both, with no window in which it holds one
//! without the other.
//!
//! Every write goes through `try_io`, including the ones with no descriptor, so
//! the connection's `UnixStream` is only ever borrowed shared. That is what lets
//! the reader keep the stream inside a `BufReader` while replies are written
//! through `get_ref`.
//!
//! The receiving half lives here too, in `recv`. It is blocking and uses no
//! tokio, because a client process is what runs it, but it is the same concern
//! read backwards and it needs the same control message sizing.
use std::io;
use std::mem;
use std::os::fd::{AsRawFd, BorrowedFd, FromRawFd, OwnedFd, RawFd};
use tokio::io::Interest;
use tokio::net::UnixStream;
/// Space for the control message, sized at runtime and checked against this.
///
/// `CMSG_SPACE(4)` is 24 bytes on Linux and 16 on Darwin, whose `cmsghdr` is
/// 12 bytes and whose alignment is 4 rather than 8. The array is `u64` so it
/// carries the alignment `cmsghdr` requires, and at 64 bytes is larger than
/// either, so a platform with a wider header is caught by the assertion rather
/// than by memory corruption. Nothing computes from the number: the send path
/// checks the runtime `CMSG_SPACE` against this buffer's size and the receive
/// path offers the whole buffer.
type CmsgBuf = [u64; 8];
/// `recvmsg` flags that make a received descriptor close-on-exec.
///
/// Linux and FreeBSD do it atomically with `MSG_CMSG_CLOEXEC`, which is the
/// only way to be certain no `fork` in another thread wins the race. macOS has
/// no equivalent flag, so there is nothing to pass and [`recv`] sets
/// `FD_CLOEXEC` on each descriptor afterwards instead.
#[cfg(not(target_os = "macos"))]
const RECV_FLAGS: libc::c_int = libc::MSG_CMSG_CLOEXEC;
#[cfg(target_os = "macos")]
const RECV_FLAGS: libc::c_int = 0;
/// Send `line` on `stream`, with `fd` in the ancillary data when given.
///
/// The whole reply goes in one datagram-shaped `sendmsg`. A short write is
/// treated as an error rather than retried: the replies are a few hundred bytes
/// into an empty socket buffer, so a partial one means something is wrong that a
/// retry loop would hide.
pub async fn reply(stream: &UnixStream, line: &[u8], fd: Option<BorrowedFd<'_>>) -> io::Result<()> {
loop {
stream.writable().await?;
match stream.try_io(Interest::WRITABLE, || {
send_once(stream.as_raw_fd(), line, fd)
}) {
Ok(written) if written == line.len() => return Ok(()),
Ok(written) => {
return Err(io::Error::other(format!(
"native API reply truncated: wrote {written} of {}",
line.len()
)));
}
Err(error) if error.kind() == io::ErrorKind::WouldBlock => continue,
Err(error) => return Err(error),
}
}
}
/// One `sendmsg` that reports a full send buffer instead of waiting for one.
///
/// The listener's task writes onto socket pairs whose client half it has not
/// handed over yet, so no process can read either one and a task that waited
/// would stop serving that listener for good. [`Seqpacket::send`] is the wrong
/// tool for exactly that reason: it treats a would-block as a reason to await
/// readiness rather than as an answer. This is the raw syscall on a descriptor
/// the reactor has already made non-blocking, so a full buffer comes back as
/// `WouldBlock` and the caller decides.
///
/// [`Seqpacket::send`]: super::seqpacket::Seqpacket::send
pub(super) fn try_send(sock: RawFd, line: &[u8], fd: Option<BorrowedFd<'_>>) -> io::Result<usize> {
send_once(sock, line, fd)
}
/// One `sendmsg`, with or without an `SCM_RIGHTS` control message.
///
/// Visible within [`super`] so the client's tests can play daemon through the
/// real ancillary framing rather than an imitation of it.
pub(super) fn send_once(sock: i32, line: &[u8], fd: Option<BorrowedFd<'_>>) -> io::Result<usize> {
let mut iov = libc::iovec {
iov_base: line.as_ptr() as *mut libc::c_void,
iov_len: line.len(),
};
// SAFETY: msghdr is a plain C struct with no invalid bit patterns; every
// field this call reads is set below.
let mut msg: libc::msghdr = unsafe { mem::zeroed() };
msg.msg_iov = &mut iov;
msg.msg_iovlen = 1;
let mut control: CmsgBuf = [0; 8];
if let Some(fd) = fd {
// SAFETY: CMSG_SPACE is a pure size computation over its argument.
let space = unsafe { libc::CMSG_SPACE(mem::size_of::<libc::c_int>() as u32) } as usize;
if space > mem::size_of::<CmsgBuf>() {
return Err(io::Error::other(
"control message buffer too small for a descriptor",
));
}
msg.msg_control = control.as_mut_ptr().cast();
msg.msg_controllen = space as _;
// SAFETY: msg_control points at `control`, which is aligned for
// cmsghdr and at least `space` bytes long, so the header the kernel
// macro returns lies inside it.
unsafe {
let header = libc::CMSG_FIRSTHDR(&msg);
if header.is_null() {
return Err(io::Error::other("control message header unavailable"));
}
(*header).cmsg_level = libc::SOL_SOCKET;
(*header).cmsg_type = libc::SCM_RIGHTS;
(*header).cmsg_len = libc::CMSG_LEN(mem::size_of::<libc::c_int>() as u32) as _;
let raw = fd.as_raw_fd();
std::ptr::copy_nonoverlapping(
std::ptr::addr_of!(raw).cast::<u8>(),
libc::CMSG_DATA(header),
mem::size_of::<libc::c_int>(),
);
}
}
// SAFETY: `sock` is the connection's open descriptor, and `msg` describes
// buffers that outlive this call.
let sent = unsafe { libc::sendmsg(sock, &msg, libc::MSG_NOSIGNAL) };
if sent < 0 {
return Err(io::Error::last_os_error());
}
Ok(sent as usize)
}
/// What one `recvmsg` on a client's RPC connection produced.
///
/// The bytes and the descriptor are reported together because they arrive
/// together. Handing them back separately would reopen the window this module
/// exists to close.
///
/// Visible within [`super`] only, like [`send_once`]: the client module is the
/// one caller, and the API this crate publishes is `FipsAddr`, `FipsStream` and
/// `FipsListener` rather than the framing underneath them.
#[derive(Debug)]
pub(super) struct Chunk {
/// How many bytes landed in the caller's buffer. Zero is end of file.
pub(super) len: usize,
/// The descriptor the message carried, where it carried one.
pub(super) fd: Option<OwnedFd>,
}
/// Receive one message into `buf`, keeping any descriptor that came with it.
///
/// Blocking and `libc`-only, with no tokio: this is the half a client process
/// runs. It lives beside [`reply`] because the two share the control message
/// sizing and the same safety argument.
///
/// **Every read on the RPC connection must come through here.** A plain `read`
/// consumes a descriptor-bearing message's bytes with no ancillary buffer, and
/// the kernel closes the descriptor rather than queueing it, so the reply looks
/// right and the flow is silently gone.
///
/// A received descriptor is kept out of a child the client forks later, by
/// [`RECV_FLAGS`] where the platform has a flag for it and by an `fcntl` on
/// each descriptor where it does not. `EINTR` is retried, because a signal
/// delivered during the wait says nothing about the connection.
pub(super) fn recv(sock: RawFd, buf: &mut [u8]) -> io::Result<Chunk> {
loop {
let mut iov = libc::iovec {
iov_base: buf.as_mut_ptr().cast(),
iov_len: buf.len(),
};
let mut control: CmsgBuf = [0; 8];
// SAFETY: as in send_once; every field this call reads is set below.
let mut msg: libc::msghdr = unsafe { mem::zeroed() };
msg.msg_iov = &mut iov;
msg.msg_iovlen = 1;
msg.msg_control = control.as_mut_ptr().cast();
msg.msg_controllen = mem::size_of::<CmsgBuf>() as _;
// SAFETY: `sock` is the caller's open socket, and `msg` describes
// buffers that outlive the call.
let received = unsafe { libc::recvmsg(sock, &mut msg, RECV_FLAGS) };
if received < 0 {
let error = io::Error::last_os_error();
if error.kind() == io::ErrorKind::Interrupted {
continue;
}
// A listener's descriptor is a socket pair half, and Darwin reports
// its closed peer as ECONNRESET where Linux returns a zero-byte
// message. They are the same event, so it is reported as the empty
// chunk every caller here already reads as the far end going away.
// Doing it here rather than in each caller keeps `accept`'s
// documented EPIPE true on both platforms.
if error.raw_os_error() == Some(libc::ECONNRESET) {
return Ok(Chunk { len: 0, fd: None });
}
return Err(error);
}
// SAFETY: recvmsg succeeded, so it filled `msg_control` within
// `msg_controllen`, and `control` is still alive.
let mut fds = unsafe { take_fds(&msg) };
// Darwin has no `MSG_CMSG_CLOEXEC`, so the flag is set here instead.
// Later than the atomic form and with the same window `seqpacket::pair`
// documents: a concurrent `fork` and `exec` in these few instructions
// would inherit the descriptor. A failure to set it is reported rather
// than ignored, because the descriptor is live either way and the
// caller must not be told the receive was clean.
#[cfg(target_os = "macos")]
for fd in &fds {
super::seqpacket::set_cloexec(fd.as_raw_fd())?;
}
// Whatever did arrive is taken before the truncation check, so nothing
// leaks on that path: dropping an `OwnedFd` closes it. The connection
// cannot continue either way, because a descriptor the kernel dropped
// is one no later read can recover.
if (msg.msg_flags & libc::MSG_CTRUNC) != 0 {
return Err(io::Error::other(
"native API control message truncated: a descriptor was lost",
));
}
// More than one descriptor is not something this protocol sends. The
// extras are dropped, and so closed, rather than leaked.
let fd = if fds.is_empty() {
None
} else {
Some(fds.swap_remove(0))
};
return Ok(Chunk {
len: received as usize,
fd,
});
}
}
/// Collect every descriptor an `SCM_RIGHTS` control message carried.
///
/// The whole control buffer is walked rather than only its first header: a
/// reader that took `CMSG_FIRSTHDR` alone would leak any descriptor behind it.
///
/// # Safety
///
/// `msg` must be a `msghdr` that a successful `recvmsg` filled in, whose
/// `msg_control` buffer is still live and unmodified since.
unsafe fn take_fds(msg: &libc::msghdr) -> Vec<OwnedFd> {
let mut fds = Vec::new();
// SAFETY: the caller guarantees `msg` came from a successful recvmsg, so
// every header these macros return lies inside its control buffer, and
// every descriptor named there is one this process now owns.
unsafe {
let mut header = libc::CMSG_FIRSTHDR(msg);
while !header.is_null() {
if (*header).cmsg_level == libc::SOL_SOCKET && (*header).cmsg_type == libc::SCM_RIGHTS {
let payload = (*header).cmsg_len as usize - libc::CMSG_LEN(0) as usize;
for index in 0..payload / mem::size_of::<libc::c_int>() {
let mut raw: libc::c_int = 0;
std::ptr::copy_nonoverlapping(
libc::CMSG_DATA(header).add(index * mem::size_of::<libc::c_int>()),
std::ptr::addr_of_mut!(raw).cast::<u8>(),
mem::size_of::<libc::c_int>(),
);
fds.push(OwnedFd::from_raw_fd(raw));
}
}
header = libc::CMSG_NXTHDR(msg, header);
}
}
fds
}
#[cfg(test)]
mod tests {
use super::*;
use crate::native::seqpacket::{Received, Seqpacket, pair};
use std::io::Write;
use std::os::fd::AsFd;
use std::os::unix::net::UnixStream as StdUnixStream;
/// Receive one line plus an optional descriptor, the way a client does.
///
/// Goes through [`recv`] rather than repeating its `recvmsg`, so these
/// tests exercise the receiving half a client actually runs.
fn recv_with_fd(sock: &StdUnixStream) -> (Vec<u8>, Option<OwnedFd>) {
let mut buf = [0u8; 4096];
let chunk = recv(sock.as_raw_fd(), &mut buf).expect("recvmsg should succeed");
(buf[..chunk.len].to_vec(), chunk.fd)
}
#[tokio::test]
async fn a_reply_without_a_descriptor_arrives_whole() {
let (ours, theirs) = StdUnixStream::pair().unwrap();
ours.set_nonblocking(true).unwrap();
let ours = UnixStream::from_std(ours).unwrap();
reply(&ours, b"{\"status\":\"ok\"}\n", None).await.unwrap();
let (line, fd) = recv_with_fd(&theirs);
assert_eq!(line, b"{\"status\":\"ok\"}\n");
assert!(fd.is_none());
}
#[tokio::test]
async fn a_descriptor_arrives_with_its_reply_and_still_works() {
let (ours, theirs) = StdUnixStream::pair().unwrap();
ours.set_nonblocking(true).unwrap();
let ours = UnixStream::from_std(ours).unwrap();
let (daemon_half, client_half) = pair().unwrap();
let daemon_half = Seqpacket::new(daemon_half).unwrap();
reply(&ours, b"{\"status\":\"ok\"}\n", Some(client_half.as_fd()))
.await
.unwrap();
// The daemon closes its copy once it is sent; the receiver holds the
// only remaining reference to the client half.
drop(client_half);
let (line, fd) = recv_with_fd(&theirs);
assert_eq!(line, b"{\"status\":\"ok\"}\n");
let fd = fd.expect("a descriptor should have arrived");
// The passed descriptor is a live half of the flow, not merely a
// number: writing on it must reach the daemon's side.
let mut received = StdUnixStream::from(fd);
received.write_all(b"through the passed fd").unwrap();
let mut buf = [0u8; 64];
assert_eq!(
daemon_half.recv(&mut buf).await.unwrap(),
Received::Datagram(21)
);
assert_eq!(&buf[..21], b"through the passed fd");
}
}
+393
View File
@@ -0,0 +1,393 @@
//! What a native API client task asks the node to do.
//!
//! The registry lives inside `Node`, so a client task cannot touch it directly.
//! It sends one of these instead and waits on the `oneshot` it carried. That is
//! the same shape the control socket uses, and it is why the receive path needs
//! no lock: only the `rx_loop` ever holds the registry.
//!
//! Every variant that can fail carries its reply channel. A dropped reply means
//! the node is shutting down, which the client task reports as such rather than
//! waiting.
use super::registry::{Arrival, Datagram, Delivery, DropCause, FlowKey, Registry, RegistryError};
use crate::identity::NodeAddr;
use secp256k1::XOnlyPublicKey;
use tokio::sync::{mpsc, oneshot};
use tracing::trace;
/// A request from a client task to the node's registry.
#[derive(Debug)]
pub enum NativeMessage {
/// Bind a listener to a local port, or to an ephemeral one.
Listen {
/// The port to hold, or `None` for an ephemeral one.
port: Option<u16>,
/// Where the node announces new peers on that port.
arrivals: mpsc::Sender<Arrival>,
/// The port actually held, or why none could be.
reply: oneshot::Sender<Result<u16, RegistryError>>,
},
/// Open a flow to a peer.
Connect {
/// The far end, by the x-only public key that is its address.
peer: XOnlyPublicKey,
/// The far end's port.
remote: u16,
/// The local port, or `None` for an ephemeral one.
local: Option<u16>,
/// Where the node delivers this flow's datagrams.
sink: mpsc::Sender<Datagram>,
/// What the flow holds, or why it could not be opened.
reply: oneshot::Sender<Result<Opened, RegistryError>>,
},
/// Take a flow a listener announced.
Accept {
/// Which announced flow.
flow: u64,
/// Where the node delivers its datagrams from now on.
sink: mpsc::Sender<Datagram>,
/// The flow's key, whatever arrived before it was accepted, and the
/// payload limit.
reply: oneshot::Sender<Result<Accepted, RegistryError>>,
},
/// Give up a flow whose descriptor closed, or a listener whose port is
/// being unbound.
///
/// Carries no reply: the sender is a task that is ending and has nothing
/// left to do with the answer. It is sent from the task rather than from a
/// `Drop`, so it can be awaited and cannot be silently lost to a full
/// channel.
Release {
/// Flows to forget.
flows: Vec<FlowKey>,
/// Listener ports to free, along with anything pending on them.
listeners: Vec<u16>,
},
/// Undo a flow the listener's task promoted but could not hand over.
///
/// Separate from [`NativeMessage::Release`] because the two count
/// differently: a release is a flow a client finished with, and this is one
/// no client ever held. Sending one message for both halves of the undo
/// keeps the registry entry and the counter from disagreeing.
Discard {
/// The flow to forget.
key: FlowKey,
/// Why it could not be handed over.
reason: DropReason,
},
/// **Debug.** Deliver a datagram as though it had arrived from the mesh.
///
/// This drives the same [`Registry::deliver`](super::registry::Registry::deliver)
/// the FSP receive path will call, so the dispatch rule is exercised before
/// the wire exists and the wire, when it lands, changes the caller rather
/// than the rule.
Arrive {
/// The peer it appears to come from, by wire address.
peer: NodeAddr,
/// That peer's key, decoded from the npub the caller named. Client
/// asserted rather than authenticated, which is one of the reasons the
/// command is gated.
pubkey: XOnlyPublicKey,
/// Its source port.
src: u16,
/// Its destination port on this node.
dst: u16,
/// The payload.
data: Datagram,
/// What the registry decided to do with it.
reply: oneshot::Sender<Outcome>,
},
}
/// One datagram a client wrote to its descriptor, on its way to the mesh.
///
/// Travels on its own channel rather than through [`NativeMessage`], so a burst
/// of client traffic cannot delay a registration and the data arm can drain in
/// batches the way the TUN arm does.
#[derive(Debug)]
pub struct Outbound {
/// The flow it belongs to, which carries both ports and the destination.
pub key: FlowKey,
/// The destination's address. Always known: a connected flow decoded it
/// from the npub its client named, and an accepted one took it from the
/// session that authenticated the peer.
pub peer: XOnlyPublicKey,
/// The payload, with no port header: the send path adds that.
pub payload: Datagram,
}
/// A flow the client opened.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub struct Opened {
/// The local port it holds.
pub local: u16,
/// The identifier the client names it by.
pub flow: u64,
/// The largest payload it may send, in bytes.
pub max: u16,
}
/// A flow the client accepted from a listener.
#[derive(Debug)]
pub struct Accepted {
/// Which flow, in both directions.
pub key: FlowKey,
/// The peer's address.
pub peer: XOnlyPublicKey,
/// Whatever arrived before the client answered.
pub held: Vec<Datagram>,
/// The largest payload it may send, in bytes.
pub max: u16,
}
/// Why a datagram was not delivered, across every path that can refuse one.
///
/// [`DropCause`] covers the refusals the registry decides. Delivery can also
/// fail after the registry has agreed, when a bounded channel to a client is
/// full, and those three cases are the remaining variants. Keeping one type
/// over the whole set is what lets a counter match be exhaustive: a match over
/// `DropCause` alone would silently miss the case a slow client actually causes.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum DropReason {
/// No listener and no flow holds the destination port.
NoPort,
/// A listener holds the port but will not hold another pending flow.
BacklogFull,
/// The node is at its flow ceiling.
TooManyFlows,
/// A flow awaiting accept will hold no more datagrams.
PendingQueueFull,
/// An established flow's client is not draining its descriptor.
FlowQueueFull,
/// A listener's client is not reading the arrivals it asked for. The flow
/// is unregistered as well, so nothing is left pending that nobody knows of.
ArrivalQueueFull,
/// A listener's client is not reading its descriptor, so the arrival could
/// not be written to it. Distinct from `ArrivalQueueFull`: that one is the
/// rx_loop refusing before anything was opened, and this one is a flow the
/// daemon had already wired and has to take apart again.
ListenerNotReading,
/// A listener's client closed its descriptor between the arrival being
/// taken off the queue and being written to it. The same cleanup as
/// `ListenerNotReading` and a different counter: this one is a race a
/// healthy client can lose, and that one is a client falling behind.
ListenerGone,
}
impl From<DropCause> for DropReason {
fn from(cause: DropCause) -> Self {
match cause {
DropCause::NoPort => DropReason::NoPort,
DropCause::BacklogFull => DropReason::BacklogFull,
DropCause::TooManyFlows => DropReason::TooManyFlows,
DropCause::QueueFull => DropReason::PendingQueueFull,
}
}
}
impl DropReason {
/// The client-facing text for this reason.
///
/// `PendingQueueFull` and `FlowQueueFull` deliberately render alike. They
/// are distinct to a counter and indistinguishable to a client, which is
/// what keeps the strings the debug `arrive` command already answers with
/// unchanged.
pub fn as_str(self) -> &'static str {
match self {
DropReason::NoPort => "no listener or flow on that port",
DropReason::BacklogFull => "listener backlog full",
DropReason::TooManyFlows => "node flow ceiling reached",
DropReason::PendingQueueFull | DropReason::FlowQueueFull => "queue full",
DropReason::ArrivalQueueFull => "arrival queue full",
DropReason::ListenerNotReading => "listener not reading arrivals",
DropReason::ListenerGone => "listener closed",
}
}
}
/// What became of a delivered datagram.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum Outcome {
/// An established flow received it.
Delivered,
/// A listener was told a new peer arrived, and the datagram was held for
/// whoever accepts it.
Announced(u64),
/// A flow already announced held it while it waits to be accepted.
Held(u64),
/// Nothing took it.
Dropped(DropReason),
}
/// What one served request did, so the shell can count it.
///
/// The core decides and answers the client on the request's own channel; this
/// exists only so the node can bump a counter without the core having to know
/// what a counter is. Only outcomes something counts are named; everything else
/// is [`Served::Untracked`].
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum Served {
/// A client opened a flow to a peer it named.
Opened,
/// A listener's task took a pending flow on a client's behalf.
Accepted,
/// A flow the daemon wired and could not hand over was taken apart again.
Discarded(DropReason),
/// This many established flows were given back when their descriptors
/// closed or their listener unbound.
Released(usize),
/// A datagram of this many bytes was dispatched by the debug arrival
/// command, which runs the same rule the wire does and is counted the same
/// way. The length rides along because `deliver` consumes the datagram.
Delivered(Outcome, usize),
/// Nothing counted: a listen, or a request the registry refused.
Untracked,
}
/// Apply one request to `registry` and answer it.
///
/// A free function over the registry rather than a method on the node: the node
/// supplies only `now`, and everything else here is a decision plus the sends
/// that decision names. That is what lets this be driven in a test without
/// building a node, and it is the same code the running daemon executes.
pub fn serve(registry: &mut Registry, message: NativeMessage, now: u64, max: u16) -> Served {
match message {
NativeMessage::Listen {
port,
arrivals,
reply,
} => {
let _ = reply.send(registry.listen(port, arrivals));
Served::Untracked
}
NativeMessage::Discard { key, reason } => {
registry.release(&key);
Served::Discarded(reason)
}
NativeMessage::Connect {
peer,
remote,
local,
sink,
reply,
} => {
let opened = registry
.connect(peer, remote, local, sink, now)
.map(|(local, flow)| Opened { local, flow, max });
let served = if opened.is_ok() {
Served::Opened
} else {
Served::Untracked
};
let _ = reply.send(opened);
served
}
NativeMessage::Accept { flow, sink, reply } => {
let accepted = registry
.accept(flow, sink, now)
.map(|(key, peer, held)| Accepted {
key,
peer,
held,
max,
});
let served = if accepted.is_ok() {
Served::Accepted
} else {
Served::Untracked
};
let _ = reply.send(accepted);
served
}
NativeMessage::Release { flows, listeners } => {
let mut closed = 0;
for key in &flows {
if registry.release(key) {
closed += 1;
}
}
for port in &listeners {
registry.release_listener(*port);
}
Served::Released(closed)
}
NativeMessage::Arrive {
peer,
pubkey,
src,
dst,
data,
reply,
} => {
let bytes = data.len();
let outcome = deliver(registry, peer, pubkey, src, dst, data, now);
let _ = reply.send(outcome);
Served::Delivered(outcome, bytes)
}
}
}
/// Deliver one inbound datagram to whatever owns its destination port.
///
/// The FSP receive path calls this once the wire is connected; until then the
/// debug arrival command is its only caller. Either way the decision is the
/// registry's and this only performs it.
pub fn deliver(
registry: &mut Registry,
peer: NodeAddr,
pubkey: XOnlyPublicKey,
src: u16,
dst: u16,
data: Datagram,
now: u64,
) -> Outcome {
match registry.deliver(peer, pubkey, src, dst, now) {
Delivery::Flow(sink) => {
// Never block the receive path on a slow client. A full queue costs
// that client a datagram, not the node its tick.
match sink.try_send(data) {
Ok(()) => Outcome::Delivered,
Err(_) => {
trace!(dst, "Native API flow queue full, dropping datagram");
Outcome::Dropped(DropReason::FlowQueueFull)
}
}
}
Delivery::Arrived(arrivals, arrival) => {
let flow = arrival.flow;
if arrivals.try_send(arrival).is_err() {
// The client is not reading its own announcements. Undo the
// registration rather than leaving a pending flow nobody will
// ever be told about.
let _ = registry.reject(flow);
return Outcome::Dropped(DropReason::ArrivalQueueFull);
}
if registry.hold(flow, data) {
Outcome::Announced(flow)
} else {
Outcome::Dropped(DropReason::PendingQueueFull)
}
}
Delivery::Pending(flow) => {
if registry.hold(flow, data) {
Outcome::Held(flow)
} else {
Outcome::Dropped(DropReason::PendingQueueFull)
}
}
Delivery::Drop(cause) => Outcome::Dropped(cause.into()),
}
}
+2244
View File
File diff suppressed because it is too large Load Diff
+566
View File
@@ -0,0 +1,566 @@
//! Native datagram API command types and the pure decisions taken over them.
//!
//! No I/O and no node state. The shell in [`super`] reads a line, calls
//! [`parse`], and acts on the typed command. Every rule that can be decided
//! from the request alone is decided here, so it is testable without a socket
//! and without a running node.
//!
//! The encoding is the control socket's: one JSON object per line, shaped
//! `{"command": "...", "params": {...}}`, answered by a
//! [`Response`](crate::control::protocol::Response). The two sockets are
//! separate listeners with separate lifetimes, but there is no reason for a
//! client to learn two envelopes.
use super::registry::RegistryError;
use crate::control::protocol::Request;
use serde::Deserialize;
use thiserror::Error;
/// Highest port reserved for protocol use, per the FSP port registry.
pub const PORT_PROTOCOL_MAX: u16 = 255;
/// Highest port reserved for FIPS standard services. Port 256 in this range is
/// the IPv6 shim ([`FSP_PORT_IPV6_SHIM`](crate::proto::fsp::FSP_PORT_IPV6_SHIM)).
pub const PORT_STANDARD_MAX: u16 = 1023;
/// Lowest port the daemon hands out when a client names no local port.
///
/// Below this the application range is available for an explicit bind, which is
/// what lets a service hold a port a peer can be told about in advance.
pub const PORT_EPHEMERAL_MIN: u16 = 49152;
/// A parsed native API command.
#[derive(Debug, Clone, PartialEq, Eq)]
pub enum Command {
/// Open a flow to a named peer.
Connect(Connect),
/// Receive flows from any peer on a local port.
Listen(Listen),
/// **Debug, gated.** Make the daemon write bytes the client chose into one
/// of that client's flows, so the receive direction is exercisable without
/// a peer.
Inject(Inject),
/// **Debug, gated.** Report what the daemon has received on a flow.
Stats(u64),
/// **Debug, gated.** Deliver a datagram as though it had arrived from the
/// mesh, which reaches any listener this node holds.
Arrive(Arrive),
}
impl Command {
/// The name a client used, when the command is one of the debug three.
///
/// The three are not a supported interface: they exist for the test
/// harness, and a node answers them only where
/// `node.native_api.debug_commands` is on. Whether that key is on is node
/// configuration and so is not decidable here; classifying the command is,
/// which is the half this module owns.
pub fn debug_name(&self) -> Option<&'static str> {
match self {
Command::Inject(_) => Some("inject"),
Command::Stats(_) => Some("stats"),
Command::Arrive(_) => Some("arrive"),
Command::Connect(_) | Command::Listen(_) => None,
}
}
}
/// Parameters of a `connect` command, after validation.
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct Connect {
/// The far end, as the npub the client supplied. Decoding it to a key is
/// the shell's job: this module holds no identity types.
pub peer: String,
/// The far end's port.
pub remote: u16,
/// The local port, or `None` to let the daemon allocate an ephemeral one.
pub local: Option<u16>,
}
/// Parameters of a `listen` command, after validation.
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct Listen {
/// The local port to receive on, or `None` to let the daemon allocate an
/// ephemeral one. Port 0 and an absent field mean the same thing, which is
/// what `bind(2)` with port 0 means and what `connect` already accepted.
pub local: Option<u16>,
}
/// Parameters of the debug `arrive` command, after validation.
///
/// This drives the node's inbound dispatch without a wire, so the rule that
/// decides between an established flow, a listener and a drop is exercised
/// before FSP is involved. The wire is a second caller of the same rule rather
/// than a new one, which is why the command outlived the work that first
/// needed it; it stays behind `node.native_api.debug_commands` because a caller that
/// reaches it can deliver to any listener on this node under any peer identity
/// it names.
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct Arrive {
/// The peer it appears to come from, as an npub.
pub peer: String,
/// Its source port.
pub src: u16,
/// Its destination port on this node.
pub dst: u16,
/// The payload, decoded from hex.
pub data: Vec<u8>,
}
/// Parameters of the debug `inject` command, after validation.
///
/// This command exists so the receive direction is testable without a peer
/// sending anything. It can only write into a flow the same connection opened,
/// so it grants that client nothing it does not already have, but it makes the
/// daemon emit bytes the client chose on a path a reader cannot distinguish
/// from the mesh. It is therefore gated on
/// `node.native_api.debug_commands` and off in a packaged node.
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct Inject {
/// Which flow to write into.
pub flow: u64,
/// The bytes to write, decoded from the hex the client sent. Hex rather
/// than text so a check can assert byte fidelity, including bytes that are
/// not valid UTF-8.
pub data: Vec<u8>,
/// How many separate datagrams to write. Sending one payload several times
/// is how a check observes that message boundaries survive.
pub repeat: u32,
}
/// Why a client's command was refused.
///
/// Each variant is a distinct refusal a client can act on, rather than one
/// string it would have to parse.
#[derive(Debug, Clone, PartialEq, Eq, Error)]
pub enum CommandError {
/// The command name is not one this API serves.
#[error("unknown command '{0}'")]
Unknown(String),
/// The command needs a `params` object and none was supplied.
#[error("command '{0}' requires params")]
NoParams(&'static str),
/// The `params` object did not match the command's shape.
#[error("invalid params for '{command}': {reason}")]
BadParams {
/// The command whose params failed to parse.
command: &'static str,
/// What serde objected to.
reason: String,
},
/// The port is in the range the protocol itself reserves.
#[error("port {0} is reserved for protocol use")]
PortProtocol(u16),
/// The port is in the range reserved for FIPS standard services, which is
/// where the IPv6 shim lives.
#[error("port {0} is reserved for FIPS standard services")]
PortStandard(u16),
/// A hex field did not decode.
#[error("invalid hex in '{field}': {reason}")]
BadHex {
/// The field that failed to decode.
field: &'static str,
/// What the decoder objected to.
reason: String,
},
/// A repeat count outside what the command will do in one call.
#[error("repeat must be between 1 and {max}, got {got}")]
BadRepeat {
/// The largest accepted value.
max: u32,
/// What the client asked for.
got: u32,
},
}
/// Largest number of datagrams one `inject` writes.
///
/// Bounded because the command runs inline on the connection task: a client
/// asking for a million would hold that task for the duration.
pub const MAX_INJECT_REPEAT: u32 = 64;
/// Raw `connect` parameters, before validation.
#[derive(Debug, Deserialize)]
struct ConnectParams {
peer: String,
remote_port: u16,
#[serde(default)]
local_port: Option<u16>,
}
/// Raw `listen` parameters, before validation.
#[derive(Debug, Deserialize)]
struct ListenParams {
#[serde(default)]
local_port: Option<u16>,
}
/// Raw `stats` parameters, before validation.
#[derive(Debug, Deserialize)]
struct FlowParams {
flow_id: u64,
}
/// Raw `arrive` parameters, before validation.
#[derive(Debug, Deserialize)]
struct ArriveParams {
peer: String,
src_port: u16,
dst_port: u16,
data: String,
}
/// Raw `inject` parameters, before validation.
#[derive(Debug, Deserialize)]
struct InjectParams {
flow_id: u64,
data: String,
#[serde(default = "InjectParams::one")]
repeat: u32,
}
impl InjectParams {
/// One datagram, when the client names no repeat count.
fn one() -> u32 {
1
}
}
/// Turn a request into a typed command, refusing anything a client may not ask
/// for.
///
/// Port policy is applied here rather than at the registry, so a refusal costs
/// no node state and reads the same whether or not a node is running.
pub fn parse(request: &Request) -> Result<Command, CommandError> {
match request.command.as_str() {
"connect" => {
let params: ConnectParams = take(request, "connect")?;
check_bindable(params.remote_port)?;
let local = match params.local_port {
// Port 0 is the UDP convention for "any port", and a client
// that reaches for it means the same thing as omitting the
// field. Accepting both spellings costs one branch and saves a
// refusal a client would find surprising.
None | Some(0) => None,
Some(port) => {
check_bindable(port)?;
Some(port)
}
};
Ok(Command::Connect(Connect {
peer: params.peer,
remote: params.remote_port,
local,
}))
}
"listen" => {
let params: ListenParams = take(request, "listen")?;
// The same rule `connect` applies to its local port, so one
// spelling of "any port" covers both commands.
let local = match params.local_port {
None | Some(0) => None,
Some(port) => {
check_bindable(port)?;
Some(port)
}
};
Ok(Command::Listen(Listen { local }))
}
"stats" => Ok(Command::Stats(
take::<FlowParams>(request, "stats")?.flow_id,
)),
"arrive" => {
let params: ArriveParams = take(request, "arrive")?;
let data = hex::decode(&params.data).map_err(|error| CommandError::BadHex {
field: "data",
reason: error.to_string(),
})?;
Ok(Command::Arrive(Arrive {
peer: params.peer,
src: params.src_port,
dst: params.dst_port,
data,
}))
}
"inject" => {
let params: InjectParams = take(request, "inject")?;
if params.repeat == 0 || params.repeat > MAX_INJECT_REPEAT {
return Err(CommandError::BadRepeat {
max: MAX_INJECT_REPEAT,
got: params.repeat,
});
}
let data = hex::decode(&params.data).map_err(|error| CommandError::BadHex {
field: "data",
reason: error.to_string(),
})?;
Ok(Command::Inject(Inject {
flow: params.flow_id,
data,
repeat: params.repeat,
}))
}
other => Err(CommandError::Unknown(other.to_string())),
}
}
/// Deserialize a command's `params` object, naming the command in any error.
fn take<T: for<'de> Deserialize<'de>>(
request: &Request,
command: &'static str,
) -> Result<T, CommandError> {
let params = request
.params
.as_ref()
.ok_or(CommandError::NoParams(command))?;
serde_json::from_value(params.clone()).map_err(|error| CommandError::BadParams {
command,
reason: error.to_string(),
})
}
/// Decide whether a client may name `port`.
///
/// The tiers come from the FSP port registry, which describes what already
/// ships on the v1 wire: 0-255 is protocol use, and 256-1023 is reserved for
/// FIPS standard services, of which 256 is the IPv6 shim. Only the application
/// range above them is a client's to use.
pub fn check_bindable(port: u16) -> Result<(), CommandError> {
if port <= PORT_PROTOCOL_MAX {
return Err(CommandError::PortProtocol(port));
}
if port <= PORT_STANDARD_MAX {
return Err(CommandError::PortStandard(port));
}
Ok(())
}
/// The `errno` name a refused setup command reports to its client.
///
/// A client must never match on the message: that is for an operator reading a
/// log, and the codes are the contract. The match is exhaustive over
/// [`RegistryError`] and over the [`CommandError`] its `Port` variant wraps, so
/// a variant added without a code is a compile error rather than a silent
/// `ECONNREFUSED`.
///
/// The names rather than the numbers, because the number belongs to the
/// platform the client is built for and this crate is not it. The client turns
/// the name into `io::Error::from_raw_os_error(libc::EADDRINUSE)` and so on.
pub fn errno(error: &RegistryError) -> &'static str {
match error {
// One port, one owner, and one key, one flow. Both are the address a
// caller asked for and cannot have.
RegistryError::PortTaken(_) | RegistryError::FlowTaken { .. } => "EADDRINUSE",
RegistryError::NoPort => "EADDRNOTAVAIL",
RegistryError::TooManyFlows(_) => "EMFILE",
// Neither reaches a setup reply today: a full backlog is decided on the
// receive path and counted rather than answered, and a missing pending
// flow is a hand-off the daemon lost to its own deadline. If either ever
// does reach a client, it is the daemon refusing for a reason of its
// own, which is what the catch-all names.
RegistryError::BacklogFull(_) | RegistryError::NoPending(_) => "ECONNREFUSED",
RegistryError::Port(command) => command_errno(command),
}
}
/// The `errno` name for a command this API refused before any node state was
/// touched.
///
/// A reserved port is `EADDRNOTAVAIL` rather than `EACCES`: `bind(2)` gives
/// `EACCES` below 1024 because that is a privilege question, and this is not
/// one. No client, however privileged, may hold port 256, because the IPv6 shim
/// has it. Everything else here is a malformed command, which is `EINVAL`.
pub fn command_errno(error: &CommandError) -> &'static str {
match error {
CommandError::PortProtocol(_) | CommandError::PortStandard(_) => "EADDRNOTAVAIL",
CommandError::Unknown(_)
| CommandError::NoParams(_)
| CommandError::BadParams { .. }
| CommandError::BadHex { .. }
| CommandError::BadRepeat { .. } => "EINVAL",
}
}
#[cfg(test)]
mod tests {
use super::*;
use serde_json::json;
/// Build a request the way the line reader would, from JSON text.
fn request(text: &str) -> Request {
serde_json::from_str(text).expect("test request should parse")
}
#[test]
fn connect_carries_the_peer_and_both_ports() {
let parsed = parse(&request(
r#"{"command":"connect","params":{"peer":"npub1abc","remote_port":4242,"local_port":5000}}"#,
))
.unwrap();
assert_eq!(
parsed,
Command::Connect(Connect {
peer: "npub1abc".to_string(),
remote: 4242,
local: Some(5000),
})
);
}
#[test]
fn an_absent_local_port_asks_for_an_ephemeral_one() {
let parsed = parse(&request(
r#"{"command":"connect","params":{"peer":"npub1abc","remote_port":4242}}"#,
))
.unwrap();
let Command::Connect(connect) = parsed else {
panic!("expected a connect");
};
assert_eq!(connect.local, None);
}
#[test]
fn a_zero_local_port_means_the_same_as_an_absent_one() {
let parsed = parse(&request(
r#"{"command":"connect","params":{"peer":"npub1abc","remote_port":4242,"local_port":0}}"#,
))
.unwrap();
let Command::Connect(connect) = parsed else {
panic!("expected a connect");
};
assert_eq!(connect.local, None);
}
#[test]
fn the_protocol_port_range_is_refused() {
let error = parse(&request(
r#"{"command":"listen","params":{"local_port":200}}"#,
))
.unwrap_err();
assert_eq!(error, CommandError::PortProtocol(200));
}
#[test]
fn the_standard_service_port_range_is_refused() {
// 256 is the IPv6 shim; refusing the whole tier keeps a client from
// binding it or anything reserved beside it.
let error = parse(&request(
r#"{"command":"listen","params":{"local_port":256}}"#,
))
.unwrap_err();
assert_eq!(error, CommandError::PortStandard(256));
assert_eq!(check_bindable(1023), Err(CommandError::PortStandard(1023)));
assert!(check_bindable(1024).is_ok());
}
#[test]
fn a_remote_port_in_a_reserved_tier_is_refused_too() {
// The far end's shim is no more addressable than our own: a client that
// could name port 256 remotely would be injecting into the peer's IPv6
// plane.
let error = parse(&request(
r#"{"command":"connect","params":{"peer":"npub1abc","remote_port":256}}"#,
))
.unwrap_err();
assert_eq!(error, CommandError::PortStandard(256));
}
#[test]
fn every_refusal_carries_the_errno_its_contract_names() {
// The rows of the table the client turns into `io::Error`. A client
// must be able to tell "that port is taken" from "that port is not
// yours to take" without matching English prose, which is the whole
// reason the code is on the wire.
assert_eq!(errno(&RegistryError::PortTaken(4242)), "EADDRINUSE");
assert_eq!(
errno(&RegistryError::FlowTaken {
local: 4242,
remote: 5000
}),
"EADDRINUSE"
);
assert_eq!(errno(&RegistryError::NoPort), "EADDRNOTAVAIL");
assert_eq!(errno(&RegistryError::TooManyFlows(256)), "EMFILE");
assert_eq!(
errno(&RegistryError::Port(CommandError::PortStandard(256))),
"EADDRNOTAVAIL"
);
assert_eq!(
errno(&RegistryError::Port(CommandError::PortProtocol(200))),
"EADDRNOTAVAIL"
);
// And the refusals that never reach the registry.
assert_eq!(
command_errno(&CommandError::Unknown("teleport".to_string())),
"EINVAL"
);
assert_eq!(command_errno(&CommandError::NoParams("stats")), "EINVAL");
}
#[test]
fn an_unknown_command_names_itself_in_the_refusal() {
let error = parse(&request(r#"{"command":"teleport"}"#)).unwrap_err();
assert_eq!(error, CommandError::Unknown("teleport".to_string()));
}
#[test]
fn a_command_that_needs_params_refuses_without_them() {
let error = parse(&request(r#"{"command":"stats"}"#)).unwrap_err();
assert_eq!(error, CommandError::NoParams("stats"));
}
#[test]
fn params_of_the_wrong_shape_name_the_command() {
let request = Request {
command: "listen".to_string(),
params: Some(json!({"local_port": "not a number"})),
};
let error = parse(&request).unwrap_err();
let CommandError::BadParams { command, .. } = error else {
panic!("expected a params error");
};
assert_eq!(command, "listen");
}
#[test]
fn accept_and_reject_are_not_commands_this_api_serves() {
// They went with the round trip they existed for: a flow is taken by
// reading the listener's descriptor and refused by closing the one that
// arrives with it. A client still sending either must be told so.
for command in ["accept", "reject"] {
let error = parse(&request(&format!(
r#"{{"command":"{command}","params":{{"flow_id":7}}}}"#
)))
.unwrap_err();
assert_eq!(error, CommandError::Unknown(command.to_string()));
}
}
#[test]
fn a_listen_may_name_no_port_and_gets_an_ephemeral_one() {
for line in [
r#"{"command":"listen","params":{"local_port":0}}"#,
r#"{"command":"listen","params":{}}"#,
] {
assert_eq!(
parse(&request(line)).unwrap(),
Command::Listen(Listen { local: None }),
"{line} should ask for an ephemeral port"
);
}
assert_eq!(
parse(&request(
r#"{"command":"listen","params":{"local_port":4242}}"#
))
.unwrap(),
Command::Listen(Listen { local: Some(4242) })
);
}
}
File diff suppressed because it is too large Load Diff
+611
View File
@@ -0,0 +1,611 @@
//! A connected `SOCK_SEQPACKET` pair, one half driven by tokio readiness.
//!
//! `SOCK_SEQPACKET` is what a datagram API wants from a local socket: it keeps
//! message boundaries, it is flow controlled, and it reports end of file when
//! the peer closes. `SOCK_STREAM` loses the boundaries and `SOCK_DGRAM` gives a
//! weaker close signal.
//!
//! Tokio ships no type for it — `UnixStream` is `SOCK_STREAM` and
//! `UnixDatagram` is `SOCK_DGRAM` — so the daemon's half is driven through
//! [`AsyncFd`], which is tokio's supported way to put an arbitrary file
//! descriptor under the reactor.
//!
//! **macOS uses `SOCK_DGRAM` instead**, because it does not implement
//! `SOCK_SEQPACKET` for `AF_UNIX`. Everything else this module needs works
//! there: `SCM_RIGHTS` is supported, `AsyncFd` is backed by kqueue, and a
//! connected `SOCK_DGRAM` pair keeps message boundaries just as `SOCK_SEQPACKET`
//! does.
//!
//! The two types are **not** interchangeable at end of file, and that is the
//! whole difficulty of the port. Measured 2026-08-20 and recorded in
//! [`super::dgram_probe`]: on Linux 6.8 a closed peer on a connected
//! `SOCK_DGRAM` pair produces no signal whatsoever, leaving `revents` empty and
//! `recv` returning `EAGAIN`, which is exactly what an idle socket with a live
//! peer does. Darwin does report the close. So the socket type is chosen per
//! platform and the close rule below accepts either signal rather than assuming
//! the one this kernel happens to use.
use std::io;
use std::os::fd::{AsRawFd, FromRawFd, OwnedFd, RawFd};
use std::time::Duration;
use tokio::io::unix::AsyncFd;
/// What one receive attempt produced.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum Received {
/// A datagram of this many bytes. Zero is a legitimate value: a client may
/// send an empty datagram, and that is not the same as closing.
Datagram(usize),
/// The peer closed its half of the pair.
Eof,
}
/// The socket type a flow's descriptor uses on this platform.
///
/// `SOCK_SEQPACKET` everywhere it exists for `AF_UNIX`. macOS does not
/// implement it there, and `SOCK_DGRAM` is the replacement: it keeps message
/// boundaries, which is the property the API's contract with its clients rests
/// on. See the module header for what changes with it and what does not.
#[cfg(not(target_os = "macos"))]
const SOCK_TYPE: libc::c_int = libc::SOCK_SEQPACKET;
#[cfg(target_os = "macos")]
const SOCK_TYPE: libc::c_int = libc::SOCK_DGRAM;
/// Create a connected socket pair of this platform's [`SOCK_TYPE`].
///
/// Both descriptors are close-on-exec so neither leaks into a child process.
/// Linux and FreeBSD take `SOCK_CLOEXEC` in `socketpair`'s type argument and
/// set it atomically; **macOS rejects it there**, so Darwin sets `FD_CLOEXEC`
/// with `fcntl` afterwards instead. The Darwin path has a window between the
/// two calls in which a concurrent `fork` and `exec` would inherit the
/// descriptors. It is accepted rather than closed because the alternative needs
/// a lock this module has no business holding, and the daemon spawns no child
/// on this path.
///
/// Neither descriptor is set non-blocking here: the two halves are independent
/// sockets, so [`Seqpacket::new`] can make the daemon's half non-blocking for
/// the reactor while the half handed to the client stays blocking, which is
/// what a client calling `recv` in a loop expects.
pub fn pair() -> io::Result<(OwnedFd, OwnedFd)> {
#[cfg(not(target_os = "macos"))]
let sock_type = SOCK_TYPE | libc::SOCK_CLOEXEC;
#[cfg(target_os = "macos")]
let sock_type = SOCK_TYPE;
let mut fds = [0 as libc::c_int; 2];
// SAFETY: `fds` is a two-element array of the type socketpair writes, and
// the return value is checked before either descriptor is read.
let rc = unsafe { libc::socketpair(libc::AF_UNIX, sock_type, 0, fds.as_mut_ptr()) };
if rc != 0 {
return Err(io::Error::last_os_error());
}
// SAFETY: socketpair reported success, so both entries are open
// descriptors this process now owns. Wrapping them before any further
// syscall means an error below closes them rather than leaking them.
let pair = unsafe { (OwnedFd::from_raw_fd(fds[0]), OwnedFd::from_raw_fd(fds[1])) };
#[cfg(target_os = "macos")]
{
set_cloexec(pair.0.as_raw_fd())?;
set_cloexec(pair.1.as_raw_fd())?;
}
Ok(pair)
}
/// Mark `fd` close-on-exec, for the platform that will not do it atomically.
///
/// Read-modify-write rather than a bare set, so that any other flag the
/// descriptor carries survives. Used both here, where `socketpair` refuses
/// `SOCK_CLOEXEC`, and by [`super::fdpass`], where `recvmsg` has no
/// `MSG_CMSG_CLOEXEC` to pass.
#[cfg(target_os = "macos")]
pub(super) fn set_cloexec(fd: RawFd) -> io::Result<()> {
// SAFETY: the descriptor is open for the call.
let flags = unsafe { libc::fcntl(fd, libc::F_GETFD) };
if flags < 0 {
return Err(io::Error::last_os_error());
}
// SAFETY: the descriptor is open and the flag word is the one just read.
let rc = unsafe { libc::fcntl(fd, libc::F_SETFD, flags | libc::FD_CLOEXEC) };
if rc < 0 {
return Err(io::Error::last_os_error());
}
Ok(())
}
/// Size the send buffer of `fd`, in bytes.
///
/// Used on the daemon's half of a listener's pair, where the buffer is the only
/// bound on arrivals a client has stopped reading. `SO_SNDBUF` on an `AF_UNIX`
/// socket accounts bytes plus per-message overhead rather than messages, so a
/// caller converting a message count to bytes is approximating and should
/// approximate generously. Linux doubles what it is given and clamps to its own
/// minimum, which is why nothing here reads the value back and asserts on it.
pub fn set_sndbuf(fd: &OwnedFd, bytes: usize) -> io::Result<()> {
let size = libc::c_int::try_from(bytes)
.map_err(|_| io::Error::new(io::ErrorKind::InvalidInput, "send buffer size too large"))?;
// SAFETY: the descriptor is open for the call, and the pointer and length
// describe one `c_int` that outlives it.
let rc = unsafe {
libc::setsockopt(
fd.as_raw_fd(),
libc::SOL_SOCKET,
libc::SO_SNDBUF,
std::ptr::addr_of!(size).cast(),
std::mem::size_of::<libc::c_int>() as libc::socklen_t,
)
};
if rc != 0 {
return Err(io::Error::last_os_error());
}
Ok(())
}
/// How long a receive waits on the reactor before retrying the syscall, on the
/// platform whose reactor cannot see a closed `SOCK_DGRAM` peer.
///
/// This is a detection latency, not a poll interval for data: an arriving
/// datagram does wake the reactor normally, so ordinary traffic is unaffected
/// and this bound is only reached on an idle flow. What it bounds is how long a
/// closed flow keeps its port and its registry entry before the daemon notices.
///
/// A quarter second is chosen against the cost, which is one timer per idle
/// flow and listener. Shortening it buys a faster reclaim of something nothing
/// is waiting on; lengthening it holds a dead flow's port longer.
#[cfg(target_os = "macos")]
const CLOSE_RETRY: Duration = Duration::from_millis(250);
/// The daemon's half of a flow's or a listener's socket pair, registered with
/// the reactor.
pub struct Seqpacket {
inner: AsyncFd<OwnedFd>,
}
impl Seqpacket {
/// Put `fd` under the reactor, making it non-blocking first.
///
/// `AsyncFd` requires a non-blocking descriptor: a blocking one would stall
/// the whole runtime thread inside a syscall the reactor believed would
/// return at once.
pub fn new(fd: OwnedFd) -> io::Result<Self> {
set_nonblocking(fd.as_raw_fd(), true)?;
Ok(Self {
inner: AsyncFd::new(fd)?,
})
}
/// The raw descriptor, for a caller that must perform its own syscall.
///
/// The one such caller is the listener's task, which needs a send that
/// reports a full buffer rather than waiting for one; see
/// [`try_send`](super::fdpass::try_send). Everything else goes through
/// [`Seqpacket::send`] and [`Seqpacket::recv`].
pub(super) fn raw(&self) -> RawFd {
self.inner.get_ref().as_raw_fd()
}
/// Receive one datagram, or report that the peer closed.
///
/// A datagram longer than `buf` is truncated and the remainder discarded,
/// which is `SOCK_SEQPACKET` behaviour. Callers size `buf` at the largest
/// payload the API accepts, so a truncation means the client exceeded it.
pub async fn recv(&self, buf: &mut [u8]) -> io::Result<Received> {
loop {
// **A read before the wait, because on one platform the reactor
// cannot see a close.** Darwin's `unp_disconnect` takes a different
// branch for `SOCK_DGRAM` than for `SOCK_STREAM`: it clears
// `SS_ISCONNECTED` and latches `so_error`, and calls none of
// `sorwakeup`, `socantrcvmore` or `soisdisconnected`. So a closed
// peer wakes no knote. The registration is edge-triggered and was
// made while the socket was healthy, so waiting first would park
// for ever on exactly the event this call exists to report. The
// latched error is visible to a syscall, and only to a syscall.
//
// On Linux this attempt costs one `recv` returning `EAGAIN` before
// the wait, and changes nothing else.
match recv_once(self.raw(), buf) {
Err(error) if error.kind() == io::ErrorKind::WouldBlock => {}
other => return other,
}
// **The wait is bounded where the close raises no event**, so a peer
// that goes away while this task is parked is noticed on the next
// attempt rather than never. `so_error` is latched until a read
// consumes it, so the bound sets detection latency and cannot lose
// the event. Everywhere else the readiness is authoritative and the
// wait stays unbounded.
#[cfg(target_os = "macos")]
let mut guard = match tokio::time::timeout(CLOSE_RETRY, self.inner.readable()).await {
Ok(ready) => ready?,
// The bound expired. The loop retries the syscall, which is the
// only thing that can see this platform's close.
Err(_elapsed) => continue,
};
#[cfg(not(target_os = "macos"))]
let mut guard = self.inner.readable().await?;
let attempt = guard.try_io(|inner| recv_once(inner.get_ref().as_raw_fd(), buf));
match attempt {
Ok(result) => return result,
// The reactor said readable and the syscall disagreed. Clear
// the readiness and wait again rather than reporting an error.
Err(_would_block) => continue,
}
}
}
/// Send one datagram.
///
/// `SOCK_SEQPACKET` delivers it whole or not at all, so a short write is
/// not a case the caller has to handle.
pub async fn send(&self, buf: &[u8]) -> io::Result<usize> {
loop {
let mut guard = self.inner.writable().await?;
let attempt = guard.try_io(|inner| {
// SAFETY: the descriptor is owned and open, and the pointer and
// length describe `buf`.
let n = unsafe {
libc::send(
inner.get_ref().as_raw_fd(),
buf.as_ptr().cast(),
buf.len(),
libc::MSG_NOSIGNAL,
)
};
if n < 0 {
Err(io::Error::last_os_error())
} else {
Ok(n as usize)
}
});
match attempt {
Ok(result) => return result,
Err(_would_block) => continue,
}
}
}
}
/// One receive, distinguishing an empty datagram from end of file.
///
/// Both produce a zero-byte read, so something else has to tell them apart.
/// `MSG_EOR` is the technique the manual pages suggest and it **does not work
/// here**: measured on Linux 6.8, `recvmsg` on an `AF_UNIX` `SOCK_SEQPACKET`
/// socket returns `msg_flags == 0` for a normal message, for an empty message
/// and at end of file alike, so the flag carries no information.
///
/// `POLLHUP` does discriminate, measured the same way. After a zero-byte read,
/// a queued empty datagram leaves the socket with no events pending, while a
/// closed peer leaves `POLLHUP` set and latched. So a zero-byte read is end of
/// file only when the peer has hung up.
///
/// This matters because reading an empty datagram as a close would let a client
/// tear down its own flow by sending nothing, and the defect would present as a
/// spurious disconnect.
///
/// **`ECONNRESET` is treated as end of file too**, for the platform whose
/// datagram sockets report a close that way rather than through `POLLHUP`. Both
/// arms are compiled everywhere rather than split by `cfg`, because a rule that
/// accepts either signal is correct on both kernels and a rule that assumes one
/// would be silently wrong on the other. `EAGAIN` is deliberately not in that
/// company: it means the socket is empty and the peer is alive, so it stays an
/// error and the caller waits for readiness again.
fn recv_once(fd: RawFd, buf: &mut [u8]) -> io::Result<Received> {
// SAFETY: the descriptor is open and the pointer and length describe `buf`.
let n = unsafe { libc::recv(fd, buf.as_mut_ptr().cast(), buf.len(), 0) };
if n < 0 {
let err = io::Error::last_os_error();
if err.raw_os_error() == Some(libc::ECONNRESET) {
return Ok(Received::Eof);
}
return Err(err);
}
if n == 0 && peer_hung_up(fd) {
return Ok(Received::Eof);
}
Ok(Received::Datagram(n as usize))
}
/// Translate a platform's spelling of "the peer is gone" into this API's.
///
/// Darwin reports a closed `SOCK_DGRAM` peer as `ECONNRESET`, on the read path
/// and the write path alike, where Linux `SOCK_SEQPACKET` reports `EPIPE` on a
/// write and a zero-byte read plus `POLLHUP` on a read. The client's contract
/// names `EPIPE` for that condition and says so in its documentation, so the
/// difference is translated once here rather than at each call site. Every
/// other errno passes through untouched, because only this one condition has
/// two spellings.
pub(super) fn peer_gone_as_epipe(error: io::Error) -> io::Error {
if error.raw_os_error() == Some(libc::ECONNRESET) {
return io::Error::from_raw_os_error(libc::EPIPE);
}
error
}
/// Whether the peer has closed its half of the pair.
///
/// `POLLIN` is requested even though only `POLLHUP` is read. `POLLHUP` is
/// reported in `revents` whether or not it was asked for, which holds on Linux
/// and was measured there; **an empty `events` is not enough on Darwin**, where
/// a poll that requests nothing registers no filter and returns 0 with an empty
/// `revents` for a peer that has in fact closed. That was measured too, in
/// [`super::dgram_probe`], and asking for `POLLIN` costs nothing on either
/// platform because the result is masked to `POLLHUP` regardless.
///
/// This function is not what detects a close on Darwin: `ECONNRESET` arrives
/// first and both callers act on it before reaching the zero-byte path. It is
/// corrected so that it answers truthfully wherever it is called from, rather
/// than being left as a function whose contract holds on one platform.
/// The poll does not block, and runs only on the zero-byte path.
///
/// Visible within [`super`] because the client half needs the same
/// discrimination on the same socket pair: the rule belongs to the pair, not to
/// the end that reads it.
pub(super) fn peer_hung_up(fd: RawFd) -> bool {
let mut poll = libc::pollfd {
fd,
events: libc::POLLIN,
revents: 0,
};
// SAFETY: `poll` points at one live pollfd and the call cannot block.
let rc = unsafe { libc::poll(&mut poll, 1, 0) };
rc > 0 && (poll.revents & libc::POLLHUP) != 0
}
/// Size the receive buffer of `fd`, in bytes.
///
/// The counterpart to [`set_sndbuf`], and needed because the two kernels charge
/// a queued `AF_UNIX` message to different ends: Linux to the sender's
/// `SO_SNDBUF`, BSD to the receiver's `so_rcv`. Sizing only one leaves the
/// queue bounded by a system default on the other platform. As with the send
/// buffer, the value is not read back and asserted, because a kernel is free to
/// double it or clamp it to its own minimum.
pub(super) fn set_rcvbuf(fd: &OwnedFd, bytes: usize) -> io::Result<()> {
let size = libc::c_int::try_from(bytes).map_err(|_| {
io::Error::new(io::ErrorKind::InvalidInput, "receive buffer size too large")
})?;
// SAFETY: the descriptor is open for the call, and the pointer and length
// describe one `c_int` that outlives it.
let rc = unsafe {
libc::setsockopt(
fd.as_raw_fd(),
libc::SOL_SOCKET,
libc::SO_RCVBUF,
std::ptr::addr_of!(size).cast(),
std::mem::size_of::<libc::c_int>() as libc::socklen_t,
)
};
if rc != 0 {
return Err(io::Error::last_os_error());
}
Ok(())
}
/// Set or clear a receive or send timeout on a socket.
///
/// `option` is `SO_RCVTIMEO` or `SO_SNDTIMEO`. `None` clears the timeout, which
/// the kernel spells as a zero `timeval`.
///
/// A zero duration is refused with `EINVAL` rather than passed through, because
/// the kernel reads a zero `timeval` as "no timeout" and a caller asking for
/// zero means the opposite. `std::net` makes the same refusal for the same
/// reason, and silently inverting the request would be worse than failing it.
pub(super) fn set_timeout(fd: RawFd, option: libc::c_int, dur: Option<Duration>) -> io::Result<()> {
let timeout = match dur {
Some(d) if d == Duration::ZERO => {
return Err(io::Error::from_raw_os_error(libc::EINVAL));
}
// Saturating rather than wrapping: a duration past `time_t` becomes the
// longest wait the kernel can express, which is the caller's intent.
//
// `tv_usec` is cast rather than converted because `suseconds_t` is
// `i64` on Linux and `i32` on Darwin. `From` does not exist for the
// Darwin width and `try_from` is a clippy error on the Linux one, so a
// cast is the only form that compiles on both. It cannot truncate:
// `subsec_micros` is below 1_000_000 by construction and fits either.
Some(d) => libc::timeval {
tv_sec: d.as_secs().min(libc::time_t::MAX as u64) as libc::time_t,
tv_usec: d.subsec_micros() as libc::suseconds_t,
},
None => libc::timeval {
tv_sec: 0,
tv_usec: 0,
},
};
// SAFETY: the descriptor is open, and the pointer and length describe a
// `timeval` this frame owns.
let rc = unsafe {
libc::setsockopt(
fd,
libc::SOL_SOCKET,
option,
std::ptr::from_ref(&timeout).cast(),
size_of::<libc::timeval>() as libc::socklen_t,
)
};
if rc < 0 {
return Err(io::Error::last_os_error());
}
Ok(())
}
/// Read back a receive or send timeout, or `None` when none is set.
pub(super) fn timeout(fd: RawFd, option: libc::c_int) -> io::Result<Option<Duration>> {
let mut timeout = libc::timeval {
tv_sec: 0,
tv_usec: 0,
};
let mut len = size_of::<libc::timeval>() as libc::socklen_t;
// SAFETY: the descriptor is open, and both pointers describe values this
// frame owns and keeps alive across the call.
let rc = unsafe {
libc::getsockopt(
fd,
libc::SOL_SOCKET,
option,
std::ptr::from_mut(&mut timeout).cast(),
&mut len,
)
};
if rc < 0 {
return Err(io::Error::last_os_error());
}
if timeout.tv_sec == 0 && timeout.tv_usec == 0 {
return Ok(None);
}
Ok(Some(
Duration::from_secs(timeout.tv_sec as u64) + Duration::from_micros(timeout.tv_usec as u64),
))
}
/// Set or clear a descriptor's non-blocking mode, preserving its other flags.
///
/// Read-modify-write rather than a bare `F_SETFL`, because the flag word also
/// carries the access mode and `O_APPEND`, and writing `O_NONBLOCK` alone would
/// drop them.
pub(super) fn set_nonblocking(fd: RawFd, nonblocking: bool) -> io::Result<()> {
// SAFETY: the descriptor is open for the duration of both calls.
let flags = unsafe { libc::fcntl(fd, libc::F_GETFL) };
if flags < 0 {
return Err(io::Error::last_os_error());
}
let wanted = if nonblocking {
flags | libc::O_NONBLOCK
} else {
flags & !libc::O_NONBLOCK
};
// SAFETY: as above; `wanted` is the value just read back, one bit changed.
if unsafe { libc::fcntl(fd, libc::F_SETFL, wanted) } < 0 {
return Err(io::Error::last_os_error());
}
Ok(())
}
#[cfg(test)]
mod tests {
use super::*;
use std::io::{Read, Write};
use std::os::unix::net::UnixStream as StdUnixStream;
/// Turn the client half into something a blocking test can drive, the way a
/// client process would after receiving it.
fn client(fd: OwnedFd) -> StdUnixStream {
StdUnixStream::from(fd)
}
/// One receive, bounded in time.
///
/// Every assertion in this module is about something arriving on the
/// descriptor, and the failure mode of each is that it never does. An
/// unbounded await against that defect parks for ever: the suite wedges
/// with no named failure and no diagnostic, and a hang is not a red. This
/// matters more since the socket type became platform-dependent, because a
/// kernel that does not report a close makes the end-of-file test the one
/// that hangs, and it would hang on a runner rather than here.
async fn recv_bounded(sock: &Seqpacket, buf: &mut [u8]) -> Received {
tokio::time::timeout(Duration::from_secs(5), sock.recv(buf))
.await
.expect("nothing arrived on the descriptor within 5s: it is still waiting")
.expect("recv failed")
}
#[tokio::test]
async fn message_boundaries_survive_in_both_directions() {
let (daemon, theirs) = pair().unwrap();
let daemon = Seqpacket::new(daemon).unwrap();
let mut theirs = client(theirs);
// Three writes must arrive as three datagrams, not one run of bytes.
// This is the property SOCK_STREAM would lose.
theirs.write_all(b"one").unwrap();
theirs.write_all(b"two").unwrap();
theirs.write_all(b"three").unwrap();
let mut buf = [0u8; 64];
assert_eq!(recv_bounded(&daemon, &mut buf).await, Received::Datagram(3));
assert_eq!(&buf[..3], b"one");
assert_eq!(recv_bounded(&daemon, &mut buf).await, Received::Datagram(3));
assert_eq!(&buf[..3], b"two");
assert_eq!(recv_bounded(&daemon, &mut buf).await, Received::Datagram(5));
assert_eq!(&buf[..5], b"three");
daemon.send(b"alpha").await.unwrap();
daemon.send(b"beta").await.unwrap();
let mut got = [0u8; 64];
assert_eq!(theirs.read(&mut got).unwrap(), 5);
assert_eq!(&got[..5], b"alpha");
assert_eq!(theirs.read(&mut got).unwrap(), 4);
assert_eq!(&got[..4], b"beta");
}
#[tokio::test]
async fn closing_the_client_half_reports_end_of_file() {
let (daemon, theirs) = pair().unwrap();
let daemon = Seqpacket::new(daemon).unwrap();
drop(client(theirs));
let mut buf = [0u8; 64];
assert_eq!(recv_bounded(&daemon, &mut buf).await, Received::Eof);
}
#[tokio::test]
async fn an_empty_datagram_is_not_end_of_file() {
// The whole reason recv_once uses recvmsg. If an empty datagram read as
// a close, a client could tear down its own flow by sending nothing,
// and the bug would look like a spurious disconnect.
let (daemon, theirs) = pair().unwrap();
let daemon = Seqpacket::new(daemon).unwrap();
let mut theirs = client(theirs);
// `write_all(b"")` is a no-op in Rust and never reaches the socket, so
// the zero-length datagram has to be sent with `send` directly. The
// first version of this test used `write_all` and asserted a behaviour
// it had not exercised.
// SAFETY: the descriptor is open and owned by `theirs`.
let sent = unsafe { libc::send(theirs.as_raw_fd(), std::ptr::null(), 0, 0) };
assert_eq!(sent, 0, "{}", io::Error::last_os_error());
theirs.write_all(b"after").unwrap();
let mut buf = [0u8; 64];
assert_eq!(recv_bounded(&daemon, &mut buf).await, Received::Datagram(0));
assert_eq!(recv_bounded(&daemon, &mut buf).await, Received::Datagram(5));
}
#[test]
fn both_halves_of_a_pair_are_close_on_exec() {
// Asserted rather than assumed because the platforms disagree on how it
// is set: Linux and FreeBSD get it atomically from socketpair's type
// argument, macOS rejects it there and has to use a second fcntl. A
// missing flag leaks both descriptors into any child the daemon spawns,
// which is silent, so nothing else would report it.
let (daemon, theirs) = pair().unwrap();
for (label, fd) in [
("daemon", daemon.as_raw_fd()),
("client", theirs.as_raw_fd()),
] {
// SAFETY: the descriptor is open and owned by this frame.
let flags = unsafe { libc::fcntl(fd, libc::F_GETFD) };
assert!(flags >= 0, "{label} half: F_GETFD failed");
assert_eq!(
flags & libc::FD_CLOEXEC,
libc::FD_CLOEXEC,
"{label} half of the pair is not close-on-exec"
);
}
}
#[tokio::test]
async fn the_client_half_is_left_blocking() {
// The daemon's half is made non-blocking for the reactor. If that flag
// reached the client's half, a client doing an ordinary blocking recv
// would get EAGAIN instead of waiting, which is a trap worth a test
// rather than a comment.
let (daemon, theirs) = pair().unwrap();
let _daemon = Seqpacket::new(daemon).unwrap();
// SAFETY: the descriptor is open and owned by `theirs`.
let flags = unsafe { libc::fcntl(theirs.as_raw_fd(), libc::F_GETFL) };
assert!(flags >= 0);
assert_eq!(flags & libc::O_NONBLOCK, 0);
}
}
+73
View File
@@ -134,6 +134,51 @@ impl Node {
// Drop unused sender to avoid keeping channel open if control is disabled
drop(control_tx);
// Native datagram API socket. Experimental, off by default, and built
// on Linux, FreeBSD and macOS: Windows has no way to pass a descriptor
// between processes at all, which is the mechanism itself. Bound
// synchronously so a bad path or a socket already in use is reported
// here, before the node starts serving, rather than at whatever later
// moment a spawned bind happened to run.
// Two channels, as the TUN plane has: registrations go one way and are
// rare, datagrams go the other and arrive in bursts. Keeping them apart
// lets the data arm drain in batches without a registration waiting
// behind a burst of traffic.
let (native_out_tx, mut native_outbound_rx) =
tokio::sync::mpsc::channel::<crate::native::link::Outbound>(1024);
let _native_out_guard = native_out_tx.clone();
let (mut native_rx, _native_guard) = {
let (tx, rx) = tokio::sync::mpsc::channel::<crate::native::link::NativeMessage>(64);
// The `cfg` sits on the binding rather than on the arm that drains
// the receiver, because `tokio::select!` does not accept one. Where
// there is no listener the channel exists and nothing ever sends,
// and the guard keeps it open so the arm never sees a closed
// receiver.
#[cfg(any(target_os = "linux", target_os = "freebsd", target_os = "macos"))]
let guard = {
let mut guard = Some(tx.clone());
if self.config().node.native_api.enabled {
match crate::native::NativeApi::bind(&self.config().node.native_api) {
Ok(socket) => {
// The accept loop owns a sender, so the dummy guard
// is dropped: the channel closes when the last
// client task ends, not while one is still serving.
guard = None;
let npub = self.npub();
tokio::spawn(socket.accept_loop(tx, native_out_tx.clone(), npub));
}
Err(e) => {
warn!(error = %e, "Failed to bind native API socket");
}
}
}
guard
};
#[cfg(not(any(target_os = "linux", target_os = "freebsd", target_os = "macos")))]
let guard = Some(tx.clone());
(rx, guard)
};
// Decrypt-worker fallback receiver. The worker pushes each
// authenticated FMP plaintext here so rx_loop can finish the
// per-peer side-effects (stats, MMP, ECN, link dispatch).
@@ -311,6 +356,30 @@ impl Node {
);
self.register_identity(identity.node_addr, identity.pubkey);
}
// Native API datagrams a client wrote to its descriptor. Drained
// in a burst like the TUN arm, for the same reason: one wake-up
// should clear what a client handed over, not one datagram.
Some(out) = native_outbound_rx.recv() => {
self.handle_native_outbound(out.key, out.peer, out.payload).await;
let mut drained = 0;
while drained < 256 {
match native_outbound_rx.try_recv() {
Ok(next) => {
self.handle_native_outbound(next.key, next.peer, next.payload).await;
drained += 1;
}
Err(_) => break,
}
}
}
// Native API registry requests. Placed after the hot inbound
// path so a burst of client registrations cannot delay packet
// processing. No `cfg` here: `tokio::select!` does not accept
// one, so the channel exists on every platform and only the
// listener that feeds it is gated.
Some(message) = native_rx.recv() => {
self.handle_native(message);
}
Some((request, response_tx)) = control_rx.recv() => {
// Only mutating COMMAND requests (`connect` / `disconnect`)
// reach the rx_loop now. Every pure-read `show_*` query is
@@ -348,6 +417,10 @@ impl Node {
instr_step!(instr_on, crate::instr::Domain::Tick, crate::instr::Step::WholeTick, {
instr_step!(instr_on, crate::instr::Domain::Tick, crate::instr::Step::CheckTimeouts,
self.check_timeouts().await);
// Discard flows the rx_loop announced and a listener's
// own task never took. Cheap: it walks only the pending
// map, which the backlog bounds.
self.native_expire();
let now_ms = Self::now_ms();
instr_step!(instr_on, crate::instr::Domain::Tick, crate::instr::Step::ReloadPeerAcl,
self.reload_peer_acl().await);
+2
View File
@@ -3,6 +3,8 @@
mod handshake;
pub(crate) mod lookup;
mod mmp;
mod native;
pub(in crate::node) use native::PendingNative;
pub(crate) mod probe;
mod rekey;
pub(in crate::node) mod session;
+228
View File
@@ -0,0 +1,228 @@
//! Serve the native API registry from inside the `rx_loop`.
//!
//! The registry lives in `Node`, so every client request arrives here and is
//! answered on the `oneshot` it carried. Nothing else touches the registry,
//! which is why the receive path takes no lock.
//!
//! The work itself is in [`crate::native::link`], as a free function over the
//! registry. This file supplies the clock and nothing else, so the logic the
//! daemon runs is the same logic a test drives without building a node.
use crate::identity::NodeAddr;
use crate::native::link::{self, NativeMessage, Outcome, Served};
use crate::native::registry::FlowKey;
use crate::node::Node;
use crate::proto::fsp::wire::FSP_PORT_HEADER_SIZE;
use crate::upper::icmp::FIPS_OVERHEAD;
use secp256k1::XOnlyPublicKey;
use tracing::debug;
/// One native datagram held while its destination's session establishes.
///
/// Carries its ports, which is exactly what the TUN pending queue lacks and
/// why native datagrams need a queue of their own.
#[derive(Debug, Clone)]
pub(in crate::node) struct PendingNative {
/// The flow this datagram belongs to.
pub key: FlowKey,
/// The payload, without any port header: `send_session_data` adds that.
pub payload: Vec<u8>,
}
impl Node {
/// Answer one native API request, and count what it did.
///
/// The counting is here rather than in [`link::serve`] because the core
/// decides and the shell observes. `serve` reports its outcome and knows
/// nothing about counters.
pub(in crate::node) fn handle_native(&mut self, message: NativeMessage) {
let max = self.native_max_payload();
let served = link::serve(&mut self.native, message, Self::now_ms(), max);
let native = &self.metrics.native;
match served {
Served::Opened => native.flows_opened.inc(),
Served::Accepted => native.flows_accepted.inc(),
Served::Discarded(reason) => native.record_drop(reason),
Served::Released(closed) => native.flows_closed.add(closed as u64),
Served::Delivered(outcome, bytes) => self.count_native_delivery(outcome, bytes),
Served::Untracked => {}
}
}
/// Count one dispatched inbound datagram.
///
/// Shared by the wire path and the debug arrival command, which run the same
/// dispatch, so a counter cannot disagree between them.
fn count_native_delivery(&self, outcome: Outcome, bytes: usize) {
let native = &self.metrics.native;
match outcome {
Outcome::Delivered | Outcome::Announced(_) | Outcome::Held(_) => {
native.record_received(bytes);
}
Outcome::Dropped(reason) => native.record_drop(reason),
}
}
/// Deliver one inbound datagram to whatever owns its destination port.
///
/// `pubkey` is the peer's address, read from the session entry that
/// authenticated it. It is captured here, at the one call site the wire
/// has, rather than resolved when a report is rendered: the node address
/// the wire carries is a truncated hash and does not invert.
pub(in crate::node) fn native_deliver(
&mut self,
peer: crate::identity::NodeAddr,
pubkey: XOnlyPublicKey,
src: u16,
dst: u16,
data: Vec<u8>,
) -> Outcome {
let bytes = data.len();
let outcome = link::deliver(
&mut self.native,
peer,
pubkey,
src,
dst,
data,
Self::now_ms(),
);
self.count_native_delivery(outcome, bytes);
outcome
}
/// The largest payload a native flow may send, in bytes.
///
/// Wire size is `FIPS_OVERHEAD` plus the four-byte port header plus the
/// payload, so the payload is whatever is left of the transport MTU. This
/// is 40 bytes more than the same application data gets through the IPv6
/// shim, which spends them on a header the application never sees.
pub(in crate::node) fn native_max_payload(&self) -> u16 {
self.transport_mtu()
.saturating_sub(FIPS_OVERHEAD)
.saturating_sub(FSP_PORT_HEADER_SIZE as u16)
}
/// Send one datagram a native API client wrote to its descriptor.
///
/// Mirrors `handle_tun_outbound` with three differences: the destination
/// comes from the flow rather than from a parsed IPv6 header, an error is
/// reported to the client rather than as an ICMPv6 message, and the flow's
/// own ports are used instead of the shim's.
pub(in crate::node) async fn handle_native_outbound(
&mut self,
key: FlowKey,
peer: XOnlyPublicKey,
payload: Vec<u8>,
) {
let max = self.native_max_payload() as usize;
if payload.len() > max {
debug!(
len = payload.len(),
max, "Native API datagram exceeds the path payload limit, dropping"
);
self.metrics.native.drop_oversize.inc();
return;
}
if let Some(entry) = self.sessions.get(&key.peer) {
if entry.is_established() {
let bytes = payload.len();
if let Err(error) = self
.send_session_data(&key.peer, key.local, key.remote, &payload)
.await
{
debug!(
peer = %self.peer_display_name(&key.peer),
error = %error,
"Failed to send a native datagram"
);
} else {
self.metrics.native.record_sent(bytes);
}
return;
}
self.queue_native(key, payload);
return;
}
// No session. Start one and hold the datagram; if there is no route,
// ask discovery for one and hold it anyway, exactly as the TUN path
// does.
//
// Every flow carries its peer's address, whether the client named it or
// the session authenticated it, so there is no cache lookup here and no
// datagram dropped for want of a key. `initiate_session` wants a full
// key, and the parity it is lifted with does not matter: the handshake
// hashes the DH output x-only and forces even parity in the XK
// premessage, so both parities derive the same material.
let pubkey = peer.public_key(secp256k1::Parity::Even);
if let Err(error) = self.initiate_session(key.peer, pubkey).await {
debug!(
peer = %self.peer_display_name(&key.peer),
error = %error,
"Failed to initiate a session for a native flow, trying discovery"
);
self.maybe_initiate_lookup(&key.peer).await;
}
self.queue_native(key, payload);
}
/// Hold a native datagram until its destination's session establishes.
///
/// Bounded the same way the TUN queue is, and by the same configured
/// numbers, so one client cannot make the node hold without limit.
fn queue_native(&mut self, key: FlowKey, payload: Vec<u8>) {
let max_dests = self.config().node.session.pending_max_destinations;
if !self.pending_native.contains_key(&key.peer) && self.pending_native.len() >= max_dests {
return;
}
let per_dest = self.config().node.session.pending_packets_per_dest;
let queue = self.pending_native.entry(key.peer).or_default();
crate::proto::fsp::push_bounded_pending(queue, PendingNative { key, payload }, per_dest);
}
/// Send the native datagrams held for a destination whose session is up.
///
/// Called beside `flush_pending_packets` at every site that flushes the TUN
/// queue. A separate queue means a separate flush, and forgetting one would
/// leave datagrams held until the session went away.
pub(in crate::node) async fn flush_pending_native(&mut self, dest: &NodeAddr) {
let Some(held) = self.pending_native.remove(dest) else {
return;
};
for entry in held {
let bytes = entry.payload.len();
if let Err(error) = self
.send_session_data(dest, entry.key.local, entry.key.remote, &entry.payload)
.await
{
debug!(
peer = %self.peer_display_name(dest),
error = %error,
"Failed to send a queued native datagram"
);
break;
}
self.metrics.native.record_sent(bytes);
}
}
/// Discard pending flows a listener's task never wired.
///
/// Called from the maintenance tick. The window it closes is the hop
/// between the rx_loop announcing an arrival and the listener's task taking
/// it, so a non-zero count is the daemon failing to complete an arrival
/// rather than a client failing to answer for one. It should be zero on a
/// healthy node.
pub(in crate::node) fn native_expire(&mut self) {
let expired = self.native.expire(Self::now_ms());
self.metrics.native.flows_expired.add(expired.len() as u64);
if !expired.is_empty() {
debug!(
count = expired.len(),
"Native API pending flows expired without an answer"
);
}
}
}
+41 -5
View File
@@ -380,6 +380,7 @@ impl Node {
debug!(len = rest.len(), "DataPacket too short for port header");
return;
}
let src_port = u16::from_le_bytes([rest[0], rest[1]]);
let dst_port = u16::from_le_bytes([rest[2], rest[3]]);
let service_payload = &rest[FSP_PORT_HEADER_SIZE..];
@@ -422,11 +423,43 @@ impl Node {
}
}
_ => {
debug!(
src = %self.peer_display_name(src_addr),
dst_port,
"Unknown FSP service port, dropping DataPacket"
);
// Every other port belongs to the native datagram API,
// which decides between an established flow, a listener
// and a drop. The same decision serves the debug
// arrival command, so the rule was exercised before
// this call site existed.
// The peer's key comes from the session entry rather
// than from the wire: the handler above refuses
// anything but an Established session, and an
// Established entry holds the key its handshake
// attested. Nothing has to invert the node address.
let payload = service_payload.to_vec();
// The entry was removed for the trial-decrypt cascade
// and re-inserted above, and this arm is reached only
// for an Established session, so the lookup finds one.
// It is written as a lookup rather than an unwrap
// because a future path that skipped the re-insert
// would otherwise panic on a peer's datagram. Skipping
// is scoped to the native dispatch alone: the idle
// timer and the pending-outbound flush at the end of
// this function are this message's bookkeeping and are
// owed whether or not it had anywhere to go.
if let Some(peer_key) = self
.sessions
.get(src_addr)
.map(|session| session.remote_pubkey().x_only_public_key().0)
{
let outcome = self
.native_deliver(*src_addr, peer_key, src_port, dst_port, payload);
if let crate::native::link::Outcome::Dropped(why) = outcome {
debug!(
src = %self.peer_display_name(src_addr),
dst_port,
why = why.as_str(),
"Unknown FSP service port, dropping DataPacket"
);
}
}
}
}
}
@@ -461,6 +494,7 @@ impl Node {
// Flush any pending outbound packets (e.g., simultaneous initiation
// where responder also had queued outbound packets)
self.flush_pending_packets(src_addr).await;
self.flush_pending_native(src_addr).await;
}
/// Handle an incoming SessionSetup (Noise XX msg1).
@@ -1083,6 +1117,7 @@ impl Node {
// Flush any queued outbound packets for this destination
self.flush_pending_packets(src_addr).await;
self.flush_pending_native(src_addr).await;
info!(src = %self.peer_display_name(src_addr), "Session established (initiator, XX)");
}
@@ -1360,6 +1395,7 @@ impl Node {
// Flush any pending packets
self.flush_pending_packets(src_addr).await;
self.flush_pending_native(src_addr).await;
info!(src = %self.peer_display_name(src_addr), "Session established (responder, XX)");
}
+145 -3
View File
@@ -16,7 +16,7 @@ use std::sync::atomic::{AtomicU64, Ordering};
use crate::node::reject::{BloomReject, DiscoveryReject, ForwardingReject, TreeReject};
use crate::node::stats::{
BloomStatsSnapshot, CongestionStatsSnapshot, ErrorSignalStatsSnapshot, ForwardingStatsSnapshot,
LookupStatsSnapshot, TreeStatsSnapshot,
LookupStatsSnapshot, NativeStatsSnapshot, TreeStatsSnapshot,
};
/// An atomic counter.
@@ -512,11 +512,104 @@ impl ErrorMetrics {
}
}
/// Native datagram API metric counters.
///
/// Every one of these is bumped on the rx_loop, in the native request handler,
/// the outbound handler and the inbound dispatch arm. No client task touches a
/// counter, so there is one path per counter rather than two.
///
/// `flows_accepted` counts flows promoted out of a listener's backlog, which
/// happens before the daemon writes the arrival that carries one to its client.
/// A hand-off that then fails is counted under `drop_listener_not_reading` or
/// `drop_listener_gone` and is not taken back off `flows_accepted`, so what a
/// client actually received is the first less the other two.
#[derive(Default)]
pub struct NativeMetrics {
pub flows_opened: Counter,
pub flows_accepted: Counter,
pub flows_closed: Counter,
pub flows_expired: Counter,
pub sent_datagrams: Counter,
pub sent_bytes: Counter,
pub received_datagrams: Counter,
pub received_bytes: Counter,
pub drop_no_port: Counter,
pub drop_backlog_full: Counter,
pub drop_too_many_flows: Counter,
pub drop_pending_queue_full: Counter,
pub drop_flow_queue_full: Counter,
pub drop_arrival_queue_full: Counter,
pub drop_listener_not_reading: Counter,
pub drop_listener_gone: Counter,
pub drop_oversize: Counter,
}
impl NativeMetrics {
/// Route a typed drop to its counter.
///
/// Exhaustive over [`DropReason`](crate::native::link::DropReason) with no
/// wildcard arm, which is what makes a variant added later a compile error
/// here rather than a datagram that vanishes uncounted. `DropReason` rather
/// than `DropCause` because delivery can also fail after the registry has
/// agreed, and those cases are the ones a slow client causes.
#[inline]
pub fn record_drop(&self, reason: crate::native::link::DropReason) {
use crate::native::link::DropReason;
match reason {
DropReason::NoPort => self.drop_no_port.inc(),
DropReason::BacklogFull => self.drop_backlog_full.inc(),
DropReason::TooManyFlows => self.drop_too_many_flows.inc(),
DropReason::PendingQueueFull => self.drop_pending_queue_full.inc(),
DropReason::FlowQueueFull => self.drop_flow_queue_full.inc(),
DropReason::ArrivalQueueFull => self.drop_arrival_queue_full.inc(),
DropReason::ListenerNotReading => self.drop_listener_not_reading.inc(),
DropReason::ListenerGone => self.drop_listener_gone.inc(),
}
}
/// Count one datagram leaving a client for the mesh.
#[inline]
pub fn record_sent(&self, bytes: usize) {
self.sent_datagrams.inc();
self.sent_bytes.add(bytes as u64);
}
/// Count one datagram reaching a client from the mesh.
#[inline]
pub fn record_received(&self, bytes: usize) {
self.received_datagrams.inc();
self.received_bytes.add(bytes as u64);
}
/// Sample every counter into a serializable snapshot.
pub fn snapshot(&self) -> NativeStatsSnapshot {
NativeStatsSnapshot {
flows_opened: self.flows_opened.get(),
flows_accepted: self.flows_accepted.get(),
flows_closed: self.flows_closed.get(),
flows_expired: self.flows_expired.get(),
sent_datagrams: self.sent_datagrams.get(),
sent_bytes: self.sent_bytes.get(),
received_datagrams: self.received_datagrams.get(),
received_bytes: self.received_bytes.get(),
drop_no_port: self.drop_no_port.get(),
drop_backlog_full: self.drop_backlog_full.get(),
drop_too_many_flows: self.drop_too_many_flows.get(),
drop_pending_queue_full: self.drop_pending_queue_full.get(),
drop_flow_queue_full: self.drop_flow_queue_full.get(),
drop_arrival_queue_full: self.drop_arrival_queue_full.get(),
drop_listener_not_reading: self.drop_listener_not_reading.get(),
drop_listener_gone: self.drop_listener_gone.get(),
drop_oversize: self.drop_oversize.get(),
}
}
}
/// Atomic counter registry shared across the node via `Arc`.
///
/// Sole storage for the forwarding, discovery, tree, bloom, congestion,
/// and error counter families; these were migrated off `NodeStats`, which
/// now holds only the session, handshake, mmp, and transport families.
/// error, and native counter families; these were migrated off `NodeStats`,
/// which now holds only the session, handshake, mmp, and transport families.
#[derive(Default)]
pub struct MetricsRegistry {
pub forwarding: ForwardingMetrics,
@@ -525,6 +618,7 @@ pub struct MetricsRegistry {
pub bloom: BloomMetrics,
pub congestion: CongestionMetrics,
pub errors: ErrorMetrics,
pub native: NativeMetrics,
}
impl MetricsRegistry {
@@ -537,6 +631,54 @@ impl MetricsRegistry {
mod tests {
use super::*;
#[test]
fn native_record_drop_routes_every_reason_to_its_own_counter() {
use crate::native::link::DropReason;
let m = NativeMetrics::default();
for reason in [
DropReason::NoPort,
DropReason::BacklogFull,
DropReason::TooManyFlows,
DropReason::PendingQueueFull,
DropReason::FlowQueueFull,
DropReason::ArrivalQueueFull,
DropReason::ListenerNotReading,
DropReason::ListenerGone,
] {
m.record_drop(reason);
}
let snap = m.snapshot();
// Each reason lands in one counter and no reason lands in two, which is
// what a shared counter would hide.
assert_eq!(snap.drop_no_port, 1);
assert_eq!(snap.drop_backlog_full, 1);
assert_eq!(snap.drop_too_many_flows, 1);
assert_eq!(snap.drop_pending_queue_full, 1);
assert_eq!(snap.drop_flow_queue_full, 1);
assert_eq!(snap.drop_arrival_queue_full, 1);
assert_eq!(snap.drop_listener_not_reading, 1);
assert_eq!(snap.drop_listener_gone, 1);
// Nothing bumped the send-side refusal, which no DropReason reaches.
assert_eq!(snap.drop_oversize, 0);
}
#[test]
fn native_pending_and_flow_queue_drops_read_alike_to_a_client_and_apart_to_a_counter() {
use crate::native::link::DropReason;
// The two render identically on purpose, so the debug arrival command
// answers the same strings it always has. A counter must still tell
// them apart, because only one of them means a slow client.
assert_eq!(
DropReason::PendingQueueFull.as_str(),
DropReason::FlowQueueFull.as_str()
);
let m = NativeMetrics::default();
m.record_drop(DropReason::FlowQueueFull);
let snap = m.snapshot();
assert_eq!(snap.drop_flow_queue_full, 1);
assert_eq!(snap.drop_pending_queue_full, 0);
}
#[test]
fn forwarding_received_tracks_packets_and_bytes() {
let m = ForwardingMetrics::default();
+103
View File
@@ -476,6 +476,18 @@ pub struct Node {
/// Keyed by destination NodeAddr, bounded per-dest and total.
pending_tun_packets: HashMap<NodeAddr, VecDeque<Vec<u8>>>,
/// Native API registry: which local ports are held, and where an inbound
/// datagram goes. Reached only from the `rx_loop`, so it takes no lock.
native: crate::native::registry::Registry,
/// Native datagrams held per destination while its session establishes.
///
/// Deliberately **not** `pending_tun_packets`: that queue is drained
/// through `send_ipv6_packet`, which compresses its bytes as an IPv6
/// header, and it records neither a port nor a kind, so nothing could tell
/// a native datagram from an IPv6 packet once it was in there.
pending_native: HashMap<NodeAddr, VecDeque<crate::node::handlers::PendingNative>>,
// === Discovery ===
/// Discovery-subsystem state: recent-request dedup cache, in-flight
/// lookups, originator-side backoff, and transit-side forward limiter.
@@ -524,6 +536,12 @@ pub struct Node {
/// see [`Self::publish_entities_snapshot`] for the rationale.
entities_snapshot: std::sync::Arc<arc_swap::ArcSwap<crate::control::snapshot::EntitySnapshot>>,
/// Read-side snapshot of the native datagram API registry (flows
/// / listeners) that the `show_native_flows` query renders off the rx_loop.
/// Published from the tick; see [`Self::publish_native_snapshot`] for why
/// the registry itself cannot be shared instead.
native_snapshot: std::sync::Arc<arc_swap::ArcSwap<crate::control::snapshot::NativeSnapshot>>,
// === TUN Interface ===
/// TUN device state.
tun_state: TunState,
@@ -786,6 +804,12 @@ impl Node {
sessions: HashMap::new(),
identity_cache: HashMap::new(),
pending_tun_packets: HashMap::new(),
pending_native: HashMap::new(),
native: crate::native::registry::Registry::new(crate::native::registry::Limits {
per_flow: config.node.native_api.pending_per_flow,
backlog: config.node.native_api.backlog,
max_flows: config.node.native_api.max_flows,
}),
next_link_id: 1,
next_transport_id: 1,
stats: stats::NodeStats::new(),
@@ -800,6 +824,9 @@ impl Node {
entities_snapshot: std::sync::Arc::new(arc_swap::ArcSwap::from_pointee(
crate::control::snapshot::EntitySnapshot::empty(),
)),
native_snapshot: std::sync::Arc::new(arc_swap::ArcSwap::from_pointee(
crate::control::snapshot::NativeSnapshot::empty(),
)),
tun_state,
tun_name: None,
index_allocator: IndexAllocator::new(),
@@ -941,6 +968,12 @@ impl Node {
sessions: HashMap::new(),
identity_cache: HashMap::new(),
pending_tun_packets: HashMap::new(),
pending_native: HashMap::new(),
native: crate::native::registry::Registry::new(crate::native::registry::Limits {
per_flow: config.node.native_api.pending_per_flow,
backlog: config.node.native_api.backlog,
max_flows: config.node.native_api.max_flows,
}),
next_link_id: 1,
next_transport_id: 1,
stats: stats::NodeStats::new(),
@@ -955,6 +988,9 @@ impl Node {
entities_snapshot: std::sync::Arc::new(arc_swap::ArcSwap::from_pointee(
crate::control::snapshot::EntitySnapshot::empty(),
)),
native_snapshot: std::sync::Arc::new(arc_swap::ArcSwap::from_pointee(
crate::control::snapshot::NativeSnapshot::empty(),
)),
tun_state,
tun_name: None,
index_allocator: IndexAllocator::new(),
@@ -1563,6 +1599,7 @@ impl Node {
self.stats_snapshot.clone(),
self.routing_snapshot.clone(),
self.entities_snapshot.clone(),
self.native_snapshot.clone(),
)
}
@@ -1711,6 +1748,72 @@ impl Node {
// Publish the per-entity read view from the same tick, with
// `Vec<Arc<Row>>` structural sharing against the previous snapshot.
self.publish_entities_snapshot();
// Publish the native datagram API read view from the same tick.
self.publish_native_snapshot();
}
/// Project the native datagram API registry into a
/// [`NativeSnapshot`](crate::control::snapshot::NativeSnapshot) and publish
/// it via `ArcSwap`, so `show_native_flows` renders off the rx_loop.
///
/// **Publisher placement.** The registry is reached only from the rx_loop,
/// which is what lets the native receive path take no lock; sharing it with
/// the control task would mean locking it and giving that property back.
/// The tick is therefore the publisher, as it is for the other three cells.
/// The peer's npub needs nothing from `&Node`: the key rides on the
/// registry entry, captured where the session authenticated it.
fn publish_native_snapshot(&self) {
use crate::control::snapshot as snap;
let flows: Vec<snap::NativeFlowRow> = self
.native
.flows()
.into_iter()
.map(|view| snap::NativeFlowRow {
flow: view.flow,
peer: view.key.peer,
peer_key: view.pubkey,
local_port: view.key.local,
remote_port: view.key.remote,
established: view.established,
queued: view.queued,
since_ms: view.at,
})
.collect();
let listeners: Vec<snap::NativeListenerRow> = self
.native
.listeners()
.into_iter()
.map(|view| snap::NativeListenerRow {
local_port: view.port,
backlog: view.backlog,
})
.collect();
self.native_snapshot
.store(std::sync::Arc::new(snap::NativeSnapshot {
flows,
listeners,
}));
}
/// Borrow the native datagram API registry, read-only.
///
/// For the on-loop `show_native_flows` oracle. Every mutation still goes
/// through the rx_loop's own handler.
pub(crate) fn native(&self) -> &crate::native::registry::Registry {
&self.native
}
/// Test-only: reach the native registry mutably, so a test can populate it
/// the way the rx_loop's handler does without standing up a client task and
/// a socket pair. A parity test over an empty registry proves nothing about
/// a publisher that drops fields, which is why this exists.
#[cfg(test)]
pub(crate) fn native_registry_for_test(&mut self) -> &mut crate::native::registry::Registry {
&mut self.native
}
/// Resolve the npub of the spanning-tree root for `show_tree`'s `root_npub`.
+32
View File
@@ -436,6 +436,38 @@ pub struct CongestionStatsSnapshot {
pub kernel_drop_events: u64,
}
/// Native datagram API counters.
///
/// Eight of the `drop_*` fields are one per
/// [`DropReason`](crate::native::link::DropReason) variant, so a variant added
/// later has nowhere to hide. `drop_oversize` is the ninth and has no reason
/// because it is refused on the send side, before the registry sees the
/// datagram.
///
/// Present on every platform even though the API builds only on Linux and
/// FreeBSD, so `show_metrics` carries one schema everywhere and a consumer does
/// not branch on the host.
#[derive(Clone, Debug, Default, Serialize)]
pub struct NativeStatsSnapshot {
pub flows_opened: u64,
pub flows_accepted: u64,
pub flows_closed: u64,
pub flows_expired: u64,
pub sent_datagrams: u64,
pub sent_bytes: u64,
pub received_datagrams: u64,
pub received_bytes: u64,
pub drop_no_port: u64,
pub drop_backlog_full: u64,
pub drop_too_many_flows: u64,
pub drop_pending_queue_full: u64,
pub drop_flow_queue_full: u64,
pub drop_arrival_queue_full: u64,
pub drop_listener_not_reading: u64,
pub drop_listener_gone: u64,
pub drop_oversize: u64,
}
#[cfg(test)]
mod tests {
use super::*;
+3 -3
View File
@@ -437,9 +437,9 @@ pub(crate) fn should_apply_path_mtu(existing: Option<u16>, candidate: u16) -> bo
/// oldest entry first when the queue is at `per_dest` capacity. Pure transform
/// over a passed-in queue (the max-destinations cap is a shell-side map-level
/// check).
pub(crate) fn push_bounded_pending(
queue: &mut alloc::collections::VecDeque<Vec<u8>>,
packet: Vec<u8>,
pub(crate) fn push_bounded_pending<T>(
queue: &mut alloc::collections::VecDeque<T>,
packet: T,
per_dest: usize,
) {
if queue.len() >= per_dest {
+3 -1
View File
@@ -1,6 +1,8 @@
//! Utility modules.
//!
//! Shared infrastructure that doesn't belong to a specific protocol layer:
//! session index allocation and other cross-cutting concerns.
//! session index allocation, Unix socket binding, and other cross-cutting
//! concerns.
pub mod index;
pub mod sockbind;
+269
View File
@@ -0,0 +1,269 @@
//! Bind a Unix domain socket under the FIPS access policy.
//!
//! Every FIPS Unix socket needs the same four things before it can listen: a
//! parent directory that exists and whose ownership is known, any stale socket
//! file removed, a listener bound, and group access applied so members of the
//! `fips` group can reach it.
//!
//! The policy lives in one place so a change to it reaches every socket rather
//! than whichever copy an author happened to be looking at. All three sockets
//! use it: the control socket, the gateway control socket and the native
//! datagram API socket. The gateway previously kept its own copy, which had
//! already drifted in that it never set the parent's mode at all.
#[cfg(unix)]
use std::path::{Path, PathBuf};
#[cfg(unix)]
use tokio::net::UnixListener;
#[cfg(unix)]
use tracing::{debug, warn};
/// Bind a Unix listener at `path` under the FIPS access policy.
///
/// `what` names the socket for diagnostics: it is the noun in the "already in
/// use" error a caller sees when another process is listening there, and a
/// structured field on the directory and stale-socket log lines. The caller
/// emits its own "listening" line, so each socket keeps its own wording.
///
/// Creates missing ancestors, removes a stale socket file, binds, then applies
/// mode `0o770` to the socket and `0o750` to the parent when this bind owns the
/// parent. An `AddrInUse` error means a live listener already holds the path.
#[cfg(unix)]
pub fn bind(path: &Path, what: &str) -> Result<UnixListener, std::io::Error> {
// Creation is useful for diagnostics, but ownership is keyed to directory
// identity as well: systemd pre-creates /run/fips on every Linux service
// start and initially owns it as root:root.
let managed_parent = match path.parent() {
Some(parent) => {
let created = ensure_socket_parent(parent)?;
if created {
debug!(path = %parent.display(), socket = what, "Created private socket directory");
}
(created || crate::config::is_managed_socket_parent(parent)).then(|| parent.to_owned())
}
None => None,
};
if path.exists() {
remove_stale_socket(path, what)?;
}
let listener = UnixListener::bind(path)?;
set_socket_access(path, managed_parent.as_deref(), chown_to_fips_group)?;
Ok(listener)
}
/// Ensure the socket's parent exists and report whether this call created the
/// leaf directory.
///
/// `create_dir` gives us an atomic ownership decision: an `AlreadyExists`
/// result means another actor owns the existing directory, while success means
/// it is safe for this bind to apply FIPS ownership and mode. Missing ancestors
/// are created recursively, but only the requested leaf is later treated as the
/// socket's private directory.
#[cfg(unix)]
fn ensure_socket_parent(parent: &Path) -> Result<bool, std::io::Error> {
if parent.as_os_str().is_empty() {
return Ok(false);
}
match std::fs::create_dir(parent) {
Ok(()) => Ok(true),
Err(error) if error.kind() == std::io::ErrorKind::AlreadyExists => {
if parent.is_dir() {
Ok(false)
} else {
Err(error)
}
}
Err(error) if error.kind() == std::io::ErrorKind::NotFound => {
let ancestor = parent.parent().ok_or(error)?;
ensure_socket_parent(ancestor)?;
ensure_socket_parent(parent)
}
Err(error) => Err(error),
}
}
/// Apply access policy to a newly bound socket.
///
/// The socket is always group-owned. `managed_parent` is either a private
/// directory this bind created or a canonical FIPS runtime directory. A shared
/// or operator-owned existing parent is omitted so it retains its ownership and
/// mode.
///
/// `chown_to_fips_group` is a parameter so the policy can be tested without
/// requiring the `fips` group to exist on the machine running the tests.
#[cfg(unix)]
fn set_socket_access(
socket_path: &Path,
managed_parent: Option<&Path>,
mut chown_to_fips_group: impl FnMut(&Path),
) -> Result<(), std::io::Error> {
use std::os::unix::fs::PermissionsExt;
std::fs::set_permissions(socket_path, std::fs::Permissions::from_mode(0o770))?;
chown_to_fips_group(socket_path);
if let Some(parent) = managed_parent {
std::fs::set_permissions(parent, std::fs::Permissions::from_mode(0o750))?;
chown_to_fips_group(parent);
}
Ok(())
}
/// Remove a stale socket file.
///
/// If the file exists but no one is listening, remove it so we can bind. This
/// handles unclean daemon exits. A live listener yields `AddrInUse` instead, so
/// two daemons cannot silently take the same path.
#[cfg(unix)]
fn remove_stale_socket(path: &Path, what: &str) -> Result<(), std::io::Error> {
match std::os::unix::net::UnixStream::connect(path) {
Ok(_) => Err(std::io::Error::new(
std::io::ErrorKind::AddrInUse,
format!("{what} socket already in use: {}", path.display()),
)),
Err(_) => {
debug!(path = %path.display(), socket = what, "Removing stale socket");
std::fs::remove_file(path)?;
Ok(())
}
}
}
/// Set group ownership of a path to the `fips` group (best-effort).
///
/// A missing group is not an error: a source build on a developer machine has
/// no `fips` group, and the socket is still usable by its owner.
#[cfg(unix)]
fn chown_to_fips_group(path: &Path) {
use std::ffi::CString;
use std::os::unix::ffi::OsStrExt;
let group_name = CString::new("fips").unwrap();
let grp = unsafe { libc::getgrnam(group_name.as_ptr()) };
if grp.is_null() {
debug!(
"'fips' group not found, skipping chown for {}",
path.display()
);
return;
}
let gid = unsafe { (*grp).gr_gid };
let c_path = match CString::new(path.as_os_str().as_bytes()) {
Ok(p) => p,
Err(_) => return,
};
let ret = unsafe { libc::chown(c_path.as_ptr(), u32::MAX, gid) };
if ret != 0 {
warn!(
path = %path.display(),
error = %std::io::Error::last_os_error(),
"Failed to chown socket to 'fips' group"
);
}
}
/// Remove a socket file at teardown, ignoring a path that is already gone.
#[cfg(unix)]
pub fn cleanup(path: &PathBuf, what: &str) {
if !path.exists() {
return;
}
match std::fs::remove_file(path) {
Ok(()) => debug!(path = %path.display(), socket = what, "Socket file removed"),
Err(error) => {
warn!(path = %path.display(), socket = what, error = %error, "Failed to remove socket file")
}
}
}
#[cfg(all(test, unix))]
mod tests {
use super::{ensure_socket_parent, set_socket_access};
use std::os::unix::fs::PermissionsExt;
#[test]
fn parent_setup_distinguishes_existing_and_created_directories() {
let temp = tempfile::tempdir().unwrap();
let existing = temp.path().join("existing");
std::fs::create_dir(&existing).unwrap();
assert!(!ensure_socket_parent(&existing).unwrap());
let nested = temp.path().join("missing").join("fips");
assert!(ensure_socket_parent(&nested).unwrap());
assert!(nested.is_dir());
assert!(!ensure_socket_parent(&nested).unwrap());
}
#[test]
fn access_setup_leaves_an_existing_shared_parent_unchanged() {
let temp = tempfile::tempdir().unwrap();
let parent = temp.path().join("shared");
std::fs::create_dir(&parent).unwrap();
std::fs::set_permissions(&parent, std::fs::Permissions::from_mode(0o711)).unwrap();
let socket = parent.join("control.sock");
std::fs::File::create(&socket).unwrap();
let mut chowned = Vec::new();
set_socket_access(&socket, None, |path| chowned.push(path.to_path_buf())).unwrap();
assert_eq!(chowned, vec![socket.clone()]);
assert_eq!(
std::fs::metadata(&parent).unwrap().permissions().mode() & 0o777,
0o711
);
assert_eq!(
std::fs::metadata(&socket).unwrap().permissions().mode() & 0o777,
0o770
);
}
#[test]
fn access_setup_secures_a_new_private_parent() {
let temp = tempfile::tempdir().unwrap();
let parent = temp.path().join("fips");
std::fs::create_dir(&parent).unwrap();
let socket = parent.join("control.sock");
std::fs::File::create(&socket).unwrap();
let mut chowned = Vec::new();
set_socket_access(&socket, Some(&parent), |path| {
chowned.push(path.to_path_buf())
})
.unwrap();
assert_eq!(chowned, vec![socket, parent.clone()]);
assert_eq!(
std::fs::metadata(&parent).unwrap().permissions().mode() & 0o777,
0o750
);
}
#[test]
fn access_setup_secures_an_existing_managed_parent() {
let temp = tempfile::tempdir().unwrap();
let parent = temp.path().join("managed");
std::fs::create_dir(&parent).unwrap();
std::fs::set_permissions(&parent, std::fs::Permissions::from_mode(0o700)).unwrap();
let socket = parent.join("control.sock");
std::fs::File::create(&socket).unwrap();
let mut chowned = Vec::new();
set_socket_access(&socket, Some(&parent), |path| {
chowned.push(path.to_path_buf())
})
.unwrap();
assert_eq!(chowned, vec![socket, parent.clone()]);
assert_eq!(
std::fs::metadata(&parent).unwrap().permissions().mode() & 0o777,
0o750
);
}
}
+37 -4
View File
@@ -212,6 +212,7 @@ NAT_SUITES=(cone symmetric lan)
NOSTR_RELAY_SUITES=(nostr-publish-consume)
STUN_FAULTS_SUITES=(stun-faults)
DNS_RESOLVER_SUITES=(dns-resolver)
NATIVE_API_SUITES=(native-api)
DEB_INSTALL_SUITES=(deb-install)
TOR_SUITES=(tor-socks5 tor-directory)
@@ -263,6 +264,9 @@ list_suites() {
echo " Sidecar:"
for s in "${SIDECAR_SUITES[@]}"; do echo " $s"; done
echo ""
echo " Native API:"
for s in "${NATIVE_API_SUITES[@]}"; do echo " $s"; done
echo ""
echo " DNS resolver:"
for s in "${DNS_RESOLVER_SUITES[@]}"; do echo " $s"; done
echo ""
@@ -511,8 +515,12 @@ run_build() {
return 1
fi
info "cargo build --release"
if cargo build --release 2>&1; then
info "cargo build --release --bins --examples"
# --bins --examples rather than the bare default: the native datagram API's
# echo server is a cargo example, and the native-api harness's image needs
# it built here rather than separately. Naming --bins keeps the daemon and
# its tools in the build, which --examples alone would drop.
if cargo build --release --bins --examples 2>&1; then
record "build" 0
else
record "build" 1
@@ -620,7 +628,13 @@ install_binaries() {
cp target/release/fipsctl "$dest/fipsctl"
[[ -f target/release/fipstop ]] && cp target/release/fipstop "$dest/fipstop" || true
[[ -f target/release/fips-gateway ]] && cp target/release/fips-gateway "$dest/fips-gateway" || true
chmod +x "$dest/fips" "$dest/fipsctl"
# Not optional: the native-api harness runs native-echo as one end of its
# echo check and native-surface as its surface walk, so a missing one must
# fail the image build rather than fail a check later with the container
# exiting on a missing entrypoint.
cp target/release/examples/native-echo "$dest/native-echo"
cp target/release/examples/native-surface "$dest/native-surface"
chmod +x "$dest/fips" "$dest/fipsctl" "$dest/native-echo" "$dest/native-surface"
[[ -f "$dest/fipstop" ]] && chmod +x "$dest/fipstop" || true
[[ -f "$dest/fips-gateway" ]] && chmod +x "$dest/fips-gateway" || true
}
@@ -974,6 +988,20 @@ run_stun_faults() {
record "stun-faults" $rc
}
# Run the native datagram API harness.
#
# Reads FIPS_TEST_IMAGE, so it exercises this run's binaries rather than a
# separately built one. Its two-node check creates and removes its own docker
# network, so nothing here has to.
run_native_api() {
info "[native-api] Running native datagram API test"
if FIPS_TEST_IMAGE="$CI_IMAGE_TEST" bash testing/native-api/test.sh 2>&1; then
record "native-api" 0
else
record "native-api" 1
fi
}
# Run dns-resolver harness (multi-distro + e2e scenarios)
run_dns_resolver() {
info "[dns-resolver] Running multi-distro test (slow — builds per-distro images)"
@@ -1031,7 +1059,7 @@ run_integration() {
local _f
for _f in "$SCRIPT_DIR"/docker/*; do
case "$(basename "$_f")" in
fips|fipsctl|fipstop|fips-gateway) continue ;;
fips|fipsctl|fipstop|fips-gateway|native-echo|native-surface) continue ;;
esac
cp -a "$_f" "$CI_BUILD_CONTEXT/" || { record "docker-build" 1; return; }
done
@@ -1154,6 +1182,9 @@ run_integration() {
# Sidecar
run_sidecar
# Native datagram API (light — one single-node run plus a two-node pair)
run_native_api
# DNS resolver multi-distro suite (heavy — per-distro systemd images)
run_dns_resolver
@@ -1208,6 +1239,8 @@ run_suite() {
run_sidecar ;;
dns-resolver)
run_dns_resolver ;;
native-api)
run_native_api ;;
deb-install)
run_deb_install ;;
tor-socks5)
+2
View File
@@ -2,3 +2,5 @@ fips
fips-gateway
fipsctl
fipstop
native-echo
native-surface
+9 -2
View File
@@ -36,8 +36,15 @@ RUN printf '%s\n' \
'no-resolv' \
>> /etc/dnsmasq.conf
COPY fips fipsctl fipstop fips-gateway /usr/local/bin/
RUN chmod +x /usr/local/bin/fips /usr/local/bin/fipsctl /usr/local/bin/fipstop /usr/local/bin/fips-gateway
# native-echo and native-surface are the native datagram API's example
# programs, built as cargo examples rather than bins. The native-api harness
# runs native-echo as one end of a two-node check and native-surface as the
# assertion walk over the client's whole public surface, so the image carries
# both beside the daemon they talk to.
COPY fips fipsctl fipstop fips-gateway native-echo native-surface /usr/local/bin/
RUN chmod +x /usr/local/bin/fips /usr/local/bin/fipsctl /usr/local/bin/fipstop \
/usr/local/bin/fips-gateway /usr/local/bin/native-echo \
/usr/local/bin/native-surface
# Mirror systemd's RuntimeDirectory=fips so the daemon's resolver picks
# /run/fips/control.sock — matches production layout and the chaos sim
+14
View File
@@ -183,6 +183,20 @@ build_one() {
chmod +x "$ctx/$bin"
done
# The shared Dockerfile COPYs the native datagram API's example programs,
# which are cargo examples that older refs do not carry and no interop
# check runs: these images exercise the wire between daemon versions. The
# cargo invocation above names only the four bins for that reason. Stage a
# stub for each so the image still builds, and so anything that does reach
# for one says why it is not there.
for example in native-echo native-surface; do
printf '%s\n' \
'#!/bin/sh' \
"echo \"$example is not built into interop images\" >&2" \
'exit 1' > "$ctx/$example"
chmod +x "$ctx/$example"
done
docker build \
--label "fips.interop.slot=$slot" \
--label "fips.interop.ref=$ref" \
+182
View File
@@ -0,0 +1,182 @@
# Native Datagram API Harness
Checks for the experimental native datagram API: a client process opens a flow
to a remote pubkey over a Unix socket, receives a file descriptor, and sends and
receives datagrams on it with no IPv6 emulation and no TUN device.
Design of record: `design/native-api/v1-datagram-experiment.md` in the project
workspace, which is a separate tree from this repository. The feature is off by
default and Unix only.
## Shape
The client runs in **its own container**, reaching the daemon through a
bind-mounted `/run/fips`. That is the real deployment shape — a separate process
with its own filesystem opening the socket — rather than a test speaking to the
daemon from inside the daemon's container. It also makes the access policy
observable: the host sees the socket file and reads its mode directly.
The step scripts are Python rather than Rust so a check changes without
rebuilding the daemon, which is what keeps the outside-in loop fast. That buys
speed at the cost of covering nothing of the Rust surface a caller links
against, so two compiled programs run here as well, both built on
`fips::native::client`: `examples/native-echo.rs`, which arrived with A5 and
serves the echo check, and `examples/native-surface.rs`, which walks the whole
public surface against a live daemon.
The table covers this directory and the two example programs the driver runs.
| File | What it is |
| ---- | ---------- |
| `test.sh` | The driver. Holds the scenarios and the pass/fail accounting. |
| `client.py` | A thin RPC client. Runs a script of steps over one connection and checks the replies. |
| `control.py` | A thin control-socket client, used to read `show_native_flows` back while a flow is open. |
| `node.yaml` | One node with the API enabled, no TUN, no DNS, no peers. Turns the debug commands on. |
| `node-api-off.yaml` | The same node with the API disabled, for the default-off check. |
| `node-debug-off.yaml` | The API enabled and the debug commands left at their default, for the gate check. |
| `../../examples/native-echo.rs` | The echo server for `check_echo_round_trip`. A program shape to copy. |
| `../../examples/native-surface.rs` | The surface walk for `check_surface_walk`. An assertion harness, not a shape to copy. |
## Running
```bash
cargo build --release --bins --examples # the driver refuses a stale binary
./testing/native-api/test.sh
```
`FIPS_TEST_IMAGE` is used when set, which is how `ci-local.sh` passes its
per-run image. There is deliberately no `fips-test:latest` to fall back on, so a
consumer that stops reading the variable fails loudly. Without it the driver
builds a minimal image from the locally compiled binary.
The driver **refuses to run against a stale binary**. Three binaries are built
or read from this tree, the daemon and the two examples, and each is probed
against what it is actually built from: `src/` plus `Cargo.toml` for all three,
this directory's `*.py` because the harness client is bind-mounted live rather
than built in, and, for an example, its own `.rs` and no other. A guard rooted
only at `src/` would let a stale example pass a check written about new code.
A stale binary is the worst outcome available here: the checks would run and
report a verdict about code that is not the working tree's.
An example is probed against its own source rather than all of `examples/`
because cargo does not relink `target/release/fips` when only an example
changes. Probing the daemon against every example would leave it permanently
older than a just-edited one, and the rebuild the refusal prescribes would not
clear the condition.
All three binaries must come from one profile directory. `resolve_image` refuses
a profile that holds only the daemon, which is what a bare
`cargo build --release` leaves behind.
## Increments
The API is built outside-in, and this harness grows with it. Each increment's
checks must pass before the next one starts.
| # | What it covers | State |
| - | -------------- | ----- |
| A1 | The socket, its access mode, the line framing, the command validation, the reserved-port refusals, and that the API is off by default | present |
| A2 | Descriptor passing over `SCM_RIGHTS`, message boundaries, `poll` readability, close reaching the daemon, flow isolation | present |
| A3 | Port ownership across clients, listening and accepting, the dispatch order, and reclaim when a descriptor closes | present |
| A4 | The end-to-end path between two nodes, and that a queued datagram is not IPv6-compressed | present |
| A5 | Counters, `show_native_flows` read back over the control socket, the Rust client module and echo example, and the debug-command gate | present |
| A6 | Every public item of `fips::native::client` walked against a live daemon: the five setup entry points, all eight `ToFipsAddr` spellings, the deadlines, non-blocking mode, the descriptor traits, and the payload limit | present |
**The `"stub": true` marker is gone.** It meant "this flow reaches no peer",
and after A4 every flow does. `max_payload` is now the real limit — the
transport MTU less the FIPS encapsulation and the port header, 1362 bytes on a
1472-byte transport — and the end-to-end check asserts that number rather than
accepting whatever is reported.
The tightening it existed for happened three times. A1's `connect` checks failed
the moment A2 began returning a descriptor, because `client.py` treats an
unannounced descriptor as a defect rather than ignoring it. A1's `accept` and
`reject` checks failed when A3 gave those commands a real registry, since a flow
no listener announced became a refusal. And the remaining `stub` assertions
failed at A4 when the field disappeared. Checks that had quietly kept passing
would have been worth nothing.
**The `accept` and `reject` commands are gone**, and so is the `incoming` event.
A listener now returns its own descriptor, so it is pollable, accepting is one
`recvmsg` on it that carries the arriving flow's descriptor, and refusing a flow
is closing that descriptor. The command socket carries replies only, in command
order. A step names a listener descriptor with `keep_listener` and takes flows
off it with an `accept` step; every descriptor a reply carries must be named, or
the run fails rather than dropping a flow silently.
**The backlog is no longer the bound a client sees.** It bounds arrivals the
daemon has announced and not yet wired, and the daemon drains that queue itself,
so a listener that never accepts is bounded by its send buffer and by
`node.native_api.max_flows` instead. `check_backlog_is_not_the_clients_bound`
asserts the change; the drop paths behind the new bound are covered by the
daemon's own tests, because neither is a number a shell check can produce.
**Flow identifiers are assigned by the node and keep counting up for its
lifetime.** A check must capture one with `keep_flow` rather than assume a
literal, or it holds only for the first flow the daemon ever made.
## The surface walk
`check_surface_walk` runs `examples/native-surface.rs` against the shared
single node, last among the single-node checks. Its subject is the Rust surface
rather than the wire: until it existed, `FipsStream::connect`, `connect_from`,
`connect_at`, `FipsListener::bind` and `bind_at` had no coverage of any kind,
and every other public item was exercised only against the hand-written stand-in
daemon in the crate's unit tests. That stand-in has already hidden a real defect
once, by being kinder than the daemon, which is why the walk talks to the real
one.
It runs last because `check_ephemeral_allocation` asserts 49152, 49153 and 49154
as the first three ports the node ever hands out and the allocator is a
forward-only cursor. The walk therefore asserts only that its own ephemeral
ports are `>= 49152`, and takes its named ports from the otherwise unused
4800-4809 band.
**The check asserts three things, not one:** that the container exited 0, that
its completion line is there, and that the count in that line equals
`SURFACE_ASSERTIONS` in `test.sh`. The third is the anti-silence measure. The
binary prints the recorder's own counter rather than a literal, so an assertion
block that stopped running — a `#[cfg]` gate that no longer matches, an early
return — still exits 0 and still prints the line, and only the count betrays it.
The number is deliberately brittle: adding an assertion must force an edit in
`test.sh`, so the two cannot drift apart quietly.
**A hang has to become a red, and has to name itself.** The walk's own subjects
fail by blocking forever: a read deadline never applied to the descriptor, a
`set_nonblocking` that did nothing. The binary arms a 30-second watchdog that
prints the assertion it was in and exits 1, and `run_surface_at` bounds the
container at 60 seconds as a backstop for a wedge before that thread is armed.
`timeout 60 docker run` is **not** that backstop, which a break-check measured
rather than a reading of the manual. `timeout` signals the docker client, the
client proxies SIGTERM to the container, and the walk is PID 1 there with no
handler for it, so the kernel discards the signal: the container was still up
five minutes after the bound passed and `docker run` never returned. The helper
runs the container detached, polls its state, and removes it by force, since
`docker rm -f` is a SIGKILL and PID 1 cannot discard that.
## The two-node check
`check_end_to_end` is the only check that runs more than one node. It derives
two identities with `testing/lib/derive_keys.py`, brings both up on their own
docker network peered by npub, and sends a datagram from a client on one to a
client on the other.
Three orderings are waited on explicitly rather than assumed, each because
assuming it produced an intermittent failure:
- **The link forms** before either client runs, watched for by the spanning tree
adopting a parent. Not by a peer-promotion log line: on this path — a
configured peer, dialled outbound — that line is never emitted.
- **The listener has bound its port** before the sender starts, watched for in
the listener's own output. Launching it first is not the same as it having
registered.
- **The client runs unbuffered** (`python3 -u`). Without it the marker above
never reaches the log file, so the wait cannot see it and every run fails at
the gate meant to make the check reliable.
The payload is deliberately not a valid IPv6 packet, and it is sent before any
session exists so it goes through the native pending queue. If a native datagram
were ever routed through the TUN pending queue it would be handed to the IPv6
compressor, which would refuse it, and this check would fail. The trap is
asserted rather than trusted.
+441
View File
@@ -0,0 +1,441 @@
#!/usr/bin/env python3
"""Native datagram API client for the increment checks.
Speaks the line-delimited JSON command protocol on the daemon's native API
socket. A run takes a script: a list of steps sent over ONE connection. The
connection owns nothing — a flow lives until its own descriptor is closed, and a
listener until its own is — so the single connection is a convenience for the
checks rather than a lifetime the daemon respects. Descriptors are what keep
things alive, and this tool holds them until the step that closes them or until
it exits.
Kinds of step:
RPC step: {"command": str, "params": {...}?, "expect": {"dotted.key": val}?,
"keep_fd": name?, "keep_listener": name?, "keep_flow": name?}
Sends a command and checks the reply. `keep_fd` stores a flow
descriptor under that name, `keep_listener` a listener descriptor;
every reply that carries one must name it, because a descriptor
nothing named is a flow or a port silently dropped. `keep_flow`
stores the reply's data.flow_id.
A parameter or expectation whose value is the string "@name" is
replaced by the flow identifier stored under `name`. Identifiers
are assigned by the node and keep counting up for its lifetime, so
a check that asserted a literal 1 would hold only for the first
flow the daemon ever made.
Accept step: {"accept": listener, "keep_fd": name, "expect": {...}?,
"keep_flow": name?}
One recvmsg on a stored listener descriptor. There is no accept
command: an arriving flow is one SOCK_SEQPACKET message on the
listener itself, carrying the flow's descriptor as ancillary data
and the arrival object as its payload. Expectations are checked
against that object, whose peer is an npub and never a hex address.
Sleep step: {"sleep": seconds}
Holds every descriptor open for a while, which is what a check that
reads the daemon's own view of a live flow needs.
Flow step: {"fd": name, ...} operating on a stored descriptor:
"write": hex, "repeat": n? send n datagrams of those bytes
"read": n, "expect_bytes": hex?, "sizes": [..]?
read n datagrams and check them
"readable": bool check poll readability now
"close": true close the descriptor
`readable` and `close` work on a listener descriptor too: a
listener is pollable, and closing it unbinds its port.
Reading is per-datagram: the descriptor is SOCK_SEQPACKET, so one recv is one
datagram. A check that reads three and gets one concatenated blob is a real
failure, not a quirk of the tool.
Usage:
client.py --socket PATH --script '<json list of steps>'
client.py --socket PATH --script-file steps.json
Exit 0 when every expectation holds, 1 otherwise, 2 on a connection failure.
"""
from __future__ import annotations
import argparse
import array
import json
import os
import select
import socket
import sys
import time
from typing import Any
# A flow's descriptor and a listener's are both AF_UNIX SOCK_SEQPACKET, so the
# wrap happens to be the same for both. Naming the roles anyway is the point:
# the next descriptor kind that is not one of these must not be wrapped
# correctly by accident.
FLOW = "flow"
LISTENER = "listener"
def recvfds(sock: socket.socket, bufsize: int, maxfds: int) -> tuple[bytes, list[int]]:
"""One recvmsg, returning its payload and whatever descriptors it carried.
Written out rather than calling `socket.recv_fds`, which takes a `flags`
argument and never forwards it to `recvmsg`: MSG_CMSG_CLOEXEC passed to that
helper does nothing, and a descriptor the harness kept would then survive
into any child process it forked. Measured on CPython 3.12 by reading
FD_CLOEXEC back with `fcntl.F_GETFD` after each of the two calls.
"""
fds = array.array("i")
data, ancillary, _flags, _addr = sock.recvmsg(
bufsize, socket.CMSG_LEN(maxfds * fds.itemsize), socket.MSG_CMSG_CLOEXEC
)
for level, kind, payload in ancillary:
if level == socket.SOL_SOCKET and kind == socket.SCM_RIGHTS:
# Truncated to whole descriptors: the kernel may cut the array
# short, and a partial one names nothing.
fds.frombytes(payload[: len(payload) - (len(payload) % fds.itemsize)])
return data, list(fds)
class Protocol(Exception):
"""The daemon broke the local protocol, so the run cannot continue."""
class Client:
"""One connection to the native API socket, plus the descriptors it holds."""
def __init__(self, path: str, timeout: float) -> None:
"""Connect to the socket at `path`, failing after `timeout` seconds."""
self.sock = socket.socket(socket.AF_UNIX, socket.SOCK_STREAM)
self.sock.settimeout(timeout)
self.sock.connect(path)
self.timeout = timeout
self.buf = b""
# Complete lines, oldest first, each with the descriptor it arrived
# with. See `fill` for the rule that decides which line that is.
self.lines: list[list[Any]] = []
self.fds: dict[str, tuple[socket.socket, str]] = {}
self.flows: dict[str, int] = {}
def call(self, command: str, params: dict | None) -> tuple[dict, int | None]:
"""Send one command; return the decoded reply and any descriptor.
The socket carries replies only, in command order, so the next complete
line is this command's answer and there is nothing to separate out.
"""
request: dict[str, Any] = {"command": command}
if params is not None:
request["params"] = params
self.sock.sendall(json.dumps(request).encode() + b"\n")
line, fd = self.line()
return json.loads(line), fd
def line(self) -> tuple[bytes, int | None]:
"""Take the next complete line, reading until one is available."""
while not self.lines:
self.fill()
line, fd = self.lines.pop(0)
return line, fd
def fill(self) -> None:
"""One recvmsg, split into lines, with any descriptor placed by the rule.
A DESCRIPTOR BELONGS TO THE LAST COMPLETE LINE OF THE READ THAT CARRIED
IT, never to the next line the reader assembles. A recvmsg returning
ancillary data ends exactly at the end of the sendmsg that carried it,
but it may begin with any amount of data written before it, so a reader
that attached the descriptor to the first line it completed would hand a
flow to the wrong reply. Both reply kinds carry a descriptor now, so
this is reachable rather than theoretical.
A read that carries a descriptor and completes no line is reported
rather than guessed at: holding it would mean choosing a later line for
it, and choosing wrong loses a flow with no error anywhere.
"""
chunk, fds = recvfds(self.sock, 65536, 4)
if not chunk:
for stray in fds:
# Closed rather than leaked: nothing can name it now.
os.close(stray)
raise ConnectionError("daemon closed the connection")
self.buf += chunk
produced = 0
while b"\n" in self.buf:
line, self.buf = self.buf.split(b"\n", 1)
self.lines.append([line, None])
produced += 1
if not fds:
return
# This protocol never sends two at once. Extras are closed rather than
# left open with no owner.
for stray in fds[1:]:
os.close(stray)
if produced == 0:
os.close(fds[0])
raise Protocol("a descriptor arrived on a read that completed no line")
self.lines[-1][1] = fds[0]
def accept(self, listener: str) -> tuple[dict, int]:
"""Take the next arriving flow off a stored listener descriptor.
One recvmsg, one arrival: SOCK_SEQPACKET means the message carries
exactly its own descriptor, so the association rule the RPC socket needs
does not arise here. The payload has no trailing newline, because the
message boundary is the framing.
"""
sock = self.held(listener, LISTENER)
data, fds = recvfds(sock, 65536, 1)
if not fds:
raise Protocol(f"{listener!r} produced an arrival with no descriptor")
if not data:
os.close(fds[0])
raise Protocol(f"{listener!r} produced a descriptor with no arrival")
return json.loads(data), fds[0]
def held(self, name: str, want: str) -> socket.socket:
"""Return a stored descriptor, refusing one of the wrong kind."""
if name not in self.fds:
raise Protocol(f"no descriptor named {name!r}")
sock, role = self.fds[name]
if role != want:
raise Protocol(f"{name!r} is a {role} descriptor, not a {want} one")
return sock
def keep(self, name: str, fd: int, role: str) -> None:
"""Store a received descriptor under `name`, wrapped for its kind."""
sock = socket.socket(socket.AF_UNIX, socket.SOCK_SEQPACKET, fileno=fd)
sock.settimeout(self.timeout)
self.fds[name] = (sock, role)
def close(self) -> None:
"""Close every descriptor, then the connection itself."""
for sock, _role in self.fds.values():
sock.close()
self.sock.close()
def substitute(value: Any, flows: dict[str, int]) -> Any:
"""Replace every "@name" with the flow identifier stored under `name`."""
if isinstance(value, str) and value.startswith("@"):
name = value[1:]
if name not in flows:
raise KeyError(f"no flow captured as {name!r}")
return flows[name]
if isinstance(value, dict):
return {key: substitute(item, flows) for key, item in value.items()}
if isinstance(value, list):
return [substitute(item, flows) for item in value]
return value
def dig(value: Any, dotted: str) -> Any:
"""Read a dotted path out of a decoded reply, or None where it is absent."""
for key in dotted.split("."):
if not isinstance(value, dict) or key not in value:
return None
value = value[key]
return value
def check(reply: dict, expect: dict) -> list[str]:
"""Return one message per expectation the reply does not satisfy."""
problems = []
for dotted, wanted in expect.items():
got = dig(reply, dotted)
if got != wanted:
problems.append(f"{dotted}: wanted {wanted!r}, got {got!r}")
return problems
def store(client: Client, step: dict, body: dict, fd: int | None) -> list[str]:
"""Store what a step asked to keep, reporting a descriptor nobody named."""
problems: list[str] = []
keep = step.get("keep_flow")
if keep is not None:
# A reply nests the identifier under `data`; an arrival message is the
# object itself. One reader for both, because a step should not have to
# know which produced it.
flow = dig(body, "data.flow_id")
if flow is None:
flow = body.get("flow_id")
if flow is None:
problems.append("keep_flow: nothing carried a flow_id")
else:
client.flows[keep] = flow
wanted = [(step.get("keep_fd"), FLOW), (step.get("keep_listener"), LISTENER)]
named = [(name, role) for name, role in wanted if name is not None]
if len(named) > 1:
if fd is not None:
os.close(fd)
problems.append("a step named both keep_fd and keep_listener")
elif named and fd is None:
problems.append(f"{named[0][0]!r}: no descriptor arrived to keep")
elif named:
client.keep(named[0][0], fd, named[0][1])
elif fd is not None:
# Leaving it unnamed would leak a flow or a held port for the rest of
# the run, with nothing to say so.
os.close(fd)
problems.append("a descriptor arrived that the step did not name")
return problems
def run_rpc(client: Client, step: dict) -> list[str]:
"""Send one command and report what did not hold."""
command = step["command"]
try:
params = substitute(step.get("params"), client.flows)
expect = substitute(step.get("expect", {}), client.flows)
except KeyError as error:
return [str(error)]
reply, fd = client.call(command, params)
problems = check(reply, expect)
problems += store(client, step, reply, fd)
if problems:
problems.append(f"reply: {json.dumps(reply)}")
return problems
def run_accept(client: Client, step: dict) -> list[str]:
"""Take one arrival off a listener and report what did not hold."""
try:
arrival, fd = client.accept(step["accept"])
except socket.timeout:
return [f"timed out waiting for an arrival on {step['accept']!r}"]
problems = check(arrival, substitute(step.get("expect", {}), client.flows))
problems += store(client, step, arrival, fd)
if problems:
problems.append(f"arrival: {json.dumps(arrival)}")
return problems
def run_sleep(step: dict) -> list[str]:
"""Hold every descriptor open for a while, failing nothing."""
time.sleep(float(step["sleep"]))
return []
def run_flow(client: Client, step: dict) -> list[str]:
"""Operate on a stored descriptor and report what did not hold."""
name = step["fd"]
if name not in client.fds:
return [f"no descriptor named {name!r}"]
flow, _role = client.fds[name]
problems: list[str] = []
if "readable" in step:
ready, _, _ = select.select([flow], [], [], 0.25)
got = bool(ready)
if got != step["readable"]:
problems.append(f"readable: wanted {step['readable']}, got {got}")
if "write" in step:
payload = bytes.fromhex(step["write"])
for _ in range(step.get("repeat", 1)):
flow.send(payload)
if "read" in step:
wanted = bytes.fromhex(step["expect_bytes"]) if "expect_bytes" in step else None
sizes = []
for index in range(step["read"]):
try:
got = flow.recv(65536)
except socket.timeout:
problems.append(f"read {index}: timed out waiting for a datagram")
break
sizes.append(len(got))
if wanted is not None and got != wanted:
problems.append(
f"read {index}: wanted {wanted.hex()}, got {got.hex()}"
)
if "sizes" in step and sizes != step["sizes"]:
problems.append(f"sizes: wanted {step['sizes']}, got {sizes}")
if step.get("close"):
flow.close()
del client.fds[name]
return problems
def label_of(step: dict) -> str:
"""The name a step is reported under, which callers wait on by substring."""
if "fd" in step:
return f"fd {step['fd']}"
if "accept" in step:
return f"accept {step['accept']}"
if "sleep" in step:
return f"sleep {step['sleep']}"
return step.get("command", "?")
def main() -> int:
"""Run the script against the socket and report every failing step."""
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--socket", required=True, help="native API socket path")
group = parser.add_mutually_exclusive_group(required=True)
group.add_argument("--script", help="steps as a JSON list")
group.add_argument("--script-file", help="file holding the steps as a JSON list")
parser.add_argument(
"--timeout",
type=float,
default=5.0,
help="socket timeout in seconds (default: 5)",
)
args = parser.parse_args()
text = args.script
if text is None:
with open(args.script_file, encoding="utf-8") as handle:
text = handle.read()
steps = json.loads(text)
try:
client = Client(args.socket, args.timeout)
except OSError as error:
print(f"connect to {args.socket} failed: {error}", file=sys.stderr)
return 2
failures = 0
try:
for index, step in enumerate(steps):
label = label_of(step)
try:
if "fd" in step:
problems = run_flow(client, step)
elif "accept" in step:
problems = run_accept(client, step)
elif "sleep" in step:
problems = run_sleep(step)
else:
problems = run_rpc(client, step)
except (OSError, ConnectionError, Protocol, json.JSONDecodeError) as error:
print(f"step {index} ({label}): {error}", file=sys.stderr)
return 2
if problems:
failures += 1
print(f"step {index} ({label}) FAILED", file=sys.stderr)
for problem in problems:
print(f" {problem}", file=sys.stderr)
else:
print(f"step {index} ({label}) ok")
finally:
client.close()
return 1 if failures else 0
if __name__ == "__main__":
sys.exit(main())
+152
View File
@@ -0,0 +1,152 @@
#!/usr/bin/env python3
"""Control socket client for the native API checks.
Speaks the control socket's line-delimited JSON protocol: one request line, one
response line, one command per connection. That is a different protocol from the
native API's — whose connection outlives its first command and carries
descriptors — which is why this is a separate tool rather than another step type
in `client.py`.
It runs in a container for the same reason the native API client does: both
sockets are bound 0o770 root:fips, so the host user running the harness cannot
open either, while a client container mounting the same directory runs as root
and can.
Usage:
control.py --socket PATH --command show_native_flows
control.py --socket PATH --command show_native_flows --expect data.flows.0.local_port=4501
An expectation is `dotted.path=value` or `dotted.path>value`. A path segment of
digits indexes a list, so `data.flows.0.state` reads the first flow's state. The
value is parsed as JSON where it parses and taken as a plain string where it
does not, so both `4501` and `established` say what they look like. `>` compares
numerically and is how a check asserts a counter moved without pinning a total
that the wire is free to reach by more than one datagram.
Exit 0 when every expectation holds, 1 when one does not, 2 on a connection or
protocol failure. The reply is printed either way, because a check that missed
one field needs the whole object to say why.
"""
from __future__ import annotations
import argparse
import json
import socket
import sys
from typing import Any
MISSING = object()
"""Distinguishes a path that is absent from one whose value is JSON null."""
def query(path: str, command: str, timeout: float) -> dict:
"""Send one command over its own connection and return the decoded reply."""
conn = socket.socket(socket.AF_UNIX, socket.SOCK_STREAM)
conn.settimeout(timeout)
try:
conn.connect(path)
conn.sendall((json.dumps({"command": command}) + "\n").encode())
buffer = b""
while b"\n" not in buffer:
chunk = conn.recv(65536)
if not chunk:
raise ConnectionError("the control socket closed before replying")
buffer += chunk
return json.loads(buffer.split(b"\n", 1)[0].decode())
finally:
conn.close()
def dig(value: Any, dotted: str) -> Any:
"""Read a dotted path out of a reply, indexing lists on a numeric segment."""
for key in dotted.split("."):
if isinstance(value, list):
if not key.isdigit() or int(key) >= len(value):
return MISSING
value = value[int(key)]
elif isinstance(value, dict) and key in value:
value = value[key]
else:
return MISSING
return value
def parse(text: str) -> Any:
"""Read an expectation's value as JSON, falling back to a plain string."""
try:
return json.loads(text)
except json.JSONDecodeError:
return text
def split(expectation: str) -> tuple[str, str, Any]:
"""Split `path=value` or `path>value` on whichever operator comes first."""
cuts = [(expectation.index(op), op) for op in ("=", ">") if op in expectation]
if not cuts:
raise ValueError(f"{expectation!r} carries neither '=' nor '>'")
at, op = min(cuts)
return expectation[:at], op, parse(expectation[at + 1:])
def check(reply: dict, expectation: str) -> str | None:
"""Return a message when the expectation does not hold, else None."""
dotted, op, wanted = split(expectation)
got = dig(reply, dotted)
if got is MISSING:
return f"{dotted}: wanted {op}{wanted!r}, but the path is absent"
if op == "=":
return None if got == wanted else f"{dotted}: wanted {wanted!r}, got {got!r}"
if isinstance(got, bool) or not isinstance(got, (int, float)) or got <= wanted:
return f"{dotted}: wanted more than {wanted!r}, got {got!r}"
return None
def main() -> int:
"""Query the control socket and report every expectation that did not hold."""
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--socket", required=True, help="control socket path")
parser.add_argument("--command", required=True, help="control command to send")
parser.add_argument(
"--expect",
action="append",
default=[],
metavar="PATH=VALUE",
help="dotted-path expectation; repeatable",
)
parser.add_argument(
"--timeout",
type=float,
default=5.0,
help="socket timeout in seconds (default: 5)",
)
args = parser.parse_args()
try:
reply = query(args.socket, args.command, args.timeout)
except (OSError, ConnectionError, json.JSONDecodeError) as error:
print(f"{args.command} on {args.socket} failed: {error}", file=sys.stderr)
return 2
print(json.dumps(reply))
if reply.get("status") != "ok":
print(f"{args.command} answered {reply.get('status')!r}", file=sys.stderr)
return 1
try:
problems = [
message
for message in (check(reply, expectation) for expectation in args.expect)
if message is not None
]
except ValueError as error:
print(f" {error}", file=sys.stderr)
return 2
for message in problems:
print(f" {message}", file=sys.stderr)
return 1 if problems else 0
if __name__ == "__main__":
sys.exit(main())
+28
View File
@@ -0,0 +1,28 @@
# Node configuration with the native API explicitly disabled.
#
# The companion to node.yaml, used by the check that no socket appears when the
# feature is off. It names the same socket_path deliberately: if the disable
# were ignored, the socket would land exactly where the check looks.
node:
identity:
nsec: "2102030405060708090a0b0c0d0e0f101112131415161718191a1b1c1d1e1f20"
native_api:
enabled: false
socket_path: "/run/fips/api.sock"
# Named for the same reason as socket_path: everything node.yaml turns on
# is turned on here too, so `enabled` is the only difference between the
# two and the only thing the check can be observing.
debug_commands: true
tun:
enabled: false
dns:
enabled: false
transports:
udp:
bind_addr: "0.0.0.0:2121"
mtu: 1472
+27
View File
@@ -0,0 +1,27 @@
# Node configuration with the native API on and the debug commands off.
#
# The companion to node.yaml, used by the check that inject, stats and arrive
# are refused where the gate is closed. It differs from node.yaml in one thing:
# `debug_commands` is not named at all, so what the check observes is the
# packaged default rather than an explicit false. A change that flipped that
# default would then red this check, which is the whole point of leaving the
# key out.
node:
identity:
nsec: "2102030405060708090a0b0c0d0e0f101112131415161718191a1b1c1d1e1f20"
native_api:
enabled: true
socket_path: "/run/fips/api.sock"
tun:
enabled: false
dns:
enabled: false
transports:
udp:
bind_addr: "0.0.0.0:2121"
mtu: 1472
+32
View File
@@ -0,0 +1,32 @@
# Node configuration for the native datagram API checks.
#
# One node, no peers, no TUN and no DNS: increment A1 exercises the API socket
# itself, so the node needs to start and bind the socket and nothing else.
# Dropping TUN is what lets the container run without NET_ADMIN or
# /dev/net/tun, which keeps the check cheap and portable.
node:
identity:
nsec: "2102030405060708090a0b0c0d0e0f101112131415161718191a1b1c1d1e1f20"
native_api:
enabled: true
# Named explicitly rather than left to the default resolver, because the
# check bind-mounts this directory from the host in order to read the
# socket's mode and to reach it from the client container.
socket_path: "/run/fips/api.sock"
# inject, stats and arrive. Off in a packaged node, and on here because
# nine of the checks below drive the receive and dispatch paths through
# them. node-debug-off.yaml is the companion that leaves them off.
debug_commands: true
tun:
enabled: false
dns:
enabled: false
transports:
udp:
bind_addr: "0.0.0.0:2121"
mtu: 1472
+1417
View File
File diff suppressed because it is too large Load Diff
+13 -4
View File
@@ -57,25 +57,34 @@ if [ "$UNAME_S" = "Darwin" ]; then
fi
echo "Building FIPS for Linux (release) using cargo-zigbuild..."
cargo zigbuild --release --target "$CARGO_TARGET" --manifest-path="$PROJECT_ROOT/Cargo.toml" "${CARGO_BUILD_ARGS[@]}"
cargo zigbuild --release --bins --examples --target "$CARGO_TARGET" --manifest-path="$PROJECT_ROOT/Cargo.toml" "${CARGO_BUILD_ARGS[@]}"
TARGET_DIR="$TARGET_ROOT/$CARGO_TARGET/release"
else
echo "Building FIPS (release)..."
cargo build --release --manifest-path="$PROJECT_ROOT/Cargo.toml" "${CARGO_BUILD_ARGS[@]}"
cargo build --release --bins --examples --manifest-path="$PROJECT_ROOT/Cargo.toml" "${CARGO_BUILD_ARGS[@]}"
TARGET_DIR="$TARGET_ROOT/release"
fi
# --bins --examples above rather than the bare default: the native datagram
# API's example programs are cargo examples, and the shared image carries them.
# Naming --bins keeps the daemon and its tools, which --examples alone would drop.
echo "Copying binaries to $DOCKER_DIR/"
cp "$TARGET_DIR/fips" "$DOCKER_DIR/fips"
cp "$TARGET_DIR/fipsctl" "$DOCKER_DIR/fipsctl"
cp "$TARGET_DIR/fips-gateway" "$DOCKER_DIR/fips-gateway"
[ -f "$TARGET_DIR/fipstop" ] && cp "$TARGET_DIR/fipstop" "$DOCKER_DIR/fipstop" || true
chmod +x "$DOCKER_DIR/fips" "$DOCKER_DIR/fipsctl" "$DOCKER_DIR/fips-gateway"
# Not optional: the Dockerfile COPYs both native API examples unconditionally,
# so a missing one must fail here, where the cause is legible, rather than
# inside docker build.
cp "$TARGET_DIR/examples/native-echo" "$DOCKER_DIR/native-echo"
cp "$TARGET_DIR/examples/native-surface" "$DOCKER_DIR/native-surface"
chmod +x "$DOCKER_DIR/fips" "$DOCKER_DIR/fipsctl" "$DOCKER_DIR/fips-gateway" \
"$DOCKER_DIR/native-echo" "$DOCKER_DIR/native-surface"
[ -f "$DOCKER_DIR/fipstop" ] && chmod +x "$DOCKER_DIR/fipstop" || true
echo "Done. Binaries at $DOCKER_DIR/{fips,fipsctl,fipstop,fips-gateway}"
echo "Done. Binaries at $DOCKER_DIR/{fips,fipsctl,fipstop,fips-gateway,native-echo,native-surface}"
if [ "$BUILD_DOCKER" = true ]; then
echo ""