mirror of
https://github.com/jmcorgan/fips.git
synced 2026-10-05 19:18:25 +00:00
macOS does not implement SOCK_SEQPACKET for AF_UNIX, so the listener was gated to Linux and FreeBSD and a Mac got no API at all. It now uses SOCK_DGRAM there, which macOS does implement and which keeps the message boundaries the API's contract with its clients rests on. Both kernels were measured rather than reasoned about, and the Linux answer alone refuted the replacement the source had proposed. On Linux 6.8 a connected SOCK_DGRAM pair reports a closed peer not at all: revents stays empty and recv returns EAGAIN, which is exactly what an idle socket with a live peer does. SOCK_SEQPACKET on the same kernel sets POLLHUP and returns a zero-byte read, which is what the receive path keyed on. Darwin does report the close, with ECONNRESET, errno 54, and does not set POLLHUP. So the receive path treats ECONNRESET as end of file alongside the existing POLLHUP rule. One rule accepting either signal is correct on both kernels, where a rule split by platform would be silently wrong on whichever one it guessed at. EAGAIN is deliberately not in that company: it means the socket is empty and the peer alive, so it stays an error and the caller waits again. The measurements are asserted rather than only written down, so a kernel that gains or loses the signal reds a test and reopens the question instead of leaving a stale comment behind. Each carries its result in the assertion message, since a passing test prints nothing and a negative result would otherwise be as uninformative as no result: the close probe reports the poll return, the whole revents bitmask broken out by flag, the recv result and the errno, which is enough to write the real rule without another round trip. A portable test walks datagram sizes upward, because Darwin bounds a unix-domain datagram with the net.local.dgram.maxdgram sysctl, whose default is small and which is a system tunable rather than something this process controls, while the API advertises 1362 bytes to its clients. Every test recv in the seqpacket suite is bounded in time. This is not tidying. Simulating the Darwin configuration on a kernel that does not report the close, the end-of-file test hung for over ten minutes rather than failing, and a hang is not a red: it would have wedged the macOS runner with no diagnostic instead of naming the assertion. The same simulation now fails by name in five seconds. Three things in the native tree compiled on one platform only, all of them in code that had never been built for Darwin before. suseconds_t is i64 on Linux and i32 there, so the timeval microseconds field is a cast, matching the tv_sec line above it; it cannot truncate, because subsec_micros is below 1_000_000 by construction, and the cast is the only form that compiles on both, since From does not exist for the narrower width and try_from is a clippy error on the wider one. MSG_CMSG_CLOEXEC does not exist on Apple, so the recvmsg flags are chosen per platform and each received descriptor is marked close-on-exec with fcntl where there is no flag to pass; a failure to set it is reported rather than ignored, since the descriptor is live either way and the caller must not be told the receive was clean. socketpair takes SOCK_CLOEXEC in its type argument on Linux and FreeBSD and rejects it on macOS, so Darwin sets FD_CLOEXEC with a second fcntl. Both windows between a call and its fcntl are stated in the code rather than closed, since the daemon spawns no child on this path, and a test asserts both halves of a pair are close-on-exec on every platform, because the failure is a silent descriptor leak into a child and nothing else would report it. Every libc item the native tree uses was then checked against the crate's own Apple definitions rather than from memory, and those two constants are the only ones absent. The close difference turned out to be unhandled in six further places, and the whole native API agrees on it now. Darwin reports a closed AF_UNIX SOCK_DGRAM peer as ECONNRESET, and a later send on the disconnected survivor as EDESTADDRREQ, where Linux SOCK_SEQPACKET gives EPIPE on a write and a zero-byte read plus POLLHUP on a read. Each site below promised one of those spellings and saw another. - The client's recv and send passed ECONNRESET through, so a closed daemon half surfaced as errno 54 against the EPIPE the documentation promises. The translation is in one function in seqpacket rather than at each call site, since only this one condition has two spellings. - accept propagated the same errno instead of its documented EPIPE. The listener's read reports a closed peer as the empty chunk both of its callers already read as the far end going away, which leaves accept's contract true on both platforms without either caller knowing which it is on. - why() classified a failed hand-off by BrokenPipe alone, so every ordinary macOS listener close was counted under the counter an operator reads to find a client that stopped reading. It recognises all three errnos now, with a test over each. - A full client buffer ended a flow's only writer. On Linux that never arrives, because the send reports EAGAIN and waits for the client to drain; Darwin has no sender-side queue to wait on and reports ENOBUFS on the send itself. Returning left the registration, the port and the reader alive while every later inbound datagram was counted as a full queue for the rest of the flow's life, and a client that resumed reading never recovered. The datagram is dropped instead, which is what a datagram API does when the far end cannot take it. - The flow pair was never sized, and the two kernels charge a queued message to different ends: Linux to the sender's SO_SNDBUF, BSD to the receiver's so_rcv. Sizing only the sender, as the listener pair does, left the flow pair bounded on Darwin by a system default small enough that a batch held for an arriving client could not fit, and the whole flow was destroyed before its client ever saw it. Both halves are sized now. - peer_hung_up polled with an empty events field, on the rule that POLLHUP is reported whether or not it is requested. That holds on Linux, where it was measured, and not on Darwin, where a poll requesting nothing registers no filter. Nothing observable depended on it, because ECONNRESET arrives first and both callers act on it earlier. The cost was elsewhere: three assertions written as tripwires for a change in Darwin's behaviour could not fail there, which is a guard that executes and proves nothing. Requesting POLLIN fixes the function and the guards together. One difference is not an errno at all, and reading the kernel source rather than a manual page is what found it. Darwin's unp_disconnect sets SS_CANTRCVMORE and runs soisdisconnected on both ends for SOCK_STREAM. For SOCK_DGRAM it removes the reflink, clears SS_ISCONNECTED and stops: no sorwakeup, no socantrcvmore, no soisdisconnected. The closing peer deposits ECONNRESET in the survivor's so_error and wakes no knote. The registration is edge-triggered and was made while the socket was healthy, so nothing re-evaluates it, and recv awaited readiness before its syscall, which left the ECONNRESET arm sitting behind an await that never returns. A client closing its descriptor left the daemon's reader parked for ever, and the flow's port and registry entry held for the node's lifetime. recv reads before it waits now, because the latched error is visible to a syscall and only to a syscall, so the attempt that precedes the wait is what sees a close that has already happened. A close can also land while the task is parked, which no first attempt can catch, so on Darwin the wait is bounded and the syscall retried; the error is latched until a read consumes it, so the bound sets how long a dead flow holds its port rather than deciding whether the close is seen at all. On Linux this is one extra recv returning EAGAIN before the wait and changes nothing else, and everywhere else the readiness is authoritative and the wait stays unbounded. The three tests this predicted are the three that had failed: end of file on a closed client half, a listener's port unbound on close, and one flow's port freed while its connection stays open. One test asserted a delivery detail rather than the rule it exists to guard. a_descriptor_lands_on_the_last_complete_line_of_the_read_that_carried_it asserted that a plain write and the sendmsg following it arrive in one recvmsg. Linux coalesces them, so the read returns both lines and the descriptor together; Darwin stops a stream read at the ancillary boundary, so the plain line arrives by itself and the descriptor-bearing line comes on the next read. The rule the module rests on is unaffected, and Darwin satisfies it more easily than Linux, because the read it arrives on holds nothing later. The test fills until both lines are queued and asserts the rule instead of the number of reads it took. The client compiled in /run/fips/api.sock on every platform, and macOS has no /run for that path to be in. The daemon never had this problem: it resolves its socket at startup by looking for a directory, and its macOS branch lands on /var/run/fips. The constant is conditional the same way now, so a client that is told nothing looks where a packaged daemon on its own platform actually is. The reference documentation described that branch as FreeBSD-only and describes both. Windows stays excluded and cannot be included: it has no SCM_RIGHTS, so there is no way to pass a descriptor to another process at all, which is the whole mechanism rather than a detail of it. The platform statements in the source and in the shipped documentation all named Linux and FreeBSD and name macOS now, including the configuration reference, the security reference, the how-to and the walkthrough. The how-to also states how far the testing goes, because the person who would meet the gap first is the one enabling the API on a Mac. The end-to-end suite drives a client container against a node container over a shared volume, which is a Linux arrangement, so the socket lifecycle, the descriptor hand-off across a process boundary and the reclaiming of a port when a client exits are covered on macOS by unit tests rather than by anything that runs a daemon and a client as two real processes. That is a gap in testing and not a known defect, and it is a coverage gap rather than a discharged risk. The same place names the socket-type difference, since a reader who knows the descriptor is SOCK_DGRAM there can make sense of a close arriving as a different errno than the Linux documentation elsewhere describes. The changelog entry for the API is revised rather than followed by a second one: it now names the socket type each platform uses and the two end-of-file signals the receive path accepts. The entry describes what the release ships rather than the order the commits landed in.
344 lines
10 KiB
Markdown
344 lines
10 KiB
Markdown
# Native Datagram API Walkthrough
|
|
|
|
A side trip. You will run two FIPS nodes on one machine, write a listening
|
|
program and a connecting program against the native datagram API, and watch one
|
|
datagram cross between them. Then you will look at the flow from the outside
|
|
with `fipsctl` while it is still open.
|
|
|
|
Nothing here touches the public mesh, and nothing needs root. Both nodes run
|
|
with no TUN device and no DNS, peered directly over loopback UDP, so the whole
|
|
session lives in one scratch directory you delete at the end.
|
|
|
|
**This is not part of the numbered progression.** Take it any time. It assumes
|
|
you can build the daemon from source and can read Rust; it does not assume you
|
|
have worked through the tutorials.
|
|
|
|
**The API is experimental.** Names, fields and the command set may change
|
|
without a deprecation cycle. It is Linux, FreeBSD and macOS only.
|
|
|
|
## What you will end up with
|
|
|
|
- Two nodes, each with its own identity, control socket and API socket.
|
|
- A listening program that holds port 4600 and echoes one datagram per flow.
|
|
- A connecting program that opens a flow to the other node's public key and
|
|
gets its datagram back.
|
|
- A reading of `fipsctl show native-flows` taken while the flow is open.
|
|
|
|
## Step 1: Build the daemon and its tools
|
|
|
|
From a checkout of the FIPS source:
|
|
|
|
```sh
|
|
cargo build --release --bins
|
|
```
|
|
|
|
That gives you `target/release/fips` and `target/release/fipsctl`. Put them on
|
|
your path for the rest of this walkthrough:
|
|
|
|
```sh
|
|
export PATH="$PWD/target/release:$PATH"
|
|
```
|
|
|
|
## Step 2: Make two identities
|
|
|
|
`keygen -s` prints a keypair to stdout and writes nothing:
|
|
|
|
```sh
|
|
fipsctl keygen -s
|
|
```
|
|
|
|
```text
|
|
nsec1...
|
|
npub1...
|
|
```
|
|
|
|
Run it twice and keep both pairs. Call them A and B. You need each node's
|
|
`nsec` for its own config, and each node's `npub` for the *other* node's peer
|
|
entry.
|
|
|
|
```sh
|
|
mkdir -p ~/napi-lab/a ~/napi-lab/b
|
|
cd ~/napi-lab
|
|
```
|
|
|
|
## Step 3: Write the two configs
|
|
|
|
Node A, at `~/napi-lab/a/fips.yaml`. Substitute A's `nsec` and B's `npub`:
|
|
|
|
```yaml
|
|
node:
|
|
identity:
|
|
nsec: "<A's nsec>"
|
|
control:
|
|
socket_path: "/home/YOU/napi-lab/a/control.sock"
|
|
native_api:
|
|
enabled: true
|
|
socket_path: "/home/YOU/napi-lab/a/api.sock"
|
|
|
|
tun:
|
|
enabled: false
|
|
|
|
dns:
|
|
enabled: false
|
|
|
|
transports:
|
|
udp:
|
|
bind_addr: "127.0.0.1:2121"
|
|
mtu: 1472
|
|
|
|
peers:
|
|
- npub: "<B's npub>"
|
|
alias: "node-b"
|
|
addresses:
|
|
- transport: udp
|
|
addr: "127.0.0.1:2122"
|
|
```
|
|
|
|
Node B, at `~/napi-lab/b/fips.yaml`, is the mirror image: B's `nsec`, A's
|
|
`npub`, its own sockets under `b/`, `bind_addr` on `2122`, and its peer address
|
|
pointing at `2121`.
|
|
|
|
> **Use absolute paths.** The daemon does not resolve a socket path relative to
|
|
> the config file. Putting an `nsec` in a config is fine for a throwaway lab
|
|
> node like this one; for anything you keep, use
|
|
> [../how-to/persistent-identity.md](../how-to/persistent-identity.md) instead.
|
|
|
|
Disabling TUN and DNS is what lets both nodes run as your own user. A node with
|
|
a TUN device needs `CAP_NET_ADMIN`, and this walkthrough does not need one:
|
|
the native API is the path that does not go through the IPv6 adapter.
|
|
|
|
## Step 4: Start both nodes
|
|
|
|
In two terminals:
|
|
|
|
```sh
|
|
fips --config ~/napi-lab/a/fips.yaml
|
|
```
|
|
|
|
```sh
|
|
fips --config ~/napi-lab/b/fips.yaml
|
|
```
|
|
|
|
Each should log that it bound its API socket:
|
|
|
|
```text
|
|
Native API socket listening on /home/YOU/napi-lab/a/api.sock
|
|
```
|
|
|
|
In a third terminal, confirm the two found each other:
|
|
|
|
```sh
|
|
fipsctl -s ~/napi-lab/a/control.sock show peers
|
|
```
|
|
|
|
Wait for B to appear with a session. The link forms over loopback UDP and
|
|
usually takes a second or two. **Wait for it before going on**: a `connect` on
|
|
a flow contacts no peer, so it will succeed whether or not the link is up, and
|
|
the datagram would simply be held and then dropped.
|
|
|
|
## Step 5: Write the listening program
|
|
|
|
Make a crate next to the lab directory:
|
|
|
|
```sh
|
|
cargo new --bin napi-listen
|
|
cd napi-listen
|
|
```
|
|
|
|
Point it at your FIPS checkout in `Cargo.toml`:
|
|
|
|
```toml
|
|
[dependencies]
|
|
fips = { path = "/path/to/your/fips/checkout" }
|
|
```
|
|
|
|
`src/main.rs`:
|
|
|
|
```rust
|
|
//! Hold a port and echo one datagram per flow.
|
|
|
|
use fips::native::client::{FipsListener, FipsStream};
|
|
use std::env;
|
|
use std::error::Error;
|
|
use std::path::Path;
|
|
use std::thread;
|
|
|
|
/// Return one datagram to where it came from, then release the flow.
|
|
fn serve(flow: FipsStream) {
|
|
// Sized at the flow's own limit, so no datagram it can carry is
|
|
// truncated on the way in and echoed short.
|
|
let mut buf = vec![0u8; flow.max_payload()];
|
|
match flow.recv(&mut buf) {
|
|
Ok(len) => {
|
|
let _ = flow.send(&buf[..len]);
|
|
println!("returned {len} bytes to {}", flow.peer_addr());
|
|
}
|
|
Err(error) => eprintln!("receiving: {error}"),
|
|
}
|
|
// Returning drops the flow, which closes its descriptor. That is what
|
|
// releases the flow at the daemon; there is no close call to make.
|
|
}
|
|
|
|
fn main() -> Result<(), Box<dyn Error>> {
|
|
let socket = env::args().nth(1).ok_or("usage: napi-listen <api-socket>")?;
|
|
let listener = FipsListener::bind_at(Path::new(&socket), 4600)?;
|
|
println!("holding {}", listener.local_addr());
|
|
|
|
for arrival in listener.incoming() {
|
|
// Detached rather than joined: the accept loop must not wait on one
|
|
// peer, and the thread owns everything it touches.
|
|
match arrival {
|
|
Ok(flow) => drop(thread::spawn(move || serve(flow))),
|
|
Err(error) => eprintln!("accepting: {error}"),
|
|
}
|
|
}
|
|
Ok(())
|
|
}
|
|
```
|
|
|
|
Run it against node B:
|
|
|
|
```sh
|
|
cargo run -- ~/napi-lab/b/api.sock
|
|
```
|
|
|
|
```text
|
|
holding npub1...:4600
|
|
```
|
|
|
|
**Note what `serve` does not do.** It does not loop reading until the flow
|
|
closes. The v1 wire carries no half-close, so nothing peer-driven would ever
|
|
end that loop; it would hold a thread and a flow slot per peer until the
|
|
process died. One exchange per flow is the program's own decision, and making
|
|
it is mandatory. See
|
|
[../how-to/use-the-native-datagram-api.md](../how-to/use-the-native-datagram-api.md#four-things-that-will-bite-you).
|
|
|
|
**Note also what `incoming()` does not do.** It never returns `None`, and a
|
|
failed accept arrives as an `Err` item rather than ending the iteration. Writing
|
|
`let flow = arrival?;` here would exit the loop on the first transient error,
|
|
which is a different shape from `TcpListener` habits.
|
|
|
|
## Step 6: Write the connecting program
|
|
|
|
```sh
|
|
cd ..
|
|
cargo new --bin napi-connect
|
|
cd napi-connect
|
|
```
|
|
|
|
Same dependency line. `src/main.rs`:
|
|
|
|
```rust
|
|
//! Open a flow to a peer, exchange one datagram, and exit.
|
|
|
|
use fips::native::client::FipsStream;
|
|
use std::env;
|
|
use std::error::Error;
|
|
use std::io;
|
|
use std::path::Path;
|
|
use std::time::Duration;
|
|
|
|
/// How long to wait for the peer's answer before giving up on it.
|
|
const REPLY: Duration = Duration::from_secs(10);
|
|
|
|
fn main() -> Result<(), Box<dyn Error>> {
|
|
let mut args = env::args().skip(1);
|
|
let (Some(socket), Some(peer)) = (args.next(), args.next()) else {
|
|
return Err("usage: napi-connect <api-socket> <peer-npub>".into());
|
|
};
|
|
|
|
// One setup call, and it contacts no peer: the daemon registers the flow
|
|
// locally and hands back the descriptor it rides on. Success here says
|
|
// nothing about the peer existing, being reachable, or listening.
|
|
let flow = FipsStream::connect_at(Path::new(&socket), 0, (peer, 4600))?;
|
|
println!("{} -> {}", flow.local_addr(), flow.peer_addr());
|
|
|
|
// Before the first recv and not after it, because the peer may never
|
|
// answer at all and the deadline is what makes that a failure rather
|
|
// than a hang.
|
|
flow.set_read_timeout(Some(REPLY))?;
|
|
|
|
flow.send(b"hello")?;
|
|
|
|
let mut buf = vec![0u8; flow.max_payload()];
|
|
match flow.recv(&mut buf) {
|
|
Ok(len) => println!("{}", String::from_utf8_lossy(&buf[..len])),
|
|
Err(error) if error.kind() == io::ErrorKind::WouldBlock => {
|
|
return Err(format!("no answer from {} in {REPLY:?}", flow.peer_addr()).into());
|
|
}
|
|
Err(error) => return Err(error.into()),
|
|
}
|
|
Ok(())
|
|
}
|
|
```
|
|
|
|
Run it against node A, naming node B's npub:
|
|
|
|
```sh
|
|
cargo run -- ~/napi-lab/a/api.sock <B's npub>
|
|
```
|
|
|
|
```text
|
|
npub1...:49152 -> npub1...:4600
|
|
hello
|
|
```
|
|
|
|
The listener's terminal reports the other half:
|
|
|
|
```text
|
|
returned 5 bytes to npub1...:49152
|
|
```
|
|
|
|
That datagram went from your connecting program, into node A over a Unix
|
|
socket, across loopback UDP inside an encrypted FSP session, into node B, and
|
|
out to your listening program on another Unix socket. No IPv6 address and no
|
|
TUN device was involved anywhere in it.
|
|
|
|
## Step 7: Watch a flow from the outside
|
|
|
|
The exchange above is over in milliseconds. To look at a live flow, make the
|
|
connector hold one open: add a `std::thread::sleep(Duration::from_secs(60));`
|
|
before the final `Ok(())` and run it again.
|
|
|
|
While it sleeps:
|
|
|
|
```sh
|
|
fipsctl -s ~/napi-lab/b/control.sock show native-flows
|
|
```
|
|
|
|
You get every flow node B holds, with its ports, its queue depth and its age,
|
|
plus every bound listener and its backlog. The counters are in:
|
|
|
|
```sh
|
|
fipsctl -s ~/napi-lab/b/control.sock stats metrics
|
|
```
|
|
|
|
under `native`, where the `drop_*` fields separate a datagram refused for
|
|
having no listening port from one dropped because a client was not reading fast
|
|
enough. Those counters are the only way to see a drop: **nothing on the API
|
|
surface reports one to your program.**
|
|
|
|
## Step 8: Clean up
|
|
|
|
Stop both daemons with Ctrl-C, then:
|
|
|
|
```sh
|
|
rm -rf ~/napi-lab
|
|
```
|
|
|
|
The identities were only ever in those config files, so removing the directory
|
|
removes them. Nothing was written outside it and nothing was published to any
|
|
relay.
|
|
|
|
## Where to go next
|
|
|
|
- [../how-to/use-the-native-datagram-api.md](../how-to/use-the-native-datagram-api.md)
|
|
— the same ground as a recipe, including enabling the API on a real node and
|
|
the security posture that grants
|
|
- [../reference/native-api.md](../reference/native-api.md)
|
|
— every type and method, the errno table, the ceilings, and what happens to
|
|
data that disappears
|
|
- [../how-to/write-a-native-api-client.md](../how-to/write-a-native-api-client.md)
|
|
— doing all of this from C, Python or Go, where there is no client library
|
|
and the obligations become yours
|