Merge branch 'maint' into master

Bring v0.5.2 up from maint. The package version stays 0.6.0-dev;
the rustls 0.23.45 update merges in. CHANGELOG.md gains the [0.5.2]
section, and master's Unreleased keeps only entries not released in
0.5.2. README.md keeps master's install section and status, naming
v0.5.2 as the current release.
This commit is contained in:
Johnathan Corgan
2026-09-29 00:11:14 +00:00
23 changed files with 1023 additions and 662 deletions
+457 -420
View File
@@ -74,19 +74,6 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
destroy-and-recreate, that an `optional` interface never moves node health,
and that absence is logged once on the edge rather than once per retry.
#### Gateway
- The gateway counts sessions on a kernel without `/proc/net/nf_conntrack`.
When the file is absent it dumps the conntrack table over netlink, as
`conntrack -L` does, so a mapping carrying traffic is pinned instead of
being reclaimed on its TTL and grace period alone. Kernels built without
`CONFIG_NF_CONNTRACK_PROCFS`, such as Ubuntu's, had session pinning off.
- The gateway says at startup whether it can read conntrack sessions. It
reads the table once, as each tick does, and logs either the source it read
or that no source is readable and session pinning is off. An operator on a
kernel with no readable source learned this only from a warning at the first
failed tick.
#### Node lifecycle
- Transport-medium change detection, controlled by the new `node.netmon.*`
@@ -359,77 +346,6 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
is still the oldest Debian-family distribution, which is no longer the oldest
distribution outright.
#### Identity and config
- An ephemeral node no longer writes `fips.key`. It wrote the private key of
an identity it discards at every restart to that file, overwriting any key
already there, including an operator's key when `persistent: true` had been
forgotten. It now writes only `fips.pub`, so the running npub stays visible,
and holds the private key in memory only. A `fips.key` found at an ephemeral
start is moved to `fips.key.unused` with a warning, so an ephemeral node logs
that warning once, on its first start after the upgrade; if that name is
taken or the rename fails, the file is left in place and a warning says so.
For a stable identity, set `node.identity.persistent: true` and restart: the
first persistent start generates and saves a key, so the npub changes once,
at that restart, and is stable from then on. Starting once in ephemeral mode
and then pinning the key it wrote no longer works.
#### Gateway
- The gateway's default DNS listen address is now `[::1]:5365`; it was
`[::1]:5353`, the mDNS port, which the daemon's LAN rendezvous, Avahi and
systemd-resolved can hold. On OpenWrt the init script now points dnsmasq at
whatever port `gateway.dns.listen` sets, and an upgrade rewrites the
previously shipped `listen: "[::1]:5353"` line. On other hosts, a resolver
you configured by hand to forward `.fips` to `[::1]:5353` must now forward
to `[::1]:5365`, or set `gateway.dns.listen: "[::1]:5353"` to keep the old
port. The gateway warns at startup when it is configured on 5353.
#### Packaging (Debian)
- An upgrade of the `.deb` now reapplies the firewall ruleset in place. Until
now an upgrade reloaded nothing, so a changed `/etc/fips/fips.nft` took
effect only at the next reboot or manual restart, and a restart deletes the
`fips` table and leaves the mesh interface unfiltered until the ruleset is
loaded again. `fips-firewall.service`, in both the Debian and the plain
systemd unit, gains a reload that replaces the ruleset in one transaction,
and the postinst reloads the unit only when it is already active, so an
upgrade never turns the firewall on for a host that has not opted in. A
reload that fails leaves the previous ruleset in place and is reported; the
upgrade goes on.
#### Packaging (AUR)
- The AUR publish on a release tag now waits until every package workflow of
that tag has succeeded. It used to push the new `pkgver` while the Linux,
macOS, Windows, OpenWrt and FreeBSD packages were still building; at v0.5.1
the AUR was updated while the release had 15 of its 17 assets. Because the
AUR package pins the tag's source archive, withdrawing a bad release after
that point left the AUR package unbuildable. A failed or cancelled package
run now stops the publish, and one that has not finished within an hour
fails it.
#### Dependencies
- The lockfile moves `chacha20` from 0.10.1 to 0.10.2, because 0.10.1 is yanked.
It arrives through `rand`, a direct dependency,
so it sits on the built path rather than off to one side. The requirement in
`Cargo.toml` already admitted 0.10.2, so this is a lockfile change and no code
changed with it. **This is not a security fix**: `cargo audit` reports nothing
against `chacha20` at either version, and 0.10.1 was withdrawn by its
maintainer rather than flagged by an advisory. What it buys is that a fresh
checkout can resolve the lockfile without reaching for a yanked version.
#### Documentation (native API)
- The native API documentation now says that on Linux an empty datagram sent
immediately before a close may read as end of file, and is then not
delivered. Linux carries the flow on `SOCK_SEQPACKET`, where a zero-length
datagram that is the last message before a close cannot be told apart from
the close; macOS and FreeBSD carry it on `SOCK_DGRAM` and are not affected.
The `Received::Datagram` rustdoc, which said an empty datagram is never a
close, now says where the exception applies.
### Removed
- **Source-breaking for consumers of the library crate**: `ActivePeer` no
@@ -450,15 +366,6 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
#### Node lifecycle
- A DNS responder or TUN thread that dies now degrades the node's published
health, and a dead responder's address is retracted. The responder's exit
report followed a loop that never returns, so it could not run, and a panic
in the responder or in either TUN thread unwound past its report. The node
kept reporting healthy with the child gone and kept publishing the DNS
address with nothing answering on it. Each child now runs inside a wrapper
that catches a panic, logs it, and reports the exit either way. A deliberate
stop still reports nothing.
- Losing an interface no longer leaves its peers in the routing table. The
peers stayed in the registry, the routes through them stayed selectable, and
the node kept advertising reachability it no longer had — so transit traffic
@@ -520,13 +427,6 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
#### Data plane and transports
- A configured `ble:` transport that this build cannot construct is now
reported. The only warning for it was compiled into test builds alone, where
logging is compiled out, so macOS, Windows, FreeBSD, OpenWrt and other musl
builds, and Android without a BLE radio armed before start, dropped the block
silently while reporting healthy. The daemon now warns once per configured
instance at startup, naming the reason.
- A peer that stops reading can no longer stall the node. TCP, Tor, Nym and
BLE wrote to their links directly from the caller's task, and a write
blocks once the peer's receive window and this node's send buffer are both
@@ -587,19 +487,318 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
platforms with the connected-socket fast path); elsewhere the heartbeat alone
carries the new address.
- Two inbound TCP connections that share a peer address but arrive on different
local addresses no longer share one pool entry. The kernel names a connection
by its four-tuple, so a listener on a wildcard address, which is what the
shipped configuration binds, can accept two connections whose peer `ip:port`
is the same on two different local addresses. The pool was keyed by the peer
address alone: the second connection's entry replaced the first's while the
inbound-connection counter counted both, the first connection's teardown then
removed the second's entry, and the second's own teardown found nothing to
remove, so the counter ended one above the connections it counts. That counter
gates the inbound connection limit, so a host repeating the collision could
hold it at the limit and lock out further inbound TCP connections until the
daemon restarted. Inbound entries now carry the accepted socket's local
address in their pool key as well as the remote one.
## [0.5.2] - 2026-09-28
### Added
#### Gateway
- The gateway counts sessions on a kernel without `/proc/net/nf_conntrack`.
When the file is absent it dumps the conntrack table over netlink, as
`conntrack -L` does, so a mapping carrying traffic is pinned instead of
being reclaimed on its TTL and grace period alone. Kernels built without
`CONFIG_NF_CONNTRACK_PROCFS`, such as Ubuntu's, had session pinning off.
- The gateway says at startup whether it can read conntrack sessions. It reads
the table once, as each tick does, and logs either the source it read or
that no source is readable and session pinning is off. In v0.5.1 an
unreadable source counted as zero sessions and nothing was logged, so an
operator on a kernel with no readable source had no way to tell.
### Changed
#### Identity and config
- The shipped `/etc/fips/hosts` no longer lists `test-us03-next`.
- An ephemeral node no longer writes `fips.key`. It wrote the private key of
an identity it discards at every restart to that file, overwriting any key
already there, including an operator's key when `persistent: true` had been
forgotten. It now writes only `fips.pub`, so the running npub stays visible,
and holds the private key in memory only. A `fips.key` found at an ephemeral
start is moved to `fips.key.unused` with a warning, so an ephemeral node logs
that warning once, on its first start after the upgrade; if that name is
taken or the rename fails, the file is left in place and a warning says so.
For a stable identity, set `node.identity.persistent: true` and restart. The
first persistent start uses the `fips.key` it finds; one already moved to
`fips.key.unused` can be renamed back to `fips.key` first, as the warning
says. With no key file, the first persistent start generates and saves one,
so the npub changes once, at that restart, and is stable from then on.
Starting once in ephemeral mode and then pinning the key it wrote no longer
works.
#### Gateway
- The gateway's default DNS listen address is now `[::1]:5365`; it was
`[::1]:5353`, the mDNS port, which the daemon's LAN rendezvous, Avahi and
systemd-resolved can hold. On OpenWrt the init script now points dnsmasq at
whatever port `gateway.dns.listen` sets, and an upgrade rewrites the
previously shipped `listen: "[::1]:5353"` line. On other hosts the upgrade
leaves `fips.yaml` alone, and a config that sets `gateway.dns.listen`, as
the v0.5.1 example config and deployment guide did, keeps its port; the
gateway warns at startup when it is configured on 5353. Where `fips.yaml`
does not set it, a resolver you configured by hand to forward `.fips` to
`[::1]:5353` must now forward to `[::1]:5365`, or set
`gateway.dns.listen: "[::1]:5353"` to keep the old port.
#### Linux packages
- An upgrade of the `.deb` now reapplies the firewall ruleset in place. Until
now an upgrade reloaded nothing, so a changed `/etc/fips/fips.nft` took
effect only at the next reboot or manual restart, and a restart deletes the
`fips` table and leaves the mesh interface unfiltered until the ruleset is
loaded again. `fips-firewall.service`, in both the Debian and the plain
systemd unit, gains a reload that replaces the ruleset in one transaction,
and the postinst reloads the unit only when it is already active, so an
upgrade never turns the firewall on for a host that has not opted in. A
reload that fails leaves the previous ruleset in place and is reported; the
upgrade goes on.
- The AUR publish on a release tag now waits until every package workflow of
that tag has succeeded. It used to push the new `pkgver` while the Linux,
macOS, Windows, OpenWrt and FreeBSD packages were still building; at v0.5.1
the AUR was updated while the release had 15 of its 17 assets. Because the
AUR package pins the tag's source archive, withdrawing a bad release after
that point left the AUR package unbuildable. A failed or cancelled package
run now stops the publish, and one that has not finished within an hour
fails it.
#### Native datagram API
- The native API documentation now says that on Linux an empty datagram sent
immediately before a close may read as end of file, and is then not
delivered. Linux carries the flow on `SOCK_SEQPACKET`, where a zero-length
datagram that is the last message before a close cannot be told apart from
the close; macOS and FreeBSD carry it on `SOCK_DGRAM` and are not affected.
The `Received::Datagram` rustdoc, which said an empty datagram is never a
close, now says where the exception applies.
#### Dependencies
- The lockfile moves `chacha20` from 0.10.1 to 0.10.2, because 0.10.1 is yanked.
It arrives through `rand`, a direct dependency,
so it sits on the built path rather than off to one side. The requirement in
`Cargo.toml` already admitted 0.10.2, so this is a lockfile change and no code
changed with it. **This is not a security fix**: `cargo audit` reports nothing
against `chacha20` at either version, and 0.10.1 was withdrawn by its
maintainer rather than flagged by an advisory. What it buys is that a fresh
checkout can resolve the lockfile without reaching for a yanked version.
### Fixed
#### Identity and config
- A persistent node whose identity key path cannot be examined now refuses to
start instead of coming up under a new identity. `Path::exists` reports false
both for a key that is absent and for one whose metadata cannot be read, so a
key symlinked onto a volume that did not mount, or one in a directory the
daemon cannot search, read as a first boot: the node generated a fresh
identity, failed to store it, and carried on under an npub that every peer
whose allowlist names the old one refuses. Only a `NotFound` result is now
treated as an absence; any other failure to stat the path aborts the start and
names the path. A dangling symlink likewise aborts rather than being replaced.
The legacy `/etc/fips/fips.key` lookup follows the same rule.
- Replacing the peer list at runtime with `Node::update_peers` now updates
everything that reads peer aliases. `.fips` names, peer ACL entries written as
an alias, and peer display names kept following the aliases the node started
with, so a new peer's alias did not resolve, a removed one still did, and a
deny entry naming an alias moved to another key kept denying the old key and
admitted the new one. They now follow the new peer list, with the hosts file
still taking precedence as it does at startup.
#### Gateway
- The NAT table is rebuilt in one netlink transaction. A rebuild deleted the
`fips_gateway` table in a batch of its own, discarded that batch's result,
and only then sent the batch that recreated the table, the chains, the
`fips0` masquerade and every per-mapping rule. Between the two sends the
gateway had no NAT at all, and a recreate the kernel refused left the table
absent for good, taking down forwarding for every existing mapping rather
than failing the one change that was being made. The delete and the recreate
now share a single batch, which the kernel applies as one transaction, so a
refused rebuild leaves the previous table in the packet path. The rules sent
are unchanged.
- The gateway's NAT rebuild no longer fails once the table holds more than
about 105 mappings. Each rebuild is one netlink batch. From about 105
mappings the default socket buffers could not hold its acknowledgements, so
rebuilds were logged as failed although they had taken effect. Past about
313 mappings the buffers could not hold the batch itself, and new `.fips`
names past that count got a virtual IP with no translation. In releases with
the gateway through 0.5.1, a rebuild past about 313 mappings also deleted the
whole `fips_gateway` table, which stopped every mapping, the `fips0`
masquerade and the port forwards. The rebuild now sizes its send buffer to
the batch and requests one acknowledgement per batch, and NAT errors now
name the kernel errno. A rebuild that still fails is logged, and the next
successful rebuild installs the mapping.
- Conntrack sessions are matched by address rather than by text, so live
traffic pins a gateway mapping again. The session count searched each
`/proc/net/nf_conntrack` line for `dst=` followed by the virtual IP in its
compressed form (`fd01::1`), while the kernel prints tuples in the full
uncompressed form (`dst=fd01:0000:0000:0000:0000:0000:0000:0001`), so the
count was zero for every mapping on every kernel. Nothing pinned an in-use
mapping, and one whose client did not re-query DNS was reclaimed about two
minutes after its last DNS reference while its traffic was still flowing.
Each `dst=` value is now parsed as an address and compared as one.
- The conntrack table is read once per tick instead of once per mapping, and
the read happens off the runtime thread. The whole file was read and scanned
for each mapping in turn, while the pool lock was held, on the same
single-threaded runtime that serves DNS. The tick now takes one snapshot with
a blocking task before it takes the lock, and the pool does a map lookup per
mapping.
- A conntrack source that cannot be read is reported. It still counts as zero
sessions for every mapping, as it always has, so reclamation keeps working
rather than pinning the whole pool; but the first failure and each change of
outcome after it are now logged, so an unreadable source is no longer
indistinguishable from an idle one. A source that fails identically every
tick is logged at debug rather than warn on a repeat.
- A DNS query that refreshes a draining mapping now cancels its old grace
period. Previously the address could be reclaimed while the client's renewed
DNS answer was still valid. The mapping now survives the full renewed TTL
and a fresh grace period before it can be reused. Contributed by Martti
Malmi (#169).
- `fips-gateway` exits when its DNS listener cannot bind, or stops while the
gateway runs, instead of staying up with `.fips` resolution dead, so systemd
or procd restarts it or reports it failed. This applies to a gateway used
only for port forwards too. An "address in use" error names the service
likely to hold the port. On an OpenWrt access point with the gateway
enabled, the gateway had lost its port to the daemon's own mDNS responder
and `.fips` names stopped resolving with nothing reported. A
`gateway.dns.upstream` written as a hostname now works: the resolver
forwards to the address the startup check resolved, where before the check
passed and the resolver then stopped on the unparsed name.
#### Linux packages
- A `.deb` upgrade whose new daemon cannot start no longer hangs apt. The
postinst started `fips.service` and then `fips-dns.service` with blocking
calls, and because `fips-dns.service` requires the daemon, a daemon that
failed on every start left the second call, apt and everything queued behind
it waiting for ever with no message. Each start is now queued and waited on
for at most 60 seconds, 90 for `fips-gateway`. A unit that does not come up
has its status printed and fails the configure step, so apt exits non-zero
and names the unit; a masked unit, or one whose condition is not met, is
reported and skipped.
- The `.deb` maintainer scripts now manage `fips-gateway` with the rest of the
package's services. An upgrade stopped the daemon, which the gateway
requires, and never brought the gateway back, so an operator who had enabled
it lost it until the next reboot; removing or purging the package left the
gateway's enablement symlink behind, pointing at a unit file that no longer
exists. The gateway is now stopped before the daemon on upgrade and
restarted afterwards only when it is enabled and the daemon came up, and it
is stopped and disabled on remove and purge. A gateway that does not come
back is reported but does not fail the upgrade.
- Purging the `.deb`, or running `uninstall.sh` from the tarball, now removes
the `.fips` DNS routing when `fips-dns` was not running at the time. The
cleanup removed the dns-delegate file from the wrong directory, never removed
the systemd-resolved global drop-in, and restarted no resolver, so the host
kept sending `.fips` queries to `[::1]:5354`, where nothing listens any more,
and `.fips` lookups timed out. Both scripts now remove all four files
`fips-dns-setup` can write, and restart systemd-resolved or reload dnsmasq or
NetworkManager when they removed that resolver's file and it is running. A
failed restart is reported and does not fail the removal.
- The `.deb` now declares `libgcc-s1 (>= 4.2)`. All four binaries link
`libgcc_s.so.1`, but cargo-deb removes every libgcc entry from the
dependencies it derives, so the package never said so. `libc6` depends on
`libgcc-s1` on Debian 12 and Ubuntu 22.04, 24.04 and 26.04, so installs there
were not affected. A new check, `testing/check-deb-depends.sh`, runs
`dpkg-shlibdeps` over the package's binaries on every build and fails the
build when the declared `Depends` leaves out a library the binaries need, or
states a floor lower or higher than the one they need. A dependency the
packaging tool drops, including one it drops after only a warning when it
cannot resolve a binary, now fails the build instead of shipping.
- The `.deb` now recommends `nftables`, and both AUR `PKGBUILD` files list it
as an optional dependency. `fips-firewall.service` runs `/usr/sbin/nft`, so
enabling it on a host without nftables failed at start. It is a
recommendation rather than a dependency because the firewall unit is opt-in
and the daemon itself does not need `nft`.
- The release `PKGBUILD` now lists `dbus` as a runtime dependency. The `fips`
binary links `libdbus-1`, and the `fips-git` package already declared it.
- `-V` on binaries built into the Linux packages now includes the source
revision, as `<version> (rev <git-hash>)`. The build image had no git, so
every container-built binary printed the version alone. A package built from
a git worktree still has no revision, because the worktree's git directory is
outside the tree the build sees. The build image's tag now includes a hash of
its Dockerfile, so a host with an older image cached builds a new one instead
of reusing it.
- `packaging/debian/build-deb-container.sh` now returns the package it just
built. It picked the most recently modified `fips_*.deb` in the output
directory that sorted last by name, so a package with a higher version left
there by an earlier run was returned instead.
#### OpenWrt
- A new OpenWrt install no longer enables and starts `fips-gateway`. The
generated postinst turned it on unconditionally, contradicting the init
script's own header, the package README and the deployment tutorial, all of
which say the service ships disabled and is enabled deliberately. The
documented `service fips-gateway enable` / `service fips-gateway start` steps
are unchanged, and the shipped `fips.yaml` still carries `gateway.enabled:
true`, so enabling the service is all that is needed.
- **The first opkg upgrade to this release re-enables and starts
`fips-gateway` on any router that has the `.ipk` installed, including one
where the gateway was disabled by hand.** opkg runs the outgoing package's
prerm, and every released `.ipk` prerm disabled the service on its way out,
leaving nothing behind that says whether the operator wanted it on, so an
upgrade cannot tell the two apart and keeps the gateway running rather than
silently turning off a working one. If you had disabled it, run
`service fips-gateway stop` and then `service fips-gateway disable` once
after upgrading. Stopping it hands dnsmasq's `.fips` forwarding back to the
daemon; disabling it alone leaves it running. An `apk` upgrade on
OpenWrt 25 runs only the incoming package's scripts, so it keeps the
gateway's enabled state from the first upgrade on. Later upgrades preserve
whatever state the service is in: the new prerm stops the services on an
upgrade but no longer disables them.
- An `apk` upgrade on OpenWrt 25 now restarts `fips`, and restarts
`fips-gateway` if it was enabled, so the new binaries run without a reboot.
apk-tools v3 runs only the incoming package's pre-upgrade and post-upgrade
scripts, and the `.apk` registered neither, so an upgrade replaced the files
on disk and left the old processes running until a reboot or a manual
restart.
- `start_service` in the `fips-gateway` init script now reads `gateway.enabled`
from `/etc/fips/fips.yaml` before doing anything. Starting a gateway that the
config disables used to hand dnsmasq's `.fips` forwarding to the gateway's
port, add the LAN prefix and advertise the pool route, and only then start a
daemon that exits immediately because the gateway is disabled, leaving `.fips`
resolution pointed at a port nothing listens on.
- The packages no longer ship `/etc/dnsmasq.d/fips.conf`. OpenWrt's dnsmasq
builds its config from UCI and never reads that directory; `.fips`
forwarding has always come from the UCI server entry, which is unchanged. An
opkg upgrade removes the old file, and an apk upgrade keeps it only if it
was modified. Either way nothing reads it.
- The package README's upgrade commands and default settings are corrected.
It now gives the `apk add` command for OpenWrt 25, where there is no opkg,
and for OpenWrt 24.10 and earlier a plain `opkg install` in place of
`--force-reinstall`, which removed and reinstalled the package and so left
`fips-gateway` disabled. Its description of the default config now matches
the shipped `fips.yaml`.
- The `.ipk` and `.apk` packages now install the same maintainer scripts. The
four script bodies live in `packaging/openwrt-ipk/scripts/` instead of inside
heredocs in the two build scripts, so the scenarios in `testing/openwrt/` run
what ships.
#### Windows
- The Windows service now writes its log to `C:\ProgramData\fips\fips.log`,
rolled at 10 MiB with four old files kept. A service has no standard output,
so everything the daemon logged in service mode was lost, including
config-load failures and panic messages. A foreground run still logs to the
console.
#### FreeBSD
- The daemon's log, `/var/log/fips.log`, is now rotated. The package ships a
newsyslog entry that keeps five compressed generations of 1000 KB, and the rc
script starts `daemon(8)` with `-H` so it reopens the log after a rotation.
The log used to grow without bound.
#### Links and transports
- A heartbeat whose send failed no longer counts as one that was delivered. The
send was recorded before it was attempted, so a peer whose heartbeat could not
go out was treated as heartbeated and was not tried again for a whole
`heartbeat_interval_secs`, although it had heard nothing and its own link-dead
timer was running. The attempt and the delivery are now recorded separately:
the interval that paces a healthy peer advances only on a send that returned
cleanly, and a peer whose send failed is retried after a shorter fixed
interval instead. That retry interval gates only a peer whose last attempt
failed, so it cannot clamp a `heartbeat_interval_secs` configured below it.
- A peer that moves to a new address now loses the per-peer `connect(2)`-ed UDP
socket pinned to the address it left. `set_current_addr` returns whether the
address actually changed so the caller can drop the stale socket, and the
@@ -612,25 +811,57 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
the peer never left the unconnected path. The flags are now set when the
socket is adopted, after its bind, so the traversal bind still receives a
port no other socket holds.
- A configured `ble:` transport that this build cannot construct is now
reported. The only warning for it was compiled into test builds alone, where
logging is compiled out, so macOS, Windows, FreeBSD, OpenWrt and other musl
builds, and Android without a BLE radio armed before start, dropped the block
silently while reporting healthy. The daemon now warns once per configured
instance at startup, naming the reason. The shipped example configs no longer
say BLE needs a `ble` Cargo feature, which does not exist: the common
`fips.yaml` names the builds that have the transport, and the OpenWrt
`fips.yaml` drops its BLE example, since its musl builds never include it.
#### Peering
#### Sessions and rekey
- A heartbeat whose send failed no longer counts as one that was delivered. The
send was recorded before it was attempted, so a peer whose heartbeat could not
go out was treated as heartbeated and was not tried again for a whole
`heartbeat_interval_secs`, although it had heard nothing and its own link-dead
timer was running. The attempt and the delivery are now recorded separately:
the interval that paces a healthy peer advances only on a send that returned
cleanly, and a peer whose send failed is retried after a shorter fixed
interval instead. That retry interval gates only a peer whose last attempt
failed, so it cannot clamp a `heartbeat_interval_secs` configured below it.
- Replacing the peer list at runtime with `Node::update_peers` now updates
everything that reads peer aliases. `.fips` names, peer ACL entries written as
an alias, and peer display names kept following the aliases the node started
with, so a new peer's alias did not resolve, a removed one still did, and a
deny entry naming an alias moved to another key kept denying the old key and
admitted the new one. They now follow the new peer list, with the hosts file
still taking precedence as it does at startup.
- A session whose last handshake message is lost no longer stays one-sided.
The initiator sent msg3 once and treated the session as established at once;
when that one datagram was lost, the responder kept waiting for it and
dropped every frame the initiator sent, and nothing sent msg3 again, because
the responder's repeated SessionAck was refused as arriving in the wrong
state. The session stayed that way until the next session rekey, or with
periodic rekey switched off, indefinitely. The initiator now keeps its msg3
and resends it on the handshake resend interval, with backoff, until a frame
from the responder authenticates or `handshake_max_resends` resends have gone
out. The wire format is unchanged: the resend carries the same msg3, and a
responder that already completed the session refuses the duplicate as before.
- A session rekey this node started no longer stays in flight forever when its
setup or the peer's ack is lost. Nothing resends a rekey setup, and the only
expiry covered a rekey the peer started, so one lost datagram left the
rekey pending and blocked every later one: the session kept its current keys
and stopped rotating them. The rekey now expires on the handshake timeout,
timed from when this node sent its setup, and the next tick starts a fresh
one. A forged ack cannot extend it. Expiries are counted as
`rekey_unanswered`. The wire format is unchanged.
- A link rekey whose reply is lost no longer splits the link. The node that
answered a rekey used to switch to the new keys on its own next tick, before
the other side had them; when the reply was lost, frames from the answering
side were dropped until the link was torn down. The answering side now
switches only when a frame on the new keys arrives from the side that started
the rekey, and drops keys that were never adopted after a hold (120 s by
default) so the next rekey can proceed.
- A node with no coordinates cached for a session's destination no longer
sends its own coordinates in their place. The lookup that supplies them falls
back to the node's own coordinates, which a first-contact SessionSetup needs
because its destination field cannot be empty, but the established data path,
the standalone CoordsWarmup and the rekey SessionSetup used the same
fallback. Every receiver files the destination coordinates it is sent under
the destination's address, so a destination reached this way cached its own
address under the sender's coordinates. On a cache miss a data frame now goes
out without coordinates and leaves the warmup budget for the first frames
after the cache is refilled, a standalone CoordsWarmup is not sent, and a
rekey SessionSetup, which can only miss for a direct peer, carries the
coordinates that peer announced. First-contact setup is unchanged. The wire
format is unchanged.
#### Routing and discovery
@@ -662,79 +893,16 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
they show a loss, and resends it once after 30 seconds when they cannot
confirm it, under the same limits.
#### Session setup
- A session whose last handshake message is lost no longer stays one-sided.
The initiator sent msg3 once and treated the session as established at once;
when that one datagram was lost, the responder kept waiting for it and
dropped every frame the initiator sent, and nothing sent msg3 again, because
the responder's repeated SessionAck was refused as arriving in the wrong
state. The session stayed that way until the next session rekey, or with
periodic rekey switched off, indefinitely. The initiator now keeps its msg3
and resends it on the handshake resend interval, with backoff, until a frame
from the responder authenticates or `handshake_max_resends` resends have gone
out. The wire format is unchanged: the resend carries the same msg3, and a
responder that already completed the session refuses the duplicate as before.
#### Link and session rekey
- A session rekey this node started no longer stays in flight forever when its
setup or the peer's ack is lost. Nothing resends a rekey setup, and the only
expiry covered a rekey the peer started, so one lost datagram left the
rekey pending and blocked every later one: the session kept its current keys
and stopped rotating them. The rekey now expires on the handshake timeout,
timed from when this node sent its setup, and the next tick starts a fresh
one. A forged ack cannot extend it. Expiries are counted as
`rekey_unanswered`. The wire format is unchanged.
- A forged rekey msg2 no longer takes the link down. The rekey initiator gave
up its handshake before reading msg2 and abandoned the cycle when the read
failed, although nothing authenticates a msg2 ahead of that read. Anyone on
the path who saw the rekey msg1 go out could answer first with a msg2 of the
right size under the index msg1 carries in cleartext. The responder has
already committed its new session by then and cuts over on its next tick, so
the two ends were left on different keys: frames from the responder were
dropped at once, frames to it failed once its drain window closed, and each
end removed the other on the link-dead timeout about 30 s later. A msg2 that
fails the read now leaves the handshake as it was before the read, along
with the msg1 resend schedule and the msg2 dispatch entry, so the
responder's genuine msg2 still completes the rekey. In exchange, every such
forgery now costs the initiator the msg2 key agreement until the cycle ends,
where before only the first one did; the msg1 resend budget bounds that. The
wire format is unchanged.
- A SessionAck that fails to read no longer ends a session rekey this node
started. The handler took the rekey handshake off the session before reading
the ack's msg2 and abandoned the rekey when the read failed, although nothing
authenticates the ack before that read and the only tie to the rekey is the
datagram's source address. The handshake now goes back rolled back to its
state before the read, so the peer's genuine ack still completes the rekey,
and the refusal is counted as `ack_handshake_failed`, as it already was for a
first-contact session. The wire format is unchanged.
- A link rekey whose reply is lost no longer splits the link. The node that
answered a rekey used to switch to the new keys on its own next tick, before
the other side had them; when the reply was lost, frames from the answering
side were dropped until the link was torn down. The answering side now
switches only when a frame on the new keys arrives from the side that started
the rekey, and drops keys that were never adopted after a hold (120 s by
default) so the next rekey can proceed.
#### Session coordinates
- A node with no coordinates cached for a session's destination no longer
sends its own coordinates in their place. The lookup that supplies them falls
back to the node's own coordinates, which a first-contact SessionSetup needs
because its destination field cannot be empty, but the established data path,
the standalone CoordsWarmup and the rekey SessionSetup used the same
fallback. Every receiver files the destination coordinates it is sent under
the destination's address, so a destination reached this way cached its own
address under the sender's coordinates. On a cache miss a data frame now goes
out without coordinates and leaves the warmup budget for the first frames
after the cache is refilled, a standalone CoordsWarmup is not sent, and a
rekey SessionSetup, which can only miss for a direct peer, carries the
coordinates that peer announced. First-contact setup is unchanged. The wire
format is unchanged.
#### Control socket
#### Node health and control socket
- A DNS responder or TUN thread that dies now degrades the node's published
health, and a dead responder's address is retracted. The responder's exit
report followed a loop that never returns, so it could not run, and a panic
in the responder or in either TUN thread unwound past its report. The node
kept reporting healthy with the child gone and kept publishing the DNS
address with nothing answering on it. Each child now runs inside a wrapper
that catches a panic, logs it, and reports the exit either way. A deliberate
stop still reports nothing.
- `show_links` (`fipsctl show links`) now reports the traffic a link has
carried. Its `packets_sent`, `packets_recv`, `bytes_sent`, `bytes_recv` and
`last_recv_ms` were read from counters on the link record that nothing on
@@ -758,68 +926,10 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
values the open-discovery tutorial described never occurred, and the
tutorial no longer lists them. The response shape is unchanged.
#### Identity and config
- A persistent node whose identity key path cannot be examined now refuses to
start instead of coming up under a new identity. `Path::exists` reports false
both for a key that is absent and for one whose metadata cannot be read, so a
key symlinked onto a volume that did not mount, or one in a directory the
daemon cannot search, read as a first boot: the node generated a fresh
identity, failed to store it, and carried on under an npub that every peer
whose allowlist names the old one refuses. Only a `NotFound` result is now
treated as an absence; any other failure to stat the path aborts the start and
names the path. A dangling symlink likewise aborts rather than being replaced.
The legacy `/etc/fips/fips.key` lookup follows the same rule.
### Security
#### Gateway
- Conntrack sessions are matched by address rather than by text, so live
traffic pins a gateway mapping again. The session count searched each
`/proc/net/nf_conntrack` line for `dst=` followed by the virtual IP in its
compressed form (`fd01::1`), while the kernel prints tuples in the full
uncompressed form (`dst=fd01:0000:0000:0000:0000:0000:0000:0001`), so the
count was zero for every mapping on every kernel. Nothing pinned an in-use
mapping, and one whose client did not re-query DNS was reclaimed about two
minutes after its last DNS reference while its traffic was still flowing.
Each `dst=` value is now parsed as an address and compared as one.
- A DNS query that refreshes a draining mapping now cancels its old grace
period. Previously the address could be reclaimed while the client's
renewed DNS answer was still valid. The mapping now survives the full
renewed TTL and a fresh grace period before it can be reused.
- The conntrack table is read once per tick instead of once per mapping, and
the read happens off the runtime thread. The whole file was read and scanned
for each mapping in turn, while the pool lock was held, on the same
single-threaded runtime that serves DNS. The tick now takes one snapshot with
a blocking task before it takes the lock, and the pool does a map lookup per
mapping.
- A conntrack source that cannot be read is reported. It still counts as zero
sessions for every mapping, as it always has, so reclamation keeps working
rather than pinning the whole pool; but the first failure and each change of
outcome after it are now logged, so an unreadable source is no longer
indistinguishable from an idle one. A source that fails identically every
tick is logged at debug rather than warn on a repeat.
- The NAT table is rebuilt in one netlink transaction. A rebuild deleted the
`fips_gateway` table in a batch of its own, discarded that batch's result,
and only then sent the batch that recreated the table, the chains, the
`fips0` masquerade and every per-mapping rule. Between the two sends the
gateway had no NAT at all, and a recreate the kernel refused left the table
absent for good, taking down forwarding for every existing mapping rather
than failing the one change that was being made. The delete and the recreate
now share a single batch, which the kernel applies as one transaction, so a
refused rebuild leaves the previous table in the packet path. The rules sent
are unchanged.
- The gateway's NAT rebuild no longer fails once the table holds more than
about 105 mappings. Each rebuild is one netlink batch. From about 105
mappings the default socket buffers could not hold its acknowledgements, so
rebuilds were logged as failed although they had taken effect. Past about
313 mappings the buffers could not hold the batch itself, and new `.fips`
names past that count got a virtual IP with no translation. In releases with
the gateway through 0.5.1, a rebuild past about 313 mappings also deleted the
whole `fips_gateway` table, which stopped every mapping, the `fips0`
masquerade and the port forwards. The rebuild now sizes its send buffer to
the batch and requests one acknowledgement per batch, and NAT errors now
name the kernel errno. A rebuild that still fails is logged, and the next
successful rebuild installs the mapping.
- A `.fips` query the gateway answers without an address no longer takes an
address from the pool. Every query type was allocated a mapping before the
code looked at what the client had asked for, and an A or HTTPS query was
@@ -838,15 +948,87 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
warning says which limit refused it. A name that already has a mapping is
answered before either limit is consulted, so names in use keep resolving
when the pool is full. The limits are compiled in, not configured.
- `fips-gateway` exits when its DNS listener cannot bind, or stops while the
gateway runs, instead of staying up with `.fips` resolution dead, so systemd
or procd restarts it or reports it failed. This applies to a gateway used
only for port forwards too. An "address in use" error names the service
likely to hold the port. On an OpenWrt access point with the gateway
enabled, the gateway had lost its port to the daemon's own mDNS responder
and `.fips` names stopped resolving with nothing reported.
#### Nostr and NAT traversal
#### Windows
- Windows now keeps its config, key, hosts and peer ACL files in
`C:\ProgramData\fips`, the directory the service installer writes to. The
config search, the key directory and the peer ACL defaults disagreed: the
config search never looked in `C:\ProgramData\fips`, so the service depended
on the `FIPS_CONFIG` the installer set, `fipsctl keygen` wrote to the
per-user `%APPDATA%\fips`, and a `peers.deny` placed beside the hosts file
was never read, so the ACL failed open. A key left in `%APPDATA%\fips` is
still used by a persistent node when the new directory has none, with a note
to move it. For one release, a `peers.allow` or `peers.deny` left at the old
`\etc\fips` location is still read when the new directory has no such file,
with a warning naming both paths.
- The Windows service installer now restricts `C:\ProgramData\fips` to SYSTEM
and Administrators. The directory inherited `C:\ProgramData`'s default ACL,
which lets any local user read the files in it and create new ones, so any
local account could read the node's key, or create a missing `fips.key`,
`fips.yaml` or `hosts` that the service then used. The installer creates the
directory with the restricted ACL, or replaces the ACL of an existing one,
resets the files already in it to inherit it, and refuses to continue if the
directory or anything in it is a link or a folder, or if the directory is
owned by another account. After moving files into the directory, stop the
service and rerun the installer; it cannot replace `fips.exe` while the
service runs. A foreground run from an unelevated prompt can no longer read
the files there.
- The Windows service installer now creates empty `peers.allow` and
`peers.deny` files in `C:\ProgramData\fips`. The service read these lists
from `\etc\fips` on the system drive, where any local user can create files,
so a planted list was enforced; for this release it still falls back there
when a file is missing from `C:\ProgramData\fips`. An empty file allows
every peer; to clear a list, empty its file rather than deleting it. The
installer stops when either file exists under `\etc\fips` and not in
`C:\ProgramData\fips`, so an upgrader's list is neither enforced from the
old location nor dropped unreviewed: review it, move it into
`C:\ProgramData\fips` or delete it, and run the installer again. Windows
upgraders should stop the service and rerun `install-service.ps1`.
#### Links and transports
- Two inbound TCP connections that share a peer address but arrive on different
local addresses no longer share one pool entry. The kernel names a connection
by its four-tuple, so a listener on a wildcard address, which is what the
shipped configuration binds, can accept two connections whose peer `ip:port`
is the same on two different local addresses. The pool was keyed by the peer
address alone: the second connection's entry replaced the first's while the
inbound-connection counter counted both, the first connection's teardown then
removed the second's entry, and the second's own teardown found nothing to
remove, so the counter ended one above the connections it counts. That counter
gates the inbound connection limit, so a host repeating the collision could
hold it at the limit and lock out further inbound TCP connections until the
daemon restarted. Inbound entries now carry the accepted socket's local
address in their pool key as well as the remote one.
#### Sessions and rekey
- A SessionAck that fails to read no longer ends a session rekey this node
started. The handler took the rekey handshake off the session before reading
the ack's msg2 and abandoned the rekey when the read failed, although nothing
authenticates the ack before that read and the only tie to the rekey is the
datagram's source address. The handshake is now rolled back to its state
before the read, so the peer's genuine ack still completes the rekey,
and the refusal is counted as `ack_handshake_failed`, as it already was for a
first-contact session. The wire format is unchanged.
- A forged rekey msg2 no longer takes the link down. The rekey initiator gave
up its handshake before reading msg2 and abandoned the cycle when the read
failed, although nothing authenticates a msg2 ahead of that read. Anyone on
the path who saw the rekey msg1 go out could answer first with a msg2 of the
right size under the index msg1 carries in cleartext. The responder had
already committed its new session by then and cut over on its next tick, so
the two ends were left on different keys: frames from the responder were
dropped at once, frames to it failed once its drain window closed, and each
end removed the other on the link-dead timeout about 30 s later. A msg2 that
fails the read now leaves the handshake as it was before the read, along
with the msg1 resend schedule and the msg2 dispatch entry, so the
responder's genuine msg2 still completes the rekey. In exchange, every such
forgery now costs the initiator the msg2 key agreement until the cycle ends,
where before only the first one did; the msg1 resend budget bounds that. The
wire format is unchanged.
#### Routing and discovery
- A node no longer publishes NIP-09 deletion requests signed with its routing
key after a NAT traversal attempt. Each request put the node's public
@@ -859,158 +1041,13 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
request, since that names an event the routing key signed itself. The
discovery and traversal design documents describe the new behaviour.
#### Packaging (OpenWrt)
#### Dependencies
- A new OpenWrt install no longer enables and starts `fips-gateway`. The
generated postinst turned it on unconditionally, contradicting the init
script's own header, the package README and the deployment tutorial, all of
which say the service ships disabled and is enabled deliberately. The
documented `service fips-gateway enable` / `service fips-gateway start` steps
are unchanged, and the shipped `fips.yaml` still carries `gateway.enabled:
true`, so enabling the service is all that is needed.
- **The first upgrade to this release re-enables and starts `fips-gateway` on
any router that has the package installed, including one where the gateway
was disabled by hand.** Every released package's prerm disabled the service
on its way out, leaving nothing behind that says whether the operator wanted
it on, so an upgrade cannot tell the two apart and keeps the gateway running
rather than silently turning off a working one. If you had disabled it, run
`service fips-gateway disable` once after upgrading. Later upgrades preserve
whatever state the service is in: the new prerm stops the services on an
upgrade but no longer disables them.
- `start_service` in the `fips-gateway` init script now reads `gateway.enabled`
from `/etc/fips/fips.yaml` before doing anything. Starting a gateway that the
config disables used to hand dnsmasq's `.fips` forwarding to the gateway's
port, add the LAN prefix and advertise the pool route, and only then start a
daemon that exits immediately because the gateway is disabled, leaving `.fips`
resolution pointed at a port nothing listens on.
- The `.ipk` and `.apk` packages now install the same maintainer scripts. The
four script bodies live in `packaging/openwrt-ipk/scripts/` instead of inside
heredocs in the two build scripts, so the scenarios in `testing/openwrt/` run
what ships.
- An `apk` upgrade on OpenWrt 25 now restarts `fips`, and restarts
`fips-gateway` if it was enabled, so the new binaries run without a reboot.
apk-tools v3 runs only the incoming package's pre-upgrade and post-upgrade
scripts, and the `.apk` registered neither, so an upgrade replaced the files
on disk and left the old processes running until a reboot or a manual
restart.
- The packages no longer ship `/etc/dnsmasq.d/fips.conf`. OpenWrt's dnsmasq
builds its config from UCI and never reads that directory; `.fips`
forwarding has always come from the UCI server entry, which is unchanged. An
opkg upgrade removes the old file, and an apk upgrade keeps it only if it
was modified. Either way nothing reads it.
- The package README's upgrade commands and default settings are corrected.
It now gives the `apk add` command for OpenWrt 25, where there is no opkg,
and for OpenWrt 24.10 and earlier a plain `opkg install` in place of
`--force-reinstall`, which removed and reinstalled the package and so left
`fips-gateway` disabled. Its description of the default config now matches
the shipped `fips.yaml`.
#### Packaging (Debian)
- A `.deb` upgrade whose new daemon cannot start no longer hangs apt. The
postinst started `fips.service` and then `fips-dns.service` with blocking
calls, and because `fips-dns.service` requires the daemon, a daemon that
failed on every start left the second call, apt and everything queued behind
it waiting for ever with no message. Each start is now queued and waited on
for at most 60 seconds. A unit that does not come up has its status printed
and fails the configure step, so apt exits non-zero and names the unit; a
masked unit, or one whose condition is not met, is reported and skipped.
- The `.deb` maintainer scripts now manage `fips-gateway` with the rest of the
package's services. An upgrade stopped the daemon, which the gateway
requires, and never brought the gateway back, so an operator who had enabled
it lost it until the next reboot; removing or purging the package left the
gateway's enablement symlink behind, pointing at a unit file that no longer
exists. The gateway is now stopped before the daemon on upgrade and
restarted afterwards only when it is enabled and the daemon came up, and it
is stopped and disabled on remove and purge. A gateway that does not come
back is reported but does not fail the upgrade.
- The `.deb` now declares `libgcc-s1 (>= 4.2)`. All four binaries link
`libgcc_s.so.1`, but cargo-deb removes every libgcc entry from the
dependencies it derives, so the package never said so. `libc6` depends on
`libgcc-s1` on Debian 12 and Ubuntu 22.04, 24.04 and 26.04, so installs there
were not affected. A new check, `testing/check-deb-depends.sh`, runs
`dpkg-shlibdeps` over the package's binaries on every build and fails the
build when the declared `Depends` leaves out a library the binaries need, or
states a floor lower or higher than the one they need. A dependency the
packaging tool drops, including one it drops after only a warning when it
cannot resolve a binary, now fails the build instead of shipping.
- `-V` on binaries built into the Linux packages now includes the source
revision, as `<version> (rev <git-hash>)`. The build image had no git, so
every container-built binary printed the version alone. A package built from
a git worktree still has no revision, because the worktree's git directory is
outside the tree the build sees. The build image's tag now includes a hash of
its Dockerfile, so a host with an older image cached builds a new one instead
of reusing it.
- `packaging/debian/build-deb-container.sh` now returns the package it just
built. It picked the most recently modified `fips_*.deb` in the output
directory that sorted last by name, so a package with a higher version left
there by an earlier run was returned instead.
- Purging the `.deb`, or running `uninstall.sh` from the tarball, now removes
the `.fips` DNS routing when `fips-dns` was not running at the time. The
cleanup removed the dns-delegate file from the wrong directory, never removed
the systemd-resolved global drop-in, and restarted no resolver, so the host
kept sending `.fips` queries to `[::1]:5354`, where nothing listens any more,
and `.fips` lookups timed out. Both scripts now remove all four files
`fips-dns-setup` can write, and restart systemd-resolved or reload dnsmasq or
NetworkManager when they removed that resolver's file and it is running. A
failed restart is reported and does not fail the removal.
- The `.deb` now recommends `nftables`. `fips-firewall.service` runs
`/usr/sbin/nft`, so enabling it on a host without nftables failed at start.
It is a recommendation rather than a dependency because the firewall unit is
opt-in and the daemon itself does not need `nft`.
#### Packaging (AUR)
- The release `PKGBUILD` now lists `dbus` as a runtime dependency. The `fips`
binary links `libdbus-1`, and the `fips-git` package already declared it.
- Both `PKGBUILD` files list `nftables` as an optional dependency, for
`fips-firewall.service`.
#### Packaging (FreeBSD)
- The daemon's log, `/var/log/fips.log`, is now rotated. The package ships a
newsyslog entry that keeps five compressed generations of 1000 KB, and the rc
script starts `daemon(8)` with `-H` so it reopens the log after a rotation.
The log used to grow without bound.
#### Windows
- Windows now keeps its config, key, hosts and peer ACL files in
`C:\ProgramData\fips`, the directory the service installer writes to. The
config search, the key directory and the peer ACL defaults disagreed: the
service found none of the installed files after a reboot and ran on
defaults, `fipsctl keygen` wrote to the per-user `%APPDATA%\fips`, and a
`peers.deny` placed beside the hosts file was never read, so the ACL failed
open. A key left in `%APPDATA%\fips` is still used when the new directory
has none, with a note to move it. For one release, a `peers.allow` or
`peers.deny` left at the old `\etc\fips` location is still read when the
new directory has no such file, with a warning naming both paths.
- The Windows service now writes its log to `C:\ProgramData\fips\fips.log`,
rolled at 10 MiB with four old files kept. A service has no standard output,
so everything the daemon logged in service mode was lost, including
config-load failures and panic messages. A foreground run still logs to the
console.
- The Windows service installer now restricts `C:\ProgramData\fips` to
SYSTEM and Administrators. The directory inherited `C:\ProgramData`'s
default ACL, which lets any local user read the files in it and create new
ones, so any local account could read the node's key, or create a missing
`fips.key`, `fips.yaml` or `hosts` that the service then used. The installer
creates the directory with the restricted ACL, or replaces the ACL of an
existing one, resets the files already in it to inherit it, and refuses to
continue if the directory or anything in it is a link or a folder, or if the
directory is owned by another account. Rerun the installer after moving files
into the directory. A foreground run from an unelevated prompt can no longer
read the files there.
- The Windows service installer now creates empty `peers.allow` and
`peers.deny` files in `C:\ProgramData\fips`. While either was missing there,
the service read that file from `\etc\fips` on the system drive, where any
local user can create files, so a planted list was enforced. An empty file
allows every peer; to clear a list, empty its file rather than deleting it.
The installer stops when either file exists under `\etc\fips` and not in
`C:\ProgramData\fips`, so an upgrader's list is neither enforced from the
old location nor dropped unreviewed: review it, move it into
`C:\ProgramData\fips` or delete it, and run the installer again. Windows
upgraders should rerun `install-service.ps1`.
- The lockfile moves `rustls` from 0.23.43 to 0.23.45, for RUSTSEC-2026-0285:
0.23.43 accepted TLS 1.3 handshake messages across encryption-level
boundaries. It is the TLS client the Nostr relay connections use, so every
default build reached it. The update is within the version range the
dependencies already allowed.
## [0.5.1] - 2026-09-06
Generated
+2 -2
View File
@@ -2960,9 +2960,9 @@ dependencies = [
[[package]]
name = "rustls"
version = "0.23.43"
version = "0.23.45"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0283386ce02abc0151e1761d08802dfe86c173b0b494af5cbc086574e453da06"
checksum = "0d41d731c7d2f962d1ccc364cec258de3c0e93b38c2fb3ba97ac74513048d634"
dependencies = [
"once_cell",
"ring",
+16 -13
View File
@@ -121,7 +121,7 @@ On Debian or Ubuntu, download `fips_<version>_amd64.deb` (or
`_arm64.deb`) and install it:
```bash
sudo dpkg -i fips_<version>_amd64.deb
sudo apt install ./fips_<version>_amd64.deb
sudo systemctl start fips fips-dns
```
@@ -152,8 +152,10 @@ does and does not share with the `.deb`.
For macOS, Windows, FreeBSD (including a pfSense build under
`packaging/pfsense/`), OpenWrt, the systemd tarball or a Nix
flake, see [docs/getting-started.md](docs/getting-started.md)
for the full multi-platform installation guide.
flake, [packaging/README.md](packaging/README.md) gives the install
commands for each package format, and
[docs/getting-started.md](docs/getting-started.md) is the full
multi-platform installation guide.
To join a live mesh and reach your first peer, follow the new-user
tutorial progression starting at
@@ -168,7 +170,7 @@ git clone https://github.com/jmcorgan/fips.git
cd fips
cargo install cargo-deb
cargo deb
sudo dpkg -i target/debian/fips_*.deb
sudo apt install ./target/debian/fips_*.deb
```
For the binaries alone, without an installer:
@@ -226,10 +228,10 @@ one and only the packaging differs. **Only the `.deb` is exercised by an install
ubuntu26; neither the AUR package nor the flake is. That suite runs on
every push and pull request, on x86_64, against a `.deb` built by the same
pinned container as the released one. The arm64 package, built the same way
on an arm64 runner, is installed and its daemon started on ubuntu22 on every
push and pull request as well; its upgrade, purge and conffile paths are not
exercised. The suite does not run at a tag: no workflow installs a published
artifact, so the released packages are checked by
on an arm64 runner, is installed, its daemon started and the package purged on
ubuntu22 on every push and pull request as well; its upgrade and conffile paths
are not exercised. The suite does not run at a tag: no workflow installs a
published artifact, so the released packages are checked by
hand. OpenWrt is a musl
target rather than glibc, and it takes an `.ipk` on 24.x and earlier or
an `.apk` on 25 and later; both carry the `fips-mesh-setup` and
@@ -295,7 +297,7 @@ Nix / NixOS section of [packaging/README.md](packaging/README.md).
then [fips-architecture.md](docs/design/fips-architecture.md) for
the protocol stack.
- **[Release notes](docs/releases/)** — per-version notes, including
[v0.5.1](docs/releases/release-notes-v0.5.1.md).
[v0.5.2](docs/releases/release-notes-v0.5.2.md).
If you want to contribute, see [CONTRIBUTING.md](CONTRIBUTING.md)
and [testing/README.md](testing/README.md).
@@ -337,10 +339,11 @@ testing/ Docker-based integration test harnesses + chaos simulation
## Status & roadmap
FIPS is at **v0.6.0-dev** on the `master` branch.
[v0.5.1](https://github.com/jmcorgan/fips/releases/tag/v0.5.1) is the
current release, a maintenance release on the v0.5.x line that makes the
Linux packages install and run on Debian 12 and Ubuntu 22.04, where every
artifact from v0.3.0 through v0.5.0 installed and then could not start.
[v0.5.2](https://github.com/jmcorgan/fips/releases/tag/v0.5.2) is the
current release, a maintenance release on the v0.5.x line that closes
security gaps in the Windows service, the gateway and the rekey
handshakes, and fixes the gateway, the Linux, OpenWrt and FreeBSD
packages, and session and discovery recovery after lost messages.
[v0.5.0](https://github.com/jmcorgan/fips/releases/tag/v0.5.0) was the last
feature release; this development line continues the testing-and-polishing
track toward v0.6.0. The core protocol works end-to-end over
+169 -157
View File
@@ -1,206 +1,218 @@
# FIPS v0.5.1
# FIPS v0.5.2
**Released**: 2026-09-06
**Released**: 2026-09-28
v0.5.1 is a maintenance release on the v0.5.x line and it exists for one
reason: **every Linux artifact from v0.3.0 through v0.5.0 installs
cleanly on Debian 12 and Ubuntu 22.04 and then cannot start.** The
package installs, the package manager reports success, and the daemon
fails to load with `GLIBC_2.39 not found`. If you run FIPS on either of
those distributions from a published package, you have never had a
working daemon, and this release is the fix.
v0.5.2 is a maintenance release on the v0.5.x line. It closes several security gaps, changes three defaults, and fixes defects that reach every node as well as the gateway, the Windows service and the Linux, OpenWrt and FreeBSD packages. There is no wire format change, so a mixed mesh works and nodes can be upgraded one at a time.
It carries two discovery fixes as well, one of them an external
contribution. There is no wire format change and no new configuration.
Read "Before you upgrade", below, if you run an ephemeral node (which is what every package's shipped config runs), the gateway, or FIPS on Windows or OpenWrt.
## At a glance
### Who should upgrade
- **Debian 12 and Ubuntu 22.04, from a package: upgrade.** The daemon on
those systems has never run. Nothing you can configure works around
it.
- **Any other Linux: upgrade at your convenience.** Your daemon was
running, and you gain the two discovery fixes.
- **macOS, Windows, FreeBSD and OpenWrt: the packaging defect never
affected you**, since it was in how the Linux artifacts were built. You
do get the two discovery fixes, which are not gated by platform.
- **From source: you were never affected.** A binary you built runs
against the C library you built it on.
- **Every node: upgrade.** The TLS library the Nostr relay connections use moves to a release that fixes RUSTSEC-2026-0285, and every default build reaches it. Four fixes also apply on every platform: a link rekey whose reply is lost no longer splits the link when the node answering the rekey runs v0.5.2; a session whose last handshake message is lost no longer stays one-sided; a session rekey whose setup or acknowledgement is lost no longer stops key rotation for the rest of the session; and a node that uses Nostr NAT traversal no longer signs deletion requests with its routing key, which linked its identity to its traversal messages on every relay it reached.
- **Windows: upgrade, and run `install-service.ps1` again.** On v0.5.1 any local user could read the node's key, or create a config, key, hosts file or peer list that the service then used, and a `peers.deny` placed beside the hosts file was never read, so the peer ACL failed open. The service log was also lost.
- **Gateway operators: upgrade.** Session pinning never worked, so a mapping carrying traffic could be reclaimed about two minutes after its last DNS reference. Any host that could reach the LAN resolver could exhaust the address pool. NAT rebuilds were reported as failed past about 105 mappings and did fail past about 313. On an OpenWrt access point with the gateway enabled, `.fips` names stopped resolving and nothing reported it.
- **Debian and Ubuntu, from the `.deb`: upgrade.** An upgrade whose new daemon could not start hung apt; an upgrade left `fips-gateway` stopped until the next reboot; a changed firewall ruleset was not reapplied; and purging the package could leave `.fips` lookups timing out.
- **OpenWrt: upgrade, and read the first-upgrade note below.** An `apk` upgrade left the old processes running until a reboot, and a fresh install enabled `fips-gateway`, which is meant to ship disabled.
- **FreeBSD: upgrade at your convenience.** The daemon's log now rotates; it grew without bound.
- **macOS: upgrade for the fixes every node gets.** Nothing in this release is specific to macOS.
### Before you upgrade
Nothing to do. No configuration key was added, removed or given a new
meaning, and no configuration that loaded under v0.5.0 fails to load
here.
Three defaults changed.
- **An ephemeral node no longer writes `fips.key`.** It used to write the private key of an identity it discards at every restart to that file, overwriting any key already there, including an operator's key when `persistent: true` had been forgotten. It now writes only `fips.pub` and holds the private key in memory. A `fips.key` found at an ephemeral start is renamed to `fips.key.unused`, with a warning, on the first start after the upgrade. **If you started a node once in ephemeral mode and then pinned the key it wrote, that no longer works**: set `node.identity.persistent: true` before upgrading, since most packages restart the daemon as they upgrade and an ephemeral start renames the key. The first persistent start uses the `fips.key` it finds. If an ephemeral start on v0.5.2 has already renamed it to `fips.key.unused`, rename it back to `fips.key` before restarting, as the daemon's warning says. With no key file, the first persistent start generates and saves one, so the npub changes once, at that restart, and stays stable from then on.
- **The gateway's default DNS listen address is `[::1]:5365`.** It was `[::1]:5353`, the mDNS port, which the daemon's own LAN rendezvous, Avahi and systemd-resolved can hold. On OpenWrt the upgrade rewrites the previously shipped `listen: "[::1]:5353"` line and the init script points dnsmasq at whatever port `gateway.dns.listen` sets, so there is nothing to do. On other hosts the upgrade leaves your `fips.yaml` as it is, so first check whether it sets `gateway.dns.listen`; the v0.5.1 example config and deployment guide set it to `"[::1]:5353"`. If it sets it, the gateway stays on that port and warns at startup when it is 5353: leave the resolver as it is, or move off the mDNS port by setting `listen: "[::1]:5365"` and forwarding `.fips` there, changing both together. If it does not set it, the gateway moves to `[::1]:5365`: a resolver you configured by hand to forward `.fips` to `[::1]:5353` must forward to `[::1]:5365` instead, or set `gateway.dns.listen: "[::1]:5353"` to keep the old port. A gateway on 5353 now exits if Avahi, systemd-resolved's MulticastDNS or the daemon's LAN rendezvous already holds that port, so that is the case to move.
- **Windows keeps its config, key, hosts and peer ACL files in `C:\ProgramData\fips`.** That is the directory the service installer already wrote to; the config search, `fipsctl keygen` and the peer ACL defaults each looked somewhere else. Upgrade by running the new `install-service.ps1` from an elevated prompt, with the service stopped. It restricts the directory to SYSTEM and Administrators and creates empty `peers.allow` and `peers.deny` files there. **It stops if it finds a `peers.allow` or `peers.deny` under `\etc\fips` on the system drive with no copy in `C:\ProgramData\fips`**: review that file, move it into `C:\ProgramData\fips` or delete it, and run the installer again. A key left in `%APPDATA%\fips` is still used by a persistent node when the new directory has none, with a note asking you to move it; stop the service and run the installer again after moving any file into the directory, since a moved file keeps its old permissions. If your v0.5.1 service read its config from `\etc\fips`, move its files before upgrading, as described in the Windows upgrade notes below.
Also worth knowing before you start:
- **`fips-gateway` now exits when its DNS listener cannot bind, or stops while the gateway runs**, where it used to stay up with `.fips` resolution dead. systemd and procd restart it or report it failed, and an "address in use" error names the service likely to hold the port. See [Port conflict on the DNS listen port](https://github.com/jmcorgan/fips/blob/v0.5.2/docs/how-to/troubleshoot-gateway.md#port-conflict-on-the-dns-listen-port).
- **The first opkg upgrade to this release re-enables and starts `fips-gateway` on any router that has the `.ipk` installed**, including one where you disabled the gateway by hand. If you had disabled it, run `service fips-gateway stop` and then `service fips-gateway disable` once after upgrading. Stopping it hands dnsmasq's `.fips` forwarding back to the daemon and removes the LAN prefix and route the gateway added; disabling it alone leaves it running, and after a reboot dnsmasq would still forward `.fips` to the gateway's port. Later upgrades keep whatever state the service is in, and an `apk` upgrade on OpenWrt 25 keeps it from the first upgrade on.
- **A `.deb` upgrade now fails if the daemon does not come up within 60 seconds**, and apt exits non-zero naming the unit, where it used to hang with no message.
- **The `.deb` now recommends `nftables`**, so `apt install` pulls it in by default. `dpkg -i` and `--no-install-recommends` do not.
- No configuration key was added or removed. Apart from `gateway.dns.listen`, no configuration key's default changed.
### What changed
- Linux packages and the systemd tarball now install and run on Debian
12 and Ubuntu 22.04.
- A node no longer relays away the answer to its own lookup.
- A returning copy of a node's own lookup request is no longer counted
against the peer that delivered it.
- Windows: files in `C:\ProgramData\fips`, the directory restricted to SYSTEM and Administrators, empty peer lists created by the installer, and a service log.
- Gateway: session pinning works, the pool has a mapping ceiling and a rate limit, the NAT rebuild is one transaction and scales past a few hundred mappings, the DNS listener moved off the mDNS port and a dead listener is fatal.
- Sessions and links: lost rekey replies, lost final handshake messages and lost rekey setups no longer leave a link or session stuck, and unauthenticated datagrams can no longer abort a rekey.
- Discovery: bloom filter and spanning-tree announces lost on a link are resent, and a node re-announces its filter when a child joins or leaves.
- Packages: the `.deb` upgrade, remove and purge paths manage every service and the DNS routing; OpenWrt `apk` upgrades restart the services; FreeBSD rotates the log.
- Dependencies: `rustls` 0.23.45 for RUSTSEC-2026-0285.
## The Linux packaging defect
## Security fixes
**What went wrong.** Every Linux artifact from v0.3.0 onward was built
on the newest available runner. That runner's C library turns the
standard library's `pidfd` references into a hard `GLIBC_2.39` version
requirement, instead of the weak, runtime-checked references they are
meant to compile to. The loader refuses an image on that entry alone, so
`fips`, `fipstop` and `fips-gateway` could not start on any system with
an older C library. **`fipsctl` was unaffected**, which is why an
install checked by running a command looked healthy while the daemon was
dead.
### Windows files readable and plantable by any local user
**No source code caused the packaging defect and none was changed to fix
it.** The defect was in the build environment. The two discovery fixes
below are this release's only behavioral change, and they are unrelated to
it.
The service installer wrote to `C:\ProgramData\fips`, which inherited `C:\ProgramData`'s default access: any local user could read the files in it and create new ones. Any local account could therefore read the node's key, or create a missing `fips.key`, `fips.yaml` or `hosts` that the service then used. Separately, the service read its peer ACL files from `\etc\fips` on the system drive, where any local user can create files, so a planted list was enforced, while a `peers.deny` placed beside the hosts file was never read.
**Why the declared dependency did not stop it.** The `.deb` said it
needed `libc6` with no version, which every glibc satisfies. So a
package whose binaries required 2.39 installed happily on a system with
2.35.
The installer now creates the directory with an ACL that admits only SYSTEM and Administrators, or replaces the ACL of an existing one, resets the files already in it to inherit that ACL, and refuses to continue if the directory or anything in it is a link or a folder, or if another account owns the directory. It creates empty `peers.allow` and `peers.deny` files, which allow every peer, so the service never falls back to the old location. To clear a list, empty its file rather than deleting it. **For this release only**, a `peers.allow` or `peers.deny` missing from `C:\ProgramData\fips` is still read from `\etc\fips`, with a warning naming both paths; the installer's stop, described above, is what keeps an unreviewed file there from being enforced. The [security reference](https://github.com/jmcorgan/fips/blob/v0.5.2/docs/reference/security.md) describes the fallback.
**What is different now.** The Linux artifacts are built in a container
pinned to the oldest supported distribution, declared in
`packaging/build-floor.env`. Every producer runs
`testing/check-glibc-floor.sh` on what it made, so an artifact that
would not load fails the build rather than reaching you. The declared
dependency is derived from the binaries themselves rather than written
by hand, so it states the floor it was actually built against and
re-derives per architecture. One script now produces the Linux
artifacts, and the systemd tarball takes its binaries out of the package
rather than from a second, unchecked set.
A foreground `fips.exe` run from an unelevated prompt can no longer read the files in that directory. Reading or editing them, `fipsctl keygen`, and `fipsctl address` with no argument need an elevated prompt.
**What was measured.** The `.deb` and the tarball were built and floor
checked on x86_64 and aarch64; both architectures carry a 2.34 floor,
which clears the 2.35 the packaging declares. Five distributions install
the package and start the daemon in CI. Separately, a project that
consumes FIPS inside an initramfs reports five machines of five
unlocking an encrypted root over the mesh, including Debian 12 and
Ubuntu 22.04, both of which fell back to a console prompt on v0.5.0.
### Gateway address pool exhausted from the LAN
**What was not measured.** No published v0.5.1 artifact existed when
this was written; the checks above ran against artifacts built by the
same scripts the release workflow uses. Reproducibility was measured
same-machine at v0.4.2 and cross-machine reproducibility has never been
tested.
Every `.fips` query was given a pool address before the gateway looked at what the client had asked for, and an A or HTTPS query was then answered with NODATA. Any host that could reach the LAN resolver could consume the pool one name at a time with a query type it is never given an address for, and could ask for one new name after another until the 65,535 addresses ran out, while every mapping made each NAT rebuild, each pool tick and shutdown slower.
## Discovery fixes
Only AAAA and ANY queries allocate now. The pool refuses a new name once it holds 1000 live mappings, and admits new names at 10 per second after a burst of 50; a refused query gets SERVFAIL, and the "Pool allocation failed" warning says which limit refused it. A name that already has a mapping is answered before either limit is consulted, so names in use keep resolving when the pool is full. The limits are compiled in, not configured.
### A node relayed away the answer to its own lookup
### Rekey aborted by an unauthenticated datagram
A lookup request is flooded to every tree peer whose bloom filter claims
the target, so a false positive can send a copy into the wider network
and circulate it back to the node that started it. The only identity
test on arrival asked whether the request named this node as the
*target*, which a lookup this node originated never satisfies. The
returning copy was therefore filed as ordinary transit under the node's
own request id. When the target answered, the reply was reverse-path
forwarded to the peer that had looped the request back, the pending
lookup was never satisfied, and discovery reported that its requests
went unanswered while the answers were in fact arriving.
Two handlers gave up a rekey handshake before reading a message that nothing authenticates ahead of that read.
An inbound response is now matched against this node's outstanding
lookups before the transit dedup record, and a returning copy of the
node's own request is dropped rather than recorded, so that id never
enters the transit cache.
- **A forged link rekey msg2 took the link down.** Anyone on the path who saw the rekey msg1 go out could answer first with a msg2 of the right size, and the two ends were left on different keys until each removed the other on the link-dead timeout, about 30 seconds later. A msg2 that fails to read now leaves the handshake as it was, so the genuine msg2 still completes the rekey.
- **A SessionAck that failed to read ended a session rekey.** The only tie between the ack and the rekey is the datagram's source address. The handshake is now rolled back to its state before the read, so the peer's genuine ack still completes the rekey, and the refusal is counted as `ack_handshake_failed`.
**This was a race rather than a hard failure**: a reply that beat the
looped copy found a clean cache and succeeded. It grew likelier as the
bloom fill ratio rose, which means it got worse as a mesh grew.
Contributed by Arjen.
### Routing key published next to traversal messages
### `req_duplicate` counted something it does not mean
After a NAT traversal attempt, a node published NIP-09 deletion requests signed with its routing key. Each one put the node's public identity next to the ids of its offer and answer gift wraps on every relay it reached, which the one-time signing keys on those wraps exist to prevent, and most of them deleted nothing, since a relay deletes a gift wrap only at its recipient's request. The node no longer sends them. A relay that stores the wraps now keeps them until their NIP-40 expiration. Withdrawing an advertisement still sends its deletion request, since that names an event the routing key signed itself.
The fix above drops the returning copy, and it first recorded that drop
under the existing `req_duplicate` rejection, whose documented meaning
is that a peer resent a request. A returning copy has a nonzero floor in
healthy operation and rises with the bloom fill ratio, so folding the
two together put a permanent number on a counter an operator reads as
neighbour misbehaviour, and made the two events indistinguishable.
### Inbound TCP limit held by a connection collision
It now has its own rejection reason and counter, `req_own_loopback`,
shown in `fipstop` as "Own Loopback". `req_duplicate` returns to meaning
only what it says.
The inbound TCP pool was keyed by the peer's address alone. A listener on a wildcard address, which is what the shipped configuration binds, can accept two connections from the same peer `ip:port` on two different local addresses, and that collision left the inbound connection counter one above the connections it counts. A host repeating it could hold the counter at the limit and lock out further inbound TCP connections until the daemon restarted. Inbound entries are now keyed by the connection's local address as well as the remote one.
**If you watch these counters**, expect a new non-zero `Own Loopback` to
appear on a node that originates lookups. That is traffic that was
previously counted elsewhere or not at all, correctly attributed, rather
than a new fault.
### rustls advisory
The lockfile moves `rustls` from 0.23.43 to 0.23.45 for RUSTSEC-2026-0285: 0.23.43 accepted TLS 1.3 handshake messages across encryption-level boundaries. It is the TLS client the Nostr relay connections use, so every default build reached it. The update is within the version range the dependencies already allowed.
## Gateway
- **Session pinning works.** The session count searched each `/proc/net/nf_conntrack` line for the virtual IP in its compressed form, while the kernel prints it uncompressed, so every mapping counted zero sessions on every kernel. A mapping whose client did not query DNS again was reclaimed about two minutes after its last DNS reference while its traffic was still flowing. Addresses are now compared as addresses. On a kernel without `/proc/net/nf_conntrack`, such as Ubuntu's, the gateway now reads conntrack over netlink, as `conntrack -L` does, where session pinning used to be off. The gateway says at startup which source it can read, or that none is readable and pinning is off, and reports an unreadable source when it happens. The table is read once per tick, off the thread that serves DNS, instead of once per mapping.
- **The NAT table is rebuilt in one netlink transaction.** A rebuild deleted the table in one batch and recreated it in another; between the two the gateway had no NAT at all, and a recreate the kernel refused left the table gone for good. A refused rebuild now leaves the previous table in place.
- **The NAT rebuild scales past a few hundred mappings.** From about 105 mappings rebuilds were logged as failed although they had taken effect; past about 313, new names got a virtual IP with no translation, and a rebuild could delete the whole table. The rebuild now sizes its buffer to the batch and asks for one acknowledgement, and NAT errors name the kernel errno.
- **A DNS query that renews a draining mapping cancels its old grace period**, so the address is no longer reclaimed while the client's renewed answer is still valid.
- **The DNS listener moved to `[::1]:5365`, and a listener that cannot bind or stops is fatal**, as described under "Before you upgrade" above. This applies to a gateway used only for port forwards too. A `gateway.dns.upstream` written as a hostname now works; before, the startup check resolved it and the resolver then stopped on the unparsed name.
The [gateway deployment guide](https://github.com/jmcorgan/fips/blob/v0.5.2/docs/how-to/deploy-gateway.md) and the [troubleshooting guide](https://github.com/jmcorgan/fips/blob/v0.5.2/docs/how-to/troubleshoot-gateway.md) describe the new limits, the listener's failure modes and the session-pinning check.
## Windows
Besides the directory move and the access restriction above, the Windows service now writes its log to `C:\ProgramData\fips\fips.log`, rolled at 10 MiB with four old files kept. A service has no standard output, so everything the daemon logged in service mode was lost, including config-load failures and panic messages. A foreground run still logs to the console.
## Linux packages
- **A `.deb` upgrade whose new daemon cannot start no longer hangs apt.** Each service start is waited on for at most 60 seconds, 90 for `fips-gateway`. A unit that does not come up has its status printed and fails the configure step; a masked unit, or one whose condition is not met, is reported and skipped.
- **The `.deb` scripts manage `fips-gateway`.** An upgrade stopped the daemon, which the gateway requires, and never brought the gateway back. It is now stopped before the daemon on upgrade and restarted afterwards when it is enabled and the daemon came up, and it is stopped and disabled on remove and purge, where its enablement link used to be left behind.
- **A `.deb` upgrade reapplies the firewall ruleset in place.** A changed `/etc/fips/fips.nft` used to take effect only at the next reboot or manual restart. `fips-firewall.service` gains a reload that replaces the ruleset in one transaction, and the `.deb` upgrade reloads it only when it is already active, so an upgrade never turns the firewall on for a host that has not opted in. A tarball or AUR upgrade does not reload it; run `systemctl reload fips-firewall` afterwards if the firewall is active.
- **Purging the `.deb`, or running `uninstall.sh` from the tarball, removes the `.fips` DNS routing** when `fips-dns` was not running at the time, and restarts or reloads the resolver that used it. The host used to keep sending `.fips` queries to a port where nothing listened.
- The `.deb` declares `libgcc-s1 (>= 4.2)`, which its binaries always needed, and every package build now checks the declared dependencies against the libraries the binaries link. It recommends `nftables`, which the opt-in firewall unit runs. The release AUR package lists `dbus` as a dependency and `nftables` as an optional one.
- `-V` on the packaged Linux binaries prints the source revision again.
- **The AUR package is published only after every package workflow for the release tag has succeeded**, so it can appear up to an hour after the release. At v0.5.1 the AUR was updated while the release had 15 of its 17 assets.
## OpenWrt
- **A fresh install no longer enables and starts `fips-gateway`.** The documented `service fips-gateway enable` and `service fips-gateway start` steps are unchanged, and the shipped `fips.yaml` still has `gateway.enabled: true`, so enabling the service is all that is needed.
- **An `apk` upgrade on OpenWrt 25 restarts `fips`**, and `fips-gateway` if it was enabled, so the new binaries run without a reboot.
- The first opkg upgrade re-enables the gateway, as described under "Before you upgrade" above. Later upgrades keep the service's state.
- Starting `fips-gateway` when the config disables the gateway no longer points dnsmasq's `.fips` forwarding at a port nothing listens on.
- The packages no longer ship `/etc/dnsmasq.d/fips.conf`, which OpenWrt's dnsmasq never read; `.fips` forwarding has always come from the UCI server entry.
- The [package README](https://github.com/jmcorgan/fips/blob/v0.5.2/packaging/openwrt-ipk/README.md) gives the `apk add` upgrade command for OpenWrt 25 and a plain `opkg install` for 24.10 and earlier, in place of `--force-reinstall`, which left the gateway disabled.
## FreeBSD
The daemon's log, `/var/log/fips.log`, is rotated by newsyslog, which keeps five compressed generations of 1000 KB, and the rc script starts `daemon(8)` with `-H` so it reopens the log after a rotation.
## Sessions, rekey and discovery
These apply on every platform.
- **A link rekey whose reply is lost no longer splits the link.** The answering side used to switch to the new keys on its own next tick, before the other side had them; when the reply was lost, frames from the answering side were dropped until the link was torn down. It now switches only when a frame on the new keys arrives from the side that started the rekey, and drops keys that were never adopted after a hold, 120 seconds by default. While it holds such keys, that node neither accepts nor starts a rekey on the link, so rekeying on that link waits for up to the hold. The fix is on the answering side: a link where a v0.5.0 or v0.5.1 node answers can still drop, as described under Compatibility below.
- **A session whose last handshake message is lost no longer stays one-sided.** The initiator sent msg3 once; when it was lost, the responder dropped every frame the initiator sent until the next session rekey, or indefinitely with periodic rekey off. The initiator now resends msg3 with backoff until the responder is heard from or the handshake resend limit is reached.
- **A session rekey whose setup or ack is lost now expires** on the handshake timeout, and the next tick starts a fresh one. It used to stay pending and block every later rekey, so the session stopped rotating its keys. Expiries are counted as `rekey_unanswered`.
- A node with no coordinates cached for a session's destination no longer sends its own coordinates in their place, which made the destination cache its own address under the sender's coordinates.
- **Lost bloom filter and spanning-tree announces are resent.** An announce counted as delivered once the transport accepted it, so a dropped datagram or a short link outage left the peer with the old filter or tree position until something else changed, and destinations could stay missing from discovery. The node now confirms each announce from the link's existing receiver reports and resends on loss, within a budget.
- A node re-announces its bloom filter when a peer starts or stops using it as parent, so destinations under a new child become discoverable and a departed child's stop being advertised.
## Links, node health and the control socket
- A heartbeat whose send failed is no longer counted as delivered, and the peer is retried sooner instead of after a whole `heartbeat_interval_secs`.
- A peer that moves to a new address loses the connected UDP socket pinned to the address it left, and a peer reached by NAT traversal now gets its connected socket.
- A configured `ble:` transport that this build cannot construct is reported at startup instead of being dropped silently.
- A DNS responder or TUN thread that dies, including by a panic, now degrades the node's published health, and a dead responder's address is retracted.
- `fipsctl show links` reports the traffic a link has carried; every link used to report zero.
- `fipsctl show peers` reports a peer that has been silent longer than `heartbeat_interval_secs` as `stale`; every peer used to read `connected` until it was removed.
- Replacing the peer list at runtime with `Node::update_peers` now updates `.fips` names and peer ACL entries written as an alias.
- A persistent node whose identity key path cannot be examined, such as a key symlinked onto a volume that did not mount, now refuses to start and names the path, instead of coming up under a new identity.
- On Linux, an empty datagram sent through the native datagram API just before a close may read as the close; the [native API reference](https://github.com/jmcorgan/fips/blob/v0.5.2/docs/reference/native-api.md) now says so.
## Compatibility
v0.5.1 is wire-compatible with v0.5.0. No frame gains, loses or resizes
a field, so a mixed mesh works and nodes can be upgraded one at a time
with no coordinated restart. The two discovery changes alter what a node
does with a message it already parsed; neither changes what is on the
wire.
v0.5.2 is wire-compatible with v0.5.1. No file that defines a frame, message, TLV or encoding changed between the two releases, so a mixed mesh works and nodes can be upgraded one at a time with no coordinated restart. The msg3 resend sends the same msg3, and the rekey and announce changes alter when a node sends or switches keys, not what it sends. Mixed-version behavior is measured by the interop run described below. Until the older nodes are upgraded, a link to a v0.5.0 or v0.5.1 node can still drop and be re-established when that node answers a link rekey and its reply is lost: the fix is on the answering side.
The library surface is unchanged. Configuration is unchanged.
No configuration key was added or removed; the gateway's DNS listen default is the only configuration-key default that changed.
## Upgrade notes
**Package upgrade, Debian and Ubuntu.** The usual upgrade replaces the
binaries and restarts the service. On Debian 12 and Ubuntu 22.04 the
daemon will start for the first time, so this is a first start rather
than a restart: check `fipsctl show status` afterwards and expect to see
peer establishment, not a resumed session.
**Debian and Ubuntu.** `apt install ./fips_0.5.2_<arch>.deb`, with `amd64` or `arm64` in place of `<arch>`. Once the new files are in place the package reloads the firewall if it is running, starts the daemon, then enables and starts `fips-dns` (every upgrade re-enables it; `systemctl mask fips-dns` keeps it off), and restarts `fips-gateway` if it is enabled. If the daemon does not come up, apt exits non-zero and prints its status; check `fipsctl show status` afterwards either way.
**Check what you actually have.** If you want to confirm the floor of an
installed binary rather than trust the version string:
**Windows.** Open an elevated PowerShell in the folder where you unzipped the new ZIP, and in it: stop the service (`Stop-Service fips`); if your v0.5.1 service read its config from `\etc\fips`, move its files as described in the next paragraph; run the installer (`powershell -ExecutionPolicy Bypass -File .\install-service.ps1`, since the execution policy can refuse an unsigned script from a downloaded ZIP); then start the service (`Start-Service fips`). In Windows PowerShell 5.1, `sc` is an alias for `Set-Content`, so use `sc.exe` if you prefer the service control tool. The installer cannot replace `fips.exe` while the service is running and stops with a file-in-use error, so stop the service before every run of it. Resolve any legacy peer list the installer reports, as described under "Before you upgrade" above. The service log is at `C:\ProgramData\fips\fips.log`.
```text
objdump -T /usr/bin/fips | grep GLIBC_ | sed 's/.*GLIBC_//' | sort -uV | tail -1
```
If your v0.5.1 service was set up by hand to read its config from `\etc\fips`, rather than by v0.5.1's `install-service.ps1` (which already used `C:\ProgramData\fips` and set `FIPS_CONFIG`), move `fips.yaml` and `fips.key` from `\etc\fips` into `C:\ProgramData\fips` after stopping the service and before running the new installer. The v0.5.2 installer sets `FIPS_CONFIG` to `C:\ProgramData\fips\fips.yaml`, and the service then reads only that file and the key beside it. Without the move, the service starts from whatever `C:\ProgramData\fips` holds. With the default config the installer places there, which does not enable a persistent identity, the node comes up under a new identity, and nothing warns about the files left in `\etc\fips`.
An artifact from this release prints `2.34) __libc_start_main`. One from
v0.5.0 or earlier prints `2.39) pidfd_spawnp`, which names the symbol that
caused this.
**OpenWrt.** Copy the package to `/tmp` on the router. On OpenWrt 25 and later run `apk add --allow-untrusted /tmp/fips_<new-version>_<arch>.apk`; on 24.10 and earlier run `opkg install /tmp/fips_<new-version>_<arch>.ipk`, with the file name you downloaded in place of the placeholders. After a first opkg upgrade, if you had disabled `fips-gateway`, run `service fips-gateway stop` and then `service fips-gateway disable`.
**Rolling upgrade.** No coordination is needed. Upgrade nodes in any
order.
**FreeBSD.** Install the new package with `pkg install ./fips-0.5.2-freebsd-amd64.pkg`, then run `service fips restart` and, if you use it, `service fips_dns restart`, and check `fipsctl show status`. The package restarts the services itself only when pkg takes its upgrade path, which `pkg add` over an installed package does not, and the log rotation fix takes effect only once the daemon restarts. Do not remove the package before adding the new one: removal stops both services and removes the `.fips` resolver drop-in.
## Getting v0.5.1
**Arch Linux and the systemd tarball.** pacman does not restart services: after the AUR package updates, run `systemctl restart fips.service`, and `systemctl restart fips-gateway.service` if you run the gateway. The tarball's `install.sh` restarts `fips` if it was running, but stopping `fips` also stops `fips-dns` and `fips-gateway`, and it starts neither again: run `systemctl start fips-dns.service`, and `systemctl start fips-gateway.service` if you run the gateway. After either, run `systemctl reload fips-firewall` if the firewall is active.
- **Linux x86_64 / aarch64**: `.deb` and tarball at the
[v0.5.1 release page](https://github.com/jmcorgan/fips/releases/tag/v0.5.1).
- **Arch Linux**: `fips` from the AUR.
- **macOS**: `.pkg` at the v0.5.1 release page.
- **Windows**: ZIP at the v0.5.1 release page.
- **FreeBSD (x86_64)**: `.pkg` at the v0.5.1 release page.
- **OpenWrt**: `.ipk` (OpenWrt 24.x and earlier) or `.apk` (OpenWrt 25+)
at the v0.5.1 release page.
- **From source**: `cargo build --release` from a checkout of the v0.5.1
tag (Rust 1.94.1 per `rust-toolchain.toml`; `libclang-dev` is a
required Linux build prerequisite).
- **Nix / NixOS**: `nix build .#fips` from a checkout of the v0.5.1 tag
builds the binaries from source with the pinned toolchain and no
manual prerequisites (see the Nix section of `packaging/README.md`).
**Gateway on other hosts.** If `fips.yaml` does not set `gateway.dns.listen` and a local resolver forwards `.fips` to `[::1]:5353`, point it at `[::1]:5365` right after the upgrade, or set `listen` to the old port. If `fips.yaml` sets it, the gateway stays on that port and nothing needs to change; see "Before you upgrade" above.
There is no Android daemon artifact. Android is supported as an embedded
crate.
**Rolling upgrade.** No coordination is needed. Upgrade nodes in any order.
The full per-commit changelog lives in
[`CHANGELOG.md`](https://github.com/jmcorgan/fips/blob/v0.5.1/CHANGELOG.md).
Issues and discussion at
[github.com/jmcorgan/fips](https://github.com/jmcorgan/fips). Security
reports have a private channel; see
[`SECURITY.md`](https://github.com/jmcorgan/fips/blob/v0.5.1/SECURITY.md).
## What was measured and what was not
**Measured, at the release's source content** (the same source as the tag, with the version string at `0.5.2-dev`):
- The full local test suite: 36 suites passed, including the chaos scenarios, the NAT traversal, gateway, firewall and DNS resolver suites, and the `.deb` install suite on five distributions.
- `cargo audit`: no vulnerabilities, with the same four allowed warnings as v0.5.1.
- The 100-node discovery test.
- The Debian package built through the release's container script, with all four binaries at a glibc floor of 2.34 against the 2.35 the package declares.
- The systemd tarball rebuilt twice on the same machine and compared byte for byte: identical.
- The OpenWrt `.ipk` cross-compiled for aarch64. The `.apk` cannot be built on the build host used here and is built by the release workflow.
- The wire format, by a diff showing that no file defining a frame, message, TLV or encoding changed since v0.5.1.
- A mixed-version mesh of v0.5.2, v0.5.1 and v0.5.0 nodes passed the interop suite, once without impairment and in eight runs with added delay and 2% packet loss: every pair passed the suite's connectivity checks, including after two rekey cycles; every pair of a v0.5.2 node and an older one completed link rekeys with each side starting them; and the logs showed no panic, error, decryption failure or handshake failure. Under loss, a link whose answering node ran v0.5.0 or v0.5.1 dropped after that node's rekey reply was lost and was re-established, three times in the eight runs. That is the defect this release fixes on the answering side: no link dropped where a v0.5.2 node answered or between two v0.5.2 nodes, and a v0.5.2-only mesh under the same loss dropped no link in four runs.
- The Windows installer and service on Windows Server 2025 (GitHub-hosted runners), under Windows PowerShell 5.1 and PowerShell 7, from a build that differs from the release's source content only in the version string, the changelog and the rustls update: the directory and peer-file permissions on a fresh install and on reruns; the installer's stop on a peer list left under `\etc\fips`; a peer list planted there by a standard user, which the service ignores once the current file exists and otherwise enforces with a warning; the upgrade from a service set up by v0.5.1's installer, which kept its node address; and the upgrade from a v0.5.1 service that read its config from `\etc\fips`, which came up under a new identity with no warning, as described in the Windows upgrade notes above.
- The gateway DNS port change on an OpenWrt 24 router running this release's source content (tested by Arjen): the gateway listened on `[::1]:5365` and dnsmasq forwarded `.fips` to it, and a gateway that cannot bind its DNS listener now exits instead of running with DNS dead. The same test found that if `fips-gateway` fails to start, dnsmasq is left forwarding `.fips` to the gateway's port until the gateway runs; this predates v0.5.2. If `.fips` stops resolving on a router, check `logread | grep fips-gateway`, and run `service fips-gateway stop` to hand `.fips` back to the daemon until the cause is fixed.
**Not measured when this was written:**
- No published v0.5.2 artifact existed. The checks above ran against artifacts built by the same scripts the release workflow uses.
- A `.deb` upgrade from an installed v0.5.1 package, which is the only case in which v0.5.1's removal scripts meet v0.5.2's install scripts. The install suite repacks one package against itself.
- The OpenWrt maintainer scripts are tested in CI in a BusyBox container against stubs, not under a real opkg or apk upgrade on a router.
- On OpenWrt router hardware, the upgrade rewriting the shipped gateway DNS listen line: the router test above ran the new port and did not report the rewrite separately.
- The Windows checks on a desktop edition of Windows, which can differ from the server runners in the owner given to new files and in when the service sees the `FIPS_CONFIG` the installer sets. The runner checks did not exercise `uninstall-service.ps1` or the TUN adapter.
- The changed rekey behavior under real mixed-version traffic on live nodes. No field soak was run for this release; the interop run described above stands in for it.
- Upgrading a running node in place, forwarding across more than one hop, loss patterns other than random 2% packet loss, and runs longer than a few minutes: the interop run started each mesh fresh, used full meshes only, and ran each repetition for about six minutes.
- Reproducibility across machines, which has never been tested.
- The FreeBSD upgrade command, and the statement that `pkg add` over an installed package does not take pkg's upgrade path: both are taken from pkg's source and were not run on a FreeBSD host.
## Getting v0.5.2
- **Linux x86_64 / aarch64**: `.deb` and tarball at the [v0.5.2 release page](https://github.com/jmcorgan/fips/releases/tag/v0.5.2).
- **Arch Linux**: `fips` from the AUR, once every package workflow for the release has finished.
- **macOS**: `.pkg` at the v0.5.2 release page.
- **Windows**: ZIP at the v0.5.2 release page.
- **FreeBSD (x86_64)**: `.pkg` at the v0.5.2 release page.
- **OpenWrt**: `.ipk` (OpenWrt 24.x and earlier) or `.apk` (OpenWrt 25+) at the v0.5.2 release page, for `aarch64_cortex-a53` and `x86_64`.
- **From source**: `cargo build --release` from a checkout of the v0.5.2 tag (Rust 1.94.1 per `rust-toolchain.toml`; `libclang-dev` is a required Linux build prerequisite, and on glibc Linux so are `libdbus-1-dev` and `pkg-config`).
- **Nix / NixOS**: `nix build .#fips` from a checkout of the v0.5.2 tag builds the binaries from source with the pinned toolchain and no manual prerequisites (see the Nix section of [`packaging/README.md`](https://github.com/jmcorgan/fips/blob/v0.5.2/packaging/README.md)). The NixOS module's documented flake input, `github:jmcorgan/fips`, follows the default branch; to run v0.5.2, set `inputs.fips.url = "github:jmcorgan/fips/v0.5.2"`, then run `nix flake update fips` and `nixos-rebuild switch`.
There is no Android daemon artifact. Android is supported as an embedded crate.
The full per-commit changelog lives in [`CHANGELOG.md`](https://github.com/jmcorgan/fips/blob/v0.5.2/CHANGELOG.md). Issues and discussion at [github.com/jmcorgan/fips](https://github.com/jmcorgan/fips). Security reports have a private channel; see [`SECURITY.md`](https://github.com/jmcorgan/fips/blob/v0.5.2/SECURITY.md).
## Contributors
Thanks to everyone who contributed code, packaging work, bug reports, or
reviews to this release.
Thanks to everyone who contributed code, packaging work, bug reports, or reviews to this release.
- [@jmcorgan](https://github.com/jmcorgan) (Johnathan Corgan): release
shepherd; the Linux build and floor-checking work, the loopback
rejection counter, and the install-suite fix that made the defect
visible instead of hanging.
- [@Origami74](https://github.com/Origami74) (Arjen): the lookup
originator fix, so a node accepts the answer to its own lookup instead
of relaying it away
([#141](https://github.com/jmcorgan/fips/pull/141)).
- [@jmcorgan](https://github.com/jmcorgan) (Johnathan Corgan): release shepherd; the gateway, Windows, packaging, session and discovery fixes.
- [@mmalmi](https://github.com/mmalmi) (Martti Malmi): the gateway fix that keeps a renewed DNS answer valid while its mapping is draining ([#169](https://github.com/jmcorgan/fips/pull/169)).
- [@Origami74](https://github.com/Origami74) (Arjen): found on router hardware that the gateway's DNS port collided with the daemon's mDNS responder on an OpenWrt access point, and reviewed the OpenWrt packaging fixes on OpenWrt 25.
- [@fr34aky](https://github.com/fr34aky): version placeholders in the FreeBSD packaging examples ([#152](https://github.com/jmcorgan/fips/pull/152), landed as `93d45191`), and a hostname test that no longer depends on the host's DNS search domain ([#154](https://github.com/jmcorgan/fips/pull/154), landed as `a1c0cd42`).
- [@shaibearary](https://github.com/shaibearary) (Sherry): a fix to the mesh test loop, which aborted its rekey setup under the bash that ships with macOS ([#155](https://github.com/jmcorgan/fips/pull/155), landed as `31ebec62`).
- [@Ghost-glitch-hub](https://github.com/Ghost-glitch-hub): reported that `fipsctl show links` showed zero traffic for active peers ([#158](https://github.com/jmcorgan/fips/issues/158)).
<!-- markdownlint-disable-file MD013 -->
+3
View File
@@ -208,6 +208,9 @@ New filter content is sent only on these events, never periodically:
- A peer's inbound filter changes (outbound filters to other peers must
be recomputed)
- Local state changes (new identity, leaf-only dependent changes)
- The tree changes around this node: it switches parent or becomes
root, or a peer starts or stops naming it as parent, which changes
the set of tree peers whose filters are merged
The one timed send is the resend of an announce that was not confirmed
delivered. The transport accepting a FilterAnnounce does not mean the peer
+42 -25
View File
@@ -96,14 +96,18 @@ fips_gateway`, with two chains:
return-path SNAT, and (when any port-forward is configured) the
LAN-side masquerade for inbound traffic.
The table is rebuilt atomically on every change. The rebuild
sequence — delete the existing table (ignore `ENOENT` on first
call), then create a new table with chains and the full rule set in
a single netlink batch — avoids reliance on kernel rule-handle
tracking, which the rustables crate does not expose. The table stays
small (one always-on masquerade plus two rules per active outbound
mapping plus one rule per inbound forward, with one extra masquerade
when any forward is present), so rebuilds are cheap.
The table is rebuilt atomically on every change, in one netlink
batch that the kernel applies as a single transaction: add the
table, delete it, add it again, then the chains and the full rule
set. Because the delete and the recreate share one transaction, the
table never leaves the packet path; the leading add gives the
delete a target when no table exists yet, and a batch the kernel
refuses leaves the previous table in place. Rebuilding the whole
table avoids reliance on kernel rule-handle tracking, which the
rustables crate does not expose. The table holds one always-on
masquerade, two rules per live outbound mapping (at most 1000
mappings), one rule per inbound forward, and one extra masquerade
when any forward is present.
### Control Socket
@@ -202,16 +206,19 @@ involving the DNS proxy or the pool.
traffic, so a `127.0.0.1:5354` upstream cannot reach a daemon
bound on `[::1]:5354`.
4. If the daemon is unreachable or times out (5 s), the gateway
replies `SERVFAIL`. If the daemon returns `NXDOMAIN` or a
non-`AAAA` answer, the gateway forwards the response unchanged.
replies `SERVFAIL`. If the daemon answers with an error such as
`NXDOMAIN`, the gateway relays that response code; if it answers
without an AAAA record, the gateway replies `SERVFAIL`.
5. The gateway extracts the AAAA (`fd00::/8`) record from the
daemon's response. This resolution primes the daemon's identity
cache as a side effect — a prerequisite for `fips0` routing,
because the daemon needs the cache entry to map the mesh address
back to a `NodeAddr` for forwarding.
6. The gateway allocates a virtual IP from the pool for that mesh
address (idempotent: an existing mapping is reused and its TTL
refreshed).
6. If the client asked for AAAA or ANY, the gateway allocates a
virtual IP from the pool for that mesh address (idempotent: an
existing mapping is reused and its TTL refreshed). Any other
query type refreshes an existing mapping's TTL, creates nothing,
and is answered with NODATA.
7. If a new mapping was created, the pool emits `MappingCreated`,
which the main loop turns into `add_mapping` calls on the NAT
manager and `add_proxy_ndp` on the network setup.
@@ -283,9 +290,13 @@ way an entry counts once toward each distinct IPv6 destination among
its original and reply tuples, so an entry counts as a session of a
virtual IP whose address is its original destination.
If the pool is exhausted, new DNS queries return `SERVFAIL`.
Existing mappings are never evicted prematurely — the correctness of
in-flight sessions takes precedence over fresh allocations.
A new mapping is refused, and the query answered `SERVFAIL`, when
the pool is exhausted, when it already holds 1000 live mappings, or
when the new-mapping rate limit is spent (a bucket of 50 that
refills at 10 per second). A name that already has a mapping keeps
resolving while new names are refused. Existing mappings are never
evicted prematurely — the correctness of in-flight sessions takes
precedence over fresh allocations.
### NAT Pipeline (Outbound)
@@ -435,22 +446,28 @@ and that table is rebuilt as one unit on every state change —
mapping added, mapping removed, port-forwards updated. The rebuild
sequence is:
1. Delete the existing table in its own batch (ignore `ENOENT`).
2. In a fresh batch: add the table; add the `prerouting` and
`postrouting` chains; add the always-on `oifname fips0`
masquerade; add per-mapping DNAT/SNAT rules for every active
pool entry; add per-port-forward DNAT rules; add the LAN-side
1. Add the table (which succeeds whether or not it exists), delete
it, and add it again, so the delete always has a target.
2. Add the `prerouting` and `postrouting` chains; the always-on
`oifname fips0` masquerade; per-mapping DNAT/SNAT rules for
every live pool entry; per-port-forward DNAT rules; the LAN-side
masquerade if any port-forwards exist.
3. Send the batch as a single netlink transaction.
3. Send all of it as one batch, which the kernel applies as a single
transaction. Only the last message before the batch end requests
an acknowledgement, the socket's send buffer is sized to the
batch, and a batch too large for any send buffer is refused
before it is sent. A batch the kernel refuses leaves the previous
table in place; the failure is logged, and the gateway keeps its
record of the change, so the next rebuild that succeeds applies
it.
The rustables crate does not expose rule-handle tracking, so
incremental update of individual rules is not available. Atomic
rebuild was chosen for simplicity and correctness: it eliminates an
entire class of partial-update inconsistency bugs at the cost of
repeating the (cheap) rule construction on every change. The total
rule count is bounded by the pool capacity (2 per mapping, capped
at 2^16) and the port-forward count, both of which are small in
practice.
rule count is bounded by the live-mapping ceiling (1000 mappings,
two rules each) and the port-forward count.
## Configuration Reference
+4 -1
View File
@@ -100,7 +100,10 @@ immediate parent reselection:
- **Periodic re-evaluation** (`reeval_interval_secs`, default 60s):
Re-evaluates parent selection using current MMP link costs, independent
of TreeAnnounce traffic. This catches link degradation after the tree
has stabilized and TreeAnnounce gossip has stopped.
has stabilized and change-driven TreeAnnounce gossip has stopped. It
runs only on a node with two or more peers, and when it finds no
reason to switch parent or become root it re-broadcasts the unchanged
declaration.
- **Flap dampening** (`flap_threshold` / `flap_window_secs` /
`flap_dampening_secs`): If a node switches parents more than
`flap_threshold` times (default 4) within `flap_window_secs` (default
+30 -17
View File
@@ -503,7 +503,11 @@ been offering a better path, but the current path to root remains intact.
and [fips-mesh-layer.md](fips-mesh-layer.md) for complete reference):
`heartbeat_interval_secs` is 10 (send heartbeat if link idle),
`link_dead_timeout_secs` is 30 (declare link dead after no traffic), and
gossip is event-driven on topology change with no periodic refresh.
gossip is event-driven on topology change, with two timed sends: an
announce that the link's receiver reports do not confirm is resent (see
[fips-spanning-tree.md](fips-spanning-tree.md)), and a node with two or
more peers re-broadcasts its unchanged declaration every
`reeval_interval_secs` (60 s by default).
### Asymmetric Failures
@@ -601,10 +605,10 @@ where all links have similar quality, effective depth tracks tree depth closely
and the algorithm produces minimum-depth trees as before.
**Periodic re-evaluation**: `evaluate_parent()` is event-driven — called on
TreeAnnounce receipt or parent loss. After the tree stabilizes and TreeAnnounce
traffic stops, link degradation goes undetected. The periodic re-evaluation
timer (`reeval_interval_secs`) calls `evaluate_parent()` from the tick handler
with current MMP link costs, independent of TreeAnnounce traffic.
TreeAnnounce receipt or parent loss. After the tree stabilizes and change-driven
TreeAnnounce traffic stops, link degradation goes undetected. The periodic
re-evaluation timer (`reeval_interval_secs`) calls `evaluate_parent()` from the
tick handler with current MMP link costs, independent of TreeAnnounce traffic.
### Design Rationale: Local-Only Cost Metrics
@@ -656,9 +660,13 @@ Once converged, what does the network look like and how does it behave?
**Quiescent gossip**:
- TreeAnnounce messages sent only on topology changes, not periodically
- No periodic root refresh — the tree is maintained purely by change-driven gossip
- In a stable network, gossip traffic drops to zero
- TreeAnnounce messages sent on topology changes, and resent, within a budget,
while the link's receiver reports do not confirm them; a node with two or more
peers also re-broadcasts its unchanged declaration every
`reeval_interval_secs` (60 s by default) as a backstop
- No periodic root refresh — the re-broadcast repeats the declaration
without changing or re-signing it
- In a stable network, gossip traffic falls to that backstop
- Bandwidth usage proportional to tree depth, not network size
**Consistent coordinates**:
@@ -670,20 +678,24 @@ Once converged, what does the network look like and how does it behave?
### Steady State Gossip Pattern
**Normal operation.** During normal operation with no topology changes, the root
does not send periodic announcements or refresh its timestamp — it announces
only when its own state changes. Every other node behaves the same way, sending
a TreeAnnounce only on parent selection change or peer link up/down. Tree gossip
is entirely change-driven: when the topology is stable, gossip traffic drops to
zero.
does not refresh its timestamp — its declaration changes only when its own
state changes. Every other node behaves the same way, announcing a new
declaration only on parent selection change or peer link up/down. When the
topology is stable, change-driven gossip stops; what remains is the resend of
any announce that the link's receiver reports have not confirmed, and, on a
node with two or more peers, a re-broadcast of the unchanged declaration every
`reeval_interval_secs`.
### Expected Steady State Properties
**Gossip volume.** Each topology change event produces an update of roughly 100
bytes for the node's own declaration, plus a variable delta for changed
ancestors, giving a total ranging from O(100 bytes) to O(depth * 100 bytes). In
steady state with no topology changes, gossip traffic is zero — there are no
periodic refreshes. Traffic resumes only when links change or nodes join and
depart, and remains negligible compared to application traffic.
steady state with no topology changes, the only gossip is the re-broadcast of
unchanged declarations by nodes with two or more peers, one TreeAnnounce per
peer every `reeval_interval_secs`. Change-driven traffic resumes only when
links change or nodes join and depart, and remains negligible compared to
application traffic.
**Memory usage.** Each node's `TreeState` stores its own entry (~100 bytes),
direct peer entries (~100 bytes each), and ancestry entries (~100 bytes each,
@@ -752,7 +764,8 @@ peer with a path to A) and sits at depth 2.
- A is root
- B, C, D are direct children of A
- E is child of C (one hop to A through C)
- No periodic gossip — TreeAnnounce only on topology changes
- No change-driven gossip; TreeAnnounce traffic is the periodic
re-broadcast backstop only (from A, B, C and D; E has one peer)
**Link failure scenario**:
+3 -2
View File
@@ -82,7 +82,8 @@ own machine, and OpenWrt is a musl target rather than glibc, so none of them
depends on that floor.
See the [project README's Quick start section](../README.md#quick-start)
for download links and per-platform invocations.
for download links, and [packaging/README.md](../packaging/README.md)
for the install commands for each package format.
### FreeBSD
@@ -150,7 +151,7 @@ make deb # or: tarball, ipk, apk, aur, pkg, freebsd, zip, all
The resulting installer lands in `deploy/` at the project root.
Apply it the same way you would a downloaded one (for example
`sudo dpkg -i deploy/fips_*.deb` on Debian/Ubuntu).
`sudo apt install ./deploy/fips_*.deb` on Debian/Ubuntu).
See [packaging/README.md](../packaging/README.md) for per-format
build details, cross-target options, and the full `make` target
+8 -5
View File
@@ -139,8 +139,10 @@ gateway:
Pick a pool CIDR that does **not** overlap with any address space in
use on the LAN or in the mesh (the FIPS mesh occupies `fd00::/8`;
pick a different `fdXX::/N`). The `/112` size yields 65 535 usable
virtual IPs, which is the gateway's hard cap regardless of CIDR
width.
addresses, the most the pool uses whatever the CIDR width. The
gateway holds at most 1000 live mappings at once and admits new
names at up to 10 per second after a burst of 50; an AAAA query for
a new name beyond either limit gets `SERVFAIL`.
This minimum config is enough to start the gateway. The `dns.*` block
is optional and defaults to `listen: "[::1]:5365"` and
@@ -187,9 +189,10 @@ Constraints:
- Must not overlap with `fd00::/8` (the FIPS mesh address space).
- Must not overlap with any LAN-side IPv6 prefix already in use.
- `/112` is the practical width — wider just wastes address space
because the pool is hard-capped at 65 535 usable entries. Narrower is
fine if you want a smaller pool, but you'll reject DNS lookups
faster under churn.
because the pool never uses more than 65 535 addresses. Narrower is
fine if you want a smaller pool; one narrower than `/118` (1023
usable addresses) runs out before the 1000-mapping ceiling is
reached, so new names are refused sooner under churn.
### Choose the DNS listen address
-1
View File
@@ -51,7 +51,6 @@ public test mesh roster:
test-us01 npub1qmc3cvfz0yu2hx96nq3gp55zdan2qclealn7xshgr448d3nh6lks7zel98
test-us02 npub10yffd020a4ag8zcy75f9pruq3rnghvvhd5hphl9s62zgp35s560qrksp9u
test-us03 npub136yqae6na688fs75g95ppps3lxe07fvxefj77938zf47uhm6074sxw8ctm
test-us03-next npub15m6c4ghuegx4pcde6tra8f7smn8vfv2wundyxwhkjynuerkrzmgsy09sh3
test-us04 npub1gd7ye2qp2lphhzx75fynnjzaxx4dqanddecet0wtt5ss5ek8h9ps62wdkf
test-de01 npub1260n42s06vzc7796w0fh3ny7zcpw6tlk4gq3940gmfrzl5c9pv2s3657q8
test-es01 npub17lpmzulpc98d8ff727k6e98atxn3phzupzsqqwe54ytduym747ws4tw5zm
+9
View File
@@ -93,6 +93,15 @@ On Windows the service writes its log to `C:\ProgramData\fips\fips.log`,
rolled at 10 MiB with four old files kept; a foreground run logs to the
console.
On Windows, `install-service.ps1` restricts `C:\ProgramData\fips` to
SYSTEM and Administrators. Reading or editing files there,
`fipsctl keygen`, `fipsctl address` with no argument, and a foreground
run that relies on `C:\ProgramData\fips\fips.yaml` need an elevated
prompt; unelevated, a foreground run may skip that file without saying
so or fail with an access error. Run the installer before `fipsctl keygen`,
and again after moving files into the directory, since a moved file
keeps its old permissions.
The control socket path is derived per
[control-socket.md](control-socket.md).
+1 -1
View File
@@ -90,7 +90,7 @@ daemon.
| Flag | Argument | Default | Description |
| ---- | -------- | ------- | ----------- |
| `-d`, `--dir` | `DIR` | `/usr/local/etc/fips` (macOS, FreeBSD), `/etc/fips` (other Unix), `C:\ProgramData\fips` (Windows) | Output directory for `fips.key` and `fips.pub`. Matches the directory the platform's packaging installs config into, which is where the daemon derives the key paths from. |
| `-d`, `--dir` | `DIR` | `/usr/local/etc/fips` (macOS, FreeBSD), `/etc/fips` (other Unix), `C:\ProgramData\fips` (Windows) | Output directory for `fips.key` and `fips.pub`. Matches the directory the platform's packaging installs config into, which is where the daemon derives the key paths from. On Windows the directory is restricted by `install-service.ps1`, so keygen needs an elevated prompt; run the installer first. |
| `-f`, `--force` | — | off | Overwrite an existing `fips.key`. |
| `-s`, `--stdout` | — | off | Print `nsec` then `npub` to stdout instead of writing files. |
+1 -1
View File
@@ -1061,7 +1061,7 @@ end-to-end design, see
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `gateway.enabled` | bool | `false` | Enable the gateway. Must be `true` for `fips-gateway` to start. |
| `gateway.pool` | string | *(required)* | Virtual IPv6 pool CIDR (e.g., `"fd01::/112"`). Must not overlap with the FIPS mesh address space (`fd00::/8`) or any address space already in use on the LAN. The `/112` size yields 65 535 usable virtual IPs (address 0 in the pool is skipped), which is the gateway's hard cap regardless of CIDR width. |
| `gateway.pool` | string | *(required)* | Virtual IPv6 pool CIDR (e.g., `"fd01::/112"`). Must not overlap with the FIPS mesh address space (`fd00::/8`) or any address space already in use on the LAN. The `/112` size yields 65 535 usable virtual IPs (address 0 in the pool is skipped), the most the pool uses whatever the CIDR width. The gateway holds at most 1000 live mappings at once and admits new names at up to 10 per second after a burst of 50; an AAAA query for a new name beyond either limit gets `SERVFAIL`. |
| `gateway.lan_interface` | string | *(required)* | LAN-facing network interface name (e.g., `"enp3s0"`). Used for proxy-NDP entry installation so LAN clients can resolve the link-layer address of allocated virtual IPs. |
| `gateway.pool_grace_period` | u64 | `60` | Seconds a virtual-IP allocation is retained after its last referencing session ends, before the address is returned to the free pool. Larger values reduce churn for short-lived flows; smaller values reclaim addresses faster. |
+25 -4
View File
@@ -117,9 +117,25 @@ and FSP layers.
## Peer ACL
Mesh-level ACL files at `/etc/fips/peers.allow` and
`/etc/fips/peers.deny` give the operator allowlist/blocklist control
over which npubs may complete the FMP Noise IK link handshake.
Mesh-level ACL files `peers.allow` and `peers.deny`, in `/etc/fips/`
on Linux and other Unix, `/usr/local/etc/fips/` on macOS and FreeBSD
and `C:\ProgramData\fips\` on Windows, give the operator
allowlist/blocklist control over which npubs may complete the FMP
Noise IK link handshake.
On Windows, v0.5.1 and earlier read both files from `\etc\fips\`,
where any local user can create files. The daemon still reads a
`peers.allow` or `peers.deny` left there, on the current drive (for
the service, normally the system drive), while the same file is
missing from `C:\ProgramData\fips\`, and logs a warning when it
does; when the file is in both places, only the `C:\ProgramData\fips\`
copy is read. `install-service.ps1` creates both files empty in
`C:\ProgramData\fips\`, which ends the fallback, and stops without
installing when it finds either file in `\etc\fips\` on the system
drive with no copy in `C:\ProgramData\fips\`, so that the old list is
reviewed and moved or deleted first. To clear a list, empty its file
rather than deleting it, or an old copy in `\etc\fips\` is read
again.
File format:
@@ -182,7 +198,7 @@ mutation; rate-limited msg1s never reach the ACL.
| ---- | ----- | ---- | ------- |
| `/etc/fips/fips.key` | root:root | `0600` | Persistent identity private key (sensitive). |
| `/etc/fips/fips.pub` | root:root | `0644` | Public key (npub). |
| `/etc/fips/fips.yaml` | root:root | `0644` | Daemon configuration (dpkg conffile). |
| `/etc/fips/fips.yaml` | root:root | `0600` | Daemon configuration (seeded by `postinst` from `/usr/share/fips/fips.yaml.example`; not a conffile). |
| `/etc/fips/fips.nft` | root:root | `0644` | nftables baseline (dpkg conffile). |
| `/etc/fips/fips.d/` | root:root | `0755` | Operator drop-in directory. |
| `/etc/fips/hosts` | root:root | `0644` | Optional hostname → npub map (dpkg conffile). |
@@ -192,6 +208,11 @@ mutation; rate-limited msg1s never reach the ACL.
| `/run/fips/api.sock` | root:fips | `0770` | Native datagram API socket, when `node.native_api.enabled` is set (experimental; absent otherwise). |
| `/run/fips/` | root:fips | `0750` | Socket parent directory. |
The `/etc/fips/` paths are the Linux ones. macOS and FreeBSD use
`/usr/local/etc/fips/`; the Windows service keeps the key, config,
hosts and ACL files in `C:\ProgramData\fips\`, which
`install-service.ps1` restricts to SYSTEM and Administrators.
Adding a user to the `fips` group grants `fipsctl` access without
requiring root. The daemon `chown`s the control socket and its parent
directory at bind time, and does the same for the native API socket when
+218
View File
@@ -0,0 +1,218 @@
# FIPS v0.5.2
**Released**: 2026-09-28
v0.5.2 is a maintenance release on the v0.5.x line. It closes several security gaps, changes three defaults, and fixes defects that reach every node as well as the gateway, the Windows service and the Linux, OpenWrt and FreeBSD packages. There is no wire format change, so a mixed mesh works and nodes can be upgraded one at a time.
Read "Before you upgrade", below, if you run an ephemeral node (which is what every package's shipped config runs), the gateway, or FIPS on Windows or OpenWrt.
## At a glance
### Who should upgrade
- **Every node: upgrade.** The TLS library the Nostr relay connections use moves to a release that fixes RUSTSEC-2026-0285, and every default build reaches it. Four fixes also apply on every platform: a link rekey whose reply is lost no longer splits the link when the node answering the rekey runs v0.5.2; a session whose last handshake message is lost no longer stays one-sided; a session rekey whose setup or acknowledgement is lost no longer stops key rotation for the rest of the session; and a node that uses Nostr NAT traversal no longer signs deletion requests with its routing key, which linked its identity to its traversal messages on every relay it reached.
- **Windows: upgrade, and run `install-service.ps1` again.** On v0.5.1 any local user could read the node's key, or create a config, key, hosts file or peer list that the service then used, and a `peers.deny` placed beside the hosts file was never read, so the peer ACL failed open. The service log was also lost.
- **Gateway operators: upgrade.** Session pinning never worked, so a mapping carrying traffic could be reclaimed about two minutes after its last DNS reference. Any host that could reach the LAN resolver could exhaust the address pool. NAT rebuilds were reported as failed past about 105 mappings and did fail past about 313. On an OpenWrt access point with the gateway enabled, `.fips` names stopped resolving and nothing reported it.
- **Debian and Ubuntu, from the `.deb`: upgrade.** An upgrade whose new daemon could not start hung apt; an upgrade left `fips-gateway` stopped until the next reboot; a changed firewall ruleset was not reapplied; and purging the package could leave `.fips` lookups timing out.
- **OpenWrt: upgrade, and read the first-upgrade note below.** An `apk` upgrade left the old processes running until a reboot, and a fresh install enabled `fips-gateway`, which is meant to ship disabled.
- **FreeBSD: upgrade at your convenience.** The daemon's log now rotates; it grew without bound.
- **macOS: upgrade for the fixes every node gets.** Nothing in this release is specific to macOS.
### Before you upgrade
Three defaults changed.
- **An ephemeral node no longer writes `fips.key`.** It used to write the private key of an identity it discards at every restart to that file, overwriting any key already there, including an operator's key when `persistent: true` had been forgotten. It now writes only `fips.pub` and holds the private key in memory. A `fips.key` found at an ephemeral start is renamed to `fips.key.unused`, with a warning, on the first start after the upgrade. **If you started a node once in ephemeral mode and then pinned the key it wrote, that no longer works**: set `node.identity.persistent: true` before upgrading, since most packages restart the daemon as they upgrade and an ephemeral start renames the key. The first persistent start uses the `fips.key` it finds. If an ephemeral start on v0.5.2 has already renamed it to `fips.key.unused`, rename it back to `fips.key` before restarting, as the daemon's warning says. With no key file, the first persistent start generates and saves one, so the npub changes once, at that restart, and stays stable from then on.
- **The gateway's default DNS listen address is `[::1]:5365`.** It was `[::1]:5353`, the mDNS port, which the daemon's own LAN rendezvous, Avahi and systemd-resolved can hold. On OpenWrt the upgrade rewrites the previously shipped `listen: "[::1]:5353"` line and the init script points dnsmasq at whatever port `gateway.dns.listen` sets, so there is nothing to do. On other hosts the upgrade leaves your `fips.yaml` as it is, so first check whether it sets `gateway.dns.listen`; the v0.5.1 example config and deployment guide set it to `"[::1]:5353"`. If it sets it, the gateway stays on that port and warns at startup when it is 5353: leave the resolver as it is, or move off the mDNS port by setting `listen: "[::1]:5365"` and forwarding `.fips` there, changing both together. If it does not set it, the gateway moves to `[::1]:5365`: a resolver you configured by hand to forward `.fips` to `[::1]:5353` must forward to `[::1]:5365` instead, or set `gateway.dns.listen: "[::1]:5353"` to keep the old port. A gateway on 5353 now exits if Avahi, systemd-resolved's MulticastDNS or the daemon's LAN rendezvous already holds that port, so that is the case to move.
- **Windows keeps its config, key, hosts and peer ACL files in `C:\ProgramData\fips`.** That is the directory the service installer already wrote to; the config search, `fipsctl keygen` and the peer ACL defaults each looked somewhere else. Upgrade by running the new `install-service.ps1` from an elevated prompt, with the service stopped. It restricts the directory to SYSTEM and Administrators and creates empty `peers.allow` and `peers.deny` files there. **It stops if it finds a `peers.allow` or `peers.deny` under `\etc\fips` on the system drive with no copy in `C:\ProgramData\fips`**: review that file, move it into `C:\ProgramData\fips` or delete it, and run the installer again. A key left in `%APPDATA%\fips` is still used by a persistent node when the new directory has none, with a note asking you to move it; stop the service and run the installer again after moving any file into the directory, since a moved file keeps its old permissions. If your v0.5.1 service read its config from `\etc\fips`, move its files before upgrading, as described in the Windows upgrade notes below.
Also worth knowing before you start:
- **`fips-gateway` now exits when its DNS listener cannot bind, or stops while the gateway runs**, where it used to stay up with `.fips` resolution dead. systemd and procd restart it or report it failed, and an "address in use" error names the service likely to hold the port. See [Port conflict on the DNS listen port](https://github.com/jmcorgan/fips/blob/v0.5.2/docs/how-to/troubleshoot-gateway.md#port-conflict-on-the-dns-listen-port).
- **The first opkg upgrade to this release re-enables and starts `fips-gateway` on any router that has the `.ipk` installed**, including one where you disabled the gateway by hand. If you had disabled it, run `service fips-gateway stop` and then `service fips-gateway disable` once after upgrading. Stopping it hands dnsmasq's `.fips` forwarding back to the daemon and removes the LAN prefix and route the gateway added; disabling it alone leaves it running, and after a reboot dnsmasq would still forward `.fips` to the gateway's port. Later upgrades keep whatever state the service is in, and an `apk` upgrade on OpenWrt 25 keeps it from the first upgrade on.
- **A `.deb` upgrade now fails if the daemon does not come up within 60 seconds**, and apt exits non-zero naming the unit, where it used to hang with no message.
- **The `.deb` now recommends `nftables`**, so `apt install` pulls it in by default. `dpkg -i` and `--no-install-recommends` do not.
- No configuration key was added or removed. Apart from `gateway.dns.listen`, no configuration key's default changed.
### What changed
- Windows: files in `C:\ProgramData\fips`, the directory restricted to SYSTEM and Administrators, empty peer lists created by the installer, and a service log.
- Gateway: session pinning works, the pool has a mapping ceiling and a rate limit, the NAT rebuild is one transaction and scales past a few hundred mappings, the DNS listener moved off the mDNS port and a dead listener is fatal.
- Sessions and links: lost rekey replies, lost final handshake messages and lost rekey setups no longer leave a link or session stuck, and unauthenticated datagrams can no longer abort a rekey.
- Discovery: bloom filter and spanning-tree announces lost on a link are resent, and a node re-announces its filter when a child joins or leaves.
- Packages: the `.deb` upgrade, remove and purge paths manage every service and the DNS routing; OpenWrt `apk` upgrades restart the services; FreeBSD rotates the log.
- Dependencies: `rustls` 0.23.45 for RUSTSEC-2026-0285.
## Security fixes
### Windows files readable and plantable by any local user
The service installer wrote to `C:\ProgramData\fips`, which inherited `C:\ProgramData`'s default access: any local user could read the files in it and create new ones. Any local account could therefore read the node's key, or create a missing `fips.key`, `fips.yaml` or `hosts` that the service then used. Separately, the service read its peer ACL files from `\etc\fips` on the system drive, where any local user can create files, so a planted list was enforced, while a `peers.deny` placed beside the hosts file was never read.
The installer now creates the directory with an ACL that admits only SYSTEM and Administrators, or replaces the ACL of an existing one, resets the files already in it to inherit that ACL, and refuses to continue if the directory or anything in it is a link or a folder, or if another account owns the directory. It creates empty `peers.allow` and `peers.deny` files, which allow every peer, so the service never falls back to the old location. To clear a list, empty its file rather than deleting it. **For this release only**, a `peers.allow` or `peers.deny` missing from `C:\ProgramData\fips` is still read from `\etc\fips`, with a warning naming both paths; the installer's stop, described above, is what keeps an unreviewed file there from being enforced. The [security reference](https://github.com/jmcorgan/fips/blob/v0.5.2/docs/reference/security.md) describes the fallback.
A foreground `fips.exe` run from an unelevated prompt can no longer read the files in that directory. Reading or editing them, `fipsctl keygen`, and `fipsctl address` with no argument need an elevated prompt.
### Gateway address pool exhausted from the LAN
Every `.fips` query was given a pool address before the gateway looked at what the client had asked for, and an A or HTTPS query was then answered with NODATA. Any host that could reach the LAN resolver could consume the pool one name at a time with a query type it is never given an address for, and could ask for one new name after another until the 65,535 addresses ran out, while every mapping made each NAT rebuild, each pool tick and shutdown slower.
Only AAAA and ANY queries allocate now. The pool refuses a new name once it holds 1000 live mappings, and admits new names at 10 per second after a burst of 50; a refused query gets SERVFAIL, and the "Pool allocation failed" warning says which limit refused it. A name that already has a mapping is answered before either limit is consulted, so names in use keep resolving when the pool is full. The limits are compiled in, not configured.
### Rekey aborted by an unauthenticated datagram
Two handlers gave up a rekey handshake before reading a message that nothing authenticates ahead of that read.
- **A forged link rekey msg2 took the link down.** Anyone on the path who saw the rekey msg1 go out could answer first with a msg2 of the right size, and the two ends were left on different keys until each removed the other on the link-dead timeout, about 30 seconds later. A msg2 that fails to read now leaves the handshake as it was, so the genuine msg2 still completes the rekey.
- **A SessionAck that failed to read ended a session rekey.** The only tie between the ack and the rekey is the datagram's source address. The handshake is now rolled back to its state before the read, so the peer's genuine ack still completes the rekey, and the refusal is counted as `ack_handshake_failed`.
### Routing key published next to traversal messages
After a NAT traversal attempt, a node published NIP-09 deletion requests signed with its routing key. Each one put the node's public identity next to the ids of its offer and answer gift wraps on every relay it reached, which the one-time signing keys on those wraps exist to prevent, and most of them deleted nothing, since a relay deletes a gift wrap only at its recipient's request. The node no longer sends them. A relay that stores the wraps now keeps them until their NIP-40 expiration. Withdrawing an advertisement still sends its deletion request, since that names an event the routing key signed itself.
### Inbound TCP limit held by a connection collision
The inbound TCP pool was keyed by the peer's address alone. A listener on a wildcard address, which is what the shipped configuration binds, can accept two connections from the same peer `ip:port` on two different local addresses, and that collision left the inbound connection counter one above the connections it counts. A host repeating it could hold the counter at the limit and lock out further inbound TCP connections until the daemon restarted. Inbound entries are now keyed by the connection's local address as well as the remote one.
### rustls advisory
The lockfile moves `rustls` from 0.23.43 to 0.23.45 for RUSTSEC-2026-0285: 0.23.43 accepted TLS 1.3 handshake messages across encryption-level boundaries. It is the TLS client the Nostr relay connections use, so every default build reached it. The update is within the version range the dependencies already allowed.
## Gateway
- **Session pinning works.** The session count searched each `/proc/net/nf_conntrack` line for the virtual IP in its compressed form, while the kernel prints it uncompressed, so every mapping counted zero sessions on every kernel. A mapping whose client did not query DNS again was reclaimed about two minutes after its last DNS reference while its traffic was still flowing. Addresses are now compared as addresses. On a kernel without `/proc/net/nf_conntrack`, such as Ubuntu's, the gateway now reads conntrack over netlink, as `conntrack -L` does, where session pinning used to be off. The gateway says at startup which source it can read, or that none is readable and pinning is off, and reports an unreadable source when it happens. The table is read once per tick, off the thread that serves DNS, instead of once per mapping.
- **The NAT table is rebuilt in one netlink transaction.** A rebuild deleted the table in one batch and recreated it in another; between the two the gateway had no NAT at all, and a recreate the kernel refused left the table gone for good. A refused rebuild now leaves the previous table in place.
- **The NAT rebuild scales past a few hundred mappings.** From about 105 mappings rebuilds were logged as failed although they had taken effect; past about 313, new names got a virtual IP with no translation, and a rebuild could delete the whole table. The rebuild now sizes its buffer to the batch and asks for one acknowledgement, and NAT errors name the kernel errno.
- **A DNS query that renews a draining mapping cancels its old grace period**, so the address is no longer reclaimed while the client's renewed answer is still valid.
- **The DNS listener moved to `[::1]:5365`, and a listener that cannot bind or stops is fatal**, as described under "Before you upgrade" above. This applies to a gateway used only for port forwards too. A `gateway.dns.upstream` written as a hostname now works; before, the startup check resolved it and the resolver then stopped on the unparsed name.
The [gateway deployment guide](https://github.com/jmcorgan/fips/blob/v0.5.2/docs/how-to/deploy-gateway.md) and the [troubleshooting guide](https://github.com/jmcorgan/fips/blob/v0.5.2/docs/how-to/troubleshoot-gateway.md) describe the new limits, the listener's failure modes and the session-pinning check.
## Windows
Besides the directory move and the access restriction above, the Windows service now writes its log to `C:\ProgramData\fips\fips.log`, rolled at 10 MiB with four old files kept. A service has no standard output, so everything the daemon logged in service mode was lost, including config-load failures and panic messages. A foreground run still logs to the console.
## Linux packages
- **A `.deb` upgrade whose new daemon cannot start no longer hangs apt.** Each service start is waited on for at most 60 seconds, 90 for `fips-gateway`. A unit that does not come up has its status printed and fails the configure step; a masked unit, or one whose condition is not met, is reported and skipped.
- **The `.deb` scripts manage `fips-gateway`.** An upgrade stopped the daemon, which the gateway requires, and never brought the gateway back. It is now stopped before the daemon on upgrade and restarted afterwards when it is enabled and the daemon came up, and it is stopped and disabled on remove and purge, where its enablement link used to be left behind.
- **A `.deb` upgrade reapplies the firewall ruleset in place.** A changed `/etc/fips/fips.nft` used to take effect only at the next reboot or manual restart. `fips-firewall.service` gains a reload that replaces the ruleset in one transaction, and the `.deb` upgrade reloads it only when it is already active, so an upgrade never turns the firewall on for a host that has not opted in. A tarball or AUR upgrade does not reload it; run `systemctl reload fips-firewall` afterwards if the firewall is active.
- **Purging the `.deb`, or running `uninstall.sh` from the tarball, removes the `.fips` DNS routing** when `fips-dns` was not running at the time, and restarts or reloads the resolver that used it. The host used to keep sending `.fips` queries to a port where nothing listened.
- The `.deb` declares `libgcc-s1 (>= 4.2)`, which its binaries always needed, and every package build now checks the declared dependencies against the libraries the binaries link. It recommends `nftables`, which the opt-in firewall unit runs. The release AUR package lists `dbus` as a dependency and `nftables` as an optional one.
- `-V` on the packaged Linux binaries prints the source revision again.
- **The AUR package is published only after every package workflow for the release tag has succeeded**, so it can appear up to an hour after the release. At v0.5.1 the AUR was updated while the release had 15 of its 17 assets.
## OpenWrt
- **A fresh install no longer enables and starts `fips-gateway`.** The documented `service fips-gateway enable` and `service fips-gateway start` steps are unchanged, and the shipped `fips.yaml` still has `gateway.enabled: true`, so enabling the service is all that is needed.
- **An `apk` upgrade on OpenWrt 25 restarts `fips`**, and `fips-gateway` if it was enabled, so the new binaries run without a reboot.
- The first opkg upgrade re-enables the gateway, as described under "Before you upgrade" above. Later upgrades keep the service's state.
- Starting `fips-gateway` when the config disables the gateway no longer points dnsmasq's `.fips` forwarding at a port nothing listens on.
- The packages no longer ship `/etc/dnsmasq.d/fips.conf`, which OpenWrt's dnsmasq never read; `.fips` forwarding has always come from the UCI server entry.
- The [package README](https://github.com/jmcorgan/fips/blob/v0.5.2/packaging/openwrt-ipk/README.md) gives the `apk add` upgrade command for OpenWrt 25 and a plain `opkg install` for 24.10 and earlier, in place of `--force-reinstall`, which left the gateway disabled.
## FreeBSD
The daemon's log, `/var/log/fips.log`, is rotated by newsyslog, which keeps five compressed generations of 1000 KB, and the rc script starts `daemon(8)` with `-H` so it reopens the log after a rotation.
## Sessions, rekey and discovery
These apply on every platform.
- **A link rekey whose reply is lost no longer splits the link.** The answering side used to switch to the new keys on its own next tick, before the other side had them; when the reply was lost, frames from the answering side were dropped until the link was torn down. It now switches only when a frame on the new keys arrives from the side that started the rekey, and drops keys that were never adopted after a hold, 120 seconds by default. While it holds such keys, that node neither accepts nor starts a rekey on the link, so rekeying on that link waits for up to the hold. The fix is on the answering side: a link where a v0.5.0 or v0.5.1 node answers can still drop, as described under Compatibility below.
- **A session whose last handshake message is lost no longer stays one-sided.** The initiator sent msg3 once; when it was lost, the responder dropped every frame the initiator sent until the next session rekey, or indefinitely with periodic rekey off. The initiator now resends msg3 with backoff until the responder is heard from or the handshake resend limit is reached.
- **A session rekey whose setup or ack is lost now expires** on the handshake timeout, and the next tick starts a fresh one. It used to stay pending and block every later rekey, so the session stopped rotating its keys. Expiries are counted as `rekey_unanswered`.
- A node with no coordinates cached for a session's destination no longer sends its own coordinates in their place, which made the destination cache its own address under the sender's coordinates.
- **Lost bloom filter and spanning-tree announces are resent.** An announce counted as delivered once the transport accepted it, so a dropped datagram or a short link outage left the peer with the old filter or tree position until something else changed, and destinations could stay missing from discovery. The node now confirms each announce from the link's existing receiver reports and resends on loss, within a budget.
- A node re-announces its bloom filter when a peer starts or stops using it as parent, so destinations under a new child become discoverable and a departed child's stop being advertised.
## Links, node health and the control socket
- A heartbeat whose send failed is no longer counted as delivered, and the peer is retried sooner instead of after a whole `heartbeat_interval_secs`.
- A peer that moves to a new address loses the connected UDP socket pinned to the address it left, and a peer reached by NAT traversal now gets its connected socket.
- A configured `ble:` transport that this build cannot construct is reported at startup instead of being dropped silently.
- A DNS responder or TUN thread that dies, including by a panic, now degrades the node's published health, and a dead responder's address is retracted.
- `fipsctl show links` reports the traffic a link has carried; every link used to report zero.
- `fipsctl show peers` reports a peer that has been silent longer than `heartbeat_interval_secs` as `stale`; every peer used to read `connected` until it was removed.
- Replacing the peer list at runtime with `Node::update_peers` now updates `.fips` names and peer ACL entries written as an alias.
- A persistent node whose identity key path cannot be examined, such as a key symlinked onto a volume that did not mount, now refuses to start and names the path, instead of coming up under a new identity.
- On Linux, an empty datagram sent through the native datagram API just before a close may read as the close; the [native API reference](https://github.com/jmcorgan/fips/blob/v0.5.2/docs/reference/native-api.md) now says so.
## Compatibility
v0.5.2 is wire-compatible with v0.5.1. No file that defines a frame, message, TLV or encoding changed between the two releases, so a mixed mesh works and nodes can be upgraded one at a time with no coordinated restart. The msg3 resend sends the same msg3, and the rekey and announce changes alter when a node sends or switches keys, not what it sends. Mixed-version behavior is measured by the interop run described below. Until the older nodes are upgraded, a link to a v0.5.0 or v0.5.1 node can still drop and be re-established when that node answers a link rekey and its reply is lost: the fix is on the answering side.
No configuration key was added or removed; the gateway's DNS listen default is the only configuration-key default that changed.
## Upgrade notes
**Debian and Ubuntu.** `apt install ./fips_0.5.2_<arch>.deb`, with `amd64` or `arm64` in place of `<arch>`. Once the new files are in place the package reloads the firewall if it is running, starts the daemon, then enables and starts `fips-dns` (every upgrade re-enables it; `systemctl mask fips-dns` keeps it off), and restarts `fips-gateway` if it is enabled. If the daemon does not come up, apt exits non-zero and prints its status; check `fipsctl show status` afterwards either way.
**Windows.** Open an elevated PowerShell in the folder where you unzipped the new ZIP, and in it: stop the service (`Stop-Service fips`); if your v0.5.1 service read its config from `\etc\fips`, move its files as described in the next paragraph; run the installer (`powershell -ExecutionPolicy Bypass -File .\install-service.ps1`, since the execution policy can refuse an unsigned script from a downloaded ZIP); then start the service (`Start-Service fips`). In Windows PowerShell 5.1, `sc` is an alias for `Set-Content`, so use `sc.exe` if you prefer the service control tool. The installer cannot replace `fips.exe` while the service is running and stops with a file-in-use error, so stop the service before every run of it. Resolve any legacy peer list the installer reports, as described under "Before you upgrade" above. The service log is at `C:\ProgramData\fips\fips.log`.
If your v0.5.1 service was set up by hand to read its config from `\etc\fips`, rather than by v0.5.1's `install-service.ps1` (which already used `C:\ProgramData\fips` and set `FIPS_CONFIG`), move `fips.yaml` and `fips.key` from `\etc\fips` into `C:\ProgramData\fips` after stopping the service and before running the new installer. The v0.5.2 installer sets `FIPS_CONFIG` to `C:\ProgramData\fips\fips.yaml`, and the service then reads only that file and the key beside it. Without the move, the service starts from whatever `C:\ProgramData\fips` holds. With the default config the installer places there, which does not enable a persistent identity, the node comes up under a new identity, and nothing warns about the files left in `\etc\fips`.
**OpenWrt.** Copy the package to `/tmp` on the router. On OpenWrt 25 and later run `apk add --allow-untrusted /tmp/fips_<new-version>_<arch>.apk`; on 24.10 and earlier run `opkg install /tmp/fips_<new-version>_<arch>.ipk`, with the file name you downloaded in place of the placeholders. After a first opkg upgrade, if you had disabled `fips-gateway`, run `service fips-gateway stop` and then `service fips-gateway disable`.
**FreeBSD.** Install the new package with `pkg install ./fips-0.5.2-freebsd-amd64.pkg`, then run `service fips restart` and, if you use it, `service fips_dns restart`, and check `fipsctl show status`. The package restarts the services itself only when pkg takes its upgrade path, which `pkg add` over an installed package does not, and the log rotation fix takes effect only once the daemon restarts. Do not remove the package before adding the new one: removal stops both services and removes the `.fips` resolver drop-in.
**Arch Linux and the systemd tarball.** pacman does not restart services: after the AUR package updates, run `systemctl restart fips.service`, and `systemctl restart fips-gateway.service` if you run the gateway. The tarball's `install.sh` restarts `fips` if it was running, but stopping `fips` also stops `fips-dns` and `fips-gateway`, and it starts neither again: run `systemctl start fips-dns.service`, and `systemctl start fips-gateway.service` if you run the gateway. After either, run `systemctl reload fips-firewall` if the firewall is active.
**Gateway on other hosts.** If `fips.yaml` does not set `gateway.dns.listen` and a local resolver forwards `.fips` to `[::1]:5353`, point it at `[::1]:5365` right after the upgrade, or set `listen` to the old port. If `fips.yaml` sets it, the gateway stays on that port and nothing needs to change; see "Before you upgrade" above.
**Rolling upgrade.** No coordination is needed. Upgrade nodes in any order.
## What was measured and what was not
**Measured, at the release's source content** (the same source as the tag, with the version string at `0.5.2-dev`):
- The full local test suite: 36 suites passed, including the chaos scenarios, the NAT traversal, gateway, firewall and DNS resolver suites, and the `.deb` install suite on five distributions.
- `cargo audit`: no vulnerabilities, with the same four allowed warnings as v0.5.1.
- The 100-node discovery test.
- The Debian package built through the release's container script, with all four binaries at a glibc floor of 2.34 against the 2.35 the package declares.
- The systemd tarball rebuilt twice on the same machine and compared byte for byte: identical.
- The OpenWrt `.ipk` cross-compiled for aarch64. The `.apk` cannot be built on the build host used here and is built by the release workflow.
- The wire format, by a diff showing that no file defining a frame, message, TLV or encoding changed since v0.5.1.
- A mixed-version mesh of v0.5.2, v0.5.1 and v0.5.0 nodes passed the interop suite, once without impairment and in eight runs with added delay and 2% packet loss: every pair passed the suite's connectivity checks, including after two rekey cycles; every pair of a v0.5.2 node and an older one completed link rekeys with each side starting them; and the logs showed no panic, error, decryption failure or handshake failure. Under loss, a link whose answering node ran v0.5.0 or v0.5.1 dropped after that node's rekey reply was lost and was re-established, three times in the eight runs. That is the defect this release fixes on the answering side: no link dropped where a v0.5.2 node answered or between two v0.5.2 nodes, and a v0.5.2-only mesh under the same loss dropped no link in four runs.
- The Windows installer and service on Windows Server 2025 (GitHub-hosted runners), under Windows PowerShell 5.1 and PowerShell 7, from a build that differs from the release's source content only in the version string, the changelog and the rustls update: the directory and peer-file permissions on a fresh install and on reruns; the installer's stop on a peer list left under `\etc\fips`; a peer list planted there by a standard user, which the service ignores once the current file exists and otherwise enforces with a warning; the upgrade from a service set up by v0.5.1's installer, which kept its node address; and the upgrade from a v0.5.1 service that read its config from `\etc\fips`, which came up under a new identity with no warning, as described in the Windows upgrade notes above.
- The gateway DNS port change on an OpenWrt 24 router running this release's source content (tested by Arjen): the gateway listened on `[::1]:5365` and dnsmasq forwarded `.fips` to it, and a gateway that cannot bind its DNS listener now exits instead of running with DNS dead. The same test found that if `fips-gateway` fails to start, dnsmasq is left forwarding `.fips` to the gateway's port until the gateway runs; this predates v0.5.2. If `.fips` stops resolving on a router, check `logread | grep fips-gateway`, and run `service fips-gateway stop` to hand `.fips` back to the daemon until the cause is fixed.
**Not measured when this was written:**
- No published v0.5.2 artifact existed. The checks above ran against artifacts built by the same scripts the release workflow uses.
- A `.deb` upgrade from an installed v0.5.1 package, which is the only case in which v0.5.1's removal scripts meet v0.5.2's install scripts. The install suite repacks one package against itself.
- The OpenWrt maintainer scripts are tested in CI in a BusyBox container against stubs, not under a real opkg or apk upgrade on a router.
- On OpenWrt router hardware, the upgrade rewriting the shipped gateway DNS listen line: the router test above ran the new port and did not report the rewrite separately.
- The Windows checks on a desktop edition of Windows, which can differ from the server runners in the owner given to new files and in when the service sees the `FIPS_CONFIG` the installer sets. The runner checks did not exercise `uninstall-service.ps1` or the TUN adapter.
- The changed rekey behavior under real mixed-version traffic on live nodes. No field soak was run for this release; the interop run described above stands in for it.
- Upgrading a running node in place, forwarding across more than one hop, loss patterns other than random 2% packet loss, and runs longer than a few minutes: the interop run started each mesh fresh, used full meshes only, and ran each repetition for about six minutes.
- Reproducibility across machines, which has never been tested.
- The FreeBSD upgrade command, and the statement that `pkg add` over an installed package does not take pkg's upgrade path: both are taken from pkg's source and were not run on a FreeBSD host.
## Getting v0.5.2
- **Linux x86_64 / aarch64**: `.deb` and tarball at the [v0.5.2 release page](https://github.com/jmcorgan/fips/releases/tag/v0.5.2).
- **Arch Linux**: `fips` from the AUR, once every package workflow for the release has finished.
- **macOS**: `.pkg` at the v0.5.2 release page.
- **Windows**: ZIP at the v0.5.2 release page.
- **FreeBSD (x86_64)**: `.pkg` at the v0.5.2 release page.
- **OpenWrt**: `.ipk` (OpenWrt 24.x and earlier) or `.apk` (OpenWrt 25+) at the v0.5.2 release page, for `aarch64_cortex-a53` and `x86_64`.
- **From source**: `cargo build --release` from a checkout of the v0.5.2 tag (Rust 1.94.1 per `rust-toolchain.toml`; `libclang-dev` is a required Linux build prerequisite, and on glibc Linux so are `libdbus-1-dev` and `pkg-config`).
- **Nix / NixOS**: `nix build .#fips` from a checkout of the v0.5.2 tag builds the binaries from source with the pinned toolchain and no manual prerequisites (see the Nix section of [`packaging/README.md`](https://github.com/jmcorgan/fips/blob/v0.5.2/packaging/README.md)). The NixOS module's documented flake input, `github:jmcorgan/fips`, follows the default branch; to run v0.5.2, set `inputs.fips.url = "github:jmcorgan/fips/v0.5.2"`, then run `nix flake update fips` and `nixos-rebuild switch`.
There is no Android daemon artifact. Android is supported as an embedded crate.
The full per-commit changelog lives in [`CHANGELOG.md`](https://github.com/jmcorgan/fips/blob/v0.5.2/CHANGELOG.md). Issues and discussion at [github.com/jmcorgan/fips](https://github.com/jmcorgan/fips). Security reports have a private channel; see [`SECURITY.md`](https://github.com/jmcorgan/fips/blob/v0.5.2/SECURITY.md).
## Contributors
Thanks to everyone who contributed code, packaging work, bug reports, or reviews to this release.
- [@jmcorgan](https://github.com/jmcorgan) (Johnathan Corgan): release shepherd; the gateway, Windows, packaging, session and discovery fixes.
- [@mmalmi](https://github.com/mmalmi) (Martti Malmi): the gateway fix that keeps a renewed DNS answer valid while its mapping is draining ([#169](https://github.com/jmcorgan/fips/pull/169)).
- [@Origami74](https://github.com/Origami74) (Arjen): found on router hardware that the gateway's DNS port collided with the daemon's mDNS responder on an OpenWrt access point, and reviewed the OpenWrt packaging fixes on OpenWrt 25.
- [@fr34aky](https://github.com/fr34aky): version placeholders in the FreeBSD packaging examples ([#152](https://github.com/jmcorgan/fips/pull/152), landed as `93d45191`), and a hostname test that no longer depends on the host's DNS search domain ([#154](https://github.com/jmcorgan/fips/pull/154), landed as `a1c0cd42`).
- [@shaibearary](https://github.com/shaibearary) (Sherry): a fix to the mesh test loop, which aborted its rekey setup under the bash that ships with macOS ([#155](https://github.com/jmcorgan/fips/pull/155), landed as `31ebec62`).
- [@Ghost-glitch-hub](https://github.com/Ghost-glitch-hub): reported that `fipsctl show links` showed zero traffic for active peers ([#158](https://github.com/jmcorgan/fips/issues/158)).
<!-- markdownlint-disable-file MD013 -->
+4 -3
View File
@@ -102,8 +102,8 @@ You should see:
- `service fips status` reports `running`.
- `fips0` exists and has one `inet6 fd97:...` address. That is the
AP's mesh-side identity.
- `fipsctl show peers` lists at least one peer with active
connectivity (not `idle` / not zero bytes).
- `fipsctl show peers` lists at least one peer whose `connectivity`
reads `connected`.
Confirm the AP can resolve a known mesh node by name:
@@ -141,7 +141,8 @@ gateway:
Three things to notice:
- `pool: "fd01::/112"` — the virtual-IP CIDR the gateway hands out
to LAN clients. 65 536 addresses, the gateway's hard cap. Pick a
to LAN clients. 65 535 usable addresses, the most the pool uses;
the gateway holds at most 1000 live mappings at once. Pick a
different `fdXX::/N` prefix if `fd01::/112` collides with anything
on your network.
- `lan_interface: "br-lan"` — the OpenWrt LAN bridge. The gateway
+10 -6
View File
@@ -114,7 +114,7 @@ never prompted for or clobbered on upgrade. To reset to defaults, remove
make deb
# Install
sudo dpkg -i deploy/fips_<version>_<arch>.deb
sudo apt install ./deploy/fips_<version>_<arch>.deb
# Remove (preserves config and keys)
sudo dpkg -r fips
@@ -261,14 +261,18 @@ installation and configuration instructions.
### OpenWrt (`.ipk`, opkg — OpenWrt 24.x and earlier)
Cross-compiled with cargo-zigbuild and assembled as a standard `.ipk`
archive. Supports aarch64, mipsel, mips, arm, and x86\_64 targets.
archive. The build script accepts aarch64, mipsel, mips, arm and
x86\_64; releases publish aarch64 and x86\_64. The MIPS targets are
not built: 32-bit MIPS has no 64-bit atomics, which fips and
`nostr-relay-pool` both use (see the comment in
`.github/workflows/package-openwrt.yml`).
```sh
# Build (default: aarch64)
make ipk
# Build for a specific architecture
bash packaging/openwrt-ipk/build-ipk.sh --arch mipsel
bash packaging/openwrt-ipk/build-ipk.sh --arch x86_64
```
See [openwrt-ipk/README.md](openwrt-ipk/README.md) for router-specific
@@ -410,13 +414,13 @@ powershell -File packaging/windows/build-zip.ps1
# Extract and install as service (requires Administrator)
Expand-Archive deploy\fips-<version>-windows-x86_64.zip -DestinationPath fips
cd fips
powershell -File install-service.ps1
powershell -ExecutionPolicy Bypass -File install-service.ps1
# Uninstall (preserves config)
powershell -File uninstall-service.ps1
powershell -ExecutionPolicy Bypass -File uninstall-service.ps1
# Uninstall and remove config
powershell -File uninstall-service.ps1 -RemoveAll
powershell -ExecutionPolicy Bypass -File uninstall-service.ps1 -RemoveAll
```
### Arch Linux (AUR)
-1
View File
@@ -25,7 +25,6 @@
test-us01 npub1qmc3cvfz0yu2hx96nq3gp55zdan2qclealn7xshgr448d3nh6lks7zel98
test-us02 npub10yffd020a4ag8zcy75f9pruq3rnghvvhd5hphl9s62zgp35s560qrksp9u
test-us03 npub136yqae6na688fs75g95ppps3lxe07fvxefj77938zf47uhm6074sxw8ctm
test-us03-next npub15m6c4ghuegx4pcde6tra8f7smn8vfv2wundyxwhkjynuerkrzmgsy09sh3
test-us04 npub1gd7ye2qp2lphhzx75fynnjzaxx4dqanddecet0wtt5ss5ek8h9ps62wdkf
test-de01 npub1260n42s06vzc7796w0fh3ny7zcpw6tlk4gq3940gmfrzl5c9pv2s3657q8
test-es01 npub17lpmzulpc98d8ff727k6e98atxn3phzupzsqqwe54ytduym747ws4tw5zm
+1 -1
View File
@@ -72,7 +72,7 @@ export APK_BIN="$PWD/build/src/apk"
```bash
# from the repo root
./packaging/openwrt-apk/build-apk.sh --arch aarch64 # or x86_64, mipsel, mips, arm
./packaging/openwrt-apk/build-apk.sh --arch aarch64 # or x86_64; releases publish these two
```
Output: `dist/fips_<version>_<openwrt-arch>.apk`. Override the version with
+6
View File
@@ -95,6 +95,12 @@ make package/fips/compile V=s
The resulting `.ipk` is placed in `bin/packages/<arch>/`.
A package built from this `Makefile` carries none of the maintainer scripts in
`scripts/`. Those scripts enable and start `fips` on install and implement the
gateway-enablement and upgrade behavior described below, so that description
does not cover a package built this way. Released packages are built by
`build-ipk.sh` (and `../openwrt-apk/build-apk.sh`), which install those scripts.
### 4. Pin the source version
For reproducible production builds, replace `PKG_SOURCE_VERSION:=master` in
+11 -1
View File
@@ -98,7 +98,17 @@ on the edge rather than once per retry.
Six nodes with per-node allowlist files mounted at the runtime ACL
paths, exercising insiders, outsiders and allowed remotes at once to
check which peer pairs are admitted and which are rejected.
check which peer pairs are admitted and which are rejected. Run by hand
only; retired from both CI runners as redundant with the unit and
in-process ACL tests (see `ci-local.sh`).
### [openwrt/](openwrt/) -- OpenWrt Maintainer Scripts
Runs the OpenWrt package scripts and the `fips-gateway` init script
under busybox `ash` against stubbed init scripts and `uci`, and checks
what the real `build-apk.sh` and `build-ipk.sh` package. No router,
opkg or apk-tools is involved. Part of both CI runners as
`openwrt-scripts`.
### [native-api/](native-api/) -- Native Datagram API
+3 -1
View File
@@ -244,7 +244,8 @@ working copy at all.
## How to read the output
The driver runs seven phases:
The driver runs eight phases (0 to 7), plus 1b and 5b when data-plane
streams are on:
| Phase | Check |
| ----- | ---------------------------------------------------------------- |
@@ -255,6 +256,7 @@ The driver runs seven phases:
| 4 | Wait out a second rekey cycle. |
| 5 | All pairs still ping after the second rekey. |
| 6 | Per-node / per-pair interop log analysis. |
| 7 | Every node's mesh-size estimate within ±25% of N after warmup. |
When data-plane streams are on (`--topology`, or `FIPS_INTEROP_STREAMS`),
Phase 1b measures stream loss over a quiet control window and Phase 5b