Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
1617f6ec1c | ||
|
|
025ab49d26 | ||
|
|
733ee512d3 | ||
|
|
b6bd28f77c | ||
|
|
d72e619c51 | ||
|
|
0e57216d98 | ||
|
|
77fdd52fe0 | ||
|
|
42b88c9bb8 | ||
|
|
eaba693b18 | ||
|
|
d52d7debb7 | ||
|
|
77ecfda1a1 | ||
|
|
9b1016ffaf | ||
|
|
5cda4a9a55 | ||
|
|
8094a51a82 | ||
|
|
0cc3de3daa | ||
|
|
2d18d019d6 | ||
|
|
64cc30df12 | ||
|
|
6ce1406664 | ||
|
|
e81fd4b477 | ||
|
|
fac4450694 | ||
|
|
253dddabe3 | ||
|
|
cd56fee7cf | ||
|
|
f0bb29ff6e | ||
|
|
e471807239 | ||
|
|
c412646498 | ||
|
|
53ad528f7d | ||
|
|
b3a1fb464f | ||
|
|
6807a3213b | ||
|
|
b547dd70f5 | ||
|
|
1d7d0d2522 | ||
|
|
67e660d813 | ||
|
|
c255e3f4a2 | ||
|
|
f32bc83034 | ||
|
|
e4f37082c2 | ||
|
|
7daca6bcf1 | ||
|
|
0fcf0f6f8f | ||
|
|
9112c8f7f0 | ||
|
|
db5b6b10bd | ||
|
|
18019bb1b5 | ||
|
|
5abf9a9325 | ||
|
|
4cdf382038 | ||
|
|
a62a0a6cf4 | ||
|
|
7fc890b7a2 | ||
|
|
a78f670a6a | ||
|
|
2f95929862 | ||
|
|
bcc9c525d3 | ||
|
|
f66be793b8 | ||
|
|
5c92fffa2b | ||
|
|
43639fecb9 | ||
|
|
33a2063672 | ||
|
|
5611e976ad | ||
|
|
d822ee8b3c | ||
|
|
ff40966832 | ||
|
|
b8b1bb03a0 | ||
|
|
616010f8c8 | ||
|
|
1d8e698b57 | ||
|
|
e08f42e3cc | ||
|
|
81c0547bdf | ||
|
|
5ed2d36464 | ||
|
|
9204888a54 | ||
|
|
00f4a4c7af | ||
|
|
9c96c9193d | ||
|
|
c86dc32197 | ||
|
|
037a965a93 | ||
|
|
953137ede7 | ||
|
|
ae607431eb | ||
|
|
996a591001 | ||
|
|
da5d23ccb7 | ||
|
|
8448e38510 | ||
|
|
a41f80a776 | ||
|
|
ab2edec2c6 | ||
|
|
239cbdc4ba | ||
|
|
96c6b7dea8 | ||
|
|
e641eb5b5f | ||
|
|
37c2973e2f | ||
|
|
674c7fe1ff | ||
|
|
23c6609a6e | ||
|
|
3092c95d54 | ||
|
|
c8502cdb97 | ||
|
|
6def31bcf6 | ||
|
|
bf77ececad | ||
|
|
34e00b9f6e | ||
|
|
1e3b2c319e | ||
|
|
f16b837a12 | ||
|
|
cbc78091ab | ||
|
|
be0708ac9b | ||
|
|
ad5ad53848 | ||
|
|
ed312ac6f2 | ||
|
|
de8db82614 | ||
|
|
03f6db58e8 | ||
|
|
db4b32110c | ||
|
|
c83e14ac97 | ||
|
|
745b523ac6 | ||
|
|
5cdcff7386 | ||
|
|
b36966be3a | ||
|
|
83b20b3078 | ||
|
|
5087ef9a95 | ||
|
|
213c0e87c3 | ||
|
|
c009eb7514 | ||
|
|
7780dffa93 | ||
|
|
7b3c2daa12 | ||
|
|
5abae0859e | ||
|
|
2d342a4e47 | ||
|
|
5029b40d49 | ||
|
|
6698c4d669 | ||
|
|
2a943e6695 | ||
|
|
7e002a3883 | ||
|
|
d16acf8cea | ||
|
|
42834b8008 | ||
|
|
5208d3222a | ||
|
|
0b81f15369 | ||
|
|
fe205e74de | ||
|
|
774e33fd27 | ||
|
|
7494ed058d | ||
|
|
19cb776216 | ||
|
|
68dafbc72a | ||
|
|
7258469b18 | ||
|
|
0d4ffc61f0 | ||
|
|
5645284893 | ||
|
|
e9da598f8a | ||
|
|
6196307f0e | ||
|
|
68bdcb2c75 | ||
|
|
48b1617497 | ||
|
|
13c0b70dc3 | ||
|
|
e693f4fb7e | ||
|
|
1e4f375dcc | ||
|
|
60e5fefb1f | ||
|
|
51119347c3 | ||
|
|
aac96510d0 | ||
|
|
a859da7748 | ||
|
|
4370441e48 | ||
|
|
864a8bcc9e | ||
|
|
15628e5b41 | ||
|
|
adfbeb2348 | ||
|
|
6633d22132 | ||
|
|
0382642d1e | ||
|
|
f6f2bea792 | ||
|
|
0ff9139b64 | ||
|
|
7d33f1f2c9 | ||
|
|
97fc29eb82 | ||
|
|
8c4455cc1c | ||
|
|
59f21ca185 | ||
|
|
5d27efb179 | ||
|
|
7224ce34f6 | ||
|
|
8e38d889fa | ||
|
|
bd08505002 | ||
|
|
4bc30d2b8a | ||
|
|
75466ae4e8 | ||
|
|
db9549885a | ||
|
|
8f1494853a | ||
|
|
cb6f263a1d | ||
|
|
d801fd0052 | ||
|
|
88fcf57067 | ||
|
|
89352d3218 | ||
|
|
d3385b902a | ||
|
|
7f33e5f867 | ||
|
|
79a10a2700 | ||
|
|
fc8c0dce15 | ||
|
|
71e2955da0 | ||
|
|
9519dc1cf4 | ||
|
|
bce5619b74 | ||
|
|
b8fbecc575 | ||
|
|
0ff3f029ed | ||
|
|
9757877c0a | ||
|
|
94884876b8 | ||
|
|
9d9e2b05a1 | ||
|
|
a16370e78d | ||
|
|
537eaf0db7 | ||
|
|
6aff490251 | ||
|
|
7a643a9ac3 | ||
|
|
c164de8808 | ||
|
|
fed6cc6987 | ||
|
|
5053cf673d | ||
|
|
8d51dbd268 | ||
|
|
9e63b42bd9 | ||
|
|
c8b7459fbc | ||
|
|
8c349f524e | ||
|
|
0e098cadad | ||
|
|
ebabb6e93c | ||
|
|
a5130b357e | ||
|
|
6e03adde72 | ||
|
|
1020f90828 | ||
|
|
0f24333c0d | ||
|
|
ef73e54316 | ||
|
|
8697897047 | ||
|
|
9e11a8e7ba | ||
|
|
e8ef15acb7 | ||
|
|
5f7fe989f3 | ||
|
|
324535e76d | ||
|
|
959134657f | ||
|
|
6c90cf6c02 | ||
|
|
c0f30d8fe8 | ||
|
|
d873d0e00e | ||
|
|
35ff4a5d0a | ||
|
|
1bfb58845a |
@@ -0,0 +1,5 @@
|
||||
# rustfmt bulk reformat (maint)
|
||||
13c0b70dc3111cf94fef217b0f8b5fdbe469d3eb
|
||||
|
||||
# rustfmt master-only code
|
||||
e9da598f8ab13de5dea3a1496531d675af6a0b94
|
||||
@@ -0,0 +1,5 @@
|
||||
# Keep snapshot test fixtures LF-only across all platforms so that
|
||||
# Windows checkouts with core.autocrlf=true don't convert them to CRLF
|
||||
# (which would mismatch the actual JSON serialization output during
|
||||
# snapshot comparison).
|
||||
src/control/snapshots/*.json text eol=lf
|
||||
@@ -0,0 +1,37 @@
|
||||
name: AUR Publish
|
||||
|
||||
on:
|
||||
workflow_dispatch:
|
||||
push:
|
||||
tags:
|
||||
- 'v*'
|
||||
|
||||
jobs:
|
||||
aur-publish-fips:
|
||||
name: Publish fips to AUR
|
||||
runs-on: ubuntu-latest
|
||||
continue-on-error: true
|
||||
if: "!contains(github.ref_name, '-')"
|
||||
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
|
||||
- name: Update pkgver in PKGBUILD
|
||||
run: |
|
||||
VERSION="${GITHUB_REF_NAME#v}"
|
||||
sed -i "s/^pkgver=.*/pkgver=${VERSION}/" packaging/aur/PKGBUILD
|
||||
|
||||
- name: Publish to AUR
|
||||
uses: KSXGitHub/github-actions-deploy-aur@v4.1.2
|
||||
with:
|
||||
pkgname: fips
|
||||
pkgbuild: packaging/aur/PKGBUILD
|
||||
updpkgsums: true
|
||||
assets: |
|
||||
packaging/aur/fips.sysusers
|
||||
packaging/aur/fips.tmpfiles
|
||||
packaging/aur/fips.install
|
||||
commit_username: ${{ github.repository_owner }}
|
||||
commit_email: ${{ secrets.AUR_EMAIL }}
|
||||
ssh_private_key: ${{ secrets.AUR_SSH_PRIVATE_KEY }}
|
||||
commit_message: "Update to ${{ github.ref_name }}"
|
||||
@@ -18,14 +18,46 @@ permissions:
|
||||
env:
|
||||
CARGO_TERM_COLOR: always
|
||||
RUST_BACKTRACE: 1
|
||||
SOURCE_DATE_EPOCH: 0 # overridden per-step after checkout
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Job 1 – Build matrix
|
||||
#
|
||||
# Builds on Linux x86_64 and Linux aarch64. macOS and Windows are in a
|
||||
# separate ci-compat.yml workflow so their failures don't mark this run red.
|
||||
# Builds on Linux x86_64, Linux aarch64, and macOS.
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
jobs:
|
||||
fmt:
|
||||
name: Format check
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
- uses: dtolnay/rust-toolchain@stable
|
||||
with:
|
||||
components: rustfmt
|
||||
- run: cargo fmt --check
|
||||
|
||||
clippy:
|
||||
name: Clippy
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
- name: Install system dependencies
|
||||
run: sudo apt-get update && sudo apt-get install -y libdbus-1-dev
|
||||
- uses: dtolnay/rust-toolchain@stable
|
||||
with:
|
||||
components: clippy
|
||||
- name: Cache Cargo registry + build
|
||||
uses: actions/cache@v4
|
||||
with:
|
||||
path: |
|
||||
~/.cargo/registry
|
||||
~/.cargo/git
|
||||
target
|
||||
key: ${{ runner.os }}-cargo-clippy-${{ hashFiles('**/Cargo.lock') }}
|
||||
restore-keys: |
|
||||
${{ runner.os }}-cargo-
|
||||
- run: cargo clippy --all-targets --all-features -- -D warnings
|
||||
|
||||
build:
|
||||
name: Build (${{ matrix.os }})
|
||||
runs-on: ${{ matrix.os }}
|
||||
@@ -36,10 +68,31 @@ jobs:
|
||||
include:
|
||||
- os: ubuntu-latest
|
||||
- os: ubuntu-24.04-arm
|
||||
- os: macos-latest
|
||||
- os: windows-latest
|
||||
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
|
||||
- name: Set SOURCE_DATE_EPOCH from git (Unix)
|
||||
if: runner.os != 'Windows'
|
||||
run: echo "SOURCE_DATE_EPOCH=$(git log -1 --format=%ct)" >> "$GITHUB_ENV"
|
||||
|
||||
- name: Set SOURCE_DATE_EPOCH from git (Windows)
|
||||
if: runner.os == 'Windows'
|
||||
shell: pwsh
|
||||
run: |
|
||||
$epoch = git log -1 --format=%ct
|
||||
echo "SOURCE_DATE_EPOCH=$epoch" >> $env:GITHUB_ENV
|
||||
|
||||
- name: Install system dependencies (Linux only)
|
||||
if: runner.os == 'Linux'
|
||||
run: sudo apt-get update && sudo apt-get install -y libdbus-1-dev nftables
|
||||
|
||||
- name: Validate fips.nft syntax (Linux only)
|
||||
if: runner.os == 'Linux'
|
||||
run: sudo nft -c -f packaging/common/fips.nft
|
||||
|
||||
- name: Install Rust toolchain
|
||||
uses: dtolnay/rust-toolchain@stable
|
||||
|
||||
@@ -57,6 +110,19 @@ jobs:
|
||||
- name: Build
|
||||
run: cargo build --release
|
||||
|
||||
- name: SHA-256 hashes (Linux)
|
||||
if: runner.os == 'Linux'
|
||||
run: sha256sum target/release/fips target/release/fipsctl target/release/fipstop target/release/fips-gateway
|
||||
|
||||
- name: SHA-256 hashes (macOS)
|
||||
if: runner.os == 'macOS'
|
||||
run: shasum -a 256 target/release/fips target/release/fipsctl target/release/fipstop
|
||||
|
||||
- name: SHA-256 hashes (Windows)
|
||||
if: runner.os == 'Windows'
|
||||
shell: pwsh
|
||||
run: Get-FileHash target\release\fips.exe, target\release\fipsctl.exe, target\release\fipstop.exe -Algorithm SHA256
|
||||
|
||||
# Upload the Linux binary so integration jobs can use it without rebuilding
|
||||
- name: Upload Linux binary
|
||||
if: matrix.os == 'ubuntu-latest'
|
||||
@@ -67,6 +133,7 @@ jobs:
|
||||
target/release/fips
|
||||
target/release/fipsctl
|
||||
target/release/fipstop
|
||||
target/release/fips-gateway
|
||||
retention-days: 1
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
@@ -82,6 +149,12 @@ jobs:
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
|
||||
- name: Set SOURCE_DATE_EPOCH from git
|
||||
run: echo "SOURCE_DATE_EPOCH=$(git log -1 --format=%ct)" >> "$GITHUB_ENV"
|
||||
|
||||
- name: Install system dependencies
|
||||
run: sudo apt-get update && sudo apt-get install -y libdbus-1-dev
|
||||
|
||||
- name: Install Rust toolchain
|
||||
uses: dtolnay/rust-toolchain@stable
|
||||
|
||||
@@ -119,14 +192,105 @@ jobs:
|
||||
check_name: Unit Tests Summary
|
||||
fail_on_failure: false
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Job 2b – Unit tests (macOS)
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
test-macos:
|
||||
name: Unit tests (macOS)
|
||||
runs-on: macos-latest
|
||||
needs: [build]
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
|
||||
- name: Set SOURCE_DATE_EPOCH from git
|
||||
run: echo "SOURCE_DATE_EPOCH=$(git log -1 --format=%ct)" >> "$GITHUB_ENV"
|
||||
|
||||
- name: Install Rust toolchain
|
||||
uses: dtolnay/rust-toolchain@stable
|
||||
|
||||
- name: Cache Cargo registry + build
|
||||
uses: actions/cache@v4
|
||||
with:
|
||||
path: |
|
||||
~/.cargo/registry
|
||||
~/.cargo/git
|
||||
target
|
||||
key: ${{ runner.os }}-cargo-${{ hashFiles('**/Cargo.lock') }}
|
||||
restore-keys: |
|
||||
${{ runner.os }}-cargo-
|
||||
|
||||
- name: Install cargo-nextest
|
||||
uses: taiki-e/install-action@nextest
|
||||
|
||||
- name: Run unit tests
|
||||
run: cargo nextest run --all --profile ci
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Job 2c – Unit tests (Windows)
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
test-windows:
|
||||
name: Unit tests (Windows)
|
||||
runs-on: windows-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
|
||||
- name: Install Rust toolchain
|
||||
uses: dtolnay/rust-toolchain@stable
|
||||
|
||||
- name: Cache Cargo registry + build
|
||||
uses: actions/cache@v4
|
||||
with:
|
||||
path: |
|
||||
~/.cargo/registry
|
||||
~/.cargo/git
|
||||
target
|
||||
key: ${{ runner.os }}-cargo-${{ hashFiles('**/Cargo.lock') }}
|
||||
restore-keys: |
|
||||
${{ runner.os }}-cargo-
|
||||
|
||||
- name: Install cargo-nextest
|
||||
uses: taiki-e/install-action@nextest
|
||||
|
||||
- name: Run unit tests
|
||||
run: cargo nextest run --all --profile ci
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Job 2d – PowerShell lint (Windows packaging scripts)
|
||||
#
|
||||
# Runs PSScriptAnalyzer against the operator-facing installer/build
|
||||
# scripts shipped in the Windows ZIP package. Settings live in
|
||||
# packaging/windows/PSScriptAnalyzerSettings.psd1 (each suppressed rule
|
||||
# is documented there). Pre-installed on windows-latest runners; no
|
||||
# Install-Module step needed.
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
windows-lint:
|
||||
name: PowerShell lint (Windows packaging)
|
||||
runs-on: windows-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
|
||||
- name: Run PSScriptAnalyzer
|
||||
shell: pwsh
|
||||
run: |
|
||||
$results = Invoke-ScriptAnalyzer `
|
||||
-Path packaging/windows/*.ps1 `
|
||||
-Settings packaging/windows/PSScriptAnalyzerSettings.psd1
|
||||
if ($results) {
|
||||
$results | Format-Table -AutoSize
|
||||
Write-Error "PSScriptAnalyzer found $($results.Count) issue(s)"
|
||||
exit 1
|
||||
} else {
|
||||
Write-Host "PSScriptAnalyzer: no issues"
|
||||
}
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Job 3 – Integration tests (static mesh + chaos simulation)
|
||||
#
|
||||
# Runs only when both build and test succeed. Each topology / scenario is a
|
||||
# separate matrix entry so they run in parallel.
|
||||
#
|
||||
# Static topologies → build Docker images, start containers, ping-test
|
||||
# Chaos scenarios → build sim image, run stochastic simulation
|
||||
# All harnesses share a single Docker image (fips-test:latest) built once
|
||||
# in the setup step from testing/docker/.
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
integration:
|
||||
name: Integration (${{ matrix.suite }})
|
||||
@@ -149,28 +313,44 @@ jobs:
|
||||
- suite: rekey
|
||||
type: rekey
|
||||
topology: rekey
|
||||
- suite: rekey-accept-off
|
||||
type: rekey-accept-off
|
||||
topology: rekey-accept-off
|
||||
- suite: rekey-outbound-only
|
||||
type: rekey-outbound-only
|
||||
topology: rekey-outbound-only
|
||||
- suite: acl-allowlist
|
||||
type: acl-allowlist
|
||||
# ── Firewall baseline (fips0 nftables default-deny) ────────────
|
||||
- suite: firewall
|
||||
type: firewall
|
||||
# ── Outbound LAN gateway integration test ──────────────────────
|
||||
- suite: gateway
|
||||
type: gateway
|
||||
topology: gateway
|
||||
# ── Chaos / stochastic scenarios ───────────────────────────────────
|
||||
- suite: chaos-smoke-10
|
||||
type: chaos
|
||||
scenario: smoke-10
|
||||
- suite: chaos-10
|
||||
- suite: churn-mixed-10
|
||||
type: chaos
|
||||
scenario: chaos-10
|
||||
scenario: churn-mixed
|
||||
chaos_flags: "--nodes 10 --duration 120"
|
||||
- suite: ethernet-mesh
|
||||
type: chaos
|
||||
scenario: ethernet-mesh
|
||||
- suite: ethernet-only
|
||||
type: chaos
|
||||
scenario: ethernet-only
|
||||
- suite: tcp-mesh
|
||||
type: chaos
|
||||
scenario: tcp-mesh
|
||||
- suite: bottleneck-parent
|
||||
type: chaos
|
||||
scenario: bottleneck-parent
|
||||
- suite: cost-avoidance
|
||||
type: chaos
|
||||
scenario: cost-avoidance
|
||||
- suite: cost-mixed-7node
|
||||
type: chaos
|
||||
scenario: cost-mixed-7node
|
||||
- suite: cost-reeval
|
||||
type: chaos
|
||||
scenario: cost-reeval
|
||||
@@ -183,9 +363,67 @@ jobs:
|
||||
- suite: mixed-technology
|
||||
type: chaos
|
||||
scenario: mixed-technology
|
||||
- suite: congestion-stress
|
||||
type: chaos
|
||||
scenario: congestion-stress
|
||||
- suite: bloom-storm
|
||||
type: chaos
|
||||
scenario: bloom-storm
|
||||
# ── Sidecar deployment ──────────────────────────────────────────
|
||||
- suite: sidecar
|
||||
type: sidecar
|
||||
# ── NAT traversal lab (Nostr/STUN UDP hole punch) ───────────────
|
||||
- suite: nat-cone
|
||||
type: nat
|
||||
scenario: cone
|
||||
- suite: nat-symmetric
|
||||
type: nat
|
||||
scenario: symmetric
|
||||
- suite: nat-lan
|
||||
type: nat
|
||||
scenario: lan
|
||||
# ── Nostr overlay advert publish/consume round-trip ─────────────
|
||||
# Two FIPS daemons + the existing strfry relay; covers Phase 1
|
||||
# (A→B publish/consume), Phase 2 (B→A reverse), and Phase 3
|
||||
# (malformed advert injected to relay; consumers must reject
|
||||
# without crashing). UDP transport baseline for v0.3.0.
|
||||
- suite: nostr-publish-consume
|
||||
type: nostr-publish-consume
|
||||
# ── STUN fault-injection ───────────────────────────────────────
|
||||
# One FIPS daemon + a netns-sharing shim that injects tc/iptables
|
||||
# faults against UDP egress to the in-lab STUN server. Three
|
||||
# phases: 100% drop, ~5s delay then clear, then full STUN
|
||||
# container kill. Asserts the daemon notices each fault,
|
||||
# recovers from delay, and never panics.
|
||||
- suite: stun-faults
|
||||
type: stun-faults
|
||||
# ── Real-deb install across target distros ─────────────────────
|
||||
# Boots a privileged systemd container per distro, runs
|
||||
# `apt install ./fips_*.deb` with the locally-built package,
|
||||
# then asserts end-to-end `.fips` resolution + the
|
||||
# gateway/daemon default-pairing. The most thorough single
|
||||
# test surface — exercises packaging, maintainer scripts,
|
||||
# systemd unit ordering, real TUN, and the DNS responder
|
||||
# filter on a per-distro resolver backend.
|
||||
- suite: deb-install-debian12
|
||||
type: deb-install
|
||||
scenario: debian12
|
||||
- suite: deb-install-ubuntu24
|
||||
type: deb-install
|
||||
scenario: ubuntu24
|
||||
- suite: deb-install-ubuntu26
|
||||
type: deb-install
|
||||
scenario: ubuntu26
|
||||
# ── DNS resolver multi-backend coverage ────────────────────────
|
||||
# Exercises every fips-dns-setup backend (resolved, dnsmasq,
|
||||
# NM+dnsmasq, dns-delegate, no-resolver) across five distros,
|
||||
# plus end-to-end scenarios that boot a real fips daemon with a
|
||||
# real TUN and assert `dig @127.0.0.53 AAAA <npub>.fips`
|
||||
# returns AAAA. Pins the production DNS bind path that
|
||||
# ISSUE-2026-0002 lived in. Single matrix entry runs all 13
|
||||
# scenarios sequentially; ~7-12 min warm, ~12-15 min cold.
|
||||
- suite: dns-resolver
|
||||
type: dns-resolver
|
||||
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
@@ -197,24 +435,24 @@ jobs:
|
||||
name: fips-linux
|
||||
path: _bin
|
||||
|
||||
# ── Static topology ────────────────────────────────────────────────────
|
||||
- name: Install binary (static)
|
||||
if: matrix.type == 'static'
|
||||
# Install binaries to unified docker context and build shared image
|
||||
- name: Install binaries and build Docker image
|
||||
run: |
|
||||
chmod +x _bin/fips _bin/fipsctl
|
||||
cp _bin/fips testing/static/fips
|
||||
cp _bin/fipsctl testing/static/fipsctl
|
||||
[ -f _bin/fipstop ] && chmod +x _bin/fipstop || true
|
||||
[ -f _bin/fips-gateway ] && chmod +x _bin/fips-gateway || true
|
||||
cp _bin/fips testing/docker/fips
|
||||
cp _bin/fipsctl testing/docker/fipsctl
|
||||
[ -f _bin/fipstop ] && cp _bin/fipstop testing/docker/fipstop || true
|
||||
[ -f _bin/fips-gateway ] && cp _bin/fips-gateway testing/docker/fips-gateway || true
|
||||
docker build -t fips-test:latest testing/docker
|
||||
docker build -t fips-test-app:latest -f testing/docker/Dockerfile.app testing/docker
|
||||
|
||||
# ── Static topology ────────────────────────────────────────────────────
|
||||
- name: Generate configs (static)
|
||||
if: matrix.type == 'static'
|
||||
run: bash testing/static/scripts/generate-configs.sh ${{ matrix.topology }}
|
||||
|
||||
- name: Build Docker images (static)
|
||||
if: matrix.type == 'static'
|
||||
run: |
|
||||
docker compose -f testing/static/docker-compose.yml \
|
||||
--profile ${{ matrix.topology }} build
|
||||
|
||||
- name: Start containers (static)
|
||||
if: matrix.type == 'static'
|
||||
run: |
|
||||
@@ -238,25 +476,12 @@ jobs:
|
||||
--profile ${{ matrix.topology }} down --volumes --remove-orphans
|
||||
|
||||
# ── Rekey integration test ──────────────────────────────────────────────
|
||||
- name: Install binary (rekey)
|
||||
if: matrix.type == 'rekey'
|
||||
run: |
|
||||
chmod +x _bin/fips _bin/fipsctl
|
||||
cp _bin/fips testing/static/fips
|
||||
cp _bin/fipsctl testing/static/fipsctl
|
||||
|
||||
- name: Generate and inject configs (rekey)
|
||||
if: matrix.type == 'rekey'
|
||||
run: |
|
||||
bash testing/static/scripts/generate-configs.sh rekey
|
||||
bash testing/static/scripts/rekey-test.sh inject-config
|
||||
|
||||
- name: Build Docker images (rekey)
|
||||
if: matrix.type == 'rekey'
|
||||
run: |
|
||||
docker compose -f testing/static/docker-compose.yml \
|
||||
--profile rekey build
|
||||
|
||||
- name: Start containers (rekey)
|
||||
if: matrix.type == 'rekey'
|
||||
run: |
|
||||
@@ -279,25 +504,115 @@ jobs:
|
||||
docker compose -f testing/static/docker-compose.yml \
|
||||
--profile rekey down --volumes --remove-orphans
|
||||
|
||||
# ── Rekey + accept_connections=false variant ──────────────────────────
|
||||
- name: Generate and inject configs (rekey-accept-off)
|
||||
if: matrix.type == 'rekey-accept-off'
|
||||
env:
|
||||
REKEY_TOPOLOGY: rekey-accept-off
|
||||
REKEY_ACCEPT_OFF_NODES: b
|
||||
run: |
|
||||
bash testing/static/scripts/generate-configs.sh rekey-accept-off
|
||||
bash testing/static/scripts/rekey-test.sh inject-config
|
||||
|
||||
- name: Start containers (rekey-accept-off)
|
||||
if: matrix.type == 'rekey-accept-off'
|
||||
run: |
|
||||
docker compose -f testing/static/docker-compose.yml \
|
||||
--profile rekey-accept-off up -d
|
||||
|
||||
- name: Run rekey test (accept-off variant)
|
||||
if: matrix.type == 'rekey-accept-off'
|
||||
env:
|
||||
REKEY_TOPOLOGY: rekey-accept-off
|
||||
REKEY_ACCEPT_OFF_NODES: b
|
||||
run: bash testing/static/scripts/rekey-test.sh
|
||||
|
||||
- name: Collect logs on failure (rekey-accept-off)
|
||||
if: matrix.type == 'rekey-accept-off' && failure()
|
||||
run: |
|
||||
docker compose -f testing/static/docker-compose.yml \
|
||||
--profile rekey-accept-off logs --no-color | tail -300
|
||||
|
||||
- name: Stop containers (rekey-accept-off)
|
||||
if: matrix.type == 'rekey-accept-off' && always()
|
||||
run: |
|
||||
docker compose -f testing/static/docker-compose.yml \
|
||||
--profile rekey-accept-off down --volumes --remove-orphans
|
||||
|
||||
# ── Rekey + udp.outbound_only=true variant ─────────────────────────────
|
||||
- name: Generate and inject configs (rekey-outbound-only)
|
||||
if: matrix.type == 'rekey-outbound-only'
|
||||
env:
|
||||
REKEY_TOPOLOGY: rekey-outbound-only
|
||||
REKEY_OUTBOUND_ONLY_NODES: b
|
||||
run: |
|
||||
bash testing/static/scripts/generate-configs.sh rekey-outbound-only
|
||||
bash testing/static/scripts/rekey-test.sh inject-config
|
||||
|
||||
- name: Start containers (rekey-outbound-only)
|
||||
if: matrix.type == 'rekey-outbound-only'
|
||||
run: |
|
||||
docker compose -f testing/static/docker-compose.yml \
|
||||
--profile rekey-outbound-only up -d
|
||||
|
||||
- name: Run rekey test (outbound-only variant)
|
||||
if: matrix.type == 'rekey-outbound-only'
|
||||
env:
|
||||
REKEY_TOPOLOGY: rekey-outbound-only
|
||||
REKEY_OUTBOUND_ONLY_NODES: b
|
||||
run: bash testing/static/scripts/rekey-test.sh
|
||||
|
||||
- name: Collect logs on failure (rekey-outbound-only)
|
||||
if: matrix.type == 'rekey-outbound-only' && failure()
|
||||
run: |
|
||||
docker compose -f testing/static/docker-compose.yml \
|
||||
--profile rekey-outbound-only logs --no-color | tail -300
|
||||
|
||||
- name: Stop containers (rekey-outbound-only)
|
||||
if: matrix.type == 'rekey-outbound-only' && always()
|
||||
run: |
|
||||
docker compose -f testing/static/docker-compose.yml \
|
||||
--profile rekey-outbound-only down --volumes --remove-orphans
|
||||
|
||||
# ── ACL allowlist integration test ─────────────────────────────────────
|
||||
- name: Run ACL allowlist integration test
|
||||
if: matrix.type == 'acl-allowlist'
|
||||
run: bash testing/acl-allowlist/test.sh --skip-build --keep-up
|
||||
|
||||
- name: Collect logs on failure (acl-allowlist)
|
||||
if: matrix.type == 'acl-allowlist' && failure()
|
||||
run: |
|
||||
docker compose -f testing/acl-allowlist/docker-compose.yml logs --no-color
|
||||
|
||||
- name: Stop containers (acl-allowlist)
|
||||
if: matrix.type == 'acl-allowlist' && always()
|
||||
run: |
|
||||
docker compose -f testing/acl-allowlist/docker-compose.yml down --volumes --remove-orphans
|
||||
|
||||
# ── Firewall baseline integration test ─────────────────────────────────
|
||||
- name: Run firewall baseline integration test
|
||||
if: matrix.type == 'firewall'
|
||||
run: bash testing/firewall/test.sh --skip-build --keep-up
|
||||
|
||||
- name: Collect logs on failure (firewall)
|
||||
if: matrix.type == 'firewall' && failure()
|
||||
run: |
|
||||
docker compose -f testing/firewall/docker-compose.yml logs --no-color
|
||||
docker exec fips-fw-container-b nft list table inet fips || true
|
||||
|
||||
- name: Stop containers (firewall)
|
||||
if: matrix.type == 'firewall' && always()
|
||||
run: |
|
||||
docker compose -f testing/firewall/docker-compose.yml down --volumes --remove-orphans
|
||||
|
||||
# ── Chaos simulation ───────────────────────────────────────────────────
|
||||
- name: Install Python deps (chaos)
|
||||
if: matrix.type == 'chaos'
|
||||
run: pip3 install --quiet pyyaml jinja2
|
||||
|
||||
- name: Install binary (chaos)
|
||||
if: matrix.type == 'chaos'
|
||||
run: |
|
||||
chmod +x _bin/fips _bin/fipsctl
|
||||
cp _bin/fips testing/chaos/fips
|
||||
cp _bin/fipsctl testing/chaos/fipsctl
|
||||
|
||||
- name: Build chaos Docker image
|
||||
if: matrix.type == 'chaos'
|
||||
run: docker build -t fips-chaos:latest testing/chaos
|
||||
|
||||
- name: Run chaos scenario
|
||||
if: matrix.type == 'chaos'
|
||||
run: bash testing/chaos/scripts/chaos.sh ${{ matrix.scenario }}
|
||||
run: bash testing/chaos/scripts/chaos.sh ${{ matrix.scenario }} ${{ matrix.chaos_flags }}
|
||||
|
||||
- name: Upload sim results on failure (chaos)
|
||||
if: matrix.type == 'chaos' && failure()
|
||||
@@ -308,17 +623,9 @@ jobs:
|
||||
retention-days: 7
|
||||
|
||||
# ── Sidecar deployment ──────────────────────────────────────────────
|
||||
- name: Install binary (sidecar)
|
||||
if: matrix.type == 'sidecar'
|
||||
run: |
|
||||
chmod +x _bin/fips _bin/fipsctl _bin/fipstop
|
||||
cp _bin/fips testing/sidecar/fips
|
||||
cp _bin/fipsctl testing/sidecar/fipsctl
|
||||
cp _bin/fipstop testing/sidecar/fipstop
|
||||
|
||||
- name: Run sidecar integration test
|
||||
if: matrix.type == 'sidecar'
|
||||
run: bash testing/sidecar/scripts/test-sidecar.sh
|
||||
run: bash testing/sidecar/scripts/test-sidecar.sh --skip-build
|
||||
|
||||
- name: Collect logs on failure (sidecar)
|
||||
if: matrix.type == 'sidecar' && failure()
|
||||
@@ -327,4 +634,145 @@ jobs:
|
||||
echo "--- sidecar-${node} logs ---"
|
||||
docker logs "sidecar-${node}-fips-1" 2>&1 || true
|
||||
echo ""
|
||||
done
|
||||
done
|
||||
|
||||
# ── NAT traversal lab ───────────────────────────────────────────────
|
||||
- name: Run NAT lab scenario
|
||||
if: matrix.type == 'nat'
|
||||
run: bash testing/nat/scripts/nat-test.sh ${{ matrix.scenario }}
|
||||
|
||||
- name: Collect logs on failure (nat)
|
||||
if: matrix.type == 'nat' && failure()
|
||||
run: |
|
||||
docker compose -f testing/nat/docker-compose.yml \
|
||||
--profile ${{ matrix.scenario }} logs --no-color
|
||||
|
||||
- name: Stop containers (nat)
|
||||
if: matrix.type == 'nat' && always()
|
||||
run: |
|
||||
docker compose -f testing/nat/docker-compose.yml \
|
||||
--profile cone --profile symmetric --profile lan \
|
||||
down --volumes --remove-orphans
|
||||
|
||||
# ── Nostr overlay advert publish/consume ───────────────────────────
|
||||
- name: Run Nostr publish/consume test
|
||||
if: matrix.type == 'nostr-publish-consume'
|
||||
run: bash testing/nat/scripts/nostr-relay-test.sh
|
||||
|
||||
- name: Collect logs on failure (nostr-publish-consume)
|
||||
if: matrix.type == 'nostr-publish-consume' && failure()
|
||||
run: |
|
||||
docker compose -f testing/nat/docker-compose.yml \
|
||||
--profile nostr-publish-consume logs --no-color | tail -300
|
||||
|
||||
- name: Stop containers (nostr-publish-consume)
|
||||
if: matrix.type == 'nostr-publish-consume' && always()
|
||||
run: |
|
||||
docker compose -f testing/nat/docker-compose.yml \
|
||||
--profile nostr-publish-consume down --volumes --remove-orphans
|
||||
|
||||
# ── STUN fault-injection ───────────────────────────────────────────
|
||||
- name: Run STUN fault-injection test
|
||||
if: matrix.type == 'stun-faults'
|
||||
run: bash testing/nat/scripts/stun-faults-test.sh
|
||||
|
||||
- name: Collect logs on failure (stun-faults)
|
||||
if: matrix.type == 'stun-faults' && failure()
|
||||
run: |
|
||||
docker compose -f testing/nat/docker-compose.yml \
|
||||
--profile stun-faults logs --no-color | tail -300
|
||||
|
||||
- name: Stop containers (stun-faults)
|
||||
if: matrix.type == 'stun-faults' && always()
|
||||
run: |
|
||||
docker compose -f testing/nat/docker-compose.yml \
|
||||
--profile stun-faults down --volumes --remove-orphans
|
||||
|
||||
# ── Outbound LAN gateway integration test ──────────────────────────
|
||||
- name: Generate configs (gateway)
|
||||
if: matrix.type == 'gateway'
|
||||
run: bash testing/static/scripts/generate-configs.sh gateway gateway-test
|
||||
|
||||
- name: Inject gateway config (gateway)
|
||||
if: matrix.type == 'gateway'
|
||||
run: bash testing/static/scripts/gateway-test.sh inject-config
|
||||
|
||||
- name: Start containers (gateway)
|
||||
if: matrix.type == 'gateway'
|
||||
run: |
|
||||
docker compose -f testing/static/docker-compose.yml \
|
||||
--profile gateway up -d
|
||||
|
||||
- name: Run gateway test
|
||||
if: matrix.type == 'gateway'
|
||||
run: bash testing/static/scripts/gateway-test.sh
|
||||
|
||||
- name: Collect logs on failure (gateway)
|
||||
if: matrix.type == 'gateway' && failure()
|
||||
run: |
|
||||
docker compose -f testing/static/docker-compose.yml \
|
||||
--profile gateway logs --no-color | tail -300
|
||||
|
||||
- name: Stop containers (gateway)
|
||||
if: matrix.type == 'gateway' && always()
|
||||
run: |
|
||||
docker compose -f testing/static/docker-compose.yml \
|
||||
--profile gateway down --volumes --remove-orphans
|
||||
|
||||
# ── Real-deb install integration ────────────────────────────────────
|
||||
# The deb-install harness builds its own .deb from source in a
|
||||
# cargo-deb builder image; the pre-built Linux binary from the
|
||||
# build job is intentionally not used here so the test exercises
|
||||
# the full packaging pipeline. ~5-7 min cold-cache on a fresh
|
||||
# runner (.deb build dominates), ~1-2 min warm-cache.
|
||||
- name: Run deb-install scenario
|
||||
if: matrix.type == 'deb-install'
|
||||
timeout-minutes: 25
|
||||
run: bash testing/deb-install/test.sh ${{ matrix.scenario }}
|
||||
|
||||
- name: Collect logs on failure (deb-install)
|
||||
if: matrix.type == 'deb-install' && failure()
|
||||
run: |
|
||||
docker ps -a --filter "name=fips-deb-test-${{ matrix.scenario }}" --format '{{.Names}}' | while read -r c; do
|
||||
echo "--- ${c} fips.service ---"
|
||||
docker exec "$c" journalctl -u fips.service --no-pager 2>&1 | tail -100 || true
|
||||
echo "--- ${c} fips-dns.service ---"
|
||||
docker exec "$c" journalctl -u fips-dns.service --no-pager 2>&1 | tail -100 || true
|
||||
echo "--- ${c} fips-gateway.service ---"
|
||||
docker exec "$c" journalctl -u fips-gateway.service --no-pager 2>&1 | tail -100 || true
|
||||
done
|
||||
|
||||
- name: Stop containers (deb-install)
|
||||
if: matrix.type == 'deb-install' && always()
|
||||
run: |
|
||||
docker ps -a --filter "name=fips-deb-test-${{ matrix.scenario }}" --format '{{.Names}}' | while read -r c; do
|
||||
docker rm -f "$c" >/dev/null 2>&1 || true
|
||||
done
|
||||
|
||||
# ── DNS resolver multi-backend integration ──────────────────────────
|
||||
# The dns-resolver harness builds its own fips binary from source in a
|
||||
# Debian 12 builder image (shared cache layout with deb-install). Runs
|
||||
# all 13 scenarios in a single job: dummy-TUN backend-detection tests
|
||||
# plus real-fips end-to-end queries through systemd-resolved across
|
||||
# five distros. ~7-12 min warm, ~12-15 min cold.
|
||||
- name: Run dns-resolver test
|
||||
if: matrix.type == 'dns-resolver'
|
||||
timeout-minutes: 30
|
||||
run: bash testing/dns-resolver/test.sh
|
||||
|
||||
- name: Collect logs on failure (dns-resolver)
|
||||
if: matrix.type == 'dns-resolver' && failure()
|
||||
run: |
|
||||
docker ps -a --filter "name=fips-dns-test-" --format '{{.Names}}' | while read -r c; do
|
||||
echo "--- ${c} fips.service ---"
|
||||
docker exec "$c" journalctl -u fips.service --no-pager 2>&1 | tail -100 || true
|
||||
echo "--- ${c} fips-dns.service ---"
|
||||
docker exec "$c" journalctl -u fips-dns.service --no-pager 2>&1 | tail -100 || true
|
||||
done
|
||||
|
||||
- name: Stop containers (dns-resolver)
|
||||
if: matrix.type == 'dns-resolver' && always()
|
||||
run: |
|
||||
docker ps -a --filter "name=fips-dns-test-" --format '{{.Names}}' | while read -r c; do
|
||||
docker rm -f "$c" >/dev/null 2>&1 || true
|
||||
done
|
||||
|
||||
@@ -0,0 +1,201 @@
|
||||
name: Linux Package
|
||||
on:
|
||||
push:
|
||||
branches:
|
||||
- master
|
||||
- maint
|
||||
- next
|
||||
tags:
|
||||
- "v*"
|
||||
pull_request:
|
||||
workflow_dispatch:
|
||||
|
||||
env:
|
||||
CARGO_TERM_COLOR: always
|
||||
|
||||
jobs:
|
||||
determine-versioning:
|
||||
runs-on: ubuntu-latest
|
||||
outputs:
|
||||
linux_package_version: ${{ steps.linux_version.outputs.linux_package_version }}
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
with:
|
||||
fetch-depth: 0
|
||||
|
||||
- name: Derive Linux package version
|
||||
id: linux_version
|
||||
shell: bash
|
||||
run: |
|
||||
: ${GITHUB_OUTPUT:=/tmp/github_output}
|
||||
|
||||
BASE_VERSION=$(grep '^version' Cargo.toml | head -1 | sed 's/.*"\(.*\)"/\1/')
|
||||
if [[ "$GITHUB_REF" == refs/tags/* ]]; then
|
||||
VERSION="${GITHUB_REF_NAME#v}"
|
||||
else
|
||||
BRANCH=$(echo "$GITHUB_REF_NAME" | sed 's|[^A-Za-z0-9]|.|g; s/\.\.+/./g; s/^\.//; s/\.$//')
|
||||
HEIGHT=$(git rev-list --count HEAD)
|
||||
HASH=$(git rev-parse --short HEAD)
|
||||
if [[ -z "$BRANCH" ]]; then
|
||||
BRANCH="ref"
|
||||
fi
|
||||
VERSION="${BASE_VERSION}+${BRANCH}.${HEIGHT}.${HASH}"
|
||||
fi
|
||||
|
||||
echo "linux_package_version=${VERSION}" >> "$GITHUB_OUTPUT"
|
||||
|
||||
build:
|
||||
name: Build Linux artifacts (${{ matrix.artifact_arch }})
|
||||
runs-on: ${{ matrix.os }}
|
||||
needs: determine-versioning
|
||||
|
||||
strategy:
|
||||
fail-fast: false
|
||||
matrix:
|
||||
include:
|
||||
- os: ubuntu-latest
|
||||
artifact_arch: x86_64
|
||||
deb_arch: amd64
|
||||
- os: ubuntu-24.04-arm
|
||||
artifact_arch: aarch64
|
||||
deb_arch: arm64
|
||||
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
with:
|
||||
fetch-depth: 0
|
||||
|
||||
- name: Set SOURCE_DATE_EPOCH from git
|
||||
run: echo "SOURCE_DATE_EPOCH=$(git log -1 --format=%ct)" >> "$GITHUB_ENV"
|
||||
|
||||
- name: Install system dependencies
|
||||
run: sudo apt-get update && sudo apt-get install -y --no-install-recommends libdbus-1-dev llvm
|
||||
|
||||
- name: Install Rust toolchain
|
||||
uses: dtolnay/rust-toolchain@stable
|
||||
|
||||
- name: Cache Cargo registry + build
|
||||
if: ${{ env.ACT != 'true' }}
|
||||
uses: actions/cache@v4
|
||||
with:
|
||||
path: |
|
||||
~/.cargo/registry
|
||||
~/.cargo/git
|
||||
target
|
||||
key: linux-release-${{ runner.os }}-${{ matrix.artifact_arch }}-${{ hashFiles('**/Cargo.lock') }}
|
||||
restore-keys: |
|
||||
linux-release-${{ runner.os }}-${{ matrix.artifact_arch }}-
|
||||
|
||||
- name: Install cargo-deb
|
||||
run: cargo install cargo-deb --version 3.6.3 --locked
|
||||
|
||||
- name: Build release binaries
|
||||
run: cargo build --release
|
||||
|
||||
- name: Build systemd tarball
|
||||
env:
|
||||
STRIP: llvm-strip
|
||||
run: |
|
||||
packaging/systemd/build-tarball.sh \
|
||||
--version "${{ needs.determine-versioning.outputs.linux_package_version }}" \
|
||||
--arch "${{ matrix.artifact_arch }}" \
|
||||
--no-build
|
||||
|
||||
- name: Build Debian package
|
||||
run: |
|
||||
packaging/debian/build-deb.sh \
|
||||
--version "${{ needs.determine-versioning.outputs.linux_package_version }}" \
|
||||
--no-build
|
||||
|
||||
- name: Resolve Linux asset paths
|
||||
id: linux-assets
|
||||
shell: bash
|
||||
run: |
|
||||
: ${GITHUB_OUTPUT:=/tmp/github_output}
|
||||
|
||||
TARBALL="deploy/fips-${{ needs.determine-versioning.outputs.linux_package_version }}-linux-${{ matrix.artifact_arch }}.tar.gz"
|
||||
if [[ ! -f "$TARBALL" ]]; then
|
||||
echo "Missing tarball: $TARBALL" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
DEB_FILE=$(find deploy -maxdepth 1 -type f -name "fips_*_${{ matrix.deb_arch }}.deb" | sort | head -n 1)
|
||||
if [[ -z "$DEB_FILE" ]]; then
|
||||
echo "Missing Debian package for ${{ matrix.deb_arch }}" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo "tarball=$TARBALL" >> "$GITHUB_OUTPUT"
|
||||
echo "deb=$DEB_FILE" >> "$GITHUB_OUTPUT"
|
||||
|
||||
- name: SHA-256 hashes
|
||||
run: |
|
||||
echo "==> Linux release assets:"
|
||||
sha256sum \
|
||||
"${{ steps.linux-assets.outputs.tarball }}" \
|
||||
"${{ steps.linux-assets.outputs.deb }}"
|
||||
|
||||
- name: Upload artifact (GitHub only)
|
||||
if: ${{ env.ACT != 'true' }}
|
||||
uses: actions/upload-artifact@v4
|
||||
with:
|
||||
name: fips_${{ needs.determine-versioning.outputs.linux_package_version }}_${{ matrix.artifact_arch }}_linux
|
||||
path: |
|
||||
${{ steps.linux-assets.outputs.tarball }}
|
||||
${{ steps.linux-assets.outputs.deb }}
|
||||
retention-days: 30
|
||||
|
||||
- name: Build Summary
|
||||
run: |
|
||||
echo "Build Summary for linux/${{ matrix.artifact_arch }}:"
|
||||
echo " Tarball: ${{ steps.linux-assets.outputs.tarball }}"
|
||||
echo " Debian: ${{ steps.linux-assets.outputs.deb }}"
|
||||
|
||||
release:
|
||||
name: Publish Linux assets to GitHub Release
|
||||
runs-on: ubuntu-latest
|
||||
needs: build
|
||||
if: startsWith(github.ref, 'refs/tags/')
|
||||
permissions:
|
||||
contents: write
|
||||
|
||||
steps:
|
||||
- name: Download Linux artifacts
|
||||
uses: actions/download-artifact@v4
|
||||
with:
|
||||
path: dist
|
||||
merge-multiple: true
|
||||
|
||||
- name: Generate Linux release checksums
|
||||
run: |
|
||||
cd dist
|
||||
find . -maxdepth 1 -type f \( -name '*.deb' -o -name '*.tar.gz' \) -printf '%P\n' \
|
||||
| LC_ALL=C sort \
|
||||
| xargs sha256sum \
|
||||
> checksums-linux.txt
|
||||
|
||||
- name: Wait for tag release
|
||||
env:
|
||||
GH_TOKEN: ${{ github.token }}
|
||||
run: |
|
||||
for attempt in $(seq 1 20); do
|
||||
if gh release view "${GITHUB_REF_NAME}" --repo "${GITHUB_REPOSITORY}" >/dev/null 2>&1; then
|
||||
exit 0
|
||||
fi
|
||||
echo "Release ${GITHUB_REF_NAME} not available yet; waiting..."
|
||||
sleep 15
|
||||
done
|
||||
|
||||
echo "Timed out waiting for release ${GITHUB_REF_NAME}" >&2
|
||||
exit 1
|
||||
|
||||
- name: Upload Linux assets
|
||||
env:
|
||||
GH_TOKEN: ${{ github.token }}
|
||||
run: |
|
||||
gh release upload "${GITHUB_REF_NAME}" \
|
||||
dist/*.deb \
|
||||
dist/*.tar.gz \
|
||||
dist/checksums-linux.txt \
|
||||
--clobber \
|
||||
--repo "${GITHUB_REPOSITORY}"
|
||||
@@ -0,0 +1,251 @@
|
||||
name: macOS Package
|
||||
on:
|
||||
push:
|
||||
branches:
|
||||
- master
|
||||
- maint
|
||||
- next
|
||||
tags:
|
||||
- "v*"
|
||||
pull_request:
|
||||
workflow_dispatch:
|
||||
|
||||
env:
|
||||
CARGO_TERM_COLOR: always
|
||||
|
||||
jobs:
|
||||
determine-versioning:
|
||||
runs-on: macos-latest
|
||||
outputs:
|
||||
macos_package_version: ${{ steps.macos_version.outputs.macos_package_version }}
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
with:
|
||||
fetch-depth: 0
|
||||
|
||||
- name: Derive macOS package version
|
||||
id: macos_version
|
||||
shell: bash
|
||||
run: |
|
||||
: ${GITHUB_OUTPUT:=/tmp/github_output}
|
||||
|
||||
BASE_VERSION=$(grep '^version' Cargo.toml | head -1 | sed 's/.*"\(.*\)"/\1/')
|
||||
if [[ "$GITHUB_REF" == refs/tags/* ]]; then
|
||||
VERSION="${GITHUB_REF_NAME#v}"
|
||||
else
|
||||
BRANCH=$(echo "$GITHUB_REF_NAME" | sed 's|[^A-Za-z0-9]|.|g; s/\.\.+/./g; s/^\.//; s/\.$//')
|
||||
HEIGHT=$(git rev-list --count HEAD)
|
||||
HASH=$(git rev-parse --short HEAD)
|
||||
if [[ -z "$BRANCH" ]]; then
|
||||
BRANCH="ref"
|
||||
fi
|
||||
VERSION="${BASE_VERSION}+${BRANCH}.${HEIGHT}.${HASH}"
|
||||
fi
|
||||
|
||||
echo "macos_package_version=${VERSION}" >> "$GITHUB_OUTPUT"
|
||||
|
||||
build:
|
||||
name: Build macOS package (${{ matrix.arch }})
|
||||
runs-on: ${{ matrix.os }}
|
||||
needs: determine-versioning
|
||||
|
||||
strategy:
|
||||
fail-fast: false
|
||||
matrix:
|
||||
include:
|
||||
- os: macos-latest
|
||||
arch: arm64
|
||||
target: aarch64-apple-darwin
|
||||
- os: macos-latest
|
||||
arch: x86_64
|
||||
target: x86_64-apple-darwin
|
||||
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
with:
|
||||
fetch-depth: 0
|
||||
|
||||
- name: Set SOURCE_DATE_EPOCH from git
|
||||
run: echo "SOURCE_DATE_EPOCH=$(git log -1 --format=%ct)" >> "$GITHUB_ENV"
|
||||
|
||||
- name: Install Rust toolchain
|
||||
uses: dtolnay/rust-toolchain@stable
|
||||
|
||||
- name: Add cross-compile target
|
||||
run: rustup target add ${{ matrix.target }}
|
||||
|
||||
- name: Cache Cargo registry + build
|
||||
uses: actions/cache@v4
|
||||
with:
|
||||
path: |
|
||||
~/.cargo/registry
|
||||
~/.cargo/git
|
||||
target
|
||||
key: macos-release-${{ matrix.arch }}-${{ hashFiles('**/Cargo.lock') }}
|
||||
restore-keys: |
|
||||
macos-release-${{ matrix.arch }}-
|
||||
|
||||
- name: Build release binaries
|
||||
run: cargo build --release --target ${{ matrix.target }}
|
||||
|
||||
- name: Build macOS package
|
||||
run: |
|
||||
packaging/macos/build-pkg.sh \
|
||||
--version "${{ needs.determine-versioning.outputs.macos_package_version }}" \
|
||||
--target ${{ matrix.target }} \
|
||||
--no-build
|
||||
|
||||
- name: Resolve macOS asset path
|
||||
id: macos-assets
|
||||
shell: bash
|
||||
run: |
|
||||
: ${GITHUB_OUTPUT:=/tmp/github_output}
|
||||
|
||||
PKG_FILE=$(find deploy -maxdepth 1 -type f -name "fips-*-macos-*.pkg" | sort | head -n 1)
|
||||
if [[ -z "$PKG_FILE" ]]; then
|
||||
echo "Missing macOS package" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo "pkg=$PKG_FILE" >> "$GITHUB_OUTPUT"
|
||||
|
||||
- name: Verify .pkg structural correctness
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
PKG="${{ steps.macos-assets.outputs.pkg }}"
|
||||
EXPAND_DIR="$(mktemp -d)/expanded"
|
||||
PAYLOAD_DIR="$(mktemp -d)/payload"
|
||||
fail=0
|
||||
|
||||
echo "==> Verifying $PKG"
|
||||
|
||||
# 1) Flat-package expansion
|
||||
if pkgutil --expand "$PKG" "$EXPAND_DIR"; then
|
||||
echo "PASS: pkgutil --expand"
|
||||
else
|
||||
echo "FAIL: pkgutil --expand"
|
||||
fail=1
|
||||
fi
|
||||
|
||||
# Extract the cpio.gz Payload so we can inspect installed file layout
|
||||
PAYLOAD_FILE="$(find "$EXPAND_DIR" -name Payload -type f | head -n 1)"
|
||||
if [[ -z "$PAYLOAD_FILE" ]]; then
|
||||
echo "FAIL: no Payload file inside expanded pkg"
|
||||
fail=1
|
||||
else
|
||||
mkdir -p "$PAYLOAD_DIR"
|
||||
(cd "$PAYLOAD_DIR" && gzip -dc "$PAYLOAD_FILE" | cpio -i --quiet)
|
||||
echo "PASS: extracted Payload to $PAYLOAD_DIR"
|
||||
fi
|
||||
|
||||
# 2) Binary at canonical install path (./usr/local/bin/fips inside payload)
|
||||
BIN_PATH="$PAYLOAD_DIR/usr/local/bin/fips"
|
||||
if [[ -f "$BIN_PATH" ]]; then
|
||||
echo "PASS: binary present at usr/local/bin/fips"
|
||||
else
|
||||
echo "FAIL: binary missing at usr/local/bin/fips"
|
||||
echo " fallback search:"
|
||||
find "$PAYLOAD_DIR" -name fips -type f -print || true
|
||||
fail=1
|
||||
fi
|
||||
for extra in fipsctl fipstop; do
|
||||
if [[ -f "$PAYLOAD_DIR/usr/local/bin/$extra" ]]; then
|
||||
echo "PASS: binary present at usr/local/bin/$extra"
|
||||
else
|
||||
echo "FAIL: binary missing at usr/local/bin/$extra"
|
||||
fail=1
|
||||
fi
|
||||
done
|
||||
|
||||
# 3) LaunchDaemon plist at canonical location
|
||||
PLIST_PATH="$PAYLOAD_DIR/Library/LaunchDaemons/com.fips.daemon.plist"
|
||||
if [[ -f "$PLIST_PATH" ]]; then
|
||||
echo "PASS: plist present at Library/LaunchDaemons/com.fips.daemon.plist"
|
||||
else
|
||||
echo "FAIL: plist missing at Library/LaunchDaemons/com.fips.daemon.plist"
|
||||
echo " fallback search:"
|
||||
find "$PAYLOAD_DIR" -name '*.plist' -print || true
|
||||
fail=1
|
||||
fi
|
||||
|
||||
# 4) plutil -lint on the plist
|
||||
if [[ -f "$PLIST_PATH" ]]; then
|
||||
if plutil -lint "$PLIST_PATH"; then
|
||||
echo "PASS: plutil -lint"
|
||||
else
|
||||
echo "FAIL: plutil -lint"
|
||||
fail=1
|
||||
fi
|
||||
fi
|
||||
|
||||
if [[ "$fail" -ne 0 ]]; then
|
||||
echo "==> .pkg verification FAILED"
|
||||
exit 1
|
||||
fi
|
||||
echo "==> .pkg verification PASSED"
|
||||
|
||||
- name: SHA-256 hash
|
||||
run: |
|
||||
echo "==> macOS release asset:"
|
||||
shasum -a 256 "${{ steps.macos-assets.outputs.pkg }}"
|
||||
|
||||
- name: Upload artifact
|
||||
uses: actions/upload-artifact@v4
|
||||
with:
|
||||
name: fips_${{ needs.determine-versioning.outputs.macos_package_version }}_${{ matrix.arch }}_macos
|
||||
path: ${{ steps.macos-assets.outputs.pkg }}
|
||||
retention-days: 30
|
||||
|
||||
- name: Build summary
|
||||
run: |
|
||||
echo "Build Summary for macOS/${{ matrix.arch }}:"
|
||||
echo " Package: ${{ steps.macos-assets.outputs.pkg }}"
|
||||
|
||||
release:
|
||||
name: Publish macOS assets to GitHub Release
|
||||
runs-on: ubuntu-latest
|
||||
needs: build
|
||||
if: startsWith(github.ref, 'refs/tags/')
|
||||
permissions:
|
||||
contents: write
|
||||
|
||||
steps:
|
||||
- name: Download macOS artifacts
|
||||
uses: actions/download-artifact@v4
|
||||
with:
|
||||
path: dist
|
||||
merge-multiple: true
|
||||
|
||||
- name: Generate macOS release checksums
|
||||
run: |
|
||||
cd dist
|
||||
find . -maxdepth 1 -type f -name '*.pkg' -printf '%P\n' \
|
||||
| LC_ALL=C sort \
|
||||
| xargs sha256sum \
|
||||
> checksums-macos.txt
|
||||
|
||||
- name: Wait for tag release
|
||||
env:
|
||||
GH_TOKEN: ${{ github.token }}
|
||||
run: |
|
||||
for attempt in $(seq 1 20); do
|
||||
if gh release view "${GITHUB_REF_NAME}" --repo "${GITHUB_REPOSITORY}" >/dev/null 2>&1; then
|
||||
exit 0
|
||||
fi
|
||||
echo "Release ${GITHUB_REF_NAME} not available yet; waiting..."
|
||||
sleep 15
|
||||
done
|
||||
|
||||
echo "Timed out waiting for release ${GITHUB_REF_NAME}" >&2
|
||||
exit 1
|
||||
|
||||
- name: Upload macOS assets
|
||||
env:
|
||||
GH_TOKEN: ${{ github.token }}
|
||||
run: |
|
||||
gh release upload "${GITHUB_REF_NAME}" \
|
||||
dist/*.pkg \
|
||||
dist/checksums-macos.txt \
|
||||
--clobber \
|
||||
--repo "${GITHUB_REPOSITORY}"
|
||||
@@ -1,7 +1,13 @@
|
||||
name: OpenWrt Package
|
||||
|
||||
on:
|
||||
push:
|
||||
branches:
|
||||
- master
|
||||
- maint
|
||||
- next
|
||||
tags:
|
||||
- "v*"
|
||||
pull_request:
|
||||
workflow_dispatch:
|
||||
inputs:
|
||||
arch:
|
||||
@@ -11,11 +17,57 @@ on:
|
||||
|
||||
env:
|
||||
CARGO_TERM_COLOR: always
|
||||
PACKAGE_NAME: "fips"
|
||||
|
||||
jobs:
|
||||
determine-versioning:
|
||||
runs-on: ubuntu-latest
|
||||
outputs:
|
||||
package_version: ${{ steps.version.outputs.package_version }}
|
||||
release_channel: ${{ steps.channel.outputs.release_channel }}
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
with:
|
||||
fetch-depth: 0
|
||||
|
||||
- name: Derive package version
|
||||
id: version
|
||||
shell: bash
|
||||
run: |
|
||||
: ${GITHUB_OUTPUT:=/tmp/github_output}
|
||||
if [[ "$GITHUB_REF" == refs/tags/* ]]; then
|
||||
echo "package_version=${GITHUB_REF_NAME}" >> "$GITHUB_OUTPUT"
|
||||
else
|
||||
BRANCH=$(echo "$GITHUB_REF_NAME" | sed 's|/|-|g')
|
||||
HEIGHT=$(git rev-list --count HEAD)
|
||||
HASH=$(git rev-parse --short HEAD)
|
||||
echo "package_version=${BRANCH}.${HEIGHT}.${HASH}" >> "$GITHUB_OUTPUT"
|
||||
fi
|
||||
|
||||
- name: Determine release channel
|
||||
id: channel
|
||||
shell: bash
|
||||
run: |
|
||||
: ${GITHUB_OUTPUT:=/tmp/github_output}
|
||||
if [[ "$GITHUB_REF" == refs/tags/* ]]; then
|
||||
TAG_NAME=${GITHUB_REF#refs/tags/}
|
||||
if [[ $TAG_NAME =~ ^v[0-9]+\.[0-9]+\.[0-9]+-alpha ]]; then
|
||||
echo "release_channel=alpha" >> "$GITHUB_OUTPUT"
|
||||
elif [[ $TAG_NAME =~ ^v[0-9]+\.[0-9]+\.[0-9]+-beta ]]; then
|
||||
echo "release_channel=beta" >> "$GITHUB_OUTPUT"
|
||||
elif [[ $TAG_NAME =~ ^v[0-9]+\.[0-9]+\.[0-9]+$ ]]; then
|
||||
echo "release_channel=stable" >> "$GITHUB_OUTPUT"
|
||||
else
|
||||
echo "release_channel=dev" >> "$GITHUB_OUTPUT"
|
||||
fi
|
||||
else
|
||||
echo "release_channel=dev" >> "$GITHUB_OUTPUT"
|
||||
fi
|
||||
|
||||
build:
|
||||
name: Build .ipk (${{ matrix.openwrt_arch }})
|
||||
runs-on: ubuntu-latest
|
||||
needs: determine-versioning
|
||||
|
||||
strategy:
|
||||
fail-fast: false
|
||||
@@ -26,7 +78,9 @@ jobs:
|
||||
rust_target: aarch64-unknown-linux-musl
|
||||
rust_channel: stable
|
||||
# MT3000, MT6000, Flint 2, RPi 3/4/5
|
||||
# MIPS disabled: 32-bit MIPS lacks AtomicU64; needs portable-atomic crate
|
||||
# MIPS disabled: nostr-relay-pool 0.44 uses std::sync::atomic::AtomicU64
|
||||
# directly (fips's own atomics already use portable_atomic). Re-enable
|
||||
# once an upstream portable-atomic patch lands (or via [patch.crates-io]).
|
||||
# - build_arch: mipsel
|
||||
# openwrt_arch: mipsel_24kc
|
||||
# rust_target: mipsel-unknown-linux-musl
|
||||
@@ -46,19 +100,13 @@ jobs:
|
||||
with:
|
||||
fetch-depth: 0
|
||||
|
||||
- name: Derive package version
|
||||
id: version
|
||||
- name: Set SOURCE_DATE_EPOCH from git
|
||||
run: echo "SOURCE_DATE_EPOCH=$(git log -1 --format=%ct)" >> "$GITHUB_ENV"
|
||||
|
||||
- name: Initialize
|
||||
run: |
|
||||
if [[ "$GITHUB_REF" == refs/tags/* ]]; then
|
||||
VERSION="${GITHUB_REF_NAME#v}"
|
||||
else
|
||||
BRANCH=$(echo "$GITHUB_REF_NAME" | sed 's|/|-|g')
|
||||
HEIGHT=$(git rev-list --count HEAD)
|
||||
HASH=$(git rev-parse --short HEAD)
|
||||
VERSION="${BRANCH}.${HEIGHT}.${HASH}"
|
||||
fi
|
||||
echo "version=$VERSION" >> "$GITHUB_OUTPUT"
|
||||
echo "filename=fips_${VERSION}_${{ matrix.openwrt_arch }}.ipk" >> "$GITHUB_OUTPUT"
|
||||
PACKAGE_FILENAME=${{ env.PACKAGE_NAME }}_${{ needs.determine-versioning.outputs.package_version }}_${{ matrix.openwrt_arch }}.ipk
|
||||
echo "PACKAGE_FILENAME=$PACKAGE_FILENAME" >> $GITHUB_ENV
|
||||
|
||||
- name: Install Rust toolchain (stable)
|
||||
if: matrix.rust_channel == 'stable'
|
||||
@@ -73,6 +121,7 @@ jobs:
|
||||
components: rust-src
|
||||
|
||||
- name: Cache Cargo registry + build
|
||||
if: ${{ env.ACT != 'true' }}
|
||||
uses: actions/cache@v4
|
||||
with:
|
||||
path: |
|
||||
@@ -84,29 +133,382 @@ jobs:
|
||||
openwrt-${{ matrix.rust_target }}-
|
||||
|
||||
- name: Install cargo-zigbuild
|
||||
run: cargo install cargo-zigbuild --locked
|
||||
run: cargo install cargo-zigbuild --version 0.19.8 --locked
|
||||
|
||||
- name: Install zig (required by cargo-zigbuild)
|
||||
uses: goto-bus-stop/setup-zig@v2
|
||||
run: |
|
||||
ZIG_VERSION="0.13.0"
|
||||
ARCH=$(uname -m)
|
||||
case "$ARCH" in
|
||||
x86_64|amd64) ZIG_ARCH="x86_64" ;;
|
||||
aarch64|arm64) ZIG_ARCH="aarch64" ;;
|
||||
*) echo "Unsupported architecture: $ARCH"; exit 1 ;;
|
||||
esac
|
||||
curl -fsSL "https://ziglang.org/download/${ZIG_VERSION}/zig-linux-${ZIG_ARCH}-${ZIG_VERSION}.tar.xz" | sudo tar xJ -C /opt
|
||||
sudo ln -sf /opt/zig-linux-${ZIG_ARCH}-${ZIG_VERSION}/zig /usr/local/bin/zig
|
||||
zig version
|
||||
|
||||
- name: Install llvm-strip
|
||||
run: sudo apt-get install -y --no-install-recommends llvm
|
||||
run: sudo apt-get update && sudo apt-get install -y --no-install-recommends llvm
|
||||
|
||||
- name: Install nak
|
||||
shell: bash
|
||||
run: |
|
||||
NAK_VERSION="0.16.2"
|
||||
ARCH=$(uname -m)
|
||||
case "$ARCH" in
|
||||
x86_64|amd64) NAK_ARCH="amd64" ;;
|
||||
aarch64|arm64) NAK_ARCH="arm64" ;;
|
||||
*) echo "Unsupported architecture: $ARCH"; exit 1 ;;
|
||||
esac
|
||||
curl -fsSL "https://github.com/fiatjaf/nak/releases/download/v${NAK_VERSION}/nak-v${NAK_VERSION}-linux-${NAK_ARCH}" \
|
||||
-o /usr/local/bin/nak
|
||||
chmod +x /usr/local/bin/nak
|
||||
nak --version
|
||||
|
||||
- name: Install jq
|
||||
run: |
|
||||
if ! command -v jq &>/dev/null; then
|
||||
sudo apt-get update && sudo apt-get install -y jq
|
||||
fi
|
||||
|
||||
# Priority: HIVE_CI_NSEC from env (loom job) > repo secret > generate ephemeral
|
||||
- name: Resolve signing key
|
||||
id: keys
|
||||
shell: bash
|
||||
env:
|
||||
SECRET_NSEC: ${{ secrets.HIVE_CI_NSEC }}
|
||||
run: |
|
||||
: ${GITHUB_OUTPUT:=/tmp/github_output}
|
||||
if [ -n "${HIVE_CI_NSEC:-}" ]; then
|
||||
echo "Using HIVE_CI_NSEC from loom job environment"
|
||||
NSEC="$HIVE_CI_NSEC"
|
||||
elif [ -n "$SECRET_NSEC" ]; then
|
||||
echo "Using HIVE_CI_NSEC from repository secrets"
|
||||
NSEC="$SECRET_NSEC"
|
||||
else
|
||||
echo "No nsec provided -- generating ephemeral keypair"
|
||||
NSEC=$(nak key generate)
|
||||
fi
|
||||
|
||||
PUBKEY=$(echo "$NSEC" | nak key public)
|
||||
echo "::add-mask::$NSEC"
|
||||
echo "nsec=$NSEC" >> "$GITHUB_OUTPUT"
|
||||
echo "pubkey=$PUBKEY" >> "$GITHUB_OUTPUT"
|
||||
echo "Publisher pubkey (hex): $PUBKEY"
|
||||
- name: Build .ipk
|
||||
env:
|
||||
PKG_VERSION: ${{ steps.version.outputs.version }}
|
||||
PKG_VERSION: ${{ needs.determine-versioning.outputs.package_version }}
|
||||
LLVM_STRIP: llvm-strip
|
||||
run: ./packaging/openwrt/build-ipk.sh --arch ${{ matrix.build_arch }}
|
||||
run: ./packaging/openwrt-ipk/build-ipk.sh --arch ${{ matrix.build_arch }}
|
||||
|
||||
- name: Upload artifact
|
||||
- name: Install shellcheck (if missing)
|
||||
shell: bash
|
||||
run: |
|
||||
if ! command -v shellcheck >/dev/null 2>&1; then
|
||||
sudo apt-get update && sudo apt-get install -y --no-install-recommends shellcheck
|
||||
fi
|
||||
shellcheck --version
|
||||
|
||||
- name: Lint shipped shell scripts
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
FILES_DIR=packaging/openwrt-ipk/files
|
||||
# Scripts shipped inside the .ipk. The init scripts use the OpenWrt
|
||||
# `#!/bin/sh /etc/rc.common` shebang; tell shellcheck to treat them
|
||||
# as POSIX sh and silence the unrecognized-shebang warning (SC1008).
|
||||
# SC2317 (unreachable command) fires on rc.common's externally-invoked
|
||||
# start_service/stop_service/reload_service hooks.
|
||||
TARGETS=(
|
||||
"$FILES_DIR/etc/init.d/fips"
|
||||
"$FILES_DIR/etc/init.d/fips-gateway"
|
||||
"$FILES_DIR/etc/fips/firewall.sh"
|
||||
"$FILES_DIR/etc/hotplug.d/net/99-fips"
|
||||
"$FILES_DIR/etc/uci-defaults/90-fips-setup"
|
||||
)
|
||||
fail=0
|
||||
for f in "${TARGETS[@]}"; do
|
||||
if [ ! -f "$f" ]; then
|
||||
echo "FAIL: missing $f"
|
||||
fail=1
|
||||
continue
|
||||
fi
|
||||
echo "==> shellcheck $f"
|
||||
if shellcheck --shell=sh --exclude=SC1008,SC2317,SC2034,SC3043,SC2086,SC2089,SC2090 "$f"; then
|
||||
echo " PASS"
|
||||
else
|
||||
echo " FAIL"
|
||||
fail=1
|
||||
fi
|
||||
done
|
||||
if [ "$fail" -ne 0 ]; then
|
||||
echo "shellcheck FAILED"
|
||||
exit 1
|
||||
fi
|
||||
echo "shellcheck PASS (${#TARGETS[@]} scripts)"
|
||||
|
||||
- name: Sysctl drop-in syntax check
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
FILES_DIR=packaging/openwrt-ipk/files
|
||||
TARGETS=(
|
||||
"$FILES_DIR/etc/sysctl.d/fips-gateway.conf"
|
||||
"$FILES_DIR/etc/sysctl.d/fips-bridge.conf"
|
||||
)
|
||||
fail=0
|
||||
for conf in "${TARGETS[@]}"; do
|
||||
if [ ! -f "$conf" ]; then
|
||||
echo "FAIL: missing $conf"
|
||||
fail=1
|
||||
continue
|
||||
fi
|
||||
echo "==> validating $conf"
|
||||
lineno=0
|
||||
file_ok=1
|
||||
while IFS= read -r line || [ -n "$line" ]; do
|
||||
lineno=$((lineno + 1))
|
||||
# Skip comments and blank lines.
|
||||
case "$line" in
|
||||
''|\#*) continue ;;
|
||||
esac
|
||||
# Match: <key> = <value> where key is dotted lower-id and value is
|
||||
# an integer (sysctl drop-ins shipped here are all numeric toggles).
|
||||
if ! [[ "$line" =~ ^[a-z0-9_.-]+[[:space:]]*=[[:space:]]*-?[0-9]+[[:space:]]*$ ]]; then
|
||||
echo " FAIL line $lineno: $line"
|
||||
file_ok=0
|
||||
fi
|
||||
done < "$conf"
|
||||
if [ "$file_ok" -eq 1 ]; then
|
||||
echo " PASS"
|
||||
else
|
||||
fail=1
|
||||
fi
|
||||
done
|
||||
if [ "$fail" -ne 0 ]; then
|
||||
echo "sysctl drop-in syntax check FAILED"
|
||||
exit 1
|
||||
fi
|
||||
echo "sysctl drop-in syntax check PASS"
|
||||
|
||||
- name: Verify ipk structural integrity
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
IPK="dist/${{ env.PACKAGE_FILENAME }}"
|
||||
if [ ! -f "$IPK" ]; then
|
||||
echo "FAIL: produced ipk not found at $IPK"
|
||||
exit 1
|
||||
fi
|
||||
echo "==> file type:"
|
||||
file "$IPK"
|
||||
|
||||
# OpenWrt .ipk = tar.gz containing debian-binary + control.tar.gz +
|
||||
# data.tar.gz (NOT an ar archive like Debian's .deb).
|
||||
WORK=$(mktemp -d)
|
||||
trap 'rm -rf "$WORK"' EXIT
|
||||
|
||||
tar -xzf "$IPK" -C "$WORK"
|
||||
echo "==> top-level entries:"
|
||||
ls -la "$WORK"
|
||||
|
||||
# Top-level structural assertions.
|
||||
fail=0
|
||||
for entry in debian-binary control.tar.gz data.tar.gz; do
|
||||
if [ ! -f "$WORK/$entry" ]; then
|
||||
echo "FAIL: missing top-level $entry"
|
||||
fail=1
|
||||
else
|
||||
echo " PASS top-level: $entry"
|
||||
fi
|
||||
done
|
||||
if [ "$fail" -ne 0 ]; then exit 1; fi
|
||||
|
||||
# debian-binary content sanity.
|
||||
dbin_content=$(cat "$WORK/debian-binary" | tr -d '[:space:]')
|
||||
if [ "$dbin_content" != "2.0" ]; then
|
||||
echo "FAIL: debian-binary content is '$dbin_content' (expected 2.0)"
|
||||
exit 1
|
||||
fi
|
||||
echo " PASS debian-binary content: 2.0"
|
||||
|
||||
# Inspect data.tar.gz contents.
|
||||
DATA_LIST="$WORK/data.list"
|
||||
tar -tzf "$WORK/data.tar.gz" > "$DATA_LIST"
|
||||
echo "==> data.tar.gz entry count: $(wc -l < "$DATA_LIST")"
|
||||
|
||||
# Required filesystem entries inside data.tar.gz. Entries are
|
||||
# produced with a leading "./" by build-ipk.sh.
|
||||
REQUIRED=(
|
||||
./usr/bin/fips
|
||||
./usr/bin/fipsctl
|
||||
./usr/bin/fipstop
|
||||
./usr/bin/fips-gateway
|
||||
./etc/init.d/fips
|
||||
./etc/init.d/fips-gateway
|
||||
./etc/fips/fips.yaml
|
||||
./etc/fips/firewall.sh
|
||||
./etc/dnsmasq.d/fips.conf
|
||||
./etc/sysctl.d/fips-gateway.conf
|
||||
./etc/sysctl.d/fips-bridge.conf
|
||||
./etc/hotplug.d/net/99-fips
|
||||
./etc/uci-defaults/90-fips-setup
|
||||
./lib/upgrade/keep.d/fips
|
||||
)
|
||||
for path in "${REQUIRED[@]}"; do
|
||||
if grep -Fxq "$path" "$DATA_LIST"; then
|
||||
echo " PASS data: $path"
|
||||
else
|
||||
echo " FAIL data: missing $path"
|
||||
fail=1
|
||||
fi
|
||||
done
|
||||
|
||||
# Inspect control.tar.gz: must contain control file + maintainer scripts.
|
||||
CTRL_LIST="$WORK/control.list"
|
||||
tar -tzf "$WORK/control.tar.gz" > "$CTRL_LIST"
|
||||
echo "==> control.tar.gz entry count: $(wc -l < "$CTRL_LIST")"
|
||||
for path in ./control ./conffiles ./postinst ./prerm; do
|
||||
if grep -Fxq "$path" "$CTRL_LIST"; then
|
||||
echo " PASS control: $path"
|
||||
else
|
||||
echo " FAIL control: missing $path"
|
||||
fail=1
|
||||
fi
|
||||
done
|
||||
|
||||
if [ "$fail" -ne 0 ]; then
|
||||
echo "ipk structural verification FAILED"
|
||||
exit 1
|
||||
fi
|
||||
echo "ipk structural verification PASS"
|
||||
|
||||
- name: SHA-256 hashes
|
||||
run: |
|
||||
echo "==> Binaries:"
|
||||
sha256sum target/${{ matrix.rust_target }}/release/fips target/${{ matrix.rust_target }}/release/fipsctl target/${{ matrix.rust_target }}/release/fipstop
|
||||
echo "==> Package:"
|
||||
sha256sum dist/${{ env.PACKAGE_FILENAME }}
|
||||
|
||||
- name: Upload artifact (GitHub only)
|
||||
if: ${{ env.ACT != 'true' }}
|
||||
uses: actions/upload-artifact@v4
|
||||
with:
|
||||
name: ${{ steps.version.outputs.filename }}
|
||||
path: dist/${{ steps.version.outputs.filename }}
|
||||
name: ${{ env.PACKAGE_FILENAME }}
|
||||
path: dist/${{ env.PACKAGE_FILENAME }}
|
||||
retention-days: 30
|
||||
|
||||
- name: Upload to Blossom
|
||||
id: blossom_upload
|
||||
shell: bash
|
||||
env:
|
||||
BLOSSOM_SERVER: "https://blossom.primal.net"
|
||||
NSEC: ${{ steps.keys.outputs.nsec }}
|
||||
run: |
|
||||
: ${GITHUB_OUTPUT:=/tmp/github_output}
|
||||
|
||||
UPLOAD_RESPONSE=$(nak blossom upload \
|
||||
--server "$BLOSSOM_SERVER" \
|
||||
--sec "$NSEC" \
|
||||
"dist/${{ env.PACKAGE_FILENAME }}" < /dev/null)
|
||||
|
||||
echo "Upload response:"
|
||||
echo "$UPLOAD_RESPONSE"
|
||||
|
||||
FILE_HASH=$(echo "$UPLOAD_RESPONSE" | jq -r '.sha256')
|
||||
if [ -z "$FILE_HASH" ] || [ "$FILE_HASH" = "null" ]; then
|
||||
echo "Failed to extract hash from upload response"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
BLOSSOM_URL="${BLOSSOM_SERVER}/${FILE_HASH}"
|
||||
echo "url=$BLOSSOM_URL" >> "$GITHUB_OUTPUT"
|
||||
echo "hash=$FILE_HASH" >> "$GITHUB_OUTPUT"
|
||||
echo "Uploaded to Blossom: $BLOSSOM_URL"
|
||||
|
||||
- name: Publish NIP-94 release event
|
||||
id: publish
|
||||
shell: bash
|
||||
env:
|
||||
RELAYS: "wss://relay.damus.io wss://nos.lol wss://nostr.mom wss://offchain.pub"
|
||||
NSEC: ${{ steps.keys.outputs.nsec }}
|
||||
run: |
|
||||
: ${GITHUB_OUTPUT:=/tmp/github_output}
|
||||
set -e
|
||||
|
||||
VERSION="${{ needs.determine-versioning.outputs.package_version }}"
|
||||
CHANNEL="${{ needs.determine-versioning.outputs.release_channel }}"
|
||||
|
||||
nak event --sec "$NSEC" -k 1063 \
|
||||
-c "FIPS Package: ${{ env.PACKAGE_NAME }} for ${{ matrix.openwrt_arch }}" \
|
||||
--tag url="${{ steps.blossom_upload.outputs.url }}" \
|
||||
--tag m="application/octet-stream" \
|
||||
--tag x="${{ steps.blossom_upload.outputs.hash }}" \
|
||||
--tag ox="${{ steps.blossom_upload.outputs.hash }}" \
|
||||
--tag filename="${{ env.PACKAGE_FILENAME }}" \
|
||||
--tag A="${{ matrix.openwrt_arch }}" \
|
||||
--tag v="$VERSION" \
|
||||
--tag n="${{ env.PACKAGE_NAME }}" \
|
||||
--tag compression="none" \
|
||||
> event.json 2> event.err
|
||||
|
||||
if [ ! -s event.json ]; then
|
||||
echo "Failed to create event"
|
||||
cat event.err 2>/dev/null || true
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo "=== Event JSON ==="
|
||||
cat event.json
|
||||
echo "=================="
|
||||
|
||||
EVENT_ID=$(jq -r '.id' event.json)
|
||||
if [ -z "$EVENT_ID" ] || [ "$EVENT_ID" = "null" ]; then
|
||||
echo "Failed to extract event ID"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Publish to relays
|
||||
cat event.json | nak event $RELAYS 2>&1
|
||||
|
||||
echo "eventId=$EVENT_ID" >> "$GITHUB_OUTPUT"
|
||||
echo "Published NIP-94 event: $EVENT_ID"
|
||||
|
||||
- name: Verify NIP-94 event on relays
|
||||
if: ${{ steps.publish.outputs.eventId != '' }}
|
||||
env:
|
||||
EVENT_ID: ${{ steps.publish.outputs.eventId }}
|
||||
RELAYS: "wss://relay.damus.io wss://nos.lol wss://nostr.mom wss://offchain.pub"
|
||||
run: |
|
||||
echo "Verifying event $EVENT_ID on relays..."
|
||||
FOUND=0
|
||||
for relay in $RELAYS; do
|
||||
echo "Checking $relay..."
|
||||
RESULT=$(nak req -i "$EVENT_ID" "$relay" 2>/dev/null || echo "")
|
||||
if [ -n "$RESULT" ]; then
|
||||
echo "Found on $relay"
|
||||
FOUND=1
|
||||
else
|
||||
echo "Not found on $relay"
|
||||
fi
|
||||
done
|
||||
|
||||
if [ $FOUND -eq 0 ]; then
|
||||
echo "Warning: Event not found on any relay yet (may still be propagating)"
|
||||
else
|
||||
echo "Event verified on at least one relay"
|
||||
fi
|
||||
|
||||
- name: Build Summary
|
||||
run: |
|
||||
echo "Build Summary for ${{ matrix.openwrt_arch }}:"
|
||||
echo " Package: ${{ env.PACKAGE_FILENAME }}"
|
||||
echo " Release EventId: ${{ steps.publish.outputs.eventId }}"
|
||||
echo " Blossom URL: ${{ steps.blossom_upload.outputs.url }}"
|
||||
|
||||
release:
|
||||
name: Publish GitHub Release
|
||||
name: Publish GitHub Release (github only)
|
||||
runs-on: ubuntu-latest
|
||||
needs: build
|
||||
if: startsWith(github.ref, 'refs/tags/')
|
||||
|
||||
@@ -0,0 +1,206 @@
|
||||
name: Windows Package
|
||||
on:
|
||||
push:
|
||||
branches:
|
||||
- master
|
||||
- maint
|
||||
- next
|
||||
tags:
|
||||
- "v*"
|
||||
pull_request:
|
||||
workflow_dispatch:
|
||||
|
||||
env:
|
||||
CARGO_TERM_COLOR: always
|
||||
|
||||
jobs:
|
||||
determine-versioning:
|
||||
runs-on: windows-latest
|
||||
outputs:
|
||||
package_version: ${{ steps.version.outputs.package_version }}
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
with:
|
||||
fetch-depth: 0
|
||||
|
||||
- name: Derive package version
|
||||
id: version
|
||||
shell: pwsh
|
||||
run: |
|
||||
$cargoToml = Get-Content Cargo.toml -Raw
|
||||
if ($cargoToml -match 'version\s*=\s*"([^"]+)"') {
|
||||
$baseVersion = $Matches[1]
|
||||
} else {
|
||||
throw "Could not determine version from Cargo.toml"
|
||||
}
|
||||
|
||||
if ($env:GITHUB_REF -like "refs/tags/*") {
|
||||
$version = $env:GITHUB_REF_NAME -replace '^v', ''
|
||||
} else {
|
||||
$branch = $env:GITHUB_REF_NAME -replace '[^A-Za-z0-9]', '.' -replace '\.\.+', '.' -replace '^\.|\.$$', ''
|
||||
$height = git rev-list --count HEAD
|
||||
$hash = git rev-parse --short HEAD
|
||||
if (-not $branch) { $branch = "ref" }
|
||||
$version = "$baseVersion+$branch.$height.$hash"
|
||||
}
|
||||
|
||||
echo "package_version=$version" >> $env:GITHUB_OUTPUT
|
||||
|
||||
build:
|
||||
name: Build Windows package
|
||||
runs-on: windows-latest
|
||||
needs: determine-versioning
|
||||
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
with:
|
||||
fetch-depth: 0
|
||||
|
||||
- name: Set SOURCE_DATE_EPOCH from git
|
||||
shell: pwsh
|
||||
run: |
|
||||
$epoch = git log -1 --format=%ct
|
||||
echo "SOURCE_DATE_EPOCH=$epoch" >> $env:GITHUB_ENV
|
||||
|
||||
- name: Install Rust toolchain
|
||||
uses: dtolnay/rust-toolchain@stable
|
||||
|
||||
- name: Cache Cargo registry + build
|
||||
uses: actions/cache@v4
|
||||
with:
|
||||
path: |
|
||||
~/.cargo/registry
|
||||
~/.cargo/git
|
||||
target
|
||||
key: windows-release-x86_64-${{ hashFiles('**/Cargo.lock') }}
|
||||
restore-keys: |
|
||||
windows-release-x86_64-
|
||||
|
||||
- name: Build release binaries
|
||||
run: cargo build --release
|
||||
|
||||
- name: Build Windows package
|
||||
shell: pwsh
|
||||
run: |
|
||||
powershell -File packaging/windows/build-zip.ps1 `
|
||||
-Version "${{ needs.determine-versioning.outputs.package_version }}" `
|
||||
-NoBuild
|
||||
|
||||
- name: Verify ZIP structural correctness
|
||||
shell: pwsh
|
||||
run: |
|
||||
$zip = Get-ChildItem deploy\fips-*-windows-*.zip | Select-Object -First 1
|
||||
if (-not $zip) { Write-Error "No ZIP artifact found in deploy\"; exit 1 }
|
||||
Write-Host "Verifying: $($zip.Name)"
|
||||
|
||||
$extractDir = "verify-extract"
|
||||
if (Test-Path $extractDir) { Remove-Item -Recurse -Force $extractDir }
|
||||
Expand-Archive -Path $zip.FullName -DestinationPath $extractDir
|
||||
|
||||
# Expected top-level files (flat ZIP, no wrapper directory).
|
||||
# Source of truth: packaging/windows/build-zip.ps1
|
||||
$expected = @(
|
||||
"fips.exe",
|
||||
"fipsctl.exe",
|
||||
"fipstop.exe",
|
||||
"fips.yaml",
|
||||
"hosts",
|
||||
"install-service.ps1",
|
||||
"uninstall-service.ps1",
|
||||
"README.txt"
|
||||
)
|
||||
|
||||
$missing = @()
|
||||
foreach ($f in $expected) {
|
||||
$path = Join-Path $extractDir $f
|
||||
if (Test-Path -LiteralPath $path -PathType Leaf) {
|
||||
Write-Host "PASS: $f"
|
||||
} else {
|
||||
Write-Host "FAIL: $f"
|
||||
$missing += $f
|
||||
}
|
||||
}
|
||||
|
||||
Write-Host ""
|
||||
Write-Host "ZIP contents:"
|
||||
Get-ChildItem -Path $extractDir -Recurse | ForEach-Object {
|
||||
Write-Host " $($_.FullName.Substring((Resolve-Path $extractDir).Path.Length + 1))"
|
||||
}
|
||||
|
||||
if ($missing.Count -gt 0) {
|
||||
Write-Error "Missing expected files: $($missing -join ', ')"
|
||||
exit 1
|
||||
}
|
||||
Write-Host ""
|
||||
Write-Host "All expected files present."
|
||||
|
||||
- name: SHA-256 hash
|
||||
shell: pwsh
|
||||
run: |
|
||||
Write-Host "==> Windows release asset:"
|
||||
Get-ChildItem deploy\fips-*-windows-*.zip | ForEach-Object {
|
||||
Get-FileHash $_.FullName -Algorithm SHA256 | Format-Table -AutoSize
|
||||
}
|
||||
|
||||
- name: Upload artifact
|
||||
uses: actions/upload-artifact@v4
|
||||
with:
|
||||
name: fips_${{ needs.determine-versioning.outputs.package_version }}_x86_64_windows
|
||||
path: deploy/fips-*-windows-*.zip
|
||||
retention-days: 30
|
||||
|
||||
- name: Build summary
|
||||
shell: pwsh
|
||||
run: |
|
||||
$pkg = Get-ChildItem deploy\fips-*-windows-*.zip | Select-Object -First 1
|
||||
Write-Host "Build Summary for Windows/x86_64:"
|
||||
Write-Host " Package: $($pkg.Name)"
|
||||
Write-Host " Size: $([math]::Round($pkg.Length / 1MB, 2)) MB"
|
||||
|
||||
release:
|
||||
name: Publish Windows assets to GitHub Release
|
||||
runs-on: ubuntu-latest
|
||||
needs: build
|
||||
if: startsWith(github.ref, 'refs/tags/')
|
||||
permissions:
|
||||
contents: write
|
||||
|
||||
steps:
|
||||
- name: Download Windows artifacts
|
||||
uses: actions/download-artifact@v4
|
||||
with:
|
||||
path: dist
|
||||
merge-multiple: true
|
||||
|
||||
- name: Generate Windows release checksums
|
||||
run: |
|
||||
cd dist
|
||||
find . -maxdepth 1 -type f -name '*.zip' -printf '%P\n' \
|
||||
| LC_ALL=C sort \
|
||||
| xargs sha256sum \
|
||||
> checksums-windows.txt
|
||||
|
||||
- name: Wait for tag release
|
||||
env:
|
||||
GH_TOKEN: ${{ github.token }}
|
||||
run: |
|
||||
for attempt in $(seq 1 20); do
|
||||
if gh release view "${GITHUB_REF_NAME}" --repo "${GITHUB_REPOSITORY}" >/dev/null 2>&1; then
|
||||
exit 0
|
||||
fi
|
||||
echo "Release ${GITHUB_REF_NAME} not available yet; waiting..."
|
||||
sleep 15
|
||||
done
|
||||
|
||||
echo "Timed out waiting for release ${GITHUB_REF_NAME}" >&2
|
||||
exit 1
|
||||
|
||||
- name: Upload Windows assets
|
||||
env:
|
||||
GH_TOKEN: ${{ github.token }}
|
||||
run: |
|
||||
gh release upload "${GITHUB_REF_NAME}" \
|
||||
dist/*.zip \
|
||||
dist/checksums-windows.txt \
|
||||
--clobber \
|
||||
--repo "${GITHUB_REPOSITORY}"
|
||||
@@ -11,13 +11,31 @@
|
||||
.vscode/
|
||||
.idea/
|
||||
|
||||
# Claude Code
|
||||
# AI Agents
|
||||
.claude/
|
||||
AGENTS.md
|
||||
CLAUDE.md
|
||||
agents/
|
||||
|
||||
deploy/
|
||||
vps.env
|
||||
|
||||
reference/
|
||||
/reference/
|
||||
|
||||
dist/
|
||||
*.ipk
|
||||
*.ipk
|
||||
|
||||
sim-results/
|
||||
|
||||
# Python
|
||||
__pycache__/
|
||||
*.py[cod]
|
||||
*.egg-info/
|
||||
*.egg
|
||||
|
||||
# Runtime artifacts from running fips in-tree during local testing.
|
||||
# Root-anchored so legitimately-tracked fips.yaml under packaging/ and
|
||||
# examples/ stays included.
|
||||
/fips.key
|
||||
/fips.pub
|
||||
/fips.yaml
|
||||
|
||||
@@ -5,7 +5,765 @@ All notable changes to this project will be documented in this file.
|
||||
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
|
||||
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
||||
|
||||
## [Unreleased]
|
||||
## [0.3.0] - 2026-05-11
|
||||
|
||||
### Added
|
||||
|
||||
#### Mesh Layer (FMP)
|
||||
|
||||
- Overlay-discovery and NAT-hole-punching path (opt-in via
|
||||
`node.discovery.nostr.enabled`). Nodes publish signed overlay adverts
|
||||
as Nostr kind `37195` parameterized replaceable events listing
|
||||
reachable transport endpoints to a configurable set of public relays,
|
||||
and consume peer adverts to populate fallback addresses for
|
||||
`via_nostr` peers or, under `policy: open`, for non-configured peers
|
||||
within a budget cap. The kind value is FIPS-specific: `37195` sits in
|
||||
the application-defined replaceable range `30000–39999`, and the
|
||||
digits visually spell `FIPS` (7=F, 1=I, 9=P, 5=S)
|
||||
- STUN-assisted UDP hole punching for `addr: "nat"` UDP endpoints. STUN
|
||||
reflexive observation, gift-wrap (NIP-59) offer/answer signaling, and
|
||||
candidate-pair punch planner (LAN-private + reflexive paths attempted in
|
||||
parallel). Successful punches hand the live socket into the standard
|
||||
FIPS UDP transport via a bootstrap-handoff API
|
||||
- New `node.discovery.nostr.*` configuration tree with operator-tunable
|
||||
resource caps, replay tracking, and punch timing; new `peers[].via_nostr`
|
||||
and per-transport `advertise_on_nostr` / `public` flags. Cross-field
|
||||
validation at startup catches mis-configured combinations
|
||||
- Docker NAT lab covering cone, symmetric (TCP-fallback), and LAN
|
||||
scenarios, wired into the integration CI matrix
|
||||
- One-shot startup advert sweep for Nostr open-discovery. On daemon
|
||||
startup under `node.discovery.nostr.policy: open`, after a short
|
||||
settle delay (`startup_sweep_delay_secs`, default 5s) the cached
|
||||
overlay-advert table is iterated once and recent adverts (newer
|
||||
than `startup_sweep_max_age_secs`, default 3600s) are queued for
|
||||
outbound retry, modulo the same skip-filters as the per-tick sweep
|
||||
(configured peer, already connected, retry-pending, connecting).
|
||||
Closes the gap where peers learned only through relay backlog at
|
||||
startup were not dialed until they republished.
|
||||
- Diagnostic logging on the open-discovery sweep. Each `queued retry`
|
||||
now logs at info-level with the peer short-npub and advert age,
|
||||
and a one-line summary (cached count, queued count, per-reason
|
||||
skip counts) is emitted on every startup sweep and on any per-tick
|
||||
sweep that queues at least one retry. Operator-facing visibility
|
||||
into what the auto-dial path is doing.
|
||||
|
||||
#### Platform Support
|
||||
|
||||
- Windows platform support: wintun TUN device, TCP control socket on
|
||||
`localhost:21210` (in place of the Unix domain socket), Windows
|
||||
Service lifecycle (`--install-service`, `--uninstall-service`,
|
||||
`--service`), ZIP packaging with PowerShell install/uninstall scripts,
|
||||
and CI build/test matrix entry
|
||||
([#45](https://github.com/jmcorgan/fips/pull/45))
|
||||
- macOS platform support: native `utun` TUN interface management, raw
|
||||
Ethernet transport via BPF, `.pkg` packaging with launchd plist and
|
||||
uninstall script, x86_64 cross-compile from arm64, and CI build/unit
|
||||
test jobs
|
||||
- MIPS atomic ABI support: `std::sync::atomic` replaced with
|
||||
`portable_atomic` so 32-bit MIPS targets without native atomics
|
||||
link cleanly
|
||||
([#62](https://github.com/jmcorgan/fips/pull/62),
|
||||
[@andrewheadricke](https://github.com/andrewheadricke)).
|
||||
|
||||
#### Mesh Peer Transports
|
||||
|
||||
- Bluetooth Low Energy (BLE) L2CAP Connection-Oriented Channel
|
||||
transport (Linux only, requires BlueZ): per-link MTU negotiation,
|
||||
continuous scan/probe peer discovery with cooldown-based
|
||||
deduplication, continuous advertising, deterministic NodeAddr
|
||||
cross-probe tie-breaker, and a configurable connection pool with
|
||||
eviction.
|
||||
- `transports.udp.outbound_only` (default `false`). When true, the UDP
|
||||
transport binds a kernel-assigned ephemeral port (`0.0.0.0:0`) instead
|
||||
of the configured `bind_addr`, refuses inbound handshakes, and is
|
||||
never advertised on Nostr regardless of `advertise_on_nostr`. Use
|
||||
this to participate in the mesh as a pure client — initiate outbound
|
||||
links without exposing an inbound listener on a known port.
|
||||
Implements the long-form fix for `udp.bind_addr: "127.0.0.1:..."`
|
||||
not actually working as a workaround (Linux pins the loopback source
|
||||
IP, dropping outbound flows to external peers at the routing layer)
|
||||
- `transports.udp.accept_connections` (default `true`). Mirrors the
|
||||
Ethernet/BLE knob; setting to `false` produces a "client" posture
|
||||
(initiate outbound, refuse inbound msg1 from new addresses). The
|
||||
Node-level handshake gate carves out msg1 from peers already
|
||||
established on this transport so rekey continues to work. Affects
|
||||
every transport via the `Transport` trait
|
||||
- Startup validation now rejects `transports.udp[*].bind_addr` set to a
|
||||
loopback address when at least one peer has a non-loopback UDP
|
||||
address. Replaces the silent "peer link won't establish" failure
|
||||
mode where Linux's source-address routing check dropped outbound
|
||||
flows from the loopback-bound socket. `outbound_only: true` is
|
||||
exempt from the check (it overrides `bind_addr` to `0.0.0.0:0`)
|
||||
|
||||
#### Security
|
||||
|
||||
- Mesh-interface nftables baseline (Linux). Ships `/etc/fips/fips.nft`
|
||||
as a documented operator conffile and `fips-firewall.service`
|
||||
(disabled by default) for default-deny inbound on the `fips0` mesh
|
||||
interface. Operators enable explicitly with
|
||||
`systemctl enable --now fips-firewall.service`. Drop-ins in
|
||||
`/etc/fips/fips.d/*.nft`. See `docs/fips-security.md`.
|
||||
- Peer access control list enforcement: optional
|
||||
`/etc/fips/peers.allow` and `/etc/fips/peers.deny` files
|
||||
(TCP-Wrappers style) gate outbound connect, inbound msg1, and
|
||||
outbound msg2 against npub, hex pubkey, host alias, or `ALL`.
|
||||
Files are reloaded automatically on mtime change. New
|
||||
`fipsctl acl show` query reports the effective rule set
|
||||
([#50](https://github.com/jmcorgan/fips/pull/50),
|
||||
[@alexxie16](https://github.com/alexxie16)).
|
||||
|
||||
#### LAN Gateway
|
||||
|
||||
- New `fips-gateway` binary that lets unmodified LAN hosts reach FIPS
|
||||
mesh destinations via DNS-allocated virtual IPs and kernel nftables
|
||||
NAT. Virtual-IP pool (`fd01::/112` by default) with state-machine
|
||||
lifecycle and TTL-based reclamation; conntrack-backed session
|
||||
tracking; proxy NDP on the LAN interface; control socket at
|
||||
`/run/fips/gateway.sock` with `show_gateway` and `show_mappings`;
|
||||
fipstop Gateway tab with pool gauge and mappings table; design doc
|
||||
at `docs/design/fips-gateway.md`; integration test harness
|
||||
- Inbound mesh port forwarding on `fips-gateway`: new
|
||||
`gateway.port_forwards` config (list of `{ listen_port, proto,
|
||||
target }` entries, IPv6 targets only) installs prerouting DNAT
|
||||
rules so mesh peers can reach a configured host:port on the
|
||||
gateway's LAN. A LAN-side masquerade is added when any forwards
|
||||
are configured so replies flow back through conntrack.
|
||||
- Gateway packaging: systemd service unit with `After=fips.service`,
|
||||
Debian and AUR package entries, OpenWrt procd init with dnsmasq
|
||||
forwarding, proxy NDP, RA route advertisements, and IPv6 forwarding
|
||||
sysctls. Gateway enabled by default on OpenWrt
|
||||
- `fips-gateway` DNS upstream probe now retries up to 5 times with a
|
||||
1-second per-attempt timeout and a 1-second delay between attempts
|
||||
(~10 second worst-case wait), instead of a single 3-second hard-fail.
|
||||
Covers the cold-boot race where the daemon's TUN is up (the systemd
|
||||
ExecStartPre wait gates on that) but the DNS responder is still
|
||||
binding `[::1]:5354`. Without retry the gateway exited and relied on
|
||||
`Restart=on-failure` for recovery (5-second blip + spurious error
|
||||
log line per cycle); with retry the gateway recovers gracefully
|
||||
without a unit restart
|
||||
|
||||
#### IPv6 Adapter
|
||||
|
||||
- Overhauled `.fips` DNS handling for systemd-based hosts. The
|
||||
default `dns.bind_addr` is `::1` (IPv6 loopback) and the setup
|
||||
script picks one of five backends in priority order: a global
|
||||
drop-in at `/etc/systemd/resolved.conf.d/fips.conf`, the systemd
|
||||
dns-delegate path, `resolvectl` per-link, standalone dnsmasq, or
|
||||
NetworkManager's dnsmasq plugin. Teardown reverses only what was
|
||||
applied. New `testing/dns-resolver/` harness exercises every
|
||||
backend across Debian 12, Debian 13, Ubuntu 22.04, Ubuntu 24.04,
|
||||
and Ubuntu 26.04
|
||||
([#58](https://github.com/jmcorgan/fips/pull/58),
|
||||
fixes [#52](https://github.com/jmcorgan/fips/issues/52),
|
||||
[#77](https://github.com/jmcorgan/fips/issues/77)).
|
||||
|
||||
#### Operator Tooling
|
||||
|
||||
- `node.log_level` config field (case-insensitive, default `info`)
|
||||
replaces the hardcoded `RUST_LOG=info` previously baked into
|
||||
systemd units and the OpenWrt procd init script. The daemon now
|
||||
loads config before initializing tracing so the configured level
|
||||
takes effect; `RUST_LOG` still overrides when set
|
||||
- `fipsctl show identity-cache` lists every cached node identity
|
||||
(npub, IPv6 address, display name, LRU age) alongside the
|
||||
configured cache capacity
|
||||
- `fipsctl show peers` extended with per-peer security signals
|
||||
(replay suppression count, consecutive decrypt failures), Noise
|
||||
session counters, session indices, and rekey lifecycle state
|
||||
- `fipsctl show sessions` extended with handshake resend count
|
||||
during establishment and rekey/session health fields when
|
||||
established (session start, K-bit epoch, coords warmup remaining,
|
||||
drain state)
|
||||
- `fipsctl show cache` now includes individual coordinate cache
|
||||
entries (tree coordinates, depth, path MTU, age). The top-level
|
||||
count field was renamed from `entries` to `count` for clarity
|
||||
- `fipsctl show routing` expands `pending_lookups` from a count to
|
||||
per-target detail (attempt, age, last sent), adds pending TUN
|
||||
packet queue depth, and adds per-peer connection retry state
|
||||
([#42](https://github.com/jmcorgan/fips/pull/42),
|
||||
[@osh](https://github.com/osh))
|
||||
- Historical node and per-peer statistics: in-memory time-series
|
||||
rings on the daemon, surfaced through new control-socket queries,
|
||||
`fipsctl stats` subcommands, and a `fipstop` Graphs tab with
|
||||
btop-style sparklines
|
||||
([#64](https://github.com/jmcorgan/fips/pull/64)).
|
||||
- `fipstop` Node tab now carries a "Listening on fips0" panel
|
||||
(right-half of the Traffic block) that lists local IPv6 listening
|
||||
sockets reachable from the mesh interface, paired with the
|
||||
`inet fips` baseline filter classification for each (proto, port).
|
||||
Rows render in default White (`OPEN` — the chain has a canonical
|
||||
unrestricted accept rule), DarkGray (`filt` — chain falls through
|
||||
to `counter drop`), or DarkGray with a `?` State suffix (`filt?` —
|
||||
the chain references the port but with matchers the panel cannot
|
||||
fully decompose, e.g. saddr filters or jumps). When the
|
||||
`fips-firewall.service` is not active, the panel renders a yellow
|
||||
banner reminding the operator that all listeners are
|
||||
mesh-exposed. Wildcard binds (`local_addr == ::`) carry a `*`
|
||||
suffix in the Process column. Powered by a new
|
||||
`show_listening_sockets` control query (Linux-only).
|
||||
|
||||
#### Packaging and Deployment
|
||||
|
||||
- Arch Linux AUR packaging for `fips` (release) and `fips-git`
|
||||
(development) packages with sysusers.d/tmpfiles.d integration
|
||||
([#21](https://github.com/jmcorgan/fips/pull/21),
|
||||
[@dskvr](https://github.com/dskvr))
|
||||
- `packaging/debian/fips-gateway.service` now waits up to 30 seconds
|
||||
for the daemon's `fips0` TUN to appear before exec'ing the gateway
|
||||
binary (`ExecStartPre` poll loop). Eliminates the cold-boot race
|
||||
where `fips-gateway` exits with `fips0 interface not found` and
|
||||
recovers via `Restart=on-failure`, producing a 5-second blip and a
|
||||
spurious error log line per restart cycle. If `fips0` never appears
|
||||
within 30 seconds, the existing error path runs as before
|
||||
- `packaging/debian/build-deb.sh` now auto-derives a per-commit Debian
|
||||
Version field for dev builds (Cargo.toml version ending in `-dev`)
|
||||
using the form `<base>~dev+git<YYYYMMDD>.<sha>[.dirty]-1`, e.g.
|
||||
`0.3.0~dev+git20260429.6def31b-1`. Each commit produces a uniquely-
|
||||
comparable Version string so `apt install ./*.deb` and
|
||||
`ansible.builtin.apt: deb:` no longer silently no-op when one dev
|
||||
build is installed on top of another. The `~dev` marker sorts
|
||||
pre-`0.3.0` so a tagged release supersedes any prior dev .deb.
|
||||
Tagged release builds (no `-dev` in Cargo.toml) keep the clean
|
||||
`<version>-1` form. Operator override via `--version` still wins
|
||||
|
||||
#### Examples
|
||||
|
||||
- macOS WireGuard sidecar: run FIPS in a local Docker container and
|
||||
route `.fips` traffic from the macOS host through a WireGuard tunnel
|
||||
to the container's `fips0` interface. Only traffic destined for
|
||||
`fd00::/8` transits the sidecar; regular internet traffic continues
|
||||
to use the host network
|
||||
([#51](https://github.com/jmcorgan/fips/pull/51))
|
||||
|
||||
#### Documentation
|
||||
|
||||
- `docs/design/port-advertisement-and-nat-traversal.md` documents
|
||||
how nodes find each other through Nostr relays and the
|
||||
STUN-assisted UDP hole punch
|
||||
|
||||
### Changed
|
||||
|
||||
- Noise session ChaCha20-Poly1305 backend switched from RustCrypto's
|
||||
`chacha20poly1305` to `ring 0.17`. ring wraps BoringSSL's
|
||||
hand-tuned ChaCha20-Poly1305 implementation, dispatching to NEON
|
||||
on aarch64 and AVX2 / AVX-512 on x86_64 — typically 3-5 GB/s/core
|
||||
vs the ~600-800 MB/s/core RustCrypto soft path on the same
|
||||
hardware. Wire format unchanged: ChaCha20-Poly1305 is
|
||||
byte-deterministic for a given `(key, nonce, plaintext, aad)`,
|
||||
so any correct AEAD produces identical ciphertext and a mixed
|
||||
pre-swap / post-swap mesh interoperates without protocol
|
||||
awareness. The keyed AEAD is now cached on `CipherState` instead
|
||||
of being re-derived per packet (the cached Poly1305 key state is
|
||||
the actual perf win); `EndToEndState` grew from ~600 B to
|
||||
~1.5 KB as a consequence and is annotated
|
||||
`#[allow(clippy::large_enum_variant)]` since boxing would re-add
|
||||
a per-packet indirection on every encrypt/decrypt. aarch64
|
||||
measurements (Apple Silicon docker, two nodes): TCP 1-stream
|
||||
437 → 1097 Mbps (~2.5×); UDP at 1000 Mbit goes from
|
||||
599 Mbps / 40 % loss to lossless line-rate; 3-node ping under
|
||||
load 7.68 ms avg / 215 ms max → 0.72 ms / 3.6 ms max as the
|
||||
relay path stops being crypto-bound
|
||||
([#80](https://github.com/jmcorgan/fips/pull/80),
|
||||
[@mmalmi](https://github.com/mmalmi))
|
||||
- Linux UDP receive path uses `recvmmsg(2)` with a 32-packet batch
|
||||
in place of single-packet `recvmsg(2)`. A single `readable()`
|
||||
wakeup drains up to 32 datagrams in one syscall before yielding
|
||||
back to the reactor, eliminating the per-packet scheduler-hop +
|
||||
futex cost that previously capped inbound rate at one event per
|
||||
scheduler quantum independent of CPU. `SO_RXQ_OVFL` is sampled
|
||||
once per batch from the cmsg chain of `msgs[0]` and surfaced
|
||||
through `AsyncUdpSocket::recv_batch` so the 1Hz
|
||||
`sample_transport_congestion()` detector continues to feed the
|
||||
per-transport `dropping` flag. macOS / Windows fall through to
|
||||
the per-packet path; `recvmmsg` is Linux-specific
|
||||
([#81](https://github.com/jmcorgan/fips/pull/81),
|
||||
[@mmalmi](https://github.com/mmalmi))
|
||||
- `Node::run_rx_loop` drains up to 256 additional ready items via
|
||||
`try_recv()` after each `tokio::select!` await fires on
|
||||
`packet_rx` / `tun_outbound_rx`, in a tight inner loop before
|
||||
yielding. Previously the select cost a full scheduler hop +
|
||||
futex per packet, capping throughput at one event per scheduler
|
||||
quantum with the worker near-idle. `biased` ordering keeps
|
||||
data-plane branches priority over tick / control / DNS under
|
||||
sustained load; the 256 cap is empirically tuned to keep the
|
||||
worker on a busy stream between yield points (≈ 400 KB of
|
||||
contiguous traffic) while still bounding the inner loop so a
|
||||
flood on one branch can't starve the periodic tick or control
|
||||
socket. Pairs with the UDP `recvmmsg` change above
|
||||
([#81](https://github.com/jmcorgan/fips/pull/81),
|
||||
[@mmalmi](https://github.com/mmalmi))
|
||||
- `PeerIdentity::pubkey_full()` now precomputes the parity-aware
|
||||
full public key at construction in `from_pubkey`. Previously the
|
||||
method fell through to a secp256k1 EC point parse (`fe_sqrt` +
|
||||
`fe_mul` + `ge_set_xo_var`) on every call when the full key
|
||||
wasn't passed at construction (i.e. for every peer constructed
|
||||
from an npub or x-only key) — ~6% of per-packet CPU on the
|
||||
bulk-data send path for a value that never changed after
|
||||
construction. The same EC point parse already runs at
|
||||
construction inside `NodeAddr::from_pubkey`, so the cost is paid
|
||||
once where it would be paid anyway
|
||||
([#81](https://github.com/jmcorgan/fips/pull/81),
|
||||
[@mmalmi](https://github.com/mmalmi))
|
||||
- Cargo feature flags `tui`, `ble`, `gateway`, and
|
||||
`nostr-discovery` removed; subsystem inclusion is now driven by
|
||||
platform `cfg` gates so plain `cargo build` compiles everything
|
||||
available on the target
|
||||
([#79](https://github.com/jmcorgan/fips/pull/79))
|
||||
- MMP link-layer report intervals retuned for constrained transports:
|
||||
steady-state floor raised from 100ms to 1000ms, ceiling from 2000ms
|
||||
to 5000ms. Cold-start uses a 200ms floor for the first 5 SRTT samples
|
||||
before switching to steady-state. Reduces BLE overhead ~10× while
|
||||
keeping reports well above the EWMA convergence threshold.
|
||||
Session-layer intervals unchanged
|
||||
- 35 info-level log messages demoted to debug (handshake
|
||||
cross-connection mechanics, periodic MMP telemetry, TUN/transport
|
||||
shutdown, retry scheduling). Info output now focuses on
|
||||
operator-relevant state changes: lifecycle events, peer promotions,
|
||||
session establishment, parent switches, transport start/stop
|
||||
- **Breaking (control socket JSON):** `show_cache` response field
|
||||
`entries` has changed type from a `u64` count to an array of entry
|
||||
objects; a new `count` field carries the previous scalar value.
|
||||
`show_routing` response field `pending_lookups` has changed type
|
||||
from a `u64` count to an array of per-target lookup objects.
|
||||
External consumers parsing these fields as numbers must be
|
||||
updated. In-tree `fipstop` is adjusted to the new schema. The
|
||||
control socket interface is still pre-1.0 and not covered by
|
||||
stability guarantees
|
||||
- Discovery rate limiting retuned to be less aggressive at cold start.
|
||||
The previous defaults (30s base post-failure suppression, doubling
|
||||
to a 300s cap, with reset only on parent change / new peer / first
|
||||
RTT / reconnection) reliably outlasted initial mesh convergence: a
|
||||
single timed-out lookup during bloom-filter propagation suppressed
|
||||
any retry for 30s while none of the reset triggers fired on a
|
||||
stable post-handshake topology. The suppression window dictated
|
||||
effective time-to-converge instead of bounding repeat traffic.
|
||||
Replaces the single-lookup-with-internal-retry model
|
||||
(`timeout_secs`/`retry_interval_secs`/`max_attempts`) with a
|
||||
per-attempt timeout sequence in
|
||||
`node.discovery.attempt_timeouts_secs` (default `[1, 2, 4, 8]`).
|
||||
Each attempt sends a fresh `LookupRequest` with a new `request_id`,
|
||||
which lets successive attempts take different forwarding paths as
|
||||
the bloom and tree state evolve. The destination is declared
|
||||
unreachable only after the full sequence is exhausted (15s total
|
||||
at the default). Disables post-failure suppression by default
|
||||
(`backoff_base_secs`/`backoff_max_secs` now both `0`); operators
|
||||
with chatty apps generating repeat lookups against unreachable
|
||||
destinations can opt back in
|
||||
- The `docs/` tree is reorganised so readers can find content by
|
||||
what they're trying to do: tutorials for new users, how-to guides
|
||||
for specific tasks, reference material for configuration and
|
||||
protocol details, and design discussion for architectural
|
||||
background. New top-level `getting-started.md` and per-section
|
||||
landing pages anchor the entry points. Content was reconciled
|
||||
against current source: protocol layer details, wire-format
|
||||
diagrams, configuration knobs, and CLI references were brought
|
||||
back into agreement with the implementation. Gateway feature-set
|
||||
documentation was rewritten end-to-end.
|
||||
- Test coverage was substantially expanded for the new release
|
||||
surface (discovery state machine, control-socket query handlers,
|
||||
decrypt-failure thresholds, STUN parser, gateway, NAT traversal,
|
||||
packaging install paths) alongside CI-side hardening for the new
|
||||
Windows and macOS platforms.
|
||||
- Gateway `dns.listen` source default changed from `[::]:53` to
|
||||
`[::1]:5353` to match the canonical deployment model (a host
|
||||
already serving DHCP/DNS to a LAN segment, where port 53 is
|
||||
taken by the existing resolver and `.fips` queries are forwarded
|
||||
to the gateway over loopback). The OpenWrt ipk previously
|
||||
overrode this in its packaged config; the override is now
|
||||
redundant and has been dropped. Operators on a host without a
|
||||
pre-existing resolver on port 53 can opt back into the wildcard
|
||||
bind by setting `dns.listen: "[::]:53"` explicitly. The new
|
||||
default binds IPv6 loopback only — forwarders that reach the
|
||||
gateway over IPv4 loopback need an explicit IPv4 listen address.
|
||||
- Generic systemd install tarball brought to feature parity with
|
||||
the `.deb` and AUR packages. The tarball now ships the
|
||||
`fips-gateway` binary with its (operator-opt-in)
|
||||
`fips-gateway.service`, a `fips-firewall.service` unit with the
|
||||
`/etc/fips/fips.nft` mesh-interface nftables baseline (also
|
||||
opt-in), an `/etc/fips/fips.d/` operator drop-in directory for
|
||||
per-service nft rules, and the multi-backend `fips-dns-setup` /
|
||||
`fips-dns-teardown` helpers. `install.sh` and `uninstall.sh`
|
||||
handle the new units and conffile (preserve-on-upgrade for
|
||||
`fips.nft`, like `fips.yaml`). `README.install.md` documents
|
||||
the gateway, firewall, and DNS-routing services. Closes the
|
||||
longest-standing parity gap for non-Debian / non-Arch systemd
|
||||
Linux distros (Fedora, RHEL/CentOS, openSUSE, etc.) installing
|
||||
from the release-distribution tarball.
|
||||
|
||||
### Fixed
|
||||
|
||||
- Generic systemd install tarball: `install.sh` now correctly
|
||||
resolves the `fips-dns-setup` and `fips-dns-teardown` helpers
|
||||
from the tarball staging directory. Previously the script
|
||||
referenced them at `${SCRIPT_DIR}/../common/`, a path that
|
||||
exists only in the source-repo layout, not in the extracted
|
||||
tarball. Bug latent since the multi-backend DNS helpers
|
||||
landed in `7260ad2`; only manifested when operators ran
|
||||
`install.sh` from an extracted tarball rather than from a
|
||||
source checkout.
|
||||
|
||||
- Adopted NAT-traversed UDP transports inherit the primary listener's
|
||||
MTU and buffer config. `Node::adopt_established_traversal`
|
||||
constructed the adopted UDP transport with `UdpConfig::default()`
|
||||
(MTU 1280, default recv/send buffer sizes, default accept/advertise
|
||||
flags) regardless of the operator's primary `[transports.udp]`
|
||||
listener. Operators who set the primary MTU higher (e.g. 1500 on
|
||||
a known-clean LAN path) silently dropped full-sized tunnel
|
||||
datagrams over the NAT-traversed link with no log explaining why
|
||||
throughput collapsed. Lookup now tries `transport_name` first (so
|
||||
multiple named listeners pick up inheritance from the matching
|
||||
one) and falls back to the unnamed `Single` listener; bind /
|
||||
external-address fields are cleared since the adopted socket is
|
||||
already bound. The 1280 default was deliberately the IPv6 minimum
|
||||
(the only value guaranteed across arbitrary middlebox paths);
|
||||
with this change, operators who raise the primary MTU accept the
|
||||
tradeoff that NAT-traversed flows initially attempt the higher
|
||||
MTU and may black-hole on tighter paths until reactive
|
||||
`MtuExceeded` recovery kicks in
|
||||
([#83](https://github.com/jmcorgan/fips/pull/83),
|
||||
[@mmalmi](https://github.com/mmalmi))
|
||||
- TreeAnnounce ancestry on self-root transitions. When a node had
|
||||
no smaller-NodeAddr peer to use as a parent, the spanning-tree
|
||||
state correctly promoted it to root, but the ancestry it
|
||||
advertised on the next `TreeAnnounce` still referenced its
|
||||
previous parent's path. Receiving peers rejected the announce
|
||||
with `invalid ancestry: advertised root X is not the minimum
|
||||
path entry Y`, blocking mesh transit on any path that needed to
|
||||
traverse the node. The self-root transition is now detected
|
||||
explicitly in `TreeState::become_root` and the advertised
|
||||
ancestry rebuilt to start from self. The MMP receive handler
|
||||
surfaces the same path so stale ancestry inherited across
|
||||
reconnect is corrected eagerly rather than waiting for the next
|
||||
observation tick
|
||||
([#82](https://github.com/jmcorgan/fips/pull/82),
|
||||
[@mmalmi](https://github.com/mmalmi))
|
||||
- Auto-connect retry refetches the cached overlay advert
|
||||
unconditionally before each retry attempt, not only when
|
||||
`fetch_advert` returns zero endpoints (`NoTransportForType`).
|
||||
The much more common stale-cache failure was: cache returned an
|
||||
endpoint that *looked* valid (the address learned before the
|
||||
peer's NAT rebound), the dial succeeded at the IP layer, the
|
||||
handshake timed out, MMP fired, the next retry hit the same
|
||||
cached endpoint, looped forever — no `NoTransportForType` ever
|
||||
fired because the cache had data, just dead data. Refetch now
|
||||
runs unconditionally before each retry attempt (one Filter query
|
||||
against `advert_relays` with a 2s per-attempt timeout, bounded
|
||||
by the retry backoff cadence). Keeps the retry loop pinned to
|
||||
relay ground truth instead of whatever the cache happened to
|
||||
learn at startup
|
||||
([#82](https://github.com/jmcorgan/fips/pull/82),
|
||||
[@mmalmi](https://github.com/mmalmi))
|
||||
- Stale overlay-advert eviction on `NoTransportForType`. Mirrors
|
||||
the existing stale-advert sweep that ran from the
|
||||
`BootstrapEvent::Failed` (NAT-traversal-streak) path, but covers
|
||||
the case where `initiate_peer_connection` / a retry tick returns
|
||||
`NodeError::NoTransportForType` — the cache had no addresses for
|
||||
the peer at all. A fire-and-forget `refetch_advert_for_stale_check`
|
||||
against the peer's npub re-fetches kind `37195` from
|
||||
`advert_relays`; if the relay has a newer advert it replaces the
|
||||
cached entry, if it has nothing it evicts the entry. Either way
|
||||
the next retry tick goes to fresh data instead of looping on the
|
||||
same dead endpoint. Resolves a deployment regression where a
|
||||
macOS daemon's view of a Linux peer would flap after NAT rebind
|
||||
with no recovery short of a daemon restart
|
||||
([#82](https://github.com/jmcorgan/fips/pull/82),
|
||||
[@mmalmi](https://github.com/mmalmi))
|
||||
- Schedule retry on startup peer-init failure. When
|
||||
`initiate_peer_connections()` ran at boot, an address-resolution
|
||||
failure (no operational transport for the configured transport
|
||||
types, all addresses unreachable, NAT rebind invalidating cached
|
||||
endpoints) was logged and silently forgotten — the peer entry
|
||||
stayed in a dead state forever, accepting incoming pings but
|
||||
unable to answer them, until the daemon was manually restarted.
|
||||
Now mirrors the `BootstrapEvent::Failed` path: on a startup
|
||||
peer-init error, parse the peer's npub and call `schedule_retry`
|
||||
so the peer recovers without operator intervention
|
||||
([#82](https://github.com/jmcorgan/fips/pull/82),
|
||||
[@mmalmi](https://github.com/mmalmi))
|
||||
- Default control-socket path resolution: daemon and client tools now
|
||||
use a shared resolver, eliminating a divergence where `fipsctl` /
|
||||
`fipstop` could connect to a socket the daemon never bound (notably
|
||||
on dev runs with `XDG_RUNTIME_DIR` set, or after a prior packaged
|
||||
install left a root-owned `/run/fips` behind). Canonical order is
|
||||
`/run/fips` → `$XDG_RUNTIME_DIR/fips/` → `/tmp/fips-<name>`. The
|
||||
`/run/fips` arm is selected by directory existence; the kernel
|
||||
enforces actual access at `connect(2)` time, surfacing a clear
|
||||
`EACCES` for users not yet in the `fips` group rather than silently
|
||||
steering them to a path the daemon never bound. `XDG_RUNTIME_DIR` is
|
||||
validated as an existing directory before being used so stale
|
||||
post-logout values are treated as missing. The deployed fleet is
|
||||
unaffected: packaged configs set `node.control.socket_path`
|
||||
explicitly.
|
||||
- UDP transport with `advertise_on_nostr: true` + `public: true` +
|
||||
a wildcard `bind_addr` (e.g. `0.0.0.0:2121`) is now advertised
|
||||
with its STUN-discovered public IPv4 instead of being silently
|
||||
dropped from the published Kind 37195 advert. Previously the
|
||||
advert builder filtered the wildcard out (since `0.0.0.0` is
|
||||
not a valid endpoint), but emitted no log explaining what
|
||||
happened — operators saw the daemon up, both flags set, and
|
||||
no UDP endpoint in the advert. The fix runs a one-shot STUN
|
||||
observation against an ephemeral socket on the daemon's
|
||||
configured `stun_servers` and combines the reflexive IPv4 with
|
||||
the configured listener port for the advert (`udp:<eip>:<port>`).
|
||||
Successful STUN observations are cached per-transport for one
|
||||
`advert_refresh_secs` cycle (default 30 min) so we don't re-STUN
|
||||
every refresh. Failed observations are cached for only 60s, so
|
||||
a transient STUN flake at startup retries within ~a minute and
|
||||
grows the advert with UDP as soon as STUN starts working —
|
||||
rather than waiting the full 30-min cycle. Per-server STUN
|
||||
response timeout is 5s for the advert-publish path (vs. 2s for
|
||||
the latency-sensitive per-traversal path), giving slow
|
||||
first-call STUN time to complete without giving up. On STUN
|
||||
failure, the wildcard-bind path still skips, but now logs a
|
||||
loud `warn!` pointing at the operator-side fixes (set
|
||||
`external_addr`, bind to a specific IP, or ensure `stun_servers`
|
||||
reachable). Restores zero-config public-IP autodiscovery on
|
||||
AWS EIP / GCP / Azure setups where binding to the public IP
|
||||
directly is impossible (1:1 NAT)
|
||||
- New `external_addr` field on `transports.udp.*` and
|
||||
`transports.tcp.*` for explicit advertise-as override. Accepts
|
||||
either a bare IP (`"198.51.100.1"` — the configured `bind_addr`
|
||||
port is appended) or a full `host:port`
|
||||
(`"198.51.100.1:8443"`). Takes precedence over both the bound
|
||||
address and any STUN-derived autodiscovery. Required for TCP
|
||||
on cloud-NAT setups (AWS EIP, GCP/Azure external IPs) where
|
||||
binding to the public IP directly fails with `EADDRNOTAVAIL`
|
||||
(the EIP isn't on a host interface). Optional but useful for
|
||||
UDP as a deterministic alternative to STUN — operators who
|
||||
want to skip STUN egress (or whose STUN is blocked) can
|
||||
specify it explicitly. Without `external_addr`, TCP with a
|
||||
wildcard `bind_addr` + `advertise_on_nostr: true` now logs a
|
||||
loud `warn!` pointing at the two fixes instead of silently
|
||||
skipping
|
||||
- Nostr-discovery now tolerates ±60s of clock skew on offer/answer
|
||||
freshness checks so a responder whose wall clock leads the
|
||||
initiator's by less than that no longer silently rejects every
|
||||
offer. Previously, a public-test daemon with un-NTP'd peers (or
|
||||
long uptime — `now_ms()` anchors to `SystemTime` once at startup,
|
||||
then advances monotonically; post-startup NTP step adjustments
|
||||
don't propagate) would see ~100% signal-timeout rate against
|
||||
skewed peers, indistinguishable from "peer is offline." New
|
||||
optional `offerReceivedAt` field on the answer payload lets the
|
||||
initiator log per-peer NTP-style skew estimates (DEBUG when ≥30s)
|
||||
for operator visibility. Backward-compatible — older responders
|
||||
that don't fill the field still produce valid answers
|
||||
- Nostr-discovery NAT-traversal failure suppression: per-npub
|
||||
consecutive-failure counter triggers a 30-min extended cooldown
|
||||
after 5 failures, preventing the daemon from hammering Nostr
|
||||
relays with offers to peers that have gone away. WARN log lines
|
||||
rate-limited to one per peer per 5 min (subsequent failures
|
||||
emit DEBUG with `consecutive_failures` + remaining `cooldown_secs`).
|
||||
Threshold-crossing also fires a one-shot active re-check of the
|
||||
peer's Kind 37195 advert against `advert_relays`; absent →
|
||||
evict cache; newer → refresh + reset streak; same → cooldown
|
||||
stands. New `failure_streak_threshold`, `extended_cooldown_secs`,
|
||||
`warn_log_interval_secs`, `failure_state_max_entries` config
|
||||
fields under `node.discovery.nostr`. Per-peer state visible in
|
||||
`fipsctl show peers` JSON under `nostr_traversal`
|
||||
- Tor onion adverts published over Nostr overlay discovery now
|
||||
include the public-facing port (`<onion>.onion:<port>`) instead of
|
||||
just the bare onion hostname. The publisher previously emitted a
|
||||
bare onion that the parser refused (`expected host:port`),
|
||||
producing a persistent retry-fail loop on any peer whose Tor
|
||||
advert was the only entry in the discovery cache. New
|
||||
`transports.tor.advertised_port` config field (default `443`,
|
||||
matching the Tor `HiddenServicePort` convention) controls the
|
||||
advertised port; operators with non-default virtual ports can
|
||||
override.
|
||||
- TCP-over-FIPS reliability on mesh paths with mixed transport
|
||||
MTUs (e.g. a UDP-1280 hop in the picker set) improved. Three
|
||||
interlocking changes: `Node::transport_mtu()` is now deterministic
|
||||
across restarts (min across operational transports rather than
|
||||
insertion-order-dependent); the TCP MSS clamp at the TUN boundary
|
||||
reads per-destination path MTU instead of a single global ceiling;
|
||||
and reactive `MtuExceeded` from forwarders is mirrored back into
|
||||
the TUN-side `path_mtu_lookup` so later flows pick up forward-path
|
||||
bottlenecks without re-discovery. Windows TUN reader receives the
|
||||
same per-destination plumbing.
|
||||
- Proactive end-to-end `PathMtuNotification` now mirrors into the
|
||||
TUN-side `path_mtu_lookup` (TCP MSS clamp store), parallel to the
|
||||
reactive `MtuExceeded` mirror that already existed. Previously the
|
||||
proactive handler only updated the session-canonical
|
||||
`MmpSessionState.path_mtu`; on stable long-lived paths where the
|
||||
destination's echo had tightened the session MTU but no transit
|
||||
router had emitted a fresh `MtuExceeded` (because all current
|
||||
traffic was already sized by the tighter session value), new TCP
|
||||
flows opened in that window kept getting clamped by the staler
|
||||
discovery-time value. The proactive mirror closes that gap with
|
||||
the same tighter-only semantics — never loosens the clamp.
|
||||
- Nostr-discovered peers running an FMP-protocol version we cannot
|
||||
speak no longer trigger an indefinite retraversal storm. Open-
|
||||
discovery NAT-traversal succeeds at the UDP layer regardless of
|
||||
protocol version, so the daemon would adopt the punched socket,
|
||||
drop every incoming packet at `Unknown FMP version`, idle out
|
||||
after 31s, and re-fire the full STUN-offer-answer-punch sequence
|
||||
~30s later — every minute, forever, against peers the handshake
|
||||
literally cannot complete with. The rx loop now detects mismatched-
|
||||
version packets arriving on adopted bootstrap transports, reverse-
|
||||
maps to the originating npub, and applies a long structural
|
||||
cooldown to the discovery layer's `failure_state` so the next
|
||||
open-discovery sweep skips the peer until either side upgrades.
|
||||
One-shot WARN per fresh observation; subsequent mismatches inside
|
||||
the cooldown window are silent. New `protocol_mismatch_cooldown_secs`
|
||||
config field under `node.discovery.nostr` (default 86400 = 24h),
|
||||
separate from the transient-failure `extended_cooldown_secs`.
|
||||
- `fipstop` now uses `ratatui::try_init()` instead of `ratatui::init()`,
|
||||
so terminal initialization failures (e.g. Docker on macOS Sequoia,
|
||||
or environments without a usable tty) produce a clean error message
|
||||
instead of a hard crash
|
||||
- Spanning-tree updates that change only the internal path between
|
||||
root and leaf — without changing the root or the depth — now
|
||||
propagate to leaves correctly. Previously a leaf could continue
|
||||
routing against a stale internal path until the parent or depth
|
||||
also changed.
|
||||
|
||||
## [0.2.1] - 2026-05-11
|
||||
|
||||
### Added
|
||||
|
||||
- Linux release artifact workflow: builds x86_64 and aarch64 tarballs
|
||||
and `.deb` packages on `v*` tag push, with SHA-256 checksums
|
||||
- AUR publish workflow for tagged stable releases
|
||||
|
||||
### Changed
|
||||
|
||||
- Validate bloom filter fill ratio on FilterAnnounce ingress.
|
||||
Inbound FilterAnnounce messages whose derived false-positive
|
||||
rate exceeds `node.bloom.max_inbound_fpr` (new config field,
|
||||
default 0.05) are rejected silently on the wire, logged at WARN,
|
||||
and counted in a new `bloom.fill_exceeded` counter. A
|
||||
rate-limited WARN also fires if our own outgoing filter's FPR
|
||||
exceeds the cap. `BloomFilter::estimated_count` now takes
|
||||
`max_fpr` and returns `Option<f64>`, returning `None` for
|
||||
saturated filters; this propagates through `compute_mesh_size`
|
||||
into `estimated_mesh_size` (already `Option<u64>`)
|
||||
|
||||
### Fixed
|
||||
|
||||
- Control socket path detection in fipsctl and fipstop now checks for
|
||||
the `/run/fips/` directory instead of the socket file inside it, so
|
||||
users not yet in the `fips` group get a clear "Permission denied"
|
||||
error instead of a misleading "No such file" fallback to
|
||||
`$XDG_RUNTIME_DIR` ([#30](https://github.com/jmcorgan/fips/issues/30),
|
||||
reported by [@Sebastix](https://github.com/Sebastix))
|
||||
- OpenWrt ipk build excluded BLE feature that requires D-Bus, which is
|
||||
unavailable on OpenWrt targets
|
||||
- IPv6 routing policy rule added at TUN setup to protect `fd00::/8`
|
||||
from interception by Tailscale's table 52 default route
|
||||
- Bloom filter routing no longer swallows traffic when no bloom
|
||||
candidate is strictly closer than the current node. `find_next_hop`
|
||||
now falls through to greedy tree routing in that case instead of
|
||||
returning `NoRoute`, which previously caused dropped packets in
|
||||
topologies where the tree parent was closer but not a bloom
|
||||
candidate
|
||||
- Auto-connect peers now reconnect after a graceful `Disconnect`
|
||||
notification from the remote side. `handle_disconnect` previously
|
||||
removed the peer without scheduling a reconnect, orphaning the
|
||||
entry on a clean upstream shutdown; the other removal paths
|
||||
(link-dead, decrypt failure, peer restart) already scheduled
|
||||
reconnect ([#60](https://github.com/jmcorgan/fips/issues/60),
|
||||
reported by [@SwapMarket](https://github.com/SwapMarket))
|
||||
- `fipsctl connect` now rejects FIPS mesh (`fd00::/8`) addresses for
|
||||
`udp`, `tcp`, and `ethernet` transports with a clear error message
|
||||
instead of echoing success while the daemon silently failed the
|
||||
bind with `EAFNOSUPPORT`
|
||||
([#61](https://github.com/jmcorgan/fips/issues/61),
|
||||
reported by [@SwapMarket](https://github.com/SwapMarket))
|
||||
- Tighten TreeAnnounce ancestry validation to match the spanning
|
||||
tree specification. The receive path now verifies that the
|
||||
ancestry is structurally consistent with the signed parent
|
||||
declaration before mutating tree state.
|
||||
- Make the tree ancestry acceptance unit test deterministic.
|
||||
`test_tree_announce_validate_semantics_accepts_valid_non_root`
|
||||
generated a random signing identity while pinning the fixed root
|
||||
to `node_addr[0] = 0x01`; about 2 in 256 random identities were
|
||||
numerically smaller than the claimed root, triggering
|
||||
`AncestryRootNotMinimum`. The test now regenerates the identity
|
||||
until its `node_addr` is strictly larger than both the fixed
|
||||
parent and root.
|
||||
|
||||
## [0.2.0] - 2026-03-22
|
||||
|
||||
### Added
|
||||
|
||||
#### Operator Tooling
|
||||
|
||||
- `fipsctl connect` and `disconnect` commands for runtime peer
|
||||
management via control socket, with hostname resolution from
|
||||
`/etc/fips/hosts`
|
||||
|
||||
#### IPv6 Adapter
|
||||
|
||||
- Pre-seed identity cache from configured peer npubs at startup, so TUN packets can be dispatched immediately without waiting for handshake completion ([@v0l](https://github.com/v0l))
|
||||
|
||||
#### Mesh Peer Transports
|
||||
|
||||
- New Tor transport with SOCKS5 and directory-mode onion service for anonymous inbound and outbound peering
|
||||
- DNS hostname support in peer addresses for UDP and TCP transports
|
||||
- Non-blocking transport connect for connection-oriented transports (TCP, Tor)
|
||||
|
||||
#### Packaging and Deployment
|
||||
|
||||
- Reproducible build infrastructure: Rust toolchain pinning via
|
||||
`rust-toolchain.toml`, `SOURCE_DATE_EPOCH` in CI and packaging
|
||||
scripts, deterministic archive timestamps
|
||||
- Top-level packaging Makefile for unified build across formats
|
||||
- Kubernetes sidecar deployment example with Nostr relay demo
|
||||
- Nostr release publishing in OpenWrt package workflow
|
||||
- SHA-256 hash output in CI build and OpenWrt workflows
|
||||
|
||||
#### Testing and CI
|
||||
|
||||
- Maelstrom chaos scenario with dynamic topology mutation and
|
||||
ephemeral node identities via connect/disconnect commands
|
||||
- Consolidated Docker test harness infrastructure
|
||||
|
||||
### Changed
|
||||
|
||||
- Discovery protocol: replace flooding with bloom-filter-guided tree
|
||||
routing. Includes originator retry (T=0/T=5s/T=10s), exponential
|
||||
backoff after timeouts and bloom misses, and transit-side per-target
|
||||
rate limiting. Removed 257-byte visited bloom filter from LookupRequest wire format. *This is a breaking change; nodes running versions prior to this release will not be compatible.*
|
||||
|
||||
### Fixed
|
||||
|
||||
- DNS responder returned NXDOMAIN for A queries on valid `.fips` names,
|
||||
causing resolvers to give up without trying AAAA. Now returns NOERROR
|
||||
with empty answers for non-AAAA queries on resolvable names.
|
||||
(#9, reported by [@alopatindev](https://github.com/alopatindev))
|
||||
- Stale end-to-end session left in session table after peer removal blocked session re-establishment on reconnect — `remove_active_peer` now cleans up `self.sessions` and `self.pending_tun_packets`. (#5, [@v0l](https://github.com/v0l))
|
||||
- `schedule_reconnect` reset exponential backoff to zero on each link-dead
|
||||
cycle instead of preserving accumulated retry count.
|
||||
(#5, [@v0l](https://github.com/v0l))
|
||||
- FMP/FSP rekey dual-initiation race on high-latency links (Tor): both
|
||||
sides' timers fired simultaneously, both msg1s crossed in flight, each
|
||||
side's responder path destroyed the initiator state. Fixed with
|
||||
deterministic tie-breaker (smaller NodeAddr wins as initiator).
|
||||
- Parent selection SRTT gate bypass: `evaluate_parent` used default cost
|
||||
1.0 for peers filtered out by `has_srtt()`, defeating the MMP eligibility
|
||||
gate. Now skips unmeasured candidates when any peer has cost data.
|
||||
- FSP rekey cutover race: initiator cut over before responder received msg3,
|
||||
causing AEAD failures. Fixed by deferring initiator cutover by 2 seconds.
|
||||
- MMP metric discontinuity after rekey: receiver state carried stale
|
||||
counters across rekey, inflating reorder counts and jitter. Fixed via
|
||||
`reset_for_rekey()`.
|
||||
- Auto-connect peers exhausted `max_retries` on initial connection failures
|
||||
and were permanently abandoned. Now retry indefinitely with exponential
|
||||
backoff capped at 300 seconds.
|
||||
- Control socket permissions: non-root users couldn't connect. Daemon now
|
||||
chowns socket and directory to `root:fips` group at bind time.
|
||||
- Post-rekey jitter spikes: old-session frames arriving via the drain window
|
||||
produced 2,000–7,000ms jitter spikes that corrupted the EWMA estimator.
|
||||
Added a 15-second grace period after rekey cutover that suppresses jitter
|
||||
updates until drain-window frames have flushed. (#10)
|
||||
- ICMPv6 Packet Too Big source was set to the local FIPS address, which
|
||||
Linux ignores (loopback PTB check). Now uses the original packet's
|
||||
destination so the kernel honors the PMTU update.
|
||||
(#16, [@v0l](https://github.com/v0l))
|
||||
- Reverse delivery ratio used lifetime cumulative counters instead of
|
||||
per-interval deltas, making ETX unresponsive to recent loss. (#14)
|
||||
- MMP delta guards used `prev_rr > 0` to detect first report, conflating
|
||||
it with a legitimate zero counter. Replaced with `has_prev_rr`. (#14)
|
||||
|
||||
## [0.1.0] - 2026-03-12
|
||||
|
||||
|
||||
@@ -2,17 +2,65 @@
|
||||
|
||||
## Getting Started
|
||||
|
||||
Clone the repo and verify your setup:
|
||||
Clone the repo:
|
||||
|
||||
```
|
||||
git clone https://github.com/jmcorgan/fips.git
|
||||
cd fips
|
||||
cargo build
|
||||
cargo test
|
||||
```
|
||||
|
||||
Read [docs/design/](docs/design/) for protocol understanding, starting with
|
||||
[fips-intro.md](docs/design/fips-intro.md).
|
||||
Before changing code, read the protocol docs in this order:
|
||||
|
||||
- [docs/design/README.md](docs/design/README.md)
|
||||
- [docs/design/fips-intro.md](docs/design/fips-intro.md)
|
||||
- the specific design doc for the behavior you are touching
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- Rust 1.94.1 and Linux with TUN support
|
||||
- Use the pinned toolchain from [rust-toolchain.toml](rust-toolchain.toml) for deterministic builds
|
||||
- For the default BLE-enabled build on Debian/Ubuntu:
|
||||
`sudo apt install bluez libdbus-1-dev pkg-config`
|
||||
- Docker is required for the integration harnesses under [testing/](testing/)
|
||||
|
||||
If you do not want BLE locally, build and test without default features:
|
||||
|
||||
```bash
|
||||
cargo build --no-default-features --features tui
|
||||
cargo test --no-default-features --features tui
|
||||
```
|
||||
|
||||
## Local Verification
|
||||
|
||||
Choose the narrowest check that matches your change:
|
||||
|
||||
- Docs-only changes:
|
||||
|
||||
```bash
|
||||
git diff --check
|
||||
```
|
||||
|
||||
- Normal code changes:
|
||||
|
||||
```bash
|
||||
cargo build
|
||||
cargo test
|
||||
cargo clippy --all -- -D warnings
|
||||
```
|
||||
|
||||
- Local CI-style unit test run:
|
||||
|
||||
```bash
|
||||
./testing/ci-local.sh --test-only
|
||||
```
|
||||
|
||||
- Narrow integration run for transport, routing, Docker, or packaging-sensitive changes:
|
||||
|
||||
```bash
|
||||
./testing/ci-local.sh --only static-mesh
|
||||
```
|
||||
|
||||
See [testing/README.md](testing/README.md) for the available integration and chaos harnesses.
|
||||
|
||||
## Filing Issues
|
||||
|
||||
@@ -22,11 +70,24 @@ Read [docs/design/](docs/design/) for protocol understanding, starting with
|
||||
|
||||
## Pull Requests
|
||||
|
||||
- All PRs must pass `cargo build`, `cargo test`, and `cargo clippy` with no
|
||||
warnings.
|
||||
- All PRs must pass `cargo build`, `cargo test`, and `cargo clippy --all -- -D warnings`.
|
||||
- Keep commits focused — one logical change per commit.
|
||||
- Add tests for new functionality.
|
||||
- Reference relevant design docs if the change touches protocol behavior.
|
||||
- Pull requests are merged via squash-merge.
|
||||
- Update docs in the same change when you modify:
|
||||
- protocol or routing behavior
|
||||
- wire formats
|
||||
- configuration shape or defaults
|
||||
- operational workflows or testing instructions
|
||||
|
||||
In practice this usually means updating one or more of:
|
||||
|
||||
- [docs/design/fips-mesh-operation.md](docs/design/fips-mesh-operation.md)
|
||||
- [docs/design/fips-wire-formats.md](docs/design/fips-wire-formats.md)
|
||||
- [docs/design/fips-configuration.md](docs/design/fips-configuration.md)
|
||||
- [README.md](README.md)
|
||||
- [testing/README.md](testing/README.md)
|
||||
|
||||
## Questions
|
||||
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
[package]
|
||||
name = "fips"
|
||||
version = "0.1.0"
|
||||
version = "0.3.0"
|
||||
edition = "2024"
|
||||
description = "A distributed, decentralized network routing protocol for mesh nodes connecting over arbitrary transports"
|
||||
license = "MIT"
|
||||
@@ -8,17 +8,13 @@ authors = ["Johnathan Corgan <jcorgan@corganlabs.com>"]
|
||||
repository = "https://github.com/jmcorgan/fips"
|
||||
readme = "README.md"
|
||||
|
||||
[features]
|
||||
default = ["tui"]
|
||||
tui = ["dep:ratatui"]
|
||||
|
||||
[dependencies]
|
||||
ratatui = { version = "0.30", optional = true }
|
||||
ratatui = "0.30"
|
||||
secp256k1 = { version = "0.30", features = ["rand", "global-context"] }
|
||||
sha2 = "0.10"
|
||||
hkdf = "0.12"
|
||||
chacha20poly1305 = "0.10"
|
||||
rand = "0.10.0"
|
||||
ring = "0.17"
|
||||
rand = "0.10.1"
|
||||
thiserror = "2.0"
|
||||
bech32 = "0.11"
|
||||
serde = { version = "1.0", features = ["derive"] }
|
||||
@@ -26,16 +22,35 @@ serde_json = "1.0"
|
||||
serde_yaml = "0.9"
|
||||
dirs = "6.0"
|
||||
hex = "0.4"
|
||||
clap = { version = "4.5", features = ["derive"] }
|
||||
clap = { version = "4.6", features = ["derive"] }
|
||||
tracing = "0.1"
|
||||
tracing-subscriber = { version = "0.3", features = ["env-filter"] }
|
||||
tun = { version = "0.8.5", features = ["async"] }
|
||||
libc = "0.2"
|
||||
rtnetlink = "0.20.0"
|
||||
tokio = { version = "1", features = ["rt", "macros", "signal", "sync", "net", "time"] }
|
||||
tokio = { version = "1", features = ["rt", "macros", "signal", "sync", "net", "time", "process", "io-util"] }
|
||||
futures = "0.3"
|
||||
simple-dns = "0.11.2"
|
||||
socket2 = { version = "0.6.2", features = ["all"] }
|
||||
tokio-socks = "0.5"
|
||||
portable-atomic = { version = "1", features = ["std"] }
|
||||
|
||||
nostr = { version = "0.44", features = ["std", "nip59"] }
|
||||
nostr-sdk = "0.44"
|
||||
|
||||
[target.'cfg(unix)'.dependencies]
|
||||
tun = { version = "0.8.7", features = ["async"] }
|
||||
libc = "0.2"
|
||||
|
||||
[target.'cfg(target_os = "linux")'.dependencies]
|
||||
rtnetlink = "0.21.0"
|
||||
rustables = "0.8.7"
|
||||
procfs = { version = "0.18", default-features = false }
|
||||
|
||||
# bluer/BlueZ needs glibc — see build.rs `bluer_available` cfg gate.
|
||||
[target.'cfg(all(target_os = "linux", not(target_env = "musl")))'.dependencies]
|
||||
bluer = { version = "0.17", features = ["bluetoothd", "l2cap"] }
|
||||
|
||||
[target.'cfg(windows)'.dependencies]
|
||||
wintun = "0.5"
|
||||
windows-service = "0.8.1"
|
||||
|
||||
[package.metadata.deb]
|
||||
maintainer = "Johnathan Corgan <jcorgan@corganlabs.com>"
|
||||
@@ -43,12 +58,14 @@ copyright = "2026 Johnathan Corgan"
|
||||
license-file = ["LICENSE", "0"]
|
||||
section = "net"
|
||||
priority = "optional"
|
||||
depends = "libc6, systemd"
|
||||
depends = "libc6, systemd, libdbus-1-3"
|
||||
recommends = "bluez"
|
||||
extended-description = """\
|
||||
FIPS is a distributed, decentralized network routing protocol for mesh \
|
||||
nodes connecting over arbitrary transports including UDP, TCP, and Ethernet. \
|
||||
It provides encrypted peer-to-peer connectivity with automatic key management, \
|
||||
TUN-based virtual networking, and .fips DNS resolution."""
|
||||
nodes connecting over arbitrary transports including UDP, TCP, Ethernet, \
|
||||
Tor, and Bluetooth (BLE). It provides encrypted peer-to-peer connectivity \
|
||||
with automatic key management, TUN-based virtual networking, and .fips DNS \
|
||||
resolution."""
|
||||
maintainer-scripts = "packaging/debian/"
|
||||
assets = [
|
||||
["target/release/fips", "/usr/bin/", "755"],
|
||||
@@ -56,21 +73,32 @@ assets = [
|
||||
["target/release/fipstop", "/usr/bin/", "755"],
|
||||
["packaging/common/fips.yaml", "/etc/fips/fips.yaml", "600"],
|
||||
["packaging/common/hosts", "/etc/fips/hosts", "644"],
|
||||
["packaging/common/fips.nft", "/etc/fips/fips.nft", "644"],
|
||||
["packaging/debian/fips.service", "/lib/systemd/system/fips.service", "644"],
|
||||
["packaging/debian/fips-dns.service", "/lib/systemd/system/fips-dns.service", "644"],
|
||||
["packaging/debian/fips-firewall.service", "/lib/systemd/system/fips-firewall.service", "644"],
|
||||
["packaging/common/fips-dns-setup", "/usr/lib/fips/fips-dns-setup", "755"],
|
||||
["packaging/common/fips-dns-teardown", "/usr/lib/fips/fips-dns-teardown", "755"],
|
||||
["packaging/debian/fips.tmpfiles", "/usr/lib/tmpfiles.d/fips.conf", "644"],
|
||||
["target/release/fips-gateway", "/usr/bin/", "755"],
|
||||
["packaging/debian/fips-gateway.service", "/lib/systemd/system/fips-gateway.service", "644"],
|
||||
["docs/design/fips-security.md", "/usr/share/doc/fips/fips-security.md", "644"],
|
||||
]
|
||||
conf-files = ["/etc/fips/fips.yaml", "/etc/fips/hosts"]
|
||||
conf-files = ["/etc/fips/fips.yaml", "/etc/fips/hosts", "/etc/fips/fips.nft"]
|
||||
|
||||
[dev-dependencies]
|
||||
tempfile = "3.15"
|
||||
criterion = { version = "0.8.2", features = ["html_reports"] }
|
||||
tokio = { version = "1", features = ["test-util"] }
|
||||
|
||||
[[bin]]
|
||||
name = "fipsctl"
|
||||
path = "src/bin/fipsctl.rs"
|
||||
|
||||
[[bin]]
|
||||
name = "fips-gateway"
|
||||
path = "src/bin/fips-gateway.rs"
|
||||
|
||||
[[bin]]
|
||||
name = "fipstop"
|
||||
path = "src/bin/fipstop/main.rs"
|
||||
required-features = ["tui"]
|
||||
|
||||
@@ -1,276 +1,241 @@
|
||||
# FIPS: Free Internetworking Peering System
|
||||
|
||||

|
||||
[](LICENSE)
|
||||
[](https://www.rust-lang.org/)
|
||||
[-yellow.svg)](#status--roadmap)
|
||||
[](#status--roadmap)
|
||||
|
||||
A distributed, decentralized network routing protocol for mesh nodes
|
||||
connecting over arbitrary transports.
|
||||
A self-organizing encrypted mesh network built on Nostr identities,
|
||||
capable of operating over arbitrary transports without central
|
||||
infrastructure.
|
||||
|
||||
> **Status: Alpha (0.1.0)**
|
||||
>
|
||||
> FIPS is under active development. The protocol and APIs are not stable.
|
||||
> Expect breaking changes. See [Status & Roadmap](#status--roadmap) below.
|
||||
> FIPS is under active development. The protocol and APIs are not
|
||||
> yet stable. See [Status & roadmap](#status--roadmap) below.
|
||||
|
||||
## Overview
|
||||
## What FIPS does
|
||||
|
||||
FIPS is a self-organizing mesh network that operates natively over a variety
|
||||
of physical and logical media — local area networks, Bluetooth, serial links,
|
||||
radio, or the existing internet as an overlay. Nodes generate their own
|
||||
identities, discover each other, and route traffic without any central
|
||||
authority or global topology knowledge.
|
||||
A machine running FIPS becomes a node in the mesh with a
|
||||
self-generated cryptographic identity (a Nostr keypair). There are
|
||||
two equally-supported deployment modes.
|
||||
|
||||
FIPS uses Nostr keypairs (secp256k1/schnorr) as native node identities,
|
||||
allowing users to generate their own persistent or ephemeral node addresses.
|
||||
Nodes address each other by npub, and the same cryptographic identity serves
|
||||
as both the routing address and the basis for end-to-end encrypted sessions
|
||||
across the mesh.
|
||||
**As an overlay** on top of existing IP networks, FIPS lets your
|
||||
node reach any other FIPS node wherever it sits — behind a NAT, on
|
||||
a different ISP, on a phone over cellular, on a laptop with only
|
||||
Bluetooth in range, or behind a Tor onion. The mesh forwards IPv6
|
||||
traffic transparently and end-to-end encrypted, with no central VPN
|
||||
concentrator or coordinating server.
|
||||
|
||||
FIPS allows existing TCP/IP based network software to use the FIPS mesh
|
||||
network by generating a local IP address from the node npub and tunnelling
|
||||
IP packets to other endpoints transparently knowing only their npub. Native
|
||||
FIPS-aware applications do not need this IP tunneling or emulation capability.
|
||||
**Ground up** over raw Ethernet, WiFi, or Bluetooth, FIPS provides
|
||||
a complete permissionless network without any pre-existing IP
|
||||
infrastructure, ISP, or DNS. Any node that joins the link gets
|
||||
routable IPv6 addresses, peer discovery, and a path to every other
|
||||
node automatically.
|
||||
|
||||
All traffic over the FIPS mesh is encrypted and authenticated both
|
||||
hop-to-hop between peers and independently end-to-end between FIPS
|
||||
endpoints.
|
||||
Either way, existing networking software runs over it unchanged —
|
||||
SSH, HTTP servers, file transfer, anything IPv6-native works the
|
||||
same way it would on a local network.
|
||||
|
||||
## Features
|
||||
|
||||
- **Self-organizing mesh routing** — spanning tree coordinates and bloom
|
||||
filter candidate selection, no global routing tables
|
||||
- **Multi-transport** — UDP, TCP, and Ethernet today; designed for
|
||||
Bluetooth, serial, radio, and Tor
|
||||
- **Noise encryption** — hop-by-hop link encryption plus independent
|
||||
end-to-end session encryption, with periodic rekey for forward secrecy
|
||||
- **Nostr-native identity** — secp256k1 keypairs as node addresses, no
|
||||
registration or central authority
|
||||
- **IPv6 adaptation** — TUN interface maps npubs to fd00::/8 addresses for
|
||||
unmodified IP applications; static hostname mapping (`/etc/fips/hosts`)
|
||||
- **Metrics Measurement Protocol** — per-link RTT, loss, jitter, and goodput
|
||||
measurement
|
||||
- **ECN congestion signaling** — hop-by-hop CE flag relay with RFC 3168 IPv6
|
||||
marking, transport kernel drop detection
|
||||
- **Operator visibility** — `fipsctl` CLI and `fipstop` TUI dashboard for
|
||||
runtime inspection of peers, links, sessions, tree state, and metrics
|
||||
- **Zero configuration** — sensible defaults; a node can start with no config
|
||||
file, though peer addresses are needed to join a network
|
||||
- **Self-organizing mesh routing.** Spanning-tree coordinates with
|
||||
bloom-filter-guided discovery; no global routing tables, no
|
||||
flooding.
|
||||
- **Multi-transport.** UDP, TCP, Ethernet, Tor, and Bluetooth (BLE
|
||||
L2CAP) ship today; transports compose on a single mesh and a
|
||||
node may run several at once.
|
||||
- **Two-layer encryption.** Noise IK between peers (hop-by-hop) and
|
||||
Noise XK between mesh endpoints (independent end-to-end), with
|
||||
periodic rekey for forward secrecy.
|
||||
- **Nostr-native identity.** secp256k1 / schnorr keypairs as node
|
||||
addresses; self-generated, no registration, no central authority.
|
||||
- **IPv6 adapter.** A TUN interface maps each remote npub to an
|
||||
`fd00::/8` address, so unmodified IPv6 software reaches mesh
|
||||
peers as `<npub>.fips`. Built-in `.fips` DNS resolver, with
|
||||
optional static name mapping via `/etc/fips/hosts`.
|
||||
- **Nostr-mediated discovery and NAT traversal.** Peers publish
|
||||
endpoint adverts on public Nostr relays, exchange candidates via
|
||||
NIP-59 gift-wrapped offers and answers, and establish direct
|
||||
paths through NATs using STUN-assisted hole punching.
|
||||
- **LAN gateway.** Optional `fips-gateway` service folds an entire
|
||||
unmodified LAN into the mesh: outbound (LAN clients reach mesh
|
||||
destinations through a DNS-allocated virtual IPv6 pool and
|
||||
nftables NAT) and inbound (LAN-side services exposed to the mesh
|
||||
through 1:1 port forwards).
|
||||
- **Per-link metrics.** RTT, loss, jitter, and goodput on every
|
||||
hop, plus mesh-size estimation, via the Metrics Measurement
|
||||
Protocol.
|
||||
- **ECN congestion signaling.** Hop-by-hop CE-flag relay with RFC
|
||||
3168 IPv6 marking and transport kernel-drop detection.
|
||||
- **Mesh-interface security baseline.** Optional default-deny
|
||||
nftables policy for `fips0` shipped as a packaged conffile
|
||||
(`/etc/fips/fips.nft`) with an operator drop-in directory
|
||||
(`/etc/fips/fips.d/`) and a disabled-by-default
|
||||
`fips-firewall.service`. The baseline polices only the mesh
|
||||
interface, leaving Docker, Tor, and the host firewall untouched.
|
||||
- **Operator visibility.** `fipsctl` CLI for control and inspection
|
||||
with time-series stats history queryable for any metric,
|
||||
`fipstop` TUI for live status with inline sparkline dashboards,
|
||||
and a JSON-line control socket on each binary for direct
|
||||
programmatic access.
|
||||
- **Reproducible builds** with toolchain pinning and
|
||||
`SOURCE_DATE_EPOCH`.
|
||||
|
||||
## Building
|
||||
## Quick start
|
||||
|
||||
The shortest path on Debian / Ubuntu:
|
||||
|
||||
```bash
|
||||
git clone https://github.com/jmcorgan/fips.git
|
||||
cd fips
|
||||
cargo build --release
|
||||
```
|
||||
|
||||
Requires Rust 1.85+ (edition 2024) and Linux with TUN support.
|
||||
|
||||
## Installation
|
||||
|
||||
After building, choose one of the following methods to install.
|
||||
|
||||
### Debian / Ubuntu (.deb)
|
||||
|
||||
Requires [cargo-deb](https://crates.io/crates/cargo-deb):
|
||||
|
||||
```bash
|
||||
cargo install cargo-deb
|
||||
cargo deb
|
||||
sudo dpkg -i target/debian/fips_*.deb
|
||||
```
|
||||
|
||||
This installs the daemon, CLI tools, systemd units, and a default
|
||||
configuration. Edit `/etc/fips/fips.yaml` before starting:
|
||||
|
||||
```bash
|
||||
sudo nano /etc/fips/fips.yaml
|
||||
sudo systemctl start fips
|
||||
```
|
||||
|
||||
The service is enabled at boot automatically. To use `fipsctl` and
|
||||
`fipstop` without sudo, add your user to the `fips` group:
|
||||
This installs the daemon, CLI tools (`fipsctl`, `fipstop`), the
|
||||
optional `fips-gateway` service, systemd units, and a default
|
||||
`/etc/fips/fips.yaml` you can edit before starting.
|
||||
|
||||
For macOS, Windows, OpenWrt, the systemd tarball, or a from-source
|
||||
build, see [docs/getting-started.md](docs/getting-started.md) for
|
||||
the full multi-platform installation guide.
|
||||
|
||||
To join a live mesh and reach your first peer, follow the new-user
|
||||
tutorial progression starting at
|
||||
[docs/tutorials/join-the-test-mesh.md](docs/tutorials/join-the-test-mesh.md).
|
||||
|
||||
### Building from source
|
||||
|
||||
```bash
|
||||
sudo usermod -aG fips $USER # log out and back in to take effect
|
||||
cargo build --release
|
||||
```
|
||||
|
||||
Remove with `sudo dpkg -r fips` (preserves config) or
|
||||
`sudo dpkg -P fips` (removes everything including identity keys).
|
||||
Requires Rust 1.94.1+ (edition 2024). Linux, macOS, and Windows are
|
||||
supported; transport availability varies by platform.
|
||||
|
||||
### Generic Linux (systemd tarball)
|
||||
| Transport | Linux | macOS | Windows | OpenWrt |
|
||||
|-----------|:-----:|:-----:|:-------:|:-------:|
|
||||
| UDP | ✅ | ✅ | ✅ | ✅ |
|
||||
| TCP | ✅ | ✅ | ✅ | ✅ |
|
||||
| Ethernet | ✅ | ✅ | ❌ | ✅ |
|
||||
| Tor | ✅ | ✅ | ✅ | ✅ |
|
||||
| BLE | ✅ | ❌ | ❌ | ❌ |
|
||||
|
||||
```bash
|
||||
./packaging/systemd/build-tarball.sh
|
||||
tar xzf deploy/fips-*-linux-*.tar.gz
|
||||
cd fips-*-linux-*/
|
||||
sudo ./install.sh
|
||||
```
|
||||
|
||||
See [packaging/systemd/README.install.md](packaging/systemd/README.install.md)
|
||||
for the full installation and configuration guide.
|
||||
|
||||
## Configuration
|
||||
|
||||
The default configuration file is installed at `/etc/fips/fips.yaml`:
|
||||
|
||||
```yaml
|
||||
# FIPS Node Configuration
|
||||
|
||||
node:
|
||||
identity:
|
||||
# By default, a new ephemeral keypair is generated on each start.
|
||||
# Uncomment persistent to keep the same identity across restarts;
|
||||
# on first start a keypair is saved to fips.key/fips.pub next to
|
||||
# this config file (mode 0600/0644).
|
||||
# persistent: true
|
||||
#
|
||||
# Or set an explicit key (overrides persistent):
|
||||
# nsec: "nsec1..."
|
||||
|
||||
tun:
|
||||
enabled: true
|
||||
name: fips0
|
||||
mtu: 1280
|
||||
|
||||
dns:
|
||||
enabled: true
|
||||
bind_addr: "127.0.0.1"
|
||||
port: 5354
|
||||
|
||||
transports:
|
||||
udp:
|
||||
bind_addr: "0.0.0.0:2121"
|
||||
|
||||
tcp:
|
||||
# Accepts inbound connections. No static outbound peers.
|
||||
bind_addr: "0.0.0.0:8443"
|
||||
|
||||
# Ethernet transport — uncomment and set your interface name.
|
||||
# ethernet:
|
||||
# interface: "eth0"
|
||||
# discovery: true
|
||||
# announce: true
|
||||
# auto_connect: true
|
||||
# accept_connections: true
|
||||
|
||||
peers:
|
||||
# Static peers for bootstrapping (UDP or TCP):
|
||||
- npub: "npub1qmc3cvfz0yu2hx96nq3gp55zdan2qclealn7xshgr448d3nh6lks7zel98"
|
||||
alias: "fips-test-node"
|
||||
addresses:
|
||||
- transport: udp
|
||||
addr: "217.77.8.91:2121"
|
||||
connect_policy: auto_connect
|
||||
```
|
||||
|
||||
See [docs/design/fips-configuration.md](docs/design/fips-configuration.md)
|
||||
for the full reference.
|
||||
|
||||
## Usage
|
||||
|
||||
### DNS Resolution
|
||||
|
||||
FIPS includes a DNS resolver (enabled by default, port 5354) that maps
|
||||
`.fips` names to fd00::/8 IPv6 addresses. With systemd-resolved:
|
||||
|
||||
```bash
|
||||
sudo resolvectl dns fips0 127.0.0.1:5354
|
||||
sudo resolvectl domain fips0 ~fips
|
||||
```
|
||||
|
||||
Then reach any FIPS node by npub with standard IPv6 tools:
|
||||
|
||||
```bash
|
||||
ping6 npub1bbb....fips
|
||||
ssh npub1bbb....fips
|
||||
```
|
||||
|
||||
### Monitoring
|
||||
|
||||
Use `fipsctl` to query a running node:
|
||||
|
||||
```bash
|
||||
fipsctl show status # Node status overview
|
||||
fipsctl show peers # Authenticated peers
|
||||
fipsctl show links # Active links
|
||||
fipsctl show tree # Spanning tree state
|
||||
fipsctl show sessions # End-to-end sessions
|
||||
fipsctl show transports # Transport instances
|
||||
fipsctl show routing # Routing table summary
|
||||
```
|
||||
|
||||
`fipstop` provides an interactive TUI dashboard with live-updating
|
||||
views of node status, peers, links, sessions, tree state, transports,
|
||||
and routing:
|
||||
|
||||
```bash
|
||||
fipstop # connect to local daemon
|
||||
fipstop -r 1 # 1-second refresh interval
|
||||
```
|
||||
|
||||
### Service Management
|
||||
|
||||
```bash
|
||||
sudo systemctl start fips
|
||||
sudo systemctl stop fips
|
||||
sudo systemctl restart fips
|
||||
sudo journalctl -u fips -f
|
||||
```
|
||||
|
||||
### Testing
|
||||
|
||||
See [testing/](testing/) for Docker-based integration test harnesses
|
||||
including static topology tests and stochastic chaos simulation.
|
||||
On Linux, BLE requires BlueZ and libdbus
|
||||
(`sudo apt install bluez libdbus-1-dev` on Debian / Ubuntu) and is
|
||||
gated on a build-script probe — install the dependencies first and
|
||||
the `cargo build` line above picks it up. The OpenWrt ipk omits
|
||||
BLE because libdbus is not available on the target.
|
||||
|
||||
## Documentation
|
||||
|
||||
Protocol design documentation is in [docs/design/](docs/design/), organized as
|
||||
a layered protocol specification. Start with
|
||||
[fips-intro.md](docs/design/fips-intro.md) for the full protocol overview.
|
||||
`docs/` is organised by reader purpose:
|
||||
|
||||
## Project Structure
|
||||
- **[Tutorials](docs/tutorials/)** — hand-held walk-throughs from
|
||||
a fresh install through to a participating mesh node, plus
|
||||
advanced deployments (gateway on OpenWrt, hosting services,
|
||||
ground-up two-device mesh).
|
||||
- **[How-to guides](docs/how-to/)** — operator recipes for
|
||||
specific tasks: firewall activation, Nostr discovery, Tor onion
|
||||
service, Bluetooth peering, LAN gateway deployment and
|
||||
troubleshooting, MTU diagnostics, host aliases, persistent
|
||||
identity, unprivileged-user setup, UDP buffer tuning.
|
||||
- **[Reference](docs/reference/)** — `fips.yaml` configuration,
|
||||
wire formats, control-socket protocol, CLI references for each
|
||||
binary, security posture matrix, Nostr events catalog, transport
|
||||
statistics inventory.
|
||||
- **[Design](docs/design/)** — protocol-level architecture and
|
||||
layer specifications. Start with
|
||||
[fips-concepts.md](docs/design/fips-concepts.md) for the framing,
|
||||
then [fips-architecture.md](docs/design/fips-architecture.md) for
|
||||
the protocol stack.
|
||||
|
||||
```
|
||||
src/ Rust source (library + fips/fipsctl/fipstop binaries)
|
||||
packaging/ Debian, systemd tarball, and shared packaging files
|
||||
docs/design/ Protocol design specifications
|
||||
testing/ Docker-based integration test harnesses
|
||||
If you want to contribute, see [CONTRIBUTING.md](CONTRIBUTING.md)
|
||||
and [testing/README.md](testing/README.md).
|
||||
|
||||
## Examples
|
||||
|
||||
- **[examples/sidecar-nostr-relay/](examples/sidecar-nostr-relay/)** —
|
||||
Run a [strfry](https://github.com/hoytech/strfry) Nostr relay
|
||||
reachable exclusively over the FIPS mesh. The relay container
|
||||
shares the FIPS sidecar's network namespace and is isolated from
|
||||
the host network.
|
||||
- **[examples/k8s-sidecar/](examples/k8s-sidecar/)** — Run FIPS as
|
||||
a Kubernetes Pod sidecar. The sidecar creates `fips0` in the
|
||||
Pod's shared network namespace so every other container in the
|
||||
Pod gets mesh access without modification.
|
||||
- **[examples/wireguard-sidecar-macos/](examples/wireguard-sidecar-macos/)** —
|
||||
Reach the FIPS mesh from a macOS host through a local Docker
|
||||
container over a WireGuard tunnel. Only traffic destined for
|
||||
`fd00::/8` transits the sidecar; regular internet traffic
|
||||
continues to use the host network.
|
||||
|
||||
## Project structure
|
||||
|
||||
```text
|
||||
src/ Rust source: library + fips, fipsctl, fipstop, fips-gateway binaries
|
||||
docs/ Documentation: tutorials, how-to, reference, design
|
||||
packaging/ Debian, macOS .pkg, Windows ZIP, OpenWrt ipk, AUR, systemd tarball
|
||||
examples/ Deployment examples (Nostr relay, K8s sidecar, macOS WireGuard)
|
||||
testing/ Docker-based integration test harnesses + chaos simulation
|
||||
```
|
||||
|
||||
## Status & Roadmap
|
||||
## Status & roadmap
|
||||
|
||||
FIPS is at **v0.1.0 (alpha)**. The core protocol works end-to-end over
|
||||
UDP, TCP, and Ethernet but has not been tested beyond small meshes.
|
||||
FIPS is at **v0.3.0**. The core protocol works end-to-end over
|
||||
UDP, TCP, Ethernet, Tor, and Bluetooth on a small live mesh of
|
||||
deployed nodes. v0.3.0 is the testing-and-polishing track for
|
||||
everything accumulated since v0.2.0 on the v0.2.x wire format —
|
||||
Nostr-mediated peer discovery, UDP NAT traversal, peer ACL, the
|
||||
DNS-responder fix, packaging hardening, and discovery rate-limit
|
||||
retuning. New wire-format work is staged on the `next` branch for
|
||||
the post-v0.3.0 release line.
|
||||
|
||||
### What works today
|
||||
|
||||
- Spanning tree construction with greedy coordinate routing
|
||||
- Bloom filter discovery for finding nodes without global state
|
||||
- Noise IK (link layer) and Noise XK (session layer) encryption
|
||||
- Periodic Noise rekey with forward secrecy (FMP + FSP)
|
||||
- Persistent node identity with key file management
|
||||
- IPv6 TUN adapter with DNS resolution of `.fips` names
|
||||
- Static hostname mapping (`/etc/fips/hosts`) with auto-reload
|
||||
- Per-link metrics (RTT, loss, jitter, goodput) and mesh size estimation
|
||||
- ECN congestion signaling (hop-by-hop CE relay, IPv6 CE marking, kernel drop detection)
|
||||
- UDP, TCP, and Ethernet transports
|
||||
- Runtime inspection via `fipsctl` and `fipstop`
|
||||
- Docker-based integration and chaos testing
|
||||
- Spanning-tree construction with greedy coordinate routing.
|
||||
- Bloom-filter-guided destination discovery (no flooding,
|
||||
single-path with retry).
|
||||
- Two-layer Noise encryption (IK at the link, XK at the session)
|
||||
with periodic hitless rekey for forward secrecy at both layers.
|
||||
- Persistent or ephemeral node identity with key-file management.
|
||||
- IPv6 TUN adapter with built-in `.fips` DNS resolver and
|
||||
multi-backend auto-configuration (systemd dns-delegate,
|
||||
systemd-resolved, dnsmasq, NetworkManager).
|
||||
- Static hostname mapping (`/etc/fips/hosts`) with auto-reload.
|
||||
- Per-link metrics (RTT, loss, jitter, goodput) and mesh size
|
||||
estimation.
|
||||
- ECN congestion signaling (hop-by-hop CE relay, IPv6 CE marking,
|
||||
kernel-drop detection).
|
||||
- UDP, TCP, Ethernet, Tor, and BLE transports (BLE via L2CAP CoC
|
||||
with per-link MTU negotiation).
|
||||
- Nostr-mediated overlay endpoint discovery and UDP hole punching
|
||||
for NAT traversal.
|
||||
- LAN gateway (`fips-gateway`) with both outbound (LAN-to-mesh)
|
||||
and inbound (mesh-to-LAN port-forwarding) modes.
|
||||
- Peer ACL: per-npub allow / deny admission control at the link
|
||||
layer; opt-in mesh-firewall baseline at `fips0` ingress.
|
||||
- Runtime inspection and peer management via `fipsctl` and
|
||||
`fipstop`.
|
||||
- Reproducible builds with toolchain pinning and
|
||||
`SOURCE_DATE_EPOCH`.
|
||||
- Linux (Debian, systemd tarball, OpenWrt, AUR), macOS (`.pkg`),
|
||||
and Windows (ZIP, service) packaging.
|
||||
- Docker-based integration and chaos testing.
|
||||
|
||||
### Near-term priorities
|
||||
|
||||
- Peer discovery via Nostr relays (bootstrap without static peer lists)
|
||||
- Additional transports (Bluetooth, Tor)
|
||||
- Improved routing resilience under churn
|
||||
- Security audit of cryptographic protocols
|
||||
- Native API for FIPS-aware applications (npub:port addressing
|
||||
without the IPv6-shim path).
|
||||
- Security audit of the cryptographic protocols.
|
||||
|
||||
### Longer-term
|
||||
|
||||
- Mobile platform support
|
||||
- Bandwidth-aware routing and QoS
|
||||
- Protocol stability and versioned wire format
|
||||
- Published crate
|
||||
- Mobile platform support.
|
||||
- Bandwidth-aware routing and QoS.
|
||||
- Protocol stability and a versioned wire format.
|
||||
- Published crate.
|
||||
|
||||
## License
|
||||
|
||||
|
||||
@@ -0,0 +1,764 @@
|
||||
# FIPS v0.3.0
|
||||
|
||||
**Released**: 2026-05-11
|
||||
|
||||
v0.3.0 is the testing-and-polishing release on the v0.2.x wire format.
|
||||
It widens the platform reach of FIPS from Linux-only to Linux, macOS,
|
||||
Windows, and OpenWrt; adds two large new mesh capabilities (Nostr-mediated
|
||||
peer discovery with UDP NAT traversal, and the `fips-gateway` LAN bridge);
|
||||
ships a default-deny security baseline for the mesh interface; introduces
|
||||
mesh-peer access control; substantially speeds up session-layer crypto and
|
||||
the Linux receive path; and tightens packaging across every supported
|
||||
distribution channel.
|
||||
|
||||
v0.3.0 is wire-compatible with v0.2.x. Mixed meshes interoperate; there
|
||||
is no flag-day upgrade.
|
||||
|
||||
v0.3.0 also rolls forward all changes from the v0.2.1 maintenance
|
||||
release. The sections below cover the cumulative v0.2.0 → v0.3.0
|
||||
delta; the per-section intros call out which entries first shipped
|
||||
in v0.2.1.
|
||||
|
||||
## At a glance
|
||||
|
||||
- 123 commits since v0.2.0 (109 non-merge), spanning 307 files with
|
||||
+44,186 / -4,078 lines.
|
||||
- 10 committers plus 3 issue reporters across feature work, fixes,
|
||||
packaging, and reviews.
|
||||
- 5 new GitHub Actions CI workflows (Linux Package, macOS Package,
|
||||
Windows Package, OpenWrt Package, AUR Publish) plus expanded
|
||||
integration matrices (gateway, NAT-cone, NAT-symmetric, NAT-LAN,
|
||||
rekey-accept-off, `.deb` install across Debian 12/13 + Ubuntu
|
||||
22/24/26, multi-backend `.fips` DNS resolver across the same five
|
||||
distros).
|
||||
- The long-standing systemd-resolved DNS-responder silent-drop is
|
||||
closed end-to-end.
|
||||
- Pre-1.0 control-socket JSON schema change for two query fields;
|
||||
see [Upgrade notes](#upgrade-notes).
|
||||
|
||||
## What's new
|
||||
|
||||
### Mesh discovery and NAT traversal
|
||||
|
||||
Previously, two FIPS nodes could only become peers if they had a way
|
||||
to find each other beforehand: a configured address, a shared LAN
|
||||
segment, or a Bluetooth radio range. v0.3.0 introduces a Nostr-based
|
||||
overlay-discovery channel that lets nodes find each other through any
|
||||
public Nostr relay set, plus a STUN-assisted UDP hole-punching path
|
||||
that connects peers across most consumer NATs.
|
||||
|
||||
Each participating node publishes a signed overlay advert as a Nostr
|
||||
**Kind 37195** parameterized replaceable event. (The kind sits in the
|
||||
application-defined replaceable range and the digits visually spell
|
||||
*FIPS*: 7=F, 1=I, 9=P, 5=S.) The advert lists reachable transport
|
||||
endpoints (UDP, TCP, Tor) and is consumed by other nodes to populate
|
||||
fallback addresses for `via_nostr` peers. Under `policy: open`, the
|
||||
advert cache is also dialed for non-configured peers within a budget
|
||||
cap.
|
||||
|
||||
When both peers are behind NAT, the daemon coordinates a UDP hole
|
||||
punch using NIP-59 gift-wrap signaling for the offer/answer exchange
|
||||
and STUN for reflexive address discovery. A candidate-pair punch
|
||||
planner attempts LAN-private and reflexive paths in parallel; on
|
||||
success the live socket is handed into the standard FIPS UDP transport
|
||||
via a bootstrap-handoff API.
|
||||
|
||||
Operators turn this on with `node.discovery.nostr.enabled: true` and
|
||||
the configured relay set. `policy: open` adds best-effort dialing of
|
||||
non-configured peers seen on the relays. New `peers[].via_nostr` and
|
||||
per-transport `advertise_on_nostr` / `public` flags control what each
|
||||
endpoint contributes to the published advert. Cross-field validation
|
||||
runs at startup to catch mis-configured combinations early.
|
||||
|
||||
A Docker NAT lab covering cone, symmetric, and LAN scenarios is wired
|
||||
into the integration CI matrix. A daemon-side failure-suppression
|
||||
layer (per-npub cooldown after consecutive failures, ±60s clock-skew
|
||||
tolerance, rate-limited WARN logs) keeps relay traffic well-mannered
|
||||
when peers come and go from the open discovery cache. A separate
|
||||
structural cooldown (`protocol_mismatch_cooldown_secs`, default 24h)
|
||||
suppresses retraversal when a punched peer turns out to be running an
|
||||
FMP version this daemon cannot handshake with: the punch completes at
|
||||
the UDP layer, the rx loop spots the version-mismatched packet,
|
||||
reverse-maps to the originating npub, and removes the peer from the
|
||||
next sweep until either side upgrades.
|
||||
|
||||
The auto-connect retry loop pins itself to relay ground truth. Each
|
||||
retry attempt refetches the cached overlay advert against the
|
||||
configured `advert_relays` (one filter query, 2s timeout) before
|
||||
dialing, so a peer whose NAT rebound to a fresh endpoint is recovered
|
||||
on the next retry rather than looping on a stale cached address.
|
||||
`NoTransportForType` triggers a fire-and-forget re-fetch that either
|
||||
replaces or evicts the cache entry. A startup peer-init failure (no
|
||||
operational transport, all addresses unreachable) now schedules a
|
||||
retry instead of leaving the peer in a dead state until the daemon is
|
||||
restarted. Adopted NAT-traversed UDP transports inherit the operator's
|
||||
primary `[transports.udp]` listener config (MTU, recv/send buffer
|
||||
sizes) instead of falling back to the 1280 IPv6-minimum default.
|
||||
|
||||
### Cross-platform reach
|
||||
|
||||
FIPS now ships first-class binaries for **Linux, macOS, Windows, and
|
||||
OpenWrt**.
|
||||
|
||||
- **macOS** support uses the native `utun` TUN interface, raw
|
||||
Ethernet via BPF, a `.pkg` installer with a launchd plist and
|
||||
uninstall script, and an x86_64 cross-compile from arm64 build
|
||||
hosts. A new CI matrix entry runs build and unit-test jobs on
|
||||
macOS hosts.
|
||||
- **Windows** support uses [wintun](https://www.wintun.net/) for the
|
||||
TUN device, a TCP control socket on `localhost:21210` (replacing
|
||||
the Unix domain socket Linux and macOS use), Windows Service
|
||||
lifecycle (`fips.exe --install-service`, `--uninstall-service`,
|
||||
`--service`), and a ZIP package with PowerShell install/uninstall
|
||||
scripts.
|
||||
- **MIPS** atomic-ABI portability lets the daemon build for 32-bit
|
||||
MIPS targets (`mips`, `mipsel`, MIPS32r2) by routing through
|
||||
`portable_atomic`. This unblocks OpenWrt deployments on
|
||||
consumer-grade MIPS routers.
|
||||
- **OpenWrt** packaging gets a procd init with dnsmasq forwarding,
|
||||
proxy NDP, RA route advertisements, and IPv6 forwarding sysctls.
|
||||
The `fips-gateway` is enabled by default in the OpenWrt build.
|
||||
|
||||
### FIPS gateway
|
||||
|
||||
The new `fips-gateway` binary lets unmodified LAN hosts reach FIPS
|
||||
mesh destinations without running the FIPS daemon themselves. Two
|
||||
flows ship together:
|
||||
|
||||
- **Outbound (LAN -> mesh)**: a virtual-IP pool (default
|
||||
`fd01::/112`) is allocated on demand from `.fips`-name DNS lookups.
|
||||
A state-machine lifecycle, conntrack-backed session tracking, proxy
|
||||
NDP on the LAN interface, and TTL-based reclamation handle the
|
||||
bookkeeping. A LAN host that resolves `peer.fips` gets a virtual
|
||||
address it can reach over IP, and the gateway translates the flow
|
||||
to the mesh.
|
||||
- **Inbound (mesh -> LAN)**: new `gateway.port_forwards` config
|
||||
installs prerouting DNAT rules so mesh peers can reach a configured
|
||||
`host:port` on the gateway's LAN. A LAN-side masquerade is added
|
||||
automatically when any forwards are configured, so replies flow
|
||||
back through conntrack.
|
||||
|
||||
A dedicated control socket at `/run/fips/gateway.sock` exposes
|
||||
`show_gateway` and `show_mappings`. `fipstop` adds a Gateway tab with
|
||||
a pool gauge and mappings table.
|
||||
|
||||
The gateway's `dns.listen` source default is now `[::1]:5353`,
|
||||
matching the canonical deployment model: the gateway sits on a host
|
||||
already serving DHCP and DNS to a LAN segment (an OpenWrt AP, a Linux
|
||||
router), port 53 there is taken by the existing resolver, and `.fips`
|
||||
queries are forwarded to the gateway over loopback. The OpenWrt ipk
|
||||
previously overrode the prior `[::]:53` source default in its packaged
|
||||
config; that override is now redundant and has been dropped.
|
||||
Operators on a host without a pre-existing resolver on port 53 can
|
||||
opt back into the wildcard bind by setting `dns.listen: "[::]:53"`
|
||||
explicitly. The new default binds IPv6 loopback only, so forwarders
|
||||
that reach the gateway over IPv4 loopback need an explicit IPv4
|
||||
listen address.
|
||||
|
||||
The cold-boot startup race between `fips.service` and
|
||||
`fips-gateway.service` is handled by a systemd `After=fips.service`
|
||||
ordering, an `ExecStartPre` poll loop that waits up to 30 seconds for
|
||||
the `fips0` interface to appear, and a DNS upstream probe in the
|
||||
gateway itself that retries up to 5 times with 1-second backoff.
|
||||
|
||||
Packaging covers systemd, Debian, AUR, and OpenWrt. The full design
|
||||
is in [`docs/design/fips-gateway.md`](../design/fips-gateway.md).
|
||||
|
||||
### Mesh-interface security baseline
|
||||
|
||||
The FIPS mesh is a flat layer-3 segment. Every authenticated peer can
|
||||
route packets to every other peer's `fips0` address. Peer identity is
|
||||
authenticated end-to-end by the FMP and FSP Noise handshakes, but
|
||||
identity is not authorization. A service on a mesh host that binds to
|
||||
a wildcard address is, by default, reachable from every peer in the
|
||||
mesh.
|
||||
|
||||
v0.3.0 ships an opt-in default-deny baseline that closes this gap on
|
||||
Linux:
|
||||
|
||||
- **`/etc/fips/fips.nft`** is installed as a documented operator
|
||||
conffile. It defines a single `inet fips` nftables table with one
|
||||
chain hooked at `input`, default-denies inbound traffic on
|
||||
`fips0`, and is a no-op for every other interface.
|
||||
- **`fips-firewall.service`** loads it. The unit ships **disabled by
|
||||
default**; activation is an explicit
|
||||
`systemctl enable --now fips-firewall.service`.
|
||||
- Per-service allowances live in **`/etc/fips/fips.d/*.nft`**
|
||||
drop-ins that the baseline includes.
|
||||
|
||||
Choosing opt-in keeps the mesh quick to bring up for evaluation while
|
||||
giving operators a documented, packaged path to lock it down for
|
||||
production. The full design (threat model, rule layout, conntrack
|
||||
handling, drop-in mechanism, and the rationale for a conffile rather
|
||||
than an auto-loaded package side-effect) is in
|
||||
[`docs/design/fips-security.md`](../design/fips-security.md).
|
||||
|
||||
`fipstop`'s Node tab gains a **"Listening on fips0" panel** that
|
||||
surfaces the answer to the operational question "what services on
|
||||
this host are reachable from the mesh, and what does the firewall
|
||||
currently say about each of them?" The panel lists every IPv6
|
||||
listening socket bound to either the wildcard address or this node's
|
||||
`fd00::/8` address, paired with its classification against the
|
||||
running `inet fips` baseline chain: `OPEN` (canonical accept rule),
|
||||
`filt` (falls through to drop), or `filt?` (referenced with matchers
|
||||
the panel cannot fully decompose, e.g. saddr filters or jumps). When
|
||||
`fips-firewall.service` is inactive, a yellow banner above the table
|
||||
reminds the operator that every listener is mesh-exposed; wildcard
|
||||
binds carry a trailing `*` in the Process column. The classifier is
|
||||
built on a new `show_listening_sockets` control query (Linux-only),
|
||||
which is also useful from `fipsctl` for scripting.
|
||||
|
||||
### Peer access control
|
||||
|
||||
Operators can now restrict which mesh peers a node will form direct
|
||||
links with. Optional `/etc/fips/peers.allow` and `/etc/fips/peers.deny`
|
||||
files (TCP-Wrappers style) match against npub, hex pubkey, host
|
||||
alias, or `ALL`. Enforcement runs at three points:
|
||||
|
||||
1. Outbound connect (before dialing).
|
||||
2. Inbound msg1 (the first FMP handshake message from a new peer).
|
||||
3. Outbound msg2 (the response).
|
||||
|
||||
Files reload automatically on mtime change; a new `fipsctl acl show`
|
||||
query reports the effective rule set. A six-node Docker integration
|
||||
harness (`testing/acl/`) exercises allowlist and denylist patterns
|
||||
end-to-end.
|
||||
|
||||
**Important scope distinction**: peer ACLs are an FMP-layer
|
||||
restriction. They control who can establish a *direct link* with this
|
||||
node. They do **not** control session-layer (FSP) reachability through
|
||||
the mesh. A node that denies peer X with an ACL can still receive FSP
|
||||
traffic from X relayed via other peers.
|
||||
|
||||
### Bluetooth Low Energy transport (experimental, Linux)
|
||||
|
||||
A new BLE L2CAP Connection-Oriented Channel transport lets FIPS nodes
|
||||
peer over Bluetooth Low Energy without any IP infrastructure in
|
||||
between. The transport handles per-link MTU negotiation, continuous
|
||||
scan/probe peer discovery with cooldown-based deduplication,
|
||||
continuous advertising, deterministic NodeAddr cross-probe
|
||||
tie-breaker, and a configurable connection pool with eviction.
|
||||
|
||||
This transport is **experimental in v0.3.0**. It is implemented and
|
||||
functional on Linux (BlueZ via `bluer`), but the reliability follow-up
|
||||
logic (probe cooldown, cross-probe tie-breaker, pubkey timeout,
|
||||
continuous advertising semantics, probe-promotion, fail-fast send) is
|
||||
not yet behaviorally tested in CI. Its maturity path is field-driven;
|
||||
please file issues with field reports. macOS BLE support is in
|
||||
development as a separate track and is not part of v0.3.0.
|
||||
|
||||
### UDP transport profiles
|
||||
|
||||
The UDP transport gains posture flags organized around deployment
|
||||
patterns:
|
||||
|
||||
- **Public-facing inbound nodes**: `bind_addr: "0.0.0.0:2121"`,
|
||||
`accept_connections: true` (default), `public: true` for advert
|
||||
publication. v0.3.0 adds STUN-based public-IP autodiscovery so
|
||||
cloud nodes (AWS EIP, GCP, Azure 1:1 NAT) advertise the right
|
||||
address even when the public IP isn't on a host interface.
|
||||
- **Ephemeral leaf nodes**: `outbound_only: true` binds an ephemeral
|
||||
port (`0.0.0.0:0`), refuses inbound msg1, and is never advertised
|
||||
on Nostr regardless of `advertise_on_nostr`. Use this for client
|
||||
postures that should connect outbound only, without exposing an
|
||||
inbound listener on a known port.
|
||||
- **General-purpose nodes**: `accept_connections: false` mirrors the
|
||||
Ethernet/BLE knob without changing the bind address. The Node-level
|
||||
handshake gate carves out msg1 from peers already established on
|
||||
this transport so rekey continues to work.
|
||||
|
||||
Startup validation now rejects `bind_addr` set to a loopback address
|
||||
when at least one peer has a non-loopback UDP address, closing a
|
||||
silent-failure trap from v0.2.0 where Linux's source-address routing
|
||||
check would drop outbound flows from the loopback-bound socket.
|
||||
|
||||
A new `external_addr` field on `transports.udp.*` and
|
||||
`transports.tcp.*` lets operators specify the advertise-as address
|
||||
explicitly. This is useful for UDP as a deterministic alternative to
|
||||
STUN, and required for TCP on cloud-NAT setups (where binding to the
|
||||
public IP fails with `EADDRNOTAVAIL` because the IP isn't on a host
|
||||
interface).
|
||||
|
||||
### `.fips` DNS resolver overhaul
|
||||
|
||||
The IPv6 adapter's `.fips` name resolution has been rebuilt around
|
||||
the constraints of contemporary systemd-based hosts. The default
|
||||
`dns.bind_addr` is now `::1` (IPv6 loopback), and a setup script
|
||||
picks one of five backends in priority order:
|
||||
|
||||
1. systemd-resolved global drop-in
|
||||
(`/etc/systemd/resolved.conf.d/fips.conf`).
|
||||
2. systemd dns-delegate (per-link configuration handed off to
|
||||
systemd-resolved).
|
||||
3. `resolvectl` per-link configuration.
|
||||
4. Standalone `dnsmasq`.
|
||||
5. NetworkManager's dnsmasq plugin.
|
||||
|
||||
Teardown reverses only what setup applied, recorded in a state file
|
||||
at `/run/fips/dns-backend`. A new `testing/dns-resolver/` harness
|
||||
exercises every backend across Debian 12, Debian 13, Ubuntu 22.04,
|
||||
Ubuntu 24.04, and Ubuntu 26.04, so a regression in any of the five
|
||||
backends shows up in CI rather than in the field.
|
||||
|
||||
This overhaul resolves the long-standing silent-drop case where the
|
||||
`resolvectl dns fips0 [<fips0_addr>]:5354` target collided with the
|
||||
daemon's mesh-interface filter on certain systemd-resolved
|
||||
deployments (typically Ubuntu 22 with systemd 249's interface-scoped
|
||||
routing).
|
||||
|
||||
### Operator tooling additions
|
||||
|
||||
A handful of additions land in `fipsctl`, `fipstop`, and the daemon's
|
||||
configuration surface:
|
||||
|
||||
- **`node.log_level`** config field replaces the hardcoded
|
||||
`RUST_LOG=info` previously baked into systemd units and the
|
||||
OpenWrt procd init. The daemon now loads config before
|
||||
initializing tracing so the configured level takes effect.
|
||||
`RUST_LOG` still overrides when set.
|
||||
- **`fipsctl show identity-cache`** is a new query that lists every
|
||||
cached node identity (npub, IPv6 address, display name, LRU age)
|
||||
alongside the configured cache capacity.
|
||||
- **`fipsctl show peers / sessions / cache / routing`** are
|
||||
substantially extended: per-peer security signals (replay
|
||||
suppression count, consecutive decrypt failures), Noise session
|
||||
counters, session indices, rekey lifecycle state, handshake resend
|
||||
counts, K-bit epoch, coords-warmup remaining, drain state, per-peer
|
||||
retry state, per-target lookup detail (attempt, age, last sent),
|
||||
and pending TUN packet queue depth.
|
||||
- **Historical statistics**: in-memory time-series rings on the
|
||||
daemon (1-second × 3600 fast, 1-minute × 1440 slow) cover per-node
|
||||
and per-peer metrics. New `show_stats_*` control-socket queries, a
|
||||
`fipsctl stats list / peers / history` subcommand with Unicode
|
||||
sparkline rendering, and a `fipstop` Graphs tab with btop-style
|
||||
sparklines surface them to the operator.
|
||||
|
||||
### Performance
|
||||
|
||||
Two independent perf threads land in v0.3.0: a session-layer crypto
|
||||
backend swap, and a Linux receive-path overhaul.
|
||||
|
||||
**Session-layer crypto backend.** The ChaCha20-Poly1305 backend used
|
||||
by every FIPS Noise session (end-to-end FSP traffic and link-layer
|
||||
FMP traffic alike) has been swapped from RustCrypto's
|
||||
`chacha20poly1305` crate to `ring 0.17`. ring wraps BoringSSL's
|
||||
hand-tuned ChaCha20-Poly1305 implementation, which dispatches to NEON
|
||||
on aarch64 and AVX2 / AVX-512 on x86_64. Typical throughput is in the
|
||||
3-5 GB/s/core range, versus the ~600-800 MB/s/core RustCrypto soft
|
||||
path on the same hardware.
|
||||
|
||||
Wire format is unchanged. ChaCha20-Poly1305 is byte-deterministic for
|
||||
a given `(key, nonce, plaintext, aad)`, so any correct AEAD
|
||||
implementation produces identical ciphertext. A mixed mesh with some
|
||||
nodes pre-swap and some post-swap interoperates without protocol
|
||||
awareness; v0.3.0 can roll out across a mesh in any order.
|
||||
|
||||
Measurements on an aarch64 Apple Silicon docker target:
|
||||
|
||||
- Two-node TCP single-stream: 437 -> 1097 Mbps (about 2.5×).
|
||||
- Two-node UDP at 1000 Mbit: 599 Mbps with 40% loss -> lossless at
|
||||
line rate.
|
||||
- Three-node ping under bulk-traffic load: 7.68 ms avg / 215 ms max
|
||||
-> 0.72 ms / 3.6 ms max as the relay path stops being crypto-bound.
|
||||
|
||||
No operator-visible action is required; the swap is internal to the
|
||||
session layer.
|
||||
|
||||
**Linux UDP receive path.** The Linux UDP receive path now uses
|
||||
`recvmmsg(2)` with a 32-packet batch in place of single-packet
|
||||
`recvmsg(2)`. A single `readable()` wakeup drains up to 32 datagrams
|
||||
in one syscall before yielding back to the reactor, eliminating the
|
||||
per-packet scheduler-hop and futex cost that previously capped
|
||||
inbound rate at one event per scheduler quantum independent of CPU.
|
||||
`SO_RXQ_OVFL` is sampled once per batch and surfaced through
|
||||
`AsyncUdpSocket::recv_batch` so the existing 1Hz transport-congestion
|
||||
detector continues to feed the per-transport `dropping` flag. macOS
|
||||
and Windows fall through to the per-packet path; `recvmmsg` is
|
||||
Linux-specific.
|
||||
|
||||
**Inner rx-loop drain batching.** `Node::run_rx_loop` drains up to
|
||||
256 additional ready items via `try_recv()` after each
|
||||
`tokio::select!` await fires on the packet and TUN-outbound branches,
|
||||
in a tight inner loop before yielding. Previously the select cost a
|
||||
full scheduler hop and futex per packet, capping throughput at one
|
||||
event per scheduler quantum with the worker near-idle. `biased`
|
||||
ordering keeps data-plane branches priority over tick / control / DNS
|
||||
under sustained load; the 256 cap keeps the worker on a busy stream
|
||||
between yield points (about 400 KB of contiguous traffic) while still
|
||||
bounding the inner loop so a flood on one branch cannot starve the
|
||||
periodic tick or control socket.
|
||||
|
||||
**Eager `pubkey_full` precompute.** `PeerIdentity::pubkey_full()`
|
||||
precomputes the parity-aware full secp256k1 public key at
|
||||
construction in `from_pubkey`. Previously the method fell through to
|
||||
an EC point parse on every call when the full key wasn't passed at
|
||||
construction (i.e. for every peer constructed from an npub or x-only
|
||||
key), about 6% of per-packet CPU on the bulk-data send path for a
|
||||
value that never changed after construction. The same parse already
|
||||
runs at construction inside `NodeAddr::from_pubkey`, so the cost is
|
||||
paid once where it would be paid anyway.
|
||||
|
||||
These three changes are a coordinated set: the syscall batching
|
||||
removes the per-packet kernel cost, the inner-loop drain removes the
|
||||
per-packet scheduler cost, and the pubkey-cache change removes the
|
||||
per-packet crypto-derivation cost. Like the AEAD swap, they are all
|
||||
internal and require no operator action.
|
||||
|
||||
### Examples
|
||||
|
||||
- **macOS WireGuard companion** ([#51](https://github.com/jmcorgan/fips/pull/51)):
|
||||
run FIPS in a local Docker container and route `.fips` traffic
|
||||
from the macOS host through a WireGuard tunnel to the container's
|
||||
`fips0`. Only traffic destined for `fd00::/8` transits the
|
||||
companion; regular internet traffic continues to use the host
|
||||
network. Persistent FIPS and WireGuard key material is generated
|
||||
on first run.
|
||||
|
||||
### Documentation
|
||||
|
||||
- **`docs/design/port-advertisement-and-nat-traversal.md`**
|
||||
documents how nodes find each other through Nostr relays and the
|
||||
STUN-assisted UDP hole punch.
|
||||
- **`docs/design/fips-gateway.md`** documents the gateway's virtual
|
||||
IP pool, lifecycle, control surface, and packaging.
|
||||
- **`docs/design/fips-security.md`** documents the mesh-interface
|
||||
security posture, threat model, default-deny baseline, and drop-in
|
||||
workflow.
|
||||
- **`CONTRIBUTING.md`** has been expanded with build prerequisites,
|
||||
Rust toolchain setup, and first-build steps.
|
||||
|
||||
The `docs/` tree has been reorganized end-to-end into four sections
|
||||
(*tutorials / how-to / reference / design*) with a new
|
||||
[`docs/getting-started.md`](../getting-started.md) and per-section
|
||||
landing pages. Content was reconciled against current source:
|
||||
protocol-layer details, wire-format diagrams, configuration knobs,
|
||||
and CLI references were brought back into agreement with the
|
||||
implementation. See [Documentation pointers](#documentation-pointers)
|
||||
below for entry points by reader intent.
|
||||
|
||||
## Behavior changes worth flagging
|
||||
|
||||
These default-config changes affect every operator on upgrade, even
|
||||
those with no explicit configuration. Two items below — bloom-filter
|
||||
fill-ratio validation and TreeAnnounce ancestry validation — first
|
||||
shipped in v0.2.1 and roll forward into v0.3.0; the rest are
|
||||
v0.3.0-net-new.
|
||||
|
||||
- **Discovery rate-limiting** has been retuned to be less aggressive
|
||||
at cold start. v0.2.0 used a single-lookup-with-internal-retry
|
||||
model where a timed-out lookup during bloom-filter propagation
|
||||
could suppress retries for 30 seconds while none of the reset
|
||||
triggers fired on a stable post-handshake topology. v0.3.0
|
||||
replaces this with a per-attempt timeout sequence
|
||||
(`node.discovery.attempt_timeouts_secs`, default `[1, 2, 4, 8]`,
|
||||
15s total). Each attempt sends a fresh `LookupRequest` with a new
|
||||
`request_id`, letting successive attempts take different
|
||||
forwarding paths as the bloom and tree state evolve. Post-failure
|
||||
suppression is **off by default**; operators with chatty
|
||||
applications can opt back in via `backoff_base_secs` /
|
||||
`backoff_max_secs`.
|
||||
- **MMP report intervals** are retuned for constrained transports.
|
||||
The steady-state floor moves from 100ms to 1000ms, the ceiling
|
||||
from 2000ms to 5000ms, with a cold-start phase running 200ms for
|
||||
the first 5 SRTT samples. This reduces BLE overhead by roughly
|
||||
10× while keeping reports well above the EWMA convergence
|
||||
threshold. Session-layer MMP intervals are unchanged.
|
||||
- **Bloom filter fill-ratio validation** runs on every inbound
|
||||
`FilterAnnounce`. Filters whose derived false-positive rate
|
||||
exceeds `node.bloom.max_inbound_fpr` (default 0.05) are rejected
|
||||
silently on the wire, logged at WARN, and counted in a new
|
||||
`bloom.fill_exceeded` counter. A rate-limited WARN also fires
|
||||
when the local outgoing filter exceeds the cap.
|
||||
- **TreeAnnounce ancestry validation** is now run before tree-state
|
||||
mutation, enforcing ancestry-self-match, root-single-entry,
|
||||
parent-second-entry, and root-is-minimum-NodeAddr. Non-conforming
|
||||
announces are rejected with a WARN. Mixed v0.2.0 / v0.2.1 / v0.3.0
|
||||
meshes may produce WARN log lines on the v0.2.1+ side until all
|
||||
peers upgrade; behavior is correct, log noise only.
|
||||
- **Log noise reduction**: 35 info-level log messages have been
|
||||
demoted to debug (handshake cross-connection mechanics, periodic
|
||||
MMP telemetry, TUN/transport shutdown, retry scheduling). The
|
||||
default `RUST_LOG` in systemd units is now `info`, where it
|
||||
previously ran at `debug`. Operator-visible info output now
|
||||
focuses on lifecycle events, peer promotions, session
|
||||
establishment, parent switches, and transport start/stop.
|
||||
|
||||
## Notable bug fixes
|
||||
|
||||
These pre-existing v0.2.0 bugs are worth singling out because they
|
||||
either affected real-world deployments or produced misleading
|
||||
operator experiences. The CHANGELOG has the exhaustive list; this is
|
||||
the operator-relevant subset. Four items below first shipped in
|
||||
v0.2.1 and roll forward into v0.3.0: auto-connect Disconnect-reconnect,
|
||||
`fipsctl connect` mesh-address rejection, `fd00::/8` routing
|
||||
protection from Tailscale interception, and bloom-filter routing
|
||||
greedy-tree fallback. The control-socket path-detection fix landed
|
||||
in v0.2.1 as well, and the unified resolver below is the v0.3.0
|
||||
refactor that builds on it.
|
||||
|
||||
- **DNS responder silent-drop on systemd-resolved** is fixed: the
|
||||
responder no longer drops queries on Ubuntu 22 / Debian 13 and
|
||||
similar deployments where systemd applies interface-scoped
|
||||
routing. Default bind moves to `::1`; new global drop-in backend
|
||||
available ([#52](https://github.com/jmcorgan/fips/issues/52),
|
||||
[#77](https://github.com/jmcorgan/fips/issues/77)).
|
||||
- **Auto-connect peers reconnect after a graceful Disconnect.**
|
||||
Previously, a clean upstream shutdown left the auto-connect peer
|
||||
orphaned; only the link-dead, decrypt-fail, and peer-restart
|
||||
paths scheduled a reconnect
|
||||
([#60](https://github.com/jmcorgan/fips/issues/60), reported by
|
||||
[@SwapMarket](https://github.com/SwapMarket)).
|
||||
- **`fipsctl connect` rejects FIPS mesh addresses** (`fd00::/8`)
|
||||
for `udp`, `tcp`, and `ethernet` transports with a clear error
|
||||
message, instead of echoing success while the daemon silently
|
||||
failed the bind with `EAFNOSUPPORT`
|
||||
([#61](https://github.com/jmcorgan/fips/issues/61), reported by
|
||||
[@SwapMarket](https://github.com/SwapMarket)).
|
||||
- **Default control-socket path resolution unified.** Daemon and
|
||||
client tools now share a single resolver, eliminating a divergence
|
||||
where `fipsctl` / `fipstop` could connect to a socket the daemon
|
||||
never bound (notably on dev runs with `XDG_RUNTIME_DIR` set, or
|
||||
after a prior packaged install left a root-owned `/run/fips`
|
||||
behind). Canonical order is
|
||||
`/run/fips` -> `$XDG_RUNTIME_DIR/fips/` -> `/tmp/fips-<name>`. The
|
||||
`/run/fips` arm is selected by directory existence; the kernel
|
||||
enforces actual access at `connect(2)` time, so users not yet in
|
||||
the `fips` group get a clear `EACCES` rather than a silent path
|
||||
mismatch and a misleading `No such file` fallback to
|
||||
`$XDG_RUNTIME_DIR`. `XDG_RUNTIME_DIR` is validated as an existing
|
||||
directory before being used so stale post-logout values are
|
||||
treated as missing. The deployed fleet is unaffected: packaged
|
||||
configs set `node.control.socket_path` explicitly
|
||||
([#30](https://github.com/jmcorgan/fips/issues/30), reported by
|
||||
[@Sebastix](https://github.com/Sebastix)).
|
||||
- **`fd00::/8` routing protected from Tailscale interception.** The
|
||||
daemon installs an IPv6 routing-policy rule
|
||||
(`ip -6 rule to fd00::/8 lookup main priority 5265`) at TUN
|
||||
setup, so Tailscale's table 52 default route can no longer divert
|
||||
mesh traffic.
|
||||
- **TCP-over-FIPS reliability on mixed-MTU paths** is markedly
|
||||
improved. Four interlocking changes ship together:
|
||||
`Node::transport_mtu()` is now deterministic across daemon
|
||||
restarts (min across operational transports rather than
|
||||
insertion-order-dependent); the TCP MSS clamp at the TUN boundary
|
||||
reads per-destination path MTU instead of a single global ceiling;
|
||||
reactive `MtuExceeded` from forwarders is mirrored back into the
|
||||
TUN-side `path_mtu_lookup` so later flows pick up forward-path
|
||||
bottlenecks without re-discovery; and the proactive end-to-end
|
||||
`PathMtuNotification` echoed by the destination is mirrored into
|
||||
the same TUN-side store. Without that fourth piece, on long-lived
|
||||
stable paths where the destination's echo had tightened the
|
||||
session MTU but no transit router had emitted a fresh
|
||||
`MtuExceeded`, new TCP flows opened in that window were clamped by
|
||||
the staler discovery-time value. The proactive mirror uses the
|
||||
same tighter-only semantics as the reactive mirror, so it never
|
||||
loosens the clamp. The Windows TUN reader receives the same
|
||||
per-destination plumbing.
|
||||
- **Bloom filter routing greedy-tree fallback.** `find_next_hop` no
|
||||
longer returns `NoRoute` when the bloom candidate set is non-empty
|
||||
but no candidate is strictly closer than the current node; it
|
||||
falls through to greedy tree routing instead. Previously, this
|
||||
caused dropped packets in topologies where the tree parent was
|
||||
closer but not a bloom candidate.
|
||||
- **`fipstop` graceful tty-init failure.** `ratatui::try_init()`
|
||||
produces a clean error message instead of a hard crash when
|
||||
terminal initialization fails (Docker on macOS Sequoia, ttyless
|
||||
environments).
|
||||
- **TreeAnnounce ancestry on self-root transitions.** When a node
|
||||
had no smaller-NodeAddr peer to use as a parent, the spanning-tree
|
||||
state correctly promoted it to root, but the ancestry advertised
|
||||
on the next `TreeAnnounce` still referenced its previous parent's
|
||||
path. Receiving peers rejected the announce as
|
||||
`invalid ancestry: advertised root X is not the minimum path entry
|
||||
Y`, blocking mesh transit on any path that needed to traverse the
|
||||
node. The self-root transition is now detected explicitly in
|
||||
`TreeState::become_root` and the advertised ancestry rebuilt to
|
||||
start from self; the MMP receive handler corrects stale ancestry
|
||||
inherited across reconnect eagerly rather than waiting for the
|
||||
next observation tick.
|
||||
- **Spanning-tree internal-path updates** that change only the
|
||||
internal path between root and leaf (without changing the root or
|
||||
the depth) now propagate to leaves correctly. Previously, a leaf
|
||||
could continue routing against a stale internal path until the
|
||||
parent or depth also changed.
|
||||
|
||||
## Upgrade notes
|
||||
|
||||
Operator-actionable items when moving from v0.2.x to v0.3.0:
|
||||
|
||||
- **Control socket JSON schema (breaking, pre-1.0).**
|
||||
- `show_cache` response field `entries` has changed type from a
|
||||
`u64` count to an array of entry objects. The previous scalar
|
||||
value is now in a new `count` field.
|
||||
- `show_routing` response field `pending_lookups` has changed
|
||||
type from a `u64` count to an array of per-target lookup
|
||||
objects.
|
||||
- External tooling parsing these fields as numbers must be
|
||||
updated. In-tree `fipstop` is adjusted to the new schema. The
|
||||
control-socket interface remains pre-1.0 and is not covered by
|
||||
stability guarantees.
|
||||
|
||||
- **Cargo feature flags removed.** `tui`, `ble`, `gateway`, and
|
||||
`nostr-discovery` are gone. Subsystem inclusion is now driven by
|
||||
platform `cfg` gates, so plain `cargo build` compiles everything
|
||||
available on the target without `--features` invocations.
|
||||
Source-build tooling that passed any of these features should be
|
||||
updated to omit them.
|
||||
|
||||
- **Discovery rate-limiting defaults changed.** Post-failure
|
||||
suppression is **off by default**
|
||||
(`node.discovery.backoff_base_secs: 0`, `backoff_max_secs: 0`).
|
||||
Operators relying on the prior 30s base / 300s cap behavior must
|
||||
set those fields explicitly. The per-attempt sequence
|
||||
(`attempt_timeouts_secs`, default `[1, 2, 4, 8]`) now governs
|
||||
cold-start lookup behavior.
|
||||
|
||||
- **`.fips` DNS bind address default changed.** The default
|
||||
`dns.bind_addr` is now `::1`. Operators with explicit overrides
|
||||
of this field should review them; many existing overrides were
|
||||
workarounds for the silent-drop bug that this release fixes
|
||||
properly.
|
||||
|
||||
- **Gateway `dns.listen` source default changed.** The
|
||||
`fips-gateway` `dns.listen` default is now `[::1]:5353` (was
|
||||
`[::]:53`), matching the canonical deployment model where a
|
||||
pre-existing resolver on the host already owns port 53. The
|
||||
OpenWrt ipk previously overrode this in its packaged config; the
|
||||
override is now redundant and has been dropped. Operators on a
|
||||
host without a pre-existing resolver on port 53 can opt back into
|
||||
the wildcard bind by setting `dns.listen: "[::]:53"` explicitly.
|
||||
The new default binds IPv6 loopback only, so forwarders that
|
||||
reach the gateway over IPv4 loopback need an explicit IPv4 listen
|
||||
address.
|
||||
|
||||
- **systemd unit log level.** The shipped systemd units no longer
|
||||
hardcode `RUST_LOG=info`; the daemon's effective log level is
|
||||
driven by `node.log_level` (default `info`). `RUST_LOG`, when
|
||||
set, still overrides.
|
||||
|
||||
- **UDP transport `bind_addr` validation.** Startup now rejects a
|
||||
`bind_addr` set to a loopback address when at least one peer has
|
||||
a non-loopback UDP address. Operators who configured a loopback
|
||||
UDP bind as a workaround should switch to `outbound_only: true`
|
||||
for the same effect, plus the correct semantics (kernel-assigned
|
||||
ephemeral port, refuses inbound, never advertised).
|
||||
|
||||
- **Tor advert port.** If the Tor `HiddenServicePort` virtual port
|
||||
isn't 443, set `transports.tor.advertised_port` to match. The
|
||||
default is 443 and matches the conventional virtual-port choice.
|
||||
|
||||
## Documentation pointers
|
||||
|
||||
v0.3.0 ships a `docs/` tree reorganized into four sections
|
||||
(*tutorials / how-to / reference / design*). A new top-level
|
||||
[`docs/getting-started.md`](../getting-started.md) and per-section
|
||||
landing pages anchor the entry points.
|
||||
|
||||
Entry points by reader intent:
|
||||
|
||||
- **New users**: [`docs/getting-started.md`](../getting-started.md)
|
||||
and [`docs/tutorials/`](../tutorials/) cover guided introductions
|
||||
for bringing up your first node, joining the test mesh,
|
||||
advertising a node over Nostr, hosting a service, deploying a
|
||||
gateway, walking through the IPv6 adapter, and resolving peers
|
||||
via Nostr.
|
||||
- **Operators with a specific task**:
|
||||
[`docs/how-to/`](../how-to/) holds task-driven guides for enabling
|
||||
Nostr discovery, deploying the gateway, troubleshooting the
|
||||
gateway, deploying a Tor onion, hosting aliases, persistent
|
||||
identity, running unprivileged, setting up a Bluetooth peer,
|
||||
enabling the mesh firewall, tuning UDP buffers, and diagnosing
|
||||
MTU issues.
|
||||
- **Reference lookups**: [`docs/reference/`](../reference/) holds
|
||||
the config field reference, control-socket query reference, the
|
||||
`fips`, `fipsctl`, `fipstop`, and `fips-gateway` CLI references,
|
||||
and the protocol diagram set.
|
||||
- **Architectural background**: [`docs/design/`](../design/) holds
|
||||
design rationale for FIPS as a whole, FMP and FSP, the spanning
|
||||
tree, bloom-filter discovery, transports, the IPv6 adapter, the
|
||||
Nostr discovery layer, and the gateway.
|
||||
- **Security**: [`docs/design/fips-security.md`](../design/fips-security.md)
|
||||
documents the mesh-interface security baseline, threat model, and
|
||||
drop-in workflow.
|
||||
|
||||
## Getting v0.3.0
|
||||
|
||||
- **Linux x86_64 / aarch64**: `.deb` and tarball at the
|
||||
[v0.3.0 release page](https://github.com/jmcorgan/fips/releases/tag/v0.3.0).
|
||||
- **Arch Linux**: `fips` from the AUR.
|
||||
- **macOS**: `.pkg` at the v0.3.0 release page.
|
||||
- **Windows**: ZIP at the v0.3.0 release page.
|
||||
- **OpenWrt**: `.ipk` at the v0.3.0 release page.
|
||||
- **From source**: `cargo build --release` from a checkout of the
|
||||
v0.3.0 tag.
|
||||
|
||||
The full per-commit changelog lives in
|
||||
[`CHANGELOG.md`](../../CHANGELOG.md). Issues and discussion at
|
||||
[github.com/jmcorgan/fips](https://github.com/jmcorgan/fips).
|
||||
|
||||
## Contributors
|
||||
|
||||
Thanks to everyone who contributed code, packaging work, bug reports,
|
||||
or reviews to this release.
|
||||
|
||||
**Code and packaging**:
|
||||
|
||||
- [@jcorgan](https://github.com/jmcorgan): release shepherd, Nostr
|
||||
discovery / NAT traversal, `fips-gateway`, ACL infrastructure,
|
||||
packaging, security baseline, BLE follow-ups.
|
||||
- [@Origami74](https://github.com/Origami74): macOS platform support,
|
||||
from-source Docker companion build and `fipstop` terminal-init
|
||||
handling, gateway co-development, OpenWrt BLE-feature build fix,
|
||||
AUR-workflow follow-ups.
|
||||
- [@jodobear](https://github.com/jodobear): Linux release-artifact
|
||||
workflow and target-aware build scripts, CONTRIBUTING.md
|
||||
expansion, rekey integration-test stabilization.
|
||||
- [@tidley](https://github.com/tidley): Nostr-mediated overlay
|
||||
discovery and UDP NAT traversal
|
||||
([#53](https://github.com/jmcorgan/fips/pull/53)).
|
||||
- [@alexxie16](https://github.com/alexxie16): peer ACL enforcement
|
||||
([#50](https://github.com/jmcorgan/fips/pull/50)),
|
||||
macOS WireGuard companion example
|
||||
([#51](https://github.com/jmcorgan/fips/pull/51)),
|
||||
follow-up ([#67](https://github.com/jmcorgan/fips/pull/67)).
|
||||
- [@osh](https://github.com/osh): diagnostic queries for security
|
||||
validation and mesh debugging
|
||||
([#42](https://github.com/jmcorgan/fips/pull/42)).
|
||||
- [@OceanSlim](https://github.com/0ceanSlim): Windows platform
|
||||
support ([#45](https://github.com/jmcorgan/fips/pull/45)).
|
||||
- [@mmalmi](https://github.com/mmalmi): ring AEAD backend
|
||||
([#80](https://github.com/jmcorgan/fips/pull/80)),
|
||||
hot-path drain batching + recvmmsg + eager pubkey_full
|
||||
([#81](https://github.com/jmcorgan/fips/pull/81)),
|
||||
TreeAnnounce self-root ancestry + overlay-advert retry hygiene
|
||||
([#82](https://github.com/jmcorgan/fips/pull/82)),
|
||||
NAT-traversal MTU inheritance
|
||||
([#83](https://github.com/jmcorgan/fips/pull/83)).
|
||||
- [@dskvr](https://github.com/dskvr): initial Arch Linux AUR
|
||||
packaging ([#21](https://github.com/jmcorgan/fips/pull/21)) and
|
||||
the AUR publish workflow.
|
||||
- [@SatsAndSports](https://github.com/SatsAndSports): rekey
|
||||
message-1 admit fix on non-accepting transports
|
||||
([#49](https://github.com/jmcorgan/fips/pull/49)),
|
||||
TreeAnnounce semantic validation, gateway test image fix
|
||||
([#69](https://github.com/jmcorgan/fips/pull/69)).
|
||||
- [@andrewheadricke](https://github.com/andrewheadricke): MIPS
|
||||
atomic-ABI portability via `portable_atomic`
|
||||
([#62](https://github.com/jmcorgan/fips/pull/62)).
|
||||
- [@sh1ftred](https://github.com/sh1ftred): Arch packaging namcap
|
||||
fixes ([#63](https://github.com/jmcorgan/fips/pull/63)).
|
||||
- [@oleksky](https://github.com/oleksky): macOS WireGuard companion
|
||||
collaboration on [#51](https://github.com/jmcorgan/fips/pull/51).
|
||||
|
||||
**Issue reports that drove fixes in this release**:
|
||||
|
||||
- [@deavmi](https://github.com/deavmi): MIPS daemon build support
|
||||
([#26](https://github.com/jmcorgan/fips/issues/26)).
|
||||
- [@Sebastix](https://github.com/Sebastix): fipsctl/fipstop
|
||||
control-socket path detection
|
||||
([#30](https://github.com/jmcorgan/fips/issues/30)).
|
||||
- [@SwapMarket](https://github.com/SwapMarket): auto-connect
|
||||
reconnect after graceful disconnect
|
||||
([#60](https://github.com/jmcorgan/fips/issues/60)) and
|
||||
fipsctl mesh-address rejection
|
||||
([#61](https://github.com/jmcorgan/fips/issues/61)).
|
||||
@@ -36,4 +36,14 @@ fn main() {
|
||||
|
||||
// Support reproducible builds (Debian packaging)
|
||||
println!("cargo:rerun-if-env-changed=SOURCE_DATE_EPOCH");
|
||||
|
||||
// bluer/BlueZ is glibc-linux only: musl cross-compiles (OpenWrt) can't
|
||||
// satisfy libdbus-sys's pkg-config cross-compile requirement, and musl
|
||||
// router targets don't run BlueZ by default anyway.
|
||||
println!("cargo:rustc-check-cfg=cfg(bluer_available)");
|
||||
let target_os = std::env::var("CARGO_CFG_TARGET_OS").unwrap_or_default();
|
||||
let target_env = std::env::var("CARGO_CFG_TARGET_ENV").unwrap_or_default();
|
||||
if target_os == "linux" && target_env != "musl" {
|
||||
println!("cargo:rustc-cfg=bluer_available");
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,5 +1,54 @@
|
||||
# FIPS Documentation
|
||||
|
||||
| Directory | Description |
|
||||
|-----------|-------------|
|
||||
| [design/](design/) | Protocol design specifications and analysis |
|
||||
FIPS (Free Internetworking Peering System) is a self-organizing
|
||||
encrypted mesh network built on Nostr identities, capable of
|
||||
operating over arbitrary transports — local networks, the public
|
||||
internet, Tor, Bluetooth, or point-to-point links — without central
|
||||
infrastructure.
|
||||
|
||||
With FIPS, your machine becomes a node in the mesh with a
|
||||
self-generated cryptographic identity. There are two ways to
|
||||
deploy it.
|
||||
|
||||
**As an overlay** on top of existing IP networks, FIPS lets
|
||||
your node reach any other FIPS node wherever it sits — behind a NAT, on a
|
||||
different ISP, on a phone over cellular, on a laptop with only
|
||||
Bluetooth in range, or behind a Tor onion. The mesh forwards
|
||||
IPv6 traffic transparently and end-to-end encrypted, with no
|
||||
central VPN concentrator or coordinating server.
|
||||
|
||||
**From the ground up** over raw Ethernet, WiFi, or Bluetooth,
|
||||
FIPS provides a complete permissionless network
|
||||
without any pre-existing IP infrastructure, ISP, or DNS. Any
|
||||
node that joins the link gets routable IPv6 addresses, peer
|
||||
discovery, and a path to every other node automatically.
|
||||
|
||||
Either way, existing networking software runs over it unchanged:
|
||||
SSH, HTTP servers, file transfer, anything IPv6-native works the
|
||||
same way it would on a local network.
|
||||
|
||||
New to FIPS? Start with the [Getting Started](getting-started.md)
|
||||
guide.
|
||||
|
||||
## Documentation Sections
|
||||
|
||||
### [Tutorials](tutorials/)
|
||||
|
||||
If you are starting from scratch and want a guided path to a
|
||||
working mesh, go here.
|
||||
|
||||
### [How-To Guides](how-to/)
|
||||
|
||||
If you have a specific task in mind — enabling a feature,
|
||||
deploying a component, diagnosing a problem — go here.
|
||||
|
||||
### [Reference](reference/)
|
||||
|
||||
If you need to look up wire formats, configuration keys, command
|
||||
flags, or counter inventories, go here.
|
||||
|
||||
### [Design](design/)
|
||||
|
||||
If you want to understand how the mesh self-organizes, why FIPS
|
||||
makes the choices it does, or how the pieces fit together, go
|
||||
here.
|
||||
|
||||
@@ -1,51 +1,65 @@
|
||||
# FIPS Design Documents
|
||||
# FIPS Design
|
||||
|
||||
Protocol design specifications for the Federated Interoperable Peering
|
||||
System — a self-organizing encrypted mesh network built on Nostr identities.
|
||||
Architectural and protocol-level explanations for FIPS — the *why*
|
||||
and the *how* behind the wire and the system. For wire formats and
|
||||
configuration keys, see [reference/](../reference/). For task
|
||||
recipes, see [how-to/](../how-to/). For end-to-end lessons, see
|
||||
[tutorials/](../tutorials/).
|
||||
|
||||
## Reading Order
|
||||
|
||||
Start with the introduction, then follow the protocol stack from bottom to
|
||||
top. After the stack, the mesh operation document explains how all the
|
||||
pieces work together. Supporting references provide deeper dives into
|
||||
specific topics.
|
||||
Start with [fips-concepts.md](fips-concepts.md) for the
|
||||
novice-friendly framing of what FIPS is and why, then move to
|
||||
[fips-architecture.md](fips-architecture.md) for the protocol stack,
|
||||
identity model, and two-layer encryption walkthrough. From there,
|
||||
follow the protocol stack from bottom to top. After the stack,
|
||||
[fips-mesh-operation.md](fips-mesh-operation.md) explains how the
|
||||
pieces work together at runtime. Cross-cutting and supporting
|
||||
documents cover specific subsystems in detail.
|
||||
|
||||
### Foundations
|
||||
|
||||
| Document | Description |
|
||||
| -------- | ----------- |
|
||||
| [fips-concepts.md](fips-concepts.md) | What FIPS is, why it exists, mental model |
|
||||
| [fips-architecture.md](fips-architecture.md) | Protocol stack, identity, two-layer encryption |
|
||||
| [fips-prior-work.md](fips-prior-work.md) | Designs and protocols FIPS builds on |
|
||||
|
||||
### Protocol Stack
|
||||
|
||||
| Document | Description |
|
||||
| -------- | ----------- |
|
||||
| [fips-intro.md](fips-intro.md) | Protocol introduction: goals, architecture, layer model |
|
||||
| [fips-transport-layer.md](fips-transport-layer.md) | Transport layer: datagram delivery over arbitrary media |
|
||||
| [fips-mesh-layer.md](fips-mesh-layer.md) | FIPS Mesh Protocol (FMP): peer authentication, link encryption, forwarding |
|
||||
| [fips-session-layer.md](fips-session-layer.md) | FIPS Session Protocol (FSP): end-to-end encryption, sessions |
|
||||
| [fips-ipv6-adapter.md](fips-ipv6-adapter.md) | IPv6 adaptation: TUN interface, DNS, MTU enforcement |
|
||||
|
||||
### Cross-Cutting
|
||||
|
||||
| Document | Description |
|
||||
| -------- | ----------- |
|
||||
| [fips-mmp.md](fips-mmp.md) | Metrics Measurement Protocol (link + session) |
|
||||
| [fips-mtu.md](fips-mtu.md) | Path MTU model, encapsulation overhead, PMTUD |
|
||||
| [fips-security.md](fips-security.md) | `fips0` interface threat model and default-deny baseline |
|
||||
|
||||
### Mesh Behavior
|
||||
|
||||
| Document | Description |
|
||||
| -------- | ----------- |
|
||||
| [fips-mesh-operation.md](fips-mesh-operation.md) | How the mesh operates: routing, discovery, error recovery |
|
||||
| [fips-wire-formats.md](fips-wire-formats.md) | Wire format reference for all message types |
|
||||
| [fips-nostr-discovery.md](fips-nostr-discovery.md) | Optional Nostr-mediated peer discovery and UDP NAT hole-punch |
|
||||
| [port-advertisement-and-nat-traversal.md](port-advertisement-and-nat-traversal.md) | Nostr-signaled port advertisement and UDP NAT-traversal protocol; generic, with FIPS as an example implementation |
|
||||
|
||||
### Supporting References
|
||||
### Deeper Dives
|
||||
|
||||
| Document | Description |
|
||||
| -------- | ----------- |
|
||||
| [fips-spanning-tree.md](fips-spanning-tree.md) | Spanning tree algorithms: root discovery, parent selection, coordinates |
|
||||
| [fips-bloom-filters.md](fips-bloom-filters.md) | Bloom filter math: FPR analysis, size classes, split-horizon |
|
||||
|
||||
### Implementation
|
||||
|
||||
| Document | Description |
|
||||
| -------- | ----------- |
|
||||
| [fips-configuration.md](fips-configuration.md) | YAML configuration reference |
|
||||
|
||||
### Supplemental
|
||||
|
||||
| Document | Description |
|
||||
| -------- | ----------- |
|
||||
| [fips-bloom-filters.md](fips-bloom-filters.md) | Bloom filter properties: FPR analysis, size classes, split-horizon |
|
||||
| [spanning-tree-dynamics.md](spanning-tree-dynamics.md) | Spanning tree walkthroughs: convergence scenarios, worked examples |
|
||||
|
||||
## Document Relationships
|
||||
### Adjacent Components
|
||||
|
||||

|
||||
| Document | Description |
|
||||
| -------- | ----------- |
|
||||
| [fips-gateway.md](fips-gateway.md) | `fips-gateway` service: outbound (LAN-to-mesh) DNS-proxy + virtual-IP NAT and inbound (mesh-to-LAN) port-forwarding, sharing one nftables table |
|
||||
|
||||
@@ -82,7 +82,7 @@
|
||||
<text x="610" y="223" text-anchor="middle" class="alabel" fill="#5080c0">LookupRequest</text>
|
||||
|
||||
<!-- Bloom filter annotation -->
|
||||
<text x="430" y="260" text-anchor="middle" class="annot">guided by bloom filters at each hop</text>
|
||||
<text x="430" y="260" text-anchor="middle" class="annot">guided by bloom filters at each hop; transits do not cache</text>
|
||||
|
||||
<!-- ═══ Phase separator ═══ -->
|
||||
<line x1="70" y1="285" x2="830" y2="285" class="sep"/>
|
||||
@@ -99,23 +99,19 @@
|
||||
<polygon points="534,335 524,340 534,345" class="resp"/>
|
||||
<text x="610" y="333" text-anchor="middle" class="alabel" fill="#d0a040">LookupResponse + coords</text>
|
||||
|
||||
<!-- Cache box at C -->
|
||||
<rect x="490" y="358" width="60" height="22" class="cache"/>
|
||||
<text x="520" y="373" text-anchor="middle" class="clabel">cache D</text>
|
||||
<!-- C → B LookupResponse (transit forwards without caching) -->
|
||||
<line x1="510" y1="380" x2="352" y2="380" class="resp"/>
|
||||
<polygon points="354,375 344,380 354,385" class="resp"/>
|
||||
<text x="430" y="373" text-anchor="middle" class="alabel" fill="#d0a040">LookupResponse + coords</text>
|
||||
|
||||
<!-- C → B LookupResponse -->
|
||||
<line x1="510" y1="395" x2="352" y2="395" class="resp"/>
|
||||
<polygon points="354,390 344,395 354,400" class="resp"/>
|
||||
<text x="430" y="388" text-anchor="middle" class="alabel" fill="#d0a040">LookupResponse + coords</text>
|
||||
<!-- B → A LookupResponse (transit forwards without caching) -->
|
||||
<line x1="330" y1="420" x2="172" y2="420" class="resp"/>
|
||||
<polygon points="174,415 164,420 174,425" class="resp"/>
|
||||
<text x="250" y="413" text-anchor="middle" class="alabel" fill="#d0a040">LookupResponse + coords</text>
|
||||
|
||||
<!-- Cache box at B -->
|
||||
<rect x="310" y="413" width="60" height="22" class="cache"/>
|
||||
<text x="340" y="428" text-anchor="middle" class="clabel">cache D</text>
|
||||
|
||||
<!-- B → A LookupResponse -->
|
||||
<line x1="330" y1="448" x2="172" y2="448" class="resp"/>
|
||||
<polygon points="174,443 164,448 174,453" class="resp"/>
|
||||
<text x="250" y="441" text-anchor="middle" class="alabel" fill="#d0a040">LookupResponse + coords</text>
|
||||
<!-- Cache box at A — only the originator caches on LookupResponse -->
|
||||
<rect x="130" y="438" width="60" height="22" class="cache"/>
|
||||
<text x="160" y="453" text-anchor="middle" class="clabel">cache D</text>
|
||||
|
||||
<!-- ═══ Phase separator ═══ -->
|
||||
<line x1="70" y1="475" x2="830" y2="475" class="sep"/>
|
||||
@@ -125,32 +121,34 @@
|
||||
<!-- ═══════════════════════════════════════════════ -->
|
||||
|
||||
<text x="40" y="508" text-anchor="middle" class="plabel">Phase 3</text>
|
||||
<text x="40" y="524" text-anchor="middle" class="annot">Routing</text>
|
||||
<text x="40" y="524" text-anchor="middle" class="annot">Data flow</text>
|
||||
|
||||
<!-- A → B Data -->
|
||||
<!-- A → B Data with coords (CP flag / SessionSetup) -->
|
||||
<line x1="170" y1="530" x2="328" y2="530" class="data"/>
|
||||
<polygon points="326,525 336,530 326,535" class="data"/>
|
||||
<text x="250" y="523" text-anchor="middle" class="alabel" fill="#40a060">Data</text>
|
||||
<text x="250" y="523" text-anchor="middle" class="alabel" fill="#40a060">Data + coords</text>
|
||||
|
||||
<!-- Cached coords note at B -->
|
||||
<text x="340" y="553" text-anchor="middle" class="annot">cached coords</text>
|
||||
<!-- Cache box at B — warmed in-band from CP-flagged data -->
|
||||
<rect x="310" y="540" width="60" height="22" class="cache"/>
|
||||
<text x="340" y="555" text-anchor="middle" class="clabel">cache D</text>
|
||||
|
||||
<!-- B → C Data -->
|
||||
<line x1="350" y1="565" x2="508" y2="565" class="data"/>
|
||||
<polygon points="506,560 516,565 506,570" class="data"/>
|
||||
<text x="430" y="558" text-anchor="middle" class="alabel" fill="#40a060">Data</text>
|
||||
<!-- B → C Data with coords -->
|
||||
<line x1="350" y1="572" x2="508" y2="572" class="data"/>
|
||||
<polygon points="506,567 516,572 506,577" class="data"/>
|
||||
<text x="430" y="565" text-anchor="middle" class="alabel" fill="#40a060">Data + coords</text>
|
||||
|
||||
<!-- Cached coords note at C -->
|
||||
<text x="520" y="588" text-anchor="middle" class="annot">cached coords</text>
|
||||
<!-- Cache box at C — warmed in-band -->
|
||||
<rect x="490" y="582" width="60" height="22" class="cache"/>
|
||||
<text x="520" y="597" text-anchor="middle" class="clabel">cache D</text>
|
||||
|
||||
<!-- C → D Data -->
|
||||
<line x1="530" y1="600" x2="688" y2="600" class="data"/>
|
||||
<polygon points="686,595 696,600 686,605" class="data"/>
|
||||
<text x="610" y="593" text-anchor="middle" class="alabel" fill="#40a060">Data</text>
|
||||
<line x1="530" y1="614" x2="688" y2="614" class="data"/>
|
||||
<polygon points="686,609 696,614 686,619" class="data"/>
|
||||
<text x="610" y="607" text-anchor="middle" class="alabel" fill="#40a060">Data + coords</text>
|
||||
|
||||
<!-- Efficient forwarding annotation -->
|
||||
<text x="430" y="622" text-anchor="middle" class="annot">cached coords enable efficient forwarding — no re-discovery needed</text>
|
||||
<text x="430" y="636" text-anchor="middle" class="annot">transits cache coords from in-flight data; subsequent traffic forwards without re-discovery</text>
|
||||
|
||||
<!-- ═══ Caption ═══ -->
|
||||
<text x="430" y="660" text-anchor="middle" class="caption">Each transit node caches coordinates from the LookupResponse return path</text>
|
||||
<text x="430" y="668" text-anchor="middle" class="caption">LookupResponse caches coords at the originator only; transit caches warm during the subsequent data flow</text>
|
||||
</svg>
|
||||
|
||||
|
Before Width: | Height: | Size: 7.8 KiB After Width: | Height: | Size: 8.1 KiB |
@@ -71,7 +71,7 @@
|
||||
<!-- Arrow down from pubkey -->
|
||||
<line x1="400" y1="120" x2="400" y2="170" class="derive"/>
|
||||
<polygon points="396,168 400,176 404,168" class="dhead"/>
|
||||
<text x="416" y="148" class="op">one-way hash</text>
|
||||
<text x="416" y="148" class="op">SHA-256, truncate to 16 bytes</text>
|
||||
|
||||
<!-- node_addr box -->
|
||||
<rect x="280" y="176" width="240" height="50" class="box derived"/>
|
||||
@@ -95,7 +95,7 @@
|
||||
<!-- Arrow down from node_addr -->
|
||||
<line x1="400" y1="226" x2="400" y2="276" class="derive"/>
|
||||
<polygon points="396,274 400,282 404,274" class="dhead"/>
|
||||
<text x="416" y="254" class="op">add fd00::/8 prefix</text>
|
||||
<text x="416" y="254" class="op">0xfd + node_addr[0..15]</text>
|
||||
|
||||
<!-- IPv6 address box -->
|
||||
<rect x="280" y="282" width="240" height="50" class="box compat"/>
|
||||
|
||||
|
Before Width: | Height: | Size: 6.2 KiB After Width: | Height: | Size: 6.3 KiB |
@@ -82,48 +82,44 @@
|
||||
<!-- === Transport layer === -->
|
||||
|
||||
<!-- Overlay transports -->
|
||||
<rect x="80" y="336" width="150" height="60" class="cat"/>
|
||||
<text x="155" y="349" text-anchor="middle" class="cat">Overlay</text>
|
||||
<rect x="80" y="336" width="216" height="60" class="cat"/>
|
||||
<text x="188" y="349" text-anchor="middle" class="cat">Overlay</text>
|
||||
|
||||
<rect x="92" y="356" width="60" height="30" class="layer xport"/>
|
||||
<text x="122" y="371" text-anchor="middle" font-size="10" font-weight="bold" fill="#e0e0e0">UDP</text>
|
||||
<text x="122" y="381" text-anchor="middle" class="sub">IP</text>
|
||||
|
||||
<rect x="158" y="356" width="60" height="30" class="layer xport"/>
|
||||
<text x="188" y="371" text-anchor="middle" font-size="10" font-weight="bold" fill="#e0e0e0">Tor</text>
|
||||
<text x="188" y="381" text-anchor="middle" class="sub">.onion</text>
|
||||
<text x="188" y="371" text-anchor="middle" font-size="10" font-weight="bold" fill="#e0e0e0">TCP</text>
|
||||
<text x="188" y="381" text-anchor="middle" class="sub">IP</text>
|
||||
|
||||
<rect x="224" y="356" width="60" height="30" class="layer xport"/>
|
||||
<text x="254" y="371" text-anchor="middle" font-size="10" font-weight="bold" fill="#e0e0e0">Tor</text>
|
||||
<text x="254" y="381" text-anchor="middle" class="sub">.onion</text>
|
||||
|
||||
<!-- Shared medium transports -->
|
||||
<rect x="240" y="336" width="300" height="60" class="cat"/>
|
||||
<text x="390" y="349" text-anchor="middle" class="cat">Shared Medium</text>
|
||||
<rect x="306" y="336" width="234" height="60" class="cat"/>
|
||||
<text x="423" y="349" text-anchor="middle" class="cat">Shared Medium</text>
|
||||
|
||||
<rect x="254" y="356" width="60" height="30" class="layer xport"/>
|
||||
<text x="284" y="371" text-anchor="middle" font-size="10" font-weight="bold" fill="#e0e0e0">Ether</text>
|
||||
<text x="284" y="381" text-anchor="middle" class="sub">802.3</text>
|
||||
<rect x="318" y="356" width="60" height="30" class="layer xport"/>
|
||||
<text x="348" y="371" text-anchor="middle" font-size="10" font-weight="bold" fill="#e0e0e0">Ether</text>
|
||||
<text x="348" y="381" text-anchor="middle" class="sub">802.3</text>
|
||||
|
||||
<rect x="320" y="356" width="60" height="30" class="layer xport"/>
|
||||
<text x="350" y="371" text-anchor="middle" font-size="10" font-weight="bold" fill="#e0e0e0">WiFi</text>
|
||||
<text x="350" y="381" text-anchor="middle" class="sub">802.11</text>
|
||||
<rect x="384" y="356" width="60" height="30" class="layer xport"/>
|
||||
<text x="414" y="371" text-anchor="middle" font-size="10" font-weight="bold" fill="#e0e0e0">BLE</text>
|
||||
<text x="414" y="381" text-anchor="middle" class="sub">L2CAP</text>
|
||||
|
||||
<rect x="386" y="356" width="60" height="30" class="layer xport"/>
|
||||
<text x="416" y="371" text-anchor="middle" font-size="10" font-weight="bold" fill="#e0e0e0">BT</text>
|
||||
<text x="416" y="381" text-anchor="middle" class="sub">RFCOMM</text>
|
||||
|
||||
<rect x="452" y="356" width="60" height="30" class="layer xport"/>
|
||||
<text x="482" y="371" text-anchor="middle" font-size="10" font-weight="bold" fill="#e0e0e0">Radio</text>
|
||||
<text x="482" y="381" text-anchor="middle" class="sub">Sat, ...</text>
|
||||
<rect x="450" y="356" width="80" height="30" class="layer xport"/>
|
||||
<text x="490" y="371" text-anchor="middle" font-size="10" font-weight="bold" fill="#e0e0e0">Radio ...</text>
|
||||
<text x="490" y="381" text-anchor="middle" class="sub">future</text>
|
||||
|
||||
<!-- Point-to-point transports -->
|
||||
<rect x="550" y="336" width="150" height="60" class="cat"/>
|
||||
<text x="625" y="349" text-anchor="middle" class="cat">Point-to-Point</text>
|
||||
|
||||
<rect x="562" y="356" width="60" height="30" class="layer xport"/>
|
||||
<text x="592" y="371" text-anchor="middle" font-size="10" font-weight="bold" fill="#e0e0e0">Serial</text>
|
||||
<text x="592" y="381" text-anchor="middle" class="sub">UART</text>
|
||||
|
||||
<rect x="628" y="356" width="60" height="30" class="layer xport"/>
|
||||
<text x="658" y="371" text-anchor="middle" font-size="10" font-weight="bold" fill="#e0e0e0">...</text>
|
||||
<text x="658" y="381" text-anchor="middle" class="sub"></text>
|
||||
<rect x="562" y="356" width="126" height="30" class="layer xport"/>
|
||||
<text x="625" y="371" text-anchor="middle" font-size="10" font-weight="bold" fill="#e0e0e0">Serial ...</text>
|
||||
<text x="625" y="381" text-anchor="middle" class="sub">future</text>
|
||||
|
||||
<!-- === Peer networks below node box === -->
|
||||
<text x="155" y="436" text-anchor="middle" class="sub">Internet / Overlay Peers</text>
|
||||
|
||||
|
Before Width: | Height: | Size: 8.1 KiB After Width: | Height: | Size: 7.8 KiB |
@@ -19,7 +19,7 @@
|
||||
<rect x="50" y="25" width="775" height="100" class="layer app"/>
|
||||
<text x="75" y="58" class="name">Application Layer Interface</text>
|
||||
<text x="75" y="78" class="desc">Native FIPS API — for FIPS-aware applications</text>
|
||||
<text x="75" y="98" class="desc">IPv6 Shim — for traditional IP application backward compatibility</text>
|
||||
<text x="75" y="98" class="desc">IPv6 adapter — for traditional IP application backward compatibility</text>
|
||||
|
||||
<!-- FSP layer -->
|
||||
<rect x="50" y="150" width="775" height="120" class="layer fsp"/>
|
||||
|
||||
|
Before Width: | Height: | Size: 2.8 KiB After Width: | Height: | Size: 2.8 KiB |
@@ -76,54 +76,54 @@
|
||||
<text x="372" y="364" class="branch">No</text>
|
||||
|
||||
<!-- ============================================================ -->
|
||||
<!-- STEP 3: Bloom filter hit? -->
|
||||
<!-- STEP 3: Coords known? -->
|
||||
<!-- ============================================================ -->
|
||||
<text x="240" y="395" text-anchor="end" class="step">3</text>
|
||||
<polygon points="360,380 440,420 360,460 280,420" class="diamond"/>
|
||||
<text x="360" y="416" text-anchor="middle" class="decision">Bloom filter</text>
|
||||
<text x="360" y="430" text-anchor="middle" class="decision">hit?</text>
|
||||
<text x="360" y="416" text-anchor="middle" class="decision">Coords</text>
|
||||
<text x="360" y="430" text-anchor="middle" class="decision">known?</text>
|
||||
|
||||
<!-- Yes → right to interim step -->
|
||||
<!-- No → right to error outcome -->
|
||||
<line x1="440" y1="420" x2="520" y2="420" class="arrow" marker-end="url(#arrowhead)"/>
|
||||
<text x="475" y="412" text-anchor="middle" class="branch">Yes</text>
|
||||
<rect x="520" y="396" width="176" height="48" class="interim"/>
|
||||
<text x="608" y="412" text-anchor="middle" class="action">Rank candidates by</text>
|
||||
<text x="608" y="426" text-anchor="middle" class="action">tree distance and</text>
|
||||
<text x="608" y="440" text-anchor="middle" class="action">link performance</text>
|
||||
<text x="475" y="412" text-anchor="middle" class="branch">No</text>
|
||||
<rect x="528" y="402" width="160" height="36" class="error"/>
|
||||
<text x="608" y="425" text-anchor="middle" class="action">No route → error signal</text>
|
||||
|
||||
<!-- Arrow from interim → final outcome -->
|
||||
<line x1="608" y1="444" x2="608" y2="470" class="arrow" marker-end="url(#arrowhead)"/>
|
||||
<rect x="528" y="478" width="160" height="36" class="outcome"/>
|
||||
<text x="608" y="501" text-anchor="middle" class="action">Forward to 'best'</text>
|
||||
|
||||
<!-- No → down -->
|
||||
<!-- Yes → down -->
|
||||
<line x1="360" y1="460" x2="360" y2="530" class="arrow" marker-end="url(#arrowhead)"/>
|
||||
<text x="372" y="500" class="branch">No</text>
|
||||
<text x="372" y="500" class="branch">Yes</text>
|
||||
|
||||
<!-- ============================================================ -->
|
||||
<!-- STEP 4: Coords known? -->
|
||||
<!-- STEP 4: Bloom filter hit? -->
|
||||
<!-- ============================================================ -->
|
||||
<text x="240" y="545" text-anchor="end" class="step">4</text>
|
||||
<polygon points="360,530 440,570 360,610 280,570" class="diamond"/>
|
||||
<text x="360" y="566" text-anchor="middle" class="decision">Coords</text>
|
||||
<text x="360" y="580" text-anchor="middle" class="decision">known?</text>
|
||||
<text x="360" y="566" text-anchor="middle" class="decision">Bloom filter</text>
|
||||
<text x="360" y="580" text-anchor="middle" class="decision">hit?</text>
|
||||
|
||||
<!-- Yes → right to outcome -->
|
||||
<!-- Yes → right to interim step -->
|
||||
<line x1="440" y1="570" x2="520" y2="570" class="arrow" marker-end="url(#arrowhead)"/>
|
||||
<text x="475" y="562" text-anchor="middle" class="branch">Yes</text>
|
||||
<rect x="528" y="552" width="160" height="36" class="outcome"/>
|
||||
<text x="608" y="575" text-anchor="middle" class="action">Greedy tree forward</text>
|
||||
<rect x="520" y="546" width="176" height="48" class="interim"/>
|
||||
<text x="608" y="562" text-anchor="middle" class="action">Rank candidates by</text>
|
||||
<text x="608" y="576" text-anchor="middle" class="action">tree distance and</text>
|
||||
<text x="608" y="590" text-anchor="middle" class="action">link performance</text>
|
||||
|
||||
<!-- Arrow from interim → final outcome -->
|
||||
<line x1="608" y1="594" x2="608" y2="620" class="arrow" marker-end="url(#arrowhead)"/>
|
||||
<rect x="528" y="628" width="160" height="36" class="outcome"/>
|
||||
<text x="608" y="651" text-anchor="middle" class="action">Forward to 'best'</text>
|
||||
|
||||
<!-- No → down -->
|
||||
<line x1="360" y1="610" x2="360" y2="660" class="arrow" marker-end="url(#arrowhead)"/>
|
||||
<text x="372" y="640" class="branch">No</text>
|
||||
|
||||
<!-- ============================================================ -->
|
||||
<!-- STEP 5: No route → error signal -->
|
||||
<!-- STEP 5: Greedy tree forward -->
|
||||
<!-- ============================================================ -->
|
||||
<text x="240" y="683" text-anchor="end" class="step">5</text>
|
||||
<rect x="260" y="668" width="200" height="36" class="error"/>
|
||||
<text x="360" y="691" text-anchor="middle" class="action">No route → error signal</text>
|
||||
<rect x="260" y="668" width="200" height="36" class="outcome"/>
|
||||
<text x="360" y="691" text-anchor="middle" class="action">Greedy tree forward</text>
|
||||
|
||||
<!-- ============================================================ -->
|
||||
<!-- Legend -->
|
||||
@@ -149,5 +149,5 @@
|
||||
<text x="442" y="835" font-size="11" fill="#a0a0b0">Control flow direction</text>
|
||||
|
||||
<!-- Caption -->
|
||||
<text x="360" y="896" text-anchor="middle" class="caption">Each hop evaluates destinations in priority order 1–4, falling through on miss</text>
|
||||
<text x="360" y="896" text-anchor="middle" class="caption">Each hop checks 1–4 in priority order, falling through to greedy tree (5) when bloom yields no candidate; missing coords is the only error path</text>
|
||||
</svg>
|
||||
|
||||
|
Before Width: | Height: | Size: 8.5 KiB After Width: | Height: | Size: 8.6 KiB |
@@ -1,100 +0,0 @@
|
||||
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 720 520" font-family="system-ui, -apple-system, sans-serif" font-size="13">
|
||||
<defs>
|
||||
<marker id="arrow" viewBox="0 0 10 10" refX="10" refY="5" markerWidth="8" markerHeight="8" orient="auto-start-reverse">
|
||||
<path d="M 0 0 L 10 5 L 0 10 z" fill="#555"/>
|
||||
</marker>
|
||||
</defs>
|
||||
|
||||
<!-- Background -->
|
||||
<rect width="720" height="520" rx="8" fill="#fafafa" stroke="#ddd" stroke-width="1"/>
|
||||
|
||||
<!-- Title -->
|
||||
<text x="360" y="32" text-anchor="middle" font-size="16" font-weight="600" fill="#333">Document Relationships</text>
|
||||
|
||||
<!-- fips-intro -->
|
||||
<rect x="260" y="50" width="200" height="32" rx="6" fill="#e3f2fd" stroke="#90caf9"/>
|
||||
<text x="360" y="71" text-anchor="middle" font-weight="500" fill="#1565c0">fips-intro.md</text>
|
||||
|
||||
<!-- Arrows from intro -->
|
||||
<line x1="310" y1="82" x2="130" y2="120" stroke="#555" marker-end="url(#arrow)"/>
|
||||
<line x1="360" y1="82" x2="360" y2="120" stroke="#555" marker-end="url(#arrow)"/>
|
||||
|
||||
<!-- fips-transport-layer -->
|
||||
<rect x="30" y="120" width="200" height="32" rx="6" fill="#e8f5e9" stroke="#a5d6a7"/>
|
||||
<text x="130" y="141" text-anchor="middle" font-weight="500" fill="#2e7d32">fips-transport-layer.md</text>
|
||||
|
||||
<!-- fips-mesh-operation -->
|
||||
<rect x="270" y="120" width="200" height="32" rx="6" fill="#fff3e0" stroke="#ffcc80"/>
|
||||
<text x="370" y="141" text-anchor="middle" font-weight="500" fill="#e65100">fips-mesh-operation.md</text>
|
||||
|
||||
<!-- Arrow: transport -> mesh-layer -->
|
||||
<line x1="130" y1="152" x2="130" y2="190" stroke="#555" marker-end="url(#arrow)"/>
|
||||
|
||||
<!-- fips-mesh-layer -->
|
||||
<rect x="30" y="190" width="200" height="32" rx="6" fill="#e8f5e9" stroke="#a5d6a7"/>
|
||||
<text x="130" y="211" text-anchor="middle" font-weight="500" fill="#2e7d32">fips-mesh-layer.md</text>
|
||||
|
||||
<!-- Arrow: mesh-operation -> mesh-layer -->
|
||||
<line x1="270" y1="145" x2="230" y2="200" stroke="#555" marker-end="url(#arrow)"/>
|
||||
|
||||
<!-- Arrow: mesh-layer -> session-layer -->
|
||||
<line x1="130" y1="222" x2="130" y2="260" stroke="#555" marker-end="url(#arrow)"/>
|
||||
|
||||
<!-- fips-session-layer -->
|
||||
<rect x="30" y="260" width="200" height="32" rx="6" fill="#e8f5e9" stroke="#a5d6a7"/>
|
||||
<text x="130" y="281" text-anchor="middle" font-weight="500" fill="#2e7d32">fips-session-layer.md</text>
|
||||
|
||||
<!-- Arrow: session-layer -> ipv6-adapter -->
|
||||
<line x1="130" y1="292" x2="130" y2="330" stroke="#555" marker-end="url(#arrow)"/>
|
||||
|
||||
<!-- fips-ipv6-adapter -->
|
||||
<rect x="30" y="330" width="200" height="32" rx="6" fill="#e8f5e9" stroke="#a5d6a7"/>
|
||||
<text x="130" y="351" text-anchor="middle" font-weight="500" fill="#2e7d32">fips-ipv6-adapter.md</text>
|
||||
|
||||
<!-- Arrow: mesh-operation -> spanning-tree -->
|
||||
<line x1="470" y1="145" x2="550" y2="190" stroke="#555" marker-end="url(#arrow)"/>
|
||||
|
||||
<!-- Arrow: mesh-operation -> bloom-filters -->
|
||||
<line x1="470" y1="148" x2="550" y2="260" stroke="#555" marker-end="url(#arrow)"/>
|
||||
|
||||
<!-- fips-spanning-tree -->
|
||||
<rect x="500" y="190" width="200" height="32" rx="6" fill="#f3e5f5" stroke="#ce93d8"/>
|
||||
<text x="600" y="211" text-anchor="middle" font-weight="500" fill="#7b1fa2">fips-spanning-tree.md</text>
|
||||
|
||||
<!-- Arrow: intro -> spanning-tree -->
|
||||
<line x1="460" y1="72" x2="560" y2="190" stroke="#555" marker-end="url(#arrow)"/>
|
||||
|
||||
<!-- fips-bloom-filters -->
|
||||
<rect x="500" y="260" width="200" height="32" rx="6" fill="#f3e5f5" stroke="#ce93d8"/>
|
||||
<text x="600" y="281" text-anchor="middle" font-weight="500" fill="#7b1fa2">fips-bloom-filters.md</text>
|
||||
|
||||
<!-- Arrow: spanning-tree -> bloom (dependency) -->
|
||||
<line x1="600" y1="222" x2="600" y2="260" stroke="#999" stroke-dasharray="4,3" marker-end="url(#arrow)"/>
|
||||
|
||||
<!-- fips-wire-formats -->
|
||||
<rect x="260" y="400" width="200" height="32" rx="6" fill="#fce4ec" stroke="#ef9a9a"/>
|
||||
<text x="360" y="421" text-anchor="middle" font-weight="500" fill="#c62828">fips-wire-formats.md</text>
|
||||
<text x="360" y="448" text-anchor="middle" font-size="11" fill="#888">(referenced by all layer docs)</text>
|
||||
|
||||
<!-- fips-configuration -->
|
||||
<rect x="30" y="400" width="200" height="32" rx="6" fill="#f5f5f5" stroke="#bdbdbd"/>
|
||||
<text x="130" y="421" text-anchor="middle" font-weight="500" fill="#424242">fips-configuration.md</text>
|
||||
<text x="130" y="448" text-anchor="middle" font-size="11" fill="#888">(standalone reference)</text>
|
||||
|
||||
<!-- spanning-tree-dynamics -->
|
||||
<rect x="500" y="330" width="200" height="32" rx="6" fill="#f3e5f5" stroke="#ce93d8"/>
|
||||
<text x="600" y="351" text-anchor="middle" font-weight="500" fill="#7b1fa2">spanning-tree-dynamics.md</text>
|
||||
<text x="600" y="378" text-anchor="middle" font-size="11" fill="#888">(companion to fips-spanning-tree)</text>
|
||||
|
||||
<!-- Legend -->
|
||||
<rect x="30" y="472" width="14" height="14" rx="3" fill="#e8f5e9" stroke="#a5d6a7"/>
|
||||
<text x="50" y="484" font-size="11" fill="#666">Protocol stack</text>
|
||||
<rect x="155" y="472" width="14" height="14" rx="3" fill="#fff3e0" stroke="#ffcc80"/>
|
||||
<text x="175" y="484" font-size="11" fill="#666">Mesh behavior</text>
|
||||
<rect x="295" y="472" width="14" height="14" rx="3" fill="#f3e5f5" stroke="#ce93d8"/>
|
||||
<text x="315" y="484" font-size="11" fill="#666">Supporting references</text>
|
||||
<rect x="460" y="472" width="14" height="14" rx="3" fill="#fce4ec" stroke="#ef9a9a"/>
|
||||
<text x="480" y="484" font-size="11" fill="#666">Wire formats</text>
|
||||
<rect x="580" y="472" width="14" height="14" rx="3" fill="#f5f5f5" stroke="#bdbdbd"/>
|
||||
<text x="600" y="484" font-size="11" fill="#666">Implementation</text>
|
||||
</svg>
|
||||
|
Before Width: | Height: | Size: 5.4 KiB |
@@ -0,0 +1,264 @@
|
||||
# FIPS Architecture
|
||||
|
||||
The protocol architecture, identity system, and two-layer encryption
|
||||
model. For the higher-level "what is FIPS and why" framing, see
|
||||
[fips-concepts.md](fips-concepts.md). For prior art and academic
|
||||
citations, see [fips-prior-work.md](fips-prior-work.md).
|
||||
|
||||
## Protocol Architecture
|
||||
|
||||
FIPS is organized in three protocol layers, each with distinct
|
||||
responsibilities and clean service boundaries. No layer depends on
|
||||
the specifics of the layers above or below it — transport plugins
|
||||
know nothing about sessions, the routing layer knows nothing about
|
||||
application addressing, and applications know nothing about which
|
||||
physical media carry their traffic. This separation means new
|
||||
transports, protocol features, and application interfaces can be
|
||||
added independently.
|
||||
|
||||

|
||||
|
||||
### Mapping to Traditional Networking
|
||||
|
||||
Readers familiar with the OSI model or TCP/IP networking may find it
|
||||
helpful to see how FIPS concepts relate to traditional layers:
|
||||
|
||||

|
||||
|
||||
Note that FMP spans what would traditionally be separate link and
|
||||
network layers. This is intentional — in a self-organizing mesh, the
|
||||
same layer that authenticates peers also makes routing decisions,
|
||||
because routing depends on authenticated peer state (spanning tree
|
||||
positions, bloom filters).
|
||||
|
||||
### Layer Responsibilities
|
||||
|
||||
**Transport layer**: Delivers datagrams between endpoints over a
|
||||
specific medium. Each transport type (UDP socket, Ethernet interface,
|
||||
radio modem) implements the same abstract interface: send and receive
|
||||
datagrams, report MTU. The transport layer knows nothing about FIPS
|
||||
identities, routing, or encryption. It provides raw datagram delivery
|
||||
to FMP above.
|
||||
|
||||
See [fips-transport-layer.md](fips-transport-layer.md) for the
|
||||
transport layer specification.
|
||||
|
||||
**FIPS Mesh Protocol (FMP)**: Manages peer connections, authenticates
|
||||
peers via Noise IK handshakes, and encrypts all traffic on each link.
|
||||
FMP is where the mesh organizes itself — nodes exchange spanning tree
|
||||
announcements and bloom filters with their direct peers, and FMP
|
||||
makes forwarding decisions for transit traffic. FMP provides
|
||||
authenticated, encrypted forwarding to FSP above.
|
||||
|
||||
See [fips-mesh-layer.md](fips-mesh-layer.md) for the FMP specification
|
||||
and [fips-mesh-operation.md](fips-mesh-operation.md) for how FMP's
|
||||
routing and self-organization work in practice.
|
||||
|
||||
**FIPS Session Protocol (FSP)**: Provides end-to-end authenticated
|
||||
encryption between any two nodes, regardless of how many intermediate
|
||||
hops separate them. FSP manages session lifecycle (setup, data
|
||||
transfer, teardown), caches destination coordinates for efficient
|
||||
routing, and handles the warmup strategy that keeps transit node
|
||||
caches populated. Session dispatch uses index-based routing inspired
|
||||
by [WireGuard](https://www.wireguard.com/), enabling O(1) packet
|
||||
demultiplexing. FSP provides a datagram service to applications above.
|
||||
|
||||
See [fips-session-layer.md](fips-session-layer.md) for the FSP
|
||||
specification.
|
||||
|
||||
**IPv6 adaptation layer**: Sits above FSP as a service on port 256,
|
||||
adapting the FIPS datagram service for unmodified IPv6 applications.
|
||||
Provides DNS resolution (npub → fd00::/8 address), identity cache
|
||||
management, IPv6 header compression, MTU enforcement, and a TUN
|
||||
interface. This is the primary way existing applications use the FIPS
|
||||
mesh.
|
||||
|
||||
See [fips-ipv6-adapter.md](fips-ipv6-adapter.md) for the IPv6 adapter.
|
||||
|
||||
### Node Architecture
|
||||
|
||||
Application services sit at the top of the stack, dispatched by FSP
|
||||
port number: the IPv6 TUN adapter (port 256) maps npubs to `fd00::/8`
|
||||
addresses with header compression so unmodified IP applications can
|
||||
use the network transparently, while the native datagram API
|
||||
addresses destinations directly by npub.
|
||||
|
||||

|
||||
|
||||
The mesh routes application traffic across heterogeneous transports
|
||||
transparently. A packet may traverse WiFi, Ethernet, UDP/IP, and Tor
|
||||
links on its way from source to destination — the application never
|
||||
needs to know which transports are involved. Each hop is independently
|
||||
encrypted at the link layer, while a single end-to-end session
|
||||
protects the payload across the entire path.
|
||||
|
||||

|
||||
|
||||

|
||||
|
||||
## Identity System
|
||||
|
||||
FIPS uses [Nostr](https://github.com/nostr-protocol/nips) keypairs
|
||||
(secp256k1) as node identities. The public key identifies the node;
|
||||
the private key signs protocol messages and establishes encrypted
|
||||
sessions.
|
||||
|
||||
The public key (or its bech32-encoded npub form) is the primary means
|
||||
for application-layer software to identify communication endpoints.
|
||||
Internally, the protocol derives a `node_addr` (a 16-byte SHA-256 hash
|
||||
of the pubkey) used as the routing identifier in packet headers, and
|
||||
an IPv6 address derived from the node_addr for the TUN adapter.
|
||||
Applications use the pubkey or npub; the routing layer uses node_addr;
|
||||
unmodified IPv6 applications use the derived `fd00::/8` address. All
|
||||
three are deterministically derived from the same keypair.
|
||||
|
||||
### FIPS Identity Handling
|
||||
|
||||

|
||||
|
||||
The pubkey is the node's cryptographic identity, used in Noise
|
||||
handshakes for both link encryption (IK) and session encryption (XK).
|
||||
It is never exposed beyond the endpoints of an encrypted channel. The node_addr, a one-way
|
||||
SHA-256 hash truncated to 16 bytes, serves as the routing identifier
|
||||
in packet headers and bloom filters. Intermediate routers see only
|
||||
node_addrs — they can forward traffic without learning the Nostr
|
||||
identities of the endpoints. An observer can verify "does this
|
||||
node_addr belong to pubkey X?" if they already know the pubkey, but
|
||||
cannot enumerate communicating identities by inspecting traffic. The
|
||||
IPv6 address prepends `fd` to the first 15 bytes of the node_addr,
|
||||
providing a ULA overlay address for unmodified IP applications via the
|
||||
TUN interface.
|
||||
|
||||
Below the FIPS identity layer, each transport uses its own native
|
||||
addressing — IP:port or hostname:port addresses, MAC addresses,
|
||||
.onion identifiers. These **link addresses** are opaque to everything
|
||||
above FMP and discarded once link authentication completes.
|
||||
|
||||
### Identity Verification
|
||||
|
||||
The Noise Protocol Framework mutually authenticates both peer-to-peer
|
||||
link connections (at FMP) and end-to-end session traffic (at FSP),
|
||||
proving each party controls the private key for their claimed
|
||||
identity.
|
||||
|
||||
See [fips-mesh-layer.md](fips-mesh-layer.md) for peer authentication
|
||||
and [fips-session-layer.md](fips-session-layer.md) for end-to-end
|
||||
session establishment.
|
||||
|
||||
Key rotation changes the node's identity — a new keypair produces a
|
||||
new node_addr and IPv6 address, requiring all sessions to be
|
||||
re-established. Migration mechanisms that allow a node to announce a
|
||||
successor key are a future consideration.
|
||||
|
||||
## Two-Layer Encryption
|
||||
|
||||
FIPS uses independent encryption at two protocol layers:
|
||||
|
||||
| Layer | Scope | Pattern | Purpose |
|
||||
| ----- | ----- | ------- | ------- |
|
||||
| **FMP (Mesh)** | Hop-by-hop | Noise IK | Encrypt all traffic on each peer link |
|
||||
| **FSP (Session)** | End-to-end | Noise XK | Encrypt application payload between endpoints |
|
||||
|
||||
### Link Layer (Hop-by-Hop)
|
||||
|
||||
When two nodes establish a direct connection, they perform a [Noise
|
||||
IK](https://noiseprotocol.org/) handshake. This authenticates both
|
||||
parties and establishes symmetric keys for encrypting all traffic on
|
||||
that link. Every packet between direct peers is encrypted — gossip
|
||||
messages, routing queries, and forwarded session datagrams alike.
|
||||
|
||||
The IK pattern is used because outbound connections know the peer's
|
||||
npub from configuration, while inbound connections learn the
|
||||
initiator's identity from the first handshake message.
|
||||
|
||||
### Session Layer (End-to-End)
|
||||
|
||||
FIPS establishes end-to-end encrypted sessions between any two
|
||||
communicating nodes using Noise XK, regardless of how many hops
|
||||
separate them. The initiator knows the destination's npub (required
|
||||
for XK's pre-message); the responder learns the initiator's identity
|
||||
from the third handshake message. Unlike the link-layer IK pattern
|
||||
where the initiator's identity is revealed in msg1, XK delays
|
||||
identity disclosure until msg3, providing stronger initiator identity
|
||||
protection for traffic traversing untrusted intermediate nodes.
|
||||
|
||||
A packet from A to D through intermediate nodes B and C:
|
||||
|
||||
1. A encrypts payload with A↔D session key (FSP)
|
||||
2. A wraps in SessionDatagram, encrypts with A↔B link key (FMP),
|
||||
sends to B
|
||||
3. B decrypts link layer, reads destination node_addr, re-encrypts
|
||||
with B↔C link key, forwards to C
|
||||
4. C decrypts link layer, re-encrypts with C↔D link key, forwards
|
||||
to D
|
||||
5. D decrypts link layer, then decrypts session layer to get payload
|
||||
|
||||
Intermediate nodes route based on destination node_addr but cannot
|
||||
read session-layer payloads. Each hop strips one link encryption and
|
||||
applies the next — the session-layer ciphertext passes through
|
||||
untouched.
|
||||
|
||||
Both layers always apply, even between adjacent peers — a packet to a
|
||||
direct neighbor is still encrypted twice. This uniform model means no
|
||||
special cases for local vs remote destinations, and topology changes
|
||||
(a direct peer becomes reachable only through intermediaries) don't
|
||||
affect existing sessions.
|
||||
|
||||
See [fips-mesh-layer.md](fips-mesh-layer.md) for link encryption and
|
||||
[fips-session-layer.md](fips-session-layer.md) for session encryption.
|
||||
|
||||
## Routing and Mesh Operation
|
||||
|
||||
Forwarding decisions are local. Each node combines spanning-tree
|
||||
coordinates with peer bloom filters to choose a next hop, falling back
|
||||
to greedy tree routing when bloom filters have not converged. Discovery
|
||||
warms transit node caches with destination coordinates, and three
|
||||
explicit error signals (CoordsRequired, PathBroken, MtuExceeded) drive
|
||||
recovery when forwarding fails. The full routing decision process,
|
||||
discovery protocol, and error-recovery integration view live in
|
||||
[fips-mesh-operation.md](fips-mesh-operation.md).
|
||||
|
||||
## Transport Abstraction
|
||||
|
||||
FIPS treats the communication medium as a pluggable component. UDP,
|
||||
TCP, raw Ethernet, Tor, and BLE all implement the same small datagram
|
||||
interface (send, receive, report MTU) and feed peers into a single FMP
|
||||
routing layer; radio and serial transports are in the planned set.
|
||||
Multi-transport nodes bridge between networks transparently. The
|
||||
transport-layer specification — including per-transport categories,
|
||||
the trait surface, the connection model, and implementation status —
|
||||
is in [fips-transport-layer.md](fips-transport-layer.md).
|
||||
|
||||
## Security
|
||||
|
||||
FIPS defends against four adversary classes (transport observers,
|
||||
active transport attackers, intermediate routers, and adversarial
|
||||
mesh nodes) through layered controls: hop-by-hop FMP link encryption,
|
||||
end-to-end FSP session encryption with stronger initiator identity
|
||||
protection, signed and replay-protected gossip, and rate-limited
|
||||
handshake processing. The threat-model details and per-layer
|
||||
mitigations are in [fips-mesh-layer.md](fips-mesh-layer.md), and the
|
||||
operator-facing controls (default-deny baseline, peer ACLs,
|
||||
filesystem permissions, cryptographic primitives) are consolidated in
|
||||
[fips-security.md](fips-security.md) and
|
||||
[../reference/security.md](../reference/security.md).
|
||||
|
||||
## MTU as a Cross-Cutting Concern
|
||||
|
||||
MTU is not owned by any single layer. The transport layer reports
|
||||
per-link MTU, FMP carries `path_mtu` in SessionDatagram and
|
||||
LookupResponse to track the minimum along a path, FSP echoes the
|
||||
observed forward-path MTU back to the source, and the IPv6 adapter
|
||||
enforces the resulting effective MTU at the TUN with ICMP Packet Too
|
||||
Big and TCP MSS clamping. The unified design — encapsulation overhead
|
||||
budget, proactive PMTUD, reactive MtuExceeded, and per-destination
|
||||
storage — is in [fips-mtu.md](fips-mtu.md).
|
||||
|
||||
## Approaches Considered but Rejected
|
||||
|
||||
One design alternative evaluated and ruled out during the architecture
|
||||
pass was onion routing, rejected because it requires the sender to
|
||||
know the full path upfront (incompatible with self-organizing
|
||||
routing) and prevents per-hop error feedback (incompatible with
|
||||
CoordsRequired/PathBroken recovery). The canonical mention lives in
|
||||
[fips-mesh-operation.md](fips-mesh-operation.md#privacy-considerations).
|
||||
@@ -153,6 +153,15 @@ this node, and this node thinks it can reach the same destination through Q.
|
||||
Split-horizon is computed per-peer: the outbound filter for peer Q merges
|
||||
all tree peer inbound filters except Q's.
|
||||
|
||||
### Filter Propagation Diagram
|
||||
|
||||

|
||||
|
||||
The outbound filter for peer Q merges this node's identity with tree
|
||||
peer inbound filters except Q's (split-horizon exclusion). Upward
|
||||
filters (child → parent) contain the child's subtree, while downward
|
||||
filters (parent → child) contain the complement.
|
||||
|
||||
### Directional Asymmetry
|
||||
|
||||
Because merge is restricted to tree peers, outgoing filters exhibit
|
||||
@@ -202,9 +211,10 @@ Filter updates are event-driven, not periodic:
|
||||
|
||||
### Rate Limiting
|
||||
|
||||
Updates are rate-limited at 500ms minimum interval per peer to prevent
|
||||
storms during topology changes. Multiple pending changes within the
|
||||
cooldown period are coalesced into a single announcement.
|
||||
Updates are rate-limited at a 500ms minimum interval per peer
|
||||
(`node.bloom.update_debounce_ms`) to prevent storms during topology
|
||||
changes. Multiple pending changes within the cooldown period are
|
||||
coalesced into a single announcement.
|
||||
|
||||
### Propagation Scope
|
||||
|
||||
@@ -249,21 +259,14 @@ Where `filter_bits = 8 × (512 << size_class)` — 8,192 for v1.
|
||||
|
||||
## Wire Format
|
||||
|
||||
FilterAnnounce messages are carried inside encrypted link-layer frames:
|
||||
|
||||
| Offset | Field | Size | Description |
|
||||
| ------ | ----- | ---- | ----------- |
|
||||
| 0 | msg_type | 1 byte | 0x20 |
|
||||
| 1 | sequence | 8 bytes LE | Monotonic counter for freshness |
|
||||
| 9 | hash_count | 1 byte | Number of hash functions (5 in v1) |
|
||||
| 10 | size_class | 1 byte | Filter size: `512 << size_class` bytes |
|
||||
| 11 | filter_bits | 1,024 bytes | Bloom filter bit array (v1) |
|
||||
|
||||
**v1 total**: 1,035 bytes payload, 1,064 bytes with link encryption
|
||||
overhead.
|
||||
|
||||
See [fips-wire-formats.md](fips-wire-formats.md) for the complete wire
|
||||
format reference.
|
||||
The FilterAnnounce byte layout (`msg_type 0x20`, sequence, hash_count,
|
||||
size_class, filter_bits) lives in
|
||||
[../reference/wire-formats.md](../reference/wire-formats.md). The
|
||||
v1 plaintext payload is 1,035 bytes (11-byte header + 1,024-byte
|
||||
filter); link encryption adds 36 bytes of FMP framing (16-byte outer
|
||||
header + 4-byte inner timestamp + 16-byte AEAD tag), bringing the
|
||||
on-the-wire size to roughly 1,071 bytes before the underlying
|
||||
transport's per-packet overhead.
|
||||
|
||||
## Scale and Size Classes
|
||||
|
||||
@@ -320,9 +323,8 @@ The envisioned approach is that hub nodes near the root — which carry the
|
||||
largest downward filters — would use larger size classes, while leaf nodes
|
||||
and resource-constrained nodes continue with smaller filters. A node
|
||||
receiving a filter larger than its own size class folds it down locally.
|
||||
The mechanism by which heterogeneous filter sizes propagate through the
|
||||
tree is a future design direction not specified in v1. See
|
||||
[IDEA-0043](../../ideas/IDEA-0043-heterogeneous-filter-propagation.md).
|
||||
The mechanism by which heterogeneous filter sizes propagate through
|
||||
the tree is a future design direction not specified in v1.
|
||||
|
||||
### Folding
|
||||
|
||||
@@ -335,6 +337,33 @@ The hash function design supports folding: membership tests at a smaller
|
||||
size use `hash(item, i) % smaller_bit_count`, which maps to the same bit
|
||||
positions that folding produces.
|
||||
|
||||
## Mesh Size Estimation
|
||||
|
||||
Each filter's saturation can be inverted into an estimated entry count
|
||||
via the standard formula `n ≈ -(m/k) · ln(1 − X/m)`, where `m` is the
|
||||
filter size in bits, `k` is the hash count, and `X` is the population
|
||||
count. Combining the parent's inbound filter with the children's
|
||||
inbound filters gives an estimate of the whole network: parent + each
|
||||
child's subtree are disjoint by construction, and adding 1 for the
|
||||
node itself yields the total. The result is cached on the node and
|
||||
exposed through the control socket and `fipstop` dashboard.
|
||||
|
||||
The estimator refuses to produce a value when any contributing filter
|
||||
is above the antipoison FPR cap (`node.bloom.max_inbound_fpr`,
|
||||
default `0.05`); a partial aggregate would silently underestimate.
|
||||
Consumers handle the resulting `None` by displaying an "unknown"
|
||||
state rather than a misleading number.
|
||||
|
||||
## Antipoison: Inbound FPR Cap
|
||||
|
||||
Inbound `FilterAnnounce` payloads are checked against
|
||||
`node.bloom.max_inbound_fpr` (default `0.05`). Filters whose
|
||||
estimated false positive rate exceeds the cap are dropped silently
|
||||
(no NACK on the wire) — they would otherwise inflate downstream
|
||||
candidate evaluation cost without contributing useful discrimination.
|
||||
The cap also gates filters from feeding into mesh size estimation,
|
||||
as described above.
|
||||
|
||||
## Implementation Status
|
||||
|
||||
| Feature | Status |
|
||||
@@ -349,6 +378,8 @@ positions that folding produces.
|
||||
| 500ms rate limiting | **Implemented** |
|
||||
| FilterAnnounce gossip (all peers) | **Implemented** |
|
||||
| Filter cardinality logging | **Implemented** |
|
||||
| Mesh size estimation (parent + children + 1) | **Implemented** |
|
||||
| Inbound FPR cap (antipoison) | **Implemented** |
|
||||
| Size class negotiation | Future direction |
|
||||
| Folding support | Future direction |
|
||||
| Adaptive filter sizing | Future direction |
|
||||
@@ -357,6 +388,7 @@ positions that folding produces.
|
||||
|
||||
- [fips-mesh-operation.md](fips-mesh-operation.md) — How bloom filters fit
|
||||
into routing
|
||||
- [fips-wire-formats.md](fips-wire-formats.md) — FilterAnnounce wire format
|
||||
- [../reference/wire-formats.md](../reference/wire-formats.md) —
|
||||
FilterAnnounce wire format
|
||||
- [fips-spanning-tree.md](fips-spanning-tree.md) — The coordinate system
|
||||
that bloom filter candidates are ranked by
|
||||
|
||||
@@ -0,0 +1,123 @@
|
||||
# FIPS Concepts
|
||||
|
||||
A novice-friendly introduction to what FIPS is, why it exists, and the
|
||||
mental model behind a self-organizing mesh. For the protocol stack,
|
||||
identity system, and encryption walkthrough, see
|
||||
[fips-architecture.md](fips-architecture.md). For prior art and
|
||||
academic citations, see [fips-prior-work.md](fips-prior-work.md).
|
||||
|
||||
## What is FIPS?
|
||||
|
||||
FIPS is a self-organizing mesh network that can operate natively over a
|
||||
variety of physical and logical media, such as local area networks,
|
||||
Bluetooth, serial links, or the existing internet as an overlay. The
|
||||
long-term goal is infrastructure that can function alongside or
|
||||
ultimately replace dependence on the Internet itself. Systems running
|
||||
FIPS establish peer connections, authenticate each other, and route
|
||||
traffic for each other without any central authority or global topology
|
||||
knowledge, and allow end-to-end encrypted sessions between any two
|
||||
nodes regardless of how many hops separate them.
|
||||
|
||||
Nodes in the mesh route traffic for each other using Nostr identities
|
||||
(npubs) as network addresses. Applications can access the mesh through
|
||||
a native FIPS datagram service, or through an IPv6 adaptation layer
|
||||
that presents each node as an IPv6 endpoint for compatibility with
|
||||
existing IP-based applications.
|
||||
|
||||
## Why FIPS?
|
||||
|
||||
**Self-sovereign identity**: FIPS nodes generate their own addresses,
|
||||
node IDs, and security credentials without coordination with any
|
||||
central authority. These identities can be long-term fixed or may be
|
||||
ephemeral, changed at any time. These identities are not visible to
|
||||
the FIPS network itself — they are used only at the application layer
|
||||
and for end-to-end session encryption.
|
||||
|
||||
**Infrastructure independence**: The internet depends on centralized
|
||||
infrastructure — ISPs, backbone providers, DNS, certificate
|
||||
authorities. FIPS works over any transport that can carry packets: a
|
||||
serial connection, onion-routed connections through Tor, local area
|
||||
networking, radio links between remote sites, or the existing internet
|
||||
as an overlay. When the internet is unavailable, unreliable, or
|
||||
untrusted, the mesh still works.
|
||||
|
||||
**Privacy by design**: FIPS provides secure, authenticated, and
|
||||
encrypted communication between any two nodes in the mesh, independent
|
||||
of the mix of transports used along the routed path between them.
|
||||
Furthermore, the mesh itself is designed to minimize metadata exposure
|
||||
— intermediate nodes route packets without learning the identities of
|
||||
the endpoints.
|
||||
|
||||
**Zero configuration**: Nodes discover each other and build routing
|
||||
automatically. Connect to one peer and you can reach the entire mesh.
|
||||
The network self-heals around failures and adapts to changing topology.
|
||||
|
||||
## A Self-Organizing Mesh
|
||||
|
||||
Traditional networks are built top-down. A central authority assigns
|
||||
addresses, configures routing tables, provisions hardware, and manages
|
||||
the topology. If the authority disappears or the infrastructure fails,
|
||||
the network fails with it. Nodes cannot reach each other without
|
||||
infrastructure mediating the connection.
|
||||
|
||||
FIPS inverts this model. There is no central authority, no address
|
||||
assignment service, no routing table pushed from above. Each node
|
||||
generates its own identity from a cryptographic keypair. Each node
|
||||
independently decides which peers to connect to and which transports
|
||||
to use. From these local decisions alone, the network self-organizes:
|
||||
|
||||
- A **spanning tree** forms through distributed parent selection,
|
||||
giving every node a coordinate in the network without any node
|
||||
knowing the full topology
|
||||
- **Bloom filters** propagate through gossip, so each node learns
|
||||
which peers can reach which destinations — again without global
|
||||
knowledge
|
||||
- **Routing decisions** are made locally at each hop, using only the
|
||||
node's immediate peers and cached coordinate information
|
||||
|
||||
Each peer link and end-to-end session actively measures RTT, loss,
|
||||
jitter, and goodput through a lightweight in-band Metrics Measurement
|
||||
Protocol (MMP), providing operator visibility and a foundation for
|
||||
quality-aware routing.
|
||||
|
||||
The result is a network that builds itself from the bottom up, heals
|
||||
around failures automatically, and scales without central coordination.
|
||||
Adding a node is as simple as connecting to one existing peer — the
|
||||
network integrates the new node through its normal mesh protocols.
|
||||
|
||||
## Specific Design Goals
|
||||
|
||||
- **Nostr-native identity and cryptography** — Use Nostr keypairs as
|
||||
node identities and leverage secp256k1, Schnorr signatures, and
|
||||
SHA-256
|
||||
- **Transport agnostic** — Support overlay, shared medium, and
|
||||
point-to-point transports transparently
|
||||
- **Self-organizing** — Automatic topology discovery and route
|
||||
optimization
|
||||
- **Privacy preserving** — Minimize metadata leakage across untrusted
|
||||
links
|
||||
- **Resilient** — Self-healing with graceful degradation
|
||||
|
||||
Non-goals include:
|
||||
|
||||
- **Reliable delivery** — FIPS provides a best-effort datagram
|
||||
service; retransmission and ordering are left to applications or
|
||||
higher-layer protocols
|
||||
- **Anonymity** — Direct peers learn each other's identity; FIPS
|
||||
minimizes metadata exposure but is not an anonymity network like Tor
|
||||
- **Congestion control** — FIPS measures link quality but does not
|
||||
implement flow control or congestion avoidance at the mesh layer
|
||||
|
||||
## Where to Read Next
|
||||
|
||||
- [fips-architecture.md](fips-architecture.md) — protocol stack,
|
||||
identity system, two-layer encryption, MTU as a cross-cutting
|
||||
concern
|
||||
- [fips-spanning-tree.md](fips-spanning-tree.md) — how the tree forms
|
||||
and reconverges
|
||||
- [fips-bloom-filters.md](fips-bloom-filters.md) — how reachability
|
||||
information propagates
|
||||
- [fips-mesh-operation.md](fips-mesh-operation.md) — how the pieces
|
||||
work together at runtime
|
||||
- [fips-prior-work.md](fips-prior-work.md) — designs and protocols
|
||||
FIPS builds on
|
||||
@@ -1,604 +0,0 @@
|
||||
# FIPS Configuration
|
||||
|
||||
FIPS uses YAML-based configuration with a cascading multi-file priority system.
|
||||
All parameters have sensible defaults; a node can run with no configuration file
|
||||
at all (it will generate an ephemeral identity and listen on default addresses).
|
||||
|
||||
## Configuration Loading
|
||||
|
||||
### Search Paths
|
||||
|
||||
When started without the `-c` flag, FIPS searches for `fips.yaml` in these
|
||||
locations, lowest to highest priority:
|
||||
|
||||
| Priority | Path | Purpose |
|
||||
|----------|------|---------|
|
||||
| 1 (lowest) | `/etc/fips/fips.yaml` | System-wide defaults |
|
||||
| 2 | `~/.config/fips/fips.yaml` | User preferences |
|
||||
| 3 | `~/.fips.yaml` | Legacy user config |
|
||||
| 4 (highest) | `./fips.yaml` | Deployment-specific overrides |
|
||||
|
||||
All found files are loaded and merged in priority order. Values from higher
|
||||
priority files override those from lower priority files. This allows a system
|
||||
administrator to set site-wide defaults in `/etc/fips/fips.yaml` while
|
||||
individual deployments override specific values in `./fips.yaml`.
|
||||
|
||||
### CLI Option
|
||||
|
||||
```text
|
||||
fips -c /path/to/config.yaml
|
||||
```
|
||||
|
||||
When `-c` is specified, only that file is loaded (search paths are skipped).
|
||||
|
||||
### Partial Configuration
|
||||
|
||||
Every field has a built-in default. A configuration file only needs to specify
|
||||
values that differ from defaults. For example, a minimal config might contain
|
||||
only the identity and peer list, inheriting all other defaults.
|
||||
|
||||
## YAML Structure
|
||||
|
||||
The configuration is organized into five top-level sections:
|
||||
|
||||
```yaml
|
||||
node: # Node behavior, protocol parameters, and tuning
|
||||
tun: # TUN virtual interface
|
||||
dns: # DNS responder for .fips domain
|
||||
transports: # Network transports (UDP, Ethernet, Bluetooth, Tor, ...)
|
||||
peers: # Static peer list
|
||||
```
|
||||
|
||||
### Control Socket (`node.control.*`)
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `node.control.enabled` | bool | `true` | Enable the Unix domain control socket |
|
||||
| `node.control.socket_path` | string | *(auto)* | Socket file path. Default: `$XDG_RUNTIME_DIR/fips/control.sock`, then `/run/fips/control.sock` (if root), then `/tmp/fips-control.sock` |
|
||||
|
||||
The control socket provides read-only access to node state via the
|
||||
`fipsctl` command-line tool. See the project
|
||||
[README](../../README.md#inspect) for the command list.
|
||||
|
||||
All tunable protocol parameters live under `node.*`, organized as sysctl-style
|
||||
dotted paths. The top-level sections (`tun`, `dns`, `transports`, `peers`)
|
||||
handle infrastructure concerns only.
|
||||
|
||||
## Node Parameters (`node.*`)
|
||||
|
||||
### Identity (`node.identity.*`)
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `node.identity.nsec` | string | *(none)* | Secret key in nsec (bech32) or hex format. If omitted, behavior depends on `persistent`. |
|
||||
| `node.identity.persistent` | bool | `false` | Persist identity across restarts via key file. |
|
||||
|
||||
Identity resolution follows a three-tier priority:
|
||||
|
||||
1. **Explicit `nsec`** in config — always used when present, regardless of `persistent`
|
||||
2. **Persistent key file** — when `persistent: true` and no `nsec`, loads from `fips.key`
|
||||
adjacent to the config file; if no key file exists, generates a new keypair and saves it
|
||||
3. **Ephemeral** — when `persistent: false` (default) and no `nsec`, generates a fresh
|
||||
keypair on each start
|
||||
|
||||
Key files (`fips.key` with mode 0600, `fips.pub` with mode 0644) are written adjacent
|
||||
to the highest-priority config file for operator visibility, even in ephemeral mode.
|
||||
|
||||
### General
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `node.leaf_only` | bool | `false` | Leaf-only mode: node does not forward traffic or participate in routing |
|
||||
| `node.tick_interval_secs` | u64 | `1` | Periodic maintenance tick interval (retry checks, timeout cleanup, tree refresh) |
|
||||
| `node.base_rtt_ms` | u64 | `100` | Initial RTT estimate for new links before measurements converge |
|
||||
| `node.heartbeat_interval_secs` | u64 | `10` | Heartbeat send interval per peer for liveness detection |
|
||||
| `node.link_dead_timeout_secs` | u64 | `30` | No-traffic timeout before a peer is declared dead and removed |
|
||||
|
||||
### Resource Limits (`node.limits.*`)
|
||||
|
||||
Controls capacity for connections, peers, and links.
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `node.limits.max_connections` | usize | `256` | Max handshake-phase connections |
|
||||
| `node.limits.max_peers` | usize | `128` | Max authenticated peers |
|
||||
| `node.limits.max_links` | usize | `256` | Max active links |
|
||||
| `node.limits.max_pending_inbound` | usize | `1000` | Max pending inbound handshakes |
|
||||
|
||||
### Rate Limiting (`node.rate_limit.*`)
|
||||
|
||||
Handshake rate limiting protects against DoS on the Noise IK handshake path.
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `node.rate_limit.handshake_burst` | u32 | `100` | Token bucket burst capacity |
|
||||
| `node.rate_limit.handshake_rate` | f64 | `10.0` | Tokens per second refill rate |
|
||||
| `node.rate_limit.handshake_timeout_secs` | u64 | `30` | Stale handshake cleanup timeout |
|
||||
| `node.rate_limit.handshake_resend_interval_ms` | u64 | `1000` | Initial handshake message resend interval |
|
||||
| `node.rate_limit.handshake_resend_backoff` | f64 | `2.0` | Resend backoff multiplier (1s, 2s, 4s, 8s, 16s with defaults) |
|
||||
| `node.rate_limit.handshake_max_resends` | u32 | `5` | Max resends per handshake attempt |
|
||||
|
||||
### Retry / Backoff (`node.retry.*`)
|
||||
|
||||
Connection retry with exponential backoff.
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `node.retry.max_retries` | u32 | `5` | Max connection retry attempts |
|
||||
| `node.retry.base_interval_secs` | u64 | `5` | Base backoff interval |
|
||||
| `node.retry.max_backoff_secs` | u64 | `300` | Cap on exponential backoff (5 minutes) |
|
||||
|
||||
Auto-reconnect (triggered by MMP link-dead removal) uses the same backoff
|
||||
parameters but bypasses `max_retries`, retrying indefinitely. See
|
||||
`peers[].auto_reconnect` below.
|
||||
|
||||
### Cache Parameters (`node.cache.*`)
|
||||
|
||||
Controls caching of tree coordinates and identity mappings.
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `node.cache.coord_size` | usize | `50000` | Max entries in coordinate cache |
|
||||
| `node.cache.coord_ttl_secs` | u64 | `300` | Coordinate cache entry TTL (5 minutes) |
|
||||
| `node.cache.identity_size` | usize | `10000` | Max entries in identity cache (LRU, no TTL) |
|
||||
|
||||
### Discovery Protocol (`node.discovery.*`)
|
||||
|
||||
Controls flood-based node discovery (LookupRequest/LookupResponse).
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `node.discovery.ttl` | u8 | `64` | Hop limit for LookupRequest flood |
|
||||
| `node.discovery.timeout_secs` | u64 | `10` | Lookup completion timeout |
|
||||
| `node.discovery.recent_expiry_secs` | u64 | `10` | Dedup cache expiry for recent request IDs |
|
||||
|
||||
### Spanning Tree (`node.tree.*`)
|
||||
|
||||
Controls tree construction and parent selection.
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|----------------------------------------|-------|---------|--------------------------------------------------|
|
||||
| `node.tree.announce_min_interval_ms` | u64 | `500` | Per-peer TreeAnnounce rate limit |
|
||||
| `node.tree.parent_hysteresis` | f64 | `0.2` | Cost improvement fraction required for same-root parent switch (0.0–1.0) |
|
||||
| `node.tree.hold_down_secs` | u64 | `30` | Suppress non-mandatory re-evaluation after parent switch |
|
||||
| `node.tree.reeval_interval_secs` | u64 | `60` | Periodic cost-based parent re-evaluation interval (0 = disabled) |
|
||||
| `node.tree.flap_threshold` | u32 | `4` | Parent switches in window before dampening engages |
|
||||
| `node.tree.flap_window_secs` | u64 | `60` | Sliding window for counting parent switches |
|
||||
| `node.tree.flap_dampening_secs` | u64 | `120` | Extended hold-down duration when flap threshold exceeded |
|
||||
|
||||
### Bloom Filter (`node.bloom.*`)
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `node.bloom.update_debounce_ms` | u64 | `500` | Debounce interval for filter update propagation |
|
||||
|
||||
Bloom filter size (1 KB), hash count (5), and size classes are protocol
|
||||
constants and not configurable.
|
||||
|
||||
### ECN Signaling (`node.ecn.*`)
|
||||
|
||||
Controls hop-by-hop ECN (Explicit Congestion Notification) signaling. When
|
||||
enabled, transit nodes detect congestion on outgoing links (via MMP loss/ETX
|
||||
metrics or kernel buffer drops) and set the CE flag on forwarded FMP frames.
|
||||
Destination nodes mark ECN-capable IPv6 packets with CE before TUN delivery
|
||||
per RFC 3168, enabling end-host TCP congestion control to react.
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `node.ecn.enabled` | bool | `true` | Enable ECN congestion signaling (CE flag relay and local congestion detection) |
|
||||
| `node.ecn.loss_threshold` | f64 | `0.05` | MMP loss rate threshold for CE marking (0.0–1.0). When the outgoing link's loss rate meets or exceeds this value, forwarded packets are CE-marked. |
|
||||
| `node.ecn.etx_threshold` | f64 | `3.0` | MMP ETX threshold for CE marking (≥1.0). When the outgoing link's ETX meets or exceeds this value, forwarded packets are CE-marked. |
|
||||
|
||||
Congestion detection triggers on any of: outgoing link loss ≥ `loss_threshold`,
|
||||
outgoing link ETX ≥ `etx_threshold`, or kernel receive buffer drops detected on
|
||||
any local transport. CE is relayed hop-by-hop: once set on any hop, the flag
|
||||
stays set for all subsequent hops to the destination.
|
||||
|
||||
### Rekey (`node.rekey.*`)
|
||||
|
||||
Controls periodic Noise rekey for forward secrecy. When enabled, both FMP
|
||||
(link-layer IK) and FSP (session-layer XK) sessions perform fresh Diffie-Hellman
|
||||
key exchanges after a time or message count threshold, whichever comes first.
|
||||
A 10-second drain window keeps the old session active for decryption during
|
||||
cutover.
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `node.rekey.enabled` | bool | `true` | Enable periodic Noise rekey on all links and sessions |
|
||||
| `node.rekey.after_secs` | u64 | `120` | Initiate rekey after this many seconds on a session |
|
||||
| `node.rekey.after_messages` | u64 | `65536` | Initiate rekey after this many messages sent on a session |
|
||||
|
||||
### Session / Data Plane (`node.session.*`)
|
||||
|
||||
Controls end-to-end session behavior and packet queuing.
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `node.session.default_ttl` | u8 | `64` | Default SessionDatagram TTL |
|
||||
| `node.session.pending_packets_per_dest` | usize | `16` | Queue depth per destination during session establishment |
|
||||
| `node.session.pending_max_destinations` | usize | `256` | Max destinations with pending packets |
|
||||
| `node.session.idle_timeout_secs` | u64 | `90` | Idle session timeout; established sessions with no application data for this duration are removed. MMP reports (SenderReport, ReceiverReport, PathMtuNotification) do not count as activity |
|
||||
| `node.session.coords_warmup_packets` | u8 | `5` | Number of initial data packets per session that include the CP flag for transit cache warmup; also the reset count on CoordsRequired/PathBroken receipt |
|
||||
| `node.session.coords_response_interval_ms` | u64 | `2000` | Minimum interval (ms) between standalone CoordsWarmup responses to CoordsRequired/PathBroken signals per destination |
|
||||
|
||||
The anti-replay window size (2048 packets) is a compile-time constant and not
|
||||
configurable.
|
||||
|
||||
### Link-Layer MMP (`node.mmp.*`)
|
||||
|
||||
Metrics Measurement Protocol for per-peer link measurement. See
|
||||
[fips-mesh-layer.md](fips-mesh-layer.md) for behavioral details.
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `node.mmp.mode` | string | `"full"` | Operating mode: `full` (sender + receiver reports), `lightweight` (receiver reports only), or `minimal` (spin bit + CE echo only, no reports) |
|
||||
| `node.mmp.log_interval_secs` | u64 | `30` | Periodic operator log interval for link metrics |
|
||||
| `node.mmp.owd_window_size` | usize | `32` | One-way delay trend ring buffer size |
|
||||
|
||||
### Session-Layer MMP (`node.session_mmp.*`)
|
||||
|
||||
Metrics Measurement Protocol for end-to-end session measurement. Configured
|
||||
independently from link-layer MMP because session reports are routed through
|
||||
every transit link, consuming bandwidth proportional to path length.
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `node.session_mmp.mode` | string | `"full"` | Operating mode: `full`, `lightweight`, or `minimal` |
|
||||
| `node.session_mmp.log_interval_secs` | u64 | `30` | Periodic operator log interval for session metrics |
|
||||
| `node.session_mmp.owd_window_size` | usize | `32` | One-way delay trend ring buffer size |
|
||||
|
||||
### Internal Buffers (`node.buffers.*`)
|
||||
|
||||
Channel sizes affecting throughput and memory. Primarily useful for performance
|
||||
tuning under high load or on memory-constrained devices.
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `node.buffers.packet_channel` | usize | `1024` | Transport to Node packet channel capacity |
|
||||
| `node.buffers.tun_channel` | usize | `1024` | TUN to Node outbound channel capacity |
|
||||
| `node.buffers.dns_channel` | usize | `64` | DNS to Node identity channel capacity |
|
||||
|
||||
## TUN Interface (`tun.*`)
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `tun.enabled` | bool | `false` | Enable TUN virtual interface |
|
||||
| `tun.name` | string | `"fips0"` | Interface name |
|
||||
| `tun.mtu` | u16 | `1280` | Interface MTU (IPv6 minimum) |
|
||||
|
||||
## DNS Responder (`dns.*`)
|
||||
|
||||
Resolves `<npub>.fips` queries to FIPS IPv6 addresses. Resolution is pure
|
||||
computation (npub to public key to address); resolved identities are registered
|
||||
with the node for routing.
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `dns.enabled` | bool | `true` | Enable DNS responder |
|
||||
| `dns.bind_addr` | string | `"127.0.0.1"` | Bind address |
|
||||
| `dns.port` | u16 | `5354` | Listen port |
|
||||
| `dns.ttl` | u32 | `300` | AAAA record TTL in seconds |
|
||||
|
||||
The `dns.ttl` value should not exceed `node.cache.coord_ttl_secs` to avoid
|
||||
stale address mappings.
|
||||
|
||||
### Host Mapping
|
||||
|
||||
The DNS resolver checks a host map before falling back to direct npub
|
||||
resolution, enabling names like `gateway.fips` instead of `npub1...fips`.
|
||||
The host map is populated from two sources:
|
||||
|
||||
1. **Peer aliases** — the `alias` field on configured peers in `peers:`.
|
||||
2. **Hosts file** — `/etc/fips/hosts`, one `hostname npub1...` per line.
|
||||
Blank lines and `#` comments are allowed.
|
||||
|
||||
The hosts file is auto-reloaded on modification (mtime change) without
|
||||
restarting the daemon. Hostnames are case-insensitive.
|
||||
|
||||
## Transports (`transports.*`)
|
||||
|
||||
### UDP (`transports.udp.*`)
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `transports.udp.bind_addr` | string | `"0.0.0.0:2121"` | UDP bind address and port |
|
||||
| `transports.udp.mtu` | u16 | `1280` | Transport MTU |
|
||||
| `transports.udp.recv_buf_size` | usize | `2097152` | UDP socket receive buffer size in bytes (2 MB). Linux kernel doubles the requested value internally. Host `net.core.rmem_max` must be >= this value. |
|
||||
| `transports.udp.send_buf_size` | usize | `2097152` | UDP socket send buffer size in bytes (2 MB). Host `net.core.wmem_max` must be >= this value. |
|
||||
|
||||
### Ethernet (`transports.ethernet.*`)
|
||||
|
||||
Ethernet transport sends raw frames via AF_PACKET SOCK_DGRAM sockets.
|
||||
Requires `CAP_NET_RAW` or running as root. Linux only.
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `interface` | string | *(required)* | Network interface name (e.g., `"eth0"`, `"enp3s0"`) |
|
||||
| `ethertype` | u16 | `0x2121` | EtherType |
|
||||
| `mtu` | u16 | *(auto)* | Override MTU. Default: interface MTU minus 3 (for frame type + length prefix) |
|
||||
| `recv_buf_size` | usize | `2097152` | Socket receive buffer size in bytes (2 MB) |
|
||||
| `send_buf_size` | usize | `2097152` | Socket send buffer size in bytes (2 MB) |
|
||||
| `discovery` | bool | `true` | Listen for discovery beacons from other nodes |
|
||||
| `announce` | bool | `false` | Broadcast announcement beacons on the LAN |
|
||||
| `auto_connect` | bool | `false` | Auto-connect to discovered peers |
|
||||
| `accept_connections` | bool | `false` | Accept incoming connection attempts from discovered peers |
|
||||
| `beacon_interval_secs` | u64 | `30` | Announcement beacon interval in seconds (minimum 10) |
|
||||
|
||||
**Named instances.** Multiple Ethernet interfaces can be configured by
|
||||
using named sub-keys instead of flat parameters:
|
||||
|
||||
```yaml
|
||||
transports:
|
||||
ethernet:
|
||||
lan:
|
||||
interface: "eth0"
|
||||
discovery: true
|
||||
announce: true
|
||||
backbone:
|
||||
interface: "eth1"
|
||||
announce: false
|
||||
```
|
||||
|
||||
Each named instance operates independently with its own socket and
|
||||
discovery state. The instance name is used in log messages and the
|
||||
`name()` method on the Transport trait.
|
||||
|
||||
### TCP (`transports.tcp.*`)
|
||||
|
||||
TCP transport enables firewall traversal on networks that block UDP but
|
||||
allow TCP (e.g., port 443). Uses FMP header-based framing with zero
|
||||
overhead.
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `transports.tcp.bind_addr` | string | *(none)* | Listen address (e.g., `"0.0.0.0:8443"`). If omitted, outbound-only mode. |
|
||||
| `transports.tcp.mtu` | u16 | `1400` | Default MTU. Per-connection MTU derived from `TCP_MAXSEG` when available. |
|
||||
| `transports.tcp.connect_timeout_ms` | u64 | `5000` | Outbound connect timeout in milliseconds |
|
||||
| `transports.tcp.nodelay` | bool | `true` | `TCP_NODELAY` (disable Nagle for low latency) |
|
||||
| `transports.tcp.keepalive_secs` | u64 | `30` | TCP keepalive interval in seconds (0 = disabled) |
|
||||
| `transports.tcp.recv_buf_size` | usize | `2097152` | Socket receive buffer size in bytes (2 MB) |
|
||||
| `transports.tcp.send_buf_size` | usize | `2097152` | Socket send buffer size in bytes (2 MB) |
|
||||
| `transports.tcp.max_inbound_connections` | usize | `256` | Maximum simultaneous inbound connections |
|
||||
| `transports.tcp.socks5_proxy` | string | *(none)* | SOCKS5 proxy for outbound connections (implementation deferred) |
|
||||
|
||||
**Named instances.** Like other transports, multiple TCP instances can
|
||||
be configured with named sub-keys:
|
||||
|
||||
```yaml
|
||||
transports:
|
||||
tcp:
|
||||
public:
|
||||
bind_addr: "0.0.0.0:443"
|
||||
tor:
|
||||
socks5_proxy: "127.0.0.1:9050"
|
||||
connect_timeout_ms: 30000
|
||||
```
|
||||
|
||||
## Peers (`peers[]`)
|
||||
|
||||
Static peer list. Each entry defines a peer to connect to.
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `peers[].npub` | string | *(required)* | Peer's Nostr public key (npub-encoded) |
|
||||
| `peers[].alias` | string | *(none)* | Human-readable name for logging |
|
||||
| `peers[].addresses[].transport` | string | *(required)* | Transport type: `udp`, `tcp`, or `ethernet` |
|
||||
| `peers[].addresses[].addr` | string | *(required)* | Transport address. UDP/TCP: `"ip:port"`. Ethernet: `"interface/mac"` (e.g., `"eth0/aa:bb:cc:dd:ee:ff"`) |
|
||||
| `peers[].addresses[].priority` | u8 | `100` | Address priority (lower = preferred) |
|
||||
| `peers[].connect_policy` | string | `"auto_connect"` | Connection policy: `auto_connect`, `on_demand`, or `manual` |
|
||||
| `peers[].auto_reconnect` | bool | `true` | Automatically reconnect after MMP link-dead removal (exponential backoff, unlimited retries) |
|
||||
|
||||
## Minimal Example
|
||||
|
||||
A typical node configuration enabling TUN, DNS, and a single peer:
|
||||
|
||||
```yaml
|
||||
node:
|
||||
identity:
|
||||
nsec: "0102030405060708090a0b0c0d0e0f101112131415161718191a1b1c1d1e1f20"
|
||||
|
||||
tun:
|
||||
enabled: true
|
||||
name: fips0
|
||||
mtu: 1280
|
||||
|
||||
dns:
|
||||
enabled: true
|
||||
bind_addr: "127.0.0.1"
|
||||
port: 53
|
||||
|
||||
transports:
|
||||
udp:
|
||||
bind_addr: "0.0.0.0:2121"
|
||||
mtu: 1472
|
||||
|
||||
peers:
|
||||
- npub: "npub1tdwa4vjrjl33pcjdpf2t4p027nl86xrx24g4d3avg4vwvayr3g8qhd84le"
|
||||
alias: "node-b"
|
||||
addresses:
|
||||
- transport: udp
|
||||
addr: "172.20.0.11:2121"
|
||||
connect_policy: auto_connect
|
||||
```
|
||||
|
||||
### Mixed UDP + Ethernet Example
|
||||
|
||||
A node bridging internet peers (UDP) and a local Ethernet segment with
|
||||
beacon discovery:
|
||||
|
||||
```yaml
|
||||
node:
|
||||
identity:
|
||||
nsec: "0102030405060708090a0b0c0d0e0f101112131415161718191a1b1c1d1e1f20"
|
||||
|
||||
tun:
|
||||
enabled: true
|
||||
|
||||
transports:
|
||||
udp:
|
||||
bind_addr: "0.0.0.0:2121"
|
||||
mtu: 1472
|
||||
ethernet:
|
||||
interface: "eth0"
|
||||
discovery: true
|
||||
announce: true
|
||||
auto_connect: true
|
||||
accept_connections: true
|
||||
|
||||
peers:
|
||||
- npub: "npub1tdwa4vjrjl33pcjdpf2t4p027nl86xrx24g4d3avg4vwvayr3g8qhd84le"
|
||||
alias: "internet-peer"
|
||||
addresses:
|
||||
- transport: udp
|
||||
addr: "203.0.113.5:2121"
|
||||
connect_policy: auto_connect
|
||||
```
|
||||
|
||||
Ethernet peers on the local segment are discovered automatically via
|
||||
beacons — no static peer entries needed. Internet peers still require
|
||||
explicit configuration.
|
||||
|
||||
All `node.*` parameters use their defaults. To override specific values, add
|
||||
only the relevant sections:
|
||||
|
||||
```yaml
|
||||
node:
|
||||
identity:
|
||||
nsec: "..."
|
||||
limits:
|
||||
max_peers: 64
|
||||
retry:
|
||||
max_retries: 10
|
||||
max_backoff_secs: 600
|
||||
cache:
|
||||
coord_size: 100000
|
||||
```
|
||||
|
||||
## Complete Reference
|
||||
|
||||
The full YAML structure with all defaults:
|
||||
|
||||
```yaml
|
||||
node:
|
||||
identity:
|
||||
nsec: null # secret key in nsec or hex (null = depends on persistent)
|
||||
persistent: false # true = load/save fips.key; false = ephemeral each start
|
||||
leaf_only: false
|
||||
tick_interval_secs: 1
|
||||
base_rtt_ms: 100
|
||||
heartbeat_interval_secs: 10
|
||||
link_dead_timeout_secs: 30
|
||||
limits:
|
||||
max_connections: 256
|
||||
max_peers: 128
|
||||
max_links: 256
|
||||
max_pending_inbound: 1000
|
||||
rate_limit:
|
||||
handshake_burst: 100
|
||||
handshake_rate: 10.0
|
||||
handshake_timeout_secs: 30
|
||||
handshake_resend_interval_ms: 1000
|
||||
handshake_resend_backoff: 2.0
|
||||
handshake_max_resends: 5
|
||||
retry:
|
||||
max_retries: 5
|
||||
base_interval_secs: 5
|
||||
max_backoff_secs: 300
|
||||
cache:
|
||||
coord_size: 50000
|
||||
coord_ttl_secs: 300
|
||||
identity_size: 10000
|
||||
discovery:
|
||||
ttl: 64
|
||||
timeout_secs: 10
|
||||
recent_expiry_secs: 10
|
||||
tree:
|
||||
announce_min_interval_ms: 500
|
||||
parent_hysteresis: 0.2 # cost improvement fraction for parent switch
|
||||
hold_down_secs: 30 # suppress re-evaluation after switch
|
||||
reeval_interval_secs: 60 # periodic cost-based re-evaluation (0 = disabled)
|
||||
flap_threshold: 4 # parent switches before dampening
|
||||
flap_window_secs: 60 # sliding window for flap detection
|
||||
flap_dampening_secs: 120 # extended hold-down on flap
|
||||
bloom:
|
||||
update_debounce_ms: 500
|
||||
session:
|
||||
default_ttl: 64
|
||||
pending_packets_per_dest: 16
|
||||
pending_max_destinations: 256
|
||||
idle_timeout_secs: 90
|
||||
coords_warmup_packets: 5
|
||||
coords_response_interval_ms: 2000
|
||||
mmp:
|
||||
mode: full # full | lightweight | minimal
|
||||
log_interval_secs: 30
|
||||
owd_window_size: 32
|
||||
session_mmp:
|
||||
mode: full # full | lightweight | minimal
|
||||
log_interval_secs: 30
|
||||
owd_window_size: 32
|
||||
ecn:
|
||||
enabled: true # ECN congestion signaling (CE flag relay)
|
||||
loss_threshold: 0.05 # MMP loss rate threshold for CE marking (5%)
|
||||
etx_threshold: 3.0 # MMP ETX threshold for CE marking
|
||||
rekey:
|
||||
enabled: true # periodic Noise rekey for forward secrecy
|
||||
after_secs: 120 # rekey interval (seconds)
|
||||
after_messages: 65536 # rekey after N messages sent
|
||||
control:
|
||||
enabled: true
|
||||
socket_path: null # null = auto ($XDG_RUNTIME_DIR → /run/fips → /tmp fallback)
|
||||
buffers:
|
||||
packet_channel: 1024
|
||||
tun_channel: 1024
|
||||
dns_channel: 64
|
||||
|
||||
tun:
|
||||
enabled: false
|
||||
name: "fips0"
|
||||
mtu: 1280
|
||||
|
||||
dns:
|
||||
enabled: true
|
||||
bind_addr: "127.0.0.1"
|
||||
port: 5354
|
||||
ttl: 300
|
||||
|
||||
transports:
|
||||
udp:
|
||||
bind_addr: "0.0.0.0:2121"
|
||||
mtu: 1280
|
||||
recv_buf_size: 2097152 # 2 MB (kernel doubles to 4 MB actual)
|
||||
send_buf_size: 2097152 # 2 MB
|
||||
# ethernet: # uncomment to enable (requires CAP_NET_RAW)
|
||||
# interface: "eth0" # required: network interface name
|
||||
# ethertype: 0x2121 # default EtherType
|
||||
# mtu: null # null = interface MTU - 3 (typically 1497)
|
||||
# recv_buf_size: 2097152 # 2 MB
|
||||
# send_buf_size: 2097152 # 2 MB
|
||||
# discovery: true # listen for beacons
|
||||
# announce: false # broadcast beacons
|
||||
# auto_connect: false # connect to discovered peers
|
||||
# accept_connections: false # accept inbound handshakes
|
||||
# beacon_interval_secs: 30 # beacon interval (min 10)
|
||||
# tcp: # uncomment to enable TCP transport
|
||||
# bind_addr: "0.0.0.0:8443" # listen address (omit for outbound-only)
|
||||
# mtu: 1400 # default MTU
|
||||
# connect_timeout_ms: 5000 # outbound connect timeout
|
||||
# nodelay: true # TCP_NODELAY
|
||||
# keepalive_secs: 30 # keepalive interval (0 = disabled)
|
||||
# recv_buf_size: 2097152 # 2 MB
|
||||
# send_buf_size: 2097152 # 2 MB
|
||||
# max_inbound_connections: 256 # resource protection limit
|
||||
# socks5_proxy: null # SOCKS5 for outbound (deferred)
|
||||
|
||||
peers: # static peer list
|
||||
# - npub: "npub1..."
|
||||
# alias: "node-b"
|
||||
# addresses:
|
||||
# - transport: udp
|
||||
# addr: "10.0.0.2:2121"
|
||||
# priority: 100
|
||||
# connect_policy: auto_connect
|
||||
# auto_reconnect: true # reconnect after link-dead removal
|
||||
```
|
||||
@@ -0,0 +1,541 @@
|
||||
# FIPS Gateway
|
||||
|
||||
The FIPS gateway lets unmodified IPv6 hosts on a LAN exchange traffic
|
||||
with the mesh without running any FIPS software themselves. It is a
|
||||
niche feature — most operators will never enable it. The gateway
|
||||
runs most conveniently on a system that is already providing network
|
||||
services (DHCP, DNS, RA) to a LAN segment, since hosts on that
|
||||
segment already get IP assignment and a default route from that box.
|
||||
The canonical example is an OpenWrt-based WiFi access point: every
|
||||
client that associates with the AP already has the AP as default
|
||||
router and DNS server, which is exactly the placement the gateway
|
||||
needs. The OpenWrt ipk ships with the `gateway:` block of
|
||||
`/etc/fips/fips.yaml` pre-populated and the integration glue
|
||||
(dnsmasq forwarding, RA route for the virtual pool, global-scope
|
||||
IPv6 prefix on `br-lan`) automated by the init script —
|
||||
[`packaging/openwrt-ipk/files/etc/init.d/fips-gateway`](https://github.com/jmcorgan/fips/blob/master/packaging/openwrt-ipk/files/etc/init.d/fips-gateway).
|
||||
The operator only needs to enable and start the service. Running the
|
||||
gateway on a non-OpenWrt LAN-edge host (a Linux router/server, for
|
||||
example) is technically possible but requires manual integration:
|
||||
distributing a route to the virtual-IP pool, wiring DNS forwarding so
|
||||
LAN clients send `.fips` queries to the gateway, configuring sysctls
|
||||
and capabilities. That path is supported but tedious; it is the
|
||||
secondary path.
|
||||
|
||||
The feature has two halves that share common machinery and have
|
||||
their own unique parts.
|
||||
|
||||
The **outbound half** carries traffic from LAN to mesh. A non-FIPS
|
||||
LAN workstation resolves `<npub>.fips` (or a `.fips` host alias) via
|
||||
the gateway's DNS proxy, which returns a virtual IPv6 address from a
|
||||
managed pool. The kernel routes the LAN packet to that virtual IP via
|
||||
a route to the pool CIDR (RA-advertised, statically distributed, or
|
||||
on-link via the default route). The gateway runs nftables NAT so the
|
||||
packet appears on the mesh as if it had originated from the gateway's
|
||||
own FIPS identity: prerouting DNAT rewrites the destination from the
|
||||
virtual IP to the real `fd00::/8` mesh address, and postrouting
|
||||
masquerade rewrites the source from the LAN host's address to the
|
||||
gateway's `fips0` address. Return traffic follows the conntrack
|
||||
reverse path back to the originating LAN host, with postrouting SNAT
|
||||
restoring the virtual IP as source so the client sees a response from
|
||||
the address it connected to.
|
||||
|
||||
The **inbound half** carries traffic from mesh to LAN. A
|
||||
configuration entry in `gateway.port_forwards[]` exposes a LAN
|
||||
service (`host:port`) on a port of the gateway's mesh-side `fips0`
|
||||
address. Mesh peers reach it as `<gateway-npub>.fips:<listen_port>`.
|
||||
A prerouting DNAT rule keyed on `(iif=fips0, l4proto, dport)`
|
||||
rewrites the destination to the LAN target; a LAN-side masquerade in
|
||||
postrouting rewrites the mesh peer's source so the LAN target sees a
|
||||
reachable LAN address and conntrack steers replies back through the
|
||||
gateway. This is the inverse of port-forwarding on a conventional NAT
|
||||
router.
|
||||
|
||||
The two halves are independent and can be configured separately.
|
||||
Inbound port-forwards work without any outbound configuration (just
|
||||
a port-forward list and the table); outbound works without any
|
||||
inbound forwards. They share the same nftables table, the same
|
||||
binary, the same control socket, and the same atomic-rebuild
|
||||
strategy. That shared machinery is what makes them halves of one
|
||||
feature rather than two separate features.
|
||||
|
||||
## Architecture
|
||||
|
||||
### The `fips-gateway` Service
|
||||
|
||||
The gateway is a separate binary, [`fips-gateway`](https://github.com/jmcorgan/fips/blob/master/src/bin/fips-gateway.rs),
|
||||
not part of the FIPS daemon. It reads the same `/etc/fips/fips.yaml`
|
||||
the daemon reads (via `--config`, or the standard search path), but
|
||||
acts on the `gateway.*` block. It needs `CAP_NET_ADMIN` to install
|
||||
nftables rules, manage proxy NDP entries, and add the pool route.
|
||||
The CLI is documented in
|
||||
[../reference/cli-fips-gateway.md](../reference/cli-fips-gateway.md).
|
||||
|
||||
The gateway connects to the daemon indirectly. The outbound half
|
||||
forwards `.fips` DNS queries to the daemon's built-in resolver
|
||||
(default `[::1]:5354`); the daemon resolves the name to a mesh
|
||||
address and primes its identity cache as a side effect. The inbound
|
||||
half does not require any daemon plumbing at all — packets that
|
||||
arrive on `fips0` after the daemon's TUN injection path are matched
|
||||
by the nftables rules on `fips0` ingress. There is no shared memory,
|
||||
no IPC channel, and no startup ordering coupling beyond "the daemon's
|
||||
DNS responder must be reachable before the gateway starts serving
|
||||
LAN queries", which the gateway enforces with a bounded reachability
|
||||
probe at startup.
|
||||
|
||||
### nftables Table Layout
|
||||
|
||||
All gateway rules live in a single nftables table, `inet
|
||||
fips_gateway`, with two chains:
|
||||
|
||||
- `prerouting` — `type nat hook prerouting priority dstnat (-100)`,
|
||||
for both LAN→mesh DNAT (per virtual-IP mapping) and mesh→LAN DNAT
|
||||
(per port-forward).
|
||||
- `postrouting` — `type nat hook postrouting priority srcnat (100)`,
|
||||
for both the always-on `oifname fips0` masquerade, the per-mapping
|
||||
return-path SNAT, and (when any port-forward is configured) the
|
||||
LAN-side masquerade for inbound traffic.
|
||||
|
||||
The table is rebuilt atomically on every change. The rebuild
|
||||
sequence — delete the existing table (ignore `ENOENT` on first
|
||||
call), then create a new table with chains and the full rule set in
|
||||
a single netlink batch — avoids reliance on kernel rule-handle
|
||||
tracking, which the rustables crate does not expose. The table stays
|
||||
small (one always-on masquerade plus two rules per active outbound
|
||||
mapping plus one rule per inbound forward, with one extra masquerade
|
||||
when any forward is present), so rebuilds are cheap.
|
||||
|
||||
### Control Socket
|
||||
|
||||
`fips-gateway` exposes a Unix-domain control socket at
|
||||
`/run/fips/gateway.sock` (`root:fips`, mode `0770`) with two
|
||||
commands: `show_gateway` and `show_mappings`. The protocol is the
|
||||
same line-delimited JSON used by the daemon's control socket. The
|
||||
shapes are documented in the
|
||||
[Gateway command catalog](../reference/control-socket.md#gateway-command-catalog).
|
||||
There is no `fipsctl gateway` subcommand; clients (including
|
||||
`fipstop`'s gateway view) talk to the socket directly.
|
||||
|
||||
### Diagram
|
||||
|
||||
```text
|
||||
LAN clients
|
||||
│
|
||||
DNS query (.fips) │ IPv6 packet
|
||||
for outbound │ to virtual IP
|
||||
│ or mesh peer
|
||||
▼
|
||||
┌───────────────────────────────────┐
|
||||
│ fips-gateway │
|
||||
│ │
|
||||
│ ┌──────────────┐ ┌───────────┐ │
|
||||
│ │ DNS proxy │ │ Virtual │ │
|
||||
│ │ ([::1]:5353) │─▶│ IP pool │ │
|
||||
│ │ .fips only │ │ (state │ │
|
||||
│ └──────┬───────┘ │ machine) │ │
|
||||
│ │ └─────┬─────┘ │
|
||||
│ │ │ │
|
||||
│ forward to │ pool │
|
||||
│ daemon resolver │ events │
|
||||
│ ([::1]:5354) ▼ │
|
||||
│ │ ┌───────────┐ │
|
||||
│ │ │ NAT │ │
|
||||
│ │ │ manager │ │
|
||||
│ │ │ (rebuild │ │
|
||||
│ │ │ inet │ │
|
||||
│ │ │ fips_ │ │
|
||||
│ │ │ gateway) │ │
|
||||
│ │ └─────┬─────┘ │
|
||||
│ │ │ │
|
||||
│ │ ┌─────▼─────┐ │
|
||||
│ │ │ net │ │
|
||||
│ │ │ setup │ │
|
||||
│ │ │ (proxy │ │
|
||||
│ │ │ NDP, lo │ │
|
||||
│ │ │ route) │ │
|
||||
│ │ └───────────┘ │
|
||||
│ │ │
|
||||
│ │ control socket │
|
||||
│ │ /run/fips/ │
|
||||
│ │ gateway.sock │
|
||||
└─────────┼─────────────────────────┘
|
||||
│
|
||||
▼
|
||||
FIPS daemon resolver
|
||||
([::1]:5354)
|
||||
│
|
||||
▼
|
||||
fips0 TUN interface
|
||||
│
|
||||
▼
|
||||
the mesh
|
||||
```
|
||||
|
||||
The DNS proxy and the virtual IP pool are exclusive to the outbound
|
||||
half. The NAT manager and the kernel-side machinery (nftables table,
|
||||
`fips0` and LAN interfaces, conntrack) are shared. The inbound half
|
||||
contributes per-port-forward rules to the same table without
|
||||
involving the DNS proxy or the pool.
|
||||
|
||||
## The Outbound Half (LAN → Mesh)
|
||||
|
||||
### DNS Resolution Flow
|
||||
|
||||
1. A LAN client sends a DNS query to the gateway's listener (default
|
||||
`[::1]:5353`, configurable via `gateway.dns.listen`). The default
|
||||
is loopback-only on an unprivileged port: the canonical deployment
|
||||
has another resolver on the host (dnsmasq, systemd-resolved, BIND)
|
||||
holding port 53 and forwarding `.fips` queries to the gateway over
|
||||
loopback. Operators on a host without a pre-existing resolver on
|
||||
53 can override the listen value to `"[::]:53"` to let LAN clients
|
||||
query the gateway directly.
|
||||
2. If the question is not for a `.fips` domain, the gateway replies
|
||||
`REFUSED`. The proxy is intentionally narrow — it does not resolve
|
||||
public DNS, and the LAN's primary resolver should hold port 53 on
|
||||
the gateway host (the OpenWrt init script wires dnsmasq to forward
|
||||
`.fips` queries to the loopback listener automatically).
|
||||
3. The gateway forwards the query to the daemon resolver
|
||||
(`gateway.dns.upstream`, default `[::1]:5354`). The daemon must
|
||||
match: an IPv6 socket bound to `[::1]` does not accept v4-mapped
|
||||
traffic, so a `127.0.0.1:5354` upstream cannot reach a daemon
|
||||
bound on `[::1]:5354`.
|
||||
4. If the daemon is unreachable or times out (5 s), the gateway
|
||||
replies `SERVFAIL`. If the daemon returns `NXDOMAIN` or a
|
||||
non-`AAAA` answer, the gateway forwards the response unchanged.
|
||||
5. The gateway extracts the AAAA (`fd00::/8`) record from the
|
||||
daemon's response. This resolution primes the daemon's identity
|
||||
cache as a side effect — a prerequisite for `fips0` routing,
|
||||
because the daemon needs the cache entry to map the mesh address
|
||||
back to a `NodeAddr` for forwarding.
|
||||
6. The gateway allocates a virtual IP from the pool for that mesh
|
||||
address (idempotent: an existing mapping is reused and its TTL
|
||||
refreshed).
|
||||
7. If a new mapping was created, the pool emits `MappingCreated`,
|
||||
which the main loop turns into `add_mapping` calls on the NAT
|
||||
manager and `add_proxy_ndp` on the network setup.
|
||||
8. The gateway returns an `AAAA` response containing the virtual IP,
|
||||
with the configured TTL (default 60 s).
|
||||
|
||||
### Virtual IP Pool
|
||||
|
||||
The pool allocates IPv6 addresses from a configured CIDR (default
|
||||
`fd01::/112`). Each address maps to one mesh destination, keyed by
|
||||
`NodeAddr` rather than by hostname — different `.fips` aliases for
|
||||
the same node share a virtual IP. Address 0 (the network-equivalent)
|
||||
is reserved; the rest are allocatable. The pool is capped at 2^16
|
||||
addresses regardless of prefix length, to bound memory.
|
||||
|
||||
The pool tracks state per address:
|
||||
|
||||
```text
|
||||
Allocated ──→ Active ──→ Draining ──→ Free
|
||||
│ ▲
|
||||
└──────────────────────────────────┘
|
||||
(TTL expired, no sessions)
|
||||
```
|
||||
|
||||
| State | Meaning |
|
||||
| ----- | ------- |
|
||||
| Allocated | DNS query created the mapping; no NAT sessions yet. |
|
||||
| Active | Conntrack reports at least one session for this virtual IP. |
|
||||
| Draining | TTL has expired; sessions may still be in progress, or grace period is running after sessions ended. |
|
||||
| Free | Reclaimed and available for new allocations. |
|
||||
|
||||
Transitions:
|
||||
|
||||
- **Allocated → Active**: conntrack sessions count goes above zero.
|
||||
- **Allocated → Free**: TTL expires before any session is ever
|
||||
observed.
|
||||
- **Active → Draining**: TTL expires (sessions may or may not still
|
||||
be present).
|
||||
- **Draining → Free**: session count is zero and the grace period
|
||||
has elapsed since draining began.
|
||||
|
||||
Timing:
|
||||
|
||||
- **TTL** (`gateway.dns.ttl`, default 60 s) is both the DNS TTL
|
||||
returned to the client and the mapping's idle lifetime. Repeated
|
||||
DNS queries for the same destination refresh the
|
||||
`last_referenced` timestamp.
|
||||
- **Grace period** (`gateway.pool_grace_period`, default 60 s) is
|
||||
the dwell time after the last session ends before the address is
|
||||
recycled. It prevents immediate reuse from confusing hosts with
|
||||
cached DNS responses.
|
||||
- **Tick interval**: the pool re-evaluates state every 10 s.
|
||||
|
||||
Active session counts come from `/proc/net/nf_conntrack`: an entry
|
||||
counts as a session if its original destination is the virtual IP.
|
||||
|
||||
If the pool is exhausted, new DNS queries return `SERVFAIL`.
|
||||
Existing mappings are never evicted prematurely — the correctness of
|
||||
in-flight sessions takes precedence over fresh allocations.
|
||||
|
||||
### NAT Pipeline (Outbound)
|
||||
|
||||
Three rule classes in `inet fips_gateway` together implement the
|
||||
LAN→mesh path:
|
||||
|
||||
**Prerouting DNAT (per mapping)** rewrites the destination from the
|
||||
virtual IP to the corresponding mesh address:
|
||||
|
||||
```text
|
||||
match: nfproto ipv6 && ip6 daddr == <virtual_ip>
|
||||
action: dnat to <mesh_addr>
|
||||
```
|
||||
|
||||
After DNAT, the kernel routes the packet through `fips0` via the
|
||||
standard routing table.
|
||||
|
||||
**Postrouting masquerade (`oifname fips0`)** rewrites the source of
|
||||
all traffic exiting via `fips0` to the gateway's own `fips0` address:
|
||||
|
||||
```text
|
||||
match: oifname == "fips0"
|
||||
action: masquerade
|
||||
```
|
||||
|
||||
This rule is critical. Without it, LAN client source addresses (for
|
||||
example `fd02::20` from the LAN's RA-advertised prefix, or virtual
|
||||
addresses from another forwarding domain) would appear as the source
|
||||
on the mesh. Those addresses are meaningless to mesh nodes, so
|
||||
return traffic would be black-holed. Masquerade ensures all mesh
|
||||
traffic appears to originate from the gateway's own FIPS identity.
|
||||
|
||||
**Postrouting SNAT (per mapping)** rewrites the source of return
|
||||
traffic from the mesh address back to the virtual IP:
|
||||
|
||||
```text
|
||||
match: nfproto ipv6 && ip6 saddr == <mesh_addr>
|
||||
action: snat to <virtual_ip>
|
||||
```
|
||||
|
||||
Without it, the LAN client would see replies from the raw
|
||||
`fd00::/8` mesh address rather than from the virtual IP it had
|
||||
originally connected to, breaking application-layer assumptions about
|
||||
the destination address.
|
||||
|
||||
### Network Requirements (Outbound)
|
||||
|
||||
The gateway host needs IPv6 forwarding enabled
|
||||
(`net.ipv6.conf.all.forwarding=1`), proxy NDP enabled on the LAN
|
||||
interface, `CAP_NET_ADMIN` for `fips-gateway`, and a `local
|
||||
<pool-cidr> dev lo` route so the kernel accepts packets to the pool
|
||||
as locally owned and runs them through the NAT chains. LAN clients
|
||||
need a route to the pool via the gateway and DNS resolution that
|
||||
forwards `.fips` queries there. On OpenWrt the init script handles
|
||||
all of this; on other Linux hosts the operator handles it manually.
|
||||
Full setup is documented in
|
||||
[../how-to/deploy-gateway.md](../how-to/deploy-gateway.md).
|
||||
|
||||
## The Inbound Half (Mesh → LAN)
|
||||
|
||||
### Configuration Shape
|
||||
|
||||
Inbound port-forwards live in `gateway.port_forwards[]`. Each entry
|
||||
is a triple:
|
||||
|
||||
| Field | Type | Notes |
|
||||
| ----- | ---- | ----- |
|
||||
| `listen_port` | `u16` | Port on the gateway's `fips0` address. Must be non-zero. |
|
||||
| `proto` | `tcp` \| `udp` | Match protocol. |
|
||||
| `target` | `[ipv6]:port` | LAN destination. IPv4 targets are rejected at parse time by `SocketAddrV6`. |
|
||||
|
||||
Validation runs at startup and on every config reload:
|
||||
`(listen_port, proto)` must be unique across the list, and zero
|
||||
listen ports are rejected. Forwards are independent of outbound
|
||||
configuration: a gateway with no `pool` consumers can still expose
|
||||
inbound services (the pool route and DNS proxy still run, since they
|
||||
are part of the same binary, but they sit idle).
|
||||
|
||||
### NAT Pipeline (Inbound)
|
||||
|
||||
For each port-forward, a single prerouting DNAT rule matches
|
||||
mesh-originated traffic landing on the gateway's `fips0` address
|
||||
and rewrites it to the LAN target:
|
||||
|
||||
```text
|
||||
match: iifname == "fips0" && nfproto ipv6
|
||||
&& l4proto == <tcp|udp> && th dport == <listen_port>
|
||||
action: dnat to <target_ip>:<target_port>
|
||||
```
|
||||
|
||||
The match clause is deliberately narrow:
|
||||
|
||||
- **`iifname == "fips0"`** restricts the rule to traffic that
|
||||
arrived from the mesh. LAN-side ingress is never subject to
|
||||
inbound forwarding.
|
||||
- **`nfproto ipv6`** is enforced both here and at config-load time
|
||||
(`SocketAddrV6` rejects IPv4 targets); FIPS is IPv6-only end to
|
||||
end.
|
||||
- **`l4proto + dport`** narrows the match to one
|
||||
`(listen_port, proto)` pair per rule. Unique-tuple validation
|
||||
ensures no two rules contend for the same packet.
|
||||
|
||||
When *any* port-forward is configured, a single LAN-side masquerade
|
||||
is added to postrouting:
|
||||
|
||||
```text
|
||||
match: iifname == "fips0" && oifname == <lan_interface>
|
||||
&& nfproto ipv6
|
||||
action: masquerade
|
||||
```
|
||||
|
||||
Without this rule, the LAN target would attempt to reply directly to
|
||||
the mesh peer's `fd00::/8` source address, which is not reachable on
|
||||
the LAN. Masquerade rewrites the source to the gateway's LAN-side
|
||||
address so the target sees a reachable peer and conntrack routes
|
||||
the reply back through the gateway.
|
||||
|
||||
This LAN-side masquerade is independent of the `oifname fips0`
|
||||
masquerade in the outbound pipeline; the two have disjoint match
|
||||
clauses (different `iifname`/`oifname` combinations) and coexist
|
||||
without interaction when both directions are active.
|
||||
|
||||
### Independence From Outbound
|
||||
|
||||
The inbound half does not require:
|
||||
|
||||
- A virtual-IP pool. Mesh peers connect directly to the gateway's
|
||||
own `fips0` address, which the FIPS daemon already owns.
|
||||
- DNS resolution. Mesh peers reach the gateway as
|
||||
`<gateway-npub>.fips:<port>` using their own resolver (or a
|
||||
numeric mesh address); the gateway's DNS proxy is not in the path.
|
||||
- A daemon-side identity cache for the LAN target. The target is a
|
||||
LAN-side IPv6 address, not a mesh address; no `fd00::/8` lookup
|
||||
happens for it.
|
||||
|
||||
A gateway configured with port-forwards but with no LAN clients ever
|
||||
issuing `.fips` DNS queries will have an empty pool and zero
|
||||
outbound mappings, but its inbound forwards work normally. The
|
||||
inverse is also true: a gateway that serves only outbound LAN→mesh
|
||||
traffic has zero entries in the port-forwards list and no LAN-side
|
||||
masquerade.
|
||||
|
||||
## Atomic Table Rebuild (Common)
|
||||
|
||||
Both halves contribute rules to the same `inet fips_gateway` table,
|
||||
and that table is rebuilt as one unit on every state change —
|
||||
mapping added, mapping removed, port-forwards updated. The rebuild
|
||||
sequence is:
|
||||
|
||||
1. Delete the existing table in its own batch (ignore `ENOENT`).
|
||||
2. In a fresh batch: add the table; add the `prerouting` and
|
||||
`postrouting` chains; add the always-on `oifname fips0`
|
||||
masquerade; add per-mapping DNAT/SNAT rules for every active
|
||||
pool entry; add per-port-forward DNAT rules; add the LAN-side
|
||||
masquerade if any port-forwards exist.
|
||||
3. Send the batch as a single netlink transaction.
|
||||
|
||||
The rustables crate does not expose rule-handle tracking, so
|
||||
incremental update of individual rules is not available. Atomic
|
||||
rebuild was chosen for simplicity and correctness: it eliminates an
|
||||
entire class of partial-update inconsistency bugs at the cost of
|
||||
repeating the (cheap) rule construction on every change. The total
|
||||
rule count is bounded by the pool capacity (2 per mapping, capped
|
||||
at 2^16) and the port-forward count, both of which are small in
|
||||
practice.
|
||||
|
||||
## Configuration Reference
|
||||
|
||||
The full `gateway.*` block — pool CIDR, LAN interface, DNS
|
||||
listen/upstream/TTL, pool grace period, conntrack timeouts, and
|
||||
inbound port-forwards — is documented in the
|
||||
[Gateway section](../reference/configuration.md#gateway-gateway)
|
||||
of the configuration reference. The same block governs both halves;
|
||||
fields specific to one half (`pool`, `dns.*` for outbound;
|
||||
`port_forwards[]` for inbound) are simply unused when the other
|
||||
half is not in play.
|
||||
|
||||
## Operations and Troubleshooting
|
||||
|
||||
- [../tutorials/deploy-fips-gateway.md](../tutorials/deploy-fips-gateway.md)
|
||||
— end-to-end walkthrough on OpenWrt.
|
||||
- [../how-to/deploy-gateway.md](../how-to/deploy-gateway.md) —
|
||||
recipe for non-OpenWrt Linux hosts and inbound-port-forwarding
|
||||
configuration.
|
||||
- [../how-to/troubleshoot-gateway.md](../how-to/troubleshoot-gateway.md)
|
||||
— diagnostic recipes (DNS failures, ping working but TCP not,
|
||||
conntrack inspection, pool exhaustion, port-53 conflicts,
|
||||
port-forward verification).
|
||||
- [../reference/cli-fips-gateway.md](../reference/cli-fips-gateway.md)
|
||||
— command-line interface.
|
||||
- [../reference/control-socket.md](../reference/control-socket.md#gateway-command-catalog)
|
||||
— `show_gateway` and `show_mappings` commands.
|
||||
|
||||
## Security Considerations
|
||||
|
||||
### Outbound
|
||||
|
||||
- **LAN trust boundary.** The DNS listener and the virtual-IP pool
|
||||
are reachable by every host on the LAN. Any LAN host that can
|
||||
resolve `.fips` and route to the pool CIDR can reach mesh
|
||||
destinations. There is no per-client authentication; access
|
||||
restriction is a network-level concern, enforced with firewall
|
||||
rules on the LAN interface or on the gateway host itself.
|
||||
- **Identity masking.** All outbound LAN traffic appears on the
|
||||
mesh under the gateway's own FIPS identity. Mesh nodes cannot
|
||||
determine which LAN host originated a connection. This provides
|
||||
privacy for LAN hosts but means the gateway's reputation covers
|
||||
all of its clients — and that abusive behavior from one LAN host
|
||||
is attributed to the gateway, not to the host.
|
||||
- **Plaintext between client and gateway.** Traffic between the LAN
|
||||
client and the gateway is unencrypted at the IP layer. FIPS
|
||||
encryption (FSP) protects the segment between the gateway and the
|
||||
destination mesh node; application-layer encryption (TLS, SSH,
|
||||
Noise) is the only thing that provides true end-to-end protection
|
||||
through the gateway.
|
||||
- **Pool addresses are ephemeral.** Virtual IPs are allocated
|
||||
dynamically and recycled. They are not authenticated and not
|
||||
bound to client identity — a LAN host connecting to a virtual IP
|
||||
is trusting the gateway's recent DNS response.
|
||||
- **DNS upstream trust.** The outbound half's correctness depends
|
||||
on the FIPS daemon's resolver returning honest `fd00::/8`
|
||||
answers; a compromised daemon could redirect LAN clients to
|
||||
arbitrary mesh nodes.
|
||||
|
||||
### Inbound
|
||||
|
||||
- **Port exposure.** Each entry in `port_forwards[]` exposes the
|
||||
matched `(listen_port, proto)` on the gateway's mesh-side
|
||||
address to every reachable mesh peer. Inbound port-forwards are
|
||||
not gated by any peer ACL beyond what FMP normally enforces;
|
||||
treat them with the same care as a public-internet port forward.
|
||||
- **Mesh peer trust.** The LAN target sees connections that have
|
||||
been masqueraded to the gateway's LAN address. The target cannot
|
||||
distinguish one mesh peer from another, and there is no
|
||||
authenticated peer identity available to the LAN target — any
|
||||
application-layer authentication or rate-limiting must run on
|
||||
the target itself.
|
||||
- **Return-path masquerade exposes the gateway's LAN address.**
|
||||
The LAN-side masquerade rewrites the mesh peer's source to the
|
||||
gateway's LAN address. A malicious or buggy LAN target can use
|
||||
this to send unsolicited traffic back at the gateway, or to
|
||||
probe other LAN hosts via the gateway's network position; LAN
|
||||
segmentation (VLANs, host firewalls) is the right control.
|
||||
|
||||
### Common
|
||||
|
||||
- **No client identity verification.** The gateway authenticates
|
||||
neither LAN clients nor mesh peers beyond what the underlying
|
||||
layers already do — `fips0` ingress carries an FSP-authenticated
|
||||
payload, the LAN side is whoever the LAN admits.
|
||||
|
||||
## References
|
||||
|
||||
- [fips-ipv6-adapter.md](fips-ipv6-adapter.md) — IPv6 adapter and
|
||||
TUN interface design.
|
||||
- [fips-architecture.md](fips-architecture.md) — protocol layer
|
||||
architecture.
|
||||
- [fips-concepts.md](fips-concepts.md) — protocol overview.
|
||||
- [../reference/configuration.md](../reference/configuration.md) —
|
||||
configuration reference.
|
||||
- [../reference/cli-fips-gateway.md](../reference/cli-fips-gateway.md)
|
||||
— `fips-gateway` CLI.
|
||||
- [../reference/control-socket.md](../reference/control-socket.md) —
|
||||
control-socket protocol and command catalog.
|
||||
- [../how-to/deploy-gateway.md](../how-to/deploy-gateway.md) —
|
||||
gateway host and LAN client setup.
|
||||
- [../how-to/troubleshoot-gateway.md](../how-to/troubleshoot-gateway.md)
|
||||
— diagnostic recipes.
|
||||
- [../tutorials/deploy-fips-gateway.md](../tutorials/deploy-fips-gateway.md)
|
||||
— OpenWrt walkthrough.
|
||||
@@ -1,817 +0,0 @@
|
||||
# FIPS: Free Internetworking Peering System
|
||||
|
||||
## What is FIPS?
|
||||
|
||||
FIPS is a self-organizing mesh network that can operate natively over a
|
||||
variety of physical and logical media, such as local area networks,
|
||||
Bluetooth, serial links, or the existing internet as an overlay. The
|
||||
long-term goal is infrastructure that can function alongside or ultimately
|
||||
replace dependence on the Internet itself. Systems running FIPS establish
|
||||
peer connections, authenticate each other, and route traffic for each other
|
||||
without any central authority or global topology knowledge, and allow
|
||||
end-to-end encrypted sessions between any two nodes regardless of how many
|
||||
hops separate them.
|
||||
|
||||
Nodes in the mesh route traffic for each other using Nostr identities
|
||||
(npubs) as network addresses. Applications can access the mesh through a
|
||||
native FIPS datagram service, or through an IPv6 adaptation layer that
|
||||
presents each node as an IPv6 endpoint for compatibility with existing
|
||||
IP-based applications.
|
||||
|
||||
## Why FIPS?
|
||||
|
||||
**Self-sovereign identity**: FIPS nodes generate their own addresses, node
|
||||
IDs, and security credentials without coordination with any central
|
||||
authority. These identities can be long-term fixed or may be ephemeral,
|
||||
changed at any time. These identities are not visible to the FIPS network
|
||||
itself — they are used only at the application layer and for end-to-end
|
||||
session encryption.
|
||||
|
||||
**Infrastructure independence**: The internet depends on centralized
|
||||
infrastructure — ISPs, backbone providers, DNS, certificate authorities.
|
||||
FIPS works over any transport that can carry packets: a serial connection,
|
||||
onion-routed connections through Tor, local area networking, radio links
|
||||
between remote sites, or the existing internet as an overlay. When the
|
||||
internet is unavailable, unreliable, or untrusted, the mesh still works.
|
||||
|
||||
**Privacy by design**: FIPS provides secure, authenticated, and encrypted
|
||||
communication between any two nodes in the mesh, independent of the mix of
|
||||
transports used along the routed path between them. Furthermore, the mesh
|
||||
itself is designed to minimize metadata exposure — intermediate nodes route
|
||||
packets without learning the identities of the endpoints.
|
||||
|
||||
**Zero configuration**: Nodes discover each other and build routing
|
||||
automatically. Connect to one peer and you can reach the entire mesh. The
|
||||
network self-heals around failures and adapts to changing topology.
|
||||
|
||||
## A Self-Organizing Mesh
|
||||
|
||||
Traditional networks are built top-down. A central authority assigns
|
||||
addresses, configures routing tables, provisions hardware, and manages the
|
||||
topology. If the authority disappears or the infrastructure fails, the
|
||||
network fails with it. Nodes cannot reach each other without infrastructure
|
||||
mediating the connection.
|
||||
|
||||
FIPS inverts this model. There is no central authority, no address
|
||||
assignment service, no routing table pushed from above. Each node generates
|
||||
its own identity from a cryptographic keypair. Each node independently
|
||||
decides which peers to connect to and which transports to use. From these
|
||||
local decisions alone, the network self-organizes:
|
||||
|
||||
- A **spanning tree** forms through distributed parent selection, giving
|
||||
every node a coordinate in the network without any node knowing the full
|
||||
topology
|
||||
- **Bloom filters** propagate through gossip, so each node learns which
|
||||
peers can reach which destinations — again without global knowledge
|
||||
- **Routing decisions** are made locally at each hop, using only the node's
|
||||
immediate peers and cached coordinate information
|
||||
|
||||
Each peer link and end-to-end session actively measures RTT, loss, jitter,
|
||||
and goodput through a lightweight in-band Metrics Measurement Protocol
|
||||
(MMP), providing operator visibility and a foundation for quality-aware
|
||||
routing.
|
||||
|
||||
The result is a network that builds itself from the bottom up, heals around
|
||||
failures automatically, and scales without central coordination. Adding a
|
||||
node is as simple as connecting to one existing peer — the network
|
||||
integrates the new node through its normal mesh protocols.
|
||||
|
||||
## Specific Design Goals
|
||||
|
||||
- **Nostr-native identity and cryptography** — Use Nostr keypairs as node
|
||||
identities and leverage secp256k1, Schnorr signatures, and SHA-256
|
||||
- **Transport agnostic** — Support overlay, shared medium, and
|
||||
point-to-point transports transparently
|
||||
- **Self-organizing** — Automatic topology discovery and route optimization
|
||||
- **Privacy preserving** — Minimize metadata leakage across untrusted links
|
||||
- **Resilient** — Self-healing with graceful degradation
|
||||
|
||||
Non-goals include:
|
||||
|
||||
- **Reliable delivery** — FIPS provides a best-effort datagram service;
|
||||
retransmission and ordering are left to applications or higher-layer
|
||||
protocols
|
||||
- **Anonymity** — Direct peers learn each other's identity; FIPS minimizes
|
||||
metadata exposure but is not an anonymity network like Tor
|
||||
- **Congestion control** — FIPS measures link quality but does not implement
|
||||
flow control or congestion avoidance at the mesh layer
|
||||
|
||||
---
|
||||
|
||||
## Protocol Architecture
|
||||
|
||||
FIPS is organized in three protocol layers, each with distinct
|
||||
responsibilities and clean service boundaries. No layer depends on the
|
||||
specifics of the layers above or below it — transport plugins know nothing
|
||||
about sessions, the routing layer knows nothing about application addressing,
|
||||
and applications know nothing about which physical media carry their traffic.
|
||||
This separation means new transports, protocol features, and application
|
||||
interfaces can be added independently.
|
||||
|
||||

|
||||
|
||||
### Mapping to Traditional Networking
|
||||
|
||||
Readers familiar with the OSI model or TCP/IP networking may find it helpful
|
||||
to see how FIPS concepts relate to traditional layers:
|
||||
|
||||

|
||||
|
||||
Note that FMP spans what would traditionally be separate link and network
|
||||
layers. This is intentional — in a self-organizing mesh, the same layer that
|
||||
authenticates peers also makes routing decisions, because routing depends on
|
||||
authenticated peer state (spanning tree positions, bloom filters).
|
||||
|
||||
### Layer Responsibilities
|
||||
|
||||
**Transport layer**: Delivers datagrams between endpoints over a specific
|
||||
medium. Each transport type (UDP socket, Ethernet interface, radio modem)
|
||||
implements the same abstract interface: send and receive datagrams, report
|
||||
MTU. The transport layer knows nothing about FIPS identities, routing, or
|
||||
encryption. It provides raw datagram delivery to FMP above.
|
||||
|
||||
See [fips-transport-layer.md](fips-transport-layer.md) for the transport layer
|
||||
specification.
|
||||
|
||||
**FIPS Mesh Protocol (FMP)**: Manages peer connections, authenticates peers
|
||||
via Noise IK handshakes, and encrypts all traffic on each link. FMP is where
|
||||
the mesh organizes itself — nodes exchange spanning tree announcements and
|
||||
bloom filters with their direct peers, and FMP makes forwarding decisions
|
||||
for transit traffic. FMP provides authenticated, encrypted forwarding to FSP
|
||||
above.
|
||||
|
||||
See [fips-mesh-layer.md](fips-mesh-layer.md) for the FMP specification and
|
||||
[fips-mesh-operation.md](fips-mesh-operation.md) for how FMP's routing and
|
||||
self-organization work in practice.
|
||||
|
||||
**FIPS Session Protocol (FSP)**: Provides end-to-end authenticated
|
||||
encryption between any two nodes, regardless of how many intermediate hops
|
||||
separate them. FSP manages session lifecycle (setup, data transfer,
|
||||
teardown), caches destination coordinates for efficient routing, and handles
|
||||
the warmup strategy that keeps transit node caches populated. Session
|
||||
dispatch uses index-based routing inspired by
|
||||
[WireGuard](https://www.wireguard.com/), enabling O(1) packet
|
||||
demultiplexing. FSP provides a datagram service to applications above.
|
||||
|
||||
See [fips-session-layer.md](fips-session-layer.md) for the FSP specification.
|
||||
|
||||
**IPv6 adaptation layer**: Sits above FSP as a service on port 256, adapting
|
||||
the FIPS datagram service for unmodified IPv6 applications. Provides DNS
|
||||
resolution (npub → fd00::/8 address), identity cache management, IPv6 header
|
||||
compression, MTU enforcement, and a TUN interface. This is the primary way
|
||||
existing applications use the FIPS mesh.
|
||||
|
||||
See [fips-ipv6-adapter.md](fips-ipv6-adapter.md) for the IPv6 adapter.
|
||||
|
||||
### Node Architecture
|
||||
|
||||
Application services sit at the top of the stack, dispatched by FSP port
|
||||
number: the IPv6 TUN adapter (port 256) maps npubs to `fd00::/8` addresses
|
||||
with header compression so unmodified IP applications can use the network
|
||||
transparently, while the native datagram API addresses destinations directly
|
||||
by npub.
|
||||
|
||||

|
||||
|
||||
The mesh routes application traffic across heterogeneous transports
|
||||
transparently. A packet may traverse WiFi, Ethernet, UDP/IP, and Tor links
|
||||
on its way from source to destination — the application never needs to know
|
||||
which transports are involved. Each hop is independently encrypted at the
|
||||
link layer, while a single end-to-end session protects the payload across
|
||||
the entire path.
|
||||
|
||||

|
||||
|
||||
---
|
||||
|
||||
## Identity System
|
||||
|
||||
FIPS uses [Nostr](https://github.com/nostr-protocol/nips) keypairs
|
||||
(secp256k1) as node identities. The public key identifies the node; the
|
||||
private key signs protocol messages and establishes encrypted sessions.
|
||||
|
||||
The public key (or its bech32-encoded npub form) is the primary means for
|
||||
application-layer software to identify communication endpoints. Internally,
|
||||
the protocol derives a `node_addr` (a 16-byte SHA-256 hash of the pubkey)
|
||||
used as the routing identifier in packet headers, and an IPv6 address derived
|
||||
from the node_addr for the TUN adapter. Applications use the pubkey or npub;
|
||||
the routing layer uses node_addr; unmodified IPv6 applications use the
|
||||
derived `fd00::/8` address. All three are deterministically derived from the
|
||||
same keypair.
|
||||
|
||||
### FIPS Identity Handling
|
||||
|
||||

|
||||
|
||||
The pubkey is the node's cryptographic identity, used in Noise IK handshakes
|
||||
for both link and session encryption. It is never exposed beyond the
|
||||
endpoints of an encrypted channel. The node_addr, a one-way SHA-256 hash
|
||||
truncated to 16 bytes, serves as the routing identifier in packet headers
|
||||
and bloom filters. Intermediate routers see only node_addrs — they can
|
||||
forward traffic without learning the Nostr identities of the endpoints. An
|
||||
observer can verify "does this node_addr belong to pubkey X?" if they already
|
||||
know the pubkey, but cannot enumerate communicating identities by inspecting
|
||||
traffic. The IPv6
|
||||
address prepends `fd` to the first 15 bytes of the node_addr, providing a
|
||||
ULA overlay address for unmodified IP applications via the TUN interface.
|
||||
|
||||
Below the FIPS identity layer, each transport uses its own native addressing
|
||||
— IP:port tuples, MAC addresses, .onion identifiers. These **link
|
||||
addresses** are opaque to everything above FMP and discarded once link
|
||||
authentication completes.
|
||||
|
||||
### Identity Verification
|
||||
|
||||
The Noise Protocol Framework mutually authenticates both peer-to-peer link
|
||||
connections (at FMP) and end-to-end session traffic (at FSP), proving each
|
||||
party controls the private key for their claimed identity.
|
||||
|
||||
See [fips-mesh-layer.md](fips-mesh-layer.md) for peer authentication and
|
||||
[fips-session-layer.md](fips-session-layer.md) for end-to-end session
|
||||
establishment.
|
||||
|
||||
Key rotation changes the node's identity — a new keypair produces a new
|
||||
node_addr and IPv6 address, requiring all sessions to be re-established.
|
||||
Migration mechanisms that allow a node to announce a successor key are a
|
||||
future consideration.
|
||||
|
||||
---
|
||||
|
||||
## Two-Layer Encryption
|
||||
|
||||
FIPS uses independent encryption at two protocol layers:
|
||||
|
||||
| Layer | Scope | Pattern | Purpose |
|
||||
| ----- | ----- | ------- | ------- |
|
||||
| **FMP (Mesh)** | Hop-by-hop | Noise IK | Encrypt all traffic on each peer link |
|
||||
| **FSP (Session)** | End-to-end | Noise XK | Encrypt application payload between endpoints |
|
||||
|
||||
### Link Layer (Hop-by-Hop)
|
||||
|
||||
When two nodes establish a direct connection, they perform a [Noise
|
||||
IK](https://noiseprotocol.org/) handshake. This authenticates both parties
|
||||
and establishes symmetric keys for encrypting all traffic on that link.
|
||||
Every packet between direct peers is encrypted — gossip messages, routing
|
||||
queries, and forwarded session datagrams alike.
|
||||
|
||||
The IK pattern is used because outbound connections know the peer's npub
|
||||
from configuration, while inbound connections learn the initiator's identity
|
||||
from the first handshake message.
|
||||
|
||||
### Session Layer (End-to-End)
|
||||
|
||||
FIPS establishes end-to-end encrypted sessions between any two communicating
|
||||
nodes using Noise XK, regardless of how many hops separate them. The
|
||||
initiator knows the destination's npub (required for XK's pre-message);
|
||||
the responder learns the initiator's identity from the third handshake
|
||||
message. Unlike the link-layer IK pattern where the initiator's identity
|
||||
is revealed in msg1, XK delays identity disclosure until msg3, providing
|
||||
stronger initiator identity protection for traffic traversing untrusted
|
||||
intermediate nodes.
|
||||
|
||||
A packet from A to D through intermediate nodes B and C:
|
||||
|
||||
1. A encrypts payload with A↔D session key (FSP)
|
||||
2. A wraps in SessionDatagram, encrypts with A↔B link key (FMP), sends to B
|
||||
3. B decrypts link layer, reads destination node_addr, re-encrypts with B↔C
|
||||
link key, forwards to C
|
||||
4. C decrypts link layer, re-encrypts with C↔D link key, forwards to D
|
||||
5. D decrypts link layer, then decrypts session layer to get payload
|
||||
|
||||
Intermediate nodes route based on destination node_addr but cannot read
|
||||
session-layer payloads. Each hop strips one link encryption and applies the
|
||||
next — the session-layer ciphertext passes through untouched.
|
||||
|
||||
Both layers always apply, even between adjacent peers — a packet to a direct
|
||||
neighbor is still encrypted twice. This uniform model means no special cases
|
||||
for local vs remote destinations, and topology changes (a direct peer
|
||||
becomes reachable only through intermediaries) don't affect existing
|
||||
sessions.
|
||||
|
||||
See [fips-mesh-layer.md](fips-mesh-layer.md) for link encryption and
|
||||
[fips-session-layer.md](fips-session-layer.md) for session encryption.
|
||||
|
||||
---
|
||||
|
||||
## Routing and Mesh Operation
|
||||
|
||||
Each node makes forwarding decisions using only local information — its
|
||||
immediate peers, their bloom filters, and cached coordinates — rather than
|
||||
centrally distributed routing tables or global topology knowledge. Two
|
||||
complementary mechanisms provide the information each node needs.
|
||||
|
||||
### Spanning Tree: The Coordinate System
|
||||
|
||||

|
||||
|
||||
Nodes self-organize into a spanning tree through gossip — each node
|
||||
exchanges announcements with its direct peers and independently selects a
|
||||
parent. Because every node applies the same rule (prefer the root with the
|
||||
smallest node_addr), the network converges on a single agreed-upon root
|
||||
without any voting or coordination. This is the same principle behind the
|
||||
[Spanning Tree
|
||||
Protocol](https://en.wikipedia.org/wiki/Spanning_Tree_Protocol) used in
|
||||
Ethernet bridging since the 1980s: purely local decisions that converge to
|
||||
consistent global state. The resulting tree gives every node a
|
||||
**coordinate** — its path from itself to the root. Using tree coordinates
|
||||
for routing is adapted from
|
||||
[Yggdrasil](https://yggdrasil-network.github.io/)'s
|
||||
[Ironwood](https://github.com/Arceliar/ironwood) routing library.
|
||||
|
||||
These coordinates enable distance calculations between any two nodes: the
|
||||
distance is the number of hops from each node to their lowest common
|
||||
ancestor in the tree. This provides a metric for routing decisions without
|
||||
any node needing to know the full network topology.
|
||||
|
||||
The tree maintains itself through gossip — nodes exchange TreeAnnounce
|
||||
messages with their peers, propagating parent selections and ancestry
|
||||
chains. Changes cascade through the tree proportional to depth, not network
|
||||
size. If the network partitions, each segment converges to its own new root
|
||||
through the same process and reconverges automatically when segments rejoin.
|
||||
|
||||
See [fips-spanning-tree.md](fips-spanning-tree.md) for the tree algorithms
|
||||
and [spanning-tree-dynamics.md](spanning-tree-dynamics.md) for detailed
|
||||
convergence walkthroughs.
|
||||
|
||||
### Bloom Filters: Candidate Selection
|
||||
|
||||
The spanning tree provides a coordinate system for distance-based routing,
|
||||
but on its own each node would only know about its immediate neighbors.
|
||||
Bloom filters complement the tree by distributing reachability knowledge
|
||||
across the entire mesh — each node learns which destinations are reachable
|
||||
through which peers, without any node needing a complete view of the
|
||||
network.
|
||||
|
||||
Each node's peer-advertised [bloom
|
||||
filter](https://en.wikipedia.org/wiki/Bloom_filter) is a compact, fixed-size
|
||||
data structure that answers one question: "can this peer possibly reach
|
||||
destination D?" The answer is either "no" (definitive) or "maybe"
|
||||
(probabilistic — false positives are possible). Because the filter size is
|
||||
constant regardless of how many destinations it represents, bloom filters
|
||||
scale efficiently as the network grows. This is candidate selection for
|
||||
routing — bloom filters narrow the set of peers worth considering, and the
|
||||
actual forwarding decision ranks those candidates by tree distance and link
|
||||
quality.
|
||||
|
||||
Filters propagate transitively through tree edges, with each node computing
|
||||
outbound filters by merging the filters received from its tree peers (parent
|
||||
and children) using a
|
||||
[split-horizon](https://en.wikipedia.org/wiki/Split_horizon_route_advertisement)
|
||||
technique borrowed from distance-vector routing. All peers — including
|
||||
non-tree mesh shortcuts — receive FilterAnnounce messages, but only tree
|
||||
peers' filters are merged into outgoing computation. This prevents filter
|
||||
saturation where mesh shortcuts would cause every filter to converge toward
|
||||
the full network.
|
||||
|
||||
See [fips-bloom-filters.md](fips-bloom-filters.md) for filter parameters and
|
||||
mathematical properties.
|
||||
|
||||

|
||||
|
||||
The outbound filter for peer Q merges this node's identity with tree peer
|
||||
inbound filters except Q's (split-horizon exclusion). This creates
|
||||
directional asymmetry: upward filters (child → parent) contain the child's
|
||||
subtree, while downward filters (parent → child) contain the complement.
|
||||
Mesh peers receive filters but their inbound filters are not merged
|
||||
transitively — they provide single-hop shortcut visibility only.
|
||||
|
||||
A node with multiple peers receives genuinely different filters from each.
|
||||
In the diagram, R receives {B, D, E} from B and {C, F} from C — two disjoint
|
||||
subtrees. When R needs to reach F, only C's filter matches. This is where
|
||||
bloom filters provide real candidate selection: a node with several peers
|
||||
can narrow the forwarding choice before consulting tree coordinates. Leaf
|
||||
nodes like D have only one peer, so their single inbound filter is
|
||||
necessarily near-complete (everything except themselves) and offers no
|
||||
selection — but leaf nodes have no choice to make anyway.
|
||||
|
||||
Bloom filter sizing (bit count and hash functions) requires further analysis
|
||||
based on actual deployment scenarios. The FMP wire format is versioned to
|
||||
accommodate future parameter changes as operational experience accumulates.
|
||||
|
||||
### Routing Decisions
|
||||
|
||||
At each hop, FMP makes a local forwarding decision using the following
|
||||
priority chain:
|
||||
|
||||
1. **Local delivery** — the destination is this node
|
||||
2. **Direct peer** — the destination is an authenticated neighbor
|
||||
3. **Bloom-guided candidate selection** — bloom filters identify peers that
|
||||
can reach the destination; tree coordinates rank them by distance and
|
||||
link quality
|
||||
4. **[Greedy routing](https://en.wikipedia.org/wiki/Greedy_embedding)** —
|
||||
fallback when bloom filters haven't converged; forward to the peer that
|
||||
minimizes tree distance to the destination
|
||||
5. **No route** — destination unreachable; send error signal to source
|
||||
|
||||
All multi-hop routing depends on knowing the destination's tree coordinates.
|
||||
These are cached at each node after being learned through discovery
|
||||
(LookupRequest/LookupResponse) or session establishment (SessionSetup). The
|
||||
coordinate cache is the critical piece that enables efficient forwarding.
|
||||
|
||||

|
||||
|
||||
### Coordinate Caching and Discovery
|
||||
|
||||
When a node first needs to reach an unknown destination, it sends a
|
||||
LookupRequest that propagates through the network guided by bloom filters
|
||||
and loop prevention. The destination responds with its coordinates, which
|
||||
the source and intermediate nodes along the return path cache. Subsequent
|
||||
traffic routes efficiently using the cached coordinates.
|
||||
|
||||
Session establishment (SessionSetup) also carries coordinates, warming
|
||||
transit node caches along the path so that data packets can be forwarded
|
||||
without individual discovery at each hop.
|
||||
|
||||

|
||||
|
||||
### Error Recovery
|
||||
|
||||
When routing fails — because cached coordinates are stale, a path has
|
||||
broken, or a packet exceeds a link's MTU — transit nodes signal the source:
|
||||
|
||||
- **CoordsRequired**: A transit node lacks the destination's coordinates.
|
||||
The source re-initiates discovery and resets its coordinate warmup
|
||||
strategy.
|
||||
- **PathBroken**: Greedy routing reached a dead end. The source re-discovers
|
||||
the destination's current coordinates.
|
||||
- **MtuExceeded**: A transit node cannot forward a packet because it exceeds
|
||||
the next-hop link MTU. The source adjusts its path MTU estimate.
|
||||
|
||||
All three signals trigger active recovery, and are rate-limited to prevent
|
||||
storms during topology changes.
|
||||
|
||||
See [fips-mesh-operation.md](fips-mesh-operation.md) for the complete
|
||||
routing and mesh behavior description.
|
||||
|
||||
### Metrics Measurement Protocol (MMP)
|
||||
|
||||
Each peer link runs an instance of the Metrics Measurement Protocol, which
|
||||
measures link quality through in-band report exchange. MMP computes smoothed
|
||||
round-trip time (SRTT), packet loss rate, interarrival jitter, goodput, and
|
||||
one-way delay trend — all derived from counter and timestamp fields already
|
||||
present in the FMP wire format, with no additional probing traffic required.
|
||||
|
||||
MMP operates in three modes. **Full** mode exchanges both SenderReports and
|
||||
ReceiverReports to compute all metrics including RTT. **Lightweight** mode
|
||||
exchanges only ReceiverReports, providing loss and jitter but not RTT — useful
|
||||
for constrained links. **Minimal** mode disables reports entirely, relying
|
||||
only on spin bit and congestion echo flags in the frame header.
|
||||
|
||||
Reports are sent at RTT-adaptive intervals (clamped to 100 ms–2 s), so
|
||||
high-latency links don't generate excessive measurement traffic while
|
||||
low-latency links converge quickly. Each metric carries both short-term and
|
||||
long-term exponentially weighted moving averages, enabling detection of
|
||||
quality changes against a stable baseline.
|
||||
|
||||
MMP serves dual roles: operator visibility and cost-based parent selection.
|
||||
Periodic log lines report per-link RTT, loss, jitter, and goodput. MMP
|
||||
computes an Expected Transmission Count (ETX) from bidirectional delivery
|
||||
ratios, which feeds into cost-based parent selection where each node
|
||||
evaluates `effective_depth = depth + link_cost` using
|
||||
`link_cost = etx * (1.0 + srtt_ms / 100.0)`. ETX is not yet used in
|
||||
`find_next_hop()` candidate ranking for data forwarding.
|
||||
|
||||
See [fips-mesh-layer.md](fips-mesh-layer.md) for MMP operating modes, report
|
||||
scheduling, and the spin bit design.
|
||||
|
||||
---
|
||||
|
||||
## Transport Abstraction
|
||||
|
||||
FIPS treats the communication medium as a pluggable component. Every transport
|
||||
— whether a UDP socket, an Ethernet interface, a Tor circuit, or a radio modem
|
||||
— implements the same simple interface: send a datagram to an address, receive
|
||||
datagrams, and report the link MTU. The rest of the protocol stack sees no
|
||||
difference between them.
|
||||
|
||||
A **transport** is a driver for a particular medium. A **link** is a peer
|
||||
connection established over a transport. Transport addresses (IP:port, MAC
|
||||
address, .onion) are opaque to all layers above FMP — they exist only to
|
||||
deliver datagrams and are discarded once FMP has authenticated the peer via
|
||||
the Noise IK handshake. From that point on, the peer is identified solely by
|
||||
its cryptographic identity.
|
||||
|
||||
Transports fall into three categories based on their connectivity model:
|
||||
|
||||
| Category | Examples | Characteristics |
|
||||
| -------- | -------- | --------------- |
|
||||
| Overlay | UDP/IP, Tor | Tunnels FIPS over existing networks |
|
||||
| Shared medium | Ethernet, WiFi, Bluetooth, Radio | Local broadcast, peer discovery |
|
||||
| Point-to-point | Serial, dialup | Fixed connections, no discovery |
|
||||
|
||||
These categories differ in addressing, MTU, reliability, and whether they
|
||||
support local discovery, but FMP handles all of them uniformly. A node
|
||||
running multiple transports simultaneously bridges between those networks
|
||||
automatically — peers from all transports feed into a single spanning tree,
|
||||
and the router selects the best path regardless of which medium carries it.
|
||||
If one transport fails, traffic reroutes through alternatives without
|
||||
application involvement.
|
||||
|
||||
Some transports support an optional discovery capability — the ability to
|
||||
broadcast and listen for announcements indicating the availability of FIPS
|
||||
endpoints on the local medium. Shared media like Ethernet, WiFi, Bluetooth,
|
||||
and radio are natural fits for this, as they can reach nearby devices without
|
||||
prior configuration. When discovery is available, nodes can automatically
|
||||
find and peer with other FIPS nodes on the same medium. Transports that
|
||||
lack discovery (such as configured UDP endpoints) simply skip this step and
|
||||
connect directly to configured addresses. Additionally, endpoint discovery
|
||||
using Nostr relays and signed events is planned, allowing internet-reachable
|
||||
nodes to publish their transport addresses for other FIPS nodes to find.
|
||||
|
||||
NAT traversal is not currently addressed by the protocol.
|
||||
Internet-connected nodes behind NAT must be reachable through port
|
||||
forwarding, a publicly addressed peer, or relay through other mesh nodes.
|
||||
UDP hole punching and relay-assisted NAT traversal are potential future
|
||||
mechanisms but are not part of the current design.
|
||||
|
||||
> **Implementation status**: UDP/IP is implemented. Ethernet and Bluetooth
|
||||
> transports are under active design and development. All others are future
|
||||
> directions.
|
||||
|
||||
See [fips-transport-layer.md](fips-transport-layer.md) for the full transport
|
||||
layer specification.
|
||||
|
||||
---
|
||||
|
||||
## Security
|
||||
|
||||
FIPS is designed around four classes of adversary, each addressed by a
|
||||
different layer of the protocol.
|
||||
|
||||
### Transport Observers
|
||||
|
||||
A passive observer on the underlying transport — someone monitoring a WiFi
|
||||
network, tapping an Ethernet segment, or inspecting UDP traffic — sees only
|
||||
encrypted packets. The FMP link-layer Noise IK session encrypts all traffic
|
||||
between direct peers, including routing gossip and forwarded session
|
||||
datagrams. The observer can infer timing, packet sizes, and which transport
|
||||
endpoints are exchanging traffic, but cannot read content or determine
|
||||
FIPS-level node identities from the encrypted packets. Traffic analysis —
|
||||
correlating timing and volume patterns across multiple vantage points to
|
||||
infer communication relationships — is not defended against (see
|
||||
[Specific Design Goals](#specific-design-goals)).
|
||||
|
||||
### Active Attackers on the Transport
|
||||
|
||||
An adversary who can inject, modify, drop, or replay packets on the
|
||||
transport is also defeated by the FMP link-layer Noise IK session. Mutual
|
||||
authentication prevents impersonation, AEAD encryption detects tampering,
|
||||
and counter-based nonces with a sliding replay window reject replayed
|
||||
packets.
|
||||
|
||||
### Other FIPS Nodes (Intermediate Routers)
|
||||
|
||||
The most important adversary class is the operators of other nodes in the
|
||||
mesh — the peers that forward your traffic. FIPS treats every intermediate
|
||||
router as potentially adversarial. The FSP session layer establishes a
|
||||
completely independent Noise XK session between the communicating endpoints,
|
||||
so intermediate nodes cannot read application payloads even though they
|
||||
decrypt and re-encrypt the link-layer envelope at each hop.
|
||||
|
||||
Routing headers expose only the destination's node_addr — an opaque
|
||||
SHA-256 hash of the actual public key. Intermediate routers can forward
|
||||
traffic without learning which Nostr identities are communicating. An
|
||||
observer can verify "does this node_addr belong to pubkey X?" if they
|
||||
already know the pubkey, but cannot enumerate communicating identities by
|
||||
inspecting routed traffic.
|
||||
|
||||
| Entity | Can See |
|
||||
| ------ | ------- |
|
||||
| Transport observer | Encrypted packets, timing, packet sizes |
|
||||
| Direct peer | Your npub, traffic volume, timing |
|
||||
| Intermediate router | Source and destination node_addrs, packet size |
|
||||
| Destination | Your npub, payload content |
|
||||
|
||||
### Adversarial Nodes Disrupting the Mesh
|
||||
|
||||
Beyond passive observation, a malicious node could attempt to disrupt
|
||||
routing by injecting false spanning tree announcements, advertising bogus
|
||||
bloom filters, or claiming invalid tree positions. FMP mitigates these
|
||||
through signed TreeAnnounce messages verified by direct peers, transitive
|
||||
ancestry chain validation, replay protection via sequence numbers, and
|
||||
discretionary peering — node operators choose who to peer with, so an
|
||||
attacker with many identities still needs real nodes to accept their
|
||||
connections. Handshake rate limiting further constrains how fast an attacker
|
||||
can establish new links. In fully open networks with automatic peer
|
||||
discovery, Sybil resistance relies primarily on rate limiting; discretionary
|
||||
peering provides stronger resistance in curated deployments where operators
|
||||
vet their peers. An attacker who controls all of a target node's direct
|
||||
peers can completely control its view of the network (an eclipse attack);
|
||||
diverse peering across independent operators and transports is the primary
|
||||
mitigation.
|
||||
|
||||
---
|
||||
|
||||
## Prior Work
|
||||
|
||||
FIPS builds on proven designs rather than inventing new cryptography or routing
|
||||
algorithms. Nearly every major design decision has deployed precedent.
|
||||
|
||||
### Spanning Tree Self-Organization
|
||||
|
||||
The idea that distributed nodes can build a spanning tree through purely local
|
||||
decisions — each node selecting a parent based on announcements from its
|
||||
neighbors — dates to the
|
||||
[IEEE 802.1D Spanning Tree Protocol](https://en.wikipedia.org/wiki/Spanning_Tree_Protocol)
|
||||
(STP, 1985). STP demonstrated that a network-wide tree emerges from a simple
|
||||
deterministic rule (lowest bridge ID wins root election) applied independently
|
||||
at each node. FIPS uses the same principle — lowest node address determines the
|
||||
root — adapted from an Ethernet bridging context to a general-purpose overlay
|
||||
mesh.
|
||||
|
||||
### Tree Coordinate Routing
|
||||
|
||||
The spanning tree coordinates, bloom filter candidate selection, and greedy
|
||||
routing algorithms are adapted from
|
||||
[Yggdrasil v0.5](https://yggdrasil-network.github.io/2023/10/22/upcoming-v05-release.html)
|
||||
and its [Ironwood](https://github.com/Arceliar/ironwood) routing library.
|
||||
Yggdrasil's key insight was using the tree path from root to node as a
|
||||
routable coordinate, enabling greedy forwarding without global routing tables.
|
||||
FIPS adapts these algorithms for multi-transport operation, Nostr identity
|
||||
integration, and constrained MTU environments.
|
||||
|
||||
The theoretical foundation for greedy routing on tree embeddings draws on
|
||||
[Kleinberg's work](https://www.cs.cornell.edu/home/kleinber/swn.pdf) on
|
||||
navigable small-world networks, which showed that greedy forwarding succeeds
|
||||
in O(log² n) steps when the network has hierarchical structure. Thorup-Zwick
|
||||
compact routing schemes separately demonstrated that sublinear routing state
|
||||
is achievable with bounded stretch, motivating the use of tree coordinates
|
||||
rather than full routing tables.
|
||||
|
||||
### Split-Horizon Bloom Filter Propagation
|
||||
|
||||
FIPS distributes reachability information using bloom filters computed with a
|
||||
split-horizon rule: when advertising to a peer, exclude that peer's own
|
||||
contributions. This technique is borrowed from distance-vector routing
|
||||
protocols — [RIP](https://en.wikipedia.org/wiki/Routing_Information_Protocol)
|
||||
(1988) and [Babel](https://www.irif.fr/~jch/software/babel/) use split-horizon
|
||||
to prevent routing loops by not advertising a route back to the neighbor it was
|
||||
learned from. FIPS applies the same principle to probabilistic set
|
||||
advertisements rather than distance-vector tables.
|
||||
|
||||
### Cryptographic Identity as Network Address
|
||||
|
||||
FIPS nodes are identified by their Nostr public keys (secp256k1). The network
|
||||
address *is* the cryptographic identity — there is no separate address
|
||||
assignment or registration step.
|
||||
[CJDNS](https://github.com/cjdelisle/cjdns) pioneered this approach in
|
||||
overlay meshes, deriving IPv6 addresses from the double-SHA-512 of each node's
|
||||
public key. Tor [.onion addresses](https://spec.torproject.org/rend-spec-v3)
|
||||
and the IETF
|
||||
[Host Identity Protocol](https://en.wikipedia.org/wiki/Host_Identity_Protocol)
|
||||
(HIP) follow the same principle. FIPS uses Nostr's existing key infrastructure
|
||||
rather than introducing a new identity scheme.
|
||||
|
||||
### Dual-Layer Encryption
|
||||
|
||||
FIPS encrypts traffic twice: FMP provides hop-by-hop link encryption
|
||||
(protecting against transport-layer observers), while FSP provides independent
|
||||
end-to-end session encryption (protecting against intermediate FIPS nodes).
|
||||
This layered approach mirrors [Tor](https://www.torproject.org/), where each
|
||||
relay peels one layer of encryption (hop-by-hop) while the innermost layer
|
||||
protects end-to-end payload. [I2P](https://geti2p.net/) uses a similar
|
||||
garlic routing scheme with tunnel-layer and end-to-end encryption. Unlike Tor
|
||||
and I2P, FIPS does not provide anonymity — its dual encryption protects
|
||||
confidentiality and integrity rather than hiding traffic patterns.
|
||||
|
||||
### Noise Protocol Framework
|
||||
|
||||
FIPS uses the [Noise Protocol Framework](https://noiseprotocol.org/) at both
|
||||
protocol layers, with different handshake patterns chosen for each layer's
|
||||
threat model. FMP link encryption uses **Noise IK**, providing mutual
|
||||
authentication with a single round trip where the initiator knows the
|
||||
responder's static key in advance.
|
||||
[WireGuard](https://www.wireguard.com/) uses the same IK base pattern
|
||||
(extended with a pre-shared key as IKpsk2) for VPN tunnels. FSP session
|
||||
encryption uses **Noise XK**, the same pattern used by the
|
||||
[Lightning Network](https://github.com/lightning/bolts/blob/master/08-transport.md),
|
||||
where the initiator's static key is transmitted in a third message rather
|
||||
than the first. XK provides stronger initiator identity hiding at the cost
|
||||
of an additional round trip — a worthwhile tradeoff for session-layer traffic
|
||||
that traverses untrusted intermediate nodes. At the link layer, where both
|
||||
peers are configured and directly connected, IK's single round trip is
|
||||
preferred.
|
||||
|
||||
### Index-Based Session Dispatch
|
||||
|
||||
FIPS uses locally-assigned 32-bit session indices to demultiplex incoming
|
||||
packets to the correct cryptographic session in O(1) time, without parsing
|
||||
source addresses or performing expensive lookups. This directly follows
|
||||
[WireGuard's](https://www.wireguard.com/papers/wireguard.pdf) receiver index
|
||||
approach, where each peer assigns a random index during handshake and the
|
||||
remote side includes it in every packet header.
|
||||
|
||||
### Transport-Agnostic Overlay Mesh
|
||||
|
||||
FIPS is designed to operate over any datagram-capable transport — UDP, raw
|
||||
Ethernet, Bluetooth, radio, serial — through a uniform transport abstraction.
|
||||
Several mesh overlays have demonstrated transport-agnostic design:
|
||||
[CJDNS](https://github.com/cjdelisle/cjdns) runs over UDP and Ethernet,
|
||||
[Yggdrasil](https://yggdrasil-network.github.io/) supports TCP and TLS
|
||||
transports, and [Tor](https://www.torproject.org/) can use pluggable
|
||||
transports to tunnel through various media. FIPS extends this pattern to
|
||||
shared-medium transports (radio, BLE) with per-transport MTU and discovery
|
||||
capabilities.
|
||||
|
||||
### Metrics Measurement Protocol
|
||||
|
||||
MMP's design assembles well-established measurement techniques into a unified
|
||||
per-link protocol. The SenderReport/ReceiverReport exchange structure follows
|
||||
[RTCP](https://www.rfc-editor.org/rfc/rfc3550) (RFC 3550), which uses the
|
||||
same report pairing for media stream quality monitoring in RTP sessions. MMP's
|
||||
jitter computation uses the RTCP interarrival jitter algorithm directly.
|
||||
|
||||
The smoothed RTT estimator uses the Jacobson/Karels algorithm
|
||||
([RFC 6298](https://www.rfc-editor.org/rfc/rfc6298)), the same SRTT
|
||||
computation used in TCP for retransmission timeout calculation since 1988.
|
||||
MMP derives RTT from timestamp-echo in ReceiverReports with dwell-time
|
||||
compensation, rather than from packet round-trips.
|
||||
|
||||
The spin bit in the FMP frame header follows the
|
||||
[QUIC](https://www.rfc-editor.org/rfc/rfc9000) spin bit
|
||||
([RFC 9312](https://www.rfc-editor.org/rfc/rfc9312)) — a single bit that
|
||||
alternates each round trip, enabling passive latency measurement. FIPS
|
||||
implements the spin bit state machine but relies on timestamp-echo for SRTT,
|
||||
as irregular mesh traffic makes spin bit RTT unreliable.
|
||||
|
||||
The Expected Transmission Count (ETX) metric, computed from bidirectional
|
||||
delivery ratios, was introduced by
|
||||
[De Couto et al. (2003)](https://pdos.csail.mit.edu/papers/grid:mobicom03/paper.pdf)
|
||||
for wireless mesh routing and is used in protocols including
|
||||
[OLSR](https://en.wikipedia.org/wiki/Optimized_Link_State_Routing_Protocol)
|
||||
and [Babel](https://www.irif.fr/~jch/software/babel/). FIPS computes ETX
|
||||
per-link from MMP loss measurements for future use in candidate ranking.
|
||||
|
||||
The CE (Congestion Experienced) echo flag provides hop-by-hop
|
||||
[ECN](https://en.wikipedia.org/wiki/Explicit_Congestion_Notification)
|
||||
signaling, following the TCP/IP ECN echo pattern (RFC 3168). Transit nodes
|
||||
detect congestion via MMP loss/ETX metrics or kernel buffer drops and set
|
||||
the CE flag on forwarded frames; destination nodes mark ECN-capable IPv6
|
||||
packets accordingly.
|
||||
|
||||
### Cryptographic Primitives
|
||||
|
||||
FIPS reuses [Nostr's](https://github.com/nostr-protocol/nips) cryptographic
|
||||
stack — secp256k1 for identity keys, Schnorr signatures for authentication,
|
||||
SHA-256 for hashing, and ChaCha20-Poly1305 for authenticated encryption. This
|
||||
is the same primitive set used across Bitcoin, Nostr, and a growing ecosystem
|
||||
of self-sovereign identity systems. No novel cryptography is introduced.
|
||||
|
||||
---
|
||||
|
||||
## Further Reading
|
||||
|
||||
### Protocol Layers
|
||||
|
||||
| Document | Description |
|
||||
| -------- | ----------- |
|
||||
| [fips-transport-layer.md](fips-transport-layer.md) | Transport layer: abstraction, types, services provided to FMP |
|
||||
| [fips-mesh-layer.md](fips-mesh-layer.md) | FMP: peer authentication, link encryption, forwarding |
|
||||
| [fips-session-layer.md](fips-session-layer.md) | FSP: end-to-end encryption, session lifecycle |
|
||||
| [fips-ipv6-adapter.md](fips-ipv6-adapter.md) | IPv6 adaptation: DNS, TUN interface, MTU enforcement |
|
||||
|
||||
### Mesh Behavior and Wire Formats
|
||||
|
||||
| Document | Description |
|
||||
| -------- | ----------- |
|
||||
| [fips-mesh-operation.md](fips-mesh-operation.md) | How the mesh operates: routing, discovery, error recovery |
|
||||
| [fips-wire-formats.md](fips-wire-formats.md) | Complete wire format reference for all protocol layers |
|
||||
|
||||
### Supporting References
|
||||
|
||||
| Document | Description |
|
||||
| -------- | ----------- |
|
||||
| [fips-spanning-tree.md](fips-spanning-tree.md) | Spanning tree algorithms and data structures |
|
||||
| [fips-bloom-filters.md](fips-bloom-filters.md) | Bloom filter parameters, math, and computation |
|
||||
| [spanning-tree-dynamics.md](spanning-tree-dynamics.md) | Scenario walkthroughs: convergence, partitions, recovery |
|
||||
|
||||
### Implementation
|
||||
|
||||
| Document | Description |
|
||||
| -------- | ----------- |
|
||||
| [fips-configuration.md](fips-configuration.md) | YAML configuration reference |
|
||||
|
||||
### External References
|
||||
|
||||
- [IEEE 802.1D Spanning Tree Protocol](https://en.wikipedia.org/wiki/Spanning_Tree_Protocol)
|
||||
- [Yggdrasil Network](https://yggdrasil-network.github.io/)
|
||||
- [Yggdrasil v0.5 Release Notes](https://yggdrasil-network.github.io/2023/10/22/upcoming-v05-release.html)
|
||||
- [Ironwood Routing Library](https://github.com/Arceliar/ironwood)
|
||||
- [Kleinberg — The Small-World Phenomenon](https://www.cs.cornell.edu/home/kleinber/swn.pdf)
|
||||
- [CJDNS](https://github.com/cjdelisle/cjdns)
|
||||
- [Tor Project](https://www.torproject.org/)
|
||||
- [I2P](https://geti2p.net/)
|
||||
- [Host Identity Protocol (HIP)](https://en.wikipedia.org/wiki/Host_Identity_Protocol)
|
||||
- [Babel Routing Protocol](https://www.irif.fr/~jch/software/babel/)
|
||||
- [Noise Protocol Framework](https://noiseprotocol.org/)
|
||||
- [WireGuard](https://www.wireguard.com/)
|
||||
- [WireGuard Whitepaper](https://www.wireguard.com/papers/wireguard.pdf)
|
||||
- [Lightning Network BOLT #8 — Transport](https://github.com/lightning/bolts/blob/master/08-transport.md)
|
||||
- [QUIC (RFC 9000)](https://www.rfc-editor.org/rfc/rfc9000)
|
||||
- [QUIC Spin Bit (RFC 9312)](https://www.rfc-editor.org/rfc/rfc9312)
|
||||
- [RTCP (RFC 3550)](https://www.rfc-editor.org/rfc/rfc3550)
|
||||
- [TCP SRTT / RTO (RFC 6298)](https://www.rfc-editor.org/rfc/rfc6298)
|
||||
- [ECN (RFC 3168)](https://www.rfc-editor.org/rfc/rfc3168)
|
||||
- [ETX — De Couto et al. 2003](https://pdos.csail.mit.edu/papers/grid:mobicom03/paper.pdf)
|
||||
- [OLSR](https://en.wikipedia.org/wiki/Optimized_Link_State_Routing_Protocol)
|
||||
- [Nostr Protocol](https://github.com/nostr-protocol/nips)
|
||||
@@ -69,6 +69,24 @@ Known cache population mechanisms:
|
||||
- **Inbound traffic**: Authenticated sessions from other nodes populate the
|
||||
cache with their identity information
|
||||
|
||||
### Mesh-Interface Query Filter
|
||||
|
||||
The DNS responder is intended for local applications resolving `.fips`
|
||||
names; queries arriving over the mesh interface itself are dropped. The
|
||||
daemon records the index of the TUN interface at startup and compares
|
||||
it against the arrival interface of each incoming UDP DNS query. When
|
||||
they match — meaning the query came from another mesh node, not from a
|
||||
local socket — the responder discards the query without replying.
|
||||
|
||||
The check is implemented in
|
||||
[`is_mesh_interface_query`](../../src/upper/dns.rs) and prevents two
|
||||
classes of misbehaviour: a peer asking the daemon to resolve `.fips`
|
||||
names on its behalf (which would let one node use another as an
|
||||
identity-cache priming proxy), and accidental query loops where a
|
||||
misconfigured resolver forwards `.fips` queries back into the mesh.
|
||||
Local applications binding to the host's loopback or non-mesh
|
||||
interfaces are unaffected.
|
||||
|
||||
## IPv6 Address Derivation
|
||||
|
||||
FIPS addresses use the IPv6 Unique Local Address (ULA) prefix `fd00::/8`:
|
||||
@@ -125,41 +143,24 @@ entry hasn't been evicted by memory pressure.
|
||||
|
||||
## MTU Enforcement
|
||||
|
||||
FIPS does not provide fragmentation or reassembly at the session or mesh
|
||||
protocol layers — every datagram must fit in a single transport-layer packet.
|
||||
Some transports may perform fragmentation and reassembly internally (e.g., BLE
|
||||
L2CAP) and can advertise a larger virtual MTU than the physical medium
|
||||
supports, but this is transparent to FIPS. The mesh layer provides two
|
||||
facilities to manage MTU across heterogeneous paths: route discovery can
|
||||
constrain results to paths that support a required minimum MTU, and transit
|
||||
nodes that cannot forward an oversized datagram send an MtuExceeded error
|
||||
signal back to the source. The adapter must ensure that IPv6 packets from
|
||||
applications fit within the FIPS encapsulation budget after all layers of
|
||||
wrapping.
|
||||
The adapter sits at the boundary between the host's IPv6 stack and the
|
||||
FIPS encapsulation budget. Its job is to keep IPv6 packets small
|
||||
enough that they fit through the FIPS protocol envelope on every link
|
||||
along the path. The cross-cutting MTU model — proactive
|
||||
SessionDatagram `path_mtu` annotation, reactive MtuExceeded signals,
|
||||
end-to-end PathMtuNotification echo, and per-destination MTU storage
|
||||
— is documented in [fips-mtu.md](fips-mtu.md). What the adapter
|
||||
contributes is the IPv6-specific overhead accounting and the TUN-side
|
||||
enforcement integration.
|
||||
|
||||
### Encapsulation Overhead
|
||||
### IPv6-Specific Overhead
|
||||
|
||||
| Layer | Overhead | Purpose |
|
||||
| ----- | -------- | ------- |
|
||||
| Link encryption | 37 bytes | 16-byte outer header + 5-byte inner header (timestamp + msg_type) + 16-byte AEAD tag |
|
||||
| SessionDatagram body | 35 bytes | ttl + path_mtu + src_addr + dest_addr (msg_type counted in inner header) |
|
||||
| FSP header | 12 bytes | 4-byte prefix + 8-byte counter (used as AEAD AAD) |
|
||||
| FSP inner header | 6 bytes | 4-byte timestamp + 1-byte msg_type + 1-byte inner_flags (inside AEAD) |
|
||||
| Session AEAD tag | 16 bytes | ChaCha20-Poly1305 tag on session-encrypted payload |
|
||||
| **Protocol envelope** | **106 bytes** | `FIPS_OVERHEAD` constant |
|
||||
| Port header | 4 bytes | src_port + dst_port (DataPacket service dispatch) |
|
||||
| IPv6 compression | −33 bytes | 40-byte IPv6 header → 7-byte format + residual |
|
||||
| **IPv6 data path total** | **77 bytes** | `FIPS_IPV6_OVERHEAD` constant |
|
||||
|
||||
Coordinate piggybacking (CP flag) adds variable overhead: `2 + entries × 16`
|
||||
per coordinate, with both src and dst coords sent. The send path skips the
|
||||
CP flag if adding coords would exceed the transport MTU.
|
||||
|
||||
The `FIPS_OVERHEAD` constant (106 bytes) represents the base protocol
|
||||
envelope overhead (link encryption + routing + session encryption). For IPv6
|
||||
traffic, FSP port multiplexing adds 4 bytes (port header) while IPv6 header
|
||||
compression saves 33 bytes (40-byte header → 7-byte format + residual),
|
||||
yielding a net `FIPS_IPV6_OVERHEAD` of 77 bytes.
|
||||
For IPv6 traffic, FSP port multiplexing adds 4 bytes (port header)
|
||||
while IPv6 header compression saves 33 bytes (40-byte header →
|
||||
7-byte format + residual), yielding a net `FIPS_IPV6_OVERHEAD` of
|
||||
77 bytes on top of the base `FIPS_OVERHEAD` (106 bytes) protocol
|
||||
envelope. The full encapsulation breakdown lives in
|
||||
[fips-mtu.md](fips-mtu.md#encapsulation-overhead).
|
||||
|
||||
### Effective IPv6 MTU
|
||||
|
||||
@@ -183,48 +184,51 @@ transport path MTU for the IPv6 adapter is therefore:
|
||||
1280 + 77 = 1357 bytes
|
||||
```
|
||||
|
||||
Transports with smaller MTUs (radio at ~250 bytes, serial at 256 bytes) cannot
|
||||
support the IPv6 adapter without some form of internal fragmentation and
|
||||
reassembly. Otherwise, applications on those transports must use the native
|
||||
FIPS datagram API.
|
||||
Transports with smaller MTUs (radio at ~250 bytes, serial at 256
|
||||
bytes) cannot support the IPv6 adapter without some form of internal
|
||||
fragmentation and reassembly. Otherwise, applications on those
|
||||
transports must use the native FIPS datagram API.
|
||||
|
||||
### ICMP Packet Too Big
|
||||
### TUN-Side ICMP Packet Too Big
|
||||
|
||||
When an outbound packet at the TUN exceeds the effective IPv6 MTU, the adapter
|
||||
generates an ICMPv6 Packet Too Big message and delivers it back to the
|
||||
application via the TUN. This triggers the kernel's Path MTU Discovery (PMTUD)
|
||||
mechanism, which adjusts TCP segment sizes for subsequent transmissions.
|
||||
When an outbound packet at the TUN exceeds the effective IPv6 MTU,
|
||||
the adapter generates an ICMPv6 Packet Too Big message and delivers
|
||||
it back to the application via the TUN. This triggers the kernel's
|
||||
Path MTU Discovery mechanism, which adjusts TCP segment sizes for
|
||||
subsequent transmissions.
|
||||
|
||||
ICMP Packet Too Big generation is rate-limited per source address (100ms
|
||||
interval) to prevent storms from applications sending many oversized packets.
|
||||
ICMP Packet Too Big generation is rate-limited per source address
|
||||
(100ms interval) to prevent storms from applications sending many
|
||||
oversized packets. The ICMP response is delivered locally back through
|
||||
the TUN; no network traversal is needed, so delivery is reliable.
|
||||
|
||||
The ICMP response is delivered locally (back through the TUN to the kernel) —
|
||||
no network traversal is needed, so delivery is reliable.
|
||||
### TUN-Side TCP MSS Clamping
|
||||
|
||||
### TCP MSS Clamping
|
||||
|
||||
The adapter intercepts TCP SYN and SYN-ACK packets at the TUN interface and
|
||||
clamps the Maximum Segment Size (MSS) option:
|
||||
The adapter intercepts TCP SYN and SYN-ACK packets at the TUN
|
||||
interface and clamps the Maximum Segment Size (MSS) option:
|
||||
|
||||
```text
|
||||
clamped_mss = effective_ipv6_mtu - 40 (IPv6 header) - 20 (TCP header)
|
||||
```
|
||||
|
||||
This prevents TCP connections from negotiating segment sizes that would exceed
|
||||
the FIPS path MTU. Clamping is applied in two places:
|
||||
Clamping is applied in two places:
|
||||
|
||||
- **TUN reader** (outbound): Clamps MSS on outbound SYN packets
|
||||
- **TUN writer** (inbound): Clamps MSS on inbound SYN-ACK packets
|
||||
|
||||
Together, these ensure both directions of a TCP connection use appropriately
|
||||
sized segments from the start, avoiding the initial oversized packet loss
|
||||
that would occur with ICMP Packet Too Big alone.
|
||||
Together, these ensure both directions of a TCP connection use
|
||||
appropriately sized segments from the start, avoiding the initial
|
||||
oversized packet loss that would occur with ICMP Packet Too Big
|
||||
alone. The conditional clamp (per-flow lookup with cold-flow
|
||||
fallback) and the rationale for `max_mss` semantics are in
|
||||
[fips-mtu.md](fips-mtu.md#tcp-mss-clamping).
|
||||
|
||||
### ICMP Rate Limiting
|
||||
|
||||
ICMPv6 error generation is rate-limited per source address using a token bucket
|
||||
(100ms interval). This matches the standard ICMP rate limiting approach and
|
||||
prevents amplification when an application sends a burst of oversized packets.
|
||||
ICMPv6 error generation is rate-limited per source address using a
|
||||
token bucket (100ms interval). This matches the standard ICMP rate
|
||||
limiting approach and prevents amplification when an application sends
|
||||
a burst of oversized packets.
|
||||
|
||||
## TUN Interface
|
||||
|
||||
@@ -287,20 +291,16 @@ path.
|
||||
|
||||
### Configuration
|
||||
|
||||
```yaml
|
||||
tun:
|
||||
enabled: true
|
||||
name: fips0
|
||||
mtu: 1280
|
||||
```
|
||||
The TUN block (`tun.*`) is documented in
|
||||
[../reference/configuration.md](../reference/configuration.md).
|
||||
|
||||
### Privileges
|
||||
|
||||
TUN device creation requires `CAP_NET_ADMIN`. Options:
|
||||
|
||||
- Run as root
|
||||
- Set capability: `sudo setcap cap_net_admin+ep ./target/debug/fips`
|
||||
- Pre-created persistent TUN device
|
||||
TUN device creation requires `CAP_NET_ADMIN`. The shipped Debian
|
||||
systemd unit runs the daemon as `root` by default; for the
|
||||
alternative — running under a dedicated unprivileged service
|
||||
account with the capability granted on the binary — see
|
||||
[../how-to/run-as-unprivileged-user.md](../how-to/run-as-unprivileged-user.md).
|
||||
|
||||
## Implementation Status
|
||||
|
||||
@@ -314,6 +314,7 @@ TUN device creation requires `CAP_NET_ADMIN`. Options:
|
||||
| ICMP rate limiting (per-source) | **Implemented** |
|
||||
| TCP MSS clamping (SYN + SYN-ACK) | **Implemented** |
|
||||
| DNS service (.fips domain) | **Implemented** |
|
||||
| DNS responder mesh-interface filter | **Implemented** |
|
||||
| Port-based service multiplexing (port 256) | **Implemented** |
|
||||
| IPv6 header compression (format 0x00) | **Implemented** |
|
||||
| Per-destination route MTU (netlink) | Planned |
|
||||
@@ -324,38 +325,27 @@ TUN device creation requires `CAP_NET_ADMIN`. Options:
|
||||
|
||||
## Design Considerations
|
||||
|
||||
### Path MTU Discovery
|
||||
### Path MTU Discovery and No-Fragmentation Policy
|
||||
|
||||
Two complementary mechanisms support full PMTUD:
|
||||
|
||||
1. **Proactive**: The `path_mtu` field (2 bytes) in the SessionDatagram envelope
|
||||
is implemented at the FMP level. The source sets it to its outbound link MTU
|
||||
minus overhead; each transit node applies
|
||||
`min(current, own_outbound_mtu - overhead)`. The destination receives the
|
||||
forward-path minimum. PathMtuNotification is handled at the session layer;
|
||||
the destination sends the observed forward-path MTU back to the source,
|
||||
which applies it with decrease-immediate / increase-requires-3-consecutive
|
||||
hysteresis.
|
||||
|
||||
2. **Reactive**: When a transit node cannot forward a packet (MTU exceeded), it
|
||||
sends an error signal back to the source. This handles the in-flight gap
|
||||
between a path MTU decrease and the source learning via the echo.
|
||||
|
||||
Both are needed: proactive handles steady state; reactive handles the transient
|
||||
window when oversized packets hit a new bottleneck before the source adapts.
|
||||
|
||||
### No Fragmentation
|
||||
|
||||
FIPS remains a pure datagram service with no fragmentation at transit nodes.
|
||||
Session-layer encryption is end-to-end — the AEAD tag authenticates the entire
|
||||
plaintext. Fragmenting encrypted datagrams would require either exposing
|
||||
plaintext structure to transit nodes (unacceptable) or reassembly before
|
||||
decryption (opens attack surface).
|
||||
Path MTU Discovery (proactive `path_mtu` annotation, reactive
|
||||
MtuExceeded, end-to-end PathMtuNotification) and the no-fragmentation
|
||||
policy that drives the design both live in the unified MTU treatment
|
||||
at [fips-mtu.md](fips-mtu.md). The adapter is a consumer of that
|
||||
model — its job is to enforce the resulting effective IPv6 MTU at the
|
||||
TUN with ICMP Packet Too Big and TCP MSS clamping.
|
||||
|
||||
## References
|
||||
|
||||
- [fips-intro.md](fips-intro.md) — Protocol overview and architecture
|
||||
- [fips-concepts.md](fips-concepts.md) — Protocol overview
|
||||
- [fips-architecture.md](fips-architecture.md) — Layer architecture and
|
||||
identity model
|
||||
- [fips-session-layer.md](fips-session-layer.md) — FSP (below the adapter)
|
||||
- [fips-wire-formats.md](fips-wire-formats.md) — FSP and SessionDatagram wire
|
||||
formats
|
||||
- [fips-configuration.md](fips-configuration.md) — TUN configuration parameters
|
||||
- [fips-mtu.md](fips-mtu.md) — Unified path MTU model (proactive,
|
||||
reactive, hysteresis, no-fragmentation)
|
||||
- [../reference/wire-formats.md](../reference/wire-formats.md) — FSP and
|
||||
SessionDatagram wire formats
|
||||
- [../reference/configuration.md](../reference/configuration.md) — TUN
|
||||
configuration parameters
|
||||
- [../how-to/run-as-unprivileged-user.md](../how-to/run-as-unprivileged-user.md)
|
||||
— privilege options for the daemon, including the unprivileged
|
||||
service-account path
|
||||
|
||||
@@ -15,7 +15,7 @@ updates, coordinate discovery, and forwarded session datagrams — all encrypted
|
||||
per-hop.
|
||||
|
||||
FMP is the boundary between opaque transport addresses and identified peers.
|
||||
Below FMP, everything is transport-specific addresses (IP:port, MAC, .onion).
|
||||
Below FMP, everything is transport-specific addresses (host:port, MAC, .onion).
|
||||
Above FMP, everything is peers identified by public keys and routable by
|
||||
node_addr. The transport layer never sees FIPS-level structure; FSP never sees
|
||||
transport addresses or routing details.
|
||||
@@ -215,8 +215,8 @@ The plaintext inside the encrypted frame begins with a 5-byte inner header
|
||||
(4-byte session-relative timestamp followed by a message type byte), then the
|
||||
message-specific payload.
|
||||
|
||||
See [fips-wire-formats.md](fips-wire-formats.md) for the complete wire format
|
||||
specification.
|
||||
See [../reference/wire-formats.md](../reference/wire-formats.md) for
|
||||
the complete wire format specification.
|
||||
|
||||
### What Encryption Provides
|
||||
|
||||
@@ -280,7 +280,7 @@ authenticated peer regardless of what transport address the packet arrived
|
||||
from. FMP updates the peer's current address to the packet's source address,
|
||||
and subsequent outbound packets use the updated address.
|
||||
|
||||
This allows peers to change transport addresses (e.g., IP:port for UDP)
|
||||
This allows peers to change transport addresses (e.g., host:port for UDP)
|
||||
without session interruption. The mechanism is:
|
||||
|
||||
1. Packet arrives from a different address than expected
|
||||
@@ -292,6 +292,12 @@ Roaming is most useful for UDP, where source addresses can change due to NAT
|
||||
rebinding or network changes. For connection-oriented transports, "roaming"
|
||||
manifests as reconnection rather than mid-session address change.
|
||||
|
||||
Roaming addresses *mid-session* NAT rebinding. Establishing the initial UDP
|
||||
path through NAT is a separate concern, addressed by the optional
|
||||
Nostr-mediated overlay discovery and STUN-assisted hole punching feature
|
||||
(see [fips-transport-layer.md](fips-transport-layer.md) and
|
||||
[../reference/configuration.md](../reference/configuration.md)).
|
||||
|
||||
## Replay Protection
|
||||
|
||||
Each link session maintains per-direction counters:
|
||||
@@ -335,7 +341,9 @@ Additional protections:
|
||||
memory usage
|
||||
- **Handshake timeout**: Stale pending handshakes are cleaned up after a
|
||||
configurable timeout
|
||||
- **Allowlist/blocklist**: Optional peer filtering before handshake processing
|
||||
- **Peer ACL**: Optional allowlist / denylist filtering of peer npubs
|
||||
before handshake processing (loaded from `/etc/fips/peers.allow` and
|
||||
`/etc/fips/peers.deny`, mtime-watched and reloaded automatically)
|
||||
|
||||
## Disconnect
|
||||
|
||||
@@ -355,6 +363,72 @@ links.
|
||||
On node shutdown, Disconnect is sent to all active peers before transports are
|
||||
stopped.
|
||||
|
||||
## Rekey
|
||||
|
||||
FMP periodically negotiates a fresh Noise session over each established link
|
||||
to bound forward-secrecy exposure: limiting AEAD nonce reuse risk, bounding
|
||||
the volume of ciphertext recoverable from a stolen long-term static key, and
|
||||
rotating the session indices that addressed packets carry on the wire.
|
||||
|
||||
A rekey is initiated when either threshold is reached on the link's current
|
||||
session: `node.rekey.after_secs` (default 120) elapsed since the link came
|
||||
up or last rekeyed, or `node.rekey.after_messages` (default 65536) frames
|
||||
sent. Either side can be the initiator independently. Rekey is on by
|
||||
default and can be disabled via `node.rekey.enabled: false` (the
|
||||
configuration tree is documented in
|
||||
[../reference/configuration.md](../reference/configuration.md)).
|
||||
|
||||
### Mechanism
|
||||
|
||||
A rekey reuses the Noise IK pattern of the initial handshake, but the two
|
||||
messages travel over the existing link as ordinary encrypted FMP frames
|
||||
rather than as plaintext bootstrap packets. The initiator builds a fresh
|
||||
`HandshakeState`, generates msg1, and sends it through the current session;
|
||||
the responder consumes msg1, builds msg2, and replies. After both sides
|
||||
have exchanged messages and finalised the new keys, traffic transitions
|
||||
from the old session to the new one.
|
||||
|
||||
Cutover is signalled in-band by the **K-bit** in the FMP flags byte. Each
|
||||
side starts emitting frames under the new session with K set; on receipt
|
||||
of the first K-marked frame the peer accepts the cutover and follows
|
||||
suit. A new pair of session indices is allocated as part of the new
|
||||
session, replacing the old indices on subsequent frames (see
|
||||
[Index Properties](#index-properties)).
|
||||
|
||||
### Drain Window
|
||||
|
||||
To absorb in-flight reordering across the cutover, the old session is not
|
||||
discarded immediately. Each peer retains it in a `previous_session` slot
|
||||
on the active-peer state for `DRAIN_WINDOW_SECS = 10` seconds (a
|
||||
compile-time constant in `src/node/handlers/rekey.rs`). During the
|
||||
window, decrypt attempts fall back to `previous_session` when the new
|
||||
session rejects a frame, so a packet sent under the old keys that
|
||||
arrives a few hundred milliseconds late still decrypts. After the
|
||||
window expires, the old session is dropped.
|
||||
|
||||
### Dual-Initiation Race
|
||||
|
||||
On high-latency links, both sides' rekey timers can fire close enough
|
||||
together that each peer's msg1 crosses the other in flight. Without
|
||||
arbitration, each side would act as both initiator and responder, end
|
||||
up with two different Noise sessions, and lose connectivity at cutover.
|
||||
FMP arbitrates with a deterministic tie-breaker: the peer with the
|
||||
**numerically smaller `NodeAddr`** wins the role of initiator and
|
||||
discards any inbound msg1 it sees during the race; the larger-`NodeAddr`
|
||||
peer abandons its own initiation and processes the inbound msg1 as
|
||||
responder. The same tie-breaker is applied to cross-connection races
|
||||
during initial handshake.
|
||||
|
||||
### Operator Visibility
|
||||
|
||||
Successful cutover is reported at INFO level on the K-bit observation;
|
||||
intermediate steps (handshake start, msg1/msg2 exchange, drain-window
|
||||
fallback decrypts) log at DEBUG/TRACE. Failures (handshake error,
|
||||
drain-window expiry without cutover) log at WARN.
|
||||
|
||||
The end-to-end rekey at the session layer follows a parallel design;
|
||||
see [fips-session-layer.md](fips-session-layer.md).
|
||||
|
||||
## Liveness Detection
|
||||
|
||||
FMP detects link liveness through a combination of explicit heartbeats and
|
||||
@@ -362,8 +436,8 @@ traffic observation.
|
||||
|
||||
### Heartbeat
|
||||
|
||||
A Heartbeat message (0x51) is sent to each active peer every
|
||||
`node.heartbeat_interval_secs` (default 10s). The heartbeat is a minimal
|
||||
A Heartbeat message (0x51) is sent to each active peer at a configurable
|
||||
interval (`node.heartbeat_interval_secs`). The heartbeat is a minimal
|
||||
encrypted frame with no payload beyond the standard inner header (timestamp +
|
||||
message type). Any successfully decrypted frame — data, gossip, MMP report,
|
||||
or heartbeat — resets the peer's last-receive timestamp tracked by the MMP
|
||||
@@ -371,8 +445,8 @@ receiver.
|
||||
|
||||
### Dead Timeout
|
||||
|
||||
When no traffic (of any kind) is received from a peer for
|
||||
`node.link_dead_timeout_secs` (default 30s), the peer is declared dead and
|
||||
When no traffic (of any kind) is received from a peer for the
|
||||
configured `node.link_dead_timeout_secs` window, the peer is declared dead and
|
||||
removed via `remove_active_peer()`. This triggers the full teardown cascade:
|
||||
spanning tree parent reselection (if the dead peer was the parent),
|
||||
TreeAnnounce propagation, coordinate cache flush, and bloom filter recompute.
|
||||
@@ -380,156 +454,74 @@ TreeAnnounce propagation, coordinate cache flush, and bloom filter recompute.
|
||||
If the dead peer is eligible for auto-reconnect (see [Auto-Reconnect]
|
||||
(#auto-reconnect)), reconnection is scheduled immediately after removal.
|
||||
|
||||
The heartbeat is independent of MMP — it is needed because idle links in
|
||||
Lightweight MMP mode have no guaranteed periodic traffic (gossip is
|
||||
event-driven, and MMP reports require at least one side running Full mode).
|
||||
The heartbeat is independent of MMP. Gossip is event-driven, Lightweight
|
||||
produces receiver reports only when traffic arrives, and Minimal emits no
|
||||
reports at all — so on a fully idle link no MMP-mode combination
|
||||
guarantees periodic activity. The heartbeat is the always-on liveness
|
||||
signal.
|
||||
|
||||
## Link Message Types
|
||||
|
||||
FMP defines eight message types carried inside encrypted frames:
|
||||
FMP defines several encrypted message types carried inside the
|
||||
established-frame envelope. They group naturally by purpose:
|
||||
|
||||
| Type | Name | Purpose |
|
||||
| ---- | ---- | ------- |
|
||||
| 0x10 | TreeAnnounce | Spanning tree state announcements between peers |
|
||||
| 0x20 | FilterAnnounce | Bloom filter reachability updates |
|
||||
| 0x30 | LookupRequest | Coordinate discovery — flood toward destination |
|
||||
| 0x31 | LookupResponse | Coordinate discovery — response with coordinates |
|
||||
| 0x00 | SessionDatagram | Encapsulated session-layer payload for forwarding |
|
||||
| 0x01 | SenderReport | MMP sender-side metrics report |
|
||||
| 0x02 | ReceiverReport | MMP receiver-side metrics report |
|
||||
| 0x50 | Disconnect | Orderly link teardown with reason code |
|
||||
| 0x51 | Heartbeat | Link liveness probe |
|
||||
- **Routing gossip**: TreeAnnounce carries spanning-tree announcements
|
||||
between direct peers; FilterAnnounce carries bloom-filter
|
||||
reachability updates between direct peers. Both are peer-to-peer
|
||||
(not forwarded).
|
||||
- **Discovery**: LookupRequest is forwarded through tree peers under
|
||||
bloom-filter guidance to find a destination's coordinates;
|
||||
LookupResponse routes back to the requester via reverse-path lookup
|
||||
in `recent_requests`.
|
||||
- **Forwarded payload**: SessionDatagram carries a session-layer
|
||||
payload hop-by-hop toward the destination.
|
||||
- **Metrics**: SenderReport and ReceiverReport carry the link-layer
|
||||
MMP report stream peer-to-peer.
|
||||
- **Liveness and lifecycle**: Heartbeat is a minimal frame sent
|
||||
peer-to-peer to keep the link alive; Disconnect carries an orderly
|
||||
teardown reason code peer-to-peer.
|
||||
|
||||
Additionally, handshake messages (phase 0x1 msg1, phase 0x2 msg2) are sent
|
||||
unencrypted before the link session is established.
|
||||
Handshake messages (phase 0x1 msg1, phase 0x2 msg2) travel before
|
||||
encryption is established and are identified by the FMP common-prefix
|
||||
`phase` field rather than a `msg_type` byte.
|
||||
|
||||
TreeAnnounce and FilterAnnounce are exchanged between direct peers only — they
|
||||
are not forwarded. LookupRequest and LookupResponse are forwarded through the
|
||||
mesh (flooded with deduplication). SessionDatagram is forwarded hop-by-hop
|
||||
toward the destination. Disconnect is peer-to-peer.
|
||||
|
||||
See [fips-mesh-operation.md](fips-mesh-operation.md) for how these messages
|
||||
work together to build and maintain the mesh, and
|
||||
[fips-wire-formats.md](fips-wire-formats.md) for byte-level message layouts.
|
||||
See [../reference/wire-formats.md](../reference/wire-formats.md) for
|
||||
byte-level message layouts and the canonical FMP message type
|
||||
catalog, and [fips-mesh-operation.md](fips-mesh-operation.md) for how
|
||||
these messages work together to build and maintain the mesh.
|
||||
|
||||
## Metrics Measurement Protocol (MMP)
|
||||
|
||||
Each active peer link runs an instance of the Metrics Measurement Protocol,
|
||||
providing per-link quality metrics to the operator and to the spanning tree
|
||||
layer for cost-based parent selection.
|
||||
MMP runs on every active link to provide per-link quality metrics
|
||||
(SRTT, loss, jitter, goodput, OWD trend, ETX) to the operator and to
|
||||
the spanning tree layer for cost-based parent selection. Reports are
|
||||
exchanged peer-to-peer between direct neighbors at RTT-adaptive
|
||||
intervals clamped to `[1s, 5s]`, with a 200 ms cold-start floor for
|
||||
the first five SRTT samples.
|
||||
|
||||
### Metrics Tracked
|
||||
The CE (Congestion Experienced) bit in the FMP flags byte carries
|
||||
hop-by-hop ECN signaling: transit nodes detect congestion on outgoing
|
||||
links (via MMP loss/ETX or `SO_RXQ_OVFL` kernel drops) and set CE on
|
||||
forwarded packets, which the destination then mirrors to the IPv6
|
||||
Traffic Class for ECN-capable flows.
|
||||
|
||||
MMP computes the following metrics from the per-frame counter and timestamp
|
||||
fields in the FMP wire format:
|
||||
|
||||
- **SRTT** — Smoothed round-trip time (Jacobson/RFC 6298, α=1/8). Derived
|
||||
from timestamp-echo in ReceiverReports with dwell-time compensation.
|
||||
- **Loss rate** — Bidirectional loss inferred from counter gaps. Tracked as
|
||||
both instantaneous (per-interval) and long-term EWMA.
|
||||
- **Jitter** — Interarrival jitter (RFC 3550 algorithm) in microseconds.
|
||||
- **Goodput** — Bytes per second of payload data (excludes MMP reports).
|
||||
- **OWD trend** — One-way delay trend (µs/s, signed). Indicates congestion
|
||||
buildup before loss occurs.
|
||||
- **ETX** — Expected Transmission Count, computed from bidirectional delivery
|
||||
ratios. Used in cost-based parent selection via
|
||||
`link_cost = etx * (1.0 + srtt_ms / 100.0)`; not yet used in
|
||||
`find_next_hop()` candidate ranking.
|
||||
- **Dual EWMA trends** — Short-term (α=1/4) and long-term (α=1/32) trend
|
||||
indicators for both RTT and loss, enabling change detection.
|
||||
|
||||
### Operating Modes
|
||||
|
||||
MMP supports three modes, configured via `node.mmp.mode`:
|
||||
|
||||
| Mode | Reports Exchanged | Metrics Available |
|
||||
| ---- | ----------------- | ----------------- |
|
||||
| **Full** (default) | SenderReport + ReceiverReport | All metrics including RTT, loss, jitter, goodput, OWD trend |
|
||||
| **Lightweight** | ReceiverReport only | Loss (from counter gaps), jitter, OWD trend. No RTT. |
|
||||
| **Minimal** | None | Spin bit and CE echo flags only. No computed metrics. |
|
||||
|
||||
### Report Scheduling
|
||||
|
||||
Reports are sent at RTT-adaptive intervals, clamped to [100ms, 2s]. A
|
||||
cold-start interval of 500ms is used before SRTT converges. The interval
|
||||
formula is `clamp(2 × SRTT, 100ms, 2000ms)`.
|
||||
|
||||
### Spin Bit and RTT
|
||||
|
||||
The SP (spin bit) flag in the FMP inner header follows the QUIC spin bit
|
||||
pattern: reflected on receive, toggled on send when the reflected value
|
||||
matches the last sent value. The spin bit state machine runs for TX
|
||||
reflection, but **RTT samples from the spin bit are discarded**. In a mesh
|
||||
protocol where frames are sent irregularly (tree announces, bloom filters,
|
||||
MMP reports on different timers), inter-frame processing delays inflate spin
|
||||
bit RTT measurements unpredictably. Timestamp-echo from ReceiverReports
|
||||
(with dwell-time compensation) is the sole SRTT source.
|
||||
|
||||
### ECN Congestion Signaling
|
||||
|
||||
The CE (Congestion Experienced) flag (bit 1 in the FMP flags byte) provides
|
||||
hop-by-hop congestion signaling through the mesh. Transit nodes detect
|
||||
congestion on outgoing links and set CE on forwarded packets; once set, the
|
||||
flag stays set for all subsequent hops to the destination.
|
||||
|
||||
**Congestion detection** (`detect_congestion()`) triggers on any of:
|
||||
|
||||
- Outgoing link MMP loss rate ≥ `node.ecn.loss_threshold` (default 5%)
|
||||
- Outgoing link MMP ETX ≥ `node.ecn.etx_threshold` (default 3.0)
|
||||
- Kernel receive buffer drops detected on any local transport (via
|
||||
`SO_RXQ_OVFL` on UDP)
|
||||
|
||||
**CE relay**: The forwarding path computes `outgoing_ce = incoming_ce ||
|
||||
local_congestion`. The `send_encrypted_link_message_with_ce()` method ORs
|
||||
`FLAG_CE` into the FMP header flags when ce is true. The original
|
||||
`send_encrypted_link_message()` delegates with `ce_flag=false`, leaving the
|
||||
20+ existing call sites unchanged.
|
||||
|
||||
**IPv6 ECN-CE marking**: When a CE-flagged DataPacket arrives at its final
|
||||
destination, the IPv6 Traffic Class ECN bits are marked CE (0b11) before
|
||||
TUN delivery — but only for ECN-capable packets (ECT(0) or ECT(1)). Not-ECT
|
||||
packets are never marked per RFC 3168. The host TCP stack then echoes ECE in
|
||||
ACKs, triggering sender cwnd reduction through standard congestion control.
|
||||
|
||||
**Session-layer tracking**: The `ecn_ce_count` field in MMP ReceiverReports
|
||||
tracks CE-flagged packets received per link, providing end-to-end visibility
|
||||
into congestion propagation.
|
||||
|
||||
**Monitoring**: `CongestionStats` tracks four counters — `ce_forwarded`,
|
||||
`ce_received`, `congestion_detected`, and `kernel_drop_events` — exposed via
|
||||
`fipsctl show routing` (congestion block) and `fipstop` (routing tab).
|
||||
Rate-limited warn logging (5s interval) alerts on congestion detection events.
|
||||
|
||||
See `node.ecn.*` in
|
||||
[fips-configuration.md](fips-configuration.md#ecn-signaling-nodeecn) for
|
||||
tuning parameters.
|
||||
|
||||
### Operator Logging
|
||||
|
||||
MMP emits periodic link metrics at info level (configurable via
|
||||
`node.mmp.log_interval_secs`, default 30s):
|
||||
|
||||
```text
|
||||
MMP link metrics peer=node-b rtt=2.3ms loss=0.2% jitter=0.1ms goodput=76.0MB/s tx_pkts=1234 rx_pkts=5678
|
||||
```
|
||||
|
||||
Teardown logs include final SRTT, loss rate, jitter, ETX, goodput, and
|
||||
cumulative tx/rx packet and byte counts.
|
||||
For the full MMP design — operating modes, report scheduling, spin
|
||||
bit interaction, ECN, and the algorithmic details shared with
|
||||
session-layer MMP — see [fips-mmp.md](fips-mmp.md). For the
|
||||
SenderReport and ReceiverReport byte layouts, see
|
||||
[../reference/wire-formats.md](../reference/wire-formats.md).
|
||||
Configuration knobs live under `node.mmp.*` and `node.ecn.*` in
|
||||
[../reference/configuration.md](../reference/configuration.md).
|
||||
|
||||
## Security Properties
|
||||
|
||||
### Threat Resistance
|
||||
|
||||
| Threat | Mitigation |
|
||||
| ------ | ---------- |
|
||||
| Connection exhaustion | Token bucket rate limit + connection count limit |
|
||||
| CPU exhaustion (msg1 flood) | Rate limit before crypto operations |
|
||||
| Replay attacks | Counter-based nonces with sliding window |
|
||||
| State confusion | Strict handshake state machine validation |
|
||||
| Spoofed encrypted packets | Index lookup + AEAD verification |
|
||||
| Spoofed msg2 | Index lookup + Noise ephemeral key binding |
|
||||
| Address spoofing | Cryptographic authority, not address-based |
|
||||
| Session correlation | Index rotation on rekey |
|
||||
The link-layer threat-resistance matrix (connection exhaustion, CPU
|
||||
exhaustion, replay, state confusion, spoofing variants, address
|
||||
spoofing, session correlation) is consolidated in
|
||||
[../reference/security.md](../reference/security.md) along with the
|
||||
session-layer matrix and operator-facing controls.
|
||||
|
||||
### Unauthenticated Attack Surface
|
||||
|
||||
@@ -574,13 +566,17 @@ an attacker sends invalid packets to elicit responses.
|
||||
| Metrics Measurement Protocol (MMP) | **Implemented** |
|
||||
| ECN congestion signaling (CE relay, IPv6 marking) | **Implemented** |
|
||||
| Rekey with index rotation | **Implemented** |
|
||||
| Allowlist/blocklist | Planned |
|
||||
| Peer ACL (allowlist / denylist) | **Implemented** |
|
||||
|
||||
## References
|
||||
|
||||
- [fips-intro.md](fips-intro.md) — Protocol overview and architecture
|
||||
- [fips-concepts.md](fips-concepts.md) — Protocol overview
|
||||
- [fips-architecture.md](fips-architecture.md) — Layer architecture and
|
||||
identity model
|
||||
- [fips-transport-layer.md](fips-transport-layer.md) — Transport layer (below FMP)
|
||||
- [fips-session-layer.md](fips-session-layer.md) — FSP (above FMP)
|
||||
- [fips-mmp.md](fips-mmp.md) — Metrics Measurement Protocol (link + session)
|
||||
- [fips-mesh-operation.md](fips-mesh-operation.md) — How FMP's routing and
|
||||
self-organization work in practice
|
||||
- [fips-wire-formats.md](fips-wire-formats.md) — Byte-level wire format reference
|
||||
- [../reference/wire-formats.md](../reference/wire-formats.md) — Byte-level
|
||||
wire format reference
|
||||
|
||||
@@ -31,183 +31,51 @@ and self-healing.
|
||||
|
||||
## Spanning Tree Formation and Maintenance
|
||||
|
||||
### What the Spanning Tree Provides
|
||||
For routing purposes, the spanning tree provides each node with a
|
||||
coordinate (its ancestry path from itself to the root) plus a way to
|
||||
compute distance between any two nodes (hops to their lowest common
|
||||
ancestor). The strictly-decreasing distance invariant gives greedy
|
||||
forwarding its loop-freedom.
|
||||
|
||||
The spanning tree gives each node a **coordinate**: its ancestry path from
|
||||
itself to the root, expressed as a sequence of node_addrs. These coordinates
|
||||
enable:
|
||||
The tree forms through distributed parent selection — root is the
|
||||
smallest node_addr (no election), and each node picks the peer with
|
||||
the lowest `effective_depth = depth + link_cost`. Cost-aware parent
|
||||
selection lets the tree trade hop count for link quality once MMP has
|
||||
accumulated SRTT and ETX metrics. Hysteresis (20% improvement
|
||||
required to switch) and hold-down (suppress non-mandatory
|
||||
re-evaluation after a switch) keep the tree stable under metric
|
||||
noise. Partitions self-resolve — each segment converges to its own
|
||||
root and reconverges to the smallest reachable root when segments
|
||||
rejoin.
|
||||
|
||||
- **Distance calculation**: The tree distance between two nodes is the number
|
||||
of hops from each to their lowest common ancestor (LCA). This provides a
|
||||
routing metric without any node knowing the full topology.
|
||||
- **Greedy routing**: At each hop, forward to the peer that minimizes tree
|
||||
distance to the destination. The strictly-decreasing distance invariant
|
||||
guarantees loop-free forwarding.
|
||||
Liveness is detected via FMP heartbeats; dead-peer removal triggers
|
||||
tree reconvergence and bloom filter recomputation for the affected
|
||||
subtree. The heartbeat and dead-timeout mechanism lives at the link
|
||||
layer; see [fips-mesh-layer.md](fips-mesh-layer.md#liveness-detection).
|
||||
|
||||
### How the Tree Forms
|
||||
|
||||
Nodes self-organize into a spanning tree through distributed parent selection:
|
||||
|
||||
1. **Root discovery**: The node with the smallest node_addr becomes the root.
|
||||
No election protocol — this is a consequence of each node independently
|
||||
preferring lower-addressed roots.
|
||||
2. **Parent selection**: Each node selects a single parent from among its
|
||||
direct peers based on which offers the lowest effective depth (tree depth
|
||||
weighted by local link cost).
|
||||
3. **Coordinate computation**: Once a node has a parent, its coordinate is
|
||||
computed from its ancestry path.
|
||||
|
||||
### How the Tree Maintains Itself
|
||||
|
||||
Nodes exchange **TreeAnnounce** messages with their direct peers (not
|
||||
forwarded — peer-to-peer only). Each TreeAnnounce carries the sender's
|
||||
current ancestry chain and a sequence number.
|
||||
|
||||
Changes cascade through the tree:
|
||||
|
||||
- A node that changes its parent recomputes its coordinates and announces to
|
||||
all peers
|
||||
- Each receiving peer evaluates whether the change affects its own parent
|
||||
selection
|
||||
- Only nodes that actually change their coordinates (root or depth changed)
|
||||
propagate further
|
||||
|
||||
TreeAnnounce propagation is rate-limited at 500ms minimum interval per peer.
|
||||
A tree of depth D reconverges in roughly D×0.5s to D×1.0s.
|
||||
|
||||
### How the Tree Adapts to Link Quality
|
||||
|
||||
The initial tree forms based on hop count alone — all links default to a
|
||||
cost of 1.0 before measurements are available. As the Metrics Measurement
|
||||
Protocol (MMP) accumulates bidirectional delivery ratios and round-trip
|
||||
time estimates, each node computes a per-link cost:
|
||||
|
||||
```text
|
||||
link_cost = ETX × (1.0 + SRTT_ms / 100.0)
|
||||
```
|
||||
|
||||
ETX (Expected Transmission Count) captures loss — a perfect link has
|
||||
ETX = 1.0, while 10% loss in each direction yields ETX ≈ 1.23. The SRTT
|
||||
term weights latency so that a low-loss but high-latency link (e.g., a
|
||||
satellite hop) costs more than a low-loss, low-latency link.
|
||||
|
||||
Parent selection uses **effective depth** rather than raw hop count:
|
||||
|
||||
```text
|
||||
effective_depth = peer.depth + link_cost_to_peer
|
||||
```
|
||||
|
||||
This allows a node to trade a shorter but lossy path for a longer but
|
||||
higher-quality one. A node two hops from the root over clean links
|
||||
(effective depth ≈ 3.0) is preferred over a node one hop away over a
|
||||
degraded link (effective depth ≈ 4.5).
|
||||
|
||||
Parent reselection is triggered by three paths:
|
||||
|
||||
1. **TreeAnnounce**: When a peer announces a new tree position, the node
|
||||
re-evaluates using current link costs
|
||||
2. **Periodic re-evaluation**: Every 60s (configurable), the node
|
||||
re-evaluates its parent choice using the latest MMP metrics, catching
|
||||
gradual link degradation that doesn't trigger TreeAnnounce
|
||||
3. **Parent loss**: When the current parent is removed, the node
|
||||
immediately selects the best alternative
|
||||
|
||||
To prevent oscillation from metric noise, parent switches are subject to
|
||||
**hysteresis**: a candidate must offer an effective depth at least 20%
|
||||
better than the current parent to trigger a switch. A **hold-down period**
|
||||
(default 30s) suppresses non-mandatory re-evaluation after a switch,
|
||||
allowing MMP metrics to stabilize on the new link before reconsidering.
|
||||
|
||||
### Flap Dampening
|
||||
|
||||
Unstable links that repeatedly connect and disconnect can cause cascading
|
||||
tree reconvergence. The spanning tree uses flap dampening with hysteresis
|
||||
and hold-down periods to suppress rapid parent oscillation. Links that flap
|
||||
above a configurable threshold are temporarily penalized, preventing them
|
||||
from being selected as parent until the link stabilizes.
|
||||
|
||||
### Link Liveness
|
||||
|
||||
Each node sends a dedicated **Heartbeat** message (0x51, 1 byte, no
|
||||
payload) to every peer at a fixed interval (default 10s). Any
|
||||
authenticated encrypted frame — heartbeat, MMP report, TreeAnnounce,
|
||||
data packet — resets the peer's liveness timer. On an idle link with no
|
||||
application data or topology changes, the heartbeat is the only traffic
|
||||
that keeps the link alive.
|
||||
|
||||
Peers that are silent for a configurable dead timeout (default 30s) are
|
||||
considered dead and removed from the peer table. With the default 10s
|
||||
heartbeat interval, a peer must miss three consecutive heartbeats before
|
||||
removal. This triggers tree reconvergence and bloom filter recomputation
|
||||
for the affected subtree.
|
||||
|
||||
### Partition Handling
|
||||
|
||||
If the network partitions, each segment independently rediscovers its own
|
||||
root (the smallest node_addr in the segment) and reconverges. When segments
|
||||
rejoin, nodes discover the globally-smallest root through TreeAnnounce
|
||||
exchange and reconverge to a single tree.
|
||||
|
||||
See [fips-spanning-tree.md](fips-spanning-tree.md) for algorithm details
|
||||
and [spanning-tree-dynamics.md](spanning-tree-dynamics.md) for convergence
|
||||
walkthroughs.
|
||||
For the parent-selection algorithm, hold-down/hysteresis details, and
|
||||
the convergence walkthroughs, see
|
||||
[fips-spanning-tree.md](fips-spanning-tree.md) and
|
||||
[spanning-tree-dynamics.md](spanning-tree-dynamics.md).
|
||||
|
||||
## Bloom Filter Gossip and Propagation
|
||||
|
||||
### What Bloom Filters Provide
|
||||
For routing purposes, each node maintains a bloom filter per peer
|
||||
that answers "can peer P possibly reach destination D?" — either "no"
|
||||
(definitive) or "maybe" (probabilistic). Because filters propagate
|
||||
along tree edges with split-horizon exclusion, a bloom hit on a tree
|
||||
peer reliably indicates which subtree contains the destination, and
|
||||
tree-coordinate distance ranks competing matches.
|
||||
|
||||
Each node maintains a bloom filter per peer, answering: "can peer P possibly
|
||||
reach destination D?" The answer is either "no" (definitive) or "maybe"
|
||||
(probabilistic — false positives are possible).
|
||||
FilterAnnounce updates are event-driven (peer changes, tree
|
||||
restructuring, local identity changes) and rate-limited to prevent
|
||||
storms. False positives at large scale never cause loops — the
|
||||
self-distance check at each hop guarantees forward progress, and
|
||||
mismatched bloom matches fall through to greedy tree routing.
|
||||
|
||||
Because filters propagate along tree edges with split-horizon exclusion,
|
||||
they encode directional reachability: a bloom hit on a tree peer reliably
|
||||
indicates which subtree contains the destination. When multiple peers match,
|
||||
tree coordinate distance ranks them.
|
||||
|
||||
### How Filters Propagate
|
||||
|
||||
Nodes exchange **FilterAnnounce** messages with all direct peers. Each
|
||||
FilterAnnounce replaces the previous filter for that peer — there is no
|
||||
incremental update.
|
||||
|
||||
Filter computation uses **tree-only merge with split-horizon exclusion**:
|
||||
the outbound filter for peer Q is computed by merging the local node's own
|
||||
identity, its leaf-only dependents (if any), and the inbound filters from
|
||||
tree peers (parent and children) *except* Q. Filters from non-tree mesh
|
||||
peers are stored locally for routing queries but are not merged into
|
||||
outgoing filters. This prevents saturation where mesh shortcuts cause
|
||||
filters to converge toward the full network.
|
||||
|
||||
The restriction creates **directional asymmetry**: upward filters
|
||||
(child → parent) contain the child's subtree, while downward filters
|
||||
(parent → child) contain the complement. Together they cover the entire
|
||||
network.
|
||||
|
||||
Filters propagate transitively through tree edges. At steady state, every
|
||||
reachable destination appears in at least one tree peer's filter.
|
||||
|
||||
### Update Triggers
|
||||
|
||||
Filter updates are event-driven, not periodic:
|
||||
|
||||
- Peer connects or disconnects
|
||||
- A peer's incoming filter changes (triggers recomputation for other peers)
|
||||
- Tree relationship changes (new parent, new child, parent switch)
|
||||
- Local state changes (new identity, leaf-only dependent changes)
|
||||
|
||||
Updates are rate-limited at 500ms to prevent storms during topology changes.
|
||||
|
||||
### Scale Properties
|
||||
|
||||
At moderate network sizes, bloom filters are highly accurate. At larger
|
||||
scales (~1M nodes), hub nodes with many peers may see elevated false positive
|
||||
rates (7–15% for nodes with 20+ peers). False positives may cause a packet
|
||||
to be forwarded toward the wrong subtree, but the self-distance check at
|
||||
each hop prevents loops and the packet falls through to greedy tree routing.
|
||||
|
||||
See [fips-bloom-filters.md](fips-bloom-filters.md) for filter parameters,
|
||||
FPR calculations, and size class folding.
|
||||
For the filter computation, split-horizon merge rules, FPR analysis,
|
||||
size classes, and folding, see
|
||||
[fips-bloom-filters.md](fips-bloom-filters.md).
|
||||
|
||||
## Routing Decision Process
|
||||
|
||||
@@ -222,34 +90,33 @@ priority chain. This is the core routing algorithm.
|
||||
2. **Direct peer** — The destination is an authenticated neighbor. Forward
|
||||
directly. No coordinates or bloom filters needed.
|
||||
|
||||
3. **Bloom-guided routing** — One or more peers' bloom filters contain the
|
||||
3. **Coordinate cache check** — Multi-hop forwarding requires the
|
||||
destination's tree coordinates to be in the local cache. On miss,
|
||||
`find_next_hop()` returns None immediately — bloom filters are never
|
||||
consulted — and the source receives a CoordsRequired error signal.
|
||||
|
||||
4. **Bloom-guided routing** — One or more peers' bloom filters contain the
|
||||
destination. Select the best peer by composite key:
|
||||
`(link_cost, tree_distance, node_addr)`. This requires the destination's
|
||||
tree coordinates to be in the local coordinate cache.
|
||||
`(link_cost, tree_distance, node_addr)`.
|
||||
|
||||
4. **Greedy tree routing** — Fallback when bloom filters haven't converged
|
||||
for this destination. Forward to the peer that minimizes tree distance.
|
||||
Also requires destination coordinates.
|
||||
5. **Greedy tree routing** — Fall-through when bloom yields no candidate.
|
||||
Forward to the peer that minimizes tree distance. If the tree has no
|
||||
next hop closer to the destination, the source receives a PathBroken
|
||||
error signal.
|
||||
|
||||
5. **No route** — Destination unreachable. Generate an error signal
|
||||
(CoordsRequired or PathBroken) back to the source.
|
||||
### Convergence Requirements
|
||||
|
||||
### The Coordinate Requirement
|
||||
|
||||
All multi-hop routing (steps 3–4) requires the destination's tree coordinates
|
||||
to be in the local coordinate cache. Without coordinates, `find_next_hop()`
|
||||
returns None immediately — bloom filters are never even consulted.
|
||||
|
||||
This creates two simultaneous convergence requirements for multi-hop routing:
|
||||
Multi-hop routing depends on two propagation processes that must run
|
||||
to convergence simultaneously:
|
||||
|
||||
1. **Bloom convergence**: Filters must propagate so peers advertise
|
||||
reachability
|
||||
2. **Coordinate availability**: Destination coordinates must be cached at
|
||||
every transit node on the path
|
||||
|
||||
Both must be satisfied simultaneously. Bloom convergence without coordinates
|
||||
causes a coordinate cache miss. Coordinates without bloom convergence falls
|
||||
through to greedy tree routing (functional but suboptimal).
|
||||
Bloom convergence without coordinates trips step 3 (coord-cache miss →
|
||||
CoordsRequired). Coordinates without bloom convergence falls through to
|
||||
greedy tree routing — functional but suboptimal.
|
||||
|
||||
### Candidate Ranking
|
||||
|
||||
@@ -270,6 +137,10 @@ A peer with a bloom filter hit but no entry in the peer ancestry table
|
||||
(missing TreeAnnounce) defaults to maximum distance and is effectively
|
||||
invisible to routing.
|
||||
|
||||
### Routing Decision Flowchart
|
||||
|
||||

|
||||
|
||||
### Loop Prevention
|
||||
|
||||
The routing decision enforces strict progress: a packet is only forwarded
|
||||
@@ -284,53 +155,19 @@ PathBroken error.
|
||||
|
||||
## Coordinate Caching
|
||||
|
||||
The coordinate cache maps `NodeAddr → TreeCoordinate` and is the critical
|
||||
data structure for multi-hop routing. Without it, forwarding decisions cannot
|
||||
be made.
|
||||
|
||||
### Unified Cache
|
||||
|
||||
The coordinate cache is a single unified cache. All sources — SessionSetup
|
||||
transit, CP-flagged data packets, LookupResponse — write to the same cache.
|
||||
|
||||
### Population Sources
|
||||
|
||||
| Source | When | What |
|
||||
| ------ | ---- | ---- |
|
||||
| SessionSetup transit | Session establishment | Both src and dest coordinates |
|
||||
| SessionAck transit | Session establishment | Both src and dest coordinates |
|
||||
| CP-flagged data packet | Warmup or recovery | Both src and dest coordinates (cleartext) |
|
||||
| LookupResponse | Discovery | Target's coordinates |
|
||||
|
||||
### Eviction
|
||||
|
||||
- **TTL-based**: Entries expire after 300s (configurable)
|
||||
- **Refresh on use**: Active routing refreshes the TTL, keeping hot entries
|
||||
alive
|
||||
- **LRU**: When full, least recently used entries are evicted first
|
||||
- **Flush on parent change**: When the local node's tree parent changes, the
|
||||
entire cache is flushed. Parent changes mean the node's own coordinates
|
||||
have changed, making relative distance calculations with cached coordinates
|
||||
potentially invalid. Flushing is preferred over stale routing: the cost of
|
||||
re-discovery is lower than routing packets to dead ends.
|
||||
|
||||
### Cache and Session Timer Ordering
|
||||
|
||||
Timer values are ordered so that idle sessions tear down before transit
|
||||
caches expire:
|
||||
|
||||
| Timer | Default | Purpose |
|
||||
| ----- | ------- | ------- |
|
||||
| Session idle | 90s | Session teardown |
|
||||
| Coordinate cache TTL | 300s | Coordinate expiration |
|
||||
|
||||
When traffic stops, the session tears down at 90s. When traffic resumes, a
|
||||
fresh SessionSetup re-warms transit caches (still within their 300s TTL).
|
||||
The coordinate cache maps `NodeAddr → TreeCoordinate` and is the
|
||||
critical data structure for multi-hop routing. The session layer owns
|
||||
this cache (its eviction policy, TTL/refresh semantics, parent-change
|
||||
flush, and timer ordering with session idle timeout); see
|
||||
[fips-session-layer.md](fips-session-layer.md#coordinate-cache) for
|
||||
the canonical treatment.
|
||||
|
||||
## Discovery Protocol
|
||||
|
||||
Discovery resolves a destination's tree coordinates so that multi-hop routing
|
||||
can proceed.
|
||||
can proceed. Requests are forwarded using **bloom-guided tree routing** —
|
||||
only to tree peers (parent + children) whose bloom filter contains the
|
||||
target — producing single-path forwarding through the spanning tree.
|
||||
|
||||
### When Discovery Is Needed
|
||||
|
||||
@@ -347,17 +184,71 @@ The source creates a LookupRequest containing:
|
||||
- **origin**: The requester's node_addr
|
||||
- **origin_coords**: The requester's current tree coordinates (so the
|
||||
response can route back)
|
||||
- **TTL**: Bounds the flood radius
|
||||
- **TTL**: Bounds the forwarding radius
|
||||
|
||||
The request floods through the mesh: each node decrements TTL, adds itself
|
||||
to a visited filter (preventing loops on a single path), and forwards to all
|
||||
peers not in the visited filter. Bloom filters may help direct the flood
|
||||
toward likely candidates.
|
||||
### Bloom-Guided Tree Routing
|
||||
|
||||
**Deduplication**: Nodes maintain a short-lived request_id dedup cache
|
||||
(default 10s window) to drop convergent duplicates (the same request
|
||||
arriving via different paths). This is a protocol requirement, not an
|
||||
optimization.
|
||||
Rather than flooding to all peers, the request is forwarded only to **tree
|
||||
peers** (parent + children) whose bloom filter contains the target. Because
|
||||
bloom filters propagate along tree edges with split-horizon exclusion,
|
||||
typically only one tree peer matches — producing a single directed path
|
||||
through the spanning tree toward the target's subtree. This reduces
|
||||
discovery traffic by roughly 90% compared to flooding.
|
||||
|
||||
If no tree peer's bloom filter matches the target, the request falls back
|
||||
to **non-tree peers** whose bloom filter contains the target. This recovers
|
||||
from dead ends caused by stale bloom filters, tree restructuring, or transit
|
||||
node failures. If no peer at all has a bloom match, the request is dropped
|
||||
at that node.
|
||||
|
||||
**Loop prevention**: The spanning tree is inherently loop-free, so tree-only
|
||||
forwarding cannot loop. The `request_id` dedup cache (default 10s window)
|
||||
provides defense-in-depth, catching edge cases during tree restructuring
|
||||
where a request might arrive via both tree and fallback paths.
|
||||
|
||||
### Retry Logic
|
||||
|
||||
Single-path forwarding is more fragile than flooding — if any transit node
|
||||
on the path has a stale bloom filter or loses a link, the request fails.
|
||||
To compensate, each discovery is a sequence of attempts with growing
|
||||
per-attempt timeouts. The default sequence is `[1s, 2s, 4s, 8s]`
|
||||
(configurable via `node.discovery.attempt_timeouts_secs`); the destination
|
||||
is declared unreachable only after the full sequence is exhausted (15s
|
||||
total at default).
|
||||
|
||||
When the current attempt's deadline elapses without a `LookupResponse`,
|
||||
the originator sends another `LookupRequest` with a **fresh `request_id`**
|
||||
and the next entry in the sequence as its deadline. Fresh `request_id`s
|
||||
let each attempt take a different forwarding path as the bloom and tree
|
||||
state evolve, which is particularly useful during cold-start convergence.
|
||||
|
||||
### Originator Backoff (optional, off by default)
|
||||
|
||||
After the per-attempt sequence is exhausted, the originator can additionally
|
||||
suppress further fresh lookups for the same target with exponential
|
||||
post-failure backoff. This is **disabled by default** (`backoff_base_secs:
|
||||
0`); the per-attempt sequence is the only retry pacing in the standard
|
||||
configuration. Operators may opt in via `node.discovery.backoff_base_secs`
|
||||
and `node.discovery.backoff_max_secs` if their deployment has chatty apps
|
||||
generating repeated lookups for genuinely unreachable destinations. When
|
||||
enabled, backoff is **reset on topology changes** that might make
|
||||
previously unreachable targets reachable: parent switch, new peer
|
||||
connection, first RTT measurement from MMP, or peer reconnection.
|
||||
|
||||
### Bloom Filter Pre-Check
|
||||
|
||||
Before initiating a lookup, the originator checks whether *any* peer's
|
||||
bloom filter contains the target. If no peer advertises reachability, the
|
||||
lookup is skipped entirely and recorded as a failure for backoff purposes.
|
||||
This avoids wasting network resources when the target is not in the mesh.
|
||||
|
||||
### Transit-Side Rate Limiting
|
||||
|
||||
Transit nodes enforce a per-target minimum interval (default 2s, configurable
|
||||
via `forward_min_interval_secs`) for forwarded lookups. This is
|
||||
defense-in-depth against misbehaving nodes that generate fresh `request_id`s
|
||||
at high rate to bypass dedup. The rate limiter collapses rapid-fire lookups
|
||||
for the same target regardless of `request_id`.
|
||||
|
||||
### LookupResponse
|
||||
|
||||
@@ -373,14 +264,21 @@ direct peer), a LookupResponse is created containing:
|
||||
authenticates that the response is genuine and the target holds the
|
||||
claimed tree position
|
||||
|
||||
The response routes back to the requester using reverse-path routing as the
|
||||
primary mechanism: each transit node looks up the request_id in its
|
||||
The response routes back to the requester using **reverse-path routing** as
|
||||
the primary mechanism: each transit node looks up the `request_id` in its
|
||||
`recent_requests` table to find the peer that forwarded the original request,
|
||||
and sends the response back through that peer. This ensures the response
|
||||
follows the same path as the request. Greedy tree routing toward the
|
||||
origin_coords is used only as a fallback if the reverse-path entry has
|
||||
`origin_coords` is used only as a fallback if the reverse-path entry has
|
||||
expired.
|
||||
|
||||
**Response-forwarded flag**: Each `recent_requests` entry tracks whether a
|
||||
response has already been forwarded for that `request_id`. If a second
|
||||
response arrives (e.g., from convergent request paths that reached the
|
||||
target via different routes), the transit node drops it. This prevents
|
||||
response routing loops where multiple responses for the same request
|
||||
circulate through the network.
|
||||
|
||||
**Proof verification**: The source verifies the Schnorr proof upon receipt,
|
||||
confirming that the target actually signed the response. The proof covers
|
||||
`(request_id || target || target_coords)` — coordinates are included because
|
||||
@@ -388,59 +286,34 @@ verification at the source confirms the target holds the claimed position.
|
||||
The `path_mtu` field is excluded from the proof because it is a transit
|
||||
annotation modified at each hop.
|
||||
|
||||
### Coordinate Discovery Sequence
|
||||
|
||||

|
||||
|
||||
### Discovery Outcome
|
||||
|
||||
On receiving a LookupResponse, the source caches the target's coordinates.
|
||||
Subsequent routing to that destination can proceed via the normal
|
||||
`find_next_hop()` priority chain.
|
||||
On receiving a verified LookupResponse, the source caches the target's
|
||||
coordinates and clears any backoff state for that target. Subsequent routing
|
||||
to that destination can proceed via the normal `find_next_hop()` priority
|
||||
chain.
|
||||
|
||||
If discovery times out (no response), queued packets receive ICMPv6
|
||||
Destination Unreachable.
|
||||
If discovery times out (no response after all retry attempts), queued
|
||||
packets receive ICMPv6 Destination Unreachable and the target enters
|
||||
backoff.
|
||||
|
||||
## SessionSetup Self-Bootstrapping
|
||||
## Coordinate Cache Warming
|
||||
|
||||
SessionSetup is the mechanism that warms transit node coordinate caches
|
||||
along a path, enabling subsequent data packets to route efficiently.
|
||||
|
||||
### How It Works
|
||||
|
||||
SessionSetup carries plaintext coordinates (outside the Noise handshake
|
||||
payload, visible to transit nodes):
|
||||
|
||||
- **src_coords**: Source's current tree coordinates
|
||||
- **dest_coords**: Destination's tree coordinates (learned from discovery)
|
||||
|
||||
As the SessionSetup transits each intermediate node:
|
||||
|
||||
1. The transit node extracts both coordinate sets
|
||||
2. Caches `src_addr → src_coords` and `dest_addr → dest_coords` in its
|
||||
coordinate cache
|
||||
3. Forwards the message using the cached destination coordinates
|
||||
|
||||
SessionAck returns along the reverse path, carrying both the responder's
|
||||
and initiator's coordinates and warming caches in the other direction. This
|
||||
ensures return-path transit nodes can route even when the reverse path
|
||||
diverges from the forward path (e.g., after tree reconvergence).
|
||||
|
||||
### Result
|
||||
|
||||
After the handshake completes, the entire forward and reverse paths have
|
||||
cached coordinates for both endpoints. Subsequent data packets use minimal
|
||||
headers (no coordinates) and route efficiently through the warmed caches.
|
||||
|
||||
## Hybrid Coordinate Warmup (CP + CoordsWarmup)
|
||||
|
||||
The CP flag in the FSP common prefix and the standalone CoordsWarmup message
|
||||
(0x14) together provide a hybrid cache-warming mechanism that complements
|
||||
SessionSetup. See [fips-session-layer.md](fips-session-layer.md) for the
|
||||
full warmup strategy.
|
||||
|
||||
Transit nodes parse the CP flag from the FSP header and extract source and
|
||||
destination coordinates from the cleartext section between the header and
|
||||
ciphertext — no decryption needed. This is the same caching operation
|
||||
performed for SessionSetup coordinates. CoordsWarmup messages use the same
|
||||
CP-flag format and are handled identically by transit nodes via the existing
|
||||
`try_warm_coord_cache()` path.
|
||||
SessionSetup carries plaintext source and destination coordinates,
|
||||
which transit nodes cache as the message travels — warming the
|
||||
forward path. SessionAck carries them back along the reverse path,
|
||||
warming return-path caches. Steady-state data packets piggyback
|
||||
coordinates via the FSP CP flag during the warmup window, falling
|
||||
back to standalone CoordsWarmup messages when piggybacking would
|
||||
exceed the transport MTU. See
|
||||
[fips-session-layer.md](fips-session-layer.md#hybrid-coordinate-warmup-strategy)
|
||||
for the canonical hybrid-warmup design (SessionSetup
|
||||
self-bootstrapping plus CP-flag piggyback plus standalone
|
||||
CoordsWarmup).
|
||||
|
||||
## Error Recovery
|
||||
|
||||
@@ -467,7 +340,7 @@ coordinates for the destination. It cannot make a forwarding decision.
|
||||
2. Reset CP warmup counter — subsequent data packets piggyback coordinates
|
||||
when possible, or trigger additional CoordsWarmup messages when
|
||||
piggybacking would exceed the transport MTU
|
||||
3. Initiate discovery (LookupRequest flood) for the destination
|
||||
3. Initiate discovery (bloom-guided LookupRequest) for the destination
|
||||
4. When discovery completes, warmup counter resets again (covers timing gap)
|
||||
|
||||
The crypto session remains active throughout — only routing state is
|
||||
@@ -553,8 +426,9 @@ sequence:
|
||||
identity cache with NodeAddr + PublicKey
|
||||
2. **Session initiation attempt**: Fails because no coordinates are cached
|
||||
for the destination
|
||||
3. **Discovery**: LookupRequest floods through the mesh; LookupResponse
|
||||
returns the destination's coordinates
|
||||
3. **Discovery**: LookupRequest routes through the spanning tree via
|
||||
bloom-guided forwarding; LookupResponse returns the destination's
|
||||
coordinates
|
||||
4. **Session establishment**: SessionSetup carries coordinates, warming
|
||||
transit caches along the path
|
||||
5. **Warmup**: First N data packets include CP flag, reinforcing transit
|
||||
@@ -639,21 +513,10 @@ routing decisions but retains its own end-to-end encryption and identity.
|
||||
|
||||
## Packet Type Summary
|
||||
|
||||
| Message | Typical Size | When | Forwarded? |
|
||||
| ------- | ------------ | ---- | ---------- |
|
||||
| TreeAnnounce | Variable (depth-dependent) | Topology changes | No (peer-to-peer) |
|
||||
| FilterAnnounce | ~1 KB | Topology changes | No (peer-to-peer) |
|
||||
| LookupRequest | ~300 bytes | First contact, recovery | Yes (flood) |
|
||||
| LookupResponse | ~400 bytes | Response to discovery | Yes (greedy routed) |
|
||||
| SessionDatagram + SessionSetup | ~232–402 bytes | Session establishment | Yes (routed) |
|
||||
| SessionDatagram + SessionAck | ~170 bytes | Session confirmation | Yes (routed) |
|
||||
| SessionDatagram + Data (minimal) | 77 bytes + IPv6 payload | Bulk IPv6 traffic (compressed) | Yes (routed) |
|
||||
| SessionDatagram + Data (with CP) | 77 + coords + IPv6 payload | Warmup/recovery (compressed) | Yes (routed) |
|
||||
| SessionDatagram + CoordsRequired | 70 bytes | Cache miss error | Yes (routed) |
|
||||
| SessionDatagram + PathBroken | 70+ bytes | Dead-end error | Yes (routed) |
|
||||
| Disconnect | 2 bytes | Link teardown | No (peer-to-peer) |
|
||||
|
||||
See [fips-wire-formats.md](fips-wire-formats.md) for byte-level layouts.
|
||||
For typical sizes, forwarding category, and the byte-level layouts
|
||||
of each FMP and FSP message type, see
|
||||
[../reference/wire-formats.md](../reference/wire-formats.md). The
|
||||
canonical Packet Type Summary table lives there.
|
||||
|
||||
## Privacy Considerations
|
||||
|
||||
@@ -696,18 +559,25 @@ recovery).
|
||||
| Flap dampening (hysteresis + hold-down) | **Implemented** |
|
||||
| Link liveness (dead timeout) | **Implemented** |
|
||||
| Discovery request deduplication | **Implemented** |
|
||||
| Discovery bloom-guided tree routing | **Implemented** |
|
||||
| Discovery retry logic | **Implemented** |
|
||||
| Discovery originator backoff | **Implemented** |
|
||||
| Discovery transit-side rate limiting | **Implemented** |
|
||||
| Discovery response-forwarded dedup | **Implemented** |
|
||||
| Leaf-only operation | Under development |
|
||||
| Link cost in parent selection (ETX) | **Implemented** |
|
||||
| Link cost in candidate ranking | **Implemented** |
|
||||
| Discovery path accumulation | Future direction |
|
||||
|
||||
## References
|
||||
|
||||
- [fips-intro.md](fips-intro.md) — Protocol overview
|
||||
- [fips-concepts.md](fips-concepts.md) — Protocol overview
|
||||
- [fips-architecture.md](fips-architecture.md) — Layer architecture and
|
||||
identity model
|
||||
- [fips-mesh-layer.md](fips-mesh-layer.md) — FMP specification
|
||||
- [fips-spanning-tree.md](fips-spanning-tree.md) — Tree algorithms and data
|
||||
structures
|
||||
- [fips-bloom-filters.md](fips-bloom-filters.md) — Filter parameters and math
|
||||
- [fips-wire-formats.md](fips-wire-formats.md) — Wire format reference
|
||||
- [../reference/wire-formats.md](../reference/wire-formats.md) — Wire
|
||||
format reference
|
||||
- [spanning-tree-dynamics.md](spanning-tree-dynamics.md) — Convergence
|
||||
walkthroughs
|
||||
|
||||
@@ -0,0 +1,218 @@
|
||||
# Metrics Measurement Protocol (MMP)
|
||||
|
||||
The Metrics Measurement Protocol provides per-link and per-session
|
||||
quality metrics — SRTT, loss, jitter, goodput, ETX, and one-way delay
|
||||
trend — using only counter and timestamp fields already present in
|
||||
the FMP and FSP wire formats. No additional probing traffic is
|
||||
required. The same algorithms and report message format are used at
|
||||
both layers; only the routing scope and configuration namespace
|
||||
differ.
|
||||
|
||||
This document is the canonical home for the MMP design. For the
|
||||
link-layer instance's role inside FMP, see
|
||||
[fips-mesh-layer.md](fips-mesh-layer.md). For the session-layer
|
||||
instance's role inside FSP, see
|
||||
[fips-session-layer.md](fips-session-layer.md). For the byte-level
|
||||
SenderReport and ReceiverReport layouts, see
|
||||
[../reference/wire-formats.md](../reference/wire-formats.md).
|
||||
|
||||
## Two Layers, One Protocol
|
||||
|
||||
MMP runs at two layers:
|
||||
|
||||
- **Link-layer MMP**: One instance per active FMP peer link. Reports
|
||||
are exchanged peer-to-peer between direct neighbors and measure the
|
||||
quality of that single hop.
|
||||
- **Session-layer MMP**: One instance per established FSP session.
|
||||
Reports are encrypted end-to-end and forwarded through every transit
|
||||
link, measuring end-to-end quality independent of hop count.
|
||||
|
||||
The algorithms (SRTT estimation, jitter computation, loss inference,
|
||||
ETX) are identical at both layers. The differences are configuration
|
||||
namespace, report intervals, and routing scope. See
|
||||
[Layer Differences](#layer-differences) below.
|
||||
|
||||
## Metrics Tracked
|
||||
|
||||
MMP computes the following metrics from the per-frame counter and
|
||||
timestamp fields:
|
||||
|
||||
- **SRTT** — Smoothed round-trip time (Jacobson/RFC 6298, α=1/8).
|
||||
Derived from timestamp-echo in ReceiverReports with dwell-time
|
||||
compensation.
|
||||
- **Loss rate** — Bidirectional loss inferred from counter gaps.
|
||||
Tracked as both instantaneous (per-interval) and long-term EWMA.
|
||||
- **Jitter** — Interarrival jitter (RFC 3550 algorithm) in
|
||||
microseconds.
|
||||
- **Goodput** — Bytes per second of payload data (excludes MMP
|
||||
reports).
|
||||
- **OWD trend** — One-way delay trend (µs/s, signed). Indicates
|
||||
congestion buildup before loss occurs.
|
||||
- **ETX** — Expected Transmission Count, computed from bidirectional
|
||||
delivery ratios. Used in cost-based parent selection via
|
||||
`link_cost = etx * (1.0 + srtt_ms / 100.0)`, and in bloom-filter
|
||||
candidate ranking inside `find_next_hop()` (the same `link_cost`
|
||||
is the primary key when choosing among bloom-filter peers, with
|
||||
tree distance as the tie-breaker).
|
||||
- **Dual EWMA trends** — Short-term (α=1/4) and long-term (α=1/32)
|
||||
trend indicators for both RTT and loss, enabling change detection.
|
||||
|
||||
Session-layer MMP additionally tracks the observed forward-path MTU;
|
||||
see [fips-mtu.md](fips-mtu.md) for the end-to-end path-MTU mechanism.
|
||||
|
||||
## Operating Modes
|
||||
|
||||
MMP supports three modes:
|
||||
|
||||
| Mode | Reports Exchanged | Metrics Available |
|
||||
| ---- | ----------------- | ----------------- |
|
||||
| **Full** (default) | SenderReport + ReceiverReport | All metrics including RTT, loss, jitter, goodput, OWD trend |
|
||||
| **Lightweight** | ReceiverReport only | Loss (from counter gaps), jitter, OWD trend. No RTT. |
|
||||
| **Minimal** | None | Spin bit and CE echo flags only. No computed metrics. |
|
||||
|
||||
The mode is configured per layer (`node.mmp.mode` and
|
||||
`node.session_mmp.mode`).
|
||||
|
||||
## Report Scheduling
|
||||
|
||||
Reports are sent at RTT-adaptive intervals computed as
|
||||
`clamp(2 × SRTT, low, high)`. A cold-start interval is used until SRTT
|
||||
has converged.
|
||||
|
||||
| Layer | Adaptive bounds | Cold-start |
|
||||
| ----- | --------------- | ---------- |
|
||||
| Link | `[1s, 5s]` | 200 ms (first 5 samples) |
|
||||
| Session | `[500ms, 10s]` | 1 s |
|
||||
|
||||
The session-layer bounds are higher because session reports are
|
||||
encrypted and forwarded through every transit link, so bandwidth cost
|
||||
is proportional to path length.
|
||||
|
||||
## Spin Bit and RTT
|
||||
|
||||
The SP (spin bit) flag in the FMP inner header follows the QUIC spin
|
||||
bit pattern: reflected on receive, toggled on send when the reflected
|
||||
value matches the last sent value. The spin bit state machine runs
|
||||
for TX reflection, but **RTT samples from the spin bit are
|
||||
discarded**. In a mesh protocol where frames are sent irregularly
|
||||
(tree announces, bloom filters, MMP reports on different timers),
|
||||
inter-frame processing delays inflate spin bit RTT measurements
|
||||
unpredictably. Timestamp-echo from ReceiverReports (with dwell-time
|
||||
compensation) is the sole SRTT source.
|
||||
|
||||
The spin bit lives in the link-layer FMP inner header, so this
|
||||
mechanism applies to link-layer MMP only. Session-layer MMP carries
|
||||
its spin bit in the FSP encrypted inner header but uses it the same
|
||||
way: reflected for diagnostic visibility, not used for SRTT.
|
||||
|
||||
## ECN Congestion Signaling
|
||||
|
||||
The CE (Congestion Experienced) flag (bit 1 in the FMP flags byte)
|
||||
provides hop-by-hop congestion signaling through the mesh. Transit
|
||||
nodes detect congestion on outgoing links and set CE on forwarded
|
||||
packets; once set, the flag stays set for all subsequent hops to the
|
||||
destination.
|
||||
|
||||
**Congestion detection** triggers on any of:
|
||||
|
||||
- Outgoing link MMP loss rate ≥ `node.ecn.loss_threshold` (default 5%)
|
||||
- Outgoing link MMP ETX ≥ `node.ecn.etx_threshold` (default 3.0)
|
||||
- Kernel receive buffer drops detected on any local transport (via
|
||||
`SO_RXQ_OVFL` on UDP)
|
||||
|
||||
**CE relay**: The forwarding path computes
|
||||
`outgoing_ce = incoming_ce || local_congestion`. Once CE is set on a
|
||||
packet, it remains set for the rest of the forward path.
|
||||
|
||||
**IPv6 ECN-CE marking**: When a CE-flagged DataPacket arrives at its
|
||||
final destination, the IPv6 Traffic Class ECN bits are marked CE
|
||||
(0b11) before TUN delivery — but only for ECN-capable packets (ECT(0)
|
||||
or ECT(1)). Not-ECT packets are never marked per RFC 3168. The host
|
||||
TCP stack then echoes ECE in ACKs, triggering sender cwnd reduction
|
||||
through standard congestion control.
|
||||
|
||||
**Session-layer tracking**: The `ecn_ce_count` field in MMP
|
||||
ReceiverReports tracks CE-flagged packets received per link, providing
|
||||
end-to-end visibility into congestion propagation.
|
||||
|
||||
ECN signaling is a link-layer mechanism. Session-layer MMP only
|
||||
observes the CE counter as part of the report stream; CE marking is
|
||||
not generated end-to-end. Tuning parameters live under `node.ecn.*`
|
||||
in [../reference/configuration.md](../reference/configuration.md).
|
||||
|
||||
## Send Failure Backoff (Session Layer Only)
|
||||
|
||||
When a session MMP report cannot be delivered (destination unreachable,
|
||||
no route), the sender applies exponential backoff to the probe
|
||||
interval — a standard distributed-systems pattern for transient
|
||||
failure handling:
|
||||
|
||||
- Each consecutive failure doubles the interval: 2x, 4x, 8x, 16x, 32x
|
||||
- Backoff caps at 32x the base interval (5 consecutive failures)
|
||||
- A successful send resets to the normal SRTT-based interval
|
||||
- Debug logging is suppressed after 3 consecutive failures; a summary
|
||||
is logged when the destination becomes reachable again
|
||||
|
||||
This prevents wasted CPU and log noise when a session's remote
|
||||
endpoint has departed the network but the local session has not yet
|
||||
timed out. Link-layer MMP has no equivalent — link-layer reports are
|
||||
peer-to-peer over an authenticated link, so delivery failure is
|
||||
indistinguishable from link death and the link-liveness mechanism
|
||||
takes over.
|
||||
|
||||
## Layer Differences
|
||||
|
||||
| Aspect | Link layer | Session layer |
|
||||
| ------ | ---------- | ------------- |
|
||||
| Routing scope | Peer-to-peer (one hop) | End-to-end (forwarded through every hop) |
|
||||
| Configuration namespace | `node.mmp.*` | `node.session_mmp.*` |
|
||||
| Report bounds | `[1s, 5s]` | `[500ms, 10s]` |
|
||||
| Cold-start interval | 200 ms (first 5 samples) | 1 s |
|
||||
| Bandwidth cost | One link | Proportional to path length |
|
||||
| Send-failure backoff | Not applicable | Yes |
|
||||
| Path-MTU echo | Not applicable | PathMtuNotification (see [fips-mtu.md](fips-mtu.md)) |
|
||||
| Idle-timeout interaction | None | Reports do **not** reset session idle timer |
|
||||
|
||||
## Idle Timeout Interaction (Session Layer Only)
|
||||
|
||||
MMP reports (SenderReport, ReceiverReport) and PathMtuNotification do
|
||||
**not** reset the session idle timer. Only application data
|
||||
(DataPacket, type 0x10) resets `last_activity`. This ensures sessions
|
||||
with no application traffic tear down after
|
||||
`node.session.idle_timeout_secs` (default 90s), while MMP continues
|
||||
providing measurement data up to the teardown moment.
|
||||
|
||||
## Operator Logging
|
||||
|
||||
Both layers emit periodic metrics at info level. The interval is
|
||||
`node.mmp.log_interval_secs` for link-layer (default 30s) and
|
||||
`node.session_mmp.log_interval_secs` for session-layer (default 30s).
|
||||
|
||||
Link-layer:
|
||||
|
||||
```text
|
||||
MMP link metrics peer=node-b rtt=2.3ms loss=0.2% jitter=0.1ms goodput=76.0MB/s tx_pkts=1234 rx_pkts=5678
|
||||
```
|
||||
|
||||
Session-layer:
|
||||
|
||||
```text
|
||||
MMP session metrics session=npub1tdwa...84le rtt=4.3ms loss=0.6% jitter=0.2ms goodput=71.3MB/s mtu=1472 tx_pkts=1234 rx_pkts=5678
|
||||
```
|
||||
|
||||
Teardown logs include final SRTT, loss rate, jitter, ETX, goodput,
|
||||
and cumulative tx/rx packet and byte counts.
|
||||
|
||||
## See also
|
||||
|
||||
- [fips-mesh-layer.md](fips-mesh-layer.md) — link-layer MMP integration
|
||||
inside FMP
|
||||
- [fips-session-layer.md](fips-session-layer.md) — session-layer MMP
|
||||
integration inside FSP
|
||||
- [fips-mtu.md](fips-mtu.md) — PathMtuNotification, the session-only
|
||||
end-to-end path-MTU echo
|
||||
- [../reference/wire-formats.md](../reference/wire-formats.md) —
|
||||
SenderReport (0x01 / 0x11) and ReceiverReport (0x02 / 0x12) byte
|
||||
layouts
|
||||
- [../reference/configuration.md](../reference/configuration.md) —
|
||||
full `node.mmp.*`, `node.session_mmp.*`, and `node.ecn.*` knob tables
|
||||
@@ -0,0 +1,316 @@
|
||||
# FIPS Path MTU and Encapsulation Overhead
|
||||
|
||||
MTU is a cross-cutting concern in FIPS. No single layer owns it: the
|
||||
transport reports per-link MTU, FMP propagates `path_mtu` along
|
||||
forward and reverse paths, FSP echoes the observed path MTU end-to-end
|
||||
back to the source, and the IPv6 adapter enforces the resulting
|
||||
effective MTU at the TUN interface. This document is the canonical
|
||||
home for the unified MTU model.
|
||||
|
||||
For operator-facing diagnostic recipes (interpreting `MtuExceeded`
|
||||
counters, tuning IPv6 application MSS, troubleshooting cold-flow
|
||||
oversize), see the relevant how-to under `docs/how-to/`.
|
||||
|
||||
## The MTU Problem in FIPS
|
||||
|
||||
A FIPS path can traverse heterogeneous link types — UDP/IP (1280
|
||||
default, IPv6 minimum), Ethernet (interface MTU − 3, typically 1497),
|
||||
BLE (negotiated ATT_MTU per link), Tor stream (1400 default), radio
|
||||
(51–222) — within a single end-to-end session.
|
||||
The minimum MTU along the path determines the largest datagram a
|
||||
session can deliver. Several properties make this harder than in
|
||||
classic IP networks:
|
||||
|
||||
- **No fragmentation.** FIPS does not fragment at transit nodes (see
|
||||
[No fragmentation policy](#no-fragmentation-policy)). A datagram
|
||||
that exceeds the next-hop link MTU is dropped, and the source is
|
||||
signaled.
|
||||
- **Forward/reverse path asymmetry.** After tree reconvergence the
|
||||
return path may diverge from the forward path, so the bottleneck
|
||||
on each direction can differ.
|
||||
- **First-flow race.** The very first SessionDatagram races
|
||||
destination discovery — the source has not yet learned the path MTU
|
||||
but must pick a payload size for the queued packet.
|
||||
- **Variable per-link MTU.** Some transports (BLE, TCP via
|
||||
`TCP_MAXSEG`) report different MTUs for different links rather than
|
||||
a single transport-wide value.
|
||||
|
||||
The unified MTU model below combines proactive and reactive
|
||||
mechanisms to converge on a working effective MTU within the first
|
||||
few packets of a session, then maintain it across topology changes.
|
||||
|
||||
## Encapsulation Overhead
|
||||
|
||||
The byte budget for a FIPS-encapsulated packet:
|
||||
|
||||
| Layer | Overhead | Purpose |
|
||||
| ----- | -------- | ------- |
|
||||
| Link encryption | 37 bytes | 16-byte outer header + 5-byte inner header (timestamp + msg_type) + 16-byte AEAD tag |
|
||||
| SessionDatagram body | 35 bytes | ttl + path_mtu + src_addr + dest_addr (msg_type counted in inner header) |
|
||||
| FSP header | 12 bytes | 4-byte prefix + 8-byte counter (used as AEAD AAD) |
|
||||
| FSP inner header | 6 bytes | 4-byte timestamp + 1-byte msg_type + 1-byte inner_flags (inside AEAD) |
|
||||
| Session AEAD tag | 16 bytes | ChaCha20-Poly1305 tag on session-encrypted payload |
|
||||
| **Protocol envelope** | **106 bytes** | `FIPS_OVERHEAD` constant — the base payload budget for any service |
|
||||
|
||||
`FIPS_OVERHEAD = 106` is the constant the rest of the system reasons
|
||||
about. Coordinate piggybacking via the CP flag adds variable extra
|
||||
overhead — `2 + entries × 16` bytes per coordinate, with both source
|
||||
and destination coordinates carried — and the send path skips the CP
|
||||
flag if adding coords would exceed the transport MTU.
|
||||
|
||||
Service-specific overheads layer on top of `FIPS_OVERHEAD`:
|
||||
|
||||
| Service | Overhead | Note |
|
||||
| ------- | -------- | ---- |
|
||||
| DataPacket port header | +4 bytes | Always present for port-multiplexed services |
|
||||
| IPv6 compression | −33 bytes | 40-byte IPv6 header → 7-byte format + residual |
|
||||
| **IPv6 effective overhead** | **77 bytes** | `FIPS_IPV6_OVERHEAD` constant |
|
||||
|
||||
See [fips-ipv6-adapter.md](fips-ipv6-adapter.md) for the IPv6
|
||||
compression scheme that lets the adapter reach `FIPS_IPV6_OVERHEAD`.
|
||||
|
||||
## Per-Link MTU Reporting
|
||||
|
||||
Each transport implements two MTU methods on its trait:
|
||||
|
||||
- `mtu() -> u16` — Transport-wide default MTU.
|
||||
- `link_mtu(addr: &TransportAddr) -> u16` — Per-link MTU for a
|
||||
specific remote address. The default implementation falls back to
|
||||
`mtu()`, so transports with uniform MTU (UDP, raw Ethernet) need
|
||||
not override it.
|
||||
|
||||
FMP uses `link_mtu()` when it needs to reason about a specific
|
||||
outbound link — typically for `path_mtu` annotation in
|
||||
SessionDatagram and LookupResponse. Per-transport defaults:
|
||||
|
||||
| Transport | Default MTU | Per-link MTU source |
|
||||
| --------- | ----------- | ------------------- |
|
||||
| UDP | 1280 (IPv6 minimum) | uniform (`mtu()` fallback) |
|
||||
| Ethernet | interface MTU − 3 (typically 1497) | uniform |
|
||||
| TCP | 1400 | derived from `TCP_MAXSEG` per connection |
|
||||
| Tor | 1400 | uniform |
|
||||
| BLE | 2048 default; negotiated ATT_MTU per link | per-link (overrides `mtu()`) |
|
||||
|
||||
For TCP, the per-connection `TCP_MAXSEG` query lets FMP discover the
|
||||
actual MSS the kernel negotiated for each connection, rather than
|
||||
assuming a single value across all TCP peers.
|
||||
|
||||
## Proactive PMTUD: SessionDatagram path_mtu
|
||||
|
||||
Every SessionDatagram and LookupResponse carries a 2-byte `path_mtu`
|
||||
field. The source initializes it to its outbound link MTU; each
|
||||
transit node applies `min(current, link_mtu(next_hop))` before
|
||||
forwarding. The destination receives the forward-path minimum.
|
||||
|
||||
For SessionDatagram, the receiver of the forward-path minimum is the
|
||||
session-layer destination, which then echoes the value back to the
|
||||
source via PathMtuNotification (see
|
||||
[End-to-end echo](#end-to-end-echo-pathmtunotification)).
|
||||
|
||||
For LookupResponse, the receiver is the original requester, and the
|
||||
annotation is reverse-path-only: the LookupResponse path is the
|
||||
return path of the lookup, so the annotated `path_mtu` reflects what
|
||||
the requester can use to reach the discovered destination over the
|
||||
discovered path.
|
||||
|
||||
Because the field is initialized by the source and mins as it travels,
|
||||
it converges to the bottleneck without any additional probing. The
|
||||
first SessionDatagram on a fresh session may carry an over-estimate
|
||||
(the source has not yet been told a smaller min), which is what makes
|
||||
the reactive MtuExceeded path necessary.
|
||||
|
||||
## Reactive PMTUD: MtuExceeded
|
||||
|
||||
When a transit node receives a SessionDatagram whose total wire size
|
||||
exceeds the next-hop `link_mtu`, it cannot forward without
|
||||
fragmentation. Instead:
|
||||
|
||||
1. The transit node generates a SessionDatagram addressed back to the
|
||||
source carrying an `MtuExceeded` payload (msg_type 0x22). The
|
||||
payload identifies the destination, the reporting router, and the
|
||||
bottleneck MTU.
|
||||
2. The error is routed via `find_next_hop(src_addr)`. If the source
|
||||
is also unreachable, the error is dropped silently (no cascading
|
||||
errors).
|
||||
3. The original oversized packet is dropped.
|
||||
|
||||
The source's FSP layer applies the reported bottleneck immediately —
|
||||
unlike the increase case (see hysteresis below), decrease is always
|
||||
take-the-lower-value because the original packet has already been
|
||||
dropped. The source can then reduce payload sizes on subsequent
|
||||
SessionDatagrams.
|
||||
|
||||
MtuExceeded is the reactive complement to the proactive `path_mtu`
|
||||
field. The proactive field tracks the minimum along the forward path
|
||||
under steady-state convergence; MtuExceeded handles the in-flight gap
|
||||
when an oversized packet hits a new bottleneck (forward path shifted,
|
||||
peer's outbound MTU dropped, BLE renegotiated) before the source has
|
||||
adapted.
|
||||
|
||||
Error generation is rate-limited at 100ms per destination at the
|
||||
transit node to prevent storms during topology changes.
|
||||
|
||||
## End-to-End Echo: PathMtuNotification
|
||||
|
||||
PathMtuNotification (msg_type 0x13, session-layer) provides
|
||||
end-to-end path MTU feedback, adapting RFC 1191 Path MTU Discovery
|
||||
for overlay networks — the transit-node `min()` propagation replaces
|
||||
ICMP Packet Too Big.
|
||||
|
||||
Mechanism:
|
||||
|
||||
1. The source sets `path_mtu` in each SessionDatagram envelope to its
|
||||
outbound link MTU.
|
||||
2. Each transit node applies `min(current, transport.link_mtu(addr))`
|
||||
before forwarding.
|
||||
3. The destination receives the forward-path minimum and sends a
|
||||
PathMtuNotification (2-byte body: `u16 LE path_mtu`) back to the
|
||||
source.
|
||||
4. The source applies the notification with hysteresis:
|
||||
- **Decrease**: immediate (take lower value).
|
||||
- **Increase**: requires 3 consecutive higher-value notifications
|
||||
spanning at least 2 × notification interval.
|
||||
5. Notifications are sent on first measurement, on any decrease, and
|
||||
periodically at `max(10s, 5 × SRTT)`.
|
||||
|
||||
The hysteresis on increase prevents oscillation when the path MTU
|
||||
fluctuates around a boundary; the immediate decrease prevents
|
||||
delivering oversized packets after a path has narrowed.
|
||||
|
||||
PathMtuNotification is wrapped in a session-layer encrypted message
|
||||
and travels back to the source via the session's normal forwarding
|
||||
path. It is part of the session-layer MMP report stream's traffic
|
||||
budget and (along with SenderReport and ReceiverReport) does not
|
||||
reset the session idle timer.
|
||||
|
||||
## Per-Destination MTU Storage
|
||||
|
||||
Two storage locations track per-destination MTU, serving different
|
||||
consumers:
|
||||
|
||||
- **Session-canonical** (`MmpSessionState.path_mtu`, type
|
||||
`PathMtuState`). Holds the running end-to-end path MTU for an
|
||||
established FSP session. Updated by both `PathMtuNotification`
|
||||
(proactive, end-to-end echo) and reactive `MtuExceeded` from
|
||||
transit routers. Read by the session layer when constructing
|
||||
outbound `SessionDatagram` envelopes.
|
||||
|
||||
- **TCP-clamp mirror** (`path_mtu_lookup`, a
|
||||
`HashMap<FipsAddress, u16>` on the Node). Read by the
|
||||
TUN-side TCP MSS clamp (`per_flow_max_mss` in
|
||||
`src/upper/tun.rs`) at first-SYN time so outbound TCP flows
|
||||
are clamped to the per-destination MTU rather than a generic
|
||||
ceiling. Written from four sites, all using tighter-only
|
||||
semantics — the clamp is never loosened:
|
||||
- Discovery's `LookupResponse` handler — reverse-path
|
||||
annotated value carried back by the discovery target.
|
||||
- `seed_path_mtu_for_link_peer` when a peer is promoted to
|
||||
an active link, seeding with the new link's `link_mtu`
|
||||
so traffic to that peer immediately uses the per-link
|
||||
value rather than a generic default.
|
||||
- The reactive `MtuExceeded` handler, mirroring the
|
||||
bottleneck reported by a transit router.
|
||||
- The proactive `PathMtuNotification` handler, mirroring
|
||||
the new effective end-to-end value so a fresh TCP flow
|
||||
benefits immediately from PMTU knowledge the session has
|
||||
already acquired.
|
||||
|
||||
All four writers apply the same tighter-only rule, so the mirror
|
||||
converges to the smallest MTU any signal has reported for that
|
||||
destination and a subsequent looser observation cannot widen it.
|
||||
|
||||
## TCP MSS Clamping
|
||||
|
||||
The IPv6 adapter intercepts TCP SYN and SYN-ACK packets at the TUN
|
||||
interface and clamps the Maximum Segment Size (MSS) option to:
|
||||
|
||||
```text
|
||||
clamped_mss = effective_ipv6_mtu - 40 (IPv6 header) - 20 (TCP header)
|
||||
```
|
||||
|
||||
Clamping is applied in two places:
|
||||
|
||||
- **TUN reader** (outbound): clamps MSS on outbound SYN packets
|
||||
- **TUN writer** (inbound): clamps MSS on inbound SYN-ACK packets
|
||||
|
||||
Together these ensure both directions of a TCP connection use
|
||||
appropriately-sized segments from the start, avoiding the initial
|
||||
oversized-packet loss that would occur if the adapter relied on ICMP
|
||||
Packet Too Big alone.
|
||||
|
||||
Clamping is **conditional**: when `per_flow_max_mss` already has an
|
||||
entry for the flow, that entry is used; otherwise the clamp falls
|
||||
back to a ceiling derived from the most pessimistic effective IPv6
|
||||
MTU the adapter knows about (1143 with the typical 1280 transport
|
||||
floor). The fallback handles cold-flow first-SYN traffic — the very
|
||||
first SYN of a flow may arrive before the MMP path-MTU echo and any
|
||||
per-flow lookup has been populated, so the conservative ceiling
|
||||
prevents the SYN-ACK chain from negotiating a too-large MSS that
|
||||
would later drop.
|
||||
|
||||
The adapter integrates with the MTU subsystem rather than owning it.
|
||||
The "why we clamp and what `max_mss` means" lives here in the MTU
|
||||
design; the "how the clamp is implemented at the TUN" lives in the
|
||||
[IPv6 adapter](fips-ipv6-adapter.md#tcp-mss-clamping) doc.
|
||||
|
||||
## ICMP Packet Too Big
|
||||
|
||||
When an outbound packet at the TUN exceeds the effective IPv6 MTU,
|
||||
the adapter generates an ICMPv6 Packet Too Big message and delivers
|
||||
it back to the application via the TUN. This triggers the kernel's
|
||||
Path MTU Discovery mechanism for non-TCP traffic and for any TCP flow
|
||||
where MSS clamping was insufficient.
|
||||
|
||||
ICMPv6 Packet Too Big generation is rate-limited per source address
|
||||
(100ms interval) to prevent storms from applications sending many
|
||||
oversized packets. The ICMP response is delivered locally back
|
||||
through the TUN; no network traversal is needed, so delivery is
|
||||
reliable.
|
||||
|
||||
## No Fragmentation Policy
|
||||
|
||||
FIPS does not perform fragmentation at transit nodes:
|
||||
|
||||
- **Why no transit fragmentation.** Session-layer encryption is
|
||||
end-to-end — the AEAD tag authenticates the entire plaintext.
|
||||
Fragmenting an encrypted SessionDatagram would require either
|
||||
exposing plaintext structure to transit nodes (unacceptable) or
|
||||
reassembling before decryption (opens an attack surface — a transit
|
||||
node could replay or withhold fragments to influence reassembly).
|
||||
- **Why no source-side fragmentation.** The source doesn't need
|
||||
fragmentation because the proactive `path_mtu` field plus the
|
||||
reactive MtuExceeded signal converge on a working size within the
|
||||
first few packets. Applications that need oversized payloads run
|
||||
TCP over the IPv6 adapter, which has its own segmentation under
|
||||
MSS clamping.
|
||||
|
||||
Some transports may perform fragmentation and reassembly internally
|
||||
(e.g., BLE L2CAP) and can advertise a larger virtual MTU than the
|
||||
physical medium supports — this is transparent to FIPS.
|
||||
|
||||
## Operational Considerations
|
||||
|
||||
Diagnosing MTU-related symptoms (handshakes succeed but bulk
|
||||
transfers stall, ssh hangs after `Welcome` banner, sporadic
|
||||
`MtuExceeded` spikes during topology changes) requires inspecting
|
||||
per-link MTU, per-session MTU, and the per-destination
|
||||
`path_mtu_lookup` table. See
|
||||
[../how-to/diagnose-mtu-issues.md](../how-to/diagnose-mtu-issues.md)
|
||||
for the operator recipes. The relevant control-socket queries are
|
||||
`fipsctl show sessions` (per-session MTU), `fipsctl show transports`
|
||||
(per-link MTU), and `fipsctl show identity-cache` (with adapter MTU
|
||||
context).
|
||||
|
||||
## See also
|
||||
|
||||
- [fips-transport-layer.md](fips-transport-layer.md) — the `mtu()` /
|
||||
`link_mtu()` trait surface and per-transport defaults
|
||||
- [fips-mesh-layer.md](fips-mesh-layer.md) — SessionDatagram and the
|
||||
MtuExceeded error signal
|
||||
- [fips-session-layer.md](fips-session-layer.md) — session-layer
|
||||
PathMtuNotification echo, applied with hysteresis
|
||||
- [fips-ipv6-adapter.md](fips-ipv6-adapter.md) — TUN-side ICMPv6 PTB
|
||||
generation, MSS clamping integration, IPv6-specific overhead table
|
||||
- [../reference/wire-formats.md](../reference/wire-formats.md) —
|
||||
SessionDatagram, LookupResponse, MtuExceeded, PathMtuNotification
|
||||
byte layouts
|
||||
@@ -0,0 +1,407 @@
|
||||
# FIPS Nostr-Mediated Discovery and NAT Traversal
|
||||
|
||||
Nostr-mediated discovery lets FIPS nodes find each other, and if
|
||||
necessary, punch through UDP NAT, using public Nostr relays as the
|
||||
signaling channel. A node publishes its reachable transport endpoints to
|
||||
a small set of relays under its own Nostr identity (which is also its
|
||||
FIPS identity), and peers resolve those endpoints at dial time by npub.
|
||||
For peers behind UDP NAT, the same relay channel carries an encrypted
|
||||
offer/answer exchange, and STUN supplies the reflexive address used for
|
||||
a coordinated hole-punch.
|
||||
|
||||
Nostr discovery is unconditionally compiled into the `fips` binary on
|
||||
every supported platform and ships in every stock packaging artifact
|
||||
(`.deb`, AUR, systemd tarball, OpenWrt `.ipk`, macOS `.pkg`, Windows
|
||||
`.zip`). It is runtime-opt-in: the YAML configuration defaults to
|
||||
disabled (`node.discovery.nostr.enabled: false`), so the discovery
|
||||
runtime stays dormant — and opens no relay connections — until an
|
||||
operator flips the flag. Default relay and STUN-server lists ship in
|
||||
the config; both are optional overrides. When disabled, nodes behave
|
||||
exactly as before: only the static `peers[]` addresses are used.
|
||||
|
||||
## Role
|
||||
|
||||
The feature adds three capabilities on top of FIPS's static peer model:
|
||||
|
||||
- **Advertising.** A node publishes the transport endpoints it wants
|
||||
peers to use (direct UDP, direct TCP, a Tor onion, or the special
|
||||
`udp:nat` rendezvous token) as a signed Nostr event. The advert is
|
||||
anchored to the node's FIPS identity key — a peer that knows the npub
|
||||
knows the advert is authentic.
|
||||
- **Lookup.** When dialing a configured peer marked `via_nostr`, or any
|
||||
peer in `policy: open` mode, the node fetches that peer's advert from
|
||||
the configured relays and appends the advertised endpoints to its
|
||||
dial list. Static addresses are always tried first.
|
||||
- **UDP NAT hole-punch.** When both sides of a connection have UDP NAT
|
||||
endpoints, the advert carries enough information to run a STUN-based
|
||||
offer/answer exchange over encrypted ([NIP-59](https://github.com/nostr-protocol/nips/blob/master/59.md))
|
||||
Nostr events. Each side observes its reflexive address via STUN,
|
||||
exchanges candidate pairs through the relay, and both sides send UDP
|
||||
probes at a shared punch time. On the first successful probe, the
|
||||
punch socket is handed to FMP and becomes a normal UDP transport.
|
||||
|
||||
## When to use it
|
||||
|
||||
- **You run a public node** and want peers who know your npub to reach
|
||||
you without you distributing an address list out-of-band.
|
||||
- **You want to reach a peer behind UDP NAT** without deploying a relay
|
||||
or running Tor on both sides. The peer advertises `udp:nat` and you
|
||||
dial by npub.
|
||||
- **You want zero-touch peer discovery** within a known application
|
||||
namespace (`policy: open`), subject to an admission budget.
|
||||
- **You want to advertise a Tor onion** so peers don't need to know the
|
||||
`.onion` address out-of-band.
|
||||
|
||||
Skip the feature when every peer is already reachable through a stable
|
||||
static address (a LAN mesh, a pre-configured test bed, or a deployment
|
||||
where operators distribute `peers[]` blocks directly). The feature adds
|
||||
relay dependencies, STUN round-trips for NAT cases, and a small ambient
|
||||
background of relay traffic; none of that is useful when you already
|
||||
know where peers are.
|
||||
|
||||
## Scenarios and configuration
|
||||
|
||||
For end-to-end operator recipes — each of the five activation scenarios
|
||||
(advertise a directly-reachable UDP node, advertise a Tor onion node,
|
||||
look up a configured peer by npub without advertising, NAT hole-punch
|
||||
between two configured peers, and open discovery within an `app`
|
||||
namespace) — see
|
||||
[../how-to/enable-nostr-discovery.md](../how-to/enable-nostr-discovery.md).
|
||||
The full configuration knob tables, per-transport keys, and startup
|
||||
validation rules live in
|
||||
[../reference/configuration.md](../reference/configuration.md) under
|
||||
`node.discovery.nostr.*`. The Kind 37195 advert event format is in
|
||||
[../reference/nostr-events.md](../reference/nostr-events.md). The rest
|
||||
of this document covers the design of the discovery runtime itself.
|
||||
|
||||
## Under the covers
|
||||
|
||||
The rest of this document describes how the feature works inside the
|
||||
node. For the generic protocol shape (event tags, NIP usage, on-the-
|
||||
wire offer/answer schema, failure-suppression machinery), see
|
||||
[port-advertisement-and-nat-traversal.md](port-advertisement-and-nat-traversal.md).
|
||||
|
||||
### Overview
|
||||
|
||||
The discovery runtime is a background task group started during node
|
||||
initialization when `nostr.enabled` is true. It maintains a single
|
||||
`nostr-sdk` client connected to the union of `advert_relays` and
|
||||
`dm_relays`, and runs four loops: advert publication, advert
|
||||
subscription (for open discovery and cache warming), DM subscription
|
||||
(for incoming offers and answers), and a periodic advert-cache prune.
|
||||
Discovery has no CLI surface; all operations are driven by the
|
||||
configuration and by connection attempts made by the rest of the node.
|
||||
|
||||
```text
|
||||
+-----------------------+
|
||||
| Discovery runtime |
|
||||
+-----------------------+
|
||||
| | |
|
||||
advert publish | | DM sub (offers, answers)
|
||||
| |
|
||||
v v
|
||||
+-------------------------+
|
||||
| Nostr relay pool | (advert_relays ∪ dm_relays)
|
||||
+-------------------------+
|
||||
^ ^
|
||||
advert fetch/cache | | encrypted signaling
|
||||
| |
|
||||
+----------------+ | | +--------------------+
|
||||
| connect_peer |--+ +->| offer / answer |
|
||||
| (node side) | | handler |
|
||||
+----------------+ +--------------------+
|
||||
| |
|
||||
v v
|
||||
+---------+ +--------------+
|
||||
| STUN |<-- same socket --->| UDP punch |
|
||||
+---------+ +--------------+
|
||||
|
|
||||
v
|
||||
adopt_established_traversal()
|
||||
|
|
||||
v
|
||||
FMP IK handshake
|
||||
on adopted socket
|
||||
```
|
||||
|
||||
### Phase 1 — Advertisement
|
||||
|
||||
Adverts are published as Nostr kind `37195` parameterized replaceable
|
||||
events (FIPS-specific, in the application-defined replaceable range
|
||||
`30000–39999`; the digits visually spell `FIPS` — 7=F, 1=I, 9=P, 5=S).
|
||||
The `d` tag is hardcoded to the wire-format identifier
|
||||
`fips-overlay-v1` (or `fips-overlay-v1-next` on the `next` branch),
|
||||
so each node has a single, in-place-updatable advert under its
|
||||
identity. The configurable `app` value populates a separate
|
||||
`protocol` tag, which scopes adverts within a relay set without
|
||||
splitting them across multiple `d`-tag streams. The event is signed
|
||||
with the node's FIPS identity key; there is no separate Nostr key. A
|
||||
NIP-40 `expiration` tag is set to now + `advert_ttl_secs`, and a
|
||||
`version` tag carries the protocol version. The advert content is a
|
||||
JSON document shaped as `OverlayAdvert` (see
|
||||
[../reference/nostr-events.md](../reference/nostr-events.md) for the
|
||||
schema).
|
||||
|
||||
Publication happens on startup, again whenever the set of advertised
|
||||
endpoints changes (for example, when a Tor onion hostname first
|
||||
becomes available), and on a refresh timer every `advert_refresh_secs`.
|
||||
If the `advertise` flag is turned off, the previous advert event is
|
||||
deleted using a NIP-9 kind 5 delete event. Advert publication is
|
||||
fan-out: the same event is sent to every relay in `advert_relays` with
|
||||
no explicit failover — relay redundancy is implicit.
|
||||
|
||||
For a UDP or TCP transport with `public: true`, the address advertised
|
||||
follows a fixed precedence: an operator-supplied `external_addr` wins;
|
||||
otherwise a non-wildcard bound `local_addr` is used directly;
|
||||
otherwise — only for UDP — the runtime asks `stun_servers` for the
|
||||
reflexive address of the bound socket and advertises that. TCP has no
|
||||
STUN equivalent, so wildcard-bound TCP without `external_addr`
|
||||
produces a loud WARN and the endpoint is omitted from the advert.
|
||||
|
||||
### Phase 2 — Lookup
|
||||
|
||||
When the node decides to dial a peer that is eligible for Nostr
|
||||
resolution (a `via_nostr` peer, or any peer under `policy: open`), it
|
||||
issues a Nostr REQ filtered by `author = peer_pubkey`, `kind = 37195`,
|
||||
`#d = fips-overlay-v1`. The fetch is time-bounded (~2 s) and runs
|
||||
against all configured `advert_relays` in parallel. The first valid
|
||||
advert wins; adverts whose `protocol` tag does not match the local
|
||||
`app` value are rejected at validation.
|
||||
|
||||
Results are kept in an in-memory cache keyed by author npub. Cache
|
||||
entries carry the advert's expiration time; a periodic prune drops
|
||||
expired entries, and an LRU-by-expiry eviction enforces
|
||||
`advert_cache_max_entries`. A parallel long-lived subscription on the
|
||||
advert relays populates the cache passively, so open-discovery
|
||||
candidates do not require per-dial fetches.
|
||||
|
||||
On cache hit, advert endpoints are appended to the peer's static
|
||||
address list with lower priority; the static list is tried first.
|
||||
|
||||
### Phase 3 — Offer/Answer signaling
|
||||
|
||||
For any endpoint shaped as `udp:nat`, dialing triggers an
|
||||
offer/answer exchange before the first packet is sent. Signaling events
|
||||
are Nostr kind `21059` (ephemeral, not stored by conforming relays),
|
||||
gift-wrapped per [NIP-59](https://github.com/nostr-protocol/nips/blob/master/59.md)
|
||||
and encrypted with [NIP-44](https://github.com/nostr-protocol/nips/blob/master/44.md),
|
||||
so only the intended recipient can decrypt the payload.
|
||||
|
||||
The initiator performs STUN first (see Phase 4), then builds a
|
||||
`TraversalOffer` containing:
|
||||
|
||||
- A unique `sessionId` and a random `nonce` (used to correlate the
|
||||
answer).
|
||||
- Its reflexive address (if STUN succeeded).
|
||||
- Its list of local (private) addresses for same-LAN paths.
|
||||
- The STUN server it used, for informational reporting only.
|
||||
- An `expiresAt` equal to now + `signal_ttl_secs`.
|
||||
|
||||
The offer is sealed to the recipient's npub and published to the peer's
|
||||
preferred signaling relays — the node first tries to resolve the peer's
|
||||
NIP-17 DM relay list (kind 10050), and falls back to `dm_relays` if
|
||||
the inbox-relays fetch fails. Each side also publishes its own inbox
|
||||
relay list on startup so dialers can discover it.
|
||||
|
||||
On the receiving side, an inbound semaphore bounds concurrent offer
|
||||
processing at `max_concurrent_incoming_offers`. When the semaphore is
|
||||
full, the offer is dropped with a warn log; this is the primary guard
|
||||
against offer-spam from a misbehaving or compromised relay. A
|
||||
`sessionId` replay cache (bounded by `seen_sessions_max_entries`, with
|
||||
entries valid for `replay_window_secs`) rejects duplicates.
|
||||
|
||||
The responder runs its own STUN query and replies with a
|
||||
`TraversalAnswer` carrying its reflexive and local addresses plus a
|
||||
`PunchHint { startAtMs, intervalMs, durationMs }` that tells both sides
|
||||
when to begin probing and how aggressively. If the responder has no
|
||||
usable addresses at all, it replies with `accepted: false` and a
|
||||
`reason` string.
|
||||
|
||||
### Phase 4 — UDP hole-punch
|
||||
|
||||
Each side runs STUN (parsing XOR-MAPPED-ADDRESS from the response, all
|
||||
other attributes ignored) on the *same* UDP socket it will later use
|
||||
for punching and for the adopted FMP transport. This is critical: NAT
|
||||
state is per-socket, so the punch has to reuse the socket that taught
|
||||
the NAT about this binding.
|
||||
|
||||
Given its own reflexive + local addresses and the peer's, each side
|
||||
builds a candidate-pair plan that tries, in priority order:
|
||||
|
||||
1. **Reflexive ↔ reflexive.** The classic STUN path. Tried first because
|
||||
it is the only candidate that's reliable across arbitrary network
|
||||
topologies — host candidates from one peer that happen to be
|
||||
reachable from the other (via a corporate VPN, a Tailscale subnet
|
||||
route, or overlapping private address space) will succeed at the
|
||||
socket layer in the punch but fail in the FMP handshake when the
|
||||
return path doesn't match.
|
||||
2. **LAN ↔ LAN.** If both sides share a /24 prefix, same-subnet private
|
||||
addresses are likely reachable directly. Only fires when both peers
|
||||
shared local host candidates (which requires `share_local_candidates`
|
||||
to be enabled — off by default).
|
||||
3. **Mixed.** Reflexive on one side, local on the other — catches
|
||||
hairpin and one-side-public scenarios.
|
||||
|
||||
At `startAtMs` both sides begin sending 24-byte probe packets on the
|
||||
candidate pair(s) at `intervalMs` cadence for up to `durationMs`. A
|
||||
probe carries a 4-byte magic (`NPTC`), a 4-byte sequence, and the
|
||||
first 16 bytes of `SHA256(sessionId)`; both sides can compute the same
|
||||
session hash independently from the public `sessionId`, so no shared
|
||||
secret is needed on the punch path itself. On receiving a valid probe,
|
||||
a side replies with an `NPTA` ack. The first valid probe or ack seen
|
||||
from the far side records the working remote address and completes the
|
||||
attempt.
|
||||
|
||||
On timeout (`attempt_timeout_secs` as overall bound,
|
||||
`punch_duration_ms` as probe window), both sides issue NIP-9 deletes
|
||||
for their offer and answer events and report failure up to the
|
||||
discovery runtime's `BootstrapEvent::Failed` channel.
|
||||
|
||||
### Phase 5 — Adoption
|
||||
|
||||
On success, the discovery runtime emits `BootstrapEvent::Established`
|
||||
carrying the session id, the punch socket, and the learned remote
|
||||
address. `adopt_established_traversal()` in the node lifecycle takes
|
||||
the socket, registers it with the UDP transport layer as a new
|
||||
transport instance, and calls `initiate_connection()` with the peer's
|
||||
FIPS identity as the expected remote. FMP's Noise IK handshake runs on
|
||||
the same socket — there is no "promote link" step between punch and
|
||||
handshake; the punch socket *is* the FMP socket.
|
||||
|
||||
From that moment on, the connection is a normal FMP link and is
|
||||
subject to the usual liveness (MMP heartbeats), rekey, and removal
|
||||
behavior. A link-dead event does not re-enter the discovery runtime
|
||||
automatically; reconnection relies on `auto_reconnect` and the same
|
||||
dial path that triggered the original punch.
|
||||
|
||||
### Auto-connect semantics
|
||||
|
||||
Discovery does not itself initiate connections. It only supplies
|
||||
addresses. Dial attempts originate from the existing peer-connection
|
||||
machinery:
|
||||
|
||||
- **Configured peers** (`peers[]` with `connect_policy: auto_connect`)
|
||||
are dialed on startup and on retry. When `via_nostr` is set, advert
|
||||
endpoints are appended to the dial list with lower priority than
|
||||
static entries.
|
||||
- **Open discovery peers** are assembled from the advert cache, fenced
|
||||
by the peer ACL, and enqueued into a bounded retry queue sized by
|
||||
`open_discovery_max_pending`. There is no event-driven
|
||||
"connect on every advert" — a peer re-enters the queue only when its
|
||||
prior attempt has drained.
|
||||
- **Manual dials** (`fipsctl connect`) can target any configured peer
|
||||
and use the same dial path, including Nostr resolution if configured.
|
||||
|
||||
### Rate limits and safeguards
|
||||
|
||||
| Mechanism | Default | What it prevents | Behavior at limit |
|
||||
| --- | --- | --- | --- |
|
||||
| Offer semaphore (`max_concurrent_incoming_offers`) | 16 | CPU and memory exhaustion from offer spam on DM relays. | Warn log, offer dropped. |
|
||||
| Advert cache (`advert_cache_max_entries`) | 2048 | Memory growth from ambient advert traffic under `policy: open`. | LRU-by-expiry eviction. |
|
||||
| Seen-sessions (`seen_sessions_max_entries`) | 2048 | Replay of stale `sessionId` values. | Oldest entry evicted. |
|
||||
| Signal TTL (`signal_ttl_secs`) | 120 s | Indefinite in-flight offers on relays. | Expired offers rejected at validation. |
|
||||
| Open discovery queue (`open_discovery_max_pending`) | 64 | Unbounded retry queue under ambient advert load. | New candidates skipped until the queue drains. |
|
||||
| Punch window (`punch_duration_ms`) | 10 s | Endless probe traffic after one side has given up. | Attempt declared failed; sockets discarded. |
|
||||
| Failure-streak threshold (`failure_streak_threshold`) | 5 | Repeated traversal attempts against a peer that keeps failing. | Peer enters extended cooldown. |
|
||||
| Extended cooldown (`extended_cooldown_secs`) | 1800 s | Tight retry loops after a failure streak. | Per-peer suppression for the cooldown window. |
|
||||
| WARN log throttle (`warn_log_interval_secs`) | 300 s | Log floods from a peer that fails on every attempt. | One WARN per peer per interval; the rest demote to debug. |
|
||||
| Failure-state cap (`failure_state_max_entries`) | 4096 | Memory growth from per-peer failure tracking. | LRU eviction. |
|
||||
|
||||
The load-shedding mechanisms (`max_concurrent_incoming_offers` and the
|
||||
failure-streak / extended-cooldown pair) are deliberately conservative
|
||||
so that a misbehaving relay cannot flood the node with offers and a
|
||||
chronically unreachable peer cannot keep the traversal pipeline
|
||||
saturated. The remaining rows are capacity bounds.
|
||||
|
||||
Adverts also undergo a stale-advert sweep: cached entries whose
|
||||
`expiresAt` has passed are evicted on the periodic prune tick. Inbound
|
||||
signaling tolerates ±60 s of clock skew between sender and receiver,
|
||||
and the runtime maintains an NTP-style skew estimate per remote so
|
||||
that consistently-skewed relays don't trip the freshness check.
|
||||
|
||||
### Relay model
|
||||
|
||||
All configured relays (advert + DM) are opened on a single
|
||||
`nostr-sdk::Client` at startup. Publication is fan-out: the same event
|
||||
is sent to every relay in the target list, with no explicit retry or
|
||||
relay selection. Redundancy is implicit — a downed relay simply means
|
||||
its copy of the advert or signal is unavailable, while other relays
|
||||
still serve the same data.
|
||||
|
||||
For signaling specifically, the node prefers the recipient's NIP-17
|
||||
DM relays when available (the recipient publishes its DM relay list as
|
||||
a kind 10050 event to its own DM relays on startup) and falls back to
|
||||
the local `dm_relays` list otherwise. This keeps the common case
|
||||
off the sender's DM relays when those are different from the
|
||||
recipient's, at the cost of one extra NIP-17 fetch per offer.
|
||||
|
||||
There is no per-relay rate limiting or health check. The relay model
|
||||
assumes that an operator chooses relays they trust to be best-effort
|
||||
available and that outright misbehavior is handled at the offer
|
||||
semaphore and replay-cache layers downstream.
|
||||
|
||||
## Security and threat model
|
||||
|
||||
- **Relay operators can observe metadata.** They see which npubs
|
||||
publish adverts, to whom offers are sent, and the timing of that
|
||||
traffic. The *contents* of offer and answer events are
|
||||
NIP-59/NIP-44 sealed — only the intended recipient decrypts them.
|
||||
Adverts are public by design.
|
||||
- **STUN servers see the node's public IP and port.** Only the STUN
|
||||
servers listed in the node's own `stun_servers` are ever contacted
|
||||
for reflexive discovery. Peer-advertised STUN values are
|
||||
informational; a malicious peer cannot steer this node to a
|
||||
chosen STUN target. See the doc comment on
|
||||
`node.discovery.nostr.stun_servers`.
|
||||
- **The FIPS identity key signs adverts.** Compromise of
|
||||
`fips.key` is compromise of the node's Nostr identity — an attacker
|
||||
can publish adverts on behalf of the node. The recovery path is
|
||||
the same as for any identity compromise: rotate the key and
|
||||
re-advertise. There is no separate Nostr keypair to rotate
|
||||
independently.
|
||||
- **Tor advertising leaks timing via clearnet relays.** When a
|
||||
Tor-only node advertises its onion address, the advert itself is
|
||||
published on clearnet WebSocket relays. Operators who want full
|
||||
unlinkability between the advertising identity and the node's
|
||||
IP must route relay traffic through Tor as well — for example by
|
||||
running `fips` inside a network namespace with a Tor SOCKS
|
||||
proxy as its only egress, or by pointing `advert_relays` and
|
||||
`dm_relays` at onion relay endpoints.
|
||||
- **Open discovery accepts anyone publishing on the same `app`.**
|
||||
Admission control is the peer ACL, not the discovery layer. Verify
|
||||
the ACL before enabling `policy: open`, and consider using a
|
||||
non-default `app` value to scope visibility.
|
||||
- **Nothing about discovery bypasses FMP.** A successful punch yields
|
||||
a UDP socket with a claimed remote identity. That identity is not
|
||||
trusted until FMP's Noise IK handshake completes. A peer whose
|
||||
advert says "I am npub X at 1.2.3.4:5678" but whose FMP handshake
|
||||
presents a different static key is rejected at the mesh layer.
|
||||
|
||||
## See also
|
||||
|
||||
- [../how-to/enable-nostr-discovery.md](../how-to/enable-nostr-discovery.md)
|
||||
— operator activation recipes grouped under three capabilities
|
||||
(resolve, advertise, open) across five scenarios.
|
||||
- [../tutorials/resolve-peers-via-nostr.md](../tutorials/resolve-peers-via-nostr.md),
|
||||
[../tutorials/advertise-your-node.md](../tutorials/advertise-your-node.md),
|
||||
and [../tutorials/open-discovery.md](../tutorials/open-discovery.md)
|
||||
— hand-held walkthroughs of the three capabilities, in
|
||||
pedagogical order.
|
||||
- [../reference/configuration.md](../reference/configuration.md) — full
|
||||
configuration reference, including all surrounding keys elided from
|
||||
the scenarios above.
|
||||
- [../reference/nostr-events.md](../reference/nostr-events.md) — Kind
|
||||
37195 (overlay advert), Kind 21059 (gift-wrapped traversal
|
||||
signaling), Kind 10050 (NIP-17 inbox relay list).
|
||||
- [../reference/security.md](../reference/security.md) — consolidated
|
||||
security reference, including how the FIPS identity key signs both
|
||||
adverts and Noise handshakes.
|
||||
- [fips-transport-layer.md](fips-transport-layer.md) — UDP, TCP, and
|
||||
Tor transport mechanics; the punch socket is adopted as a normal
|
||||
UDP transport after handoff.
|
||||
- [fips-mesh-layer.md](fips-mesh-layer.md) — FMP Noise IK handshake
|
||||
that runs on the adopted socket.
|
||||
- [port-advertisement-and-nat-traversal.md](port-advertisement-and-nat-traversal.md)
|
||||
— generic protocol reference (event tags, NIP usage, on-the-wire
|
||||
offer/answer schema, failure-suppression machinery), with the
|
||||
FIPS-specific values called out as worked examples.
|
||||
@@ -0,0 +1,341 @@
|
||||
# FIPS Prior Work and References
|
||||
|
||||
FIPS builds on proven designs rather than inventing new cryptography or
|
||||
routing algorithms. Nearly every major design decision has deployed
|
||||
precedent. This document collects the relevant prior art, organized by
|
||||
the FIPS subsystem that draws on it, and gathers the academic and
|
||||
standards references cited from the per-subsystem design docs.
|
||||
|
||||
## Spanning Tree Self-Organization
|
||||
|
||||
The idea that distributed nodes can build a spanning tree through
|
||||
purely local decisions — each node selecting a parent based on
|
||||
announcements from its neighbors — dates to the
|
||||
[IEEE 802.1D Spanning Tree Protocol](https://en.wikipedia.org/wiki/Spanning_Tree_Protocol)
|
||||
(STP, 1985). STP demonstrated that a network-wide tree emerges from a
|
||||
simple deterministic rule (lowest bridge ID wins root election)
|
||||
applied independently at each node. FIPS uses the same principle —
|
||||
lowest node address determines the root — adapted from an Ethernet
|
||||
bridging context to a general-purpose overlay mesh.
|
||||
|
||||
## Tree Coordinate Routing
|
||||
|
||||
The spanning tree coordinates, bloom filter candidate selection, and
|
||||
greedy routing algorithms are adapted from
|
||||
[Yggdrasil v0.5](https://yggdrasil-network.github.io/2023/10/22/upcoming-v05-release.html)
|
||||
and its [Ironwood](https://github.com/Arceliar/ironwood) routing
|
||||
library. Yggdrasil's key insight was using the tree path from root to
|
||||
node as a routable coordinate, enabling greedy forwarding without
|
||||
global routing tables. FIPS adapts these algorithms for
|
||||
multi-transport operation, Nostr identity integration, and constrained
|
||||
MTU environments.
|
||||
|
||||
The theoretical foundation for greedy routing on tree embeddings draws
|
||||
on [Kleinberg's work](https://www.cs.cornell.edu/home/kleinber/swn.pdf)
|
||||
on navigable small-world networks, which showed that greedy forwarding
|
||||
succeeds in O(log² n) steps when the network has hierarchical
|
||||
structure. Thorup-Zwick compact routing schemes separately demonstrated
|
||||
that sublinear routing state is achievable with bounded stretch,
|
||||
motivating the use of tree coordinates rather than full routing tables.
|
||||
|
||||
## Split-Horizon Bloom Filter Propagation
|
||||
|
||||
FIPS distributes reachability information using bloom filters computed
|
||||
with a split-horizon rule: when advertising to a peer, exclude that
|
||||
peer's own contributions. This technique is borrowed from
|
||||
distance-vector routing protocols —
|
||||
[RIP](https://en.wikipedia.org/wiki/Routing_Information_Protocol)
|
||||
(1988) and [Babel](https://www.irif.fr/~jch/software/babel/) use
|
||||
split-horizon to prevent routing loops by not advertising a route back
|
||||
to the neighbor it was learned from. FIPS applies the same principle
|
||||
to probabilistic set advertisements rather than distance-vector tables.
|
||||
|
||||
## Cryptographic Identity as Network Address
|
||||
|
||||
FIPS nodes are identified by their Nostr public keys (secp256k1). The
|
||||
network address *is* the cryptographic identity — there is no separate
|
||||
address assignment or registration step.
|
||||
[CJDNS](https://github.com/cjdelisle/cjdns) pioneered this approach in
|
||||
overlay meshes, deriving IPv6 addresses from the double-SHA-512 of
|
||||
each node's public key. Tor [.onion
|
||||
addresses](https://spec.torproject.org/rend-spec-v3) and the IETF
|
||||
[Host Identity Protocol](https://en.wikipedia.org/wiki/Host_Identity_Protocol)
|
||||
(HIP) follow the same principle. FIPS uses Nostr's existing key
|
||||
infrastructure rather than introducing a new identity scheme.
|
||||
|
||||
## Dual-Layer Encryption
|
||||
|
||||
FIPS encrypts traffic twice: FMP provides hop-by-hop link encryption
|
||||
(protecting against transport-layer observers), while FSP provides
|
||||
independent end-to-end session encryption (protecting against
|
||||
intermediate FIPS nodes). This layered approach mirrors
|
||||
[Tor](https://www.torproject.org/), where each relay peels one layer
|
||||
of encryption (hop-by-hop) while the innermost layer protects
|
||||
end-to-end payload. [I2P](https://geti2p.net/) uses a similar garlic
|
||||
routing scheme with tunnel-layer and end-to-end encryption. Unlike Tor
|
||||
and I2P, FIPS does not provide anonymity — its dual encryption
|
||||
protects confidentiality and integrity rather than hiding traffic
|
||||
patterns.
|
||||
|
||||
## Noise Protocol Framework
|
||||
|
||||
FIPS uses the [Noise Protocol Framework](https://noiseprotocol.org/)
|
||||
at both protocol layers, with different handshake patterns chosen for
|
||||
each layer's threat model. FMP link encryption uses **Noise IK**,
|
||||
providing mutual authentication with a single round trip where the
|
||||
initiator knows the responder's static key in advance.
|
||||
[WireGuard](https://www.wireguard.com/) uses the same IK base pattern
|
||||
(extended with a pre-shared key as IKpsk2) for VPN tunnels. FSP
|
||||
session encryption uses **Noise XK**, the same pattern used by the
|
||||
[Lightning Network](https://github.com/lightning/bolts/blob/master/08-transport.md),
|
||||
where the initiator's static key is transmitted in a third message
|
||||
rather than the first. XK provides stronger initiator identity hiding
|
||||
at the cost of an additional round trip — a worthwhile tradeoff for
|
||||
session-layer traffic that traverses untrusted intermediate nodes. At
|
||||
the link layer, where both peers are configured and directly
|
||||
connected, IK's single round trip is preferred.
|
||||
|
||||
Specific Noise references and adapted constructions:
|
||||
|
||||
- Perrin, T. ["The Noise Protocol Framework"](https://noiseprotocol.org/noise.html).
|
||||
Revision 34, 2018. *Framework for building crypto protocols using
|
||||
Diffie-Hellman key agreement and AEAD ciphers. FSP uses the XK
|
||||
handshake pattern.*
|
||||
|
||||
- Donenfeld, J.A. ["WireGuard: Next Generation Kernel Network Tunnel"](https://www.wireguard.com/papers/wireguard.pdf).
|
||||
NDSS 2017. *Transport-independent cryptographic sessions bound to
|
||||
identity keys rather than network addresses; AEAD-only authentication
|
||||
model.*
|
||||
|
||||
## Index-Based Session Dispatch
|
||||
|
||||
FIPS uses locally-assigned 32-bit session indices to demultiplex
|
||||
incoming packets to the correct cryptographic session in O(1) time,
|
||||
without parsing source addresses or performing expensive lookups.
|
||||
This directly follows
|
||||
[WireGuard's](https://www.wireguard.com/papers/wireguard.pdf) receiver
|
||||
index approach, where each peer assigns a random index during
|
||||
handshake and the remote side includes it in every packet header.
|
||||
|
||||
## Replay Protection Over Unreliable Transports
|
||||
|
||||
FSP and FMP both use explicit per-packet counters with a sliding
|
||||
bitmap window for replay protection — the standard DTLS approach,
|
||||
chosen because implicit nonce counters desynchronize permanently under
|
||||
UDP packet loss or reordering.
|
||||
|
||||
- Rescorla, E., Modadugu, N. [RFC 6347](https://datatracker.ietf.org/doc/html/rfc6347):
|
||||
"Datagram Transport Layer Security Version 1.2". 2012. *Explicit
|
||||
sequence numbers with sliding bitmap window for replay protection
|
||||
over unreliable transports.*
|
||||
|
||||
## Transport-Agnostic Overlay Mesh
|
||||
|
||||
FIPS is designed to operate over any datagram-capable transport — UDP,
|
||||
raw Ethernet, Bluetooth, radio, serial — through a uniform transport
|
||||
abstraction. Several mesh overlays have demonstrated transport-agnostic
|
||||
design: [CJDNS](https://github.com/cjdelisle/cjdns) runs over UDP and
|
||||
Ethernet, [Yggdrasil](https://yggdrasil-network.github.io/) supports
|
||||
TCP and TLS transports, and [Tor](https://www.torproject.org/) can use
|
||||
pluggable transports to tunnel through various media. FIPS extends
|
||||
this pattern to shared-medium transports (radio, BLE) with
|
||||
per-transport MTU and discovery capabilities.
|
||||
|
||||
## Metrics Measurement Protocol
|
||||
|
||||
MMP's design assembles well-established measurement techniques into a
|
||||
unified per-link protocol. The SenderReport/ReceiverReport exchange
|
||||
structure follows [RTCP](https://www.rfc-editor.org/rfc/rfc3550)
|
||||
(RFC 3550), which uses the same report pairing for media stream
|
||||
quality monitoring in RTP sessions. MMP's jitter computation uses the
|
||||
RTCP interarrival jitter algorithm directly.
|
||||
|
||||
The smoothed RTT estimator uses the Jacobson/Karels algorithm
|
||||
([RFC 6298](https://www.rfc-editor.org/rfc/rfc6298)), the same SRTT
|
||||
computation used in TCP for retransmission timeout calculation since
|
||||
1988. MMP derives RTT from timestamp-echo in ReceiverReports with
|
||||
dwell-time compensation, rather than from packet round-trips.
|
||||
|
||||
The spin bit in the FMP frame header follows the
|
||||
[QUIC](https://www.rfc-editor.org/rfc/rfc9000) spin bit
|
||||
([RFC 9312](https://www.rfc-editor.org/rfc/rfc9312)) — a single bit
|
||||
that alternates each round trip, enabling passive latency measurement.
|
||||
FIPS implements the spin bit state machine but relies on
|
||||
timestamp-echo for SRTT, as irregular mesh traffic makes spin bit RTT
|
||||
unreliable.
|
||||
|
||||
The Expected Transmission Count (ETX) metric, computed from
|
||||
bidirectional delivery ratios, was introduced by
|
||||
[De Couto et al. (2003)](https://pdos.csail.mit.edu/papers/grid:mobicom03/paper.pdf)
|
||||
for wireless mesh routing and is used in protocols including
|
||||
[OLSR](https://en.wikipedia.org/wiki/Optimized_Link_State_Routing_Protocol)
|
||||
and [Babel](https://www.irif.fr/~jch/software/babel/). FIPS computes
|
||||
ETX per-link from MMP loss measurements for future use in candidate
|
||||
ranking.
|
||||
|
||||
The CE (Congestion Experienced) echo flag provides hop-by-hop
|
||||
[ECN](https://en.wikipedia.org/wiki/Explicit_Congestion_Notification)
|
||||
signaling, following the TCP/IP ECN echo pattern (RFC 3168). Transit
|
||||
nodes detect congestion via MMP loss/ETX metrics or kernel buffer
|
||||
drops and set the CE flag on forwarded frames; destination nodes mark
|
||||
ECN-capable IPv6 packets accordingly.
|
||||
|
||||
## Path MTU Discovery
|
||||
|
||||
FSP adapts RFC 1191 Path MTU Discovery for overlay networks. The
|
||||
classic ICMP Packet Too Big mechanism is replaced by a transit-node
|
||||
`min()` propagation in SessionDatagram and LookupResponse plus an
|
||||
end-to-end PathMtuNotification echo back to the source.
|
||||
|
||||
- Mogul, J., Deering, S. [RFC 1191](https://datatracker.ietf.org/doc/html/rfc1191):
|
||||
"Path MTU Discovery". 1990. *End-to-end path MTU discovery; FSP
|
||||
adapts this for overlay networks using transit-node min()
|
||||
propagation.*
|
||||
|
||||
## Session Restart and Simultaneous Initiation
|
||||
|
||||
FSP's epoch-based peer restart detection mirrors IKEv2's
|
||||
INITIAL_CONTACT notification, and its lowest-address-wins
|
||||
simultaneous-initiation tie-breaker mirrors IKEv2's resolution rule.
|
||||
|
||||
- Kaufman, C., Hoffman, P., Nir, Y., Eronen, P., Kivinen, T.
|
||||
[RFC 7296](https://datatracker.ietf.org/doc/html/rfc7296):
|
||||
"Internet Key Exchange Protocol Version 2 (IKEv2)". 2014.
|
||||
*Simultaneous initiation resolution (§2.8) and INITIAL_CONTACT peer
|
||||
restart detection (§2.4).*
|
||||
|
||||
## Hybrid Coordinate Warmup
|
||||
|
||||
FSP's hybrid coordinate warmup (CP flag piggybacking + standalone
|
||||
CoordsWarmup) draws on Yggdrasil's approach of embedding coordinates
|
||||
in session traffic to keep transit caches populated.
|
||||
|
||||
- [Yggdrasil Network](https://yggdrasil-network.github.io/).
|
||||
*Coordinate-based overlay routing with session traffic used to warm
|
||||
transit node coordinate caches.*
|
||||
|
||||
## Cryptographic Primitives
|
||||
|
||||
FIPS reuses [Nostr's](https://github.com/nostr-protocol/nips)
|
||||
cryptographic stack — secp256k1 for identity keys, Schnorr signatures
|
||||
for authentication, SHA-256 for hashing, and ChaCha20-Poly1305 for
|
||||
authenticated encryption. This is the same primitive set used across
|
||||
Bitcoin, Nostr, and a growing ecosystem of self-sovereign identity
|
||||
systems. No novel cryptography is introduced.
|
||||
|
||||
## Spanning-Tree Dynamics: Foundations
|
||||
|
||||
The CRDT framing, gossip dissemination, failure detection, link
|
||||
metrics, and route stability mechanisms in
|
||||
[spanning-tree-dynamics.md](spanning-tree-dynamics.md) draw on a body
|
||||
of academic and standards work, summarized below.
|
||||
|
||||
### Virtual Coordinate Routing
|
||||
|
||||
- Rao, A., Ratnasamy, S., Papadimitriou, C., Shenker, S., Stoica, I.
|
||||
["Geographic Routing without Location Information"](https://people.eecs.berkeley.edu/~sylvia/papers/p327-rao.pdf).
|
||||
MobiCom 2003. *Established virtual coordinate routing using network
|
||||
topology.*
|
||||
|
||||
### Greedy Embedding Theory
|
||||
|
||||
- Kleinberg, R.
|
||||
["Geographic Routing Using Hyperbolic Space"](https://www.semanticscholar.org/paper/Geographic-Routing-Using-Hyperbolic-Space-Kleinberg/f506b2ddb142d2ec539400297ba53383d958abef).
|
||||
IEEE INFOCOM 2007. *Proved every connected graph has a greedy
|
||||
embedding in hyperbolic space; showed spanning trees enable
|
||||
coordinate assignment.*
|
||||
|
||||
- Cvetkovski, A., Crovella, M.
|
||||
["Hyperbolic Embedding and Routing for Dynamic Graphs"](https://www.cs.bu.edu/faculty/crovella/paper-archive/infocom09-hyperbolic.pdf).
|
||||
IEEE INFOCOM 2009. *Dynamic embedding for nodes joining/leaving;
|
||||
introduced Gravity-Pressure routing for failure recovery.*
|
||||
|
||||
- Crovella, M. et al.
|
||||
["On the Choice of a Spanning Tree for Greedy Embedding"](https://www.cs.bu.edu/faculty/crovella/paper-archive/networking-science13.pdf).
|
||||
Networking Science 2013. *Analysis of how tree structure affects
|
||||
routing stretch.*
|
||||
|
||||
- Bläsius, T. et al.
|
||||
["Hyperbolic Embeddings for Near-Optimal Greedy Routing"](https://dl.acm.org/doi/10.1145/3381751).
|
||||
ACM Journal of Experimental Algorithmics 2020. *Achieved 100%
|
||||
success ratio with 6% stretch on Internet graph.*
|
||||
|
||||
### Link Metrics
|
||||
|
||||
- De Couto, D., Aguayo, D., Bicket, J., Morris, R.
|
||||
"A High-Throughput Path Metric for Multi-Hop Wireless Routing".
|
||||
MobiCom 2003. *Introduced ETX (Expected Transmission Count) as a
|
||||
link quality metric for wireless mesh networks.*
|
||||
|
||||
### Routing Protocol Stability
|
||||
|
||||
- IEEE 802.1D. "IEEE Standard for Local and Metropolitan Area
|
||||
Networks: Media Access Control (MAC) Bridges". *Spanning Tree
|
||||
Protocol (STP) — root election via bridge ID, BPDU exchange.*
|
||||
|
||||
- Moy, J. [RFC 2328](https://datatracker.ietf.org/doc/html/rfc2328):
|
||||
"OSPF Version 2". 1998. *Link-state routing with cumulative path
|
||||
costs and SPF computation. FIPS's local-only cost approach is
|
||||
contrasted with OSPF's cumulative model in
|
||||
[spanning-tree-dynamics.md §8](spanning-tree-dynamics.md#8-parent-selection).*
|
||||
|
||||
### Distributed Systems Primitives
|
||||
|
||||
- Shapiro, M., Preguiça, N., Baquero, C., Zawirski, M.
|
||||
"Conflict-free Replicated Data Types". SSS 2011. *Formal definition
|
||||
of CRDTs enabling coordination-free consistency.*
|
||||
|
||||
- Das, A., Gupta, I., Motivala, A.
|
||||
["SWIM: Scalable Weakly-consistent Infection-style Process Group Membership"](https://www.cs.cornell.edu/projects/Quicksilver/public_pdfs/SWIM.pdf).
|
||||
IPDPS 2002. *O(1) failure detection, O(log N) dissemination via
|
||||
gossip.*
|
||||
|
||||
- Kermarrec, A-M.
|
||||
["Gossiping in Distributed Systems"](https://www.distributed-systems.net/my-data/papers/2007.osr.pdf).
|
||||
ACM SIGOPS Operating Systems Review 2007. *Framework for
|
||||
gossip-based protocols achieving O(log N) propagation.*
|
||||
|
||||
## FIPS Contributions
|
||||
|
||||
The protocol builds on these foundations and adds several new elements:
|
||||
|
||||
- Cost-aware parent selection using local-only link metrics
|
||||
(`effective_depth = depth + link_cost`), replacing Yggdrasil's
|
||||
depth-only selection
|
||||
- Combined ETX + SRTT link cost formula with MMP-measured components
|
||||
- Flap dampening with mandatory switch bypass
|
||||
- Announcement suppression for transient state changes
|
||||
- Tree-only bloom filter merge with split-horizon exclusion
|
||||
- Hybrid coordinate warmup (CP flag piggybacking plus standalone
|
||||
CoordsWarmup) layered on top of SessionSetup self-bootstrapping
|
||||
- Bloom-guided tree routing for discovery (vs. flooding)
|
||||
- Reverse-path routing for LookupResponse via `recent_requests`
|
||||
|
||||
## External Reference Index
|
||||
|
||||
| Reference | Used by |
|
||||
| --------- | ------- |
|
||||
| [IEEE 802.1D STP](https://en.wikipedia.org/wiki/Spanning_Tree_Protocol) | spanning tree, root election |
|
||||
| [Yggdrasil v0.5](https://yggdrasil-network.github.io/2023/10/22/upcoming-v05-release.html) | tree coordinates, greedy routing |
|
||||
| [Ironwood](https://github.com/Arceliar/ironwood) | tree coordinates, candidate ranking |
|
||||
| [Kleinberg, Small-world](https://www.cs.cornell.edu/home/kleinber/swn.pdf) | greedy routing on tree embeddings |
|
||||
| [CJDNS](https://github.com/cjdelisle/cjdns) | cryptographic-identity-as-address |
|
||||
| [Tor](https://www.torproject.org/) | onion address scheme, dual-layer encryption |
|
||||
| [I2P](https://geti2p.net/) | dual-layer encryption (garlic routing) |
|
||||
| [HIP](https://en.wikipedia.org/wiki/Host_Identity_Protocol) | identity-as-address |
|
||||
| [Babel](https://www.irif.fr/~jch/software/babel/) | split-horizon, ETX |
|
||||
| [RIP](https://en.wikipedia.org/wiki/Routing_Information_Protocol) | split-horizon |
|
||||
| [Noise Framework](https://noiseprotocol.org/) | FMP IK, FSP XK |
|
||||
| [WireGuard](https://www.wireguard.com/) | IK pattern, receiver-index dispatch, identity-bound sessions |
|
||||
| [Lightning BOLT #8](https://github.com/lightning/bolts/blob/master/08-transport.md) | XK pattern |
|
||||
| [QUIC (RFC 9000)](https://www.rfc-editor.org/rfc/rfc9000) | spin bit, transport design |
|
||||
| [QUIC Spin Bit (RFC 9312)](https://www.rfc-editor.org/rfc/rfc9312) | passive RTT measurement |
|
||||
| [RTCP (RFC 3550)](https://www.rfc-editor.org/rfc/rfc3550) | sender/receiver report structure, jitter algorithm |
|
||||
| [TCP SRTT/RTO (RFC 6298)](https://www.rfc-editor.org/rfc/rfc6298) | Jacobson/Karels SRTT |
|
||||
| [ECN (RFC 3168)](https://www.rfc-editor.org/rfc/rfc3168) | CE echo |
|
||||
| [DTLS 1.2 (RFC 6347)](https://datatracker.ietf.org/doc/html/rfc6347) | replay window |
|
||||
| [IKEv2 (RFC 7296)](https://datatracker.ietf.org/doc/html/rfc7296) | INITIAL_CONTACT, simultaneous-initiation tie-breaker |
|
||||
| [PMTUD (RFC 1191)](https://datatracker.ietf.org/doc/html/rfc1191) | adapted PMTUD |
|
||||
| [ETX paper, De Couto et al.](https://pdos.csail.mit.edu/papers/grid:mobicom03/paper.pdf) | ETX metric |
|
||||
| [OLSR](https://en.wikipedia.org/wiki/Optimized_Link_State_Routing_Protocol) | ETX in mesh routing |
|
||||
| [Nostr](https://github.com/nostr-protocol/nips) | identity stack |
|
||||
@@ -0,0 +1,215 @@
|
||||
# FIPS Mesh-Interface Security
|
||||
|
||||
This document describes the threat model and design rationale for the
|
||||
operator-facing security posture of the `fips0` mesh interface on Linux.
|
||||
The default-deny nftables baseline shipped as `/etc/fips/fips.nft` is the
|
||||
artifact discussed below; for the operator activation steps and drop-in
|
||||
extension recipes, see [enable-mesh-firewall.md](../how-to/enable-mesh-firewall.md).
|
||||
|
||||
The baseline is a documented operator conffile, not an auto-loaded
|
||||
package side-effect. Activation is an explicit one-liner. The
|
||||
rationale for that design follows.
|
||||
|
||||
## Threat Model for `fips0`
|
||||
|
||||
The mesh is a flat layer-3 segment. Every mesh node that can route to
|
||||
you can deliver packets to your `fips0` address — your direct peers
|
||||
forward traffic from non-peer mesh nodes onto your `fips0` the same
|
||||
way any router forwards transit traffic. Identity on the mesh is the
|
||||
originating node's npub — the FMP link layer authenticates direct
|
||||
peers with Noise IK and the FSP session layer authenticates session
|
||||
endpoints with Noise XK — but identity is **not** authorization.
|
||||
Knowing who sent a packet does not, by itself, decide whether the
|
||||
local host should accept it.
|
||||
|
||||
That means: any service on a mesh host that binds to a wildcard
|
||||
address (`0.0.0.0`, `[::]`, or any IPv6 address that includes the
|
||||
`fips0` interface in its scope) is reachable from every mesh node
|
||||
that can route to you by default, not only from your direct peers.
|
||||
There is no NAT, no perimeter firewall, no "local-only" address
|
||||
space between you and an arbitrary mesh node. The mesh is closer
|
||||
to a shared LAN than to the public internet.
|
||||
|
||||
Compare to the corresponding internet trust assumptions:
|
||||
|
||||
| Surface | Public internet | FIPS mesh (no baseline) |
|
||||
|---|---|---|
|
||||
| Reachability from arbitrary mesh node | Mediated by NAT, firewalls, ISPs | Direct |
|
||||
| Default identity | None | Originating node's npub (authenticated) |
|
||||
| Default authorization | None | None |
|
||||
| Accidental exposure cost | Low (NAT hides you) | High (every mesh node sees you) |
|
||||
|
||||
The third row is the gap this document closes. The default-deny
|
||||
baseline removes "accidental exposure" from the failure modes an
|
||||
operator has to think about.
|
||||
|
||||
## The Default-Deny Baseline
|
||||
|
||||
The shipped baseline is `/etc/fips/fips.nft`. It defines a single
|
||||
nftables table, `inet fips`, with one chain hooked at `input`. The
|
||||
chain:
|
||||
|
||||
1. Returns immediately for any packet not arriving on `fips0`. This
|
||||
makes the table a no-op for every other interface — Docker, Tor,
|
||||
the host's main filter table, OPNsense, anything.
|
||||
2. Accepts packets that conntrack identifies as `established` or
|
||||
`related`. Replies to outbound flows initiated from the mesh host
|
||||
come back; ICMPv6 errors related to existing flows (Packet Too
|
||||
Big, Destination Unreachable) come back.
|
||||
3. Accepts ICMPv6 echo-request, so `ping6` reachability tests work.
|
||||
4. Includes operator drop-ins from `/etc/fips/fips.d/*.nft`. An empty
|
||||
directory is fine — the include glob simply matches nothing.
|
||||
5. Falls through to `counter drop`. Every dropped packet increments
|
||||
the counter, visible via `nft list table inet fips`.
|
||||
|
||||
Outbound from `fips0` is unrestricted. The baseline is concerned only
|
||||
with what the mesh host accepts, not what it sends.
|
||||
|
||||
The file is a documented dpkg conffile. Operator edits to
|
||||
`/etc/fips/fips.nft` are preserved across upgrades, the same way
|
||||
edits to `/etc/fips/fips.yaml` and `/etc/fips/hosts` are preserved.
|
||||
If the packaged baseline is ever updated upstream, dpkg prompts the
|
||||
operator on upgrade rather than silently overwriting local changes.
|
||||
|
||||
The canonical artifact is the file itself; read it for the inline
|
||||
documentation that the rest of this document references.
|
||||
|
||||
## Why no auto-load on package install
|
||||
|
||||
The `postinst` script does **not** enable `fips-firewall.service`.
|
||||
This is deliberate. Quietly mutating host firewall state on package
|
||||
install is hostile on every axis that matters: it surprises operators
|
||||
who already have their own nftables ruleset, it can collide with
|
||||
podman/Docker/OPNsense integrations even though the early-return
|
||||
makes it technically safe, and it converts an explicit security
|
||||
decision into an invisible one. The mesh-interface filter belongs to
|
||||
the operator, not to the package's `postinst`.
|
||||
|
||||
The activation gesture is one short, well-formed command. The
|
||||
rationale is documented in the file's inline header and in this
|
||||
document. That is enough; auto-loading would trade discoverability
|
||||
for no real gain.
|
||||
|
||||
## Coexistence with other firewalls
|
||||
|
||||
The `inet fips` table only matches packets arriving on `fips0`.
|
||||
Anything else returns from the chain on the first rule. Specifically:
|
||||
|
||||
- **Docker / containerd** install nftables rules in the `ip` and `ip6`
|
||||
families and operate on `docker0`, `br-*`, and `veth*`
|
||||
interfaces. They do not touch `fips0`. The two tables coexist
|
||||
without interference.
|
||||
- **Tor** runs in user space and does not install firewall rules. The
|
||||
baseline is independent of Tor's onion-service and SOCKS listeners.
|
||||
- **OPNsense** is an upstream perimeter device. The baseline runs on
|
||||
the local host and applies only to traffic that has already reached
|
||||
the host's `fips0` interface. They do not interact.
|
||||
- **The host's main `/etc/nftables.conf`** typically defines a
|
||||
separate `inet filter` table. nftables allows multiple tables in
|
||||
the same family to coexist; both run in parallel at hook
|
||||
`input`/priority 0 and the `iifname != "fips0" return` rule keeps
|
||||
the `inet fips` table from interfering with anything outside the
|
||||
mesh interface.
|
||||
- **`inet fips_gateway`**, when `fips-gateway` is running, manages
|
||||
DNAT/SNAT on the LAN-facing interface to translate virtual IPs to
|
||||
mesh addresses. It is a separate concern owned by the gateway
|
||||
binary and is unrelated to this baseline. See the section below.
|
||||
|
||||
## Coexistence with `inet fips_gateway`
|
||||
|
||||
When `fips-gateway` is running, it manages a separate nftables
|
||||
table, `inet fips_gateway`, containing the DNAT and masquerade rules
|
||||
that translate between the gateway's virtual-IP pool and mesh
|
||||
addresses on the LAN-facing interface. That table is created and
|
||||
torn down by the gateway binary at runtime and is not an operator
|
||||
artifact in the same sense as `inet fips`.
|
||||
|
||||
The two tables do not interfere:
|
||||
|
||||
- `inet fips` filters inbound on `fips0`.
|
||||
- `inet fips_gateway` performs NAT on the LAN interface.
|
||||
|
||||
They operate on different interfaces and at different hook points
|
||||
(`input` filter vs. `prerouting`/`postrouting` NAT). Both can be
|
||||
loaded simultaneously on a gateway host, and that is the intended
|
||||
deployment shape. See [fips-gateway.md](fips-gateway.md) for the
|
||||
gateway table's structure.
|
||||
|
||||
## What the Baseline Does Not Cover
|
||||
|
||||
The baseline is one half of a defense-in-depth posture. It is
|
||||
explicitly not:
|
||||
|
||||
- **Outbound filtering.** Anything the mesh host originates on
|
||||
`fips0` is unrestricted. If you need to constrain what the host
|
||||
can send to the mesh, add rules to a separate chain hooked at
|
||||
`output` — out of scope for the baseline.
|
||||
- **Application-layer authorization.** The baseline decides whether
|
||||
a packet reaches a service. It does not decide whether the
|
||||
originating mesh node's npub is allowed to use that service. That
|
||||
is the application's responsibility (e.g., an `authorized_keys`
|
||||
file for SSH, an ACL in the application's configuration).
|
||||
- **ACL on the mesh handshake.** The FMP Noise IK handshake
|
||||
authenticates the peer's npub and, on both inbound and outbound
|
||||
paths, consults the peer ACL (`peers.allow` / `peers.deny`) before
|
||||
promoting the connection. The ACL evaluates in TCP-Wrappers order:
|
||||
an `allow` match permits, otherwise a `deny` match rejects,
|
||||
otherwise the connection is permitted. A strict allowlist posture
|
||||
therefore requires an explicit `ALL` entry in `peers.deny`; a
|
||||
populated `peers.allow` alone does not turn the ACL into a strict
|
||||
allowlist. Mesh-level ACLs are a separate concern from the inbound
|
||||
packet filter described here; see the peer ACL section in
|
||||
[../reference/security.md](../reference/security.md).
|
||||
- **Compromised peers.** A peer whose key has been stolen or whose
|
||||
host has been taken over is, by mesh-level identity, still that
|
||||
peer. Source-address filtering in drop-ins operates on the source
|
||||
mesh address of inbound traffic regardless of whether that source
|
||||
is a direct peer or a multi-hop mesh node, and so can limit damage
|
||||
from a known-compromised mesh address; but the baseline cannot
|
||||
revoke trust on its own.
|
||||
|
||||
Treat the baseline as removing the "wide-open by default" failure
|
||||
mode. Higher-layer authorization decisions are the operator's and
|
||||
the application's, the same as on any other shared network.
|
||||
|
||||
## Future Work
|
||||
|
||||
The current baseline is Linux-only. Parallel work for other targets:
|
||||
|
||||
- **macOS PF baseline.** macOS uses Packet Filter (PF), inherited
|
||||
from OpenBSD. PF maps cleanly onto the same conceptual model as
|
||||
nftables: stateful inspection (`keep state` ≈ `ct state
|
||||
established,related`), default policy, anchor-based modular rule
|
||||
loading. A `packaging/macos/fips.pf` will land alongside the
|
||||
Linux baseline with the same posture: documented asset, no
|
||||
auto-load, operator opts in via launchd. The macOS interface name
|
||||
is `utunN` rather than `fips0`, so the rule template needs runtime
|
||||
substitution or a PF interface group assigned at TUN bring-up;
|
||||
this is being worked through with the macOS port.
|
||||
- **OpenWrt fw4 path.** OpenWrt's fw4 already drives nftables under
|
||||
the hood, but rules go into `/etc/nftables.d/` includes or UCI
|
||||
entries in `/etc/config/firewall`, not a free-standing
|
||||
`fips.nft`. The ipk will ship a layout-compatible variant or
|
||||
document the operator setup separately, decided when the OpenWrt
|
||||
packaging is updated.
|
||||
- **Cross-OS gateway abstraction.** `fips-gateway` is currently
|
||||
Linux-only because `src/gateway/nat.rs` uses the `rustables`
|
||||
netlink API directly. macOS gateway support requires a PF-backed
|
||||
equivalent behind a shared backend trait. This is a larger lift
|
||||
than the static baseline and is tracked separately under the same
|
||||
cross-OS thread.
|
||||
|
||||
When those land, this document will grow per-OS sections describing
|
||||
each baseline's load mechanism and extension points. The threat
|
||||
model and the operator-extension principle are the same on every OS;
|
||||
only the filter syntax and the activation gesture differ.
|
||||
|
||||
## See also
|
||||
|
||||
- [enable-mesh-firewall.md](../how-to/enable-mesh-firewall.md) — operator
|
||||
activation steps, drop-in recipes, drop visibility and debugging
|
||||
- [../reference/security.md](../reference/security.md) — consolidated
|
||||
security reference (cryptographic primitives, peer ACL format,
|
||||
filesystem permissions, default network exposures)
|
||||
- [fips-gateway.md](fips-gateway.md) — `fips-gateway` service and the
|
||||
separate `inet fips_gateway` table
|
||||
@@ -1,10 +1,10 @@
|
||||
# FIPS Session Protocol (FSP)
|
||||
|
||||
The FIPS Session Protocol is the top protocol layer in the FIPS stack. It sits
|
||||
above the FIPS Mesh Protocol (FMP) and below applications (native FIPS API or
|
||||
IPv6 adapter). FSP provides end-to-end authenticated, encrypted datagram
|
||||
delivery between any two FIPS nodes, regardless of how many intermediate hops
|
||||
separate them.
|
||||
The FIPS Session Protocol is the topmost layer of the FIPS protocol stack.
|
||||
It sits above the FIPS Mesh Protocol (FMP) and below applications (native
|
||||
FIPS API or IPv6 adapter). FSP provides end-to-end authenticated, encrypted
|
||||
datagram delivery between any two FIPS nodes, regardless of how many
|
||||
intermediate hops separate them.
|
||||
|
||||
## Role
|
||||
|
||||
@@ -99,9 +99,9 @@ FMP signals routing failures asynchronously:
|
||||
a standalone CoordsWarmup (rate-limited), re-discovering the destination's
|
||||
current coordinates, and resetting the warmup counter.
|
||||
- **MtuExceeded**: A transit node cannot forward a SessionDatagram because
|
||||
the packet exceeds the next-hop link MTU. FSP uses the reported bottleneck
|
||||
MTU to adjust its session-layer path MTU estimate. MtuExceeded is the
|
||||
reactive complement to the proactive `path_mtu` field in SessionDatagram.
|
||||
the packet exceeds the next-hop link MTU; FSP adjusts its
|
||||
session-layer path MTU estimate from the reported bottleneck. See
|
||||
[fips-mtu.md](fips-mtu.md) for the full forward/reverse MTU model.
|
||||
|
||||
All three signals are generated by transit nodes (not the destination) and
|
||||
travel back to the source inside a new SessionDatagram. They are plaintext
|
||||
@@ -123,9 +123,10 @@ a destination with no existing session.
|
||||
FSP uses Noise XK for session key agreement (Noise Protocol Framework;
|
||||
Perrin 2018). The initiator knows the destination's npub (required for
|
||||
XK's pre-message `s` token); the responder learns the initiator's
|
||||
identity from msg3 (not msg1, unlike IK at the link layer). This provides stronger initiator identity hiding
|
||||
— the initiator's static key is encrypted under the established shared
|
||||
secret rather than under only the responder's static key.
|
||||
identity from msg3 (not msg1, unlike IK at the link layer). This
|
||||
provides stronger initiator identity hiding — the initiator's static
|
||||
key is encrypted under the established shared secret rather than under
|
||||
only the responder's static key.
|
||||
|
||||
The handshake is a three-message flow carried in SessionSetup, SessionAck,
|
||||
and SessionMsg3:
|
||||
@@ -244,17 +245,9 @@ operations) rather than under only the responder's static key.
|
||||
|
||||
### Cryptographic Primitives
|
||||
|
||||
| Component | Choice | Notes |
|
||||
| --------- | ------ | ----- |
|
||||
| Curve | secp256k1 | Nostr-native |
|
||||
| DH | ECDH on secp256k1 | Standard EC Diffie-Hellman |
|
||||
| Cipher | ChaCha20-Poly1305 | AEAD, same as NIP-44 |
|
||||
| Hash | SHA-256 | Nostr-native |
|
||||
| Key derivation | HKDF-SHA256 | Standard Noise KDF |
|
||||
|
||||
These choices prioritize compatibility with the Nostr cryptographic stack —
|
||||
secp256k1 + ChaCha20-Poly1305 + SHA-256 aligns with the NIP-44 encrypted
|
||||
messaging standard.
|
||||
FSP uses ChaCha20-Poly1305 with secp256k1 ECDH; see
|
||||
[../reference/security.md](../reference/security.md) for the full
|
||||
primitive table shared with the link layer.
|
||||
|
||||
### secp256k1 Parity Normalization
|
||||
|
||||
@@ -391,22 +384,10 @@ signal generation (100ms per destination).
|
||||
|
||||
## Identity Cache
|
||||
|
||||
The identity cache maps FIPS address prefix (15 bytes, the `fd00::/8` IPv6
|
||||
address minus the `fd` prefix) to `(NodeAddr, PublicKey)`. This cache is
|
||||
needed only when using the IPv6 adapter — the native FIPS API provides the
|
||||
public key directly.
|
||||
|
||||
The mapping is deterministic (derived from the public key via SHA-256) and
|
||||
never becomes stale. The cache uses LRU-only eviction bounded by a
|
||||
configurable size (default 10K entries). There is no TTL — entries are evicted
|
||||
only when the cache is full and space is needed for a new entry.
|
||||
|
||||
Cache population mechanisms:
|
||||
|
||||
- **DNS lookup**: The primary path. Resolving `npub1xxx...xxx.fips` derives
|
||||
the IPv6 address and populates the identity cache.
|
||||
- **Inbound traffic**: Authenticated sessions from other nodes populate the
|
||||
cache with their identity information.
|
||||
The IPv6 adapter requires an identity cache to map `fd00::/8` addresses
|
||||
back to `(NodeAddr, PublicKey)` for routing; see
|
||||
[fips-ipv6-adapter.md](fips-ipv6-adapter.md#identity-cache) for the
|
||||
cache rationale, eviction policy, and population mechanics.
|
||||
|
||||
## Coordinate Cache
|
||||
|
||||
@@ -451,84 +432,30 @@ node caches (still within their 300s TTL) are re-warmed.
|
||||
|
||||
## Session-Layer MMP
|
||||
|
||||
Each established session runs its own Metrics Measurement Protocol instance,
|
||||
providing end-to-end quality metrics independent of the number of hops.
|
||||
FSP runs an MMP instance per established session for end-to-end metrics
|
||||
independent of hop count. Reports are encrypted and forwarded through
|
||||
every transit link, so bandwidth cost is proportional to path length;
|
||||
the session-layer report intervals are correspondingly higher than the
|
||||
link-layer intervals (clamped to `[500ms, 10s]` vs. `[1s, 5s]`).
|
||||
|
||||
### Relationship to Link-Layer MMP
|
||||
The session-layer instance shares its algorithms (SRTT, jitter, loss,
|
||||
ETX) and report wire format with link-layer MMP. The differences —
|
||||
configuration namespace (`node.session_mmp.*`), routing scope,
|
||||
send-failure backoff, idle-timeout interaction, and the
|
||||
PathMtuNotification mechanism — are documented in the unified MMP
|
||||
treatment at [fips-mmp.md](fips-mmp.md). For the end-to-end path-MTU
|
||||
echo specifically, see [fips-mtu.md](fips-mtu.md). Reports and
|
||||
PathMtuNotification do **not** reset the session idle timer, so a
|
||||
session carrying only MMP traffic still tears down at the configured
|
||||
idle threshold.
|
||||
|
||||
Session-layer MMP uses the same report wire format (SenderReport 0x11,
|
||||
ReceiverReport 0x12) and identical algorithms as link-layer MMP, but with
|
||||
two key differences:
|
||||
### MtuExceeded Handling
|
||||
|
||||
1. **End-to-end routing**: Session reports are encrypted and forwarded through
|
||||
every transit link. A 3-hop session generates report traffic on all 3 links,
|
||||
making bandwidth cost proportional to path length.
|
||||
2. **Independent configuration**: The `node.session_mmp.*` parameters are
|
||||
separate from `node.mmp.*`, allowing operators to run a lighter mode for
|
||||
sessions (e.g., Lightweight) while keeping Full mode on links.
|
||||
|
||||
### Metrics
|
||||
|
||||
The same metrics as link-layer MMP: SRTT, loss rate, jitter, goodput, OWD
|
||||
trend, ETX, and dual EWMA trends. Session MMP additionally tracks observed
|
||||
path MTU.
|
||||
|
||||
### Report Intervals
|
||||
|
||||
Session-layer report intervals are higher than link-layer to account for
|
||||
bandwidth cost: clamped to [500ms, 10s] with a cold-start interval of 1s
|
||||
(vs. link-layer [100ms, 2s] with 500ms cold-start).
|
||||
|
||||
### Path MTU Tracking
|
||||
|
||||
PathMtuNotification (message type 0x13) provides end-to-end path MTU
|
||||
feedback, adapting RFC 1191 Path MTU Discovery for overlay networks — the
|
||||
transit-node `min()` propagation replaces ICMP Packet Too Big:
|
||||
|
||||
1. The source sets `path_mtu` in each SessionDatagram envelope to its
|
||||
outbound link MTU.
|
||||
2. Each transit node applies `min(current, transport.link_mtu(addr))` before
|
||||
forwarding.
|
||||
3. The destination receives the forward-path minimum and sends a
|
||||
PathMtuNotification (2-byte body: u16 LE path_mtu) back to the source.
|
||||
4. The source applies the notification with hysteresis:
|
||||
- **Decrease**: immediate (take lower value).
|
||||
- **Increase**: requires 3 consecutive higher-value notifications spanning
|
||||
at least 2 × notification interval.
|
||||
5. Notifications are sent on first measurement, on any decrease, and
|
||||
periodically at `max(10s, 5 × SRTT)`.
|
||||
|
||||
### Send Failure Backoff
|
||||
|
||||
When a session MMP report cannot be delivered (destination unreachable, no
|
||||
route), the sender applies exponential backoff to the probe interval (a
|
||||
standard distributed systems pattern for transient failure handling):
|
||||
|
||||
- Each consecutive failure doubles the interval: 2x, 4x, 8x, 16x, 32x
|
||||
- Backoff caps at 32x the base interval (5 consecutive failures)
|
||||
- A successful send resets to the normal SRTT-based interval
|
||||
- Debug logging is suppressed after 3 consecutive failures; a summary is
|
||||
logged when the destination becomes reachable again
|
||||
|
||||
This prevents wasted CPU and log noise when a session's remote endpoint has
|
||||
departed the network but the local session has not yet timed out.
|
||||
|
||||
### Idle Timeout Interaction
|
||||
|
||||
MMP reports (SenderReport, ReceiverReport) and PathMtuNotification do **not**
|
||||
reset the session idle timer. Only application data (DataPacket, type 0x10)
|
||||
resets `last_activity`. This ensures sessions with no application traffic
|
||||
tear down after `node.session.idle_timeout_secs` (default 90s), while MMP
|
||||
continues providing measurement data up to the teardown moment.
|
||||
|
||||
### Operator Logging
|
||||
|
||||
Session metrics are logged at info level (configurable via
|
||||
`node.session_mmp.log_interval_secs`, default 30s):
|
||||
|
||||
```text
|
||||
MMP session metrics session=npub1tdwa...84le rtt=4.3ms loss=0.6% jitter=0.2ms goodput=71.3MB/s mtu=1472 tx_pkts=1234 rx_pkts=5678
|
||||
```
|
||||
When FMP signals MtuExceeded (a transit node could not forward a
|
||||
SessionDatagram because it exceeded the next-hop link MTU), FSP uses
|
||||
the reported bottleneck MTU to adjust its session-layer path MTU
|
||||
estimate immediately. See [fips-mtu.md](fips-mtu.md) for the full
|
||||
reactive PMTUD mechanism.
|
||||
|
||||
## Implementation Status
|
||||
|
||||
@@ -559,36 +486,27 @@ MMP session metrics session=npub1tdwa...84le rtt=4.3ms loss=0.6% jitter=0.2ms go
|
||||
|
||||
### FIPS Internal Documentation
|
||||
|
||||
- [fips-intro.md](fips-intro.md) — Protocol overview and architecture
|
||||
- [fips-concepts.md](fips-concepts.md) — Protocol overview
|
||||
- [fips-architecture.md](fips-architecture.md) — Layer architecture and
|
||||
identity model
|
||||
- [fips-mesh-layer.md](fips-mesh-layer.md) — FMP specification (below FSP)
|
||||
- [fips-ipv6-adapter.md](fips-ipv6-adapter.md) — IPv6 adaptation layer (above FSP)
|
||||
- [fips-mesh-operation.md](fips-mesh-operation.md) — Routing, discovery, and
|
||||
error recovery
|
||||
- [fips-wire-formats.md](fips-wire-formats.md) — Wire format reference for all
|
||||
session message types
|
||||
- [fips-ipv6-adapter.md](fips-ipv6-adapter.md) — IPv6 adaptation layer
|
||||
(above FSP)
|
||||
- [fips-mesh-operation.md](fips-mesh-operation.md) — Routing, discovery,
|
||||
and error recovery
|
||||
- [fips-mmp.md](fips-mmp.md) — Metrics Measurement Protocol (link + session)
|
||||
- [fips-mtu.md](fips-mtu.md) — Path MTU model (PathMtuNotification,
|
||||
MtuExceeded, hysteresis)
|
||||
- [fips-prior-work.md](fips-prior-work.md) — Noise XK, WireGuard,
|
||||
DTLS replay window, IKEv2 simultaneous initiation, hybrid coordinate
|
||||
warmup citations
|
||||
- [../reference/wire-formats.md](../reference/wire-formats.md) — Wire
|
||||
format reference for all session message types
|
||||
- [../reference/security.md](../reference/security.md) — Cryptographic
|
||||
primitives and rekey defaults
|
||||
|
||||
### External References
|
||||
|
||||
- Perrin, T. ["The Noise Protocol Framework"](https://noiseprotocol.org/noise.html).
|
||||
Revision 34, 2018. *Framework for building crypto protocols using Diffie-Hellman
|
||||
key agreement and AEAD ciphers. FSP uses the XK handshake pattern.*
|
||||
|
||||
- Donenfeld, J.A. ["WireGuard: Next Generation Kernel Network Tunnel"](https://www.wireguard.com/papers/wireguard.pdf).
|
||||
NDSS 2017. *Transport-independent cryptographic sessions bound to identity keys
|
||||
rather than network addresses; AEAD-only authentication model.*
|
||||
|
||||
- Rescorla, E., Modadugu, N. [RFC 6347](https://datatracker.ietf.org/doc/html/rfc6347):
|
||||
"Datagram Transport Layer Security Version 1.2". 2012. *Explicit sequence numbers
|
||||
with sliding bitmap window for replay protection over unreliable transports.*
|
||||
|
||||
- Kaufman, C., Hoffman, P., Nir, Y., Eronen, P., Kivinen, T.
|
||||
[RFC 7296](https://datatracker.ietf.org/doc/html/rfc7296):
|
||||
"Internet Key Exchange Protocol Version 2 (IKEv2)". 2014. *Simultaneous
|
||||
initiation resolution (§2.8) and INITIAL_CONTACT peer restart detection (§2.4).*
|
||||
|
||||
- Mogul, J., Deering, S. [RFC 1191](https://datatracker.ietf.org/doc/html/rfc1191):
|
||||
"Path MTU Discovery". 1990. *End-to-end path MTU discovery; FSP adapts this for
|
||||
overlay networks using transit-node min() propagation.*
|
||||
|
||||
- [Yggdrasil Network](https://yggdrasil-network.github.io/). *Coordinate-based
|
||||
overlay routing with session traffic used to warm transit node coordinate caches.*
|
||||
|
||||
@@ -71,8 +71,11 @@ quality.
|
||||
2. Compute **effective depth** for each candidate peer:
|
||||
`effective_depth = peer.depth + link_cost`, where
|
||||
`link_cost = etx * (1.0 + srtt_ms / 100.0)` using locally measured MMP
|
||||
metrics. When MMP metrics have not yet converged, `link_cost` defaults to
|
||||
1.0, preserving pure depth-based behavior as a graceful fallback.
|
||||
metrics. During cold start (no peer has MMP data yet), candidates without
|
||||
measurements default to `link_cost = 1.0`, preserving pure depth-based
|
||||
behavior. Once any peer has MMP data, unmeasured candidates are excluded
|
||||
so that a freshly connected peer cannot win parent selection on the
|
||||
default cost alone.
|
||||
3. Apply **hysteresis**: switch parents only when the best candidate's
|
||||
effective depth is significantly better than the current parent's:
|
||||
`best_eff_depth < current_eff_depth * (1.0 - parent_hysteresis)`
|
||||
@@ -109,6 +112,9 @@ immediate parent reselection:
|
||||
(ETX and SRTT). No cumulative path costs are propagated, avoiding
|
||||
the trust problems inherent in self-reported cost metrics in a
|
||||
permissionless network.
|
||||
- **Loop rejection**: Candidates whose advertised ancestry already contains
|
||||
the local node are skipped, preventing two nodes from selecting each
|
||||
other as parent and entering an alternating coordinate loop.
|
||||
|
||||
### After Parent Change
|
||||
|
||||
@@ -180,9 +186,14 @@ A node re-announces (propagates) only when its own state changes:
|
||||
|
||||
- **Root changed**: Always propagate — this is a significant topology event
|
||||
- **Depth changed**: Always propagate — affects routing distance calculations
|
||||
- **Mid-chain ancestor swap**: A reroute that replaces an interior ancestor
|
||||
without changing the root or the path length still alters the node's
|
||||
coordinate path, so it propagates. Without this, downstream peers would
|
||||
route into a phantom intermediate that no longer appears on the parent's
|
||||
tree.
|
||||
- **Sequence-only refresh**: Does NOT propagate beyond depth 1 — peers that
|
||||
receive a sequence-only update do not re-announce, because their own root
|
||||
and depth have not changed
|
||||
receive a sequence-only update do not re-announce, because their own root,
|
||||
depth, and address path have not changed
|
||||
|
||||
This means TreeAnnounce cascades through the tree proportional to depth,
|
||||
not network size. A change at depth D affects at most D nodes along the
|
||||
@@ -303,12 +314,15 @@ Example: In a 1000-node network with depth 10 and 5 peers, a node stores
|
||||
| Rate limiting (500ms per peer) | **Implemented** |
|
||||
| Coord cache flush on parent change | **Implemented** |
|
||||
| Flap dampening (extended hold-down on rapid switches) | **Implemented** |
|
||||
| Loop rejection (ancestry self-check in `evaluate_parent`) | **Implemented** |
|
||||
| Mid-chain ancestor swap propagation | **Implemented** |
|
||||
| Per-ancestry-entry signatures | Future direction |
|
||||
|
||||
## References
|
||||
|
||||
- [fips-mesh-operation.md](fips-mesh-operation.md) — How the spanning tree
|
||||
fits into mesh routing
|
||||
- [fips-wire-formats.md](fips-wire-formats.md) — TreeAnnounce wire format
|
||||
- [../reference/wire-formats.md](../reference/wire-formats.md) —
|
||||
TreeAnnounce wire format
|
||||
- [spanning-tree-dynamics.md](spanning-tree-dynamics.md) — Convergence
|
||||
scenario walkthroughs
|
||||
|
||||
@@ -1,5 +1,7 @@
|
||||
# FIPS Transport Layer
|
||||
|
||||
<!-- markdownlint-disable MD024 -->
|
||||
|
||||
The transport layer is the bottom of the FIPS protocol stack. It delivers
|
||||
datagrams between transport-specific endpoints over arbitrary physical or
|
||||
logical media. Everything above — peer authentication, routing, encryption,
|
||||
@@ -14,7 +16,7 @@ address, deliver the datagram to that address, and push inbound datagrams up
|
||||
to the FIPS Mesh Protocol (FMP) above.
|
||||
|
||||
The transport layer deals exclusively in **transport addresses** — IP:port
|
||||
tuples, MAC addresses, .onion identifiers, radio device addresses. These are
|
||||
or hostname:port addresses, MAC addresses, .onion identifiers, radio device addresses. These are
|
||||
opaque to every layer above FMP. The mapping from transport address to FIPS
|
||||
identity happens at the link layer after the Noise IK link handshake completes.
|
||||
The word "peer" belongs to the link layer and above; the transport layer
|
||||
@@ -50,7 +52,7 @@ determine how much payload can fit in a single packet after link-layer
|
||||
encryption overhead.
|
||||
|
||||
MTU is fundamentally a per-link property. A transport with a fixed MTU
|
||||
(Ethernet: 1500, UDP configured at 1472) returns the same value for every
|
||||
(Ethernet effective 1499, UDP default 1280) returns the same value for every
|
||||
link — this is the degenerate case. Transports that negotiate MTU
|
||||
per-connection (e.g., BLE ATT_MTU) report the negotiated value for each
|
||||
link individually.
|
||||
@@ -68,11 +70,31 @@ forwarding and LookupResponse transit annotation.
|
||||
### Connection Lifecycle
|
||||
|
||||
For connection-oriented transports, manage the underlying connection: TCP
|
||||
handshake, Tor circuit establishment, Bluetooth pairing. FMP cannot begin
|
||||
the Noise IK link handshake until the transport-layer connection is established.
|
||||
handshake, Tor circuit establishment, BLE pairing. FMP cannot begin
|
||||
the Noise IK link handshake until the transport-layer connection is
|
||||
established.
|
||||
|
||||
Connectionless transports (UDP, raw Ethernet) skip this — datagrams can flow
|
||||
immediately to any reachable address.
|
||||
Connection-oriented transports expose a non-blocking connect interface.
|
||||
`connect(addr)` initiates the connection in a background task and returns
|
||||
immediately. `connection_state(addr)` reports the current status:
|
||||
|
||||
```text
|
||||
ConnectionState {
|
||||
None No connection attempt in progress
|
||||
Connecting Background task running
|
||||
Connected Ready for send()
|
||||
Failed(msg) Error message from failed attempt
|
||||
}
|
||||
```
|
||||
|
||||
Connectionless transports (UDP, raw Ethernet) return `Connected`
|
||||
immediately — no async work needed.
|
||||
|
||||
At the node level, `PendingConnect` entries track links waiting for
|
||||
transport connection. `poll_pending_connects()` runs each tick, checks
|
||||
`connection_state()`, and calls `start_handshake()` on success or
|
||||
`schedule_retry()` on failure. This decouples transport-layer connection
|
||||
(which may take seconds for Tor circuits) from the FMP event loop.
|
||||
|
||||
### Discovery (Optional)
|
||||
|
||||
@@ -95,9 +117,8 @@ for internet connectivity:
|
||||
|
||||
| Transport | Addressing | MTU | Reliability | Notes |
|
||||
| --------- | ---------- | --- | ----------- | ----- |
|
||||
| UDP/IP | IP:port | 1280–1472 | Unreliable | Primary internet transport |
|
||||
| TCP/IP | IP:port | Stream | Reliable | Requires length-prefix framing |
|
||||
| WebSocket | URL | Stream | Reliable | Browser-compatible |
|
||||
| UDP/IP | host:port | 1280–1472 | Unreliable | Primary internet transport |
|
||||
| TCP/IP | host:port | Stream | Reliable | Requires length-prefix framing |
|
||||
| Tor | .onion | Stream | Reliable | High latency, strong anonymity |
|
||||
|
||||
**Shared medium transports** operate over broadcast- or multicast-capable
|
||||
@@ -107,7 +128,6 @@ media:
|
||||
| --------- | ---------- | --- | ----------- | ----- |
|
||||
| Ethernet | MAC | 1500 | Unreliable | Raw AF_PACKET frames |
|
||||
| WiFi | MAC | 1500 | Unreliable | Infrastructure mode = Ethernet |
|
||||
| Bluetooth | BD_ADDR | 672–64K | Reliable | L2CAP |
|
||||
| BLE | BD_ADDR | 23–517 | Reliable | Negotiated ATT_MTU |
|
||||
| Radio | Device addr | 51–222 | Unreliable | Low bandwidth, long range |
|
||||
|
||||
@@ -137,9 +157,9 @@ require connection setup before FMP can begin the Noise IK link handshake,
|
||||
adding startup latency.
|
||||
|
||||
**Stream vs. datagram**: Datagram transports have natural packet boundaries.
|
||||
Stream transports (TCP, WebSocket, Tor) require framing to delineate FIPS
|
||||
packets within the byte stream. The FMP common prefix includes a payload
|
||||
length field that provides this framing directly, replacing the need for a
|
||||
Stream transports (TCP, Tor) require framing to delineate FIPS packets
|
||||
within the byte stream. The FMP common prefix includes a payload length
|
||||
field that provides this framing directly, replacing the need for a
|
||||
separate length-prefix layer.
|
||||
|
||||
**Addressing opacity**: Transport addresses are opaque byte vectors. FMP
|
||||
@@ -169,9 +189,7 @@ proceed.
|
||||
| Transport | Connection Setup |
|
||||
| --------- | ---------------- |
|
||||
| TCP/IP | TCP three-way handshake |
|
||||
| WebSocket | HTTP upgrade + TCP |
|
||||
| Tor | Circuit establishment (500ms–5s) |
|
||||
| Bluetooth | L2CAP connection |
|
||||
| Tor | Circuit establishment (typically 10–60s, default timeout 120s) |
|
||||
| BLE | L2CAP CoC or GATT connection |
|
||||
| Serial | Physical connection (static) |
|
||||
|
||||
@@ -183,9 +201,9 @@ Connected → Disconnected. Failure can occur during connection setup, adding
|
||||
error handling paths that connectionless transports don't have.
|
||||
|
||||
**Startup latency**: Connection-oriented transports add delay before a peer
|
||||
becomes usable. This ranges from milliseconds (TCP) to seconds (Tor
|
||||
circuit). Peer timeout configuration must account for transport-specific
|
||||
setup times.
|
||||
becomes usable. This ranges from milliseconds (TCP) to tens of seconds
|
||||
(Tor circuit). Peer timeout configuration must account for
|
||||
transport-specific setup times.
|
||||
|
||||
**Framing**: Stream transports must delimit FIPS packets within the byte
|
||||
stream. The FMP common prefix includes a payload length field that provides
|
||||
@@ -209,45 +227,29 @@ NAT devices and firewalls, limiting deployment to networks without NAT.
|
||||
|
||||
### Socket Buffer Sizing
|
||||
|
||||
The default Linux UDP receive buffer (`net.core.rmem_default`, typically
|
||||
212 KB) is insufficient for high-throughput forwarding. At ~85 MB/s, a 212 KB
|
||||
buffer fills in ~2.5 ms; any stall in the async receive loop (decryption,
|
||||
routing, forwarding overhead) causes the kernel to silently drop incoming
|
||||
datagrams.
|
||||
The default Linux UDP receive buffer (`net.core.rmem_default`,
|
||||
typically 212 KB) is insufficient for high-throughput forwarding. At
|
||||
~85 MB/s, a 212 KB buffer fills in ~2.5 ms; any stall in the async
|
||||
receive loop (decryption, routing, forwarding overhead) causes the
|
||||
kernel to silently drop incoming datagrams.
|
||||
|
||||
FIPS uses `socket2::Socket` wrapped in `tokio::io::unix::AsyncFd` for the
|
||||
UDP receive path. This replaces `tokio::UdpSocket` and enables direct
|
||||
`libc::recvmsg()` calls with ancillary data parsing — specifically the
|
||||
`SO_RXQ_OVFL` socket option, which delivers a cumulative kernel receive
|
||||
buffer drop counter on every received packet. The drop counter feeds into
|
||||
the ECN congestion detection system (see
|
||||
[fips-mesh-layer.md](fips-mesh-layer.md#ecn-congestion-signaling)).
|
||||
FIPS uses `socket2::Socket` wrapped in `tokio::io::unix::AsyncFd` for
|
||||
the UDP receive path. This replaces `tokio::UdpSocket` and enables
|
||||
direct `libc::recvmsg()` calls with ancillary data parsing —
|
||||
specifically the `SO_RXQ_OVFL` socket option, which delivers a
|
||||
cumulative kernel receive buffer drop counter on every received
|
||||
packet. The drop counter feeds into the ECN congestion detection
|
||||
system (see [fips-mmp.md](fips-mmp.md#ecn-congestion-signaling)).
|
||||
|
||||
Socket buffers are configured at bind time via `socket2`:
|
||||
|
||||
| Parameter | Default | Description |
|
||||
| ---------------- | ------- | ------------------------------------ |
|
||||
| `recv_buf_size` | 2 MB | `SO_RCVBUF` — kernel receive buffer |
|
||||
| `send_buf_size` | 2 MB | `SO_SNDBUF` — kernel send buffer |
|
||||
|
||||
Linux internally doubles the requested value (to account for kernel
|
||||
bookkeeping overhead), so requesting 2 MB yields 4 MB actual buffer space.
|
||||
The kernel silently clamps to `net.core.rmem_max` if the request exceeds it.
|
||||
|
||||
**Host requirement**: `net.core.rmem_max` and `net.core.wmem_max` must be
|
||||
set to at least the requested buffer size on the host. For Docker containers,
|
||||
this must be configured on the Docker host (containers share the host kernel).
|
||||
Verify with:
|
||||
|
||||
```text
|
||||
sysctl net.core.rmem_max net.core.wmem_max
|
||||
```
|
||||
|
||||
Actual buffer sizes are logged at startup:
|
||||
|
||||
```text
|
||||
UDP transport started local_addr=0.0.0.0:2121 recv_buf=4194304 send_buf=4194304
|
||||
```
|
||||
Socket buffers (`recv_buf_size`, `send_buf_size`) are configured at
|
||||
bind time via `socket2`. Linux internally doubles the requested value
|
||||
(to account for kernel bookkeeping overhead) and silently clamps to
|
||||
`net.core.rmem_max` / `net.core.wmem_max` if the request exceeds the
|
||||
host kernel limits. The full UDP transport configuration is in
|
||||
[../reference/configuration.md](../reference/configuration.md). The
|
||||
host-side sysctl requirements and how to set them persistently live
|
||||
in
|
||||
[../how-to/tune-udp-buffers.md](../how-to/tune-udp-buffers.md).
|
||||
|
||||
## Ethernet: The Local Network Transport
|
||||
|
||||
@@ -297,18 +299,16 @@ x-only public key. Receiving nodes extract the MAC source address from the
|
||||
frame and the public key from the payload, then report the discovered peer
|
||||
to FMP.
|
||||
|
||||
Four configuration flags control discovery behavior:
|
||||
Four configuration flags control discovery behavior — `discovery`
|
||||
(listen for beacons), `announce` (broadcast beacons), `auto_connect`
|
||||
(initiate handshakes to discovered peers), and `accept_connections`
|
||||
(accept inbound handshakes). The flag table and per-flag defaults
|
||||
live in [../reference/configuration.md](../reference/configuration.md)
|
||||
under `transports.ethernet.*`.
|
||||
|
||||
| Flag | Default | Description |
|
||||
| ---- | ------- | ----------- |
|
||||
| `discovery` | true | Listen for beacons from other nodes |
|
||||
| `announce` | false | Broadcast beacons periodically |
|
||||
| `auto_connect` | false | Initiate handshakes to discovered peers |
|
||||
| `accept_connections` | false | Accept inbound handshake attempts |
|
||||
|
||||
A typical discoverable node sets `announce: true`, `auto_connect: true`, and
|
||||
`accept_connections: true`. A passive listener uses just `discovery: true` to
|
||||
observe the network without announcing itself.
|
||||
A typical discoverable node sets `announce`, `auto_connect`, and
|
||||
`accept_connections` all true. A passive listener uses just
|
||||
`discovery: true` to observe the network without announcing itself.
|
||||
|
||||
### WiFi Compatibility
|
||||
|
||||
@@ -323,11 +323,12 @@ Startup logging:
|
||||
Ethernet transport started name=eth0 interface=eth0 mac=aa:bb:cc:dd:ee:ff mtu=1499 if_mtu=1500
|
||||
```
|
||||
|
||||
## TCP/IP: Firewall Traversal Transport
|
||||
## TCP/IP: Transport for UDP-Filtered Networks
|
||||
|
||||
For networks where UDP is blocked but TCP port 443 is open, the TCP
|
||||
transport provides an alternative path. It also serves as the foundation
|
||||
for the future Tor transport.
|
||||
For peers whose networks filter outbound UDP, the TCP transport
|
||||
provides an alternative datagram path between public endpoints. TCP
|
||||
is not a NAT-traversal mechanism — there is no `tcp:nat` analogue to
|
||||
the UDP hole-punch flow.
|
||||
|
||||
FIPS protocols (FMP, FSP, MMP) are all unreliable datagrams. Running them
|
||||
over TCP introduces head-of-line blocking, which adds latency jitter. MMP
|
||||
@@ -338,16 +339,18 @@ penalizes TCP links (higher SRTT leads to higher link cost). ETX will be
|
||||
### Architecture
|
||||
|
||||
Unlike UDP (one socket serves all peers), TCP requires one `TcpStream` per
|
||||
peer. The transport maintains a connection pool (`HashMap<TransportAddr,
|
||||
TcpConnection>`) plus an optional `TcpListener` for inbound connections.
|
||||
peer. The transport maintains two pools: a `ConnectingPool` for background
|
||||
connection attempts in progress, and an established connection pool
|
||||
(`HashMap<TransportAddr, TcpConnection>`) for active connections, plus an
|
||||
optional `TcpListener` for inbound connections.
|
||||
|
||||
| Property | Value |
|
||||
| -------- | ----- |
|
||||
| Addressing | IP:port (same as UDP) |
|
||||
| Addressing | host:port — IP address or DNS hostname |
|
||||
| Default MTU | 1400 bytes |
|
||||
| Per-link MTU | Derived from `TCP_MAXSEG` socket option |
|
||||
| Framing | FMP header-based (zero overhead) |
|
||||
| Connection model | Connect-on-send, optional listener |
|
||||
| Connection model | Non-blocking connect, connect-on-send fallback, optional listener |
|
||||
| Platform | Cross-platform (no `#[cfg]` gates) |
|
||||
|
||||
### FMP Header-Based Framing
|
||||
@@ -364,29 +367,36 @@ packet boundaries:
|
||||
|
||||
This provides zero framing overhead and built-in phase validation. The
|
||||
stream reader is implemented in a separate module (`stream.rs`) for reuse
|
||||
by the future Tor transport.
|
||||
by the Tor transport.
|
||||
|
||||
### Connect-on-Send
|
||||
### Connection Establishment
|
||||
|
||||
When `send(addr, data)` is called with no existing connection:
|
||||
TCP connections use a non-blocking connect model. When FMP needs to reach
|
||||
a configured peer address, the node calls `connect(addr)` on the transport,
|
||||
which spawns a background tokio task to perform the TCP handshake and socket
|
||||
configuration (TCP_NODELAY, keepalive, buffer sizes, TCP_MAXSEG query). The
|
||||
call returns immediately without blocking the event loop.
|
||||
|
||||
1. Connect with configurable timeout (default 5s)
|
||||
2. Configure socket: `TCP_NODELAY`, keepalive, buffer sizes
|
||||
3. Read `TCP_MAXSEG` for per-connection MTU
|
||||
4. Split stream into read/write halves
|
||||
5. Spawn per-connection receive task
|
||||
6. Store connection in pool
|
||||
7. Write packet directly to stream
|
||||
The node tracks each pending connection in a `PendingConnect` entry. On
|
||||
every tick, `poll_pending_connects()` calls `connection_state(addr)` to
|
||||
check progress. When the transport reports `Connected`, the completed
|
||||
connection is promoted to the established pool (stream split into
|
||||
read/write halves, per-connection receive task spawned), and the node
|
||||
initiates the Noise IK link handshake. If the transport reports `Failed`,
|
||||
the node schedules a retry with exponential backoff.
|
||||
|
||||
If connect fails, return error. The node's handshake retry mechanism
|
||||
handles re-attempts.
|
||||
As a fallback, `send(addr, data)` still performs synchronous
|
||||
connect-on-send if no connection exists — this handles the case where a
|
||||
send arrives before the node-level connect path runs. The non-blocking
|
||||
path is the primary mechanism for configured peers.
|
||||
|
||||
### Session Independence
|
||||
|
||||
TCP connection loss does **not** tear down the FIPS peer. Noise keys, MMP
|
||||
state, and FSP sessions are bound to the peer's npub, not the TCP
|
||||
connection. The transport reconnects transparently on the next send via
|
||||
connect-on-send. MMP liveness timeout is the sole authority for peer death.
|
||||
connection. The transport reconnects transparently via the non-blocking
|
||||
connect path or connect-on-send fallback. MMP liveness timeout is the sole
|
||||
authority for peer death.
|
||||
|
||||
### Connection Deduplication
|
||||
|
||||
@@ -397,23 +407,197 @@ removes it from the pool and aborts its receive task.
|
||||
|
||||
### Configuration
|
||||
|
||||
```yaml
|
||||
transports:
|
||||
tcp:
|
||||
bind_addr: "0.0.0.0:8443" # Listen address (omit for outbound-only)
|
||||
mtu: 1400 # Default MTU
|
||||
connect_timeout_ms: 5000 # Outbound connect timeout
|
||||
nodelay: true # TCP_NODELAY (disable Nagle)
|
||||
keepalive_secs: 30 # TCP keepalive interval (0 = disabled)
|
||||
recv_buf_size: 2097152 # SO_RCVBUF (2 MB)
|
||||
send_buf_size: 2097152 # SO_SNDBUF (2 MB)
|
||||
max_inbound_connections: 256 # Resource protection limit
|
||||
socks5_proxy: "127.0.0.1:9050" # SOCKS5 for outbound (deferred)
|
||||
The TCP transport configuration block (`transports.tcp.*` — bind
|
||||
address, MTU, connect timeout, TCP_NODELAY, keepalive, socket buffer
|
||||
sizes, max inbound connections) is documented in
|
||||
[../reference/configuration.md](../reference/configuration.md). If
|
||||
`bind_addr` is configured, the transport accepts inbound connections;
|
||||
without it, the transport operates in outbound-only mode (no listener
|
||||
socket is created).
|
||||
|
||||
## Tor: The Anonymity Transport
|
||||
|
||||
The Tor transport routes FIPS traffic through the Tor network, hiding
|
||||
a node's IP address from its peers. A node behind Tor connects outbound
|
||||
through a local Tor SOCKS5 proxy; the remote peer sees the Tor exit
|
||||
node's IP, not the initiator's. After the Noise IK handshake, the remote
|
||||
peer knows the initiator's FIPS identity (npub) but not its network
|
||||
location.
|
||||
|
||||
Like TCP, Tor is connection-oriented and reliable. The same TCP-over-TCP
|
||||
considerations apply — MMP correctly measures the elevated latency and
|
||||
cost-based parent selection naturally deprioritizes Tor links.
|
||||
|
||||
### Architecture
|
||||
|
||||
The Tor transport is a separate `TorTransport` implementation, not a TCP
|
||||
variant, because it manages SOCKS5 proxy negotiation, has different
|
||||
address semantics (.onion vs IP:port), and has significantly different
|
||||
latency characteristics. It reuses the FMP header-based stream reader
|
||||
(`tcp/stream.rs`) for packet framing on the underlying TCP connection.
|
||||
|
||||
The transport maintains two pools (same pattern as TCP): a
|
||||
`ConnectingPool` for background SOCKS5 connection attempts, and an
|
||||
established pool of `TorConnection` entries. Each `TorConnection` holds
|
||||
a write half, a per-connection receive task, the negotiated MTU, and
|
||||
a connection timestamp.
|
||||
|
||||
| Property | Value |
|
||||
| -------- | ----- |
|
||||
| Addressing | .onion:port or IP:port |
|
||||
| Default MTU | 1400 bytes |
|
||||
| Framing | FMP header-based (shared with TCP) |
|
||||
| Connection model | Non-blocking connect, outbound SOCKS5 + inbound via onion service |
|
||||
| Platform | Cross-platform (requires external Tor daemon) |
|
||||
|
||||
### Address Types
|
||||
|
||||
The Tor transport accepts three address formats, parsed into a `TorAddr`
|
||||
enum:
|
||||
|
||||
- **Onion**: `.onion:port` — connects to a Tor hidden service. Both
|
||||
sides anonymous. (e.g., `abcdef...xyz.onion:8443`)
|
||||
- **Clearnet IP**: `IP:port` — connects through a Tor exit node to a
|
||||
remote TCP listener. Hides the initiator's IP; the remote peer sees
|
||||
the exit node's IP.
|
||||
- **Clearnet Hostname**: `hostname:port` — hostname is passed through
|
||||
SOCKS5 for Tor-side DNS resolution, avoiding local DNS leaks. Compatible
|
||||
with SafeSocks 1. (e.g., `fips.example.com:8443`)
|
||||
|
||||
All address types are routed through the same SOCKS5 proxy.
|
||||
|
||||
### Connection Establishment
|
||||
|
||||
Connection setup follows the same non-blocking pattern as TCP. When FMP
|
||||
needs to reach a peer, the node calls `connect(addr)` on the transport.
|
||||
The transport spawns a background tokio task that:
|
||||
|
||||
1. Opens a SOCKS5 connection through the local Tor proxy
|
||||
2. Configures the socket: `TCP_NODELAY`, keepalive (30s)
|
||||
3. Returns the connected stream
|
||||
|
||||
The call returns immediately. `connection_state(addr)` reports progress.
|
||||
Tor circuit establishment typically takes 10–60 seconds (vs milliseconds
|
||||
for TCP), making non-blocking connect essential — a blocking connect
|
||||
would stall the entire FMP event loop.
|
||||
|
||||
The connect timeout defaults to 120 seconds (vs 5 seconds for TCP),
|
||||
accounting for Tor circuit setup time. As a fallback, `send(addr, data)`
|
||||
performs synchronous connect-on-send if no connection exists.
|
||||
|
||||
### Inbound via Onion Service (Directory Mode)
|
||||
|
||||
In `directory` mode (recommended for production), Tor manages the onion
|
||||
service via `HiddenServiceDir` in `torrc`. FIPS reads the `.onion` address
|
||||
from the hostname file at startup and binds a local TCP listener that the
|
||||
Tor daemon forwards inbound connections to.
|
||||
|
||||
This mode enables Tor's `Sandbox 1` (seccomp-bpf) — the strongest single
|
||||
hardening option — because no control port interaction is required for
|
||||
onion service management. Tor handles key generation and persistence
|
||||
directly through the `HiddenServiceDir`.
|
||||
|
||||
The inbound accept loop mirrors the TCP transport's pattern: accept
|
||||
connection, configure socket (TCP_NODELAY, keepalive), spawn a
|
||||
per-connection receive loop using the shared FMP stream reader. Inbound
|
||||
connections arrive from `127.0.0.1` (Tor daemon's local forwarding); peer
|
||||
identity is resolved during the Noise IK handshake, not from the transport
|
||||
address.
|
||||
|
||||
Configuration requires coordinating `torrc` and `fips.yaml`. The
|
||||
operator setup — torrc directives, `fips.yaml` `tor` section,
|
||||
HiddenServiceDir permissions, and `Sandbox 1` notes — is in
|
||||
[../how-to/deploy-tor-onion.md](../how-to/deploy-tor-onion.md). In
|
||||
brief: the `HiddenServicePort` external port is what peers connect
|
||||
to, and `tor.directory_service.bind_addr` must match the
|
||||
`HiddenServicePort` target address.
|
||||
|
||||
### Session Independence
|
||||
|
||||
Same as TCP: Tor connection loss does **not** tear down the FIPS peer.
|
||||
Noise keys, MMP state, and FSP sessions survive reconnection.
|
||||
|
||||
### Bridge Node Pattern
|
||||
|
||||
A node running both Tor and UDP transports acts as a bridge between
|
||||
anonymous and clearnet portions of the mesh:
|
||||
|
||||
```text
|
||||
[Anonymous node] --tor--> [Bridge node] --udp--> [Clearnet node]
|
||||
```
|
||||
|
||||
If `bind_addr` is configured, the transport accepts inbound connections.
|
||||
Without it, the transport operates in outbound-only mode (no listener
|
||||
socket is created).
|
||||
No special code is needed — FIPS multi-transport routing handles it.
|
||||
Anonymous nodes connect to the bridge via Tor; the bridge forwards
|
||||
traffic to clearnet peers over UDP. Clearnet peers never see the
|
||||
anonymous node's IP.
|
||||
|
||||
### Latency Characteristics
|
||||
|
||||
Tor adds 200ms–2s RTT per circuit. MMP measures this elevated latency,
|
||||
and cost-based parent selection penalizes Tor links (high SRTT → high
|
||||
link cost). ETX is 1.0 since TCP handles retransmission.
|
||||
|
||||
Tor throughput is typically 1–5 Mbps — adequate for control plane and
|
||||
moderate data transfer, not for bulk transfer.
|
||||
|
||||
### Monitoring
|
||||
|
||||
In `control_port` mode and optionally in `directory` mode (when
|
||||
`control_addr` is configured), the transport spawns a background
|
||||
monitoring task that polls the Tor daemon every 10 seconds via the
|
||||
control port. The cached monitoring data is exposed through the
|
||||
`show_transports` control socket query and displayed in fipstop.
|
||||
|
||||
Monitoring data includes:
|
||||
|
||||
- **Bootstrap progress** (0–100%) with INFO logging at milestones
|
||||
(25/50/75/100%) and WARN if stalled >60s
|
||||
- **Circuit status** (whether Tor has a working circuit)
|
||||
- **Network liveness** (up/down) with WARN on transitions
|
||||
- **Dormant mode** detection with WARN on entry
|
||||
- **Tor daemon version** and **traffic counters** (bytes read/written)
|
||||
|
||||
The control port connection uses cookie authentication by default
|
||||
(reading from `/var/run/tor/control.authcookie`). Unix socket
|
||||
connections (`/run/tor/control`) are preferred over TCP for security.
|
||||
|
||||
### Configuration
|
||||
|
||||
The Tor transport block (`transports.tor.*`) is documented in
|
||||
[../reference/configuration.md](../reference/configuration.md). Three
|
||||
modes are available:
|
||||
|
||||
- **`socks5`** (default): Outbound-only through a SOCKS5 proxy. No
|
||||
control port, no inbound connections.
|
||||
- **`control_port`**: Outbound via SOCKS5 plus control port connection
|
||||
for Tor daemon monitoring. No inbound connections.
|
||||
- **`directory`** (recommended for inbound): Outbound via SOCKS5 plus
|
||||
inbound via Tor-managed `HiddenServiceDir` onion service.
|
||||
Optionally connects to the control port for monitoring when
|
||||
`control_addr` is set. Enables Tor's `Sandbox 1` for maximum
|
||||
security.
|
||||
|
||||
The Tor transport requires an external Tor daemon. Named instances
|
||||
are supported for multiple proxy endpoints.
|
||||
|
||||
### Implementation Roadmap
|
||||
|
||||
- Outbound SOCKS5 connections to .onion, clearnet IP, and clearnet
|
||||
hostname addresses *(implemented)*
|
||||
- Inbound connections via Tor onion service using `HiddenServiceDir`
|
||||
directory mode *(implemented)*
|
||||
- Operator visibility: cached monitoring snapshot, control socket
|
||||
exposure, fipstop display, bootstrap/liveness logging *(implemented)*
|
||||
- Embedded `arti` (Rust Tor implementation) for self-contained operation
|
||||
without an external Tor daemon *(future)*
|
||||
|
||||
### Statistics
|
||||
|
||||
The Tor transport exposes per-instance counters covering successful
|
||||
send/receive, send/receive errors, connection establishment,
|
||||
SOCKS5-level errors, MTU rejections, accepted/rejected inbound
|
||||
connections, and Tor control-port errors. The full counter table
|
||||
lives in [../reference/transports.md](../reference/transports.md).
|
||||
|
||||
## Discovery
|
||||
|
||||
@@ -423,7 +607,7 @@ detection — a new TCP connection or UDP packet from an unknown source is not
|
||||
discovery; a FIPS-specific announcement or response is.
|
||||
|
||||
Discovery is an optional transport capability. Transports that don't support
|
||||
it (configured UDP endpoints, TCP) simply don't provide discovery events.
|
||||
it (configured UDP endpoints, TCP, Tor) simply don't provide discovery events.
|
||||
FMP handles both cases uniformly: with discovery, it waits for events then
|
||||
initiates link setup; without discovery, it initiates link setup directly to
|
||||
configured addresses.
|
||||
@@ -449,11 +633,11 @@ X." FMP does not need to distinguish beacons from query responses.
|
||||
| Radio | Beacon | Shared RF channel, natural fit |
|
||||
| BLE | Advertising | GATT service UUID |
|
||||
|
||||
### Nostr Relay Discovery *(future direction)*
|
||||
### Nostr Relay Discovery
|
||||
|
||||
For internet-reachable transports, a node publishes a signed Nostr event
|
||||
containing its FIPS discovery information — public key and reachable
|
||||
transport endpoints (UDP IP:port, TCP IP:port, .onion address). Other FIPS
|
||||
transport endpoints (UDP host:port, TCP host:port, .onion address). Other FIPS
|
||||
nodes subscribing on the same relays learn about available peers.
|
||||
|
||||
Nostr relay discovery is not a transport — it is a discovery service that
|
||||
@@ -461,6 +645,12 @@ feeds addresses to other transports. A node discovers via Nostr that a peer
|
||||
is reachable at UDP 1.2.3.4:9735, then establishes the link over the UDP
|
||||
transport.
|
||||
|
||||
For NAT'd UDP endpoints, a node may advertise `addr: "nat"` instead of a
|
||||
concrete address, signaling that peers should initiate STUN-assisted UDP
|
||||
hole punching. Offer/answer exchange uses Nostr gift-wrap (NIP-59) events
|
||||
on the configured DM relays; the resulting punched socket is adopted into
|
||||
the standard UDP transport via the bootstrap handoff path.
|
||||
|
||||
Key properties:
|
||||
|
||||
- Identity is built in — Nostr events are signed, so discovery information
|
||||
@@ -472,13 +662,16 @@ Key properties:
|
||||
|
||||
### Current State
|
||||
|
||||
> **Implemented**: UDP and TCP peers are configured via YAML. Ethernet
|
||||
> peers are discovered via beacon broadcast — the `discover()` trait
|
||||
> method returns newly seen endpoints, and per-transport `auto_connect()`
|
||||
> / `accept_connections()` policies control whether discovered peers are
|
||||
> connected automatically or require explicit configuration. TCP has no
|
||||
> discovery mechanism (peers are configured). Nostr relay discovery is
|
||||
> not yet implemented.
|
||||
> **Implemented**: UDP, TCP, Tor, and Ethernet peers can be configured
|
||||
> statically via YAML. Ethernet peers can also be discovered via beacon
|
||||
> broadcast — the `discover()` trait method returns newly seen endpoints,
|
||||
> and per-transport `auto_connect()` / `accept_connections()` policies
|
||||
> control whether discovered peers are connected automatically or require
|
||||
> explicit configuration. TCP and Tor have no built-in discovery mechanism.
|
||||
> Nostr relay discovery and STUN-assisted UDP hole punching are
|
||||
> implemented and toggled via configuration; see
|
||||
> [../reference/configuration.md](../reference/configuration.md) for the
|
||||
> `node.discovery.nostr.*` configuration tree.
|
||||
|
||||
## Transport Interface
|
||||
|
||||
@@ -496,6 +689,8 @@ link_mtu(addr) → u16 Per-link MTU (defaults to mtu())
|
||||
start() → lifecycle Bring transport up (bind socket, open device)
|
||||
stop() → lifecycle Bring transport down
|
||||
send(addr, data) → delivery Send datagram to transport address
|
||||
connect(addr) → () Initiate non-blocking connection (connection-oriented only)
|
||||
connection_state(addr)→ ConnectionState Poll connection status (None/Connecting/Connected/Failed)
|
||||
close_connection(addr)→ () Close a specific connection (no-op for connectionless)
|
||||
congestion() → TransportCongestion Local congestion indicators (optional)
|
||||
discover() → Vec<DiscoveredPeer> Report discovered FIPS endpoints (optional)
|
||||
@@ -553,14 +748,16 @@ on all forwarded datagrams.
|
||||
| Transport | Congestion Source | Mechanism |
|
||||
| --------- | ----------------- | --------- |
|
||||
| UDP | `SO_RXQ_OVFL` kernel drop counter | `recvmsg()` ancillary data on every packet |
|
||||
| TCP | Not yet implemented | Returns `None` (TCP handles congestion internally) |
|
||||
| Ethernet | Not yet implemented | Returns `None` |
|
||||
| TCP | Not implemented | Returns `None` (TCP handles congestion internally) |
|
||||
| Tor | Not implemented | Returns `None` (TCP handles congestion internally) |
|
||||
| Ethernet | Not implemented | Returns `None` |
|
||||
|
||||
### Transport Addresses
|
||||
|
||||
Transport addresses (`TransportAddr`) are opaque byte vectors. The transport
|
||||
layer interprets them (e.g., UDP parses "ip:port" strings); all layers above
|
||||
treat them as opaque handles passed back to the transport for sending.
|
||||
layer interprets them — e.g. UDP and TCP resolve `host:port` strings (IP
|
||||
fast path, DNS fallback with a 60s cache on UDP). All layers above treat
|
||||
them as opaque handles passed back to the transport for sending.
|
||||
|
||||
### Transport State Machine
|
||||
|
||||
@@ -579,11 +776,11 @@ transitions through `Starting` to `Up` (operational). `stop()` moves to
|
||||
| Transport | Status | Notes |
|
||||
| --------- | ------ | ----- |
|
||||
| UDP/IP | **Implemented** | Primary transport, AsyncFd/recvmsg, SO_RXQ_OVFL kernel drop detection |
|
||||
| TCP/IP | **Implemented** | FMP header-based framing, connect-on-send, per-connection MSS MTU |
|
||||
| TCP/IP | **Implemented** | FMP header-based framing, non-blocking connect, per-connection MSS MTU |
|
||||
| Ethernet | **Implemented** | AF_PACKET SOCK_DGRAM, EtherType 0x2121, beacon discovery, Linux only |
|
||||
| WiFi | Future direction | Infrastructure mode = Ethernet driver |
|
||||
| Tor | Future direction | High latency, .onion addressing |
|
||||
| BLE | Future direction | ATT_MTU negotiation, per-link MTU |
|
||||
| WiFi | **Implemented** (via Ethernet transport, infrastructure mode) | mac80211 translates 802.11↔802.3; broadcast beacons unreliable through APs |
|
||||
| Tor | **Implemented** | Outbound SOCKS5, inbound via onion service, .onion and clearnet addressing |
|
||||
| BLE | **Implemented** (Linux/glibc only; experimental) | L2CAP CoC, ATT_MTU negotiation, per-link MTU; musl/macOS/Windows skip |
|
||||
| Radio | Future direction | Constrained MTU (51–222 bytes) |
|
||||
| Serial | Future direction | SLIP/COBS framing, point-to-point |
|
||||
|
||||
@@ -591,7 +788,7 @@ transitions through `Starting` to `Up` (operational). `stop()` moves to
|
||||
|
||||
### TCP-over-TCP Avoidance
|
||||
|
||||
Running TCP application traffic over a reliable transport (TCP, WebSocket)
|
||||
Running TCP application traffic over a reliable transport (TCP, Tor)
|
||||
creates a layering violation where retransmission and congestion control
|
||||
operate at both levels. When the inner TCP detects loss (which may just be
|
||||
transport-layer retransmission delay), it retransmits, creating more traffic
|
||||
@@ -626,6 +823,15 @@ quality difference is significant. Link cost is not yet used in
|
||||
|
||||
## References
|
||||
|
||||
- [fips-intro.md](fips-intro.md) — Protocol overview and layer architecture
|
||||
- [fips-mesh-layer.md](fips-mesh-layer.md) — FMP specification (the layer above)
|
||||
- [fips-wire-formats.md](fips-wire-formats.md) — Transport framing details
|
||||
- [fips-concepts.md](fips-concepts.md) — Protocol overview
|
||||
- [fips-architecture.md](fips-architecture.md) — Layer architecture
|
||||
- [fips-mesh-layer.md](fips-mesh-layer.md) — FMP specification (the
|
||||
layer above)
|
||||
- [fips-mtu.md](fips-mtu.md) — How transport-reported `link_mtu`
|
||||
feeds the unified path-MTU model
|
||||
- [../reference/wire-formats.md](../reference/wire-formats.md) —
|
||||
Transport framing details
|
||||
- [../reference/configuration.md](../reference/configuration.md) —
|
||||
Per-transport configuration blocks
|
||||
- [../reference/transports.md](../reference/transports.md) —
|
||||
Per-transport statistics counter inventory
|
||||
|
||||
@@ -0,0 +1,594 @@
|
||||
# Port Advertisement and NAT Traversal via Nostr
|
||||
|
||||
## Abstract
|
||||
|
||||
This document describes two related-but-independent mechanisms that an
|
||||
application protocol can build on top of Nostr relays:
|
||||
|
||||
1. **Port advertisement.** A node publishes a parameterized replaceable
|
||||
event describing the application protocol it speaks, the version, and
|
||||
the endpoint(s) at which it can be reached. Other nodes discover the
|
||||
advert by querying relays.
|
||||
2. **NAT traversal.** When the advertised endpoint indicates that the
|
||||
responder is behind NAT, the two peers exchange ephemeral
|
||||
gift-wrapped offer/answer events through Nostr relays, run STUN
|
||||
against a public server to learn their reflexive addresses, and
|
||||
coordinate UDP hole punching so they can exchange application traffic
|
||||
over a direct UDP path.
|
||||
|
||||
The two mechanisms compose naturally — an advert that includes a
|
||||
`<protocol>:nat` endpoint signals "reach me by running the traversal
|
||||
protocol" — but they are independently useful. An advert with only
|
||||
public-IP endpoints needs no traversal. A pair of peers that already
|
||||
know each other's pubkeys but want to coordinate a traversal can do so
|
||||
without ever publishing a public advert.
|
||||
|
||||
The protocol described here is generic. Any application protocol can
|
||||
adopt it by picking its own kind number, `d`-tag scope, and endpoint
|
||||
schema. [FIPS](https://github.com/jmcorgan/fips) (the Free
|
||||
Internetworking Peering System) is used throughout the document as an
|
||||
example implementation; FIPS-specific values appear in clearly marked
|
||||
example blocks and do not affect the generic protocol shape.
|
||||
|
||||
No WebRTC, DTLS, or ICE stack is required. The protocol operates at
|
||||
the raw UDP level, using Nostr solely for ephemeral signaling and STUN
|
||||
solely for reflexive address discovery.
|
||||
|
||||
---
|
||||
|
||||
## Terminology
|
||||
|
||||
- **Application protocol.** The protocol that runs on top of the
|
||||
punched UDP channel after this document's procedures complete.
|
||||
- **Initiator.** The peer that discovers the responder's advert and
|
||||
begins the traversal exchange.
|
||||
- **Responder.** The peer that publishes a service advertisement and
|
||||
is willing to be dialled.
|
||||
- **Reflexive address.** The public `IP:port` tuple that a STUN server
|
||||
observes for a UDP socket — i.e., the NAT's external mapping for
|
||||
that socket.
|
||||
- **Punch socket.** The single UDP socket a peer uses for STUN, for
|
||||
the offer/answer exchange's address fields, for the punch packets
|
||||
themselves, and for the application traffic that follows. The same
|
||||
socket must be used across all phases of one traversal attempt.
|
||||
|
||||
### Socket lifecycle
|
||||
|
||||
The protocol assumes **per-peer, per-attempt punch sockets**:
|
||||
|
||||
- Each outbound traversal attempt allocates a fresh UDP socket bound
|
||||
to `0.0.0.0:0` (OS-assigned port).
|
||||
- That socket is owned by exactly one remote peer and exactly one
|
||||
traversal session.
|
||||
- STUN, the offer/answer reflexive-address fields, the punch packets,
|
||||
and the eventual adopted application transport all share that
|
||||
socket for the lifetime of the attempt.
|
||||
- If the attempt fails, the socket is discarded. A retry allocates a
|
||||
new socket and obtains a fresh reflexive address.
|
||||
- A long-lived application listener (for example, a fixed UDP port
|
||||
shared across peers) must **not** be reused as the punch socket —
|
||||
doing so couples NAT mappings and retry state across peers.
|
||||
|
||||
This rule is not optional: closing or rebinding the socket between
|
||||
phases invalidates the NAT mapping that the rest of the protocol
|
||||
depends on.
|
||||
|
||||
---
|
||||
|
||||
## Part 1: Service Advertisement
|
||||
|
||||
### Event shape
|
||||
|
||||
The advert is a NIP-01 parameterized replaceable event whose kind
|
||||
falls in the application-defined replaceable range
|
||||
`30000–39999`. The event carries:
|
||||
|
||||
- A `d` tag scoping the advert (so the same pubkey can publish
|
||||
multiple distinct adverts under different scopes).
|
||||
- A `protocol` tag carrying the application protocol's name, used as
|
||||
a discovery filter for peers that don't already know the
|
||||
responder's pubkey.
|
||||
- A `version` tag carrying the application protocol version.
|
||||
- An optional `expiration` tag (NIP-40) so a relay garbage-collects
|
||||
the advert when the responder goes offline without explicitly
|
||||
deleting it.
|
||||
- An optional `relays` tag listing relays where the responder
|
||||
subscribes for incoming signaling messages (used by Part 2).
|
||||
- An optional `stun` tag listing STUN servers the responder
|
||||
recommends.
|
||||
- A `content` field carrying the application-specific payload —
|
||||
typically the endpoint set, capability flags, and any encryption
|
||||
keys the application layer needs. The content may be plaintext or
|
||||
NIP-44-encrypted; encryption requires the consumer to already know
|
||||
the responder's pubkey.
|
||||
|
||||
The replaceable semantics let the responder update the advert in
|
||||
place under the same `d` tag. A NIP-09 deletion event removes the
|
||||
advert when the responder permanently retires.
|
||||
|
||||
```json
|
||||
{
|
||||
"kind": <application-specific>,
|
||||
"pubkey": "<responder_pubkey>",
|
||||
"created_at": <unix_seconds>,
|
||||
"tags": [
|
||||
["d", "<application-defined-scope>"],
|
||||
["protocol", "<application_protocol_name>"],
|
||||
["version", "<protocol_version>"],
|
||||
["relays", "wss://relay1.example.com", "wss://relay2.example.com"],
|
||||
["stun", "stun.l.google.com:19302"],
|
||||
["expiration", "<unix_seconds + ttl>"]
|
||||
],
|
||||
"content": "<application payload, optionally NIP-44 encrypted>",
|
||||
"sig": "<signature>"
|
||||
}
|
||||
```
|
||||
|
||||
### Endpoint schema
|
||||
|
||||
The `content` field is application-defined. Its structure typically
|
||||
includes a list of endpoints describing how the responder can be
|
||||
reached. Endpoint entries should distinguish:
|
||||
|
||||
- **Direct public endpoints** (transport + address + port) where any
|
||||
initiator can connect without traversal.
|
||||
- **NAT-mapped endpoints** that signal "I can be reached by running
|
||||
the traversal protocol against this transport on my pubkey."
|
||||
- **Anonymity-network endpoints** (e.g. Tor onion services) where
|
||||
the addressing scheme implies its own connection semantics.
|
||||
|
||||
#### FIPS example: kind 37195 advertisement
|
||||
|
||||
FIPS uses **kind `37195`** (the digits visually spell `FIPS` —
|
||||
7=F, 1=I, 9=P, 5=S). The `d` tag is hardcoded to
|
||||
`fips-overlay-v1`; the configurable `app` value populates the
|
||||
separate `protocol` tag, scoping adverts within a relay set
|
||||
without splitting them across multiple `d`-tag streams.
|
||||
|
||||
The advert content is a JSON document carrying a list of endpoint
|
||||
entries, each shaped as `{transport, addr}`. The transport string
|
||||
takes one of:
|
||||
|
||||
- `udp:host:port` — direct public UDP endpoint.
|
||||
- `udp:nat` — NAT-mapped UDP endpoint; reach via Part 2 traversal.
|
||||
- `tcp:host:port` — direct public TCP endpoint, for peers whose
|
||||
networks filter outbound UDP. Public-only; there is no
|
||||
`tcp:nat` analogue.
|
||||
- `tor:<onion>:<port>` — Tor onion-service endpoint.
|
||||
|
||||
FIPS publishes the advert with `expiration` set to `now +
|
||||
advert_ttl_secs` (default 1 hour) and refreshes it every
|
||||
`advert_refresh_secs` (default 30 minutes).
|
||||
|
||||
### Public-IP discovery on advertisement
|
||||
|
||||
A responder behind a NAT or wildcard-bound to a non-routable address
|
||||
needs to determine what external address to put in its advert. The
|
||||
responder uses a fixed precedence:
|
||||
|
||||
1. An operator-supplied external address override (FIPS:
|
||||
`transports.{udp,tcp}.external_addr`) wins.
|
||||
2. A non-wildcard `local_addr` is used directly.
|
||||
3. For a wildcard-bound UDP listener with an explicit "publish this"
|
||||
flag (FIPS: `public: true`), the runtime queries STUN against
|
||||
the configured servers and publishes the reflexive address.
|
||||
4. For a wildcard-bound TCP listener, no STUN equivalent exists.
|
||||
Implementations should refuse to silently advertise an unreachable
|
||||
endpoint; FIPS emits a loud WARN and omits the endpoint.
|
||||
|
||||
This precedence keeps adverts honest: an endpoint that appears in
|
||||
the published content is one the responder believes is reachable.
|
||||
|
||||
### Discovery (consumer side)
|
||||
|
||||
A consumer queries one or more relays for an advert it can act on.
|
||||
Two filter shapes are typical:
|
||||
|
||||
By author, when the responder's pubkey is already known:
|
||||
|
||||
```json
|
||||
["REQ", "<sub_id>", {
|
||||
"kinds": [<advert_kind>],
|
||||
"authors": ["<responder_pubkey>"],
|
||||
"#d": ["<application-defined-scope>"]
|
||||
}]
|
||||
```
|
||||
|
||||
By application protocol, for "open discovery" of any peer running
|
||||
the same application:
|
||||
|
||||
```json
|
||||
["REQ", "<sub_id>", {
|
||||
"kinds": [<advert_kind>],
|
||||
"#protocol": ["<application_protocol_name>"]
|
||||
}]
|
||||
```
|
||||
|
||||
Adverts whose `protocol` tag does not match the consumer's expected
|
||||
value, or whose `expiration` tag has elapsed, are rejected at
|
||||
validation. Consumers cache adverts in memory keyed by author npub
|
||||
and respect the embedded expiration.
|
||||
|
||||
#### FIPS example: discovery filters
|
||||
|
||||
The FIPS daemon issues both filter shapes: by-author for peers it
|
||||
intends to dial directly, and by-`#protocol` when an operator has
|
||||
opted into open discovery against the same application namespace.
|
||||
Cached adverts persist until their `expiration` lapses; a periodic
|
||||
prune drops expired entries.
|
||||
|
||||
---
|
||||
|
||||
## Part 2: NAT Traversal
|
||||
|
||||
The traversal protocol coordinates UDP hole punching between two
|
||||
peers via gift-wrapped Nostr signaling. It is invoked when the
|
||||
initiator decides to dial a NAT-mapped endpoint advertised by the
|
||||
responder.
|
||||
|
||||
### Signaling event shape
|
||||
|
||||
Signaling messages are ephemeral kinds in the range `20000–29999`,
|
||||
NIP-44-encrypted to the recipient, and NIP-59 gift-wrapped so the
|
||||
outer event is signed by an ephemeral keypair rather than the
|
||||
sender's long-term identity. The wrap carries a `p` tag pointing at
|
||||
the recipient's pubkey and an NIP-40 `expiration` tag bounding how
|
||||
long the relay should retain it.
|
||||
|
||||
#### FIPS example: signaling kind 21059
|
||||
|
||||
FIPS signaling uses **kind `21059`**. Wraps are addressed by `p`
|
||||
tag and published to the responder's NIP-17 inbox relay list (kind
|
||||
`10050`) when one is available, falling back to the local
|
||||
`dm_relays` configuration otherwise. Each side publishes its own
|
||||
inbox relay list on startup so dialers can discover it.
|
||||
|
||||
### Phase 1: Initiator STUN binding
|
||||
|
||||
Before constructing any signaling message, the initiator:
|
||||
|
||||
1. Allocates a fresh UDP punch socket bound to `0.0.0.0:0`.
|
||||
2. Sends a STUN Binding Request (RFC 8489) to one of its locally
|
||||
configured STUN servers.
|
||||
3. Parses the Binding Response, extracts the
|
||||
`XOR-MAPPED-ADDRESS` attribute, and records that as its
|
||||
reflexive address. Other STUN attributes are ignored.
|
||||
4. Records local-candidate addresses for the same socket port:
|
||||
active private non-loopback interface addresses (RFC1918 IPv4,
|
||||
IPv6 ULA) and probed local egress addresses.
|
||||
|
||||
The punch socket must remain open across all subsequent phases.
|
||||
Closing or rebinding it discards the NAT mapping.
|
||||
|
||||
### Phase 2: Initiator sends offer
|
||||
|
||||
The initiator constructs an offer payload containing its reflexive
|
||||
address, its local-candidate addresses, an opaque session
|
||||
identifier, freshness timestamps, and any application-specific
|
||||
parameters. The payload is NIP-44-encrypted to the responder's
|
||||
pubkey, wrapped with NIP-59, and published to the responder's
|
||||
signaling relays. The initiator also subscribes by `p` tag on
|
||||
those relays to receive the answer.
|
||||
|
||||
```json
|
||||
{
|
||||
"type": "offer",
|
||||
"sessionId": "<random_hex_32>",
|
||||
"issuedAt": <unix_millis>,
|
||||
"expiresAt": <unix_millis>,
|
||||
"nonce": "<random_nonce>",
|
||||
"senderNpub": "<initiator_npub>",
|
||||
"recipientNpub": "<responder_npub>",
|
||||
"reflexiveAddress": {"protocol":"udp","ip":"<ip>","port":<port>},
|
||||
"localAddresses": [{"protocol":"udp","ip":"<ip>","port":<port>}],
|
||||
"stunServer": "<host>:<port>",
|
||||
"app_params": { ... }
|
||||
}
|
||||
```
|
||||
|
||||
- `sessionId` is a random identifier correlating offer and answer.
|
||||
- `reflexiveAddress` is the address STUN observed in Phase 1.
|
||||
- `localAddresses` enables a same-LAN fast path when both peers
|
||||
happen to share a private subnet.
|
||||
- `stunServer` is informational, recording which server the
|
||||
initiator used.
|
||||
- `issuedAt` / `expiresAt` bound the freshness window — the
|
||||
responder rejects stale offers, since a NAT mapping that has not
|
||||
been refreshed in tens of seconds may already be gone.
|
||||
|
||||
### Phase 3: Responder validates and answers
|
||||
|
||||
The responder maintains a standing `p`-tagged subscription on its
|
||||
advertised signaling relays. On receiving an offer:
|
||||
|
||||
1. Decrypts the wrap and recovers the offer payload.
|
||||
2. Validates freshness (rejects if outside the configured window;
|
||||
see *Skew tolerance* below).
|
||||
3. Rejects replays — if the `sessionId` is in a recently-seen
|
||||
cache, drop the offer.
|
||||
4. Allocates its own punch socket (`0.0.0.0:0`) and runs its own
|
||||
STUN query.
|
||||
5. Constructs an answer payload that echoes `sessionId`, carries
|
||||
the responder's reflexive and local addresses, includes a
|
||||
`PunchHint { startAtMs, intervalMs, durationMs }` telling both
|
||||
sides when to begin probing and how aggressively, and is
|
||||
wrapped, encrypted, and published the same way as the offer.
|
||||
|
||||
```json
|
||||
{
|
||||
"type": "answer",
|
||||
"sessionId": "<same as offer>",
|
||||
"issuedAt": <unix_millis>,
|
||||
"expiresAt": <unix_millis>,
|
||||
"nonce": "<random_nonce>",
|
||||
"senderNpub": "<responder_npub>",
|
||||
"recipientNpub": "<initiator_npub>",
|
||||
"inReplyTo": "<offer_event_id>",
|
||||
"accepted": true,
|
||||
"reflexiveAddress": {"protocol":"udp","ip":"<ip>","port":<port>},
|
||||
"localAddresses": [{"protocol":"udp","ip":"<ip>","port":<port>}],
|
||||
"stunServer": "<host>:<port>",
|
||||
"punch": {"startAtMs": <ms>, "intervalMs": <ms>, "durationMs": <ms>},
|
||||
"offerReceivedAt": <unix_millis>,
|
||||
"app_params": { ... }
|
||||
}
|
||||
```
|
||||
|
||||
If the responder has no usable addresses, it returns
|
||||
`accepted: false` with an explanatory `reason` and no `punch`.
|
||||
|
||||
The optional `offerReceivedAt` field carries the responder's
|
||||
wall-clock at the moment the offer arrived. The initiator can
|
||||
combine its own `T1` (offer-publish time), `T2 = offerReceivedAt`,
|
||||
`T3` (answer's `issuedAt`), and `T4` (answer-receive time) into the
|
||||
NTP-style estimate `((T2 − T1) + (T3 − T4)) / 2`, giving a per-peer
|
||||
clock-skew measurement that's useful for tuning freshness windows
|
||||
and for telemetry.
|
||||
|
||||
**Immediately after publishing the answer**, the responder begins
|
||||
Phase 4 punching without waiting for any acknowledgement that the
|
||||
initiator received the answer. NAT mappings are decaying and time
|
||||
is the binding constraint.
|
||||
|
||||
The responder must bind the inner JSON `senderNpub` /
|
||||
`recipientNpub` fields to the actual Nostr pubkeys that delivered
|
||||
the gift wrap, rather than treating those JSON fields as
|
||||
independently trustworthy. The wrap pubkey is the authentication
|
||||
ground-truth.
|
||||
|
||||
### Phase 4: Hole punching
|
||||
|
||||
Both peers now know each other's reflexive and local addresses.
|
||||
Both begin sending UDP packets from their respective punch sockets:
|
||||
|
||||
1. Send punch packets every **`intervalMs`** (typically 200 ms)
|
||||
across each planned target path:
|
||||
- reflexive-to-reflexive
|
||||
- private-subnet local-address paths (when subnet-compatible)
|
||||
- mixed local/reflexive fallbacks
|
||||
2. Each punch packet carries a fixed magic header so transit and
|
||||
peer code can distinguish it from stray UDP traffic:
|
||||
|
||||
```text
|
||||
Bytes 0–3: <PROBE_MAGIC> (application-defined u32)
|
||||
Bytes 4–7: sequence number (u32, big-endian, starting at 0)
|
||||
Bytes 8–23: first 16 bytes of SHA-256(sessionId)
|
||||
```
|
||||
|
||||
3. On receiving a valid punch packet (magic matches, session-id
|
||||
hash matches), the peer records the source address as the
|
||||
confirmed peer address and replies with an acknowledgement
|
||||
packet under a different magic value:
|
||||
|
||||
```text
|
||||
Bytes 0–3: <ACK_MAGIC> (application-defined u32)
|
||||
Bytes 4–7: echoed sequence number
|
||||
Bytes 8–23: first 16 bytes of SHA-256(sessionId)
|
||||
```
|
||||
|
||||
4. On receiving an acknowledgement, the peer considers the path
|
||||
punched and transitions to Phase 5.
|
||||
|
||||
If both peers advertised compatible local-subnet candidates, the
|
||||
local-address path will typically punch through faster than the
|
||||
reflexive path. The first path to acknowledge wins.
|
||||
|
||||
### Phase 5: Application protocol takeover
|
||||
|
||||
Once the path has acknowledged in both directions:
|
||||
|
||||
- The application protocol takes over the punch socket.
|
||||
- The signaling subscription can be closed.
|
||||
- The application is responsible for sending keepalive traffic at
|
||||
least every 15 seconds to refresh the NAT mapping. A flow that
|
||||
goes idle longer risks losing its mapping and having to retraverse.
|
||||
|
||||
### Phase 6: Cleanup
|
||||
|
||||
After the attempt completes (success or failure):
|
||||
|
||||
1. Close the relay subscription used for signaling.
|
||||
2. Optionally publish a NIP-09 deletion event referencing any
|
||||
signaling events the peer published. Because the wraps were
|
||||
ephemeral kinds with NIP-40 expiration tags, well-behaved relays
|
||||
will discard them automatically without explicit deletion.
|
||||
3. Discard the per-attempt punch socket if the attempt failed; a
|
||||
retry must allocate a new socket and a fresh reflexive address.
|
||||
|
||||
If the responder is going offline permanently it should also
|
||||
delete its kind-37195 (or equivalent) advert.
|
||||
|
||||
### Timeouts and retries
|
||||
|
||||
- If the initiator publishes an offer and receives no answer
|
||||
within a configured window (e.g. 10 s from offer publish), the
|
||||
attempt has failed. Causes: responder offline, advert stale,
|
||||
responder relay unreachable.
|
||||
- If the answer arrives but no valid punch acknowledgement is
|
||||
observed within `durationMs` (typically 10 s), the attempt has
|
||||
failed. Causes: symmetric NAT on either side, firewall
|
||||
interference, stale reflexive addresses.
|
||||
|
||||
The initiator may retry with a fresh STUN query, a fresh punch
|
||||
socket, and a new offer. Repeated failures against the same
|
||||
responder should be suppressed by the application layer; see
|
||||
*Application-specific failure handling* below.
|
||||
|
||||
---
|
||||
|
||||
## Security
|
||||
|
||||
### Authentication
|
||||
|
||||
Offer and answer payloads are NIP-44-encrypted to the recipient and
|
||||
NIP-59 gift-wrapped, so only the intended recipient can decrypt.
|
||||
Authentication of the sender comes from the inner-wrap signature
|
||||
(the rumour signed by the sender's long-term identity inside the
|
||||
NIP-59 seal), **not** from the outer wrap signature (which is the
|
||||
ephemeral pubkey).
|
||||
|
||||
The inner JSON `senderNpub` / `recipientNpub` fields must be bound
|
||||
to the actual signing pubkey of the inner rumour. Treating those
|
||||
JSON fields as independently trustworthy is a vulnerability —
|
||||
implementations must compare them against the unwrapped signature.
|
||||
|
||||
Once the UDP path is punched, the raw UDP channel has **no inherent
|
||||
authentication or encryption**. The application layer is responsible
|
||||
for establishing its own security on the punched channel — for
|
||||
example, a Noise Protocol handshake keyed from the Nostr identity,
|
||||
or an application-specific authenticated-encryption layer. FIPS
|
||||
runs its FMP Noise IK handshake immediately after adoption; the
|
||||
identity proven by the Noise handshake is the same Nostr pubkey
|
||||
that signed the inner offer/answer rumour, so a man-in-the-middle on
|
||||
the relay cannot impersonate the responder.
|
||||
|
||||
### Replay protection
|
||||
|
||||
The `sessionId` and `issuedAt` / `expiresAt` fields together
|
||||
defeat replays at the signaling layer. The responder must keep a
|
||||
bounded cache of recently-seen `sessionId` values and reject
|
||||
duplicates within the freshness window.
|
||||
|
||||
### Skew tolerance
|
||||
|
||||
Strict freshness checks fail under modest clock skew between
|
||||
peers. Implementations should accept offers and answers whose
|
||||
timestamps are off by a small absolute amount (FIPS uses ±60 s),
|
||||
and feed observed skew into a per-peer estimate for telemetry and
|
||||
tuning. Outright rejection should be reserved for grossly stale or
|
||||
future-dated messages.
|
||||
|
||||
### Metadata exposure
|
||||
|
||||
Even though signaling content is encrypted, the gift-wrap metadata
|
||||
reveals that the initiator's ephemeral pubkey contacted the
|
||||
responder's pubkey at a particular time, through a particular
|
||||
relay. The advert itself is public and reveals the responder's
|
||||
pubkey and the application protocol it speaks.
|
||||
|
||||
If metadata privacy is required, the advert content can be
|
||||
encrypted (consumers must already know the responder's pubkey),
|
||||
both peers can use ephemeral Nostr identities rather than their
|
||||
long-term keys, and the operator can run a private relay.
|
||||
|
||||
### NAT mapping integrity
|
||||
|
||||
If too much wall-clock time elapses between STUN discovery and the
|
||||
hole-punch attempt, the reflexive address goes stale. Both peers
|
||||
should complete the entire signaling exchange within tens of
|
||||
seconds of their respective STUN queries. Relay latency is the
|
||||
primary risk factor. Implementations targeting flaky relays should
|
||||
prefer relays known to deliver ephemeral events sub-second.
|
||||
|
||||
---
|
||||
|
||||
## Relay requirements
|
||||
|
||||
The protocol works best with relays that:
|
||||
|
||||
- Support ephemeral event kinds (`20000–29999`) and do not persist
|
||||
them.
|
||||
- Honor NIP-40 `expiration` tags and garbage-collect expired
|
||||
events.
|
||||
- Deliver events with low latency (sub-second WebSocket push).
|
||||
- Support NIP-09 deletion requests.
|
||||
|
||||
Relays that do not support ephemeral kinds will store the
|
||||
signaling events as regular events. The encrypted content remains
|
||||
opaque, but persisted wraps are wasteful and expose metadata
|
||||
unnecessarily. Operators deploying this protocol at scale should
|
||||
prefer relays that handle ephemeral kinds correctly, or run their
|
||||
own.
|
||||
|
||||
---
|
||||
|
||||
## Failure modes
|
||||
|
||||
| Failure | Symptom | Mitigation |
|
||||
| --- | --- | --- |
|
||||
| Symmetric NAT (one side) | Punch timeout | Retry with port-prediction heuristics; otherwise fall back to a relay or different transport |
|
||||
| Symmetric NAT (both sides) | Punch timeout | Application-level relay required |
|
||||
| Relay latency > 60 s | Stale reflexive address | Use low-latency relays; consider self-hosted relay |
|
||||
| Relay does not support ephemeral kinds | Signaling events persist | Use NIP-40 expiration + NIP-09 deletion as fallback |
|
||||
| Responder offline | No answer received | Initiator times out after configurable period |
|
||||
| Stale advert (responder no longer up) | Offer reaches no listener | Application-level failure suppression (see below) |
|
||||
| STUN server unreachable | No reflexive address | Fall back to alternate STUN server; fail if none reachable |
|
||||
| Firewall blocks outbound UDP | STUN fails entirely | NAT-traversal does not apply; reachable peers are limited to those that publish a non-UDP transport (e.g. TCP) and accept inbound |
|
||||
|
||||
### Application-specific failure handling
|
||||
|
||||
Repeated traversal failures against the same responder are common
|
||||
in practice — the responder may be offline, the advert may be
|
||||
stale, or the responder may be on a network that doesn't admit
|
||||
incoming UDP. A naive implementation that retries on every dial
|
||||
attempt floods the relay layer and the operator's logs.
|
||||
|
||||
Implementations should layer per-peer suppression on top of the
|
||||
basic retry. The shape of that suppression is application-specific.
|
||||
|
||||
#### FIPS example: failure suppression
|
||||
|
||||
FIPS layers the following suppression machinery on the basic retry
|
||||
loop:
|
||||
|
||||
- **Per-npub WARN log rate-limit** (`warn_log_interval_secs`,
|
||||
default 5 minutes). Subsequent failures inside the window log
|
||||
at debug level instead.
|
||||
- **Per-npub consecutive-failure counter and extended cooldown.**
|
||||
After `failure_streak_threshold` (default 5) consecutive
|
||||
failures, the per-peer retry deadline is pushed past
|
||||
`extended_cooldown_secs` (default 30 minutes). Open-discovery
|
||||
sweeps consult the cooldown so they don't immediately re-enqueue
|
||||
the same peer.
|
||||
- **Stale-advert eviction on streak transition.** When a peer
|
||||
hits the failure-streak threshold, the daemon actively
|
||||
re-fetches its advert from the configured advert relays. If the
|
||||
advert has been removed or replaced, the cache entry is evicted
|
||||
and the streak resets; if the advert is unchanged, the cooldown
|
||||
applies.
|
||||
- **Per-peer skew estimate.** The NTP-style skew computed from
|
||||
`offerReceivedAt` is recorded so consistently-skewed peers don't
|
||||
trip the freshness check on every attempt.
|
||||
- **Bounded failure-state cache** (`failure_state_max_entries`,
|
||||
default 4096) with LRU eviction so the suppression machinery
|
||||
itself does not grow unbounded.
|
||||
|
||||
These knobs are documented in
|
||||
[FIPS configuration reference](https://github.com/jmcorgan/fips/blob/master/docs/reference/configuration.md)
|
||||
under `node.discovery.nostr`.
|
||||
|
||||
---
|
||||
|
||||
## References
|
||||
|
||||
- **RFC 8489** — Session Traversal Utilities for NAT (STUN)
|
||||
- **RFC 8445** — Interactive Connectivity Establishment (ICE)
|
||||
- **RFC 4787** — NAT Behavioral Requirements for Unicast UDP
|
||||
- **NIP-01** — Basic Nostr protocol flow
|
||||
- **NIP-09** — Event deletion request
|
||||
- **NIP-17** — Inbox relay list (kind `10050`) for direct-message
|
||||
routing
|
||||
- **NIP-40** — Expiration timestamp
|
||||
- **NIP-44** — Versioned encryption
|
||||
- **NIP-59** — Gift wrap
|
||||
- **NIP-78** — Application-specific data
|
||||
@@ -1,14 +1,20 @@
|
||||
# FIPS Spanning Tree Protocol Dynamics
|
||||
|
||||
A detailed study of the gossip-based spanning tree protocol, focusing on
|
||||
operational behavior under various mesh conditions. This document complements
|
||||
[fips-intro.md](fips-intro.md) with step-by-step walkthroughs of protocol
|
||||
dynamics rather than message formats and data structures.
|
||||
A detailed study of the gossip-based spanning tree protocol, focusing
|
||||
on operational behavior under various mesh conditions. This document
|
||||
complements [fips-concepts.md](fips-concepts.md) and
|
||||
[fips-architecture.md](fips-architecture.md) with step-by-step
|
||||
walkthroughs of protocol dynamics rather than message formats and
|
||||
data structures.
|
||||
|
||||
For wire formats, see [fips-wire-formats.md](fips-wire-formats.md) (TreeAnnounce section).
|
||||
For spanning tree algorithms and data structures, see
|
||||
[fips-spanning-tree.md](fips-spanning-tree.md). For how the spanning tree fits
|
||||
into mesh routing, see [fips-mesh-operation.md](fips-mesh-operation.md).
|
||||
For wire formats, see
|
||||
[../reference/wire-formats.md](../reference/wire-formats.md)
|
||||
(TreeAnnounce section). For spanning tree algorithms and data
|
||||
structures, see [fips-spanning-tree.md](fips-spanning-tree.md). For
|
||||
how the spanning tree fits into mesh routing, see
|
||||
[fips-mesh-operation.md](fips-mesh-operation.md). For the academic
|
||||
foundations and references that underpin this document, see
|
||||
[fips-prior-work.md](fips-prior-work.md).
|
||||
|
||||
## Contents
|
||||
|
||||
@@ -93,7 +99,7 @@ When a node starts with no peers, it bootstraps as a single-node network.
|
||||
**T0: Node A starts.**
|
||||
|
||||
- Generates or loads keypair `(npub_A, nsec_A)`
|
||||
- Computes `node_addr_A = SHA-256(npub_A)`
|
||||
- Computes `node_addr_A = SHA-256(pubkey_A)[..16]` (128 bits)
|
||||
- Initializes empty TreeState
|
||||
- Sets `parent = self` (A is its own root), `sequence = 1`
|
||||
- Records current timestamp
|
||||
@@ -518,6 +524,11 @@ converge to the same "link failed" state, though B detects it up to
|
||||
## 8. Parent Selection
|
||||
|
||||
Parent selection determines tree structure and routing efficiency.
|
||||
The algorithm itself (effective-depth ranking, hold-down, hysteresis,
|
||||
mandatory-switch bypass) is canonically documented in
|
||||
[fips-spanning-tree.md](fips-spanning-tree.md); this section walks
|
||||
through what re-selection looks like under specific dynamic
|
||||
conditions and the rationale for the local-only cost metric.
|
||||
|
||||
### Cost-Based Selection with Effective Depth
|
||||
|
||||
@@ -537,10 +548,15 @@ purely by tree depth without link quality consideration.
|
||||
2. **Compute effective depth for each candidate.** For every peer whose
|
||||
announced root matches the smallest root, the algorithm calculates
|
||||
`effective_depth = peer.depth + link_cost`, where `link_cost` comes from
|
||||
`peer_costs` (MMP-derived) or defaults to 1.0 when metrics have not yet
|
||||
converged. The best candidate is the peer with the lowest effective depth,
|
||||
with ties broken by numerically smallest `NodeAddr`. If the best candidate
|
||||
is already the current parent, no switch is needed.
|
||||
`peer_costs` (MMP-derived). During cold start, when no peer has MMP data
|
||||
yet (`peer_costs` is empty), unmeasured candidates default to 1.0; once
|
||||
any peer has MMP data, unmeasured candidates are skipped so a freshly
|
||||
connected peer cannot win on its default cost. Candidates whose ancestry
|
||||
already contains the local node are also rejected, preventing an
|
||||
alternating two-node loop. The best candidate is the peer with the
|
||||
lowest effective depth, with ties broken by numerically smallest
|
||||
`NodeAddr`. If the best candidate is already the current parent, no
|
||||
switch is needed.
|
||||
|
||||
3. **Check for mandatory switches.** Two conditions bypass all stability
|
||||
mechanisms and trigger an immediate parent change: the current parent is no
|
||||
@@ -571,10 +587,10 @@ re-evaluation independent of TreeAnnounce traffic).
|
||||
|
||||
Where ETX (Expected Transmission Count, from De Couto et al., "A
|
||||
High-Throughput Path Metric for Multi-Hop Wireless Routing", 2003) comes from
|
||||
bidirectional MMP delivery ratios and SRTT (Smoothed Round-Trip Time) from MMP
|
||||
timestamp-echo. When MMP
|
||||
metrics have not yet converged, `link_cost` defaults to 1.0, preserving
|
||||
depth-only behavior as a graceful fallback.
|
||||
bidirectional MMP delivery ratios and SRTT (Smoothed Round-Trip Time) from
|
||||
MMP timestamp-echo. During cold start, before any peer has MMP data, the
|
||||
default cost of 1.0 is used and the algorithm reduces to depth-only
|
||||
selection.
|
||||
|
||||
**What this means for tree structure**: The algorithm can prefer a deeper parent
|
||||
with a better link over a shallower parent with a poor link, when the effective
|
||||
@@ -942,28 +958,16 @@ costs to form efficient tree structures.
|
||||
|
||||
### Prior Art and FIPS Contributions
|
||||
|
||||
The protocol builds on established foundations and adds several new elements:
|
||||
|
||||
**Derived from prior work**:
|
||||
|
||||
- Spanning tree coordinate routing (Yggdrasil/Ironwood, building on Kleinberg
|
||||
2007 and Cvetkovski/Crovella 2009)
|
||||
- Deterministic root discovery via smallest identifier (Yggdrasil; echoes
|
||||
IEEE 802.1D STP bridge ID selection)
|
||||
- CRDT-based distributed state (Shapiro et al. 2011)
|
||||
- Gossip dissemination (epidemic model; Kermarrec 2007)
|
||||
- Heartbeat-based failure detection (SWIM; Das et al. 2002)
|
||||
- ETX link metric (De Couto et al. 2003)
|
||||
- Hysteresis and hold-down for route stability (OSPF, BGP, IS-IS)
|
||||
|
||||
**FIPS additions**:
|
||||
|
||||
- Cost-aware parent selection using local-only link metrics (effective depth =
|
||||
tree depth + link cost), replacing Yggdrasil's depth-only selection
|
||||
- Combined ETX + SRTT link cost formula with MMP-measured components
|
||||
- Flap dampening with mandatory switch bypass
|
||||
- Announcement suppression for transient state changes
|
||||
- Tree-only bloom filter merge with split-horizon exclusion
|
||||
The protocol builds on established foundations (Yggdrasil/Ironwood
|
||||
tree-coordinate routing, IEEE 802.1D STP root election, CRDT-based
|
||||
distributed state, SWIM-style failure detection, ETX, OSPF-style
|
||||
hysteresis and hold-down) and adds several new elements (cost-aware
|
||||
parent selection on local-only metrics, the combined ETX + SRTT cost
|
||||
formula, flap dampening with mandatory-switch bypass, announcement
|
||||
suppression, and tree-only bloom filter merge with split-horizon).
|
||||
Both the prior-art map and the FIPS contributions list are
|
||||
consolidated in
|
||||
[fips-prior-work.md](fips-prior-work.md#fips-contributions).
|
||||
|
||||
---
|
||||
|
||||
@@ -971,75 +975,17 @@ The protocol builds on established foundations and adds several new elements:
|
||||
|
||||
### FIPS Internal Documentation
|
||||
|
||||
- [fips-spanning-tree.md](fips-spanning-tree.md) — Spanning tree algorithms and data structures
|
||||
- [fips-mesh-operation.md](fips-mesh-operation.md) — How the spanning tree fits into mesh routing
|
||||
- [fips-wire-formats.md](fips-wire-formats.md) — TreeAnnounce wire format
|
||||
- [fips-spanning-tree.md](fips-spanning-tree.md) — Spanning tree
|
||||
algorithms and data structures
|
||||
- [fips-mesh-operation.md](fips-mesh-operation.md) — How the spanning
|
||||
tree fits into mesh routing
|
||||
- [../reference/wire-formats.md](../reference/wire-formats.md) —
|
||||
TreeAnnounce wire format
|
||||
|
||||
### Yggdrasil Documentation
|
||||
### Prior Art and Academic Foundations
|
||||
|
||||
- [Yggdrasil v0.5 Release Notes](https://yggdrasil-network.github.io/2023/10/22/upcoming-v05-release.html)
|
||||
- [Ironwood Routing Library](https://github.com/Arceliar/ironwood)
|
||||
- [The World Tree (Yggdrasil Blog)](https://yggdrasil-network.github.io/2018/07/17/world-tree.html)
|
||||
- [Yggdrasil Implementation Overview](https://yggdrasil-network.github.io/implementation.html)
|
||||
|
||||
### Academic Foundations
|
||||
|
||||
#### Virtual Coordinate Routing
|
||||
|
||||
- Rao, A., Ratnasamy, S., Papadimitriou, C., Shenker, S., Stoica, I.
|
||||
["Geographic Routing without Location Information"](https://people.eecs.berkeley.edu/~sylvia/papers/p327-rao.pdf).
|
||||
MobiCom 2003. *Established virtual coordinate routing using network topology.*
|
||||
|
||||
#### Greedy Embedding Theory
|
||||
|
||||
- Kleinberg, R.
|
||||
["Geographic Routing Using Hyperbolic Space"](https://www.semanticscholar.org/paper/Geographic-Routing-Using-Hyperbolic-Space-Kleinberg/f506b2ddb142d2ec539400297ba53383d958abef).
|
||||
IEEE INFOCOM 2007. *Proved every connected graph has a greedy embedding in
|
||||
hyperbolic space; showed spanning trees enable coordinate assignment.*
|
||||
|
||||
- Cvetkovski, A., Crovella, M.
|
||||
["Hyperbolic Embedding and Routing for Dynamic Graphs"](https://www.cs.bu.edu/faculty/crovella/paper-archive/infocom09-hyperbolic.pdf).
|
||||
IEEE INFOCOM 2009. *Dynamic embedding for nodes joining/leaving; introduced
|
||||
Gravity-Pressure routing for failure recovery.*
|
||||
|
||||
- Crovella, M. et al.
|
||||
["On the Choice of a Spanning Tree for Greedy Embedding"](https://www.cs.bu.edu/faculty/crovella/paper-archive/networking-science13.pdf).
|
||||
Networking Science 2013. *Analysis of how tree structure affects routing stretch.*
|
||||
|
||||
- Bläsius, T. et al.
|
||||
["Hyperbolic Embeddings for Near-Optimal Greedy Routing"](https://dl.acm.org/doi/10.1145/3381751).
|
||||
ACM Journal of Experimental Algorithmics 2020. *Achieved 100% success ratio
|
||||
with 6% stretch on Internet graph.*
|
||||
|
||||
#### Link Metrics
|
||||
|
||||
- De Couto, D., Aguayo, D., Bicket, J., Morris, R.
|
||||
"A High-Throughput Path Metric for Multi-Hop Wireless Routing".
|
||||
MobiCom 2003. *Introduced ETX (Expected Transmission Count) as a link
|
||||
quality metric for wireless mesh networks.*
|
||||
|
||||
#### Routing Protocol Stability
|
||||
|
||||
- IEEE 802.1D. "IEEE Standard for Local and Metropolitan Area
|
||||
Networks: Media Access Control (MAC) Bridges". *Spanning Tree
|
||||
Protocol (STP) — root election via bridge ID, BPDU exchange.*
|
||||
|
||||
- Moy, J. [RFC 2328](https://datatracker.ietf.org/doc/html/rfc2328):
|
||||
"OSPF Version 2". 1998. *Link-state routing with cumulative path
|
||||
costs and SPF computation. FIPS's local-only cost approach is
|
||||
contrasted with OSPF's cumulative model in §8.*
|
||||
|
||||
#### Distributed Systems Primitives
|
||||
|
||||
- Shapiro, M., Preguiça, N., Baquero, C., Zawirski, M.
|
||||
"Conflict-free Replicated Data Types". SSS 2011.
|
||||
*Formal definition of CRDTs enabling coordination-free consistency.*
|
||||
|
||||
- Das, A., Gupta, I., Motivala, A.
|
||||
["SWIM: Scalable Weakly-consistent Infection-style Process Group Membership"](https://www.cs.cornell.edu/projects/Quicksilver/public_pdfs/SWIM.pdf).
|
||||
IPDPS 2002. *O(1) failure detection, O(log N) dissemination via gossip.*
|
||||
|
||||
- Kermarrec, A-M.
|
||||
["Gossiping in Distributed Systems"](https://www.distributed-systems.net/my-data/papers/2007.osr.pdf).
|
||||
ACM SIGOPS Operating Systems Review 2007. *Framework for gossip-based
|
||||
protocols achieving O(log N) propagation.*
|
||||
The Yggdrasil documentation and the academic-foundations bibliography
|
||||
(virtual coordinate routing, greedy embedding theory, link metrics,
|
||||
routing-protocol stability, and distributed systems primitives) are
|
||||
collected in
|
||||
[fips-prior-work.md](fips-prior-work.md#spanning-tree-dynamics-foundations).
|
||||
|
||||
@@ -0,0 +1,242 @@
|
||||
# Getting Started with FIPS
|
||||
|
||||
FIPS (Free Internetworking Peering System) is a self-organizing
|
||||
encrypted mesh network built on Nostr identities. Your machine
|
||||
becomes a node in the mesh with a self-generated cryptographic
|
||||
identity, and existing networking software — SSH, web servers,
|
||||
file transfer, anything IPv6-native — runs over the mesh
|
||||
unchanged.
|
||||
|
||||
There are two common ways to deploy FIPS, and the rest of this
|
||||
guide and the linked docs branch accordingly:
|
||||
|
||||
- **As an overlay** on top of existing IP networks (Ethernet,
|
||||
WiFi, the public internet, Tor), FIPS lets your node reach
|
||||
any other peer regardless of NAT, ISP, or physical location.
|
||||
- **From the ground up** over non-IP transports — raw Ethernet,
|
||||
WiFi, Bluetooth — FIPS provides a complete permissionless
|
||||
network without any pre-existing IP infrastructure, ISP, or
|
||||
DNS.
|
||||
|
||||
The two paths share a lot of common ground — install, identity,
|
||||
configuration. They diverge mainly in transport setup and the
|
||||
deployment topology you choose.
|
||||
|
||||
There is no central server. Any node can run; any pair of
|
||||
running nodes can mesh.
|
||||
|
||||
## What you'll need
|
||||
|
||||
- A Linux, macOS, or Windows host. Linux is the most exercised
|
||||
platform; macOS and Windows installers are available.
|
||||
- The pre-built installer for your platform (see the project
|
||||
README's [Quick start](../README.md#quick-start) section for
|
||||
download links), **or** a source checkout if you want to build
|
||||
the installer yourself.
|
||||
- For the source-build path only: a working Rust toolchain (the
|
||||
version pinned in `rust-toolchain.toml` is auto-installed by
|
||||
rustup), and the platform-specific build dependencies listed in
|
||||
[packaging/README.md](../packaging/README.md).
|
||||
|
||||
## Install
|
||||
|
||||
FIPS is installed by running a binary installer for your
|
||||
platform. The installer drops the daemon and CLI tools into
|
||||
system locations, installs systemd / launchd / Windows-service
|
||||
unit files, places a default `fips.yaml`, and creates the `fips`
|
||||
system group. There is no `cargo install` path: the daemon needs
|
||||
more than just binaries copied into place.
|
||||
|
||||
You can either build the installer yourself from source, or
|
||||
download a pre-built one from the release distribution. Both
|
||||
paths produce the same installer artifacts and the same
|
||||
post-install state.
|
||||
|
||||
### From the release distribution
|
||||
|
||||
The most direct path. The release distribution carries a
|
||||
per-platform installer:
|
||||
|
||||
- Debian/Ubuntu — `.deb` package
|
||||
- Arch Linux — `fips` AUR package
|
||||
- OpenWrt — `.ipk` package
|
||||
- macOS — `.pkg` installer
|
||||
- Windows — `.zip` with service-install scripts
|
||||
- Generic systemd Linux — `.tar.gz` with an `install.sh` script
|
||||
|
||||
See the [project README's Quick start section](../README.md#quick-start)
|
||||
for download links and per-platform invocations.
|
||||
|
||||
### From source
|
||||
|
||||
For development, custom builds, or unsupported architectures.
|
||||
The `packaging/` tree builds the same installer formats locally;
|
||||
you then apply the resulting installer the same way you would a
|
||||
downloaded one.
|
||||
|
||||
```sh
|
||||
git clone https://github.com/jmcorgan/fips.git
|
||||
cd fips/packaging
|
||||
make deb # or: tarball, ipk, aur, pkg, zip, all
|
||||
```
|
||||
|
||||
The resulting installer lands in `deploy/` at the project root.
|
||||
Apply it the same way you would a downloaded one (for example
|
||||
`sudo dpkg -i deploy/fips_*.deb` on Debian/Ubuntu).
|
||||
|
||||
See [packaging/README.md](../packaging/README.md) for per-format
|
||||
build details, cross-target options, and the full `make` target
|
||||
list.
|
||||
|
||||
## What's installed and running
|
||||
|
||||
Here's what the installer leaves on your machine, what's
|
||||
running, and what you'll need to set up yourself.
|
||||
|
||||
**Binaries installed system-wide:**
|
||||
|
||||
- `fips` (daemon)
|
||||
- `fipsctl` (control-socket client)
|
||||
- `fipstop` (live-status TUI)
|
||||
- `fips-gateway`
|
||||
|
||||
**Files placed on disk:**
|
||||
|
||||
- `/etc/fips/fips.yaml` — default daemon config (preserved on
|
||||
upgrade).
|
||||
- `/etc/fips/fips.nft` — mesh-interface nftables baseline (used
|
||||
only when the firewall service is enabled).
|
||||
- `/etc/fips/fips.d/` — empty drop-in directory for operator
|
||||
nftables additions.
|
||||
- Systemd, launchd, or Windows-service unit files for the four
|
||||
fips services.
|
||||
|
||||
**System changes:**
|
||||
|
||||
- A `fips` system group is created. Add your user to it
|
||||
(`sudo usermod -aG fips $USER`, then re-login) to run
|
||||
`fipsctl` and `fipstop` without `sudo`.
|
||||
- The runtime directory `/run/fips/` exists with mode
|
||||
`0750 root:fips`.
|
||||
|
||||
**Services enabled and started on boot:**
|
||||
|
||||
- `fips.service` — the daemon. Brings up the `fips0` TUN
|
||||
adapter, listens on the configured transports, and exposes
|
||||
the control socket at `/run/fips/control.sock`.
|
||||
- `fips-dns.service` — wires `.fips` hostname resolution into
|
||||
the host resolver (a `/etc/systemd/resolved.conf.d/` drop-in
|
||||
pointing at `[::1]:5354` on systemd hosts).
|
||||
|
||||
**Services installed but not enabled** (operator opt-in):
|
||||
|
||||
- `fips-firewall.service` — applies `/etc/fips/fips.nft` to
|
||||
the mesh interface. See
|
||||
[how-to/enable-mesh-firewall.md](how-to/enable-mesh-firewall.md).
|
||||
|
||||
**What's working out of the box:**
|
||||
|
||||
- The daemon is running with a fresh **ephemeral** identity —
|
||||
a new Nostr keypair is generated on every start.
|
||||
- The `fips0` TUN adapter exists with the daemon's mesh address.
|
||||
- The daemon's transport listeners are up: UDP `0.0.0.0:2121`
|
||||
and TCP `0.0.0.0:8443`. They are inert at this point because
|
||||
no other node knows your daemon's npub yet — see "What's not
|
||||
yet configured" below.
|
||||
- `.fips` hostname resolution is plumbed into the host
|
||||
resolver.
|
||||
|
||||
**What's not yet configured** — these are what guide your next
|
||||
steps:
|
||||
|
||||
- **No peers.** The daemon has nobody to talk to until you add
|
||||
a static peer entry, enable Nostr-mediated discovery, or
|
||||
bring up a transport (Ethernet, Bluetooth) where peers find
|
||||
each other automatically on the same physical link.
|
||||
- **Ephemeral identity.** Your node's npub changes every
|
||||
restart. The
|
||||
[persistent-identity tutorial](tutorials/persistent-identity.md)
|
||||
walks through pinning the daemon to a stable Nostr keypair
|
||||
for any node others will reference by name.
|
||||
- **Mesh firewall not active.** Inbound exposure on `fips0`
|
||||
follows the host's existing firewall rules until you enable
|
||||
the baseline service.
|
||||
|
||||
## Reaching mesh nodes by name
|
||||
|
||||
A FIPS node is identified by its Nostr public key (`npub1...`).
|
||||
For ordinary IP software running over the mesh — SSH, web
|
||||
browsers, `ping`, file transfer — use the form `<npub>.fips`
|
||||
as the destination; the local `.fips` resolver translates that
|
||||
to the corresponding mesh IPv6 address so the FIPS node can be
|
||||
found. The resolver runs entirely on your machine and does not
|
||||
generate any external DNS traffic.
|
||||
|
||||
For shorter forms, the resolver also consults two host maps
|
||||
before falling back to direct npub lookup: `/etc/fips/hosts`
|
||||
(shipped pre-populated with the public test mesh roster, and
|
||||
freely editable for your own entries) and the `alias:` field
|
||||
on configured peers in `fips.yaml`. So `test-us01.fips`,
|
||||
`my-laptop.fips`, or any other shortname you map resolves the
|
||||
same way `<npub>.fips` does. See
|
||||
[how-to/host-aliases.md](how-to/host-aliases.md) for the full
|
||||
mechanics.
|
||||
|
||||
## Join the test mesh
|
||||
|
||||
The fastest way to see FIPS in action is to connect your daemon
|
||||
to the public FIPS test mesh. The
|
||||
[Join the Test Mesh](tutorials/join-the-test-mesh.md) tutorial
|
||||
walks through adding a single static peer entry, watching the
|
||||
link come up, and reaching both that peer and a second mesh node
|
||||
forwarded through it — a ten-minute exercise that demonstrates
|
||||
the central FIPS guarantee that one good peer connects you to
|
||||
the rest of the mesh.
|
||||
|
||||
## Where to go next
|
||||
|
||||
Documentation is organised into four sections, each with a different
|
||||
job. Pick the one that matches what you want to do.
|
||||
|
||||
### [Tutorials](tutorials/)
|
||||
|
||||
Step-by-step lessons that take you from zero to a working setup.
|
||||
Read these end-to-end. Start with
|
||||
[Join the Test Mesh](tutorials/join-the-test-mesh.md) and follow
|
||||
with
|
||||
[ipv6-adapter-walkthrough](tutorials/ipv6-adapter-walkthrough.md)
|
||||
to understand what each piece does, then move on to
|
||||
[persistent-identity](tutorials/persistent-identity.md) and
|
||||
the three Nostr-discovery tutorials —
|
||||
[resolve-peers-via-nostr](tutorials/resolve-peers-via-nostr.md),
|
||||
[advertise-your-node](tutorials/advertise-your-node.md), and
|
||||
[open-discovery](tutorials/open-discovery.md) — to give your
|
||||
node a stable npub, look up peer endpoints, publish your
|
||||
own, and join the ambient discovery namespace. Then [host-a-service](tutorials/host-a-service.md) for hosting
|
||||
a service on your node, and [ground-up-mesh](tutorials/ground-up-mesh.md)
|
||||
for the second deployment mode where two devices peer over
|
||||
Ethernet, WiFi, or Bluetooth with no IP between them.
|
||||
|
||||
### [How-To Guides](how-to/)
|
||||
|
||||
Task-oriented recipes for operators with a specific goal: enable a
|
||||
firewall, deploy the LAN gateway, set up Bluetooth peering,
|
||||
diagnose an MTU problem, configure persistent identity. Each guide
|
||||
takes the shortest correct path from "I want to do X" to "X is done".
|
||||
|
||||
### [Reference](reference/)
|
||||
|
||||
Lookup material consulted on demand: wire formats, configuration
|
||||
keys, command-line flags, control-socket commands. Austere by
|
||||
design; no guidance on when to use a feature.
|
||||
|
||||
### [Design](design/)
|
||||
|
||||
Architectural and protocol-level explanations: the mesh layer, the
|
||||
session layer, the spanning tree, Bloom-filter discovery, the
|
||||
unified MTU model, the IPv6 adapter. Read these to understand *why*
|
||||
FIPS makes the choices it does.
|
||||
|
||||
The design section's
|
||||
[fips-concepts.md](design/fips-concepts.md) is a good entry point if
|
||||
you want the mental model before touching any commands.
|
||||
@@ -0,0 +1,27 @@
|
||||
# How-To Guides
|
||||
|
||||
Task-oriented, step-by-step recipes for operators with a specific
|
||||
goal in mind. Each guide assumes the reader already knows what FIPS
|
||||
is and wants to get a particular thing done — enable a feature,
|
||||
deploy a component, troubleshoot a class of problem.
|
||||
|
||||
How-to guides do not teach concepts (that is the role of design/)
|
||||
and do not enumerate options (that is the role of reference/). They
|
||||
take the reader along the shortest correct path from "I want to do
|
||||
X" to "X is done".
|
||||
|
||||
## Available Guides
|
||||
|
||||
| Guide | Goal |
|
||||
| ----- | ---- |
|
||||
| [enable-mesh-firewall.md](enable-mesh-firewall.md) | Activate the default-deny nftables baseline on `fips0` |
|
||||
| [enable-nostr-discovery.md](enable-nostr-discovery.md) | Turn on Nostr-mediated discovery (3 capabilities — resolve, advertise, open — across 5 scenarios) |
|
||||
| [deploy-tor-onion.md](deploy-tor-onion.md) | Run a Tor onion service for inbound FIPS connections |
|
||||
| [tune-udp-buffers.md](tune-udp-buffers.md) | Set host sysctls so FIPS UDP sockets don't get clamped |
|
||||
| [run-as-unprivileged-user.md](run-as-unprivileged-user.md) | Run the daemon under a dedicated unprivileged service account (drops the default-root posture) |
|
||||
| [deploy-gateway.md](deploy-gateway.md) | Manually deploy `fips-gateway` on a non-OpenWrt Linux host (LAN-to-mesh outbound + mesh-to-LAN inbound port-forwards). For the OpenWrt path, see the gateway tutorial. |
|
||||
| [troubleshoot-gateway.md](troubleshoot-gateway.md) | Diagnostic recipes for the gateway, organised by half (outbound, inbound, common) |
|
||||
| [persistent-identity.md](persistent-identity.md) | Provision a stable Nostr keypair so the node keeps the same npub across restarts |
|
||||
| [host-aliases.md](host-aliases.md) | Use shortnames (`test-us01.fips`, `my-laptop.fips`) instead of full npubs by editing `/etc/fips/hosts` or setting peer aliases |
|
||||
| [set-up-bluetooth-peer.md](set-up-bluetooth-peer.md) | Configure a Bluetooth Low Energy peer link |
|
||||
| [diagnose-mtu-issues.md](diagnose-mtu-issues.md) | Triage MTU-shaped failures and rule out their imposters (bufferbloat, transport saturation) |
|
||||
@@ -0,0 +1,458 @@
|
||||
# Deploy `fips-gateway` (Manual Linux-Host Setup)
|
||||
|
||||
`fips-gateway` is a separate service that runs alongside the FIPS
|
||||
daemon and bridges a non-FIPS LAN to the FIPS mesh in two
|
||||
independent directions: **outbound** (LAN clients reach mesh
|
||||
services through DNS proxy + virtual-IP NAT) and **inbound** (mesh
|
||||
peers reach LAN services through 1:1 port forwards on `fips0`).
|
||||
This guide covers the **manual Linux-host** deployment path —
|
||||
wiring DNS forwarding, route distribution, and firewall integration
|
||||
on a server or non-OpenWrt router by hand.
|
||||
|
||||
> **Running OpenWrt?** Use the
|
||||
> [tutorial](../tutorials/deploy-fips-gateway.md) instead. The OpenWrt
|
||||
> ipk ships with the `gateway:` block pre-populated and the init
|
||||
> script automates dnsmasq forwarding, RA route distribution, and the
|
||||
> global IPv6 prefix on `br-lan`. The OpenWrt path is the canonical
|
||||
> deployment of this feature; this how-to is the secondary path for
|
||||
> operators with a different LAN-edge box (a Linux server already
|
||||
> serving DHCP/DNS, a custom router distribution, etc.).
|
||||
|
||||
For the gateway design (NAT pipeline, virtual IP pool lifecycle, DNS
|
||||
resolution flow), see [../design/fips-gateway.md](../design/fips-gateway.md).
|
||||
For the full `gateway.*` configuration block, see the
|
||||
[Gateway section](../reference/configuration.md#gateway-gateway) of
|
||||
the configuration reference. For the `fips-gateway` binary's CLI
|
||||
flags, see [../reference/cli-fips-gateway.md](../reference/cli-fips-gateway.md).
|
||||
|
||||
## The two halves
|
||||
|
||||
The gateway exposes two independent features that share a common
|
||||
control plane (the same binary, the same nftables table `inet
|
||||
fips_gateway`, the same control socket `/run/fips/gateway.sock`, the
|
||||
same `gateway.*` config block). You can configure either half on its
|
||||
own or both together.
|
||||
|
||||
- **Outbound gateway** (LAN → mesh). Non-FIPS LAN workstations resolve
|
||||
`<npub>.fips` names against the gateway's DNS listener and receive
|
||||
AAAA answers from the gateway's virtual-IP pool. Outbound traffic
|
||||
to those addresses is DNAT'd to the real mesh address and SNAT'd
|
||||
(masqueraded) onto `fips0` under the gateway's mesh identity. The
|
||||
audience is unmodified LAN clients.
|
||||
|
||||
- **Inbound gateway** (mesh → LAN). A static `(listen_port, proto)
|
||||
→ [target_addr]:target_port` table — configured in
|
||||
`gateway.port_forwards[]` — exposes selected LAN services to the
|
||||
mesh as `<gateway-npub>.fips:<listen_port>`. Mesh peers connect to
|
||||
the gateway's mesh address; the gateway DNATs to the LAN target
|
||||
and masquerades on the LAN side so return traffic flows through
|
||||
conntrack. The audience is mesh peers reaching a service that
|
||||
happens to live on this LAN.
|
||||
|
||||
The two halves are independent. Configure the outbound half if you
|
||||
want LAN clients to *reach* the mesh; configure the inbound half if
|
||||
you want mesh peers to *reach into* the LAN; configure both if you
|
||||
want both.
|
||||
|
||||
## Common gateway-host setup
|
||||
|
||||
Both halves require the same host preparation. Work through this
|
||||
section first, then jump to whichever half (or both) you need.
|
||||
|
||||
### FIPS daemon prerequisites
|
||||
|
||||
The gateway runs alongside a `fips` daemon on the same host:
|
||||
|
||||
- The daemon must be running with the TUN adapter enabled (the
|
||||
`fips0` interface must exist).
|
||||
- The daemon's DNS resolver must be enabled (`dns.enabled: true`,
|
||||
default) and reachable from `fips-gateway`. By default that means
|
||||
`[::1]:5354` (IPv6 loopback). The gateway's default
|
||||
`dns.upstream` matches this; a v4 upstream like `127.0.0.1:5354`
|
||||
cannot reach a daemon bound on `[::1]:5354` because Linux IPv6
|
||||
sockets bound to explicit `::1` do not accept v4-mapped traffic.
|
||||
|
||||
If the daemon is not yet running with these features, set up the
|
||||
daemon first — see [persistent-identity.md](persistent-identity.md)
|
||||
and [../reference/configuration.md](../reference/configuration.md).
|
||||
|
||||
### Kernel sysctls
|
||||
|
||||
```sh
|
||||
sudo sysctl -w net.ipv6.conf.all.forwarding=1
|
||||
sudo sysctl -w net.ipv6.conf.all.proxy_ndp=1
|
||||
```
|
||||
|
||||
`forwarding` lets the host route IPv6 packets between the LAN
|
||||
interface and `fips0`. `proxy_ndp` lets the gateway answer Neighbor
|
||||
Solicitation requests for virtual-pool addresses so LAN clients can
|
||||
resolve their link-layer addresses (only relevant for the outbound
|
||||
half, but harmless if you only run the inbound half).
|
||||
|
||||
Persist via a drop-in:
|
||||
|
||||
```sh
|
||||
sudo tee /etc/sysctl.d/60-fips-gateway.conf <<'EOF'
|
||||
net.ipv6.conf.all.forwarding = 1
|
||||
net.ipv6.conf.all.proxy_ndp = 1
|
||||
EOF
|
||||
sudo sysctl --system
|
||||
```
|
||||
|
||||
### Capability
|
||||
|
||||
`fips-gateway` requires `CAP_NET_ADMIN` to manage its nftables table
|
||||
(`inet fips_gateway`) and proxy-NDP entries. The packaged systemd
|
||||
unit (`fips-gateway.service`) runs as root, which satisfies this. For
|
||||
non-package installs, set the file capability:
|
||||
|
||||
```sh
|
||||
sudo setcap cap_net_admin+ep /usr/bin/fips-gateway
|
||||
```
|
||||
|
||||
### Pool route
|
||||
|
||||
At startup `fips-gateway` adds `local <pool-cidr> dev lo` to the
|
||||
local routing table. This tells the kernel to accept packets
|
||||
destined for pool addresses as locally-owned, enabling the NAT
|
||||
processing path. The route is cleaned up on shutdown. You do not
|
||||
need to install it manually; if you see "destination unreachable"
|
||||
errors for pool addresses on the gateway host, verify the route is
|
||||
present:
|
||||
|
||||
```sh
|
||||
ip -6 route show table local | grep <pool-cidr>
|
||||
```
|
||||
|
||||
### Minimum configuration
|
||||
|
||||
In `/etc/fips/fips.yaml`, populate the `gateway` block with at minimum
|
||||
`enabled: true`, `pool`, and `lan_interface`:
|
||||
|
||||
```yaml
|
||||
gateway:
|
||||
enabled: true
|
||||
pool: "fd01::/112"
|
||||
lan_interface: "enp3s0"
|
||||
```
|
||||
|
||||
Pick a pool CIDR that does **not** overlap with any address space in
|
||||
use on the LAN or in the mesh (the FIPS mesh occupies `fd00::/8`;
|
||||
pick a different `fdXX::/N`). The `/112` size yields 65 536 virtual
|
||||
IPs, which is the gateway's hard cap regardless of CIDR width.
|
||||
|
||||
This minimum config is enough to start the gateway. The `dns.*` block
|
||||
is optional and defaults to `listen: "[::1]:5353"` and
|
||||
`upstream: "[::1]:5354"`. The full block — including `dns.*`,
|
||||
`pool_grace_period`, `conntrack.*`, and `port_forwards[]` — is
|
||||
documented in
|
||||
[../reference/configuration.md#gateway-gateway](../reference/configuration.md#gateway-gateway).
|
||||
|
||||
### Start the service
|
||||
|
||||
```sh
|
||||
sudo systemctl enable --now fips-gateway
|
||||
```
|
||||
|
||||
Verify the unit came up:
|
||||
|
||||
```sh
|
||||
sudo systemctl status fips-gateway
|
||||
sudo journalctl -u fips-gateway -e
|
||||
```
|
||||
|
||||
The startup log will report `Gateway config loaded`,
|
||||
`DNS upstream is reachable`, `Created nftables table 'fips_gateway'`,
|
||||
and finally `fips-gateway running`. The unit's `ExecStartPre` waits up
|
||||
to 30 s for `fips0` to appear, which covers the cold-boot race where
|
||||
the daemon is still bringing up its TUN.
|
||||
|
||||
## Configure the outbound half
|
||||
|
||||
The outbound half lets LAN clients resolve `.fips` names and reach
|
||||
mesh destinations. Three operator decisions are involved: pool CIDR,
|
||||
DNS listen address, and how LAN clients learn the route to the pool
|
||||
and the resolver address.
|
||||
|
||||
### Choose the pool CIDR
|
||||
|
||||
```yaml
|
||||
gateway:
|
||||
pool: "fd01::/112"
|
||||
```
|
||||
|
||||
Constraints:
|
||||
|
||||
- Must not overlap with `fd00::/8` (the FIPS mesh address space).
|
||||
- Must not overlap with any LAN-side IPv6 prefix already in use.
|
||||
- `/112` is the practical width — wider just wastes address space
|
||||
because the pool is hard-capped at 65 536 entries. Narrower is
|
||||
fine if you want a smaller pool, but you'll reject DNS lookups
|
||||
faster under churn.
|
||||
|
||||
### Choose the DNS listen address
|
||||
|
||||
```yaml
|
||||
gateway:
|
||||
dns:
|
||||
listen: "[::1]:5353"
|
||||
upstream: "[::1]:5354"
|
||||
ttl: 60
|
||||
```
|
||||
|
||||
Common cases:
|
||||
|
||||
- **Another resolver on the host (the canonical case):** the default
|
||||
`listen: "[::1]:5353"` is loopback-only on an unprivileged port,
|
||||
so it never conflicts with dnsmasq, systemd-resolved, or BIND
|
||||
holding 53. Configure the existing resolver to forward `.fips`
|
||||
queries to `[::1]:5353` and you are done — this is what the
|
||||
OpenWrt ipk does automatically.
|
||||
- **No other resolver on the host:** set `listen: "[::]:53"`
|
||||
explicitly and LAN clients can query the gateway directly.
|
||||
- **systemd-resolved is on port 53:** the default already side-steps
|
||||
this — leave the listen address at `[::1]:5353` and configure the
|
||||
stub or a small forwarder to delegate `.fips` to the gateway. If
|
||||
you would rather have the gateway on 53 directly, disable the
|
||||
systemd stub listener (`DNSStubListener=no` in
|
||||
`/etc/systemd/resolved.conf`) and switch `listen` to `"[::]:53"`.
|
||||
See
|
||||
[troubleshoot-gateway.md](troubleshoot-gateway.md#port-conflict-on-the-dns-listen-port).
|
||||
- **Bind on the LAN address only:** `listen: "192.168.1.1:53"`
|
||||
exposes the resolver only to LAN clients, not loopback.
|
||||
|
||||
The gateway returns `REFUSED` for any non-`.fips` query — clients
|
||||
that point at it directly need a fallback resolver, or you should
|
||||
front it with a stub forwarder.
|
||||
|
||||
### Distribute the route to LAN clients
|
||||
|
||||
Each LAN client must route the gateway's pool CIDR to the gateway's
|
||||
LAN-side IPv6 address. Three options, in order of preference for
|
||||
production:
|
||||
|
||||
- **RA Route Information Option** (RFC 4191). If the LAN's RA daemon
|
||||
(`radvd`, `dnsmasq --enable-ra`, OpenWrt's `odhcpd`) supports
|
||||
publishing route options, configure it to advertise the pool CIDR
|
||||
with the gateway as next-hop. Clients pick this up automatically.
|
||||
|
||||
- **Static route on the LAN router**. If clients route through a
|
||||
central LAN router, add a static route entry there — the router
|
||||
then handles forwarding to the gateway. The exact syntax depends
|
||||
on the router OS.
|
||||
|
||||
- **Per-host static route** (testing or single-client deployments):
|
||||
|
||||
```sh
|
||||
sudo ip -6 route add fd01::/112 via fe80::<gateway-link-local>%<iface>
|
||||
# or, if the gateway has a stable global LAN address:
|
||||
sudo ip -6 route add fd01::/112 via <gateway-lan-addr>
|
||||
```
|
||||
|
||||
### Distribute the resolver to LAN clients
|
||||
|
||||
LAN clients also need to send `.fips` queries to the gateway. Two
|
||||
patterns:
|
||||
|
||||
- **Forward `.fips` from the LAN's main resolver.** If the LAN runs
|
||||
Pi-hole, Unbound, dnsmasq, or systemd-resolved as the central
|
||||
resolver, configure a conditional forward for `fips.`. Unbound
|
||||
example:
|
||||
|
||||
```text
|
||||
forward-zone:
|
||||
name: "fips."
|
||||
forward-addr: <gateway-lan-addr>@53
|
||||
```
|
||||
|
||||
dnsmasq example:
|
||||
|
||||
```text
|
||||
server=/fips/<gateway-lan-addr>
|
||||
```
|
||||
|
||||
Clients keep their existing DNS settings; only `.fips` queries are
|
||||
diverted.
|
||||
|
||||
- **Point clients directly at the gateway.** Simpler for testing,
|
||||
but the gateway returns `REFUSED` for non-`.fips` queries, so each
|
||||
client must also have a fallback resolver configured.
|
||||
|
||||
### Verify the outbound path
|
||||
|
||||
From a LAN client:
|
||||
|
||||
```sh
|
||||
dig @<gateway-lan-addr> hostname.fips AAAA
|
||||
# Expect an AAAA from the pool CIDR
|
||||
|
||||
ping6 hostname.fips
|
||||
# Should succeed via the gateway
|
||||
```
|
||||
|
||||
If either step fails, see
|
||||
[troubleshoot-gateway.md](troubleshoot-gateway.md#outbound-half-diagnostics).
|
||||
|
||||
## Configure the inbound half
|
||||
|
||||
The inbound half exposes a LAN-side service to mesh peers. Configured
|
||||
under `gateway.port_forwards[]`:
|
||||
|
||||
```yaml
|
||||
gateway:
|
||||
port_forwards:
|
||||
- listen_port: 8080
|
||||
proto: tcp
|
||||
target: "[fd12:3456::10]:80"
|
||||
- listen_port: 2222
|
||||
proto: tcp
|
||||
target: "[fd12:3456::20]:22"
|
||||
- listen_port: 5353
|
||||
proto: udp
|
||||
target: "[fd12:3456::10]:53"
|
||||
```
|
||||
|
||||
Field reference:
|
||||
|
||||
- `listen_port` — port on the gateway's `fips0` mesh-side address
|
||||
that mesh peers connect to. Must be non-zero. Each
|
||||
`(listen_port, proto)` pair must be unique across the list (the
|
||||
same port on TCP and UDP is allowed; the same port twice on the
|
||||
same proto is rejected at config-load time).
|
||||
- `proto` — `tcp` or `udp`.
|
||||
- `target` — IPv6 LAN destination as `[addr]:port`. IPv4 targets are
|
||||
rejected at parse time by the YAML deserializer (the field is
|
||||
typed `SocketAddrV6`). If the LAN host is reachable only by IPv4,
|
||||
put a small IPv6-aware reverse proxy in front of it on the gateway
|
||||
itself.
|
||||
|
||||
### Worked example: HTTP and DNS
|
||||
|
||||
Suppose the gateway runs on a LAN with an HTTP server at
|
||||
`[fd12:3456::10]:80` and a recursive resolver at
|
||||
`[fd12:3456::10]:53`, and you want mesh peers to reach them as
|
||||
`<gateway-npub>.fips:8080` (HTTP) and `<gateway-npub>.fips:5353`
|
||||
(DNS). Add to the gateway's `fips.yaml`:
|
||||
|
||||
```yaml
|
||||
gateway:
|
||||
port_forwards:
|
||||
- listen_port: 8080
|
||||
proto: tcp
|
||||
target: "[fd12:3456::10]:80"
|
||||
- listen_port: 5353
|
||||
proto: udp
|
||||
target: "[fd12:3456::10]:53"
|
||||
```
|
||||
|
||||
Reload:
|
||||
|
||||
```sh
|
||||
sudo systemctl restart fips-gateway
|
||||
```
|
||||
|
||||
From any mesh peer (the host name `gateway` is whatever the gateway's
|
||||
npub maps to in the local `hosts` file or via Nostr advert):
|
||||
|
||||
```sh
|
||||
curl http://gateway.fips:8080/
|
||||
dig @gateway.fips -p 5353 example.com A
|
||||
```
|
||||
|
||||
Each mesh-side request enters `fips0` on the listen port, gets DNAT'd
|
||||
to the LAN target, and the LAN-side masquerade rule rewrites the
|
||||
source to the gateway's LAN address so return traffic flows back
|
||||
through conntrack.
|
||||
|
||||
### Compose with the mesh firewall
|
||||
|
||||
`gateway.port_forwards[]` opens *mesh-side* listeners on `fips0`. If
|
||||
the host's mesh firewall is enabled (see
|
||||
[enable-mesh-firewall.md](enable-mesh-firewall.md)), inbound TCP/UDP
|
||||
on `fips0` for these ports must be permitted in the baseline or via
|
||||
a drop-in. The default baseline allows established/related and
|
||||
ICMPv6 only, so without an explicit allow rule, mesh peers will see
|
||||
TCP RSTs or silent drops on the listen port.
|
||||
|
||||
A typical drop-in for the worked example:
|
||||
|
||||
```nft
|
||||
# /etc/fips/fips.d/gateway-inbound.nft
|
||||
tcp dport 8080 accept
|
||||
udp dport 5353 accept
|
||||
```
|
||||
|
||||
Reload the firewall:
|
||||
|
||||
```sh
|
||||
sudo systemctl reload-or-restart fips-firewall.service
|
||||
```
|
||||
|
||||
If the inbound half doesn't need access control beyond the listen
|
||||
port itself, no source filter is needed. To restrict to specific
|
||||
mesh peers, follow the `ip6 saddr <addr> tcp dport <port> accept`
|
||||
pattern from the firewall guide.
|
||||
|
||||
### Verify the inbound path
|
||||
|
||||
From a mesh peer (any FIPS node):
|
||||
|
||||
```sh
|
||||
curl -v http://<gateway-npub>.fips:8080/
|
||||
```
|
||||
|
||||
A successful response confirms the full path: mesh ingress on
|
||||
`fips0`, DNAT to the LAN target, LAN-side masquerade, and conntrack-
|
||||
tracked return. If it fails, see
|
||||
[troubleshoot-gateway.md](troubleshoot-gateway.md#inbound-half-diagnostics).
|
||||
|
||||
## Operate and verify
|
||||
|
||||
`fips-gateway` exposes its own control socket at
|
||||
`/run/fips/gateway.sock`, separate from the daemon's
|
||||
`/run/fips/control.sock`. There is no `fipsctl gateway` subcommand —
|
||||
talk to it directly:
|
||||
|
||||
```sh
|
||||
echo '{"command":"show_gateway"}' | sudo nc -U /run/fips/gateway.sock
|
||||
echo '{"command":"show_mappings"}' | sudo nc -U /run/fips/gateway.sock
|
||||
```
|
||||
|
||||
`show_gateway` returns pool counters (`pool_total`, `pool_allocated`,
|
||||
`pool_active`, `pool_draining`, `pool_free`), `nat_mappings`,
|
||||
`dns_listen`, `uptime_secs`, and the active config snapshot.
|
||||
`show_mappings` returns the per-allocation list with virtual IP, mesh
|
||||
address, npub-derived `node_addr`, dns name, state (`Allocated`,
|
||||
`Active`, `Draining`), session count, and ages. For the full schema
|
||||
see [../reference/control-socket.md#gateway-command-catalog](../reference/control-socket.md#gateway-command-catalog).
|
||||
|
||||
The journal is the other primary signal:
|
||||
|
||||
```sh
|
||||
sudo systemctl status fips-gateway
|
||||
sudo journalctl -u fips-gateway -e
|
||||
```
|
||||
|
||||
Expect `MappingCreated`/`MappingRemoved` debug lines as DNS-driven
|
||||
allocations come and go (run with `--log-level debug` to see them),
|
||||
and `Final pool status` on shutdown. Errors in adding NAT rules or
|
||||
proxy-NDP entries surface here.
|
||||
|
||||
## See also
|
||||
|
||||
- [../tutorials/deploy-fips-gateway.md](../tutorials/deploy-fips-gateway.md) —
|
||||
the canonical, package-driven OpenWrt deployment path.
|
||||
- [../design/fips-gateway.md](../design/fips-gateway.md) — gateway
|
||||
design, NAT pipeline, virtual IP pool lifecycle, security
|
||||
considerations.
|
||||
- [Gateway section](../reference/configuration.md#gateway-gateway) of
|
||||
the configuration reference — full `gateway.*` block.
|
||||
- [../reference/cli-fips-gateway.md](../reference/cli-fips-gateway.md) —
|
||||
`fips-gateway` binary CLI flags.
|
||||
- [Gateway command catalog](../reference/control-socket.md#gateway-command-catalog)
|
||||
in the control-socket reference — JSON schema for `show_gateway`
|
||||
and `show_mappings`.
|
||||
- [troubleshoot-gateway.md](troubleshoot-gateway.md) — diagnostic
|
||||
recipes grouped by half.
|
||||
- [enable-mesh-firewall.md](enable-mesh-firewall.md) — mesh-firewall
|
||||
baseline and drop-ins (needed when exposing inbound ports).
|
||||
@@ -0,0 +1,217 @@
|
||||
# Deploy a Tor Onion Service for FIPS
|
||||
|
||||
This guide covers running a Tor onion service that accepts inbound
|
||||
FIPS peer connections.
|
||||
|
||||
For the Tor transport's design and the bridge-node pattern (running
|
||||
Tor and UDP simultaneously), see
|
||||
[../design/fips-transport-layer.md](../design/fips-transport-layer.md).
|
||||
For the full `transports.tor.*` config knob inventory, see
|
||||
[../reference/configuration.md](../reference/configuration.md).
|
||||
|
||||
## Inbound modes
|
||||
|
||||
FIPS supports two inbound Tor modes. (A third mode, `socks5`, is
|
||||
outbound-only and not covered here.)
|
||||
|
||||
- **`directory` mode** *(recommended)*. Tor manages the onion
|
||||
service via `HiddenServiceDir` and `HiddenServicePort` directives
|
||||
in `torrc`. FIPS reads the resulting `.onion` hostname from a
|
||||
file and binds a local TCP listener for Tor to forward inbound
|
||||
connections to. No control-port interaction is required, which
|
||||
makes this mode compatible with Tor's `Sandbox 1` seccomp-bpf
|
||||
hardening.
|
||||
- **`torrc` requires:** `HiddenServiceDir` + `HiddenServicePort`.
|
||||
- **`control_port` mode**. FIPS speaks to Tor's control port to
|
||||
create an ephemeral onion service at startup (`ADD_ONION`). The
|
||||
onion key lives only for the lifetime of the FIPS daemon's
|
||||
control-port session. This mode is **incompatible** with
|
||||
`Sandbox 1` — the sandbox forbids control-port-driven onion
|
||||
service management.
|
||||
- **`torrc` requires:** `ControlPort` (typically the Unix socket
|
||||
`/run/tor/control`) and a usable auth method
|
||||
(`CookieAuthentication 1` is the common choice).
|
||||
|
||||
Pick `directory` unless you have a specific reason to prefer
|
||||
`control_port`. The rest of this guide covers `directory` mode
|
||||
end-to-end.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- Tor daemon installed and running (Debian/Ubuntu: `apt install tor`)
|
||||
- FIPS daemon configured and able to start
|
||||
- Operator access to `/etc/tor/torrc` (or a drop-in under
|
||||
`/etc/tor/torrc.d/`)
|
||||
|
||||
## Step 1: Configure Tor's HiddenServiceDir
|
||||
|
||||
Add the following to `/etc/tor/torrc`:
|
||||
|
||||
```text
|
||||
HiddenServiceDir /var/lib/tor/fips
|
||||
HiddenServicePort 8443 127.0.0.1:8444
|
||||
```
|
||||
|
||||
`HiddenServiceDir` tells Tor where to store the onion service's
|
||||
private key and `hostname` file. `HiddenServicePort` declares that
|
||||
inbound TCP traffic to port 8443 of the onion address should be
|
||||
forwarded to `127.0.0.1:8444` on the local host — that is where FIPS
|
||||
will bind its listener.
|
||||
|
||||
The external port (`8443` here) is what peers will connect to over
|
||||
Tor; the internal target (`127.0.0.1:8444`) is purely local and is
|
||||
not directly reachable from the network.
|
||||
|
||||
## Step 2: Reload Tor and read the onion hostname
|
||||
|
||||
```sh
|
||||
sudo systemctl reload tor@default # or `tor` on systems without instance support
|
||||
```
|
||||
|
||||
After Tor processes the new config, the hostname file appears:
|
||||
|
||||
```sh
|
||||
sudo cat /var/lib/tor/fips/hostname
|
||||
# xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx.onion
|
||||
```
|
||||
|
||||
Tor regenerates the onion key only on first run (or if you remove
|
||||
`HiddenServiceDir`). The `hostname` value is stable across daemon
|
||||
restarts as long as `HiddenServiceDir` is preserved.
|
||||
|
||||
## Step 3: Verify HiddenServiceDir permissions
|
||||
|
||||
The directory must be readable only by the Tor user (Tor refuses to
|
||||
start otherwise):
|
||||
|
||||
```sh
|
||||
ls -la /var/lib/tor/fips
|
||||
# drwx------ debian-tor debian-tor ...
|
||||
```
|
||||
|
||||
With the shipped Debian systemd unit, FIPS runs as root and reads
|
||||
the `hostname` file directly — no permission adjustment is needed.
|
||||
|
||||
### Non-default deployments
|
||||
|
||||
If you run FIPS as an unprivileged user (custom packaging,
|
||||
hardened deployment, etc.), the FIPS daemon user needs read access
|
||||
to `hostname`. Options:
|
||||
|
||||
- Add the FIPS user to the `debian-tor` group and loosen group
|
||||
read on `HiddenServiceDir` (Tor still requires the directory
|
||||
itself to be `0700`, so this typically means making `hostname`
|
||||
itself group-readable rather than the directory).
|
||||
- Read `hostname` once at startup as root, then drop privileges.
|
||||
- Copy the hostname into a path the FIPS user can read, refreshed
|
||||
whenever the onion key changes.
|
||||
|
||||
## Step 4: Configure the FIPS Tor transport
|
||||
|
||||
In `/etc/fips/fips.yaml`, configure `transports.tor` with `mode:
|
||||
directory`:
|
||||
|
||||
```yaml
|
||||
transports:
|
||||
tor:
|
||||
mode: directory
|
||||
socks5_addr: "127.0.0.1:9050"
|
||||
connect_timeout_ms: 120000
|
||||
mtu: 1400
|
||||
advertised_port: 8443
|
||||
directory_service:
|
||||
hostname_file: "/var/lib/tor/fips/hostname"
|
||||
bind_addr: "127.0.0.1:8444"
|
||||
```
|
||||
|
||||
The `bind_addr` must match the *target* of the `HiddenServicePort`
|
||||
directive in `torrc`. The `hostname_file` path must match
|
||||
`HiddenServiceDir` plus `/hostname`.
|
||||
|
||||
`advertised_port` is the *virtual* onion port peers dial — i.e. the
|
||||
first number on the `HiddenServicePort` line, **not** the local
|
||||
target. The default is `443`; this guide uses `8443` on both sides
|
||||
to match the `HiddenServicePort 8443 127.0.0.1:8444` example
|
||||
above. Setting this explicitly is important if you ever flip
|
||||
`advertise_on_nostr: true`: the published advert otherwise
|
||||
defaults to `tor:<hash>.onion:443`, which won't match the actual
|
||||
onion port.
|
||||
|
||||
The `socks5_addr` is the Tor SOCKS5 proxy used for *outbound*
|
||||
connections to other onion services or clearnet endpoints (separate
|
||||
from inbound onion service handling).
|
||||
|
||||
Optional monitoring knobs: `control_addr` and `control_auth` (e.g.
|
||||
`/run/tor/control` and `cookie`) let the daemon read Tor's status
|
||||
through the control port even in `directory` mode. They are
|
||||
non-fatal on failure — the onion service still works without them.
|
||||
See [../reference/configuration.md](../reference/configuration.md)
|
||||
for the full key list and examples.
|
||||
|
||||
## Step 5: Reload the FIPS daemon
|
||||
|
||||
```sh
|
||||
sudo systemctl reload-or-restart fips
|
||||
```
|
||||
|
||||
At startup the daemon reads the `.onion` hostname from
|
||||
`hostname_file`, binds `127.0.0.1:8444`, and announces the onion
|
||||
endpoint internally. From this point inbound connections to
|
||||
`<your-onion>.onion:8443` arrive at FIPS over Tor.
|
||||
|
||||
## Step 6: Verify
|
||||
|
||||
Check that the FIPS daemon log shows the onion endpoint at startup:
|
||||
|
||||
```sh
|
||||
sudo journalctl -u fips -e | grep -i 'onion\|directory'
|
||||
```
|
||||
|
||||
You should see a line indicating the onion address FIPS will accept
|
||||
inbound connections on, and that the local bind on `127.0.0.1:8444`
|
||||
succeeded.
|
||||
|
||||
From another node configured with the Tor transport in `socks5` or
|
||||
`directory` mode, attempt to dial:
|
||||
|
||||
```sh
|
||||
fipsctl connect <peer-npub-or-hostname> <your-onion>.onion:8443 tor
|
||||
```
|
||||
|
||||
A successful `fipsctl show peers` afterwards on the inbound side
|
||||
shows the new peer with `transport=tor`.
|
||||
|
||||
## Optional: advertise the onion endpoint via Nostr discovery
|
||||
|
||||
If `node.discovery.nostr.enabled: true`, set
|
||||
`transports.tor.advertise_on_nostr: true` so the onion endpoint
|
||||
appears in this node's published advert. See
|
||||
[enable-nostr-discovery.md](enable-nostr-discovery.md) Scenario 2.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
- **Tor refuses to start with `Sandbox 1` and onion-service errors.**
|
||||
`Sandbox 1` requires `directory` mode and forbids creating onion
|
||||
services through the control port. Verify your `torrc` uses
|
||||
`HiddenServiceDir` (this guide), not `ADD_ONION` via control port.
|
||||
- **FIPS daemon fails to bind `127.0.0.1:8444`.** Another process is
|
||||
already bound to that port. Either stop the conflicting process or
|
||||
pick a different port and update both `torrc`'s
|
||||
`HiddenServicePort` target and `fips.yaml`'s `bind_addr` to match.
|
||||
- **Onion hostname is empty or missing.** Check `journalctl -u tor`
|
||||
for permission errors on `HiddenServiceDir`. The directory must be
|
||||
owned by the Tor user with mode `0700`.
|
||||
- **FIPS daemon cannot read `hostname_file`.** File is owned by the
|
||||
Tor user and not readable by the FIPS daemon user. Adjust
|
||||
permissions, or copy the hostname into a path the FIPS user can
|
||||
read.
|
||||
|
||||
## See also
|
||||
|
||||
- [../design/fips-transport-layer.md](../design/fips-transport-layer.md)
|
||||
— Tor transport design, three modes (`socks5`, `control_port`,
|
||||
`directory`), bridge-node pattern
|
||||
- [../reference/configuration.md](../reference/configuration.md) —
|
||||
full `transports.tor.*` configuration knob table
|
||||
- [enable-nostr-discovery.md](enable-nostr-discovery.md) — Scenario 2
|
||||
for advertising the onion endpoint to peers via Nostr
|
||||
@@ -0,0 +1,211 @@
|
||||
# Diagnose MTU Issues
|
||||
|
||||
MTU symptoms in FIPS look like ordinary network failures: handshakes
|
||||
succeed but bulk transfers hang, ssh connects but stalls after the
|
||||
banner, an HTTP request times out on the first response. This guide
|
||||
walks through the diagnostic surfaces that FIPS exposes so you can
|
||||
distinguish a real MTU problem from its frequent imposters
|
||||
(bufferbloat, transport saturation, transient packet loss).
|
||||
|
||||
For the underlying model — encapsulation overhead, proactive vs
|
||||
reactive PMTUD, the per-destination MTU storage layout — read
|
||||
[../design/fips-mtu.md](../design/fips-mtu.md) first.
|
||||
|
||||
## Symptom map
|
||||
|
||||
| Application symptom | Likely cause |
|
||||
| ------------------- | ------------ |
|
||||
| `iperf3 -c <host.fips>` control socket closes immediately after `Connecting to host`. | Forward-path MTU smaller than the negotiated MSS on the control connection. |
|
||||
| `ssh user@<host.fips>` shows the SSH banner then hangs forever. | First post-banner exchange exceeds the path MTU; SYN MSS clamp did not engage in time, or the path narrowed mid-session. |
|
||||
| `curl http://<host.fips>/` connects, then times out before the first response byte. | Same shape as the SSH-banner case, applied to the first server-to-client large packet. |
|
||||
| Throughput bursts then drops to zero, recovers, drops again, in seconds-long cycles. | Bufferbloat masquerading as MTU failure — usually the upload of the underlay link is saturated. See [Distinguishing bufferbloat](#distinguishing-bufferbloat-from-mtu-drops). |
|
||||
| `MtuExceeded` counters tick up under topology change but settle in seconds. | Normal: the reactive MTU mechanism doing its job. No action needed. |
|
||||
| `MtuExceeded` counters tick continuously under steady state. | Forward-path MTU smaller than what the source learned via `path_mtu` echo. After `mmp.path_mtu` has settled, this is a bug — see [File a bug](#file-a-bug). |
|
||||
|
||||
The first three are MTU candidates; the fourth is usually not. The
|
||||
fifth is benign. The sixth is the bug shape worth filing.
|
||||
|
||||
## Diagnostic toolkit
|
||||
|
||||
### `fipsctl show sessions`
|
||||
|
||||
The authoritative end-to-end MTU for an established session:
|
||||
|
||||
```sh
|
||||
fipsctl show sessions | jq '.sessions[] | {display_name, state, mmp: .mmp.path_mtu}'
|
||||
```
|
||||
|
||||
`mmp.path_mtu` is the value the session-layer MMP currently believes
|
||||
is in force end-to-end. It updates on each PathMtuNotification echo
|
||||
from the destination — immediately on decrease, with hysteresis on
|
||||
increase. A field that starts at `1280` (the IPv6 floor) and then
|
||||
climbs to a higher value as echoes arrive is healthy; one that
|
||||
oscillates between two values may indicate a flapping path.
|
||||
|
||||
### `fipsctl show transports`
|
||||
|
||||
Per-transport MTU. The `mtu` field reports the transport-wide
|
||||
default; for BLE, individual links may have a smaller negotiated
|
||||
ATT_MTU.
|
||||
|
||||
```sh
|
||||
fipsctl show transports | jq '.transports[] | {type, mtu}'
|
||||
```
|
||||
|
||||
### `fipsctl show cache`
|
||||
|
||||
The coordinate cache carries reverse-path-annotated MTU per
|
||||
destination — the freshest "what fit on the way back from the
|
||||
discovery target" estimate, consulted before the session has any
|
||||
PathMtuNotification feedback.
|
||||
|
||||
```sh
|
||||
fipsctl show cache | jq '.entries[] | {display_name, depth, path_mtu}'
|
||||
```
|
||||
|
||||
Entries without a `path_mtu` field are pre-discovery or were
|
||||
populated through a path that did not annotate the MTU.
|
||||
|
||||
### `fipsctl show peers`
|
||||
|
||||
Per-peer link state, including the link-layer MMP metrics. Useful
|
||||
mostly for ruling out underlying loss (loss rate near zero, SRTT
|
||||
sane) before chasing an MTU explanation.
|
||||
|
||||
```sh
|
||||
fipsctl show peers | jq '.peers[] | {display_name, mmp: .mmp}'
|
||||
```
|
||||
|
||||
### Trace logging
|
||||
|
||||
Module-scoped trace logging on the TUN reader and the MMP handler
|
||||
shows the per-packet decisions. The `tracing` macros default the
|
||||
target to the emitting module path, so the filter targets are the
|
||||
fully-qualified module paths under the `fips` crate.
|
||||
|
||||
```sh
|
||||
sudo systemctl edit fips
|
||||
# Add:
|
||||
# [Service]
|
||||
# Environment=RUST_LOG=info,fips::upper::tun=trace,fips::node::handlers::mmp=debug
|
||||
sudo systemctl restart fips
|
||||
sudo journalctl -u fips -f
|
||||
```
|
||||
|
||||
### tcpdump on `fips0`
|
||||
|
||||
Capturing on the TUN reveals the IPv6 packets the daemon hands the
|
||||
kernel and vice-versa. Two important caveats live in the design doc
|
||||
and are worth restating here:
|
||||
|
||||
- TX direction (outbound from a local app): tcpdump sees the packet
|
||||
**before** the daemon's TCP MSS clamp at the TUN boundary. The
|
||||
packet may be larger than the daemon will let leave the node.
|
||||
- RX direction (inbound to a local app): tcpdump sees the packet
|
||||
**after** the daemon's MSS clamp on inbound SYN-ACKs. The clamp
|
||||
fires only when `max_mss < kernel-natural-MSS`; otherwise it is a
|
||||
silent no-op.
|
||||
|
||||
```sh
|
||||
sudo tcpdump -ni fips0 -w /tmp/fips0.pcap port 22 or port 80
|
||||
# in another terminal, reproduce the symptom, then Ctrl-C
|
||||
```
|
||||
|
||||
Open the pcap in Wireshark and check segment sizes against what the
|
||||
session's `path_mtu` reports.
|
||||
|
||||
## Distinguishing bufferbloat from MTU drops
|
||||
|
||||
WAN bufferbloat (sustained upload saturation on a cable or DSL link)
|
||||
produces a retransmit signature that looks remarkably like
|
||||
oversized-packet drops. Both manifest as long stalls in TCP flows,
|
||||
both clear when you stop pushing data, both can ramp the loss-rate
|
||||
counter without obvious cause.
|
||||
|
||||
Two ways to disambiguate:
|
||||
|
||||
1. **Saturate the underlay first.** Run a reference upload outside
|
||||
FIPS (`iperf3 -c <internet-target>`) until it stabilises, then
|
||||
measure latency to the underlay's first hop with a separate `ping`.
|
||||
If RTT shoots up by hundreds of ms during the upload, the
|
||||
underlay buffer is the culprit, not FIPS MTU. Apply CAKE / fq_codel
|
||||
on the underlay router before continuing.
|
||||
|
||||
2. **Watch the FIPS counters during the symptom.** A real MTU
|
||||
problem ticks `MtuExceeded` (visible in `fipsctl show routing`'s
|
||||
`error_signals` block) and shifts the session's `mmp.path_mtu`
|
||||
downward. Bufferbloat ticks loss rate and RTT but leaves
|
||||
`path_mtu` and `MtuExceeded` alone.
|
||||
|
||||
If both signatures fire together, you have both problems.
|
||||
|
||||
## Cold-flow first-SYN
|
||||
|
||||
The MMP echo populates path-MTU state only after the first
|
||||
end-to-end exchange, but the TUN reader has to size the very first
|
||||
SYN before any echo has arrived. The cold-flow ceiling is the
|
||||
1143-byte conservative fallback derived from the 1280-byte IPv6
|
||||
floor. The first SYN may therefore be smaller than what the path
|
||||
ultimately supports; once MMP echoes arrive, subsequent flows use
|
||||
the larger learned value.
|
||||
|
||||
If the first SYN of a flow is still oversized relative to the path,
|
||||
the receiving transit node generates an `MtuExceeded`, the source
|
||||
shrinks immediately, and the next packet of the flow fits. This is
|
||||
expected for one round trip; it becomes a problem only if it
|
||||
persists.
|
||||
|
||||
## Fixes
|
||||
|
||||
The operator's choices, in rough order of preference:
|
||||
|
||||
### Pin a per-transport MTU floor in config
|
||||
|
||||
If a known link in the path has a small MTU that discovery does not
|
||||
pick up promptly (e.g., a Tor hop with an unusually tight cap), set
|
||||
a transport-level MTU floor on the relevant `transports.*` block.
|
||||
See [../reference/configuration.md](../reference/configuration.md)
|
||||
for the per-transport MTU keys.
|
||||
|
||||
### Tune host UDP buffers
|
||||
|
||||
For UDP transports specifically, undersized kernel buffers can drop
|
||||
oversized datagrams in a way that looks identical to MTU failure.
|
||||
See [tune-udp-buffers.md](tune-udp-buffers.md).
|
||||
|
||||
### Accept the floor on intrinsically small links
|
||||
|
||||
Tor and BLE link MTUs are properties of the medium, not tunables.
|
||||
For sessions that cross those links, the path MTU will be small; the
|
||||
fix is to design applications around it (smaller TCP windows, fewer
|
||||
large RTTs) rather than fight the transport.
|
||||
|
||||
### File a bug
|
||||
|
||||
The bug shape worth filing is session `mmp.path_mtu` itself
|
||||
oscillating, or `MtuExceeded` ticking *within* an established
|
||||
session after `mmp.path_mtu` has settled. The TCP-clamp mirror
|
||||
(`path_mtu_lookup`) is now updated on every successful proactive
|
||||
`PathMtuNotification` apply (tighter-only) as well as by the
|
||||
reactive `MtuExceeded` handler, so a steady-state divergence
|
||||
between the per-session `mmp.path_mtu` and the mirror used for
|
||||
new TCP flows is itself a defect, not an expected behavior.
|
||||
|
||||
Capture `fipsctl show sessions`, `fipsctl show cache`, `fipsctl
|
||||
show routing` (for the `error_signals` block), and a tcpdump from
|
||||
`fips0` covering the symptom window. See
|
||||
[../design/fips-mtu.md](../design/fips-mtu.md#per-destination-mtu-storage)
|
||||
for the per-destination MTU storage layout.
|
||||
|
||||
## See also
|
||||
|
||||
- [../design/fips-mtu.md](../design/fips-mtu.md) — encapsulation
|
||||
overhead, the proactive `path_mtu` field, the reactive
|
||||
`MtuExceeded` mechanism, MSS clamping, the no-fragmentation
|
||||
policy.
|
||||
- [../design/fips-mmp.md](../design/fips-mmp.md) — what the MMP
|
||||
metrics mean and how they are computed.
|
||||
- [../design/fips-ipv6-adapter.md](../design/fips-ipv6-adapter.md) —
|
||||
TUN-side ICMPv6 PTB generation and the MSS clamp.
|
||||
- [tune-udp-buffers.md](tune-udp-buffers.md) — host sysctl recipes
|
||||
that rule out kernel-buffer drops as a confounder.
|
||||
@@ -0,0 +1,192 @@
|
||||
# Enable the Mesh-Interface Firewall
|
||||
|
||||
FIPS ships a default-deny nftables baseline at `/etc/fips/fips.nft` that
|
||||
restricts inbound traffic on the `fips0` mesh interface to conntrack
|
||||
replies and ICMPv6 echo. The baseline is **not** enabled by default — see
|
||||
[../design/fips-security.md](../design/fips-security.md) for the threat
|
||||
model and the rationale behind keeping activation explicit. This guide
|
||||
covers the operator steps to load the baseline, extend it with per-host
|
||||
allowances, and inspect drops.
|
||||
|
||||
## Activate the baseline
|
||||
|
||||
The package ships `fips-firewall.service`, a systemd oneshot that runs
|
||||
`nft -f /etc/fips/fips.nft` on start and removes the `inet fips` table
|
||||
on stop. To activate:
|
||||
|
||||
```sh
|
||||
sudo systemctl enable --now fips-firewall.service
|
||||
```
|
||||
|
||||
This loads the table now and arranges for it to load on every subsequent
|
||||
boot. To disable and tear it down:
|
||||
|
||||
```sh
|
||||
sudo systemctl disable --now fips-firewall.service
|
||||
```
|
||||
|
||||
To reload after editing `/etc/fips/fips.nft` or adding a drop-in under
|
||||
`/etc/fips/fips.d/`:
|
||||
|
||||
```sh
|
||||
sudo systemctl reload-or-restart fips-firewall.service
|
||||
```
|
||||
|
||||
The file is idempotent — it begins with `add table inet fips; flush
|
||||
table inet fips;` so re-running it replaces the live ruleset atomically.
|
||||
Equivalently:
|
||||
|
||||
```sh
|
||||
sudo nft -f /etc/fips/fips.nft
|
||||
```
|
||||
|
||||
## Folding the baseline into the host's main nftables
|
||||
|
||||
If you prefer to load the baseline from your existing
|
||||
`/etc/nftables.conf` rather than via the systemd unit, include it
|
||||
directly:
|
||||
|
||||
```nft
|
||||
# in /etc/nftables.conf
|
||||
include "/etc/fips/fips.nft"
|
||||
```
|
||||
|
||||
In that case do **not** enable `fips-firewall.service` — the host's main
|
||||
nftables setup owns the loading. The two paths are mutually exclusive.
|
||||
|
||||
## Extend with per-host allowances via drop-ins
|
||||
|
||||
The baseline drops everything inbound on `fips0` except conntrack
|
||||
replies and ICMPv6 echo. To open specific services to specific mesh
|
||||
nodes, drop a file into `/etc/fips/fips.d/` ending in `.nft`. Each
|
||||
file is included inline into the `inbound` chain at the marked point
|
||||
and may contain any nftables rule lines valid in that context.
|
||||
|
||||
Reload after editing:
|
||||
|
||||
```sh
|
||||
sudo systemctl reload-or-restart fips-firewall.service
|
||||
# or: sudo nft -f /etc/fips/fips.nft
|
||||
```
|
||||
|
||||
### Allow inbound SSH from a specific mesh node
|
||||
|
||||
```nft
|
||||
# /etc/fips/fips.d/ssh-from-bastion.nft
|
||||
ip6 saddr fd97:1234:5678:9abc:def0:1234:5678:9abc tcp dport 22 accept
|
||||
```
|
||||
|
||||
The source filter is the node's mesh address. To find a node's mesh
|
||||
address, look in their `fips.pub` (which contains the npub) and derive
|
||||
the `fd97:...` address from it, or query the running daemon:
|
||||
|
||||
```sh
|
||||
fipsctl show identity-cache
|
||||
fipsctl show peers
|
||||
```
|
||||
|
||||
### Allow inbound DNS broadly
|
||||
|
||||
Some services need to be reachable from any mesh node (a public DNS
|
||||
resolver, a public bootstrap node):
|
||||
|
||||
```nft
|
||||
# /etc/fips/fips.d/dns-public.nft
|
||||
udp dport 53 accept
|
||||
tcp dport 53 accept
|
||||
```
|
||||
|
||||
Omit the source filter only when the service is intended to be
|
||||
universally reachable on the mesh. The baseline's purpose is to make
|
||||
"universally reachable" an explicit decision rather than the default.
|
||||
|
||||
### Multiple nodes, one service
|
||||
|
||||
```nft
|
||||
# /etc/fips/fips.d/git-from-trusted.nft
|
||||
ip6 saddr {
|
||||
fd97:1111:2222:3333:4444:5555:6666:7777,
|
||||
fd97:8888:9999:aaaa:bbbb:cccc:dddd:eeee
|
||||
} tcp dport 9418 accept
|
||||
```
|
||||
|
||||
Set syntax keeps multi-node rules readable and is more efficient than a
|
||||
chain of individual rules.
|
||||
|
||||
## Verify with fipstop
|
||||
|
||||
`fipstop`'s Node tab carries a **Listening on fips0** panel
|
||||
(right-half of the Traffic block) that pairs each local IPv6
|
||||
listener with its current baseline-filter classification. After
|
||||
adding or editing a drop-in and reloading, this is the fastest
|
||||
way to confirm the rule landed correctly without manually
|
||||
parsing `nft list table inet fips`.
|
||||
|
||||
| Panel state | Reading |
|
||||
| ----------- | ------- |
|
||||
| Service row in **default White** with `OPEN` in the State column | The chain has a canonical, unrestricted accept rule for this (proto, port). The service is reachable from any mesh node. |
|
||||
| Service row in **DarkGray** with `filt` | No matching accept rule; the chain falls through to `counter drop`. The service is not reachable from the mesh. |
|
||||
| Service row in **DarkGray** with `filt?` | A rule references the port but uses matchers the panel cannot fully decompose (saddr filter, jump, daddr filter). The intent is operator-defined; inspect with `sudo nft list table inet fips` to see the actual rule. |
|
||||
| **Yellow banner** above the panel: "fips-firewall.service inactive — all listeners exposed" | The `inet fips` table is not loaded. Every listener is mesh-reachable (subject only to whatever ACL you have at the peer layer). |
|
||||
|
||||
A common workflow when extending the baseline is to keep `fipstop`
|
||||
open on the Node tab in one terminal while editing
|
||||
`/etc/fips/fips.d/` in another. After each
|
||||
`sudo systemctl reload-or-restart fips-firewall.service`, the panel
|
||||
re-classifies on the next poll tick and the affected row's State
|
||||
column flips. A row staying `filt` after you expected `OPEN`
|
||||
usually means the drop-in failed to load (syntax error in any file
|
||||
under `/etc/fips/fips.d/` aborts the whole reload) or carries a
|
||||
saddr filter that triggers `filt?` rather than `OPEN`.
|
||||
|
||||
The classifier is conservative: it recognizes only the canonical
|
||||
unrestricted shapes (`tcp dport N accept`, `udp dport N accept`,
|
||||
`dport { ... } accept`, `dport A-B accept`). Source-restricted
|
||||
accepts intentionally render as `filt?` rather than `OPEN` —
|
||||
the panel is a security screen, and any rule that varies by
|
||||
source is an operator decision the panel will not silently bless
|
||||
as fully open.
|
||||
|
||||
## Inspect drops
|
||||
|
||||
The baseline counter increments on every dropped packet. Inspect it:
|
||||
|
||||
```sh
|
||||
sudo nft list table inet fips
|
||||
```
|
||||
|
||||
Look for the `counter packets N bytes M drop` line at the bottom of the
|
||||
`inbound` chain. A non-zero counter means mesh nodes are sending
|
||||
traffic that hits the default-deny — usually benign (probes, neighbor
|
||||
discovery) but occasionally a misconfigured drop-in.
|
||||
|
||||
To see which packets are being dropped, uncomment the `log` line near
|
||||
the bottom of `/etc/fips/fips.nft`:
|
||||
|
||||
```nft
|
||||
log prefix "fips drop: " level info limit rate 10/minute
|
||||
```
|
||||
|
||||
Reload:
|
||||
|
||||
```sh
|
||||
sudo nft -f /etc/fips/fips.nft
|
||||
```
|
||||
|
||||
Then tail the kernel log:
|
||||
|
||||
```sh
|
||||
sudo journalctl -k -f -g "fips drop:"
|
||||
```
|
||||
|
||||
The rate-limit prevents flooding the journal under sustained probing.
|
||||
Adjust the rate, log level, or prefix as needed for the situation.
|
||||
Re-comment the rule when you are done; production hosts do not need
|
||||
the log line on by default.
|
||||
|
||||
## See also
|
||||
|
||||
- [../design/fips-security.md](../design/fips-security.md) — threat
|
||||
model, baseline design, and coexistence with other firewalls
|
||||
- [../reference/security.md](../reference/security.md) — consolidated
|
||||
security reference
|
||||
@@ -0,0 +1,361 @@
|
||||
# Enable Nostr-Mediated Discovery and NAT Traversal
|
||||
|
||||
Nostr-mediated discovery lets FIPS nodes find each other (and punch
|
||||
through UDP NAT) using public Nostr relays as the signaling channel.
|
||||
The feature ships in every stock packaging artifact but is **off by
|
||||
default** — it activates when an operator sets
|
||||
`node.discovery.nostr.enabled: true`. Default relay and STUN-server
|
||||
lists ship in the config; both are optional overrides. See
|
||||
[../design/fips-nostr-discovery.md](../design/fips-nostr-discovery.md)
|
||||
for the design and rationale; see
|
||||
[../reference/configuration.md](../reference/configuration.md) for the
|
||||
full knob inventory.
|
||||
|
||||
Nostr discovery provides three independent capabilities. They can be
|
||||
enabled separately; most deployments end up using two or three of
|
||||
them together.
|
||||
|
||||
1. **Resolve a known peer's address by npub.** Your daemon consumes
|
||||
adverts from the relays to look up the current network endpoint
|
||||
for a peer you have configured by npub. You don't have to know
|
||||
their IP / port / transport in advance.
|
||||
2. **Publish your own endpoint so others can resolve you.** Your
|
||||
daemon publishes a signed advert listing the transports it will
|
||||
accept connections on. Has two sub-shapes depending on your
|
||||
network topology: UDP (using NAT traversal if needed) or TCP.
|
||||
Running a Tor onion service is a separate deployment mode,
|
||||
covered in its own section below.
|
||||
3. **Discover peers without prior configuration.** Your daemon
|
||||
subscribes to all adverts on a chosen application namespace and
|
||||
treats any publisher as a connection candidate. The most
|
||||
permissive posture; useful for ambient mesh participation.
|
||||
|
||||
Each capability is covered below as one or more scenarios with the
|
||||
minimal YAML fragment that enables it. Only keys relevant to Nostr
|
||||
discovery are shown; surrounding node, transport, TUN, DNS, and peer
|
||||
configuration follows the usual shape.
|
||||
|
||||
All scenarios assume `node.identity` is set to a persistent key — an
|
||||
ephemeral identity would invalidate any advert the moment the node
|
||||
restarts. See [persistent-identity.md](persistent-identity.md) for
|
||||
the persistent-key setup.
|
||||
|
||||
For hand-held walkthroughs of each capability, see the
|
||||
[resolve-peers-via-nostr](../tutorials/resolve-peers-via-nostr.md),
|
||||
[advertise-your-node](../tutorials/advertise-your-node.md),
|
||||
and [open-discovery](../tutorials/open-discovery.md)
|
||||
tutorials.
|
||||
|
||||
## Capability 1: Resolve a known peer's address by npub
|
||||
|
||||
The node does not publish any advert of its own. It only consumes
|
||||
adverts for peers it has explicitly listed with `via_nostr: true`.
|
||||
This is the right shape for a client that wants Nostr-mediated
|
||||
resolution without becoming a rendezvous target itself.
|
||||
|
||||
```yaml
|
||||
node:
|
||||
identity:
|
||||
persistent: true
|
||||
discovery:
|
||||
nostr:
|
||||
enabled: true
|
||||
advertise: false
|
||||
policy: configured_only
|
||||
|
||||
transports:
|
||||
udp:
|
||||
bind_addr: "0.0.0.0:2121"
|
||||
|
||||
peers:
|
||||
- npub: "npub1peer..."
|
||||
alias: "remote-node"
|
||||
via_nostr: true
|
||||
connect_policy: auto_connect
|
||||
```
|
||||
|
||||
What this achieves: dial endpoints for this peer are taken from
|
||||
the peer's published Nostr advert. `configured_only` is the
|
||||
default — it is shown here for clarity.
|
||||
|
||||
> **Note:** You can also supply a static address alongside
|
||||
> `via_nostr: true` (for example, while testing, or as a
|
||||
> known-good fallback if the advert is stale). Add an `addresses`
|
||||
> block to the peer entry; static addresses are tried first on
|
||||
> dial and Nostr-resolved endpoints are appended as additional
|
||||
> candidates.
|
||||
|
||||
## Capability 2: Publish your own endpoint so others can resolve you
|
||||
|
||||
This capability has three sub-scenarios depending on the network
|
||||
shape your node sits behind.
|
||||
|
||||
### Sub-scenario 2a: UDP (using NAT traversal if needed)
|
||||
|
||||
The node has a public IP (or a stable port-forward) and binds UDP on
|
||||
a known port. It publishes `udp:host:port` to the advert relays. Any
|
||||
peer that knows this node's npub and has Nostr discovery enabled can
|
||||
dial it without knowing the address out-of-band.
|
||||
|
||||
When UDP is wildcard-bound (`0.0.0.0:2121`, the default), the daemon
|
||||
needs help knowing what IP to put in the advert. There are two ways:
|
||||
STUN auto-discovery (`public: true`) or an explicit override
|
||||
(`external_addr`). Both are first-class options; pick the one that
|
||||
fits the deployment.
|
||||
|
||||
```yaml
|
||||
node:
|
||||
identity:
|
||||
persistent: true
|
||||
discovery:
|
||||
nostr:
|
||||
enabled: true
|
||||
advertise: true
|
||||
|
||||
transports:
|
||||
udp:
|
||||
bind_addr: "0.0.0.0:2121"
|
||||
advertise_on_nostr: true
|
||||
public: true # ← STUN auto-discovery
|
||||
```
|
||||
|
||||
Or, when the public IP is known up front (static residential IP,
|
||||
cloud Elastic IP behind 1:1 NAT, etc.):
|
||||
|
||||
```yaml
|
||||
transports:
|
||||
udp:
|
||||
bind_addr: "0.0.0.0:2121"
|
||||
advertise_on_nostr: true
|
||||
public: true # ← required, master switch
|
||||
external_addr: "203.0.113.45:2121" # ← explicit address
|
||||
```
|
||||
|
||||
`external_addr` accepts a bare IP (combined with the bind port) or a
|
||||
full `host:port`. `public: true` is the master switch that gates UDP
|
||||
advertisement; inside that branch, the daemon picks the advertised
|
||||
address in precedence order: explicit `external_addr` (no STUN
|
||||
observation), a non-wildcard `bind_addr`, or STUN auto-discovery.
|
||||
Setting `external_addr` alongside `public: true` skips STUN entirely
|
||||
— there is no logging cross-check. If UDP is bound directly to a
|
||||
public IP rather than to a wildcard, neither `external_addr` nor STUN
|
||||
is needed — but `advertise_on_nostr: true` and `public: true` are
|
||||
still both required for the daemon to publish the endpoint.
|
||||
|
||||
What this achieves: the node publishes a single
|
||||
`udp:<public-ip>:2121` endpoint to the three default advert relays
|
||||
(`wss://relay.damus.io`, `wss://nos.lol`, `wss://offchain.pub`).
|
||||
|
||||
What the other side needs: either a static `addresses` entry for this
|
||||
peer, or a peer entry with `via_nostr: true` and an empty (or
|
||||
omitted) `addresses` list — the advert-resolved endpoint will be used
|
||||
at dial time. Static and Nostr-resolved addresses can also be
|
||||
combined: when both are present, static addresses are tried first and
|
||||
Nostr-resolved endpoints are appended as fallback.
|
||||
|
||||
#### When the node is behind NAT
|
||||
|
||||
If this node doesn't have a stable public UDP endpoint, advertise
|
||||
`udp:nat`. The daemon runs the STUN + offer/answer exchange with
|
||||
the peer and punches through the NAT to establish a direct UDP
|
||||
link. The peer can either have a public endpoint of its own or
|
||||
also be behind NAT — both shapes work, as long as at least one
|
||||
side has a NAT type compatible with hole-punching.
|
||||
|
||||
```yaml
|
||||
node:
|
||||
identity:
|
||||
persistent: true
|
||||
discovery:
|
||||
nostr:
|
||||
enabled: true
|
||||
advertise: true
|
||||
dm_relays: # overrides the default three-relay
|
||||
- "wss://relay.damus.io" # set with two for demonstration;
|
||||
- "wss://nos.lol" # omit this block to keep the defaults
|
||||
stun_servers:
|
||||
- "stun:stun.l.google.com:19302"
|
||||
- "stun:stun.cloudflare.com:3478"
|
||||
|
||||
transports:
|
||||
udp:
|
||||
bind_addr: "0.0.0.0:2121"
|
||||
advertise_on_nostr: true
|
||||
public: false
|
||||
|
||||
peers:
|
||||
- npub: "npub1peer..."
|
||||
alias: "nat-peer"
|
||||
addresses:
|
||||
- transport: udp
|
||||
addr: "nat"
|
||||
via_nostr: true
|
||||
connect_policy: auto_connect
|
||||
```
|
||||
|
||||
What this achieves: the node publishes a `udp:nat` endpoint plus its
|
||||
signaling relays in the advert. When either side initiates, an
|
||||
encrypted offer is sealed to the peer's npub, a matching answer
|
||||
comes back, and both sides punch at the negotiated time. On success,
|
||||
the punch socket is adopted as an FMP UDP transport and Noise IK
|
||||
proceeds normally.
|
||||
|
||||
> **Validation:** `advertise_on_nostr: true` with `public: false` on
|
||||
> UDP requires `dm_relays` and `stun_servers` to be non-empty. Both
|
||||
> ship with non-empty defaults (three relays and three STUN servers
|
||||
> respectively), so the default config passes. The node fails
|
||||
> startup only if the operator has explicitly emptied either list —
|
||||
> a `udp:nat` advert without signaling relays or STUN servers is
|
||||
> unreachable by construction.
|
||||
|
||||
Hole-punching is best-effort. It works reliably when both sides are
|
||||
full-cone or port-restricted NATs. Symmetric NAT on either side
|
||||
typically defeats the punch — the public port a peer sees varies per
|
||||
remote endpoint, so the address learned via STUN does not match the
|
||||
mapping the peer actually needs. The punch attempt times out after
|
||||
`punch_duration_ms`. `udp:nat` is the only NAT-traversal mechanism
|
||||
in FIPS; when it can't succeed, there's no in-protocol substitute.
|
||||
Being reachable then becomes a deployment-prerequisite question
|
||||
rather than a transport question — a publicly reachable port (UDP
|
||||
or TCP — both require the same kind of network resource) published
|
||||
as a direct advert per Sub-scenario 2a or 2b.
|
||||
|
||||
### Sub-scenario 2b: TCP
|
||||
|
||||
The node has a public IP (or a stable port-forward) and accepts
|
||||
inbound TCP. It publishes `tcp:host:port` to the advert relays.
|
||||
|
||||
TCP endpoints exist to serve peers whose networks filter outbound
|
||||
UDP (corporate LANs, restrictive guest WiFi). NAT traversal does
|
||||
not apply: the publishing node is publicly reachable on TCP, and
|
||||
the dialing peer's network only needs to permit outbound TCP to
|
||||
the advertised port.
|
||||
|
||||
```yaml
|
||||
node:
|
||||
identity:
|
||||
persistent: true
|
||||
discovery:
|
||||
nostr:
|
||||
enabled: true
|
||||
advertise: true
|
||||
|
||||
transports:
|
||||
tcp:
|
||||
bind_addr: "0.0.0.0:8443"
|
||||
advertise_on_nostr: true
|
||||
external_addr: "203.0.113.45:8443"
|
||||
```
|
||||
|
||||
`external_addr` is typically required on cloud setups (AWS Elastic
|
||||
IP, etc.) where binding directly to the public IP returns
|
||||
`EADDRNOTAVAIL`. When TCP is bound directly to a public IP, the
|
||||
override is unnecessary.
|
||||
|
||||
What this achieves: the node publishes a `tcp:<public-ip>:8443`
|
||||
endpoint to the advert relays. Peers with Nostr discovery enabled
|
||||
dial by npub without out-of-band address exchange.
|
||||
|
||||
### Tor onion node
|
||||
|
||||
A separate deployment mode for nodes that want anonymity and
|
||||
censorship-resistance properties on the data plane. Functionally
|
||||
this still uses Capability 2 (publishing an endpoint to advert
|
||||
relays) — the difference is that the published endpoint is a Tor
|
||||
hidden service rather than a public IP.
|
||||
|
||||
The node runs a Tor onion service in directory mode (Tor-managed
|
||||
`HiddenServiceDir`) and advertises the `.onion` address. Peers dial
|
||||
via their local Tor SOCKS5 proxy without ever knowing the onion
|
||||
string out-of-band. For the Tor daemon side of this setup, including the inbound-mode
|
||||
trade-offs and the `torrc` directives each requires, see
|
||||
[deploy-tor-onion.md](deploy-tor-onion.md).
|
||||
|
||||
```yaml
|
||||
node:
|
||||
identity:
|
||||
persistent: true
|
||||
discovery:
|
||||
nostr:
|
||||
enabled: true
|
||||
advertise: true
|
||||
|
||||
transports:
|
||||
tor:
|
||||
mode: directory
|
||||
socks5_addr: "127.0.0.1:9050"
|
||||
directory_service:
|
||||
hostname_file: "/var/lib/tor/fips/hostname"
|
||||
bind_addr: "127.0.0.1:8444"
|
||||
advertise_on_nostr: true
|
||||
```
|
||||
|
||||
What this achieves: the node publishes a `tor:<hash>.onion:8443`
|
||||
endpoint alongside any other advertised transports. The advert itself
|
||||
is still published over clearnet WebSocket relays — Tor protects the
|
||||
data plane, not the discovery plane. See the security and threat
|
||||
model section in
|
||||
[../design/fips-nostr-discovery.md](../design/fips-nostr-discovery.md#security-and-threat-model)
|
||||
for the trade-off and how to route relay traffic through Tor as well.
|
||||
|
||||
## Capability 3: Discover peers without prior configuration
|
||||
|
||||
Under `policy: open`, any node that publishes an advert under the
|
||||
same `app` namespace becomes a candidate. Discovered peers are queued
|
||||
for connection attempts subject to `open_discovery_max_pending`.
|
||||
|
||||
```yaml
|
||||
node:
|
||||
identity:
|
||||
persistent: true
|
||||
discovery:
|
||||
nostr:
|
||||
enabled: true
|
||||
advertise: true
|
||||
policy: open
|
||||
open_discovery_max_pending: 64
|
||||
app: "my-experiment.v1"
|
||||
|
||||
transports:
|
||||
udp:
|
||||
bind_addr: "0.0.0.0:2121"
|
||||
advertise_on_nostr: true
|
||||
public: true
|
||||
|
||||
peers: []
|
||||
```
|
||||
|
||||
What this achieves: peers are discovered entirely through ambient
|
||||
advert traffic on the configured relays. Setting a non-default `app`
|
||||
value (replacing `fips-overlay-v1`) scopes the discovery set to
|
||||
participants who opt into the same experiment and avoids being joined
|
||||
to unrelated overlays that happen to share the default namespace.
|
||||
|
||||
> **Scope warning:** Open discovery is an admission-free mode. Any
|
||||
> node that publishes on the same `app` name and passes the peer-ACL
|
||||
> check becomes a connection candidate. If you rely on peer ACLs for
|
||||
> admission control, verify that list is set correctly before
|
||||
> enabling this mode. See
|
||||
> [../reference/security.md](../reference/security.md) for the peer
|
||||
> ACL format.
|
||||
|
||||
## See also
|
||||
|
||||
- [../tutorials/resolve-peers-via-nostr.md](../tutorials/resolve-peers-via-nostr.md)
|
||||
— hand-held walkthrough of capability 1
|
||||
- [../tutorials/advertise-your-node.md](../tutorials/advertise-your-node.md)
|
||||
— hand-held walkthrough of capability 2 (publish, plus a
|
||||
short section on `udp:nat` NAT traversal)
|
||||
- [../tutorials/open-discovery.md](../tutorials/open-discovery.md)
|
||||
— hand-held walkthrough of capability 3 (open ambient
|
||||
discovery, the additive policy: open mode)
|
||||
- [../design/fips-nostr-discovery.md](../design/fips-nostr-discovery.md)
|
||||
— discovery runtime design, security model
|
||||
- [../reference/configuration.md](../reference/configuration.md) —
|
||||
full `node.discovery.nostr.*` and per-transport
|
||||
`advertise_on_nostr`/`public` table
|
||||
- [../reference/nostr-events.md](../reference/nostr-events.md) — Kind
|
||||
37195 advert format, Kind 21059 traversal signaling, Kind 10050
|
||||
inbox relay list
|
||||
- [deploy-tor-onion.md](deploy-tor-onion.md) — Tor daemon-side setup
|
||||
for advertising onion endpoints
|
||||
@@ -0,0 +1,173 @@
|
||||
# Use Shortnames Instead of Long Npubs
|
||||
|
||||
A FIPS node's canonical address is `<npub>.fips`. The npub is
|
||||
63 characters of bech32 — fine for the daemon, awkward to type
|
||||
or fit in a docs example. The local DNS resolver consults a
|
||||
host map before falling back to direct-npub resolution, so
|
||||
short names like `test-us01.fips` work as substitutes wherever
|
||||
`<npub>.fips` would.
|
||||
|
||||
This guide covers the two ways to populate that map and when
|
||||
to use which.
|
||||
|
||||
## When to use which
|
||||
|
||||
Two independent mechanisms feed the same DNS responder:
|
||||
|
||||
| Mechanism | Source | Scope | Reload |
|
||||
|-----------|--------|-------|--------|
|
||||
| Hosts file | `/etc/fips/hosts` | Node-local, intended for shared rosters | Auto on mtime change |
|
||||
| Peer alias | `alias:` field on a `peers:` entry | Node-local, scoped to configured peers | Daemon restart |
|
||||
|
||||
Pick the hosts file when:
|
||||
|
||||
- The shortname refers to a peer your operator-team agrees to
|
||||
call by that name across machines (the public test mesh
|
||||
ships this way).
|
||||
- You want the destination's name to resolve in DNS or appear
|
||||
in `fipsctl show peers` display even though it isn't in your
|
||||
`peers:` block — e.g., a mesh node you reach transitively
|
||||
through your direct peers. The hosts-file entry is for name
|
||||
resolution and display only; it does not stand in for the
|
||||
npub a peer-config entry requires.
|
||||
|
||||
Pick the peer alias when:
|
||||
|
||||
- The shortname is just a label *you* use locally for a peer
|
||||
that's already in your `peers:` block.
|
||||
- You want the alias to live with the rest of the peer config
|
||||
(one place to look) rather than in a separate file.
|
||||
|
||||
The two coexist. If both reference the same shortname,
|
||||
`/etc/fips/hosts` wins — the file is treated as the
|
||||
authoritative shared roster.
|
||||
|
||||
## What ships in the default `/etc/fips/hosts`
|
||||
|
||||
The installer drops `/etc/fips/hosts` populated with the
|
||||
public test mesh roster:
|
||||
|
||||
```text
|
||||
test-us01 npub1qmc3cvfz0yu2hx96nq3gp55zdan2qclealn7xshgr448d3nh6lks7zel98
|
||||
test-us02 npub10yffd020a4ag8zcy75f9pruq3rnghvvhd5hphl9s62zgp35s560qrksp9u
|
||||
test-us03 npub136yqae6na688fs75g95ppps3lxe07fvxefj77938zf47uhm6074sxw8ctm
|
||||
test-us03-next npub15m6c4ghuegx4pcde6tra8f7smn8vfv2wundyxwhkjynuerkrzmgsy09sh3
|
||||
test-us04 npub1gd7ye2qp2lphhzx75fynnjzaxx4dqanddecet0wtt5ss5ek8h9ps62wdkf
|
||||
test-de01 npub1260n42s06vzc7796w0fh3ny7zcpw6tlk4gq3940gmfrzl5c9pv2s3657q8
|
||||
test-es01 npub17lpmzulpc98d8ff727k6e98atxn3phzupzsqqwe54ytduym747ws4tw5zm
|
||||
test-uk01 npub1u0z26dc4qeneu5rvwvmpfhtwh3522ed6rlgxr9jarrfnjrc6ew4qxjysrs
|
||||
```
|
||||
|
||||
These resolve out of the box — `ping6 test-us01.fips` works
|
||||
even before you've added any peer to your config, as long as
|
||||
the destination is reachable through your mesh links.
|
||||
|
||||
If you don't intend to interact with the public test mesh,
|
||||
the entries are safe to comment out or delete. They are
|
||||
plain hosts-file lines, not protocol participants — removing
|
||||
them only changes name resolution on your machine.
|
||||
|
||||
## Add an entry to `/etc/fips/hosts`
|
||||
|
||||
Append a line to `/etc/fips/hosts`:
|
||||
|
||||
```text
|
||||
my-laptop npub1abc...xyz
|
||||
```
|
||||
|
||||
Format rules:
|
||||
|
||||
- One hostname and one npub per line, separated by
|
||||
whitespace.
|
||||
- Hostnames are lowercase letters, digits, and hyphens; max
|
||||
63 characters.
|
||||
- Comments start with `#` and continue to end of line; blank
|
||||
lines are ignored.
|
||||
- On duplicate hostnames, the last entry wins.
|
||||
|
||||
The daemon picks up the change on the next DNS query — no
|
||||
restart required (the file's mtime is checked on each query).
|
||||
Verify:
|
||||
|
||||
```sh
|
||||
dig my-laptop.fips AAAA +short
|
||||
```
|
||||
|
||||
Expect one `fd97:...` AAAA record.
|
||||
|
||||
`/etc/fips/hosts` is shipped as a dpkg conffile (and the AUR
|
||||
equivalent), so package upgrades preserve your edits. The
|
||||
file is `0644 root:root` — readable by anyone, writable by
|
||||
root.
|
||||
|
||||
## Add a peer alias
|
||||
|
||||
In `/etc/fips/fips.yaml`, set the `alias:` field on the peer
|
||||
entry:
|
||||
|
||||
```yaml
|
||||
peers:
|
||||
- npub: "npub1abc...xyz"
|
||||
alias: "my-laptop"
|
||||
addresses:
|
||||
- transport: udp
|
||||
addr: "192.0.2.10:2121"
|
||||
connect_policy: auto_connect
|
||||
```
|
||||
|
||||
Restart the daemon for the alias to take effect:
|
||||
|
||||
```sh
|
||||
sudo systemctl restart fips
|
||||
dig my-laptop.fips AAAA +short
|
||||
```
|
||||
|
||||
The alias also shows up in `fipsctl show peers` `display_name`
|
||||
column, so log entries and CLI output reference the peer by
|
||||
shortname instead of truncated npub.
|
||||
|
||||
## Resolution order
|
||||
|
||||
When the DNS responder receives a query for `<name>.fips`:
|
||||
|
||||
1. **Hosts file lookup.** If `<name>` matches an entry in
|
||||
`/etc/fips/hosts`, the daemon returns the AAAA record
|
||||
derived from that entry's npub.
|
||||
2. **Peer alias lookup.** If `<name>` matches the `alias`
|
||||
field on a configured peer, return that peer's AAAA.
|
||||
3. **Direct npub resolution.** If `<name>` is itself a valid
|
||||
bech32 npub (the canonical 63-char `npub1...` form), the
|
||||
daemon returns the AAAA derived from that npub directly.
|
||||
4. **NXDOMAIN.** If none of the above match, the query
|
||||
returns no answer.
|
||||
|
||||
The order means the hosts file overrides peer aliases on
|
||||
conflict. That's deliberate: the file represents
|
||||
operator-shared naming, the peer alias is a node-local label.
|
||||
|
||||
## Cross-references and ACLs
|
||||
|
||||
Aliases interact with the peer ACL — if you maintain
|
||||
`peers.allow` or `peers.deny` lists keyed on hostnames rather
|
||||
than npubs, those names go through the same hosts-file
|
||||
resolution. See
|
||||
[../reference/security.md](../reference/security.md) for the
|
||||
ACL format and the alias-resolution semantics.
|
||||
|
||||
`fipsctl connect` and `fipsctl disconnect` accept a shortname
|
||||
where they expect an npub. Resolution for these commands goes
|
||||
through `/etc/fips/hosts` only — peer-config `alias:` entries
|
||||
are not loaded by `fipsctl`, so a shortname that exists only as
|
||||
a peer alias must still be referenced by full npub on the CLI.
|
||||
See [../reference/cli-fipsctl.md](../reference/cli-fipsctl.md).
|
||||
|
||||
## See also
|
||||
|
||||
- [../reference/configuration.md § Host Mapping](../reference/configuration.md#host-mapping)
|
||||
— minimal reference entry for the host-map mechanism.
|
||||
- [../reference/cli-fipsctl.md](../reference/cli-fipsctl.md)
|
||||
— `fipsctl` arguments that accept shortnames.
|
||||
- [../reference/security.md](../reference/security.md)
|
||||
— peer ACL semantics with aliased entries.
|
||||
- [../design/fips-ipv6-adapter.md](../design/fips-ipv6-adapter.md)
|
||||
— the DNS resolver design and the npub-to-IPv6 derivation.
|
||||
@@ -0,0 +1,232 @@
|
||||
# Provision a Persistent Identity
|
||||
|
||||
A FIPS node's identity is a Nostr keypair. Its public key (npub)
|
||||
determines the node's `fd00::/8` mesh address; peers and configs
|
||||
reference the node by that npub. Out of the box the daemon generates
|
||||
a fresh identity on every start (`node.identity.persistent: false`),
|
||||
which is fine for one-off testing but useless when other nodes need
|
||||
to refer to this one across restarts.
|
||||
|
||||
This guide covers the three ways to give a node a stable identity.
|
||||
For the configuration keys involved, see
|
||||
[../reference/configuration.md](../reference/configuration.md).
|
||||
|
||||
> **First time?** If you have just installed FIPS and want a
|
||||
> hand-held walkthrough of the package-default path (set
|
||||
> `persistent: true`, restart, observe the keys land), the
|
||||
> [persistent-identity tutorial](../tutorials/persistent-identity.md)
|
||||
> is the gentler entry point. This guide assumes an operator
|
||||
> picking among Options A/B/C for a deployment.
|
||||
|
||||
## When to use
|
||||
|
||||
Use a persistent identity for any node that:
|
||||
|
||||
- Other operators reference by npub (in their `peers` lists, `hosts`
|
||||
files, or ACL allow-lists).
|
||||
- Acts as a discoverable bootstrap or rendezvous (Nostr advert,
|
||||
static peer entry, gateway).
|
||||
- Is expected to keep its `fd00::/8` mesh address across restarts.
|
||||
|
||||
Stay with the ephemeral default for throw-away clients, sandbox
|
||||
nodes, and tests where you actively want a fresh identity per run.
|
||||
|
||||
## Option A: Let the package do it
|
||||
|
||||
The Debian/Ubuntu `.deb` and the Arch `fips` AUR package both ship a
|
||||
default `/etc/fips/fips.yaml` with `node.identity.persistent` left as
|
||||
the upstream default (false), so the daemon writes a fresh keypair to
|
||||
`/etc/fips/fips.{key,pub}` on every start until you set
|
||||
`persistent: true`. To pin the current keypair:
|
||||
|
||||
1. Install the package and start the daemon once so it generates
|
||||
`fips.key` / `fips.pub`:
|
||||
|
||||
```sh
|
||||
sudo systemctl start fips
|
||||
sudo systemctl status fips # confirm it came up
|
||||
```
|
||||
|
||||
2. Edit `/etc/fips/fips.yaml` and set:
|
||||
|
||||
```yaml
|
||||
node:
|
||||
identity:
|
||||
persistent: true
|
||||
```
|
||||
|
||||
3. Restart the daemon and verify the identity is reused:
|
||||
|
||||
```sh
|
||||
sudo systemctl restart fips
|
||||
fipsctl show status | grep -E '"npub"|"node_addr"'
|
||||
cat /etc/fips/fips.pub
|
||||
```
|
||||
|
||||
The npub printed by `fipsctl show status` should match
|
||||
`/etc/fips/fips.pub` and remain stable across subsequent restarts.
|
||||
|
||||
The package's `postinst` script does **not** generate the keypair —
|
||||
the daemon does, on first start. This means the keypair is only
|
||||
present after the first successful daemon start. If the daemon never
|
||||
came up cleanly (config error, permission problem), the key files
|
||||
will be missing.
|
||||
|
||||
### File layout and permissions
|
||||
|
||||
| Path | Mode | Owner | Contents |
|
||||
| ---- | ---- | ----- | -------- |
|
||||
| `/etc/fips/fips.key` | `0600` | `root:root` | Bech32 `nsec` (one line). |
|
||||
| `/etc/fips/fips.pub` | `0644` | `root:root` | Bech32 `npub` (one line). |
|
||||
|
||||
Both files live next to the highest-priority `fips.yaml` the daemon
|
||||
loaded. For non-systemd installs that use a different config path,
|
||||
the key files are placed in that config's directory.
|
||||
|
||||
## Option B: Generate manually
|
||||
|
||||
For from-source installs, custom config paths, or any deployment
|
||||
where you want to mint the keypair before the daemon ever runs.
|
||||
|
||||
### With `fipsctl keygen`
|
||||
|
||||
```sh
|
||||
sudo fipsctl keygen --dir /etc/fips
|
||||
```
|
||||
|
||||
This writes `/etc/fips/fips.key` (mode `0600`) and
|
||||
`/etc/fips/fips.pub` (mode `0644`), prints the new npub on stderr,
|
||||
and reminds you to set `persistent: true`. Add `--force` to overwrite
|
||||
an existing `fips.key`. Add `--stdout` to print `nsec` then `npub`
|
||||
to stdout instead of writing files.
|
||||
|
||||
To put the keypair in a non-default directory (e.g., a per-deployment
|
||||
config tree), pass `--dir` and point your `fips.yaml` search at the
|
||||
matching directory.
|
||||
|
||||
### Without the daemon installed
|
||||
|
||||
If you cannot run `fipsctl` (e.g., scripting on a build host), any
|
||||
nostr-tools-equivalent that emits a bech32 `nsec` works. Write the
|
||||
nsec to `fips.key` (mode `0600`) and the corresponding `npub` to
|
||||
`fips.pub` (mode `0644`).
|
||||
|
||||
### Hooking the keypair into the config
|
||||
|
||||
```yaml
|
||||
node:
|
||||
identity:
|
||||
persistent: true
|
||||
```
|
||||
|
||||
`persistent: true` plus a `fips.key` next to the loaded config is the
|
||||
intended steady-state setup.
|
||||
|
||||
## Option C: Provision from an existing nsec
|
||||
|
||||
To migrate an existing Nostr identity into a FIPS node — for example,
|
||||
re-using a personal npub for a node you operate.
|
||||
|
||||
1. Obtain the bech32 `nsec` for the identity.
|
||||
2. Write it to the config-adjacent key file:
|
||||
|
||||
```sh
|
||||
sudo install -m 0600 -o root -g root /dev/null /etc/fips/fips.key
|
||||
sudo bash -c 'printf "%s\n" nsec1... > /etc/fips/fips.key'
|
||||
```
|
||||
|
||||
3. Derive the matching `npub` and write `fips.pub`:
|
||||
|
||||
```sh
|
||||
# compute the npub with any nostr tool, then:
|
||||
sudo bash -c 'printf "%s\n" npub1... > /etc/fips/fips.pub'
|
||||
sudo chmod 0644 /etc/fips/fips.pub
|
||||
```
|
||||
|
||||
4. Set `persistent: true` and restart:
|
||||
|
||||
```yaml
|
||||
node:
|
||||
identity:
|
||||
persistent: true
|
||||
```
|
||||
|
||||
```sh
|
||||
sudo systemctl restart fips
|
||||
fipsctl show status | grep '"npub"'
|
||||
```
|
||||
|
||||
The reported npub should match the one you wrote to `fips.pub`.
|
||||
|
||||
## Verifying
|
||||
|
||||
The daemon prints the resolved identity at startup; the same value is
|
||||
queryable via the control socket:
|
||||
|
||||
```sh
|
||||
fipsctl show status | jq '{npub, node_addr, ipv6_addr}'
|
||||
cat /etc/fips/fips.pub
|
||||
```
|
||||
|
||||
The `npub` field of `show status` and the contents of `fips.pub`
|
||||
should match. The `node_addr` is the SHA-256 prefix used internally
|
||||
by FMP/FSP; the `ipv6_addr` is the routable `fd00::/8` mesh address
|
||||
derived from the node addr. Together they are stable for the lifetime
|
||||
of the keypair.
|
||||
|
||||
The journal also records the source on every start:
|
||||
|
||||
```text
|
||||
INFO Loaded persistent identity from key file path=/etc/fips/fips.key
|
||||
```
|
||||
|
||||
(`Generated persistent identity, saved to key file` on the first
|
||||
start; `Using ephemeral identity (new keypair each start)` when
|
||||
persistence is off.)
|
||||
|
||||
## Rotating
|
||||
|
||||
Key rotation is a destructive operation: every cached
|
||||
`(node_addr → npub)` mapping on every other node points at the old
|
||||
key, every Nostr advert and every static peer entry references the
|
||||
old npub, and every existing FSP session was authenticated under the
|
||||
old keypair. There is no in-protocol "key change" message.
|
||||
|
||||
To rotate:
|
||||
|
||||
1. Stop the daemon.
|
||||
|
||||
```sh
|
||||
sudo systemctl stop fips
|
||||
```
|
||||
|
||||
2. Remove the existing key files.
|
||||
|
||||
```sh
|
||||
sudo rm /etc/fips/fips.key /etc/fips/fips.pub
|
||||
```
|
||||
|
||||
3. Start the daemon. With `persistent: true`, the daemon generates a
|
||||
new keypair and writes new `fips.key` / `fips.pub`.
|
||||
|
||||
```sh
|
||||
sudo systemctl start fips
|
||||
cat /etc/fips/fips.pub # the new npub
|
||||
```
|
||||
|
||||
4. Update every downstream reference: peer configs that name this
|
||||
node by npub, `hosts` files, ACL allow-lists, Nostr adverts
|
||||
pinned by other operators.
|
||||
|
||||
There is no recovery from a lost `fips.key` — the npub is gone with
|
||||
the secret. Treat key rotation as a coordinated event; do not rotate
|
||||
production identities ad hoc.
|
||||
|
||||
## See also
|
||||
|
||||
- [../reference/configuration.md](../reference/configuration.md) —
|
||||
`node.identity.*` keys.
|
||||
- [../reference/cli-fipsctl.md](../reference/cli-fipsctl.md) —
|
||||
`fipsctl keygen`.
|
||||
- [../design/fips-architecture.md](../design/fips-architecture.md) —
|
||||
identity model, npub-to-NodeAddr derivation.
|
||||
@@ -0,0 +1,195 @@
|
||||
# Run the FIPS Daemon as an Unprivileged User
|
||||
|
||||
By default, the FIPS daemon runs as `root` — the shipped Debian
|
||||
systemd unit configures this, and no further setup is required.
|
||||
The trade-off is that the daemon has full root authority,
|
||||
including outside its actual network needs. Acceptable for
|
||||
single-purpose hosts; less desirable for shared hosts.
|
||||
|
||||
This guide covers the alternative: drop privileges and run the
|
||||
daemon under a dedicated unprivileged user account. The TUN
|
||||
device that the FIPS IPv6 adapter creates requires
|
||||
`CAP_NET_ADMIN` on Linux; the recipe below grants that privilege
|
||||
via a file capability on the binary, plus everything else the
|
||||
daemon needs to keep working without root: a service user
|
||||
account, file permissions on the config directory, and a systemd
|
||||
unit override to drop privileges.
|
||||
|
||||
For the design context (why the adapter needs a TUN, how the
|
||||
adapter integrates with the kernel routing table), see
|
||||
[../design/fips-ipv6-adapter.md](../design/fips-ipv6-adapter.md).
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- FIPS package installed (the postinst already creates the `fips`
|
||||
system group used for control-socket access).
|
||||
- `setcap` available (`apt install libcap2-bin` on Debian/Ubuntu;
|
||||
it is a standard utility on most distributions).
|
||||
- Operator access to systemd unit overrides (`systemctl edit`).
|
||||
|
||||
## Step 1: Create a `fips` system user
|
||||
|
||||
The package creates a `fips` system *group* but no matching user.
|
||||
Add a system user that belongs to the `fips` group:
|
||||
|
||||
```sh
|
||||
sudo useradd --system --gid fips --no-create-home --shell /usr/sbin/nologin fips
|
||||
```
|
||||
|
||||
The user has no home directory and no login shell — this account
|
||||
exists only to run the daemon.
|
||||
|
||||
## Step 2: Grant `CAP_NET_ADMIN` to the binary
|
||||
|
||||
Apply the file capability so the daemon can create the TUN device
|
||||
without root authority:
|
||||
|
||||
```sh
|
||||
sudo setcap cap_net_admin+ep /usr/bin/fips
|
||||
```
|
||||
|
||||
Verify:
|
||||
|
||||
```sh
|
||||
getcap /usr/bin/fips
|
||||
# /usr/bin/fips cap_net_admin=ep
|
||||
```
|
||||
|
||||
The binary can now create TUN devices when run by any user.
|
||||
|
||||
**File-capability caveats:**
|
||||
|
||||
- The capability is attached to the binary file. **Re-applying
|
||||
the capability after every package upgrade is required**,
|
||||
because package upgrades replace the binary file and lose the
|
||||
cap. The systemd override in Step 4 includes an `ExecStartPre`
|
||||
line that automates this.
|
||||
- File capabilities are stripped when the binary is copied across
|
||||
most filesystems and when it is downloaded via web tooling. If
|
||||
you build from source and install manually, remember to
|
||||
re-`setcap` after each rebuild.
|
||||
- `LD_LIBRARY_PATH` and similar environment-driven loader
|
||||
controls are stripped at exec time when file capabilities are
|
||||
present; this is normally what you want, but development
|
||||
workflows that rely on custom library paths may be surprised.
|
||||
|
||||
## Step 3: Adjust config-file permissions
|
||||
|
||||
The shipped `/etc/fips/fips.yaml` is mode `0600` and owned by
|
||||
`root:root`. The daemon needs to read it and, if persistent
|
||||
identity is enabled, write `/etc/fips/fips.key` into the same
|
||||
directory.
|
||||
|
||||
```sh
|
||||
sudo chown -R fips:fips /etc/fips
|
||||
sudo chmod 0640 /etc/fips/fips.yaml
|
||||
```
|
||||
|
||||
If `node.identity.persistent: true` is set and `fips.key` does
|
||||
not exist yet, leave `/etc/fips` itself writable by the `fips`
|
||||
user so the daemon can create it on first start. After the key
|
||||
file exists, you can tighten further:
|
||||
|
||||
```sh
|
||||
sudo chmod 0600 /etc/fips/fips.key
|
||||
```
|
||||
|
||||
## Step 4: Drop privileges in the systemd unit
|
||||
|
||||
Create an override:
|
||||
|
||||
```sh
|
||||
sudo systemctl edit fips.service
|
||||
```
|
||||
|
||||
Add:
|
||||
|
||||
```ini
|
||||
[Service]
|
||||
User=fips
|
||||
Group=fips
|
||||
AmbientCapabilities=CAP_NET_ADMIN
|
||||
NoNewPrivileges=no
|
||||
ExecStartPre=/sbin/setcap cap_net_admin+ep /usr/bin/fips
|
||||
```
|
||||
|
||||
`User=` / `Group=` set the service identity.
|
||||
`AmbientCapabilities=` ensures the file capability granted in
|
||||
Step 2 actually carries into the daemon's process tree.
|
||||
`NoNewPrivileges=no` is required for file-capability execution
|
||||
to work — systemd defaults this to `yes` for hardened units,
|
||||
which would block the `setcap` from taking effect.
|
||||
`ExecStartPre=` re-applies the capability before each start,
|
||||
which makes the package-upgrade path self-heal.
|
||||
|
||||
The unit's `RuntimeDirectory=fips` directive already arranges
|
||||
for `/run/fips/` to be created with the right ownership at
|
||||
service start, now as `fips:fips 0750` instead of
|
||||
`root:fips 0750`.
|
||||
|
||||
Reload and restart:
|
||||
|
||||
```sh
|
||||
sudo systemctl daemon-reload
|
||||
sudo systemctl restart fips
|
||||
```
|
||||
|
||||
## Step 5: Verify
|
||||
|
||||
Confirm the daemon is running as `fips`:
|
||||
|
||||
```sh
|
||||
ps -eo user,cmd | grep '[/]usr/bin/fips'
|
||||
# fips /usr/bin/fips --config /etc/fips/fips.yaml
|
||||
```
|
||||
|
||||
Confirm the TUN device came up (the `setcap` worked):
|
||||
|
||||
```sh
|
||||
ip link show fips0
|
||||
# fips0: <POINTOPOINT,UP,...> mtu 1280 ...
|
||||
```
|
||||
|
||||
Confirm the control socket is bound and accessible to the `fips`
|
||||
group:
|
||||
|
||||
```sh
|
||||
ls -la /run/fips/control.sock
|
||||
# srwxrwx--- 1 fips fips ... /run/fips/control.sock
|
||||
```
|
||||
|
||||
Add yourself to the `fips` group so you can use `fipsctl` /
|
||||
`fipstop` without `sudo`:
|
||||
|
||||
```sh
|
||||
sudo usermod -aG fips $USER
|
||||
# log out and back in for the group change to take effect
|
||||
```
|
||||
|
||||
Then:
|
||||
|
||||
```sh
|
||||
fipsctl show node
|
||||
```
|
||||
|
||||
## Caveats
|
||||
|
||||
- **`fips-firewall.service` still runs as root.** Loading nftables
|
||||
rules into the kernel requires root regardless. The firewall
|
||||
unit is intentionally separate from the daemon unit.
|
||||
- **Bluetooth peers (`transports.ble.*`)** require additional
|
||||
privileges the `CAP_NET_ADMIN` setcap doesn't cover. If you use
|
||||
the BLE transport, you'll likely need to keep running as root
|
||||
or layer additional capability/D-Bus configuration; that path
|
||||
is not covered here.
|
||||
|
||||
## See also
|
||||
|
||||
- [persistent-identity.md](persistent-identity.md) — how the
|
||||
daemon manages `/etc/fips/fips.key`
|
||||
- [../design/fips-ipv6-adapter.md](../design/fips-ipv6-adapter.md)
|
||||
— IPv6 adapter design, TUN interface architecture
|
||||
- [../reference/security.md](../reference/security.md) —
|
||||
consolidated security surface
|
||||
- [../reference/configuration.md](../reference/configuration.md)
|
||||
— `tun.*` configuration block
|
||||
@@ -0,0 +1,299 @@
|
||||
# Set Up a Bluetooth (BLE) Peer Link
|
||||
|
||||
FIPS supports Bluetooth Low Energy as a transport for short-range
|
||||
mesh extension — same room, same building, no IP infrastructure
|
||||
between the two endpoints. The BLE transport runs as L2CAP
|
||||
Connection-Oriented Channels on a configurable PSM and reports
|
||||
per-link MTU back to the mesh layer for path-MTU computation.
|
||||
|
||||
For the design rationale and per-link MTU model, see
|
||||
[../design/fips-transport-layer.md](../design/fips-transport-layer.md).
|
||||
For all `transports.ble.*` configuration keys, see
|
||||
[../reference/configuration.md](../reference/configuration.md).
|
||||
|
||||
> **Experimental.** The BLE transport works but is still maturing.
|
||||
> Expect rougher edges than UDP or TCP — particularly around link
|
||||
> stability under interference and MTU negotiation on older
|
||||
> controllers. Treat it as you would any experimental transport in a
|
||||
> production deployment.
|
||||
|
||||
## When to use
|
||||
|
||||
BLE is the right transport when:
|
||||
|
||||
- Two nodes are within roughly 10 metres line-of-sight (more with
|
||||
external antennas, less through walls).
|
||||
- You want a self-contained mesh segment with no shared WiFi or
|
||||
Ethernet between the participants.
|
||||
- You can work within practical L2CAP CoC throughput (1-2 Mbps in
|
||||
good conditions, often substantially less under interference or
|
||||
at range) and the higher latency variance compared to WiFi.
|
||||
|
||||
It is **not** the right transport for backbone links between rooms
|
||||
where WiFi or Ethernet exists, for high-throughput data, or for any
|
||||
deployment where range matters more than infrastructure-freedom.
|
||||
|
||||
## Platform support
|
||||
|
||||
The BLE transport is **Linux-only** in the current implementation.
|
||||
The runtime depends on BlueZ via the `bluer` crate, which in turn
|
||||
needs `glibc` (musl builds skip BLE; the build script gates the
|
||||
crate accordingly).
|
||||
|
||||
| Platform | BLE transport |
|
||||
| -------- | -------------- |
|
||||
| Linux (glibc) | Supported. |
|
||||
| Linux (musl, OpenWrt) | Disabled at build time. |
|
||||
| macOS | Not supported. |
|
||||
| Windows | Not supported. |
|
||||
|
||||
The Debian package `Recommends: bluez`; install it explicitly if you
|
||||
opted out:
|
||||
|
||||
```sh
|
||||
sudo apt install bluez
|
||||
```
|
||||
|
||||
## Prerequisites
|
||||
|
||||
Both endpoints need:
|
||||
|
||||
1. A BLE-capable HCI adapter visible to BlueZ. Confirm with:
|
||||
|
||||
```sh
|
||||
sudo bluetoothctl show
|
||||
```
|
||||
|
||||
Note the controller name (typically `hci0`).
|
||||
|
||||
2. The `bluetoothd` service running and the adapter powered on:
|
||||
|
||||
```sh
|
||||
sudo systemctl enable --now bluetooth
|
||||
sudo bluetoothctl power on
|
||||
```
|
||||
|
||||
3. Sufficient privileges for the FIPS daemon. There are two
|
||||
independent privilege concerns; the BLE-only deployment case
|
||||
(mesh router with `tun.enabled: false`) needs only the second.
|
||||
|
||||
- **TUN adapter (always required when `tun.enabled: true`).**
|
||||
The daemon needs `CAP_NET_ADMIN` to create and configure the
|
||||
TUN device. The shipped systemd unit handles this by running
|
||||
as root; if you prefer to drop privileges, see
|
||||
[run-as-unprivileged-user.md](run-as-unprivileged-user.md).
|
||||
|
||||
- **BLE access (required for this how-to).** BlueZ exposes
|
||||
L2CAP and D-Bus paths under either group membership or
|
||||
`CAP_NET_RAW`. Pick one:
|
||||
|
||||
- Run the daemon as root. The shipped systemd unit takes
|
||||
this route.
|
||||
- Run as an unprivileged user that is a member of the
|
||||
`bluetooth` group. No additional capability is needed for
|
||||
the BLE side.
|
||||
- Run as an unprivileged user with no group membership, and
|
||||
grant the binary `CAP_NET_RAW`:
|
||||
|
||||
```sh
|
||||
sudo setcap cap_net_raw+ep $(which fips)
|
||||
```
|
||||
|
||||
This bypasses BlueZ's polkit/group check by holding
|
||||
`CAP_NET_RAW` directly. If you also need `CAP_NET_ADMIN`
|
||||
for TUN, combine them:
|
||||
|
||||
```sh
|
||||
sudo setcap cap_net_admin,cap_net_raw+ep $(which fips)
|
||||
```
|
||||
|
||||
4. The same L2CAP PSM on both endpoints. The default is `0x0085`
|
||||
(133); override only if you need to coexist with another L2CAP
|
||||
service on that PSM.
|
||||
|
||||
## Configuration
|
||||
|
||||
Add a `ble` block under `transports` in `fips.yaml`. A minimum BLE-
|
||||
active node looks like this:
|
||||
|
||||
```yaml
|
||||
transports:
|
||||
ble:
|
||||
adapter: "hci0"
|
||||
advertise: true
|
||||
scan: true
|
||||
auto_connect: true
|
||||
accept_connections: true
|
||||
```
|
||||
|
||||
Note: `auto_connect: true` is intentionally non-default (the default
|
||||
is `false`). For a symmetric ground-up discovery flow where either
|
||||
side may dial, both ends must opt in explicitly.
|
||||
|
||||
| Key | Purpose |
|
||||
| --- | ------- |
|
||||
| `adapter` | HCI controller name. Default: `hci0`. |
|
||||
| `psm` | L2CAP PSM. Default: `0x0085` (must match on both ends). |
|
||||
| `mtu` | Default L2CAP CoC MTU. Default: `2048`. The kernel may negotiate lower per link. |
|
||||
| `max_connections` | Concurrent BLE connections. Default: `7` (Bluetooth controllers typically support up to ~7 simultaneous L2CAP CoCs). |
|
||||
| `advertise` | Broadcast our BLE adverts so other FIPS nodes discover us. Default: `true`. |
|
||||
| `scan` | Listen for other FIPS nodes' BLE adverts. Default: `true`. |
|
||||
| `auto_connect` | Initiate a BLE connection to discovered FIPS adverts. Default: `false`. |
|
||||
| `accept_connections` | Accept inbound L2CAP connections. Default: `true`. |
|
||||
| `connect_timeout_ms` | Outbound L2CAP connect timeout. Default: `10000`. |
|
||||
| `probe_cooldown_secs` | After probing a BD_ADDR (success or failure), wait this long before probing it again. Default: `30`. |
|
||||
|
||||
Two pairing patterns are common:
|
||||
|
||||
**Symmetric auto-discovery.** Both nodes advertise, scan, and
|
||||
auto-connect. Whichever side completes the L2CAP connection first
|
||||
wins; the other side aborts its in-flight attempt. This is the
|
||||
"toss two devices in the same room" setup.
|
||||
|
||||
```yaml
|
||||
# Both nodes
|
||||
transports:
|
||||
ble:
|
||||
adapter: "hci0"
|
||||
advertise: true
|
||||
scan: true
|
||||
auto_connect: true
|
||||
accept_connections: true
|
||||
```
|
||||
|
||||
**Asymmetric peripheral / central.** One node only listens
|
||||
(peripheral), the other actively dials (central). Useful when one
|
||||
endpoint is a dedicated bootstrap and the other is mobile.
|
||||
|
||||
```yaml
|
||||
# Listener
|
||||
transports:
|
||||
ble:
|
||||
adapter: "hci0"
|
||||
advertise: true
|
||||
scan: false
|
||||
auto_connect: false
|
||||
accept_connections: true
|
||||
```
|
||||
|
||||
```yaml
|
||||
# Dialer
|
||||
transports:
|
||||
ble:
|
||||
adapter: "hci0"
|
||||
advertise: false
|
||||
scan: true
|
||||
auto_connect: true
|
||||
accept_connections: false
|
||||
```
|
||||
|
||||
After editing, restart the daemon on each side:
|
||||
|
||||
```sh
|
||||
sudo systemctl restart fips
|
||||
```
|
||||
|
||||
## Verify
|
||||
|
||||
On each endpoint, confirm the transport came up:
|
||||
|
||||
```sh
|
||||
fipsctl show transports
|
||||
```
|
||||
|
||||
Look for an entry of type `ble` in the `state: Running` (or
|
||||
equivalent) state. The `mtu` field reports the configured default;
|
||||
per-link MTU is reported separately.
|
||||
|
||||
Confirm the link is established:
|
||||
|
||||
```sh
|
||||
fipsctl show peers
|
||||
```
|
||||
|
||||
The peer entry for the BLE-attached neighbour should report
|
||||
`transport_type: "ble"` and a non-zero `last_seen_ms`.
|
||||
|
||||
BLE peering is auto-discovery only: there is no `fipsctl connect`
|
||||
path for BLE (the command accepts `udp`, `tcp`, `tor`, and
|
||||
`ethernet` only). Links come up via advert/scan; if you don't see
|
||||
the peer here, the configuration above is the only knob.
|
||||
|
||||
To watch the link in real time, use `fipstop`'s **Peers** and
|
||||
**Transports** tabs:
|
||||
|
||||
```sh
|
||||
fipstop
|
||||
```
|
||||
|
||||
The Performance tab reports the per-link MMP metrics — SRTT, loss
|
||||
rate, ETX — which on BLE typically run an order of magnitude worse
|
||||
than over UDP, with much higher jitter.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Transport never comes up
|
||||
|
||||
Check the BlueZ side first:
|
||||
|
||||
```sh
|
||||
systemctl status bluetooth
|
||||
sudo bluetoothctl show
|
||||
```
|
||||
|
||||
If `bluetoothctl show` reports `Powered: no`, fix that before
|
||||
debugging FIPS. The FIPS daemon will log a warning if it cannot
|
||||
acquire the adapter.
|
||||
|
||||
If the FIPS log contains `bluer` D-Bus errors, the daemon usually
|
||||
lacks permission. Run as root or grant `CAP_NET_ADMIN` and add the
|
||||
fips user to the `bluetooth` group.
|
||||
|
||||
### Peers see each other but never connect
|
||||
|
||||
Verify `accept_connections` is true on at least one side and
|
||||
`auto_connect` is true on at least one side. Two listen-only nodes
|
||||
will discover each other but never establish an L2CAP connection.
|
||||
|
||||
Check `psm` matches on both ends. A mismatch presents as adverts
|
||||
visible (in `fipstop` discovery counters) but every connect attempt
|
||||
fails.
|
||||
|
||||
### Link comes up but throughput is poor
|
||||
|
||||
Practical L2CAP CoC throughput in good conditions reaches
|
||||
1-2 Mbps, but interference, range, and controller capability all
|
||||
push it lower. If throughput is well below that range, check the
|
||||
negotiated ATT_MTU — a small ATT_MTU (default 23 bytes when
|
||||
extended ATT MTU is not negotiated) caps per-PDU payload
|
||||
regardless of radio conditions. The per-link MTU reported in
|
||||
`fipsctl show transports` reveals what was negotiated.
|
||||
|
||||
If MTU is unexpectedly low, both endpoints must support and have
|
||||
negotiated the BlueZ L2CAP `cocmode=2` extension. Older Bluetooth
|
||||
controllers cap MTU regardless.
|
||||
|
||||
### Unstable links / repeated reconnects
|
||||
|
||||
Bluetooth in busy 2.4 GHz environments suffers from WiFi
|
||||
interference. Switch the adapter to a less crowded channel (kernel
|
||||
side, not configurable from FIPS) or add an external antenna. The
|
||||
`probe_cooldown_secs` tunable backs off retry attempts; raise it if
|
||||
the daemon log shows many short-lived probes.
|
||||
|
||||
### Permission errors on socket open
|
||||
|
||||
Most modern systemd installs do not allow non-root processes to
|
||||
open raw L2CAP sockets without an explicit policy. Run the daemon
|
||||
as root (the shipped systemd unit does this) or add a `polkit`
|
||||
rule for the `bluetooth` group.
|
||||
|
||||
## See also
|
||||
|
||||
- [../design/fips-transport-layer.md](../design/fips-transport-layer.md)
|
||||
— per-transport MTU reporting and the BLE row of the supported-
|
||||
transports table.
|
||||
- [../reference/configuration.md](../reference/configuration.md) —
|
||||
full `transports.ble.*` reference.
|
||||
- [run-as-unprivileged-user.md](run-as-unprivileged-user.md) —
|
||||
adjacent privilege handling for the daemon process.
|
||||
@@ -0,0 +1,453 @@
|
||||
# Troubleshoot `fips-gateway`
|
||||
|
||||
Diagnostic recipes for `fips-gateway`, grouped by which half of the
|
||||
gateway is failing. For gateway design and deployment, see
|
||||
[../design/fips-gateway.md](../design/fips-gateway.md) and
|
||||
[deploy-gateway.md](deploy-gateway.md). For OpenWrt-specific
|
||||
deployment problems, see the
|
||||
[OpenWrt deployment tutorial](../tutorials/deploy-fips-gateway.md);
|
||||
most of the recipes below apply on OpenWrt as well, but paths and
|
||||
service names differ.
|
||||
|
||||
## Inspect gateway state via the control socket
|
||||
|
||||
Before digging into nftables or conntrack, ask the gateway directly
|
||||
whether it has the mapping or session you expect. `fips-gateway`
|
||||
exposes a separate control socket (`/run/fips/gateway.sock`) with its
|
||||
own command set; there is no `fipsctl gateway` subcommand — talk to
|
||||
the socket directly with `nc -U`. Each request is a single line of
|
||||
JSON terminated with a newline; the connection is closed after one
|
||||
response.
|
||||
|
||||
Pool summary, listen address, NAT counters, uptime, and the loaded
|
||||
config snapshot:
|
||||
|
||||
```sh
|
||||
echo '{"command":"show_gateway"}' | sudo nc -U /run/fips/gateway.sock
|
||||
```
|
||||
|
||||
Per-mapping virtual-IP state (allocated, active, draining):
|
||||
|
||||
```sh
|
||||
echo '{"command":"show_mappings"}' | sudo nc -U /run/fips/gateway.sock
|
||||
```
|
||||
|
||||
If either command returns `gateway not yet initialized`, the gateway
|
||||
is still in early startup; wait a moment and retry. If a mapping you
|
||||
expect is not in the list, the DNS path didn't allocate it — fall
|
||||
through to the outbound DNS recipes below. If the mapping exists in
|
||||
`state: Active` but mesh traffic still fails, the problem is
|
||||
downstream of the allocation (firewall, route, masquerade); see the
|
||||
recipes that follow.
|
||||
|
||||
For the full command catalog and JSON shapes, see
|
||||
[../reference/control-socket.md#gateway-command-catalog](../reference/control-socket.md#gateway-command-catalog).
|
||||
|
||||
## Common (either-half) issues
|
||||
|
||||
These break both halves at once because they affect the gateway
|
||||
process itself or the shared NAT machinery.
|
||||
|
||||
### "No gateway section in configuration"
|
||||
|
||||
`fips-gateway` is normally launched by the systemd unit shipped with
|
||||
the package:
|
||||
|
||||
```sh
|
||||
sudo systemctl restart fips-gateway
|
||||
sudo journalctl -u fips-gateway -e
|
||||
```
|
||||
|
||||
The unit reads the standard FIPS config search paths (typically
|
||||
`/etc/fips/fips.yaml`). If the unit logs "no gateway section in
|
||||
configuration" or "Gateway section exists but is not enabled", confirm
|
||||
the section is present and `enabled: true`:
|
||||
|
||||
```sh
|
||||
grep -A1 '^gateway:' /etc/fips/fips.yaml
|
||||
```
|
||||
|
||||
For one-off debugging outside systemd, run the binary directly and
|
||||
point it at a specific config file:
|
||||
|
||||
```sh
|
||||
sudo fips-gateway --config /etc/fips/fips.yaml --log-level debug
|
||||
```
|
||||
|
||||
This is useful to capture stderr in a terminal, but the systemd unit
|
||||
is the supported entry point in production. See
|
||||
[../reference/cli-fips-gateway.md](../reference/cli-fips-gateway.md)
|
||||
for the full flag list.
|
||||
|
||||
### Port conflict on the DNS listen port
|
||||
|
||||
Symptom: gateway fails to start with "address already in use" on
|
||||
the configured `gateway.dns.listen` address.
|
||||
|
||||
The default `[::1]:5353` is loopback-only on an unprivileged port and
|
||||
should not collide with any standard resolver. If you have overridden
|
||||
`dns.listen` to bind port 53 (or a LAN-side address) and another DNS
|
||||
server (systemd-resolved, dnsmasq, BIND) is already bound there,
|
||||
identify it:
|
||||
|
||||
```sh
|
||||
sudo ss -tulnp | grep ':53'
|
||||
```
|
||||
|
||||
Two options:
|
||||
|
||||
- **Stay on the loopback default.** Drop the override and let the
|
||||
gateway use `[::1]:5353`. Configure the existing resolver to
|
||||
forward `.fips` queries to it (the canonical OpenWrt deployment
|
||||
works this way out of the box).
|
||||
|
||||
- **Relocate the conflicting resolver.** Move it to a different port
|
||||
(or disable it if not needed) and let the gateway bind 53.
|
||||
Practical for systemd-resolved (set `DNSStubListener=no` in
|
||||
`/etc/systemd/resolved.conf`); rarely worth it for production
|
||||
resolvers.
|
||||
|
||||
### IPv6 forwarding disabled
|
||||
|
||||
Symptom: gateway exits at startup with
|
||||
"IPv6 forwarding is disabled. Enable with: sysctl -w
|
||||
net.ipv6.conf.all.forwarding=1".
|
||||
|
||||
The gateway is completely non-functional without forwarding — packets
|
||||
cannot traverse the NAT pipeline. Enable it:
|
||||
|
||||
```sh
|
||||
sudo sysctl -w net.ipv6.conf.all.forwarding=1
|
||||
```
|
||||
|
||||
Persist via the drop-in shown in
|
||||
[deploy-gateway.md](deploy-gateway.md#kernel-sysctls). The same
|
||||
section lists `proxy_ndp`, which is also required for the outbound
|
||||
half.
|
||||
|
||||
### nftables table missing or not loaded
|
||||
|
||||
Symptom: `show_gateway` reports an active gateway but
|
||||
`nft list table inet fips_gateway` errors with "No such file or
|
||||
directory".
|
||||
|
||||
The table is created by the gateway at startup and rebuilt atomically
|
||||
on every mapping change and on every `set_port_forwards` call. If the
|
||||
table is missing while the gateway claims to be running, something
|
||||
else (a host firewall script, a `nft flush ruleset` from another
|
||||
service) deleted it after creation. Restart the gateway to recreate
|
||||
it:
|
||||
|
||||
```sh
|
||||
sudo systemctl restart fips-gateway
|
||||
```
|
||||
|
||||
If a peer service is repeatedly clobbering the table, switch that
|
||||
service to use `add table` / `flush table <name>` for its own table
|
||||
rather than `flush ruleset`, which destroys every table on the host.
|
||||
|
||||
### Control socket permission errors
|
||||
|
||||
Symptom: `nc -U /run/fips/gateway.sock` fails with "Permission
|
||||
denied" or "No such file or directory".
|
||||
|
||||
The socket is owned by root with mode `0660` (group `fips`). Either
|
||||
run `nc` as root (`sudo nc -U ...`) or add your user to the `fips`
|
||||
group and re-login. If the file does not exist at all, the gateway
|
||||
either failed to start (check `journalctl -u fips-gateway`) or
|
||||
failed to bind the socket and continued without it (the warning
|
||||
`Failed to bind gateway control socket — continuing without it` is
|
||||
in the journal in that case).
|
||||
|
||||
## Outbound-half diagnostics
|
||||
|
||||
Symptoms in this section all involve a LAN client trying to reach a
|
||||
mesh destination through the gateway and failing.
|
||||
|
||||
### DNS queries fail
|
||||
|
||||
Symptom: LAN clients get `SERVFAIL` or no response when querying
|
||||
`.fips` names; or the gateway log shows DNS upstream timeouts.
|
||||
|
||||
**Step 1.** Verify the daemon resolver is running and reachable from
|
||||
the gateway host:
|
||||
|
||||
```sh
|
||||
dig @::1 -p 5354 hostname.fips AAAA
|
||||
```
|
||||
|
||||
If this returns no answer or fails, the FIPS daemon's DNS resolver is
|
||||
not running or not enabled. Check that the daemon config has
|
||||
`dns.enabled: true` (the default) and the daemon is healthy:
|
||||
`fipsctl show status`.
|
||||
|
||||
**Step 2.** Verify the gateway is listening on its DNS port:
|
||||
|
||||
```sh
|
||||
sudo ss -tulnp | grep -E ':(53|5353)\b'
|
||||
```
|
||||
|
||||
If nothing is listening on the configured `dns.listen` address, the
|
||||
gateway either failed to start or is bound to a different address.
|
||||
Check the gateway log: `sudo journalctl -u fips-gateway -e`.
|
||||
|
||||
**Step 3.** Verify the LAN client can reach the gateway's DNS port:
|
||||
|
||||
```sh
|
||||
# from the LAN client
|
||||
dig @<gateway-lan-addr> hostname.fips AAAA
|
||||
```
|
||||
|
||||
If this hangs, the LAN-side firewall is blocking DNS, or the LAN
|
||||
route to the gateway is missing.
|
||||
|
||||
### Ping works but TCP does not
|
||||
|
||||
Symptom: `ping6 <virtual-ip>` succeeds from a LAN client, but TCP
|
||||
connections (SSH, HTTP) hang or time out.
|
||||
|
||||
This usually means the `fips0`-side masquerade rule is missing or
|
||||
misconfigured. Inspect the gateway's nftables table:
|
||||
|
||||
```sh
|
||||
sudo nft list table inet fips_gateway
|
||||
```
|
||||
|
||||
In the `postrouting` chain, look for a rule matching
|
||||
`oifname "fips0"` with a `masquerade` verdict. Without masquerade,
|
||||
the destination mesh node sees a source address (from the virtual
|
||||
pool) it cannot route replies to, and return packets are
|
||||
black-holed.
|
||||
|
||||
If the rule is missing, restart the gateway — the table is rebuilt
|
||||
atomically on every mapping change and on startup.
|
||||
|
||||
### Connection timeout to a virtual IP
|
||||
|
||||
Symptom: any traffic to a virtual pool address times out, including
|
||||
ping.
|
||||
|
||||
**Step 1.** Verify IPv6 forwarding is still enabled:
|
||||
|
||||
```sh
|
||||
sysctl net.ipv6.conf.all.forwarding
|
||||
# Expect: net.ipv6.conf.all.forwarding = 1
|
||||
```
|
||||
|
||||
**Step 2.** Verify the pool route exists:
|
||||
|
||||
```sh
|
||||
ip -6 route show table local | grep <pool-cidr>
|
||||
```
|
||||
|
||||
If the route is missing, the kernel does not recognize pool
|
||||
addresses as locally-owned and drops the packets before NAT can
|
||||
process them. The gateway adds this route at startup; if it's
|
||||
missing, check the gateway log for startup errors.
|
||||
|
||||
**Step 3.** Verify the destination mesh address actually exists in
|
||||
the FIPS daemon's identity cache:
|
||||
|
||||
```sh
|
||||
fipsctl show identity-cache | grep <fd00-mesh-addr>
|
||||
```
|
||||
|
||||
If the entry is missing, the DNS-side mapping never primed the
|
||||
identity cache, which means the daemon resolver did not actually
|
||||
resolve the `.fips` name. Re-test the DNS path:
|
||||
|
||||
```sh
|
||||
dig @::1 -p 5354 hostname.fips AAAA
|
||||
```
|
||||
|
||||
### Virtual IP unreachable from a LAN client
|
||||
|
||||
Symptom: client cannot reach the virtual IP at all (no ping, no
|
||||
ARP/ND response).
|
||||
|
||||
**Step 1.** Verify the client has a route to the pool via the
|
||||
gateway:
|
||||
|
||||
```sh
|
||||
# from the LAN client
|
||||
ip -6 route get <virtual-ip>
|
||||
```
|
||||
|
||||
The output should show the gateway as the next-hop. If it shows
|
||||
something else (or "unreachable"), fix the LAN-side route — see
|
||||
[deploy-gateway.md](deploy-gateway.md#distribute-the-route-to-lan-clients).
|
||||
|
||||
**Step 2.** On the gateway, verify proxy NDP entries exist for
|
||||
allocated virtual IPs:
|
||||
|
||||
```sh
|
||||
ip -6 neigh show proxy
|
||||
```
|
||||
|
||||
If proxy NDP entries are missing, the gateway cannot answer Neighbor
|
||||
Solicitation requests for virtual IPs on the LAN, so clients cannot
|
||||
resolve the link-layer address and packets never leave the client's
|
||||
NIC.
|
||||
|
||||
The gateway adds these entries when a mapping is created (i.e., when
|
||||
a `.fips` DNS query allocates a virtual IP). If they're absent,
|
||||
trigger a DNS query first:
|
||||
|
||||
```sh
|
||||
dig @<gateway-lan-addr> hostname.fips AAAA
|
||||
```
|
||||
|
||||
Then re-check `ip -6 neigh show proxy`.
|
||||
|
||||
**Step 3.** Verify `proxy_ndp` is enabled in the kernel:
|
||||
|
||||
```sh
|
||||
sysctl net.ipv6.conf.all.proxy_ndp
|
||||
# Expect: net.ipv6.conf.all.proxy_ndp = 1
|
||||
```
|
||||
|
||||
If 0, enable it (see
|
||||
[deploy-gateway.md](deploy-gateway.md#kernel-sysctls)).
|
||||
|
||||
## Inbound-half diagnostics
|
||||
|
||||
Symptoms in this section all involve a mesh peer trying to reach a
|
||||
LAN-side service through the gateway and failing.
|
||||
|
||||
### Mesh peer can't reach `<gateway-npub>.fips:<listen_port>`
|
||||
|
||||
Walk the path from the mesh-side ingress to the LAN target:
|
||||
|
||||
**Step 1.** Verify the port-forward rule is loaded. On the gateway:
|
||||
|
||||
```sh
|
||||
sudo nft list table inet fips_gateway
|
||||
```
|
||||
|
||||
Look in the `prerouting` chain for a rule of the form
|
||||
|
||||
```text
|
||||
iif "fips0" meta nfproto ipv6 meta l4proto <tcp|udp> \
|
||||
<th> dport <listen_port> dnat ip6 to [<target_addr>]:<target_port>
|
||||
```
|
||||
|
||||
and, in the `postrouting` chain, a rule of the form
|
||||
|
||||
```text
|
||||
iif "fips0" oif "<lan_interface>" meta nfproto ipv6 masquerade
|
||||
```
|
||||
|
||||
The port-forward DNAT and the LAN-side masquerade come from
|
||||
`gateway.port_forwards[]` and the active `lan_interface` setting.
|
||||
The masquerade is emitted only when at least one port-forward exists.
|
||||
If either rule is missing, restart the gateway — the table is rebuilt
|
||||
atomically on config load.
|
||||
|
||||
**Step 2.** Verify the mesh firewall is not blocking the listen
|
||||
port. If `fips-firewall.service` is enabled, the default baseline
|
||||
drops everything inbound on `fips0` except established/related and
|
||||
ICMPv6. Add an explicit allow rule under `/etc/fips/fips.d/`:
|
||||
|
||||
```nft
|
||||
# /etc/fips/fips.d/gateway-inbound.nft
|
||||
tcp dport <listen_port> accept
|
||||
```
|
||||
|
||||
(See [enable-mesh-firewall.md](enable-mesh-firewall.md) for the
|
||||
full drop-in pattern, including source-address restrictions.)
|
||||
Without an allow rule, mesh peers see TCP RSTs (the firewall drops
|
||||
on the way in) or silent UDP loss.
|
||||
|
||||
**Step 3.** Verify the LAN target is reachable from the gateway
|
||||
itself:
|
||||
|
||||
```sh
|
||||
ping6 <target_addr>
|
||||
curl -v http://[<target_addr>]:<target_port>/ # for TCP HTTP
|
||||
```
|
||||
|
||||
If the target is unreachable from the gateway, the DNAT rule will
|
||||
fire but the inner connection attempt will fail. Fix LAN-side
|
||||
routing or the target service before going further.
|
||||
|
||||
**Step 4.** Verify conntrack is tracking the inbound flow. Try the
|
||||
connection from a mesh peer once:
|
||||
|
||||
```sh
|
||||
curl -v http://<gateway-npub>.fips:<listen_port>/
|
||||
```
|
||||
|
||||
Then on the gateway:
|
||||
|
||||
```sh
|
||||
sudo conntrack -L | grep -E '<listen_port>|<target_port>'
|
||||
```
|
||||
|
||||
You should see a flow tuple in both directions (orig and reply) with
|
||||
the mesh peer's source on `fips0` and the gateway's LAN address as
|
||||
the masqueraded source on the LAN side. No conntrack entry suggests
|
||||
the prerouting DNAT didn't match — recheck step 1.
|
||||
|
||||
**Step 5.** Check the gateway log for nftables or rule install
|
||||
errors:
|
||||
|
||||
```sh
|
||||
sudo journalctl -u fips-gateway -e | grep -E 'port_forward|nftables'
|
||||
```
|
||||
|
||||
A "Failed to install port-forward rules" log line at startup means
|
||||
the rule batch was rejected by netlink — usually a transient
|
||||
condition during a config edit, but persistent failures warrant
|
||||
inspecting the rule with `nft -d`.
|
||||
|
||||
### Config rejected: IPv4 target
|
||||
|
||||
Symptom: `fips-gateway` exits at startup with a deserialization
|
||||
error referencing `port_forwards[N].target` and an invalid IPv6
|
||||
literal.
|
||||
|
||||
The `target` field is typed as `SocketAddrV6` and rejects IPv4
|
||||
literals at parse time:
|
||||
|
||||
```yaml
|
||||
# fails at config load
|
||||
- listen_port: 8080
|
||||
proto: tcp
|
||||
target: "192.168.1.10:80"
|
||||
```
|
||||
|
||||
Either re-address the LAN service to be reachable on IPv6, or front
|
||||
it with a small IPv6-aware reverse proxy on the gateway and point
|
||||
the `target` at that proxy.
|
||||
|
||||
### Config rejected: zero or duplicate listen_port
|
||||
|
||||
Symptom: `fips-gateway` exits at startup with
|
||||
"Invalid gateway.port_forwards: …".
|
||||
|
||||
`validate_port_forwards()` enforces:
|
||||
|
||||
- `listen_port` must be non-zero.
|
||||
- The pair `(listen_port, proto)` must be unique across the list
|
||||
(the same port on TCP and UDP simultaneously is allowed; the same
|
||||
port twice on the same proto is not).
|
||||
|
||||
Fix the offending entry and reload.
|
||||
|
||||
## See also
|
||||
|
||||
- [../tutorials/deploy-fips-gateway.md](../tutorials/deploy-fips-gateway.md) —
|
||||
canonical OpenWrt deployment.
|
||||
- [../design/fips-gateway.md](../design/fips-gateway.md) — gateway
|
||||
design, NAT pipeline, virtual IP pool lifecycle, security
|
||||
considerations.
|
||||
- [deploy-gateway.md](deploy-gateway.md) — manual Linux-host setup.
|
||||
- [Gateway section](../reference/configuration.md#gateway-gateway) of
|
||||
the configuration reference — full `gateway.*` block.
|
||||
- [../reference/cli-fips-gateway.md](../reference/cli-fips-gateway.md) —
|
||||
`fips-gateway` binary CLI flags.
|
||||
- [Gateway command catalog](../reference/control-socket.md#gateway-command-catalog)
|
||||
in the control-socket reference — JSON schema for `show_gateway`
|
||||
and `show_mappings`.
|
||||
- [enable-mesh-firewall.md](enable-mesh-firewall.md) — mesh-firewall
|
||||
baseline and drop-ins.
|
||||
@@ -0,0 +1,125 @@
|
||||
# Tune Host UDP Socket Buffers for FIPS
|
||||
|
||||
The FIPS UDP transport requests larger send and receive socket
|
||||
buffers (default 2 MB each, doubled by the kernel to 4 MB actual)
|
||||
than the Linux defaults provide. The kernel silently clamps the
|
||||
request to `net.core.rmem_max` and `net.core.wmem_max` if those
|
||||
sysctls are smaller than the requested size — which causes silent
|
||||
packet drops under high throughput. For the design context (why FIPS
|
||||
requests larger buffers and how `SO_RXQ_OVFL` feeds ECN congestion
|
||||
detection), see
|
||||
[../design/fips-transport-layer.md](../design/fips-transport-layer.md#socket-buffer-sizing).
|
||||
|
||||
This guide covers the host-side sysctl setup needed before deploying
|
||||
a high-throughput FIPS node.
|
||||
|
||||
## Why this matters
|
||||
|
||||
The default Linux UDP receive buffer (`net.core.rmem_default`,
|
||||
typically 212 KB) fills in roughly 2.5 ms at ~85 MB/s. Any stall in
|
||||
the FIPS receive loop (decryption, routing, forwarding) causes the
|
||||
kernel to drop incoming datagrams without notification — they don't
|
||||
appear in `recv` errors, they don't trigger any application-visible
|
||||
event. The drops show up only in `SO_RXQ_OVFL` on subsequent
|
||||
packets, where FIPS surfaces them as congestion-detection events.
|
||||
|
||||
Setting `rmem_max` and `wmem_max` to at least the requested buffer
|
||||
size prevents the kernel clamp and the silent drop loss it causes.
|
||||
|
||||
## Step 1: Check current limits
|
||||
|
||||
```sh
|
||||
sysctl net.core.rmem_max net.core.wmem_max
|
||||
```
|
||||
|
||||
Typical defaults on stock Linux distributions are 212992 bytes
|
||||
(212 KB). FIPS requests 2 MB by default, which the kernel doubles
|
||||
internally to 4 MB; for the request to succeed without clamping, both
|
||||
sysctls must be at least 4194304 (4 MB).
|
||||
|
||||
## Step 2: Set the limits temporarily
|
||||
|
||||
```sh
|
||||
sudo sysctl -w net.core.rmem_max=4194304
|
||||
sudo sysctl -w net.core.wmem_max=4194304
|
||||
```
|
||||
|
||||
Verify:
|
||||
|
||||
```sh
|
||||
sysctl net.core.rmem_max net.core.wmem_max
|
||||
```
|
||||
|
||||
These changes take effect immediately for new socket binds but do
|
||||
not survive a reboot.
|
||||
|
||||
## Step 3: Make the limits persistent
|
||||
|
||||
Drop a file under `/etc/sysctl.d/`:
|
||||
|
||||
```sh
|
||||
sudo tee /etc/sysctl.d/60-fips.conf <<'EOF'
|
||||
# FIPS UDP transport requests 2 MB socket buffers, kernel doubles to 4 MB.
|
||||
# Avoid silent receive-buffer drops under load.
|
||||
net.core.rmem_max = 4194304
|
||||
net.core.wmem_max = 4194304
|
||||
EOF
|
||||
```
|
||||
|
||||
Apply:
|
||||
|
||||
```sh
|
||||
sudo sysctl --system
|
||||
```
|
||||
|
||||
The drop-in is loaded automatically on every boot.
|
||||
|
||||
## Step 4: Restart FIPS and verify the actual buffer size
|
||||
|
||||
After raising the host limits, restart the FIPS daemon so the next
|
||||
socket bind picks up the new ceiling:
|
||||
|
||||
```sh
|
||||
sudo systemctl restart fips
|
||||
```
|
||||
|
||||
The daemon logs the actual buffer sizes at startup:
|
||||
|
||||
```text
|
||||
UDP transport started local_addr=0.0.0.0:2121 recv_buf=4194304 send_buf=4194304
|
||||
```
|
||||
|
||||
If `recv_buf` or `send_buf` shows a smaller number than expected, the
|
||||
host sysctl is still clamping. Recheck `sysctl net.core.rmem_max
|
||||
net.core.wmem_max` and confirm the drop-in file is being loaded
|
||||
(`sudo sysctl --system` prints the loaded files).
|
||||
|
||||
## Docker and other container hosts
|
||||
|
||||
Containers share the host kernel, so sysctls apply to the host, not
|
||||
the container. If you run FIPS inside Docker, set
|
||||
`net.core.rmem_max` / `net.core.wmem_max` on the **Docker host**, not
|
||||
inside the container. Container privileges (cap_sys_admin) and
|
||||
`--sysctl` flags do not let you raise these particular limits from
|
||||
inside a container — they are global to the host network namespace.
|
||||
|
||||
For Kubernetes deployments, the host-level sysctl tuning is the same;
|
||||
node-level configuration (DaemonSet with `privileged: true`, or a
|
||||
node-init script) is the typical mechanism.
|
||||
|
||||
## Tuning higher
|
||||
|
||||
The 4 MB ceiling is a conservative starting point. For very high
|
||||
throughput (multi-gigabit per second), raise both sysctls and the
|
||||
corresponding `transports.udp.recv_buf_size` /
|
||||
`transports.udp.send_buf_size` config values together. Setting a config value
|
||||
larger than the host ceiling silently clamps to the ceiling, so both
|
||||
must move in lockstep.
|
||||
|
||||
## See also
|
||||
|
||||
- [../design/fips-transport-layer.md](../design/fips-transport-layer.md)
|
||||
— UDP transport design, why FIPS requests larger buffers,
|
||||
`SO_RXQ_OVFL` and ECN integration
|
||||
- [../reference/configuration.md](../reference/configuration.md) —
|
||||
`transports.udp.recv_buf_size` and `send_buf_size` defaults
|
||||
|
After Width: | Height: | Size: 1.2 MiB |
|
After Width: | Height: | Size: 1.2 MiB |
@@ -0,0 +1,25 @@
|
||||
# Reference
|
||||
|
||||
Information-oriented technical descriptions for lookup on demand.
|
||||
Reference content describes *what is*: wire formats, configuration
|
||||
keys, command-line flags, control-socket commands, default values,
|
||||
file paths, exit codes. It is consulted, not read end-to-end.
|
||||
|
||||
Reference is austere by design: minimal narrative, no opinions, no
|
||||
guidance on when to use a feature. The "why" lives in design/; the
|
||||
"how do I accomplish X" lives in how-to/.
|
||||
|
||||
## Available Reference
|
||||
|
||||
| Document | Scope |
|
||||
| -------- | ----- |
|
||||
| [wire-formats.md](wire-formats.md) | All FMP and FSP message byte layouts, encapsulation walkthrough |
|
||||
| [configuration.md](configuration.md) | Full YAML configuration reference for the daemon and gateway |
|
||||
| [security.md](security.md) | nftables baseline, peer ACL, cryptographic primitives, rekey defaults, threat-resistance matrix |
|
||||
| [nostr-events.md](nostr-events.md) | Kind 37195 advert, Kind 21059 traversal signaling, Kind 10050 inbox relays |
|
||||
| [transports.md](transports.md) | Per-transport statistics counter inventory |
|
||||
| [control-socket.md](control-socket.md) | Line-delimited JSON control protocol for the daemon and gateway |
|
||||
| [cli-fips.md](cli-fips.md) | `fips` daemon CLI: options, exit codes, environment, files |
|
||||
| [cli-fipsctl.md](cli-fipsctl.md) | `fipsctl` control-client: subcommands, options, exit codes |
|
||||
| [cli-fipstop.md](cli-fipstop.md) | `fipstop` live-status TUI: tabs, keybindings |
|
||||
| [cli-fips-gateway.md](cli-fips-gateway.md) | `fips-gateway` service CLI: options, exit codes, files |
|
||||
@@ -0,0 +1,125 @@
|
||||
# `fips-gateway`
|
||||
|
||||
Long-running service that bridges a LAN segment into the FIPS mesh.
|
||||
|
||||
## Synopsis
|
||||
|
||||
```text
|
||||
fips-gateway [-c FILE] [-l LEVEL]
|
||||
```
|
||||
|
||||
## Description
|
||||
|
||||
`fips-gateway` runs alongside `fips` on the same host, reads the same
|
||||
`fips.yaml`, and exposes two complementary functions to the LAN it
|
||||
fronts:
|
||||
|
||||
- **Outbound (LAN -> mesh).** Allocates a virtual IPv6 from a managed
|
||||
pool when a LAN client resolves `<npub>.fips`, installs nftables
|
||||
DNAT/SNAT/masquerade rules so the client's traffic is rewritten and
|
||||
carried into the mesh through the daemon's `fips0` adapter.
|
||||
- **Inbound (mesh -> LAN).** Installs nftables DNAT and LAN-side
|
||||
masquerade rules so mesh-side traffic arriving on `fips0` for the
|
||||
configured listen ports is rewritten to a LAN `host:port`, per the
|
||||
`gateway.port_forwards[]` block.
|
||||
|
||||
The service runs alongside `fips`, not as a replacement for it:
|
||||
the daemon must be running on the same host with the TUN adapter
|
||||
and DNS resolver enabled. The gateway is read-only with respect to the
|
||||
daemon's state, and connects to the daemon's resolver only — it is
|
||||
not a peer. For the architecture, see
|
||||
[../design/fips-gateway.md](../design/fips-gateway.md).
|
||||
|
||||
`fips-gateway` is **Linux-only**. The binary errors out and exits with
|
||||
status `1` on any other platform, since the NAT pipeline is built on
|
||||
nftables and proxy NDP. See
|
||||
[Configuration](#configuration) for the platform notes that follow
|
||||
from this.
|
||||
|
||||
## Options
|
||||
|
||||
| Flag | Argument | Default | Description |
|
||||
| ---- | -------- | ------- | ----------- |
|
||||
| `-c`, `--config` | `FILE` | *(default search paths)* | Use `FILE` as the configuration. Skips the default search paths. |
|
||||
| `-l`, `--log-level` | `LEVEL` | `info` | Tracing level: `trace`, `debug`, `info`, `warn`, `error`. Overridden by `RUST_LOG` if set (see [Environment](#environment)). |
|
||||
| `-V` | — | — | Print the short version. |
|
||||
| `--version` | — | — | Print the long version (short version plus build target triple). |
|
||||
| `-h`, `--help` | — | — | Print usage and exit. |
|
||||
|
||||
## Configuration
|
||||
|
||||
`fips-gateway` reads the same `fips.yaml` as `fips`; the gateway is
|
||||
configured under the top-level `gateway:` block. The block must
|
||||
include at minimum `enabled: true`, `pool`, and `lan_interface`. For
|
||||
each field — pool, LAN interface, DNS listener, conntrack overrides,
|
||||
and inbound `port_forwards[]` — see the
|
||||
[Gateway section](configuration.md#gateway-gateway) of the
|
||||
configuration reference.
|
||||
|
||||
The same default search paths apply as for `fips`
|
||||
(see [`fips`](cli-fips.md#files)); `-c FILE` overrides the search.
|
||||
The gateway must be able to read the same configuration file the
|
||||
daemon is reading, or the two will disagree about pool, DNS port,
|
||||
and LAN interface.
|
||||
|
||||
For deployment recipes, see
|
||||
[../how-to/deploy-gateway.md](../how-to/deploy-gateway.md) (manual
|
||||
Linux host) and
|
||||
[../tutorials/deploy-fips-gateway.md](../tutorials/deploy-fips-gateway.md)
|
||||
(OpenWrt walk-through).
|
||||
|
||||
## Exit Codes
|
||||
|
||||
| Code | Meaning |
|
||||
| ---- | ------- |
|
||||
| `0` | Clean shutdown after `SIGINT` / `SIGTERM`. |
|
||||
| `1` | Non-Linux platform, configuration load failure, missing or invalid `gateway:` block, NAT/network setup failure, or control-socket bind failure. The reason is printed to stderr or the log before exit. |
|
||||
|
||||
## Environment
|
||||
|
||||
| Variable | Description |
|
||||
| -------- | ----------- |
|
||||
| `RUST_LOG` | Tracing filter directive. Takes precedence over `--log-level`. Examples: `info`, `debug`, `fips=trace,fips::gateway=debug`. |
|
||||
|
||||
## Files
|
||||
|
||||
| Path | Purpose |
|
||||
| ---- | ------- |
|
||||
| `/etc/fips/fips.yaml` | Gateway configuration (top-level `gateway:` block). Same file the daemon reads. |
|
||||
| `/run/fips/gateway.sock` | Gateway control socket. Hardcoded path; chowned to group `fips` (mode `0770`) at startup so members of that group can query without sudo. |
|
||||
| `inet fips_gateway` (nftables) | NAT table the gateway installs and tears down. View with `nft list table inet fips_gateway`. |
|
||||
|
||||
The gateway also adds and removes a `local <pool-cidr> dev lo` route
|
||||
in the local routing table so the kernel accepts pool addresses as
|
||||
locally-owned.
|
||||
|
||||
## Control Socket
|
||||
|
||||
`fips-gateway` exposes a JSON line-protocol control socket separate
|
||||
from the daemon's. The command set (`show_gateway`, `show_mappings`)
|
||||
and JSON shapes are documented in the
|
||||
[Gateway Command Catalog](control-socket.md#gateway-command-catalog).
|
||||
|
||||
There is no `fipsctl` subcommand for the gateway — query the socket
|
||||
directly with `nc -U`, or watch the **Gateway** tab in
|
||||
[`fipstop`](cli-fipstop.md), which polls the gateway socket
|
||||
automatically.
|
||||
|
||||
## See also
|
||||
|
||||
- [`fips`](cli-fips.md) — the daemon. Required to be running on the
|
||||
same host.
|
||||
- [`fipstop`](cli-fipstop.md) — the live-status TUI; its Gateway tab
|
||||
polls the gateway control socket.
|
||||
- [configuration.md § Gateway](configuration.md#gateway-gateway) —
|
||||
full `gateway.*` block reference.
|
||||
- [control-socket.md § Gateway Command Catalog](control-socket.md#gateway-command-catalog)
|
||||
— wire protocol for the gateway socket.
|
||||
- [../design/fips-gateway.md](../design/fips-gateway.md) — design,
|
||||
NAT pipeline, virtual IP pool lifecycle.
|
||||
- [../how-to/deploy-gateway.md](../how-to/deploy-gateway.md) — manual
|
||||
Linux deployment.
|
||||
- [../how-to/troubleshoot-gateway.md](../how-to/troubleshoot-gateway.md)
|
||||
— diagnostic recipes.
|
||||
- [../tutorials/deploy-fips-gateway.md](../tutorials/deploy-fips-gateway.md)
|
||||
— OpenWrt walk-through.
|
||||
@@ -0,0 +1,93 @@
|
||||
# `fips`
|
||||
|
||||
The FIPS mesh network daemon.
|
||||
|
||||
## Synopsis
|
||||
|
||||
```text
|
||||
fips [-c FILE]
|
||||
```
|
||||
|
||||
On Windows the same binary additionally accepts `--install-service`,
|
||||
`--uninstall-service`, and (used internally by the service control
|
||||
manager) `--service`.
|
||||
|
||||
## Description
|
||||
|
||||
`fips` is the FIPS daemon. It loads a YAML configuration, resolves an
|
||||
identity, brings up the TUN adapter, listens on configured transports,
|
||||
authenticates peers, maintains the spanning tree, and forwards mesh
|
||||
traffic. There is one daemon per node.
|
||||
|
||||
The daemon stays in the foreground, logging to stderr, until it
|
||||
receives `SIGINT` or `SIGTERM`. On Windows, the service variant is
|
||||
controlled through the standard service control manager.
|
||||
|
||||
## Options
|
||||
|
||||
| Flag | Argument | Description |
|
||||
| ---- | -------- | ----------- |
|
||||
| `-c`, `--config` | `FILE` | Use `FILE` as the configuration. Skips the default search paths. |
|
||||
| `-V` | — | Print the short version (e.g. `0.3.0-dev (rev abcdef1)`). |
|
||||
| `--version` | — | Print the long version: short version plus build target triple. |
|
||||
| `-h`, `--help` | — | Print usage and exit. |
|
||||
| `--install-service` | — | (Windows only) Install `fips` as a Windows service. Requires Administrator. |
|
||||
| `--uninstall-service` | — | (Windows only) Uninstall the Windows service. Requires Administrator. |
|
||||
| `--service` | — | (Windows only, internal) Run as a Windows service. Invoked by the service control manager — not for direct use. |
|
||||
|
||||
There are no other CLI flags; all daemon behaviour is governed by the
|
||||
YAML configuration. See [configuration.md](configuration.md).
|
||||
|
||||
## Exit Codes
|
||||
|
||||
| Code | Meaning |
|
||||
| ---- | ------- |
|
||||
| `0` | Clean shutdown after `SIGINT` / `SIGTERM`. |
|
||||
| `1` | Failed to load configuration, resolve identity, construct the node, or start the node. The reason is printed to stderr before exit. |
|
||||
|
||||
## Environment
|
||||
|
||||
| Variable | Description |
|
||||
| -------- | ----------- |
|
||||
| `RUST_LOG` | Tracing filter directive. Overrides `node.log_level` from the config. Examples: `info`, `debug`, `fips=trace,fips::node::handlers::mmp=debug`. |
|
||||
| `XDG_RUNTIME_DIR` | Used to derive the default control-socket path when `/run/fips` does not exist. See [control-socket.md](control-socket.md). |
|
||||
| `FIPS_CONFIG` | (Windows service mode only) Path to the configuration file when the daemon runs under the service control manager. |
|
||||
|
||||
The daemon also clamps the `nostr_relay_pool`, `nostr_sdk`, and `nostr`
|
||||
log targets to `info` whenever the effective log level is below
|
||||
`trace`, so that `RUST_LOG=debug` does not flood the journal with raw
|
||||
relay frames. To see those frames, set the level to `trace`.
|
||||
|
||||
## Files
|
||||
|
||||
`fips` looks for `fips.yaml` in the following locations, lowest to
|
||||
highest priority. All present files are merged in priority order; the
|
||||
highest-priority value wins.
|
||||
|
||||
| Priority | Path | Purpose |
|
||||
| -------- | ---- | ------- |
|
||||
| 1 | `/etc/fips/fips.yaml` | System-wide defaults |
|
||||
| 2 | `~/.config/fips/fips.yaml` | User preferences |
|
||||
| 3 | `~/.fips.yaml` | Legacy user config |
|
||||
| 4 | `./fips.yaml` | Deployment-specific overrides |
|
||||
|
||||
Adjacent to the highest-priority config file the daemon reads (or
|
||||
writes, on first start) the identity files:
|
||||
|
||||
| File | Mode | Purpose |
|
||||
| ---- | ---- | ------- |
|
||||
| `fips.key` | `0600` | Bech32 nsec for the persistent identity (Unix only; Windows inherits parent ACLs). |
|
||||
| `fips.pub` | `0644` | Bech32 npub corresponding to `fips.key`. |
|
||||
|
||||
When `node.identity.persistent` is `false` (the default), a fresh
|
||||
keypair is written to these files on every start.
|
||||
|
||||
The control socket path is derived per
|
||||
[control-socket.md](control-socket.md).
|
||||
|
||||
## See also
|
||||
|
||||
- [`fipsctl`](cli-fipsctl.md) — control-socket client.
|
||||
- [`fipstop`](cli-fipstop.md) — live-status TUI.
|
||||
- [configuration.md](configuration.md) — YAML reference.
|
||||
- [control-socket.md](control-socket.md) — control-socket protocol.
|
||||
@@ -0,0 +1,148 @@
|
||||
# `fipsctl`
|
||||
|
||||
Command-line client for the FIPS daemon's control socket.
|
||||
|
||||
## Synopsis
|
||||
|
||||
```text
|
||||
fipsctl [-s SOCKET] <subcommand> [args...]
|
||||
```
|
||||
|
||||
## Description
|
||||
|
||||
`fipsctl` connects to a running daemon over its control socket
|
||||
(Unix domain socket on Linux/macOS, TCP loopback on Windows), sends
|
||||
one JSON request, and pretty-prints the response. Exits with a
|
||||
non-zero status if the socket cannot be reached, the daemon returns an
|
||||
error, or the request times out.
|
||||
|
||||
`fipsctl keygen` is a special case: it does not contact the daemon and
|
||||
operates purely on local files.
|
||||
|
||||
For the line-delimited JSON wire protocol, see
|
||||
[control-socket.md](control-socket.md). For the YAML configuration
|
||||
that defines the socket location, see
|
||||
[configuration.md](configuration.md).
|
||||
|
||||
## Global Options
|
||||
|
||||
| Flag | Argument | Description |
|
||||
| ---- | -------- | ----------- |
|
||||
| `-s`, `--socket` | `PATH` | Override the control-socket path (Linux/macOS) or TCP port (Windows). |
|
||||
| `-V`, `--version` | — | Print the short version. |
|
||||
| `--version` | — | Print the long version. |
|
||||
| `-h`, `--help` | — | Print usage and exit. Per-subcommand help via `fipsctl <subcommand> --help`. |
|
||||
|
||||
## Subcommands
|
||||
|
||||
### `show <what>`
|
||||
|
||||
Read-only queries against the daemon. Each subcommand maps 1:1 to a
|
||||
control-socket query (see [control-socket.md](control-socket.md)) and
|
||||
prints the response's `data` object as pretty JSON.
|
||||
|
||||
| Subcommand | Control-socket command | Returns |
|
||||
| ---------- | ---------------------- | ------- |
|
||||
| `show status` | `show_status` | Node-level status: identity, version, peer/link/session counts, TUN state, recent sparklines. |
|
||||
| `show peers` | `show_peers` | Authenticated peer list with link IDs, transport addresses, MMP metrics, Noise/rekey state. |
|
||||
| `show links` | `show_links` | Active links (one per FMP-authenticated peer): direction, state, byte counters. |
|
||||
| `show tree` | `show_tree` | Spanning-tree state: root, my coordinates, parent, peer declarations. |
|
||||
| `show sessions` | `show_sessions` | End-to-end FSP sessions: state, traffic counters, session-MMP metrics, path MTU. |
|
||||
| `show bloom` | `show_bloom` | Bloom-filter state: own filter sequence, leaf dependents, per-peer filter summaries. |
|
||||
| `show mmp` | `show_mmp` | MMP metrics summary: per-peer link-layer metrics and per-session session-layer metrics. |
|
||||
| `show cache` | `show_cache` | Coordinate cache: TTL, fill ratio, per-destination coords and path MTU. |
|
||||
| `show connections` | `show_connections` | Pending handshake connections: state, idle time, resend count. |
|
||||
| `show transports` | `show_transports` | Transport instances: type, state, MTU, local address, per-transport stats. |
|
||||
| `show routing` | `show_routing` | Routing summary: pending lookups, retry state, forwarding/discovery/error/congestion counters. |
|
||||
| `show identity-cache` | `show_identity_cache` | Cached `(node_addr → npub)` entries with last-seen timestamps. |
|
||||
|
||||
### `acl <what>`
|
||||
|
||||
| Subcommand | Control-socket command | Returns |
|
||||
| ---------- | ---------------------- | ------- |
|
||||
| `acl show` | `show_acl` | Loaded peer-ACL state: allow/deny files, effective mode, default decision, entry counts. |
|
||||
|
||||
### `stats <what>`
|
||||
|
||||
Time-series metrics from the in-process history rings.
|
||||
|
||||
| Subcommand | Control-socket command | Description |
|
||||
| ---------- | ---------------------- | ----------- |
|
||||
| `stats list` | `show_stats_list` | Enumerate available metrics, their units, and the per-ring retention windows. |
|
||||
| `stats peers` | `show_stats_peers` | List peers tracked in stats history (active or recently active). |
|
||||
| `stats history <metric> [options]` | `show_stats_history` | Fetch a time-series window for one metric. |
|
||||
|
||||
`stats history` options:
|
||||
|
||||
| Flag | Argument | Default | Description |
|
||||
| ---- | -------- | ------- | ----------- |
|
||||
| `--peer` | `npub` or hostname | *(none)* | Required for per-peer metrics; resolves through `/etc/fips/hosts` if not an npub. |
|
||||
| `--window` | `<N>s` / `<N>m` / `<N>h` | `10m` | Window duration. |
|
||||
| `--granularity` | `1s` or `1m` | `1s` | Ring resolution. `1s` uses the fast ring; `1m` uses the slow ring. |
|
||||
| `--plot` | — | off | Render a Unicode-block sparkline to stdout instead of JSON. |
|
||||
|
||||
### `keygen [options]`
|
||||
|
||||
Generate a new FIPS identity keypair locally. Does not contact the
|
||||
daemon.
|
||||
|
||||
| Flag | Argument | Default | Description |
|
||||
| ---- | -------- | ------- | ----------- |
|
||||
| `-d`, `--dir` | `DIR` | `/etc/fips` (Unix), `%APPDATA%\fips` (Windows) | Output directory for `fips.key` and `fips.pub`. |
|
||||
| `-f`, `--force` | — | off | Overwrite an existing `fips.key`. |
|
||||
| `-s`, `--stdout` | — | off | Print `nsec` then `npub` to stdout instead of writing files. |
|
||||
|
||||
`fips.key` is written with mode `0600` and `fips.pub` with mode `0644`
|
||||
on Unix. After running `keygen`, set `node.identity.persistent: true`
|
||||
in `fips.yaml` or the daemon will overwrite the keys on next start.
|
||||
|
||||
### `connect <peer> <address> <transport>`
|
||||
|
||||
Tell the daemon to dial a peer over a specific transport.
|
||||
|
||||
| Argument | Description |
|
||||
| -------- | ----------- |
|
||||
| `peer` | npub (bech32) or hostname from `/etc/fips/hosts`. |
|
||||
| `address` | Transport endpoint, e.g. `192.168.1.10:2121`, `[2001:db8::1]:2121`, or a Tor onion. FIPS-mesh ULAs (`fd00::/8`) are rejected for the IP-based transports (udp, tcp, ethernet). |
|
||||
| `transport` | One of `udp`, `tcp`, `tor`, `ethernet`. |
|
||||
|
||||
### `disconnect <peer>`
|
||||
|
||||
Tell the daemon to drop a peer link.
|
||||
|
||||
| Argument | Description |
|
||||
| -------- | ----------- |
|
||||
| `peer` | npub (bech32) or hostname from `/etc/fips/hosts`. |
|
||||
|
||||
## Exit Codes
|
||||
|
||||
| Code | Meaning |
|
||||
| ---- | ------- |
|
||||
| `0` | Daemon returned `{"status":"ok",...}`. |
|
||||
| `1` | Argument parse failure, control-socket connection failure, daemon returned `{"status":"error",...}`, or local I/O failure (keygen). The error message is printed to stderr. |
|
||||
|
||||
## Environment
|
||||
|
||||
| Variable | Description |
|
||||
| -------- | ----------- |
|
||||
| `XDG_RUNTIME_DIR` | Used to derive the default control-socket path when `/run/fips` is absent. |
|
||||
|
||||
`fipsctl` does not consume `RUST_LOG`; logging is for the daemon.
|
||||
|
||||
## Files
|
||||
|
||||
| Path | Purpose |
|
||||
| ---- | ------- |
|
||||
| `/etc/fips/hosts` | Maps hostnames to npubs for the `connect`, `disconnect`, and `--peer` arguments. See [configuration.md](configuration.md). |
|
||||
| Control socket (default) | Same resolution as the daemon: `/run/fips/control.sock` if present, else `$XDG_RUNTIME_DIR/fips/control.sock`, else `/tmp/fips-control.sock` (Unix); TCP `localhost:21210` (Windows). |
|
||||
|
||||
If you get `Permission denied` connecting to the socket on Linux,
|
||||
add your user to the `fips` group (`sudo usermod -aG fips $USER`)
|
||||
and log out and back in.
|
||||
|
||||
## See also
|
||||
|
||||
- [`fips`](cli-fips.md) — the daemon.
|
||||
- [`fipstop`](cli-fipstop.md) — live-status TUI.
|
||||
- [control-socket.md](control-socket.md) — wire protocol.
|
||||
- [configuration.md](configuration.md) — YAML reference.
|
||||
@@ -0,0 +1,155 @@
|
||||
# `fipstop`
|
||||
|
||||
Live-status terminal UI for a running FIPS daemon.
|
||||
|
||||
## Synopsis
|
||||
|
||||
```text
|
||||
fipstop [-s SOCKET] [--gateway-socket PATH] [-r SECONDS]
|
||||
```
|
||||
|
||||
## Description
|
||||
|
||||
`fipstop` is a `ratatui`-based dashboard. It opens the daemon control
|
||||
socket, polls a small set of `show_*` queries on a timer, and renders
|
||||
the state in a tabbed full-screen UI. A separate poll runs against the
|
||||
gateway control socket when the Gateway tab is active.
|
||||
|
||||
`fipstop` is read-only — it cannot mutate daemon state. Use
|
||||
[`fipsctl`](cli-fipsctl.md) for `connect` / `disconnect` and friends.
|
||||
|
||||
## Options
|
||||
|
||||
| Flag | Argument | Default | Description |
|
||||
| ---- | -------- | ------- | ----------- |
|
||||
| `-s`, `--socket` | `PATH` | (auto) | Daemon control-socket path / port. Same default as `fipsctl`. |
|
||||
| `--gateway-socket` | `PATH` | (auto) | `fips-gateway` control-socket path / port. Default: `/run/fips/gateway.sock` (Unix), TCP port `21211` (Windows). |
|
||||
| `-r`, `--refresh` | `SECONDS` | `2` | Poll interval. |
|
||||
| `-V`, `--version` | — | — | Print short version. |
|
||||
| `--version` | — | — | Print long version. |
|
||||
| `-h`, `--help` | — | — | Print usage and exit. |
|
||||
|
||||
## Tabs
|
||||
|
||||
Tabs cycle in this order. Each tab issues the listed control-socket
|
||||
query on its first activation and on every refresh tick while active.
|
||||
|
||||
| Tab | Query | Shows |
|
||||
| --- | ----- | ----- |
|
||||
| **Node** | `show_status` (+ `show_listening_sockets`) | Identity, version, uptime, peer/link/session counts, sparklines for mesh size, tree depth, peer count, bytes, loss. The Traffic block on this tab is split: TUN counters on the left, the **Listening on fips0** panel on the right (see below). |
|
||||
| **Peers** | `show_peers` (+ `show_links`, `show_transports` cross-refs) | Authenticated peers in a table. Selecting a row and pressing Enter opens a detail view. |
|
||||
| **Transports** | `show_transports` (+ `show_links`, `show_peers` cross-refs) | Tree of transport instances with per-link children when expanded. |
|
||||
| **Sessions** | `show_sessions` | End-to-end FSP sessions. |
|
||||
| **Tree** | `show_tree` | Spanning-tree state and per-peer coordinates. |
|
||||
| **Filters** | `show_bloom` | Per-peer Bloom-filter state. |
|
||||
| **Performance** | `show_mmp` | Link-layer and session-layer MMP metrics. |
|
||||
| **Routing** | `show_routing` (+ `show_cache` cross-ref) | Forwarding/discovery counters, pending lookups, retry state. |
|
||||
| **Graphs** | `show_stats_history` family + `show_stats_peers` | Stacked time-series plots. Three modes: node-level metrics, one metric across peers, all metrics for one peer. |
|
||||
| **Gateway** | `show_gateway` and `show_mappings` against the gateway socket | Pool utilisation and per-mapping state when `fips-gateway` is running. Empty when the gateway socket is unreachable. |
|
||||
|
||||
The cycle order in the UI is: Node → Peers → Transports → Sessions →
|
||||
Tree → Filters → Performance → Routing → Graphs → Gateway. The Links
|
||||
and Cache tabs are not in the cycle but are fetched as cross-references
|
||||
to populate Peers, Transports, and Routing detail views.
|
||||
|
||||
## Listening on fips0 panel (Node tab)
|
||||
|
||||
The right half of the Node tab's Traffic block lists local IPv6
|
||||
listening sockets reachable from `fips0`, paired with the current
|
||||
`inet fips` baseline filter classification for each (proto, port).
|
||||
The panel exists to remind the operator which local services are
|
||||
exposed to the mesh and which of those are admitted by the
|
||||
default-deny firewall.
|
||||
|
||||
| Column | Meaning |
|
||||
| ------ | ------- |
|
||||
| **Proto** | `tcp` or `udp`. IPv4 listeners are not enumerated; `fips0` is IPv6-only. |
|
||||
| **Port** | Listening port number. |
|
||||
| **Process** | `comm(pid)` resolved by walking `/proc/<pid>/fd/`. A trailing `*` marks wildcard binds (`local_addr == ::`) — the bind is not fips0-specific, so the operator sees that the service is exposed across every interface, not just the mesh. |
|
||||
| **State** | `OPEN` (default White) — the baseline filter has a canonical accept rule for this (proto, port). `filt` (DarkGray) — chain falls through to `counter drop`. `filt?` (DarkGray) — a rule references the port but uses matchers (saddr filter, jump, daddr) the panel cannot fully decompose; operator should `nft list table inet fips` to confirm. |
|
||||
|
||||
When `fips-firewall.service` is **not** active, the `inet fips`
|
||||
table is absent. The panel renders every row in default White and
|
||||
replaces the title with a yellow banner reading
|
||||
"`Listening on fips0 fips-firewall.service inactive — all listeners exposed`".
|
||||
|
||||
The panel is read-only and unselectable. It refreshes on the same
|
||||
poll tick as the rest of the Node tab. Sockets owned by other users
|
||||
that the daemon could not resolve to a PID render as `?` in the
|
||||
Process column; this only happens if the daemon itself is running
|
||||
without root privileges (an unusual dev setup), since walking
|
||||
`/proc/<pid>/fd/` for processes the daemon does not own requires
|
||||
elevated capabilities.
|
||||
|
||||
The panel is Linux-only; on non-Linux daemons the query returns an
|
||||
empty list and the panel hides.
|
||||
|
||||
## Keybindings
|
||||
|
||||
### Global
|
||||
|
||||
| Key | Action |
|
||||
| --- | ------ |
|
||||
| `q`, `Ctrl-C` | Quit. |
|
||||
| `Tab` | Next tab. |
|
||||
| `Shift-Tab` | Previous tab. |
|
||||
| `g` | Jump to the Graphs tab. |
|
||||
| `Esc` | Close detail view (if open). |
|
||||
|
||||
### Table tabs (Peers, Sessions, Transports, Gateway)
|
||||
|
||||
| Key | Action |
|
||||
| --- | ------ |
|
||||
| `Up`, `Down` | Move row selection. |
|
||||
| `Enter` | Open detail view for the selected row. |
|
||||
|
||||
### Transports tab (extra)
|
||||
|
||||
| Key | Action |
|
||||
| --- | ------ |
|
||||
| `Right`, `Space` | Expand the selected transport row to show its links. |
|
||||
| `Left` | Collapse the selected transport row. |
|
||||
| `e` | Expand all transports. |
|
||||
| `c` | Collapse all transports. |
|
||||
|
||||
### Graphs tab (extra)
|
||||
|
||||
| Key | Action |
|
||||
| --- | ------ |
|
||||
| `Up`, `Down` | Scroll within the stacked plots. |
|
||||
| `Right`, `Space` | Next time window. Cycles `1m / 1s` → `10m / 1s` → `1h / 1s` → `24h / 1m`. |
|
||||
| `Left` | Previous time window. |
|
||||
| `m` | Cycle view mode: `Node` (stacked node metrics) → `MetricByPeer` (one per-peer metric across all peers) → `PeerByMetric` (all per-peer metrics for one peer). |
|
||||
| `n` | Next selector (next per-peer metric in MetricByPeer; next peer in PeerByMetric). |
|
||||
| `Shift-N` | Previous selector. |
|
||||
|
||||
## Exit Codes
|
||||
|
||||
| Code | Meaning |
|
||||
| ---- | ------- |
|
||||
| `0` | Normal quit. |
|
||||
| `1` | Failed to initialise the terminal. The reason is printed to stderr. |
|
||||
|
||||
A failure to reach the daemon socket is **not** fatal: the dashboard
|
||||
displays "Disconnected" in the status bar and retries on every refresh
|
||||
tick.
|
||||
|
||||
## Environment
|
||||
|
||||
| Variable | Description |
|
||||
| -------- | ----------- |
|
||||
| `XDG_RUNTIME_DIR` | Used to derive the default control-socket and gateway-socket paths when `/run/fips` is absent. |
|
||||
|
||||
## Files
|
||||
|
||||
Same control-socket resolution rules as
|
||||
[`fipsctl`](cli-fipsctl.md#files). The gateway socket follows the same
|
||||
pattern with `gateway.sock` in place of `control.sock`, falling back
|
||||
to `/tmp/fips-gateway.sock` if neither system path nor
|
||||
`XDG_RUNTIME_DIR` is available.
|
||||
|
||||
## See also
|
||||
|
||||
- [`fipsctl`](cli-fipsctl.md) — issue mutating commands.
|
||||
- [`fips`](cli-fips.md) — the daemon.
|
||||
- [control-socket.md](control-socket.md) — wire protocol fipstop polls.
|
||||
@@ -0,0 +1,167 @@
|
||||
# Control Socket Protocol
|
||||
|
||||
The FIPS daemon and `fips-gateway` each expose a local control socket
|
||||
that accepts line-delimited JSON requests and returns line-delimited
|
||||
JSON responses. `fipsctl` and `fipstop` are clients of this protocol;
|
||||
operators can also drive it directly with any tool that can speak
|
||||
length-bounded JSON over a stream socket.
|
||||
|
||||
## Connection
|
||||
|
||||
### Linux / macOS
|
||||
|
||||
A Unix domain socket. The default path is resolved in this order:
|
||||
|
||||
1. `/run/fips/control.sock` (or `/run/fips/gateway.sock` for the
|
||||
gateway), if `/run/fips` exists. This is what the `fips.service`
|
||||
systemd unit creates.
|
||||
2. `$XDG_RUNTIME_DIR/fips/control.sock` otherwise.
|
||||
3. `/tmp/fips-control.sock` if neither of the above is available.
|
||||
|
||||
The daemon `chown`s the socket file and its parent directory to the
|
||||
`fips` group at bind time and sets mode `0770`. Members of the `fips`
|
||||
group can therefore connect without root. Add a user with
|
||||
`sudo usermod -aG fips $USER` (re-login required).
|
||||
|
||||
The path can be overridden at the daemon side via
|
||||
`node.control.socket_path` in the YAML config, and at the client side
|
||||
via `fipsctl -s PATH` or `fipstop -s PATH`.
|
||||
|
||||
### Windows
|
||||
|
||||
A TCP listener bound to `127.0.0.1`. The daemon's port is `21210` by
|
||||
default; the gateway's is `21211`. Only loopback connections are
|
||||
accepted. Override via `node.control.socket_path` (which takes a port
|
||||
number string on Windows).
|
||||
|
||||
Windows TCP does not provide filesystem-level ACLs — any local user
|
||||
can connect. See the security note in
|
||||
[configuration.md](configuration.md#control-socket-nodecontrol).
|
||||
|
||||
## Request Format
|
||||
|
||||
One JSON object per line, terminated by `\n`. Maximum request size is
|
||||
4096 bytes; longer requests are dropped with `request too large`.
|
||||
|
||||
```json
|
||||
{"command": "<name>", "params": {<object>}}
|
||||
```
|
||||
|
||||
| Field | Type | Required | Description |
|
||||
| ----- | ---- | -------- | ----------- |
|
||||
| `command` | string | yes | Command name. See [Daemon command catalog](#daemon-command-catalog) and [Gateway command catalog](#gateway-command-catalog). |
|
||||
| `params` | object | only for commands that take parameters | Parameter object. Unknown fields are ignored; missing required fields produce an error response. |
|
||||
|
||||
Unknown top-level fields in the request are silently ignored.
|
||||
|
||||
## Response Format
|
||||
|
||||
One JSON object per line.
|
||||
|
||||
```json
|
||||
{"status": "ok", "data": {<object>}}
|
||||
{"status": "error", "message": "<reason>"}
|
||||
```
|
||||
|
||||
| Field | Type | When present |
|
||||
| ----- | ---- | ------------ |
|
||||
| `status` | string | always; one of `"ok"` or `"error"`. |
|
||||
| `data` | object | on `ok` responses. |
|
||||
| `message` | string | on `error` responses. |
|
||||
|
||||
### I/O timeouts
|
||||
|
||||
The daemon enforces a 5-second timeout for both the request read and
|
||||
the response write. If the connection idles longer than that, the
|
||||
daemon closes it with no response.
|
||||
|
||||
### Common error messages
|
||||
|
||||
| Message | Cause |
|
||||
| ------- | ----- |
|
||||
| `empty request` | Connection closed before a newline was received. |
|
||||
| `invalid request: <serde error>` | Malformed JSON or missing `command`. |
|
||||
| `request too large` | Request exceeded 4096 bytes. |
|
||||
| `read timeout` / `read error: ...` | Slow client or transport failure. |
|
||||
| `unknown command: <name>` | Command not registered with this daemon. |
|
||||
| `missing params for <name>` | Command requires `params` but none were provided. |
|
||||
| `missing '<field>' parameter` | Required parameter missing. |
|
||||
| `query timeout` | Internal handler did not respond within 5 seconds. |
|
||||
| `node shutting down` | Daemon is exiting. |
|
||||
| `gateway not yet initialized` | (Gateway socket only) snapshot has not been published yet. |
|
||||
|
||||
## Daemon Command Catalog
|
||||
|
||||
Read-only queries are dispatched in `src/control/queries.rs`;
|
||||
mutating commands are dispatched in `src/control/commands.rs`. The
|
||||
table below lists every command currently registered.
|
||||
|
||||
### Read-only queries
|
||||
|
||||
| Command | Params | `data` shape (top-level keys) |
|
||||
| ------- | ------ | ----------------------------- |
|
||||
| `show_status` | — | `version`, `npub`, `node_addr`, `ipv6_addr`, `state`, `is_leaf_only`, `peer_count`, `session_count`, `link_count`, `transport_count`, `connection_count`, `tun_state`, `tun_name`, `effective_ipv6_mtu`, `control_socket`, `pid`, `exe_path`, `uptime_secs`, `estimated_mesh_size`, `forwarding`, `sparklines`. |
|
||||
| `show_acl` | — | `allow_file`, `deny_file`, `enforcement_active`, `effective_mode`, `default_decision`, `allow_all`, `deny_all`, `allow_file_entries`, `deny_file_entries`, `allow_entries`, `deny_entries`. |
|
||||
| `show_peers` | — | `peers[]` — per-peer object: `node_addr`, `npub`, `display_name`, `ipv6_addr`, `connectivity`, `link_id`, `direction`, `transport_addr`, `transport_type`, `is_parent`, `is_child`, `tree_depth`, `stats`, `noise`, `current_k_bit`, `mmp`, plus optional `nostr_traversal`, `rekey_in_progress`, `rekey_draining`. |
|
||||
| `show_links` | — | `links[]` — `link_id`, `transport_id`, `remote_addr`, `direction`, `state`, `created_at_ms`, `stats`. |
|
||||
| `show_tree` | — | `my_node_addr`, `root`, `is_root`, `depth`, `my_coords[]`, `parent`, `parent_display_name`, `declaration_sequence`, `declaration_signed`, `peer_tree_count`, `peers[]`, `stats`. |
|
||||
| `show_sessions` | — | `sessions[]` — `remote_addr`, `npub`, `display_name`, `state` (`established`, `initiating`, `awaiting_msg3`, `unknown`), `is_initiator`, `last_activity_ms`, `stats`, optional `mmp`, `current_k_bit`, `is_draining`. |
|
||||
| `show_bloom` | — | `own_node_addr`, `is_leaf_only`, `sequence`, `leaf_dependent_count`, `leaf_dependents[]`, `peer_filters[]`, `stats`. |
|
||||
| `show_mmp` | — | `peers[]` (link-layer per peer), `sessions[]` (session-layer per session). Each entry includes loss/RTT/ETX/goodput, smoothed values, trends. |
|
||||
| `show_cache` | — | `count`, `max_entries`, `fill_ratio`, `default_ttl_ms`, `expired`, `avg_age_ms`, `entries[]` — per-destination coords, depth, age, last-used, optional `path_mtu`. |
|
||||
| `show_connections` | — | `connections[]` — pending handshakes: `link_id`, `direction`, `handshake_state`, `started_at_ms`, `idle_ms`, `resend_count`, optional `expected_peer`. |
|
||||
| `show_transports` | — | `transports[]` — `transport_id`, `type`, `state`, `mtu`, `name`, `local_addr`, optional `tor_mode`, `onion_address`, `tor_monitoring`, `stats`. |
|
||||
| `show_routing` | — | `coord_cache_entries`, `identity_cache_entries`, `pending_lookups[]`, `pending_tun_destinations`, `pending_tun_packets`, `recent_requests`, `retries[]`, `forwarding`, `discovery`, `error_signals`, `congestion`. |
|
||||
| `show_identity_cache` | — | `entries[]`, `count`, `max_entries`. Each entry: `node_addr`, `npub`, `display_name`, `ipv6_addr`, `last_seen_ms`, `age_ms`. |
|
||||
| `show_listening_sockets` | — | `fips0_addr`, `firewall_active` (bool — `inet fips` table loaded), `sockets[]`. Each entry: `proto` (`tcp` / `udp`), `local_addr` (`::` or the node's fd00::/8 address), `port`, `pid` (nullable), `process` (nullable), `wildcard_bind` (bool — `local_addr == ::`), `filter` (`accept` / `drop` / `unknown` / `no_firewall`). Linux-only; returns an empty `sockets[]` on other platforms. |
|
||||
| `show_stats_list` | — | `metrics[]` (each with `name`, `unit`, `scope`), `fast_ring_seconds`, `slow_ring_minutes`, `peer_retention_seconds`. |
|
||||
| `show_stats_history` | `metric` (req), `peer` (req for per-peer metrics), `window` (`<N>s` / `<N>m` / `<N>h`, default `10m`), `granularity` (`1s` / `1m`, default `1s`) | A single `Series`: `metric`, `unit`, `granularity_seconds`, `values[]`. |
|
||||
| `show_stats_all_history` | `peer` (optional npub), `window`, `granularity` | `granularity_seconds`, `window_seconds`, `peer`, `series[]` (one per metric). |
|
||||
| `show_stats_peers` | — | `peers[]`, `count`. Each entry: `npub`, `node_addr`, `display_name`, `is_active`, `first_seen_secs_ago`, `last_contact_secs_ago`. |
|
||||
| `show_stats_history_all_peers` | `metric` (req per-peer name), `window`, `granularity` | `metric`, `unit`, `granularity_seconds`, `window_seconds`, `peers[]` (each with `node_addr`, `display_name`, `is_active`, `values[]`). |
|
||||
|
||||
The schema of each query response is pinned by snapshot tests in
|
||||
`src/control/snapshots/`; intentional schema changes regenerate those
|
||||
fixtures.
|
||||
|
||||
### Mutating commands
|
||||
|
||||
| Command | Required params | Behaviour |
|
||||
| ------- | --------------- | --------- |
|
||||
| `connect` | `npub` (bech32), `address` (transport endpoint), `transport` (`udp`, `tcp`, `tor`, `ethernet`) | Asks the node to dial the peer over the named transport. Returns the API result on success or an error string on failure. |
|
||||
| `disconnect` | `npub` (bech32) | Asks the node to drop the link to the named peer. |
|
||||
|
||||
Both commands run on the daemon's main task and may block briefly
|
||||
while the node mutates its state.
|
||||
|
||||
## Gateway Command Catalog
|
||||
|
||||
`fips-gateway` exposes a separate control socket with its own command
|
||||
set. Dispatch lives in `src/gateway/control.rs`.
|
||||
|
||||
| Command | Params | `data` shape |
|
||||
| ------- | ------ | ------------ |
|
||||
| `show_gateway` | — | `pool_total`, `pool_allocated`, `pool_active`, `pool_draining`, `pool_free`, `nat_mappings`, `dns_listen`, `uptime_secs`, `pool_cidr`, `lan_interface`, `dns_upstream`, `dns_ttl`, `pool_grace_period`. |
|
||||
| `show_mappings` | — | `mappings[]` — `virtual_ip`, `mesh_addr`, `node_addr`, `dns_name`, `state` (`Allocated`, `Active`, `Draining`), `sessions`, `age_secs`, `last_ref_secs`. |
|
||||
|
||||
Until the first snapshot has been published (very early in startup),
|
||||
both commands return `gateway not yet initialized`.
|
||||
|
||||
## Driving the Socket Directly
|
||||
|
||||
```sh
|
||||
# Linux / macOS
|
||||
echo '{"command":"show_status"}' | sudo nc -U /run/fips/control.sock
|
||||
|
||||
# Windows (PowerShell with a TCP-capable tool of your choice)
|
||||
```
|
||||
|
||||
The newline at the end of the request is required: the daemon reads
|
||||
one line per connection. The connection is closed after the single
|
||||
response is written.
|
||||
|
||||
## See also
|
||||
|
||||
- [`fipsctl`](cli-fipsctl.md) — full-featured client.
|
||||
- [`fipstop`](cli-fipstop.md) — read-only TUI.
|
||||
- [configuration.md](configuration.md) — `node.control.*` keys.
|
||||
|
Before Width: | Height: | Size: 1.6 KiB After Width: | Height: | Size: 1.6 KiB |
|
Before Width: | Height: | Size: 2.3 KiB After Width: | Height: | Size: 2.3 KiB |
|
Before Width: | Height: | Size: 2.0 KiB After Width: | Height: | Size: 2.0 KiB |
|
Before Width: | Height: | Size: 1.0 KiB After Width: | Height: | Size: 1.0 KiB |
|
Before Width: | Height: | Size: 5.6 KiB After Width: | Height: | Size: 5.6 KiB |
|
Before Width: | Height: | Size: 1.8 KiB After Width: | Height: | Size: 1.8 KiB |
|
Before Width: | Height: | Size: 3.2 KiB After Width: | Height: | Size: 3.2 KiB |
|
Before Width: | Height: | Size: 2.6 KiB After Width: | Height: | Size: 2.6 KiB |
|
Before Width: | Height: | Size: 6.0 KiB After Width: | Height: | Size: 6.0 KiB |
|
Before Width: | Height: | Size: 4.2 KiB After Width: | Height: | Size: 4.2 KiB |
@@ -1,9 +1,9 @@
|
||||
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 580 478" font-family="monospace" font-size="13">
|
||||
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 580 430" font-family="monospace" font-size="13">
|
||||
<!-- Background -->
|
||||
<rect width="580" height="478" fill="#1a1a2e" rx="4"/>
|
||||
<rect width="580" height="430" fill="#1a1a2e" rx="4"/>
|
||||
|
||||
<!-- Title -->
|
||||
<text x="310" y="26" fill="#e0e0e0" text-anchor="middle" font-size="14" font-weight="bold">LookupRequest (0x30) — 303 + 16n bytes</text>
|
||||
<text x="310" y="26" fill="#e0e0e0" text-anchor="middle" font-size="14" font-weight="bold">LookupRequest (0x30) — 46 + 16n bytes</text>
|
||||
|
||||
<!-- Row 0 (0–3): msg_type + request_id starts -->
|
||||
<text x="50" y="62" fill="#666" font-size="10" text-anchor="end">0–3</text>
|
||||
@@ -58,15 +58,6 @@
|
||||
<rect x="55" y="348" width="520" height="40" fill="#1f1f3a" stroke="#4ad9d9" stroke-width="1" stroke-dasharray="4,3" rx="3"/>
|
||||
<text x="315" y="372" fill="#8af8f8" text-anchor="middle" font-size="11">origin_coords × n (16 bytes each — NodeAddr only)</text>
|
||||
|
||||
<!-- Row 7: visited bloom -->
|
||||
<text x="50" y="414" fill="#666" font-size="10" text-anchor="end">+1+256</text>
|
||||
|
||||
<rect x="55" y="394" width="130" height="40" fill="#3d3d2d" stroke="#a0a04a" stroke-width="1.5" rx="3"/>
|
||||
<text x="120" y="418" fill="#d8d88a" text-anchor="middle" font-size="10">hash_cnt</text>
|
||||
|
||||
<rect x="185" y="394" width="390" height="40" fill="#1f1f3a" stroke="#a0a04a" stroke-width="1" stroke-dasharray="4,3" rx="3"/>
|
||||
<text x="380" y="418" fill="#d8d88a" text-anchor="middle" font-size="11">visited_bits (256 bytes)</text>
|
||||
|
||||
<!-- Total -->
|
||||
<text x="310" y="462" fill="#777" font-size="10" text-anchor="middle">total: 303 + (n × 16) bytes, where n = origin depth + 1</text>
|
||||
<text x="310" y="414" fill="#777" font-size="10" text-anchor="middle">total: 46 + (n × 16) bytes, where n = origin depth + 1</text>
|
||||
</svg>
|
||||
|
Before Width: | Height: | Size: 4.3 KiB After Width: | Height: | Size: 3.8 KiB |
|
Before Width: | Height: | Size: 3.4 KiB After Width: | Height: | Size: 3.4 KiB |
|
Before Width: | Height: | Size: 2.5 KiB After Width: | Height: | Size: 2.5 KiB |
|
Before Width: | Height: | Size: 4.3 KiB After Width: | Height: | Size: 4.3 KiB |
|
Before Width: | Height: | Size: 4.1 KiB After Width: | Height: | Size: 4.1 KiB |
|
Before Width: | Height: | Size: 2.6 KiB After Width: | Height: | Size: 2.6 KiB |
|
Before Width: | Height: | Size: 1000 B After Width: | Height: | Size: 1000 B |
|
Before Width: | Height: | Size: 6.8 KiB After Width: | Height: | Size: 6.8 KiB |
|
Before Width: | Height: | Size: 4.1 KiB After Width: | Height: | Size: 4.1 KiB |
|
Before Width: | Height: | Size: 2.9 KiB After Width: | Height: | Size: 2.9 KiB |
|
Before Width: | Height: | Size: 2.9 KiB After Width: | Height: | Size: 2.9 KiB |
|
Before Width: | Height: | Size: 2.2 KiB After Width: | Height: | Size: 2.2 KiB |
|
Before Width: | Height: | Size: 2.9 KiB After Width: | Height: | Size: 2.9 KiB |
|
Before Width: | Height: | Size: 3.5 KiB After Width: | Height: | Size: 3.5 KiB |
@@ -0,0 +1,197 @@
|
||||
# Nostr Event Reference
|
||||
|
||||
The Nostr-protocol surface FIPS uses for discovery and signaling. For
|
||||
the design of the discovery runtime and the rationale behind these
|
||||
event shapes, see
|
||||
[../design/fips-nostr-discovery.md](../design/fips-nostr-discovery.md).
|
||||
For operator activation recipes, see
|
||||
[../how-to/enable-nostr-discovery.md](../how-to/enable-nostr-discovery.md).
|
||||
|
||||
FIPS uses three Nostr event kinds:
|
||||
|
||||
| Kind | Name | Encryption | Storage | Purpose |
|
||||
| ---- | ---- | ---------- | ------- | ------- |
|
||||
| 37195 | Overlay advert | None (signed only) | Replaceable | Publish reachable transport endpoints |
|
||||
| 21059 | Traversal signaling | NIP-44 inside NIP-59 gift wrap | Ephemeral | Carry `TraversalOffer`/`TraversalAnswer` payloads |
|
||||
| 10050 | NIP-17 inbox relay list | None (signed only) | Replaceable | Tell dialers where to publish offers |
|
||||
|
||||
All three are signed with the node's FIPS identity key (the same
|
||||
secp256k1 keypair Nostr uses); there is no separate Nostr key.
|
||||
|
||||
## Kind 37195 — Overlay Advert
|
||||
|
||||
A parameterized replaceable event in the application-defined
|
||||
replaceable range `30000–39999` (the digits visually spell `FIPS`:
|
||||
7=F, 1=I, 9=P, 5=S). Each node has a single in-place-updatable advert
|
||||
under its identity.
|
||||
|
||||
### Tags
|
||||
|
||||
- `d` — fixed to the literal `fips-overlay-v1` (the application
|
||||
identifier baked into the binary). Together with `pubkey`, this
|
||||
identifies the unique replaceable event slot.
|
||||
- `protocol` — the configured `node.discovery.nostr.app` value
|
||||
(default `fips-overlay-v1`). Distinct from the `d` tag so the
|
||||
application string can evolve without breaking the replaceable
|
||||
event slot.
|
||||
- `version` — protocol version string (currently `"1"`).
|
||||
- `expiration` — NIP-40 expiration timestamp set to now +
|
||||
`node.discovery.nostr.advert_ttl_secs` (default 3600 seconds).
|
||||
Conforming relays stop serving the event after this time.
|
||||
|
||||
### Content
|
||||
|
||||
The event content is a JSON document shaped as `OverlayAdvert`:
|
||||
|
||||
```json
|
||||
{
|
||||
"identifier": "fips-overlay-v1",
|
||||
"version": 1,
|
||||
"endpoints": [
|
||||
{"transport": "udp", "addr": "203.0.113.45:2121"},
|
||||
{"transport": "tor", "addr": "xxxxx.onion:8443"},
|
||||
{"transport": "udp", "addr": "nat"}
|
||||
],
|
||||
"signalRelays": ["wss://relay.damus.io", "wss://nos.lol"],
|
||||
"stunServers": ["stun:stun.l.google.com:19302"]
|
||||
}
|
||||
```
|
||||
|
||||
Field semantics:
|
||||
|
||||
| Field | Type | Description |
|
||||
| ----- | ---- | ----------- |
|
||||
| `identifier` | string | Application namespace; must match the `d` tag. |
|
||||
| `version` | integer | Advert schema version (currently 1). |
|
||||
| `endpoints` | array | List of transport endpoints. Each is `{transport, addr}` where `transport` is `"udp"`, `"tcp"`, or `"tor"`, and `addr` is `"host:port"`, `".onion:port"`, or the literal `"nat"` (for UDP NAT-punch). |
|
||||
| `signalRelays` | array? | Optional. Relays the publisher prefers for offer/answer signaling. Present only when at least one endpoint is `udp:nat`. |
|
||||
| `stunServers` | array? | Optional. STUN servers the publisher uses for reflexive discovery. Present only when at least one endpoint is `udp:nat`. Informational — peers do not use these to choose their own STUN targets. |
|
||||
|
||||
### Signature scope
|
||||
|
||||
The Nostr event signature covers the standard Nostr event ID
|
||||
(serialized `[0, pubkey, created_at, kind, tags, content]`), so the
|
||||
content JSON, tags, kind, and timestamp are all bound to the signing
|
||||
identity.
|
||||
|
||||
### Replacement and deletion
|
||||
|
||||
Because kind 37195 is replaceable, publishing a new advert replaces
|
||||
the prior one in the same `(pubkey, d-tag)` slot. To withdraw an
|
||||
advert without publishing a successor, the node publishes a NIP-9
|
||||
kind 5 delete event referencing the prior advert.
|
||||
|
||||
## Kind 21059 — Traversal Signaling
|
||||
|
||||
An ephemeral event (kinds in the 20000–29999 range are not stored by
|
||||
conforming relays). Used to deliver gift-wrapped, NIP-44-encrypted
|
||||
`TraversalOffer` and `TraversalAnswer` payloads between dialer and
|
||||
responder during a UDP NAT hole-punch.
|
||||
|
||||
### Encryption envelope
|
||||
|
||||
The wire shape is the standard NIP-59 gift wrap:
|
||||
|
||||
1. **Rumor** — the unsigned `TraversalOffer`/`TraversalAnswer`
|
||||
payload (JSON), authored by the actual sender's identity.
|
||||
2. **Seal** — a kind 13 event whose content is the rumor
|
||||
NIP-44-encrypted to the recipient's pubkey, signed by the sender.
|
||||
3. **Gift wrap** — a kind 21059 event whose content is the seal
|
||||
NIP-44-encrypted to the recipient under an ephemeral key, signed
|
||||
by that ephemeral key. The outer `pubkey` of the kind 21059 event
|
||||
is the ephemeral identity, not the sender's real identity.
|
||||
|
||||
Only the intended recipient can decrypt the wrap to recover the seal,
|
||||
and only the recipient can decrypt the seal to recover the rumor.
|
||||
|
||||
### Wrapped payloads
|
||||
|
||||
The `TraversalOffer` carries:
|
||||
|
||||
- `type` — message-type tag.
|
||||
- `sessionId` — unique identifier correlating offer and answer.
|
||||
- `senderNpub` / `recipientNpub` — bech32-encoded pubkeys, repeated
|
||||
inside the encrypted payload (the outer wrap pubkey is ephemeral).
|
||||
- `issuedAt` / `expiresAt` — Unix-ms timestamps; `expiresAt` is
|
||||
`issuedAt + signal_ttl_secs * 1000`.
|
||||
- `nonce` — random per-offer value.
|
||||
- `reflexiveAddress` — `{protocol, ip, port}` observed via STUN, or
|
||||
`null` if STUN failed or returned no usable address.
|
||||
- `localAddresses` — array of `{protocol, ip, port}` private
|
||||
candidates, populated when `share_local_candidates` is enabled.
|
||||
- `stunServer` — the STUN server actually used (informational).
|
||||
|
||||
The `TraversalAnswer` echoes `sessionId` and carries:
|
||||
|
||||
- `type`, `senderNpub`, `recipientNpub`, `issuedAt`, `expiresAt`,
|
||||
`nonce` — same shape as the offer.
|
||||
- `inReplyTo` — the offer's event id.
|
||||
- `accepted` — boolean; false when the responder has no usable
|
||||
addresses.
|
||||
- `reflexiveAddress` and `localAddresses` — the responder's
|
||||
candidates, in the same shape as the offer.
|
||||
- `stunServer` — informational.
|
||||
- `punch` — a `PunchHint { startAtMs, intervalMs, durationMs }`
|
||||
telling both sides when to begin probing and how aggressively.
|
||||
Absent on rejected offers.
|
||||
- `reason` — optional rejection string when `accepted` is false.
|
||||
- `offerReceivedAt` — optional responder wall-clock (Unix ms) at
|
||||
the moment it received the offer; the initiator uses this to
|
||||
derive a clock-skew estimate.
|
||||
|
||||
### Relay selection
|
||||
|
||||
Dialer publishes offers to the recipient's NIP-17 inbox relays (kind
|
||||
10050) when available; otherwise to the local
|
||||
`node.discovery.nostr.dm_relays` list. The responder publishes the
|
||||
answer back through the same relay channel.
|
||||
|
||||
## Kind 10050 — NIP-17 Inbox Relay List
|
||||
|
||||
A standard NIP-17 event used by FIPS to advertise which relays this
|
||||
node prefers for receiving direct-message-style signaling — for FIPS,
|
||||
the gift-wrapped traversal offers (kind 21059).
|
||||
|
||||
This is **NIP-17** (`kind 10050`, inbox relays for DM delivery), not
|
||||
NIP-65 (`kind 10002`, general read/write relay list). The two serve
|
||||
different purposes:
|
||||
|
||||
- Kind 10002 (NIP-65) — general read/write relays for ordinary event
|
||||
publication and subscription.
|
||||
- Kind 10050 (NIP-17) — relays the recipient prefers for receiving
|
||||
DM-shaped (NIP-59 wrapped) events.
|
||||
|
||||
FIPS publishes its own kind 10050 on startup so dialers can discover
|
||||
where to send traversal offers. When dialing a peer, FIPS first
|
||||
fetches the peer's kind 10050 from the peer's `advert_relays`; on
|
||||
fetch failure it falls back to the local `dm_relays` list.
|
||||
|
||||
### Tags
|
||||
|
||||
Standard NIP-17 form: each relay is encoded as an `r` tag whose
|
||||
single value is the relay URL.
|
||||
|
||||
```text
|
||||
["r", "wss://relay.damus.io"]
|
||||
["r", "wss://nos.lol"]
|
||||
```
|
||||
|
||||
### Content
|
||||
|
||||
Empty per NIP-17.
|
||||
|
||||
## See also
|
||||
|
||||
- [../design/fips-nostr-discovery.md](../design/fips-nostr-discovery.md)
|
||||
— discovery runtime design and the five activation scenarios
|
||||
- [../how-to/enable-nostr-discovery.md](../how-to/enable-nostr-discovery.md)
|
||||
— operator recipes
|
||||
- [../tutorials/resolve-peers-via-nostr.md](../tutorials/resolve-peers-via-nostr.md),
|
||||
[../tutorials/advertise-your-node.md](../tutorials/advertise-your-node.md),
|
||||
[../tutorials/open-discovery.md](../tutorials/open-discovery.md)
|
||||
— hand-held tutorial walkthroughs of the three capabilities
|
||||
- [../design/port-advertisement-and-nat-traversal.md](../design/port-advertisement-and-nat-traversal.md)
|
||||
— generic protocol reference (event tags, NIP usage, on-the-wire
|
||||
offer/answer schema), with FIPS values as worked examples
|
||||
- [security.md](security.md) — how the FIPS identity key signs both
|
||||
adverts and Noise handshakes
|
||||
@@ -0,0 +1,235 @@
|
||||
# Security Reference
|
||||
|
||||
Consolidated security reference covering the nftables baseline, peer
|
||||
ACL file format, cryptographic primitives, rekey defaults, replay
|
||||
window, filesystem permissions, threat-resistance matrix, and default
|
||||
network exposures per transport. For the threat-model design and
|
||||
rationale, see [../design/fips-security.md](../design/fips-security.md).
|
||||
For the operator activation steps and drop-in recipes, see
|
||||
[../how-to/enable-mesh-firewall.md](../how-to/enable-mesh-firewall.md).
|
||||
|
||||
## nftables Baseline
|
||||
|
||||
The shipped baseline is `/etc/fips/fips.nft`. It defines a single
|
||||
nftables table `inet fips` with one chain hooked at `input`, structured
|
||||
as follows:
|
||||
|
||||
| Step | Rule | Effect |
|
||||
| ---- | ---- | ------ |
|
||||
| 1 | `iifname != "fips0" return` | Match only traffic arriving on `fips0`; everything else short-circuits. |
|
||||
| 2 | `ct state established,related accept` | Allow conntrack replies and related ICMPv6 errors. |
|
||||
| 3 | `icmpv6 type echo-request accept` | Allow IPv6 echo (ping6 reachability). |
|
||||
| 4 | `include "/etc/fips/fips.d/*.nft"` | Splice in operator drop-ins (empty matches nothing). |
|
||||
| 5 | `counter drop` | Default-deny everything else; counter increments on every drop. |
|
||||
|
||||
Outbound from `fips0` is unrestricted. The baseline is a documented
|
||||
dpkg conffile — operator edits to `/etc/fips/fips.nft` are preserved
|
||||
across upgrades.
|
||||
|
||||
The systemd unit is `fips-firewall.service` (oneshot). It is **not**
|
||||
enabled by default; activation is an explicit operator gesture
|
||||
documented in
|
||||
[../how-to/enable-mesh-firewall.md](../how-to/enable-mesh-firewall.md).
|
||||
|
||||
## Drop-In File Format
|
||||
|
||||
Operator extensions live under `/etc/fips/fips.d/` with the `.nft`
|
||||
suffix. Each file is included inline into the `inbound` chain at the
|
||||
marked point and may contain any nftables rule lines valid in that
|
||||
context.
|
||||
|
||||
Naming convention: `<purpose>-from-<source>.nft` keeps drop-ins easy
|
||||
to scan. Examples shipped in the design discussion:
|
||||
|
||||
- `ssh-from-bastion.nft` — accept TCP/22 from a single mesh-node address
|
||||
- `http-from-cluster.nft` — accept TCP/80 from a `/64` mesh-address prefix
|
||||
- `dns-public.nft` — accept UDP/53 and TCP/53 from any mesh node
|
||||
- `git-from-trusted.nft` — accept TCP/9418 from a set of mesh-node addresses
|
||||
|
||||
After editing, reload via
|
||||
`sudo systemctl reload-or-restart fips-firewall.service` (or
|
||||
equivalently `sudo nft -f /etc/fips/fips.nft` since the file is
|
||||
idempotent).
|
||||
|
||||
## Cryptographic Primitives
|
||||
|
||||
| Component | Choice | Where Used |
|
||||
| --------- | ------ | ---------- |
|
||||
| Curve | secp256k1 | FMP IK, FSP XK, Schnorr signatures |
|
||||
| Diffie-Hellman | ECDH on secp256k1 (x-only normalized) | Noise IK, Noise XK |
|
||||
| AEAD | ChaCha20-Poly1305 | FMP link encryption, FSP session encryption |
|
||||
| Hash | SHA-256 | NodeAddr derivation, Noise transcript |
|
||||
| Key derivation | HKDF-SHA256 | Noise key schedule |
|
||||
| Signatures | secp256k1 Schnorr | TreeAnnounce, LookupResponse proof, Nostr adverts |
|
||||
| Noise pattern (link) | `Noise_IK_secp256k1_ChaChaPoly_SHA256` | FMP link layer (IK with epoch payload) |
|
||||
| Noise pattern (session) | `Noise_XK_secp256k1_ChaChaPoly_SHA256` | FSP session layer (XK with epoch payload) |
|
||||
|
||||
These choices align with the Nostr cryptographic stack
|
||||
(secp256k1 + ChaCha20-Poly1305 + SHA-256) and the NIP-44 encrypted
|
||||
messaging standard.
|
||||
|
||||
## Rekey Defaults
|
||||
|
||||
Both link-layer and session-layer Noise sessions rekey under one of
|
||||
two triggers, configurable under `node.rekey.*`:
|
||||
|
||||
| Parameter | Default | Description |
|
||||
| --------- | ------- | ----------- |
|
||||
| `enabled` | `true` | Master switch. |
|
||||
| `after_secs` | `120` | Time-based rekey threshold. |
|
||||
| `after_messages` | `65536` | Message-count rekey threshold. |
|
||||
|
||||
In addition to the configurable triggers, the daemon retains the old
|
||||
session keys for a fixed **10-second drain window** after each
|
||||
cutover (compile-time constant `DRAIN_WINDOW_SECS` in
|
||||
`src/node/handlers/rekey.rs`). Rekey rotates the Noise key schedule
|
||||
and the session indices; old session keys are kept in
|
||||
`previous_session` for the drain window so in-flight packets
|
||||
encrypted under the old keys still decrypt.
|
||||
|
||||
## Replay Window
|
||||
|
||||
Both layers use explicit per-packet counters with a sliding bitmap
|
||||
window for replay protection. The bitmap is **2048 entries** at both
|
||||
layers — large enough to accommodate UDP reordering and packet loss
|
||||
without false-positive replay rejection. Counters older than the
|
||||
window are rejected. The same `ReplayWindow` and
|
||||
`decrypt_with_replay_check()` implementation is used at both the FMP
|
||||
and FSP layers.
|
||||
|
||||
## Peer ACL
|
||||
|
||||
Mesh-level ACL files at `/etc/fips/peers.allow` and
|
||||
`/etc/fips/peers.deny` give the operator allowlist/blocklist control
|
||||
over which npubs may complete the FMP Noise IK link handshake.
|
||||
|
||||
File format:
|
||||
|
||||
- One entry per line. An entry is either a bech32 `npub1...`,
|
||||
an alias defined in `/etc/fips/hosts`, or the literal `ALL`
|
||||
wildcard (case-insensitive).
|
||||
- Lines beginning with `#` are comments.
|
||||
- Blank lines are ignored.
|
||||
|
||||
Evaluation order (first match wins, default-allow on no match):
|
||||
|
||||
1. `peers.allow` — if the peer matches an entry here (or `ALL` is
|
||||
in `peers.allow`), the handshake is admitted, regardless of any
|
||||
`peers.deny` entry.
|
||||
2. `peers.deny` — if the peer matches an entry here (or `ALL` is
|
||||
in `peers.deny`), the handshake is refused.
|
||||
3. Otherwise the peer is admitted.
|
||||
|
||||
`peers.allow` is **not** an exclusive gate on its own: an unlisted
|
||||
peer falls through to step 3 and is admitted unless it appears in
|
||||
`peers.deny`. To turn `peers.allow` into a strict allowlist, place
|
||||
`ALL` in `peers.deny` so every unlisted peer is rejected at step 2.
|
||||
|
||||
The `ALL` wildcard makes the operator's posture explicit:
|
||||
|
||||
- `ALL` in `peers.allow` admits every peer (same effect as the
|
||||
default-allow behavior, but documented in the file).
|
||||
- `ALL` in `peers.deny` blocks every peer except those listed in
|
||||
`peers.allow` — the "allowlist-strict" posture.
|
||||
|
||||
In practice this collapses to a few common postures:
|
||||
|
||||
- **Default-allow with denylist**: leave `peers.allow` empty;
|
||||
populate `peers.deny`. All npubs may peer except those listed.
|
||||
- **Allowlist-strict**: populate `peers.allow` and put `ALL`
|
||||
in `peers.deny`. Only the listed npubs may peer; everyone else
|
||||
is rejected at step 2.
|
||||
|
||||
A populated `peers.allow` with an empty `peers.deny` is not a
|
||||
strict allowlist — it is equivalent to default-allow plus an
|
||||
explicit "always-admit" set. The strict variant requires `ALL`
|
||||
in `peers.deny`.
|
||||
|
||||
Aliases are resolved through `/etc/fips/hosts` at file-load
|
||||
time. If `peers.allow` lists `core-vm` and `/etc/fips/hosts`
|
||||
maps `core-vm` to a specific npub, that npub is admitted. If
|
||||
`core-vm` is later remapped to a different npub, the ACL
|
||||
re-resolves on the next mtime change. Operators should be aware
|
||||
that ACL semantics follow the `hosts`-file aliasing, not just
|
||||
the literal npubs visible in the file.
|
||||
|
||||
Both files are reloaded automatically when their mtime changes
|
||||
— no daemon restart or signal is needed. ACL evaluation runs
|
||||
after msg1 decryption but before any further peer-state
|
||||
mutation; rate-limited msg1s never reach the ACL.
|
||||
|
||||
## Filesystem Permissions
|
||||
|
||||
| Path | Owner | Mode | Purpose |
|
||||
| ---- | ----- | ---- | ------- |
|
||||
| `/etc/fips/fips.key` | root:root | `0600` | Persistent identity private key (sensitive). |
|
||||
| `/etc/fips/fips.pub` | root:root | `0644` | Public key (npub). |
|
||||
| `/etc/fips/fips.yaml` | root:root | `0644` | Daemon configuration (dpkg conffile). |
|
||||
| `/etc/fips/fips.nft` | root:root | `0644` | nftables baseline (dpkg conffile). |
|
||||
| `/etc/fips/fips.d/` | root:root | `0755` | Operator drop-in directory. |
|
||||
| `/etc/fips/hosts` | root:root | `0644` | Optional hostname → npub map (dpkg conffile). |
|
||||
| `/etc/fips/peers.allow` | root:root | `0644` | Optional peer allowlist. |
|
||||
| `/etc/fips/peers.deny` | root:root | `0644` | Optional peer denylist. |
|
||||
| `/run/fips/control.sock` | root:fips | `0770` | Control socket (members of `fips` group can use `fipsctl`). |
|
||||
| `/run/fips/` | root:fips | `0750` | Control socket parent directory. |
|
||||
|
||||
Adding a user to the `fips` group grants `fipsctl` access without
|
||||
requiring root. The daemon `chown`s the control socket and its parent
|
||||
directory at bind time.
|
||||
|
||||
## Threat-Resistance Matrix
|
||||
|
||||
The link layer's threat-resistance matrix is consolidated here from
|
||||
the FMP design document:
|
||||
|
||||
| Threat | Mitigation |
|
||||
| ------ | ---------- |
|
||||
| Connection exhaustion | Token-bucket rate limit + connection count limit |
|
||||
| CPU exhaustion (msg1 flood) | Rate limit before crypto operations |
|
||||
| Replay attacks | Counter-based nonces with sliding window (2048 entries) |
|
||||
| State confusion | Strict handshake state machine validation |
|
||||
| Spoofed encrypted packets | Index lookup + AEAD verification |
|
||||
| Spoofed msg2 | Index lookup + Noise ephemeral key binding |
|
||||
| Address spoofing | Cryptographic authority, not address-based |
|
||||
| Session correlation | Index rotation on rekey |
|
||||
| Inbound exposure on `fips0` | Default-deny nftables baseline (operator opt-in) |
|
||||
| Sybil identities | Discretionary peering + handshake rate limiting + optional peer ACL |
|
||||
| Eclipse attack | Diverse peering across independent operators and transports |
|
||||
| Unauthorized peer admission | Optional `peers.allow` allowlist consulted before handshake |
|
||||
|
||||
See [../design/fips-mesh-layer.md](../design/fips-mesh-layer.md) for
|
||||
the unauthenticated-attack-surface analysis (only handshake msg1 is
|
||||
reachable by unauthenticated parties), and
|
||||
[../design/fips-mesh-operation.md](../design/fips-mesh-operation.md#privacy-considerations)
|
||||
for the metadata-privacy model and the rejection of onion routing.
|
||||
|
||||
## Default Network Exposures by Transport
|
||||
|
||||
| Transport | Default Inbound | Default Bind | Opt-in |
|
||||
| --------- | --------------- | ------------ | ------ |
|
||||
| UDP | None until `bind_addr` set | `0.0.0.0:2121` typical | Operator sets `transports.udp.bind_addr` |
|
||||
| TCP | None until `bind_addr` set | None — outbound-only without bind | Operator sets `transports.tcp.bind_addr` |
|
||||
| Ethernet | Listens on configured interface (raw `AF_PACKET`) | EtherType 0x2121 on selected interface | Per-flag `discovery`, `announce`, `auto_connect`, `accept_connections` |
|
||||
| Tor | None until `directory_service` configured | `127.0.0.1:8443` (loopback only) | Operator sets `transports.tor.directory_service` and configures `HiddenServiceDir` in `torrc` |
|
||||
| BLE | Off by default | n/a | Operator enables `transports.ble.*` |
|
||||
| Nostr discovery | Off by default | n/a (relay client, not a listener) | Operator sets `node.discovery.nostr.enabled: true` |
|
||||
|
||||
The mesh-layer `fips0` interface is reachable from any mesh node that
|
||||
can route to you, not only direct peers — your direct peers forward
|
||||
traffic from any reachable mesh node onto your `fips0`. The
|
||||
default-deny nftables baseline (operator opt-in) is the recommended
|
||||
way to restrict inbound traffic on `fips0`. See
|
||||
[../how-to/enable-mesh-firewall.md](../how-to/enable-mesh-firewall.md).
|
||||
|
||||
## See also
|
||||
|
||||
- [../design/fips-security.md](../design/fips-security.md) — threat
|
||||
model and design rationale for the `fips0` baseline
|
||||
- [../design/fips-mesh-layer.md](../design/fips-mesh-layer.md) — FMP
|
||||
link encryption, replay protection, rate limiting
|
||||
- [../design/fips-session-layer.md](../design/fips-session-layer.md)
|
||||
— FSP end-to-end encryption, Noise XK, replay window
|
||||
- [../how-to/enable-mesh-firewall.md](../how-to/enable-mesh-firewall.md)
|
||||
— operator activation and drop-in recipes
|
||||
- [configuration.md](configuration.md) — full `node.rekey.*`,
|
||||
`node.rate_limit.*` parameter tables
|
||||
@@ -0,0 +1,92 @@
|
||||
# Transport Statistics Reference
|
||||
|
||||
Per-transport statistics counter inventories. Counters are exposed
|
||||
through the daemon control socket (`fipsctl show transports`) and the
|
||||
`fipstop` operator UI. For the transport-layer design (services
|
||||
provided to FMP, transport categories, the trait surface, connection
|
||||
model), see
|
||||
[../design/fips-transport-layer.md](../design/fips-transport-layer.md).
|
||||
|
||||
All transports report counters via `fipsctl show transports`; the
|
||||
tables below are source-extracted from each transport's `stats.rs`
|
||||
module.
|
||||
|
||||
## UDP
|
||||
|
||||
| Counter | Description |
|
||||
| ------- | ----------- |
|
||||
| `packets_sent` / `bytes_sent` | Successful sends |
|
||||
| `packets_recv` / `bytes_recv` | Successful receives |
|
||||
| `send_errors` / `recv_errors` | Send/receive failures |
|
||||
| `mtu_exceeded` | Packets rejected for MTU violation |
|
||||
| `kernel_drops` | Kernel `SO_RXQ_OVFL` drop count (feeds ECN congestion detection) |
|
||||
|
||||
## TCP
|
||||
|
||||
| Counter | Description |
|
||||
| ------- | ----------- |
|
||||
| `packets_sent` / `bytes_sent` | Successful sends |
|
||||
| `packets_recv` / `bytes_recv` | Successful receives |
|
||||
| `send_errors` / `recv_errors` | Send/receive failures |
|
||||
| `mtu_exceeded` | Packets rejected for MTU violation |
|
||||
| `connections_established` | Successful outbound connections |
|
||||
| `connections_accepted` | Accepted inbound connections |
|
||||
| `connections_rejected` | Rejected inbound connections (limit exceeded) |
|
||||
| `connect_timeouts` | Connection timeout count |
|
||||
| `connect_refused` | Connection refused count |
|
||||
|
||||
## Ethernet
|
||||
|
||||
| Counter | Description |
|
||||
| ------- | ----------- |
|
||||
| `frames_sent` / `frames_recv` | Successful frame send/receive |
|
||||
| `bytes_sent` / `bytes_recv` | Byte counters |
|
||||
| `send_errors` / `recv_errors` | Send/receive failures |
|
||||
| `beacons_sent` / `beacons_recv` | Peer-discovery beacon traffic |
|
||||
| `frames_too_short` | Frames below minimum length, dropped |
|
||||
| `frames_too_long` | Frames above transport MTU, dropped |
|
||||
|
||||
## Tor
|
||||
|
||||
| Counter | Description |
|
||||
| ------- | ----------- |
|
||||
| `packets_sent` / `bytes_sent` | Successful sends |
|
||||
| `packets_recv` / `bytes_recv` | Successful receives |
|
||||
| `send_errors` / `recv_errors` | Send/receive failures |
|
||||
| `connections_established` | Successful SOCKS5 connections |
|
||||
| `connect_timeouts` | Connection timeout count |
|
||||
| `connect_refused` | Connection refused count |
|
||||
| `socks5_errors` | SOCKS5 protocol errors |
|
||||
| `mtu_exceeded` | Packets rejected for MTU violation |
|
||||
| `connections_accepted` | Accepted inbound connections via onion service |
|
||||
| `connections_rejected` | Rejected inbound connections (limit exceeded) |
|
||||
| `control_errors` | Tor control port errors |
|
||||
|
||||
## Bluetooth
|
||||
|
||||
| Counter | Description |
|
||||
| ------- | ----------- |
|
||||
| `packets_sent` / `bytes_sent` | Successful L2CAP CoC sends |
|
||||
| `packets_recv` / `bytes_recv` | Successful L2CAP CoC receives |
|
||||
| `send_errors` / `recv_errors` | Send/receive failures |
|
||||
| `mtu_exceeded` | Packets rejected for MTU violation |
|
||||
| `connections_established` | Successful outbound L2CAP connections |
|
||||
| `connections_accepted` | Accepted inbound L2CAP connections |
|
||||
| `connections_rejected` | Rejected inbound (limit exceeded) |
|
||||
| `connect_timeouts` | Connection timeout count |
|
||||
| `pool_evictions` | Connection-pool entries evicted |
|
||||
| `advertisements_sent` | BLE advertisements emitted |
|
||||
| `scan_results` | BLE scan results observed |
|
||||
|
||||
## See also
|
||||
|
||||
- [../design/fips-transport-layer.md](../design/fips-transport-layer.md)
|
||||
— transport-layer design, trait surface, per-transport sections
|
||||
- [configuration.md](configuration.md) — `transports.*` configuration
|
||||
blocks
|
||||
- [../how-to/tune-udp-buffers.md](../how-to/tune-udp-buffers.md) —
|
||||
host-side `net.core.rmem_max` / `net.core.wmem_max` setup for UDP
|
||||
- [../how-to/deploy-tor-onion.md](../how-to/deploy-tor-onion.md) —
|
||||
Tor `directory` mode operator setup
|
||||
- [../how-to/set-up-bluetooth-peer.md](../how-to/set-up-bluetooth-peer.md)
|
||||
— Linux BLE peer config
|
||||
@@ -5,6 +5,45 @@ protocol layers. It covers transport framing, link-layer message formats,
|
||||
and session-layer message formats, with an encapsulation walkthrough showing
|
||||
how application data is wrapped through each layer.
|
||||
|
||||
## FMP Message Type Catalog
|
||||
|
||||
The FMP link layer defines the following message types, dispatched by the
|
||||
`msg_type` byte in the encrypted inner header:
|
||||
|
||||
| Type | Name | Forwarding |
|
||||
| ---- | ---- | ---------- |
|
||||
| 0x00 | SessionDatagram | Routed hop-by-hop toward the destination |
|
||||
| 0x01 | SenderReport | Peer-to-peer (MMP, link-layer instance) |
|
||||
| 0x02 | ReceiverReport | Peer-to-peer (MMP, link-layer instance) |
|
||||
| 0x10 | TreeAnnounce | Peer-to-peer (spanning-tree gossip) |
|
||||
| 0x20 | FilterAnnounce | Peer-to-peer (bloom-filter gossip) |
|
||||
| 0x30 | LookupRequest | Forwarded — bloom-guided through tree peers |
|
||||
| 0x31 | LookupResponse | Forwarded — reverse-path via `recent_requests` |
|
||||
| 0x50 | Disconnect | Peer-to-peer (orderly link teardown) |
|
||||
| 0x51 | Heartbeat | Peer-to-peer (link liveness) |
|
||||
|
||||
Handshake messages travel before encryption is established and are identified
|
||||
by the FMP common-prefix `phase` field rather than a `msg_type` byte
|
||||
(phase 0x1 = Noise IK msg1, phase 0x2 = Noise IK msg2).
|
||||
|
||||
## Packet Type Summary
|
||||
|
||||
A higher-level summary that includes typical sizes and forwarding category:
|
||||
|
||||
| Message | Typical Size | When | Forwarded? |
|
||||
| ------- | ------------ | ---- | ---------- |
|
||||
| TreeAnnounce | Variable (depth-dependent) | Topology changes | No (peer-to-peer) |
|
||||
| FilterAnnounce | ~1 KB | Topology changes | No (peer-to-peer) |
|
||||
| LookupRequest | ~300 bytes | First contact, recovery | Yes (bloom-guided tree) |
|
||||
| LookupResponse | ~400 bytes | Response to discovery | Yes (reverse-path) |
|
||||
| SessionDatagram + SessionSetup | ~232–402 bytes | Session establishment | Yes (routed) |
|
||||
| SessionDatagram + SessionAck | ~170 bytes | Session confirmation | Yes (routed) |
|
||||
| SessionDatagram + Data (minimal) | 77 bytes + IPv6 payload | Bulk IPv6 traffic (compressed) | Yes (routed) |
|
||||
| SessionDatagram + Data (with CP) | 77 + coords + IPv6 payload | Warmup/recovery (compressed) | Yes (routed) |
|
||||
| SessionDatagram + CoordsRequired | 70 bytes | Cache miss error | Yes (routed) |
|
||||
| SessionDatagram + PathBroken | 70+ bytes | Dead-end error | Yes (routed) |
|
||||
| Disconnect | 2 bytes | Link teardown | No (peer-to-peer) |
|
||||
|
||||
## Encoding Rules
|
||||
|
||||
- All multi-byte integers are **little-endian** (LE)
|
||||
@@ -19,9 +58,10 @@ how application data is wrapped through each layer.
|
||||
|
||||
Datagram-oriented transports (UDP, raw Ethernet, radio) preserve natural
|
||||
packet boundaries and require no additional framing. Stream-oriented
|
||||
transports (TCP, WebSocket, Tor) must delineate FIPS packets within the
|
||||
byte stream; the common prefix `payload_len` field provides this
|
||||
framing directly.
|
||||
transports (TCP, Tor) must delineate FIPS packets within the byte
|
||||
stream; the common prefix `payload_len` field provides this framing
|
||||
directly. TCP and Tor share a common stream reader (`tcp/stream.rs`)
|
||||
that implements this framing.
|
||||
|
||||
**Ethernet data frame header.** The Ethernet transport prepends a 3-byte
|
||||
header before the FMP payload on data frames: a 1-byte frame type
|
||||
@@ -254,7 +294,11 @@ With link overhead: 1,072 bytes.
|
||||
|
||||
### LookupRequest (0x30)
|
||||
|
||||
Coordinate discovery request, flooded through the mesh.
|
||||
Coordinate discovery request, routed through the spanning tree via
|
||||
bloom-filter-guided forwarding. Each transit node forwards only to tree
|
||||
peers (parent + children) whose bloom filter contains the target.
|
||||
Request_id dedup in recent_requests handles edge cases from tree
|
||||
restructuring.
|
||||
|
||||

|
||||
|
||||
@@ -268,20 +312,19 @@ Coordinate discovery request, flooded through the mesh.
|
||||
| 42 | min_mtu | 2 bytes LE | Minimum transport MTU the origin requires (0 = no requirement) |
|
||||
| 44 | origin_coords_cnt | 2 bytes LE | Number of coordinate entries |
|
||||
| 46 | origin_coords | 16 x n bytes | Requester's ancestry (NodeAddr only) |
|
||||
| 46 + 16n | visited_hash_cnt | 1 byte | Hash count for visited filter |
|
||||
| 47 + 16n | visited_bits | 256 bytes | Compact bloom of visited nodes |
|
||||
|
||||
**Size**: `303 + (n x 16)` bytes, where n = origin depth + 1
|
||||
**Size**: `46 + (n x 16)` bytes, where n = origin depth + 1
|
||||
|
||||
| Origin Depth | Payload |
|
||||
| ------------ | ------- |
|
||||
| 3 | 351 bytes |
|
||||
| 5 | 383 bytes |
|
||||
| 10 | 463 bytes |
|
||||
| 3 | 110 bytes |
|
||||
| 5 | 142 bytes |
|
||||
| 10 | 222 bytes |
|
||||
|
||||
### LookupResponse (0x31)
|
||||
|
||||
Coordinate discovery response, greedy-routed back to requester.
|
||||
Coordinate discovery response, reverse-path routed back to the
|
||||
requester via the transit nodes that forwarded the request.
|
||||
|
||||

|
||||
|
||||
@@ -501,10 +544,21 @@ Message types 0x10-0x14 are carried inside the AEAD ciphertext (dispatched
|
||||
by the `msg_type` field in the encrypted inner header). Types 0x20-0x22 are
|
||||
plaintext error signals (U flag set, no encryption).
|
||||
|
||||
Session-layer SenderReport (0x11) and ReceiverReport (0x12) use the same
|
||||
body format as their link-layer counterparts (0x01 and 0x02). The msg_type
|
||||
byte in the body matches the link-layer value; dispatch to the correct layer
|
||||
happens at the session level based on the FSP message type.
|
||||
Session-layer SenderReport (0x11) and ReceiverReport (0x12) carry the same
|
||||
metric fields as their link-layer counterparts (0x01 and 0x02), but the
|
||||
body framing differs because the FSP encrypted inner header already
|
||||
carries the message-type byte. The session body therefore omits the
|
||||
msg_type byte and uses 2 reserved bytes (not 3) before the fields:
|
||||
|
||||
| Layer | Wire size | Header inside body |
|
||||
| ----- | --------- | ------------------ |
|
||||
| Link SenderReport (0x01) | 48 bytes | msg_type(1) + reserved(3) + fields(44) |
|
||||
| Session SenderReport (0x11) | 46 bytes | reserved(2) + fields(44) |
|
||||
| Link ReceiverReport (0x02) | 68 bytes | msg_type(1) + reserved(3) + fields(64) |
|
||||
| Session ReceiverReport (0x12) | 66 bytes | reserved(2) + fields(64) |
|
||||
|
||||
Dispatch happens at the session level via the `msg_type` byte in the FSP
|
||||
encrypted inner header.
|
||||
|
||||
### SessionSetup (phase 0x1)
|
||||
|
||||
@@ -840,7 +894,7 @@ endpoint session keys).
|
||||
| ------- | ---- | ----- |
|
||||
| TreeAnnounce | 100 + 32n bytes | n = depth + 1 |
|
||||
| FilterAnnounce | 1,035 bytes | v1 (1KB filter) |
|
||||
| LookupRequest | 303 + 16n bytes | n = origin depth + 1 |
|
||||
| LookupRequest | 46 + 16n bytes | n = origin depth + 1 |
|
||||
| LookupResponse | 93 + 16n bytes | n = target depth + 1 |
|
||||
| SessionDatagram | 36 + payload bytes | Fixed 36-byte header |
|
||||
| Disconnect | 2 bytes | |
|
||||
@@ -875,8 +929,10 @@ endpoint session keys).
|
||||
|
||||
## References
|
||||
|
||||
- [fips-mesh-layer.md](fips-mesh-layer.md) — FMP behavioral specification
|
||||
- [fips-session-layer.md](fips-session-layer.md) — FSP behavioral specification
|
||||
- [fips-transport-layer.md](fips-transport-layer.md) — Transport framing
|
||||
- [fips-mesh-operation.md](fips-mesh-operation.md) — How messages work together
|
||||
- [fips-ipv6-adapter.md](fips-ipv6-adapter.md) — MTU enforcement
|
||||
- [../design/fips-mesh-layer.md](../design/fips-mesh-layer.md) — FMP behavioral specification
|
||||
- [../design/fips-session-layer.md](../design/fips-session-layer.md) — FSP behavioral specification
|
||||
- [../design/fips-transport-layer.md](../design/fips-transport-layer.md) — Transport framing
|
||||
- [../design/fips-mesh-operation.md](../design/fips-mesh-operation.md) — How messages work together
|
||||
- [../design/fips-ipv6-adapter.md](../design/fips-ipv6-adapter.md) — MTU enforcement
|
||||
- [../design/fips-bloom-filters.md](../design/fips-bloom-filters.md) — FilterAnnounce parameters and FPR analysis
|
||||
- [../design/fips-mtu.md](../design/fips-mtu.md) — How `path_mtu` and MtuExceeded fit together
|
||||
@@ -0,0 +1,141 @@
|
||||
# FIPS v0.2.1
|
||||
|
||||
**Released**: 2026-05-11
|
||||
|
||||
v0.2.1 is a maintenance release on the v0.2.x line. No new features
|
||||
and no wire-format changes; operators running v0.2.0 can upgrade in
|
||||
place. The release rolls up bug fixes and operational hardening for
|
||||
issues surfaced in v0.2.0 deployments, plus a bloom-filter fill-ratio
|
||||
validation that protects mesh-size estimates from saturated-filter
|
||||
inputs.
|
||||
|
||||
## At a glance
|
||||
|
||||
- 22 commits since v0.2.0, 5 committers plus 2 issue reporters.
|
||||
- All changes are backwards-compatible with v0.2.0 on the wire.
|
||||
- Bloom filter fill-ratio validation hardens the FilterAnnounce
|
||||
ingress path.
|
||||
- TreeAnnounce ancestry validation tightened to match the
|
||||
spanning-tree specification.
|
||||
- Signed-tarball + `.deb` artifact workflow added for tagged
|
||||
releases; AUR auto-publish on stable tags.
|
||||
|
||||
## Behavior changes worth flagging
|
||||
|
||||
- **Bloom filter fill-ratio validation** runs on every inbound
|
||||
`FilterAnnounce`. Filters whose derived false-positive rate exceeds
|
||||
`node.bloom.max_inbound_fpr` (new config field, default `0.05`) are
|
||||
rejected silently on the wire, logged at WARN, and counted in a
|
||||
new `bloom.fill_exceeded` counter. A rate-limited WARN also fires
|
||||
when the local outgoing filter exceeds the cap.
|
||||
`BloomFilter::estimated_count` now takes `max_fpr` and returns
|
||||
`Option<f64>`, returning `None` for saturated filters; this
|
||||
propagates through `compute_mesh_size` into `estimated_mesh_size`.
|
||||
- **TreeAnnounce ancestry validation** is now run before tree-state
|
||||
mutation, enforcing ancestry-self-match, root-single-entry,
|
||||
parent-second-entry, and root-is-minimum-NodeAddr. Non-conforming
|
||||
announces are rejected with a WARN. Mixed v0.2.0 / v0.2.1 meshes
|
||||
may produce WARN log lines on the v0.2.1 side until all peers
|
||||
upgrade; behavior is correct, log noise only.
|
||||
|
||||
## Notable bug fixes
|
||||
|
||||
- **Control socket path detection** in `fipsctl` and `fipstop` now
|
||||
checks for the `/run/fips/` directory instead of the socket file
|
||||
inside it. Users not yet in the `fips` group get a clear
|
||||
"Permission denied" error instead of a misleading "No such file"
|
||||
fallback to `$XDG_RUNTIME_DIR`
|
||||
([#30](https://github.com/jmcorgan/fips/issues/30), reported by
|
||||
[@Sebastix](https://github.com/Sebastix)).
|
||||
- **`fd00::/8` routing protected from Tailscale interception.** The
|
||||
daemon installs an IPv6 routing-policy rule
|
||||
(`ip -6 rule to fd00::/8 lookup main priority 5265`) at TUN setup,
|
||||
so Tailscale's table 52 default route can no longer divert mesh
|
||||
traffic.
|
||||
- **Bloom filter routing greedy-tree fallback.** `find_next_hop` no
|
||||
longer returns `NoRoute` when the bloom candidate set is non-empty
|
||||
but no candidate is strictly closer than the current node; it
|
||||
falls through to greedy tree routing instead. Previously, this
|
||||
caused dropped packets in topologies where the tree parent was
|
||||
closer but not a bloom candidate.
|
||||
- **Auto-connect peers reconnect after a graceful Disconnect.**
|
||||
Previously, a clean upstream shutdown left the auto-connect peer
|
||||
orphaned; only the link-dead, decrypt-fail, and peer-restart paths
|
||||
scheduled a reconnect
|
||||
([#60](https://github.com/jmcorgan/fips/issues/60), reported by
|
||||
[@SwapMarket](https://github.com/SwapMarket)).
|
||||
- **`fipsctl connect` rejects FIPS mesh addresses** (`fd00::/8`) for
|
||||
`udp`, `tcp`, and `ethernet` transports with a clear error message
|
||||
instead of echoing success while the daemon silently failed the
|
||||
bind with `EAFNOSUPPORT`
|
||||
([#61](https://github.com/jmcorgan/fips/issues/61), reported by
|
||||
[@SwapMarket](https://github.com/SwapMarket)).
|
||||
- **OpenWrt ipk** cross-compiles cleanly again after excluding the
|
||||
BLE feature that requires D-Bus, which is unavailable on OpenWrt
|
||||
targets.
|
||||
|
||||
## Packaging
|
||||
|
||||
- **Linux release artifact workflow** builds x86_64 and aarch64
|
||||
tarballs and `.deb` packages on `v*` tag push, with SHA-256
|
||||
checksums, and publishes them to the GitHub release page.
|
||||
- **AUR publish workflow** auto-publishes the `fips` PKGBUILD on
|
||||
stable `v*` tags.
|
||||
|
||||
## Upgrade notes
|
||||
|
||||
Operator-actionable items when moving from v0.2.0 to v0.2.1:
|
||||
|
||||
- **Bloom filter fill-ratio cap (default 0.05).** Inbound
|
||||
`FilterAnnounce` messages whose derived FPR exceeds the cap are
|
||||
rejected silently on the wire. Operators with unusually saturated
|
||||
filters in the field may want to confirm that the default applies
|
||||
cleanly to their deployment; check the new `bloom.fill_exceeded`
|
||||
counter if mesh-size estimates drift after upgrade.
|
||||
- **TreeAnnounce ancestry tightening.** Mixed v0.2.0 / v0.2.1 meshes
|
||||
may produce WARN log lines on the v0.2.1 side until all peers
|
||||
upgrade. Behavior is correct, log noise only.
|
||||
|
||||
## Getting v0.2.1
|
||||
|
||||
- **Linux x86_64 / aarch64**: `.deb` and tarball at the
|
||||
[v0.2.1 release page](https://github.com/jmcorgan/fips/releases/tag/v0.2.1).
|
||||
- **Arch Linux**: `fips` from the AUR.
|
||||
- **OpenWrt**: `.ipk` at the v0.2.1 release page.
|
||||
- **From source**: `cargo build --release` from a checkout of the
|
||||
v0.2.1 tag.
|
||||
|
||||
The full per-commit changelog lives in
|
||||
[`CHANGELOG.md`](../../CHANGELOG.md). Issues and discussion at
|
||||
[github.com/jmcorgan/fips](https://github.com/jmcorgan/fips).
|
||||
|
||||
## Contributors
|
||||
|
||||
Thanks to everyone who contributed code or bug reports to this
|
||||
release.
|
||||
|
||||
**Code and packaging**:
|
||||
|
||||
- [@jcorgan](https://github.com/jmcorgan): release shepherd, bloom
|
||||
fill-ratio validation, auto-connect reconnect fix, `fipsctl`
|
||||
mesh-address rejection, control-socket path detection,
|
||||
Tailscale-vs-`fd00::/8` routing policy, bloom routing greedy
|
||||
fallback, rustfmt baseline.
|
||||
- [@Origami74](https://github.com/Origami74): OpenWrt ipk
|
||||
BLE-feature build fix.
|
||||
- [@jodobear](https://github.com/jodobear): Linux release-artifact
|
||||
workflow and target-aware build scripts.
|
||||
- [@dskvr](https://github.com/dskvr): AUR publish workflow.
|
||||
- [@SatsAndSports](https://github.com/SatsAndSports): TreeAnnounce
|
||||
semantic validation.
|
||||
|
||||
**Issue reports that drove fixes in this release**:
|
||||
|
||||
- [@Sebastix](https://github.com/Sebastix): `fipsctl` / `fipstop`
|
||||
control-socket path detection
|
||||
([#30](https://github.com/jmcorgan/fips/issues/30)).
|
||||
- [@SwapMarket](https://github.com/SwapMarket): auto-connect
|
||||
reconnect after graceful disconnect
|
||||
([#60](https://github.com/jmcorgan/fips/issues/60)) and
|
||||
`fipsctl` mesh-address rejection
|
||||
([#61](https://github.com/jmcorgan/fips/issues/61)).
|
||||
@@ -0,0 +1,764 @@
|
||||
# FIPS v0.3.0
|
||||
|
||||
**Released**: 2026-05-11
|
||||
|
||||
v0.3.0 is the testing-and-polishing release on the v0.2.x wire format.
|
||||
It widens the platform reach of FIPS from Linux-only to Linux, macOS,
|
||||
Windows, and OpenWrt; adds two large new mesh capabilities (Nostr-mediated
|
||||
peer discovery with UDP NAT traversal, and the `fips-gateway` LAN bridge);
|
||||
ships a default-deny security baseline for the mesh interface; introduces
|
||||
mesh-peer access control; substantially speeds up session-layer crypto and
|
||||
the Linux receive path; and tightens packaging across every supported
|
||||
distribution channel.
|
||||
|
||||
v0.3.0 is wire-compatible with v0.2.x. Mixed meshes interoperate; there
|
||||
is no flag-day upgrade.
|
||||
|
||||
v0.3.0 also rolls forward all changes from the v0.2.1 maintenance
|
||||
release. The sections below cover the cumulative v0.2.0 → v0.3.0
|
||||
delta; the per-section intros call out which entries first shipped
|
||||
in v0.2.1.
|
||||
|
||||
## At a glance
|
||||
|
||||
- 123 commits since v0.2.0 (109 non-merge), spanning 307 files with
|
||||
+44,186 / -4,078 lines.
|
||||
- 10 committers plus 3 issue reporters across feature work, fixes,
|
||||
packaging, and reviews.
|
||||
- 5 new GitHub Actions CI workflows (Linux Package, macOS Package,
|
||||
Windows Package, OpenWrt Package, AUR Publish) plus expanded
|
||||
integration matrices (gateway, NAT-cone, NAT-symmetric, NAT-LAN,
|
||||
rekey-accept-off, `.deb` install across Debian 12/13 + Ubuntu
|
||||
22/24/26, multi-backend `.fips` DNS resolver across the same five
|
||||
distros).
|
||||
- The long-standing systemd-resolved DNS-responder silent-drop is
|
||||
closed end-to-end.
|
||||
- Pre-1.0 control-socket JSON schema change for two query fields;
|
||||
see [Upgrade notes](#upgrade-notes).
|
||||
|
||||
## What's new
|
||||
|
||||
### Mesh discovery and NAT traversal
|
||||
|
||||
Previously, two FIPS nodes could only become peers if they had a way
|
||||
to find each other beforehand: a configured address, a shared LAN
|
||||
segment, or a Bluetooth radio range. v0.3.0 introduces a Nostr-based
|
||||
overlay-discovery channel that lets nodes find each other through any
|
||||
public Nostr relay set, plus a STUN-assisted UDP hole-punching path
|
||||
that connects peers across most consumer NATs.
|
||||
|
||||
Each participating node publishes a signed overlay advert as a Nostr
|
||||
**Kind 37195** parameterized replaceable event. (The kind sits in the
|
||||
application-defined replaceable range and the digits visually spell
|
||||
*FIPS*: 7=F, 1=I, 9=P, 5=S.) The advert lists reachable transport
|
||||
endpoints (UDP, TCP, Tor) and is consumed by other nodes to populate
|
||||
fallback addresses for `via_nostr` peers. Under `policy: open`, the
|
||||
advert cache is also dialed for non-configured peers within a budget
|
||||
cap.
|
||||
|
||||
When both peers are behind NAT, the daemon coordinates a UDP hole
|
||||
punch using NIP-59 gift-wrap signaling for the offer/answer exchange
|
||||
and STUN for reflexive address discovery. A candidate-pair punch
|
||||
planner attempts LAN-private and reflexive paths in parallel; on
|
||||
success the live socket is handed into the standard FIPS UDP transport
|
||||
via a bootstrap-handoff API.
|
||||
|
||||
Operators turn this on with `node.discovery.nostr.enabled: true` and
|
||||
the configured relay set. `policy: open` adds best-effort dialing of
|
||||
non-configured peers seen on the relays. New `peers[].via_nostr` and
|
||||
per-transport `advertise_on_nostr` / `public` flags control what each
|
||||
endpoint contributes to the published advert. Cross-field validation
|
||||
runs at startup to catch mis-configured combinations early.
|
||||
|
||||
A Docker NAT lab covering cone, symmetric, and LAN scenarios is wired
|
||||
into the integration CI matrix. A daemon-side failure-suppression
|
||||
layer (per-npub cooldown after consecutive failures, ±60s clock-skew
|
||||
tolerance, rate-limited WARN logs) keeps relay traffic well-mannered
|
||||
when peers come and go from the open discovery cache. A separate
|
||||
structural cooldown (`protocol_mismatch_cooldown_secs`, default 24h)
|
||||
suppresses retraversal when a punched peer turns out to be running an
|
||||
FMP version this daemon cannot handshake with: the punch completes at
|
||||
the UDP layer, the rx loop spots the version-mismatched packet,
|
||||
reverse-maps to the originating npub, and removes the peer from the
|
||||
next sweep until either side upgrades.
|
||||
|
||||
The auto-connect retry loop pins itself to relay ground truth. Each
|
||||
retry attempt refetches the cached overlay advert against the
|
||||
configured `advert_relays` (one filter query, 2s timeout) before
|
||||
dialing, so a peer whose NAT rebound to a fresh endpoint is recovered
|
||||
on the next retry rather than looping on a stale cached address.
|
||||
`NoTransportForType` triggers a fire-and-forget re-fetch that either
|
||||
replaces or evicts the cache entry. A startup peer-init failure (no
|
||||
operational transport, all addresses unreachable) now schedules a
|
||||
retry instead of leaving the peer in a dead state until the daemon is
|
||||
restarted. Adopted NAT-traversed UDP transports inherit the operator's
|
||||
primary `[transports.udp]` listener config (MTU, recv/send buffer
|
||||
sizes) instead of falling back to the 1280 IPv6-minimum default.
|
||||
|
||||
### Cross-platform reach
|
||||
|
||||
FIPS now ships first-class binaries for **Linux, macOS, Windows, and
|
||||
OpenWrt**.
|
||||
|
||||
- **macOS** support uses the native `utun` TUN interface, raw
|
||||
Ethernet via BPF, a `.pkg` installer with a launchd plist and
|
||||
uninstall script, and an x86_64 cross-compile from arm64 build
|
||||
hosts. A new CI matrix entry runs build and unit-test jobs on
|
||||
macOS hosts.
|
||||
- **Windows** support uses [wintun](https://www.wintun.net/) for the
|
||||
TUN device, a TCP control socket on `localhost:21210` (replacing
|
||||
the Unix domain socket Linux and macOS use), Windows Service
|
||||
lifecycle (`fips.exe --install-service`, `--uninstall-service`,
|
||||
`--service`), and a ZIP package with PowerShell install/uninstall
|
||||
scripts.
|
||||
- **MIPS** atomic-ABI portability lets the daemon build for 32-bit
|
||||
MIPS targets (`mips`, `mipsel`, MIPS32r2) by routing through
|
||||
`portable_atomic`. This unblocks OpenWrt deployments on
|
||||
consumer-grade MIPS routers.
|
||||
- **OpenWrt** packaging gets a procd init with dnsmasq forwarding,
|
||||
proxy NDP, RA route advertisements, and IPv6 forwarding sysctls.
|
||||
The `fips-gateway` is enabled by default in the OpenWrt build.
|
||||
|
||||
### FIPS gateway
|
||||
|
||||
The new `fips-gateway` binary lets unmodified LAN hosts reach FIPS
|
||||
mesh destinations without running the FIPS daemon themselves. Two
|
||||
flows ship together:
|
||||
|
||||
- **Outbound (LAN -> mesh)**: a virtual-IP pool (default
|
||||
`fd01::/112`) is allocated on demand from `.fips`-name DNS lookups.
|
||||
A state-machine lifecycle, conntrack-backed session tracking, proxy
|
||||
NDP on the LAN interface, and TTL-based reclamation handle the
|
||||
bookkeeping. A LAN host that resolves `peer.fips` gets a virtual
|
||||
address it can reach over IP, and the gateway translates the flow
|
||||
to the mesh.
|
||||
- **Inbound (mesh -> LAN)**: new `gateway.port_forwards` config
|
||||
installs prerouting DNAT rules so mesh peers can reach a configured
|
||||
`host:port` on the gateway's LAN. A LAN-side masquerade is added
|
||||
automatically when any forwards are configured, so replies flow
|
||||
back through conntrack.
|
||||
|
||||
A dedicated control socket at `/run/fips/gateway.sock` exposes
|
||||
`show_gateway` and `show_mappings`. `fipstop` adds a Gateway tab with
|
||||
a pool gauge and mappings table.
|
||||
|
||||
The gateway's `dns.listen` source default is now `[::1]:5353`,
|
||||
matching the canonical deployment model: the gateway sits on a host
|
||||
already serving DHCP and DNS to a LAN segment (an OpenWrt AP, a Linux
|
||||
router), port 53 there is taken by the existing resolver, and `.fips`
|
||||
queries are forwarded to the gateway over loopback. The OpenWrt ipk
|
||||
previously overrode the prior `[::]:53` source default in its packaged
|
||||
config; that override is now redundant and has been dropped.
|
||||
Operators on a host without a pre-existing resolver on port 53 can
|
||||
opt back into the wildcard bind by setting `dns.listen: "[::]:53"`
|
||||
explicitly. The new default binds IPv6 loopback only, so forwarders
|
||||
that reach the gateway over IPv4 loopback need an explicit IPv4
|
||||
listen address.
|
||||
|
||||
The cold-boot startup race between `fips.service` and
|
||||
`fips-gateway.service` is handled by a systemd `After=fips.service`
|
||||
ordering, an `ExecStartPre` poll loop that waits up to 30 seconds for
|
||||
the `fips0` interface to appear, and a DNS upstream probe in the
|
||||
gateway itself that retries up to 5 times with 1-second backoff.
|
||||
|
||||
Packaging covers systemd, Debian, AUR, and OpenWrt. The full design
|
||||
is in [`docs/design/fips-gateway.md`](../design/fips-gateway.md).
|
||||
|
||||
### Mesh-interface security baseline
|
||||
|
||||
The FIPS mesh is a flat layer-3 segment. Every authenticated peer can
|
||||
route packets to every other peer's `fips0` address. Peer identity is
|
||||
authenticated end-to-end by the FMP and FSP Noise handshakes, but
|
||||
identity is not authorization. A service on a mesh host that binds to
|
||||
a wildcard address is, by default, reachable from every peer in the
|
||||
mesh.
|
||||
|
||||
v0.3.0 ships an opt-in default-deny baseline that closes this gap on
|
||||
Linux:
|
||||
|
||||
- **`/etc/fips/fips.nft`** is installed as a documented operator
|
||||
conffile. It defines a single `inet fips` nftables table with one
|
||||
chain hooked at `input`, default-denies inbound traffic on
|
||||
`fips0`, and is a no-op for every other interface.
|
||||
- **`fips-firewall.service`** loads it. The unit ships **disabled by
|
||||
default**; activation is an explicit
|
||||
`systemctl enable --now fips-firewall.service`.
|
||||
- Per-service allowances live in **`/etc/fips/fips.d/*.nft`**
|
||||
drop-ins that the baseline includes.
|
||||
|
||||
Choosing opt-in keeps the mesh quick to bring up for evaluation while
|
||||
giving operators a documented, packaged path to lock it down for
|
||||
production. The full design (threat model, rule layout, conntrack
|
||||
handling, drop-in mechanism, and the rationale for a conffile rather
|
||||
than an auto-loaded package side-effect) is in
|
||||
[`docs/design/fips-security.md`](../design/fips-security.md).
|
||||
|
||||
`fipstop`'s Node tab gains a **"Listening on fips0" panel** that
|
||||
surfaces the answer to the operational question "what services on
|
||||
this host are reachable from the mesh, and what does the firewall
|
||||
currently say about each of them?" The panel lists every IPv6
|
||||
listening socket bound to either the wildcard address or this node's
|
||||
`fd00::/8` address, paired with its classification against the
|
||||
running `inet fips` baseline chain: `OPEN` (canonical accept rule),
|
||||
`filt` (falls through to drop), or `filt?` (referenced with matchers
|
||||
the panel cannot fully decompose, e.g. saddr filters or jumps). When
|
||||
`fips-firewall.service` is inactive, a yellow banner above the table
|
||||
reminds the operator that every listener is mesh-exposed; wildcard
|
||||
binds carry a trailing `*` in the Process column. The classifier is
|
||||
built on a new `show_listening_sockets` control query (Linux-only),
|
||||
which is also useful from `fipsctl` for scripting.
|
||||
|
||||
### Peer access control
|
||||
|
||||
Operators can now restrict which mesh peers a node will form direct
|
||||
links with. Optional `/etc/fips/peers.allow` and `/etc/fips/peers.deny`
|
||||
files (TCP-Wrappers style) match against npub, hex pubkey, host
|
||||
alias, or `ALL`. Enforcement runs at three points:
|
||||
|
||||
1. Outbound connect (before dialing).
|
||||
2. Inbound msg1 (the first FMP handshake message from a new peer).
|
||||
3. Outbound msg2 (the response).
|
||||
|
||||
Files reload automatically on mtime change; a new `fipsctl acl show`
|
||||
query reports the effective rule set. A six-node Docker integration
|
||||
harness (`testing/acl/`) exercises allowlist and denylist patterns
|
||||
end-to-end.
|
||||
|
||||
**Important scope distinction**: peer ACLs are an FMP-layer
|
||||
restriction. They control who can establish a *direct link* with this
|
||||
node. They do **not** control session-layer (FSP) reachability through
|
||||
the mesh. A node that denies peer X with an ACL can still receive FSP
|
||||
traffic from X relayed via other peers.
|
||||
|
||||
### Bluetooth Low Energy transport (experimental, Linux)
|
||||
|
||||
A new BLE L2CAP Connection-Oriented Channel transport lets FIPS nodes
|
||||
peer over Bluetooth Low Energy without any IP infrastructure in
|
||||
between. The transport handles per-link MTU negotiation, continuous
|
||||
scan/probe peer discovery with cooldown-based deduplication,
|
||||
continuous advertising, deterministic NodeAddr cross-probe
|
||||
tie-breaker, and a configurable connection pool with eviction.
|
||||
|
||||
This transport is **experimental in v0.3.0**. It is implemented and
|
||||
functional on Linux (BlueZ via `bluer`), but the reliability follow-up
|
||||
logic (probe cooldown, cross-probe tie-breaker, pubkey timeout,
|
||||
continuous advertising semantics, probe-promotion, fail-fast send) is
|
||||
not yet behaviorally tested in CI. Its maturity path is field-driven;
|
||||
please file issues with field reports. macOS BLE support is in
|
||||
development as a separate track and is not part of v0.3.0.
|
||||
|
||||
### UDP transport profiles
|
||||
|
||||
The UDP transport gains posture flags organized around deployment
|
||||
patterns:
|
||||
|
||||
- **Public-facing inbound nodes**: `bind_addr: "0.0.0.0:2121"`,
|
||||
`accept_connections: true` (default), `public: true` for advert
|
||||
publication. v0.3.0 adds STUN-based public-IP autodiscovery so
|
||||
cloud nodes (AWS EIP, GCP, Azure 1:1 NAT) advertise the right
|
||||
address even when the public IP isn't on a host interface.
|
||||
- **Ephemeral leaf nodes**: `outbound_only: true` binds an ephemeral
|
||||
port (`0.0.0.0:0`), refuses inbound msg1, and is never advertised
|
||||
on Nostr regardless of `advertise_on_nostr`. Use this for client
|
||||
postures that should connect outbound only, without exposing an
|
||||
inbound listener on a known port.
|
||||
- **General-purpose nodes**: `accept_connections: false` mirrors the
|
||||
Ethernet/BLE knob without changing the bind address. The Node-level
|
||||
handshake gate carves out msg1 from peers already established on
|
||||
this transport so rekey continues to work.
|
||||
|
||||
Startup validation now rejects `bind_addr` set to a loopback address
|
||||
when at least one peer has a non-loopback UDP address, closing a
|
||||
silent-failure trap from v0.2.0 where Linux's source-address routing
|
||||
check would drop outbound flows from the loopback-bound socket.
|
||||
|
||||
A new `external_addr` field on `transports.udp.*` and
|
||||
`transports.tcp.*` lets operators specify the advertise-as address
|
||||
explicitly. This is useful for UDP as a deterministic alternative to
|
||||
STUN, and required for TCP on cloud-NAT setups (where binding to the
|
||||
public IP fails with `EADDRNOTAVAIL` because the IP isn't on a host
|
||||
interface).
|
||||
|
||||
### `.fips` DNS resolver overhaul
|
||||
|
||||
The IPv6 adapter's `.fips` name resolution has been rebuilt around
|
||||
the constraints of contemporary systemd-based hosts. The default
|
||||
`dns.bind_addr` is now `::1` (IPv6 loopback), and a setup script
|
||||
picks one of five backends in priority order:
|
||||
|
||||
1. systemd-resolved global drop-in
|
||||
(`/etc/systemd/resolved.conf.d/fips.conf`).
|
||||
2. systemd dns-delegate (per-link configuration handed off to
|
||||
systemd-resolved).
|
||||
3. `resolvectl` per-link configuration.
|
||||
4. Standalone `dnsmasq`.
|
||||
5. NetworkManager's dnsmasq plugin.
|
||||
|
||||
Teardown reverses only what setup applied, recorded in a state file
|
||||
at `/run/fips/dns-backend`. A new `testing/dns-resolver/` harness
|
||||
exercises every backend across Debian 12, Debian 13, Ubuntu 22.04,
|
||||
Ubuntu 24.04, and Ubuntu 26.04, so a regression in any of the five
|
||||
backends shows up in CI rather than in the field.
|
||||
|
||||
This overhaul resolves the long-standing silent-drop case where the
|
||||
`resolvectl dns fips0 [<fips0_addr>]:5354` target collided with the
|
||||
daemon's mesh-interface filter on certain systemd-resolved
|
||||
deployments (typically Ubuntu 22 with systemd 249's interface-scoped
|
||||
routing).
|
||||
|
||||
### Operator tooling additions
|
||||
|
||||
A handful of additions land in `fipsctl`, `fipstop`, and the daemon's
|
||||
configuration surface:
|
||||
|
||||
- **`node.log_level`** config field replaces the hardcoded
|
||||
`RUST_LOG=info` previously baked into systemd units and the
|
||||
OpenWrt procd init. The daemon now loads config before
|
||||
initializing tracing so the configured level takes effect.
|
||||
`RUST_LOG` still overrides when set.
|
||||
- **`fipsctl show identity-cache`** is a new query that lists every
|
||||
cached node identity (npub, IPv6 address, display name, LRU age)
|
||||
alongside the configured cache capacity.
|
||||
- **`fipsctl show peers / sessions / cache / routing`** are
|
||||
substantially extended: per-peer security signals (replay
|
||||
suppression count, consecutive decrypt failures), Noise session
|
||||
counters, session indices, rekey lifecycle state, handshake resend
|
||||
counts, K-bit epoch, coords-warmup remaining, drain state, per-peer
|
||||
retry state, per-target lookup detail (attempt, age, last sent),
|
||||
and pending TUN packet queue depth.
|
||||
- **Historical statistics**: in-memory time-series rings on the
|
||||
daemon (1-second × 3600 fast, 1-minute × 1440 slow) cover per-node
|
||||
and per-peer metrics. New `show_stats_*` control-socket queries, a
|
||||
`fipsctl stats list / peers / history` subcommand with Unicode
|
||||
sparkline rendering, and a `fipstop` Graphs tab with btop-style
|
||||
sparklines surface them to the operator.
|
||||
|
||||
### Performance
|
||||
|
||||
Two independent perf threads land in v0.3.0: a session-layer crypto
|
||||
backend swap, and a Linux receive-path overhaul.
|
||||
|
||||
**Session-layer crypto backend.** The ChaCha20-Poly1305 backend used
|
||||
by every FIPS Noise session (end-to-end FSP traffic and link-layer
|
||||
FMP traffic alike) has been swapped from RustCrypto's
|
||||
`chacha20poly1305` crate to `ring 0.17`. ring wraps BoringSSL's
|
||||
hand-tuned ChaCha20-Poly1305 implementation, which dispatches to NEON
|
||||
on aarch64 and AVX2 / AVX-512 on x86_64. Typical throughput is in the
|
||||
3-5 GB/s/core range, versus the ~600-800 MB/s/core RustCrypto soft
|
||||
path on the same hardware.
|
||||
|
||||
Wire format is unchanged. ChaCha20-Poly1305 is byte-deterministic for
|
||||
a given `(key, nonce, plaintext, aad)`, so any correct AEAD
|
||||
implementation produces identical ciphertext. A mixed mesh with some
|
||||
nodes pre-swap and some post-swap interoperates without protocol
|
||||
awareness; v0.3.0 can roll out across a mesh in any order.
|
||||
|
||||
Measurements on an aarch64 Apple Silicon docker target:
|
||||
|
||||
- Two-node TCP single-stream: 437 -> 1097 Mbps (about 2.5×).
|
||||
- Two-node UDP at 1000 Mbit: 599 Mbps with 40% loss -> lossless at
|
||||
line rate.
|
||||
- Three-node ping under bulk-traffic load: 7.68 ms avg / 215 ms max
|
||||
-> 0.72 ms / 3.6 ms max as the relay path stops being crypto-bound.
|
||||
|
||||
No operator-visible action is required; the swap is internal to the
|
||||
session layer.
|
||||
|
||||
**Linux UDP receive path.** The Linux UDP receive path now uses
|
||||
`recvmmsg(2)` with a 32-packet batch in place of single-packet
|
||||
`recvmsg(2)`. A single `readable()` wakeup drains up to 32 datagrams
|
||||
in one syscall before yielding back to the reactor, eliminating the
|
||||
per-packet scheduler-hop and futex cost that previously capped
|
||||
inbound rate at one event per scheduler quantum independent of CPU.
|
||||
`SO_RXQ_OVFL` is sampled once per batch and surfaced through
|
||||
`AsyncUdpSocket::recv_batch` so the existing 1Hz transport-congestion
|
||||
detector continues to feed the per-transport `dropping` flag. macOS
|
||||
and Windows fall through to the per-packet path; `recvmmsg` is
|
||||
Linux-specific.
|
||||
|
||||
**Inner rx-loop drain batching.** `Node::run_rx_loop` drains up to
|
||||
256 additional ready items via `try_recv()` after each
|
||||
`tokio::select!` await fires on the packet and TUN-outbound branches,
|
||||
in a tight inner loop before yielding. Previously the select cost a
|
||||
full scheduler hop and futex per packet, capping throughput at one
|
||||
event per scheduler quantum with the worker near-idle. `biased`
|
||||
ordering keeps data-plane branches priority over tick / control / DNS
|
||||
under sustained load; the 256 cap keeps the worker on a busy stream
|
||||
between yield points (about 400 KB of contiguous traffic) while still
|
||||
bounding the inner loop so a flood on one branch cannot starve the
|
||||
periodic tick or control socket.
|
||||
|
||||
**Eager `pubkey_full` precompute.** `PeerIdentity::pubkey_full()`
|
||||
precomputes the parity-aware full secp256k1 public key at
|
||||
construction in `from_pubkey`. Previously the method fell through to
|
||||
an EC point parse on every call when the full key wasn't passed at
|
||||
construction (i.e. for every peer constructed from an npub or x-only
|
||||
key), about 6% of per-packet CPU on the bulk-data send path for a
|
||||
value that never changed after construction. The same parse already
|
||||
runs at construction inside `NodeAddr::from_pubkey`, so the cost is
|
||||
paid once where it would be paid anyway.
|
||||
|
||||
These three changes are a coordinated set: the syscall batching
|
||||
removes the per-packet kernel cost, the inner-loop drain removes the
|
||||
per-packet scheduler cost, and the pubkey-cache change removes the
|
||||
per-packet crypto-derivation cost. Like the AEAD swap, they are all
|
||||
internal and require no operator action.
|
||||
|
||||
### Examples
|
||||
|
||||
- **macOS WireGuard companion** ([#51](https://github.com/jmcorgan/fips/pull/51)):
|
||||
run FIPS in a local Docker container and route `.fips` traffic
|
||||
from the macOS host through a WireGuard tunnel to the container's
|
||||
`fips0`. Only traffic destined for `fd00::/8` transits the
|
||||
companion; regular internet traffic continues to use the host
|
||||
network. Persistent FIPS and WireGuard key material is generated
|
||||
on first run.
|
||||
|
||||
### Documentation
|
||||
|
||||
- **`docs/design/port-advertisement-and-nat-traversal.md`**
|
||||
documents how nodes find each other through Nostr relays and the
|
||||
STUN-assisted UDP hole punch.
|
||||
- **`docs/design/fips-gateway.md`** documents the gateway's virtual
|
||||
IP pool, lifecycle, control surface, and packaging.
|
||||
- **`docs/design/fips-security.md`** documents the mesh-interface
|
||||
security posture, threat model, default-deny baseline, and drop-in
|
||||
workflow.
|
||||
- **`CONTRIBUTING.md`** has been expanded with build prerequisites,
|
||||
Rust toolchain setup, and first-build steps.
|
||||
|
||||
The `docs/` tree has been reorganized end-to-end into four sections
|
||||
(*tutorials / how-to / reference / design*) with a new
|
||||
[`docs/getting-started.md`](../getting-started.md) and per-section
|
||||
landing pages. Content was reconciled against current source:
|
||||
protocol-layer details, wire-format diagrams, configuration knobs,
|
||||
and CLI references were brought back into agreement with the
|
||||
implementation. See [Documentation pointers](#documentation-pointers)
|
||||
below for entry points by reader intent.
|
||||
|
||||
## Behavior changes worth flagging
|
||||
|
||||
These default-config changes affect every operator on upgrade, even
|
||||
those with no explicit configuration. Two items below — bloom-filter
|
||||
fill-ratio validation and TreeAnnounce ancestry validation — first
|
||||
shipped in v0.2.1 and roll forward into v0.3.0; the rest are
|
||||
v0.3.0-net-new.
|
||||
|
||||
- **Discovery rate-limiting** has been retuned to be less aggressive
|
||||
at cold start. v0.2.0 used a single-lookup-with-internal-retry
|
||||
model where a timed-out lookup during bloom-filter propagation
|
||||
could suppress retries for 30 seconds while none of the reset
|
||||
triggers fired on a stable post-handshake topology. v0.3.0
|
||||
replaces this with a per-attempt timeout sequence
|
||||
(`node.discovery.attempt_timeouts_secs`, default `[1, 2, 4, 8]`,
|
||||
15s total). Each attempt sends a fresh `LookupRequest` with a new
|
||||
`request_id`, letting successive attempts take different
|
||||
forwarding paths as the bloom and tree state evolve. Post-failure
|
||||
suppression is **off by default**; operators with chatty
|
||||
applications can opt back in via `backoff_base_secs` /
|
||||
`backoff_max_secs`.
|
||||
- **MMP report intervals** are retuned for constrained transports.
|
||||
The steady-state floor moves from 100ms to 1000ms, the ceiling
|
||||
from 2000ms to 5000ms, with a cold-start phase running 200ms for
|
||||
the first 5 SRTT samples. This reduces BLE overhead by roughly
|
||||
10× while keeping reports well above the EWMA convergence
|
||||
threshold. Session-layer MMP intervals are unchanged.
|
||||
- **Bloom filter fill-ratio validation** runs on every inbound
|
||||
`FilterAnnounce`. Filters whose derived false-positive rate
|
||||
exceeds `node.bloom.max_inbound_fpr` (default 0.05) are rejected
|
||||
silently on the wire, logged at WARN, and counted in a new
|
||||
`bloom.fill_exceeded` counter. A rate-limited WARN also fires
|
||||
when the local outgoing filter exceeds the cap.
|
||||
- **TreeAnnounce ancestry validation** is now run before tree-state
|
||||
mutation, enforcing ancestry-self-match, root-single-entry,
|
||||
parent-second-entry, and root-is-minimum-NodeAddr. Non-conforming
|
||||
announces are rejected with a WARN. Mixed v0.2.0 / v0.2.1 / v0.3.0
|
||||
meshes may produce WARN log lines on the v0.2.1+ side until all
|
||||
peers upgrade; behavior is correct, log noise only.
|
||||
- **Log noise reduction**: 35 info-level log messages have been
|
||||
demoted to debug (handshake cross-connection mechanics, periodic
|
||||
MMP telemetry, TUN/transport shutdown, retry scheduling). The
|
||||
default `RUST_LOG` in systemd units is now `info`, where it
|
||||
previously ran at `debug`. Operator-visible info output now
|
||||
focuses on lifecycle events, peer promotions, session
|
||||
establishment, parent switches, and transport start/stop.
|
||||
|
||||
## Notable bug fixes
|
||||
|
||||
These pre-existing v0.2.0 bugs are worth singling out because they
|
||||
either affected real-world deployments or produced misleading
|
||||
operator experiences. The CHANGELOG has the exhaustive list; this is
|
||||
the operator-relevant subset. Four items below first shipped in
|
||||
v0.2.1 and roll forward into v0.3.0: auto-connect Disconnect-reconnect,
|
||||
`fipsctl connect` mesh-address rejection, `fd00::/8` routing
|
||||
protection from Tailscale interception, and bloom-filter routing
|
||||
greedy-tree fallback. The control-socket path-detection fix landed
|
||||
in v0.2.1 as well, and the unified resolver below is the v0.3.0
|
||||
refactor that builds on it.
|
||||
|
||||
- **DNS responder silent-drop on systemd-resolved** is fixed: the
|
||||
responder no longer drops queries on Ubuntu 22 / Debian 13 and
|
||||
similar deployments where systemd applies interface-scoped
|
||||
routing. Default bind moves to `::1`; new global drop-in backend
|
||||
available ([#52](https://github.com/jmcorgan/fips/issues/52),
|
||||
[#77](https://github.com/jmcorgan/fips/issues/77)).
|
||||
- **Auto-connect peers reconnect after a graceful Disconnect.**
|
||||
Previously, a clean upstream shutdown left the auto-connect peer
|
||||
orphaned; only the link-dead, decrypt-fail, and peer-restart
|
||||
paths scheduled a reconnect
|
||||
([#60](https://github.com/jmcorgan/fips/issues/60), reported by
|
||||
[@SwapMarket](https://github.com/SwapMarket)).
|
||||
- **`fipsctl connect` rejects FIPS mesh addresses** (`fd00::/8`)
|
||||
for `udp`, `tcp`, and `ethernet` transports with a clear error
|
||||
message, instead of echoing success while the daemon silently
|
||||
failed the bind with `EAFNOSUPPORT`
|
||||
([#61](https://github.com/jmcorgan/fips/issues/61), reported by
|
||||
[@SwapMarket](https://github.com/SwapMarket)).
|
||||
- **Default control-socket path resolution unified.** Daemon and
|
||||
client tools now share a single resolver, eliminating a divergence
|
||||
where `fipsctl` / `fipstop` could connect to a socket the daemon
|
||||
never bound (notably on dev runs with `XDG_RUNTIME_DIR` set, or
|
||||
after a prior packaged install left a root-owned `/run/fips`
|
||||
behind). Canonical order is
|
||||
`/run/fips` -> `$XDG_RUNTIME_DIR/fips/` -> `/tmp/fips-<name>`. The
|
||||
`/run/fips` arm is selected by directory existence; the kernel
|
||||
enforces actual access at `connect(2)` time, so users not yet in
|
||||
the `fips` group get a clear `EACCES` rather than a silent path
|
||||
mismatch and a misleading `No such file` fallback to
|
||||
`$XDG_RUNTIME_DIR`. `XDG_RUNTIME_DIR` is validated as an existing
|
||||
directory before being used so stale post-logout values are
|
||||
treated as missing. The deployed fleet is unaffected: packaged
|
||||
configs set `node.control.socket_path` explicitly
|
||||
([#30](https://github.com/jmcorgan/fips/issues/30), reported by
|
||||
[@Sebastix](https://github.com/Sebastix)).
|
||||
- **`fd00::/8` routing protected from Tailscale interception.** The
|
||||
daemon installs an IPv6 routing-policy rule
|
||||
(`ip -6 rule to fd00::/8 lookup main priority 5265`) at TUN
|
||||
setup, so Tailscale's table 52 default route can no longer divert
|
||||
mesh traffic.
|
||||
- **TCP-over-FIPS reliability on mixed-MTU paths** is markedly
|
||||
improved. Four interlocking changes ship together:
|
||||
`Node::transport_mtu()` is now deterministic across daemon
|
||||
restarts (min across operational transports rather than
|
||||
insertion-order-dependent); the TCP MSS clamp at the TUN boundary
|
||||
reads per-destination path MTU instead of a single global ceiling;
|
||||
reactive `MtuExceeded` from forwarders is mirrored back into the
|
||||
TUN-side `path_mtu_lookup` so later flows pick up forward-path
|
||||
bottlenecks without re-discovery; and the proactive end-to-end
|
||||
`PathMtuNotification` echoed by the destination is mirrored into
|
||||
the same TUN-side store. Without that fourth piece, on long-lived
|
||||
stable paths where the destination's echo had tightened the
|
||||
session MTU but no transit router had emitted a fresh
|
||||
`MtuExceeded`, new TCP flows opened in that window were clamped by
|
||||
the staler discovery-time value. The proactive mirror uses the
|
||||
same tighter-only semantics as the reactive mirror, so it never
|
||||
loosens the clamp. The Windows TUN reader receives the same
|
||||
per-destination plumbing.
|
||||
- **Bloom filter routing greedy-tree fallback.** `find_next_hop` no
|
||||
longer returns `NoRoute` when the bloom candidate set is non-empty
|
||||
but no candidate is strictly closer than the current node; it
|
||||
falls through to greedy tree routing instead. Previously, this
|
||||
caused dropped packets in topologies where the tree parent was
|
||||
closer but not a bloom candidate.
|
||||
- **`fipstop` graceful tty-init failure.** `ratatui::try_init()`
|
||||
produces a clean error message instead of a hard crash when
|
||||
terminal initialization fails (Docker on macOS Sequoia, ttyless
|
||||
environments).
|
||||
- **TreeAnnounce ancestry on self-root transitions.** When a node
|
||||
had no smaller-NodeAddr peer to use as a parent, the spanning-tree
|
||||
state correctly promoted it to root, but the ancestry advertised
|
||||
on the next `TreeAnnounce` still referenced its previous parent's
|
||||
path. Receiving peers rejected the announce as
|
||||
`invalid ancestry: advertised root X is not the minimum path entry
|
||||
Y`, blocking mesh transit on any path that needed to traverse the
|
||||
node. The self-root transition is now detected explicitly in
|
||||
`TreeState::become_root` and the advertised ancestry rebuilt to
|
||||
start from self; the MMP receive handler corrects stale ancestry
|
||||
inherited across reconnect eagerly rather than waiting for the
|
||||
next observation tick.
|
||||
- **Spanning-tree internal-path updates** that change only the
|
||||
internal path between root and leaf (without changing the root or
|
||||
the depth) now propagate to leaves correctly. Previously, a leaf
|
||||
could continue routing against a stale internal path until the
|
||||
parent or depth also changed.
|
||||
|
||||
## Upgrade notes
|
||||
|
||||
Operator-actionable items when moving from v0.2.x to v0.3.0:
|
||||
|
||||
- **Control socket JSON schema (breaking, pre-1.0).**
|
||||
- `show_cache` response field `entries` has changed type from a
|
||||
`u64` count to an array of entry objects. The previous scalar
|
||||
value is now in a new `count` field.
|
||||
- `show_routing` response field `pending_lookups` has changed
|
||||
type from a `u64` count to an array of per-target lookup
|
||||
objects.
|
||||
- External tooling parsing these fields as numbers must be
|
||||
updated. In-tree `fipstop` is adjusted to the new schema. The
|
||||
control-socket interface remains pre-1.0 and is not covered by
|
||||
stability guarantees.
|
||||
|
||||
- **Cargo feature flags removed.** `tui`, `ble`, `gateway`, and
|
||||
`nostr-discovery` are gone. Subsystem inclusion is now driven by
|
||||
platform `cfg` gates, so plain `cargo build` compiles everything
|
||||
available on the target without `--features` invocations.
|
||||
Source-build tooling that passed any of these features should be
|
||||
updated to omit them.
|
||||
|
||||
- **Discovery rate-limiting defaults changed.** Post-failure
|
||||
suppression is **off by default**
|
||||
(`node.discovery.backoff_base_secs: 0`, `backoff_max_secs: 0`).
|
||||
Operators relying on the prior 30s base / 300s cap behavior must
|
||||
set those fields explicitly. The per-attempt sequence
|
||||
(`attempt_timeouts_secs`, default `[1, 2, 4, 8]`) now governs
|
||||
cold-start lookup behavior.
|
||||
|
||||
- **`.fips` DNS bind address default changed.** The default
|
||||
`dns.bind_addr` is now `::1`. Operators with explicit overrides
|
||||
of this field should review them; many existing overrides were
|
||||
workarounds for the silent-drop bug that this release fixes
|
||||
properly.
|
||||
|
||||
- **Gateway `dns.listen` source default changed.** The
|
||||
`fips-gateway` `dns.listen` default is now `[::1]:5353` (was
|
||||
`[::]:53`), matching the canonical deployment model where a
|
||||
pre-existing resolver on the host already owns port 53. The
|
||||
OpenWrt ipk previously overrode this in its packaged config; the
|
||||
override is now redundant and has been dropped. Operators on a
|
||||
host without a pre-existing resolver on port 53 can opt back into
|
||||
the wildcard bind by setting `dns.listen: "[::]:53"` explicitly.
|
||||
The new default binds IPv6 loopback only, so forwarders that
|
||||
reach the gateway over IPv4 loopback need an explicit IPv4 listen
|
||||
address.
|
||||
|
||||
- **systemd unit log level.** The shipped systemd units no longer
|
||||
hardcode `RUST_LOG=info`; the daemon's effective log level is
|
||||
driven by `node.log_level` (default `info`). `RUST_LOG`, when
|
||||
set, still overrides.
|
||||
|
||||
- **UDP transport `bind_addr` validation.** Startup now rejects a
|
||||
`bind_addr` set to a loopback address when at least one peer has
|
||||
a non-loopback UDP address. Operators who configured a loopback
|
||||
UDP bind as a workaround should switch to `outbound_only: true`
|
||||
for the same effect, plus the correct semantics (kernel-assigned
|
||||
ephemeral port, refuses inbound, never advertised).
|
||||
|
||||
- **Tor advert port.** If the Tor `HiddenServicePort` virtual port
|
||||
isn't 443, set `transports.tor.advertised_port` to match. The
|
||||
default is 443 and matches the conventional virtual-port choice.
|
||||
|
||||
## Documentation pointers
|
||||
|
||||
v0.3.0 ships a `docs/` tree reorganized into four sections
|
||||
(*tutorials / how-to / reference / design*). A new top-level
|
||||
[`docs/getting-started.md`](../getting-started.md) and per-section
|
||||
landing pages anchor the entry points.
|
||||
|
||||
Entry points by reader intent:
|
||||
|
||||
- **New users**: [`docs/getting-started.md`](../getting-started.md)
|
||||
and [`docs/tutorials/`](../tutorials/) cover guided introductions
|
||||
for bringing up your first node, joining the test mesh,
|
||||
advertising a node over Nostr, hosting a service, deploying a
|
||||
gateway, walking through the IPv6 adapter, and resolving peers
|
||||
via Nostr.
|
||||
- **Operators with a specific task**:
|
||||
[`docs/how-to/`](../how-to/) holds task-driven guides for enabling
|
||||
Nostr discovery, deploying the gateway, troubleshooting the
|
||||
gateway, deploying a Tor onion, hosting aliases, persistent
|
||||
identity, running unprivileged, setting up a Bluetooth peer,
|
||||
enabling the mesh firewall, tuning UDP buffers, and diagnosing
|
||||
MTU issues.
|
||||
- **Reference lookups**: [`docs/reference/`](../reference/) holds
|
||||
the config field reference, control-socket query reference, the
|
||||
`fips`, `fipsctl`, `fipstop`, and `fips-gateway` CLI references,
|
||||
and the protocol diagram set.
|
||||
- **Architectural background**: [`docs/design/`](../design/) holds
|
||||
design rationale for FIPS as a whole, FMP and FSP, the spanning
|
||||
tree, bloom-filter discovery, transports, the IPv6 adapter, the
|
||||
Nostr discovery layer, and the gateway.
|
||||
- **Security**: [`docs/design/fips-security.md`](../design/fips-security.md)
|
||||
documents the mesh-interface security baseline, threat model, and
|
||||
drop-in workflow.
|
||||
|
||||
## Getting v0.3.0
|
||||
|
||||
- **Linux x86_64 / aarch64**: `.deb` and tarball at the
|
||||
[v0.3.0 release page](https://github.com/jmcorgan/fips/releases/tag/v0.3.0).
|
||||
- **Arch Linux**: `fips` from the AUR.
|
||||
- **macOS**: `.pkg` at the v0.3.0 release page.
|
||||
- **Windows**: ZIP at the v0.3.0 release page.
|
||||
- **OpenWrt**: `.ipk` at the v0.3.0 release page.
|
||||
- **From source**: `cargo build --release` from a checkout of the
|
||||
v0.3.0 tag.
|
||||
|
||||
The full per-commit changelog lives in
|
||||
[`CHANGELOG.md`](../../CHANGELOG.md). Issues and discussion at
|
||||
[github.com/jmcorgan/fips](https://github.com/jmcorgan/fips).
|
||||
|
||||
## Contributors
|
||||
|
||||
Thanks to everyone who contributed code, packaging work, bug reports,
|
||||
or reviews to this release.
|
||||
|
||||
**Code and packaging**:
|
||||
|
||||
- [@jcorgan](https://github.com/jmcorgan): release shepherd, Nostr
|
||||
discovery / NAT traversal, `fips-gateway`, ACL infrastructure,
|
||||
packaging, security baseline, BLE follow-ups.
|
||||
- [@Origami74](https://github.com/Origami74): macOS platform support,
|
||||
from-source Docker companion build and `fipstop` terminal-init
|
||||
handling, gateway co-development, OpenWrt BLE-feature build fix,
|
||||
AUR-workflow follow-ups.
|
||||
- [@jodobear](https://github.com/jodobear): Linux release-artifact
|
||||
workflow and target-aware build scripts, CONTRIBUTING.md
|
||||
expansion, rekey integration-test stabilization.
|
||||
- [@tidley](https://github.com/tidley): Nostr-mediated overlay
|
||||
discovery and UDP NAT traversal
|
||||
([#53](https://github.com/jmcorgan/fips/pull/53)).
|
||||
- [@alexxie16](https://github.com/alexxie16): peer ACL enforcement
|
||||
([#50](https://github.com/jmcorgan/fips/pull/50)),
|
||||
macOS WireGuard companion example
|
||||
([#51](https://github.com/jmcorgan/fips/pull/51)),
|
||||
follow-up ([#67](https://github.com/jmcorgan/fips/pull/67)).
|
||||
- [@osh](https://github.com/osh): diagnostic queries for security
|
||||
validation and mesh debugging
|
||||
([#42](https://github.com/jmcorgan/fips/pull/42)).
|
||||
- [@OceanSlim](https://github.com/0ceanSlim): Windows platform
|
||||
support ([#45](https://github.com/jmcorgan/fips/pull/45)).
|
||||
- [@mmalmi](https://github.com/mmalmi): ring AEAD backend
|
||||
([#80](https://github.com/jmcorgan/fips/pull/80)),
|
||||
hot-path drain batching + recvmmsg + eager pubkey_full
|
||||
([#81](https://github.com/jmcorgan/fips/pull/81)),
|
||||
TreeAnnounce self-root ancestry + overlay-advert retry hygiene
|
||||
([#82](https://github.com/jmcorgan/fips/pull/82)),
|
||||
NAT-traversal MTU inheritance
|
||||
([#83](https://github.com/jmcorgan/fips/pull/83)).
|
||||
- [@dskvr](https://github.com/dskvr): initial Arch Linux AUR
|
||||
packaging ([#21](https://github.com/jmcorgan/fips/pull/21)) and
|
||||
the AUR publish workflow.
|
||||
- [@SatsAndSports](https://github.com/SatsAndSports): rekey
|
||||
message-1 admit fix on non-accepting transports
|
||||
([#49](https://github.com/jmcorgan/fips/pull/49)),
|
||||
TreeAnnounce semantic validation, gateway test image fix
|
||||
([#69](https://github.com/jmcorgan/fips/pull/69)).
|
||||
- [@andrewheadricke](https://github.com/andrewheadricke): MIPS
|
||||
atomic-ABI portability via `portable_atomic`
|
||||
([#62](https://github.com/jmcorgan/fips/pull/62)).
|
||||
- [@sh1ftred](https://github.com/sh1ftred): Arch packaging namcap
|
||||
fixes ([#63](https://github.com/jmcorgan/fips/pull/63)).
|
||||
- [@oleksky](https://github.com/oleksky): macOS WireGuard companion
|
||||
collaboration on [#51](https://github.com/jmcorgan/fips/pull/51).
|
||||
|
||||
**Issue reports that drove fixes in this release**:
|
||||
|
||||
- [@deavmi](https://github.com/deavmi): MIPS daemon build support
|
||||
([#26](https://github.com/jmcorgan/fips/issues/26)).
|
||||
- [@Sebastix](https://github.com/Sebastix): fipsctl/fipstop
|
||||
control-socket path detection
|
||||
([#30](https://github.com/jmcorgan/fips/issues/30)).
|
||||
- [@SwapMarket](https://github.com/SwapMarket): auto-connect
|
||||
reconnect after graceful disconnect
|
||||
([#60](https://github.com/jmcorgan/fips/issues/60)) and
|
||||
fipsctl mesh-address rejection
|
||||
([#61](https://github.com/jmcorgan/fips/issues/61)).
|
||||
@@ -0,0 +1,73 @@
|
||||
# Tutorials
|
||||
|
||||
If you have just installed FIPS, this is where to start. The
|
||||
tutorials below take you from a freshly-installed daemon to a
|
||||
node that:
|
||||
|
||||
- Has joined the public test mesh and can reach other nodes on it.
|
||||
- Carries a stable identity that other operators can address.
|
||||
- Discovers peers — and is discoverable — over Nostr.
|
||||
- Hosts and consumes real services across the mesh.
|
||||
|
||||
Each tutorial is a complete, working session at the keyboard. You
|
||||
configure something, restart the daemon, watch it come up, and
|
||||
verify the result. The point is to build muscle memory, not to
|
||||
cover every option.
|
||||
|
||||
> **Read them in order.** Each tutorial assumes the state the
|
||||
> previous one left you in. If you skip ahead, the cross-references
|
||||
> that lead you back may not match what you have on disk.
|
||||
|
||||
## The new-user progression
|
||||
|
||||
| # | Tutorial | What you'll do |
|
||||
| - | -------- | -------------- |
|
||||
| 1 | [join-the-test-mesh.md](join-the-test-mesh.md) | Add one public test peer to your config, watch the link come up, ping that peer and a second mesh node it routes you to. The starting point for everything else. |
|
||||
| 2 | [persistent-identity.md](persistent-identity.md) | Pin your daemon to a stable Nostr keypair so your address stops changing on every restart. Other operators can now add you to their `peers:` lists; the services you host get a fixed name. |
|
||||
| 3 | [resolve-peers-via-nostr.md](resolve-peers-via-nostr.md) | Stop hard-coding peer addresses. Drop the address line from your peer entry and let the daemon look up the current endpoint from public Nostr relays at dial time. |
|
||||
| 4 | [advertise-your-node.md](advertise-your-node.md) | Publish your own UDP endpoint to Nostr so any operator who knows your npub can reach you, with a short final section on `udp:nat` best-effort hole-punching for nodes without a directly reachable UDP endpoint. |
|
||||
| 5 | [open-discovery.md](open-discovery.md) | Switch to `policy: open` and let your peer list populate itself from the ambient `fips-overlay-v1` namespace. Hands-off mesh participation. |
|
||||
| 6 | [reach-mesh-services.md](reach-mesh-services.md) | Drive ordinary IPv6 tools — `ping6`, `nc`, `traceroute6`, `curl`, `ssh` — at mesh nodes by `<npub>.fips`. Get a feel for the daemon's IPv6 adapter, which makes unmodified IPv6 software work over the mesh. |
|
||||
| 7 | [host-a-service.md](host-a-service.md) | Bring up an HTTP server bound to `fips0` so mesh nodes can reach it, with a deliberate exposure decision (mesh-only vs every interface), and the mesh firewall as a default-deny baseline. The peer ACL (a separate, transport-layer control over which npubs may peer with your node) is briefly mentioned alongside. |
|
||||
| 8 | [ground-up-mesh.md](ground-up-mesh.md) | Bring up a second deployment mode: two devices joined by Ethernet (or WiFi, or BLE) with no IP infrastructure between them. The mesh emerges from layer 2 up. Coexists with overlay peers — the same daemon can carry both. |
|
||||
|
||||
After tutorial 8 you have a fully participating mesh node that
|
||||
reaches services hosted by other mesh nodes and hosts services of its own,
|
||||
with identity, discovery, reachability, an explicit exposure
|
||||
policy, and an understanding of both deployment modes — overlay
|
||||
on top of existing IP, and ground-up where the mesh is the
|
||||
network.
|
||||
|
||||
There is also a side trip you can take any time after tutorial 1:
|
||||
|
||||
- [ipv6-adapter-walkthrough.md](ipv6-adapter-walkthrough.md) —
|
||||
trace one `ssh` from DNS query through session setup to the
|
||||
far-side TUN, using `fipstop` and `fipsctl` to watch each step.
|
||||
Optional, but if you like seeing how the pieces fit together,
|
||||
this is the doc that shows you.
|
||||
|
||||
## Advanced
|
||||
|
||||
These are not part of the new-user progression. They assume you
|
||||
have already worked through the tutorials above and now want to
|
||||
fold FIPS into a wider network deployment.
|
||||
|
||||
- [deploy-fips-gateway.md](deploy-fips-gateway.md) — Stand up a
|
||||
`fips-gateway` on an OpenWrt access point so unmodified LAN
|
||||
hosts can reach `<npub>.fips` destinations through a DNS-
|
||||
allocated virtual IPv6 pool and kernel nftables NAT, with no
|
||||
per-host FIPS install. Also walks through one inbound port
|
||||
forward exposing a LAN service to mesh peers. Aimed at
|
||||
operators bridging a LAN segment into the overlay from the
|
||||
edge router. For a non-OpenWrt host the same deployment is
|
||||
in [../how-to/deploy-gateway.md](../how-to/deploy-gateway.md).
|
||||
|
||||
## When to use the how-to guides instead
|
||||
|
||||
The tutorials here walk through one specific path each. The
|
||||
how-to guides under [../how-to/](../how-to/) are the operator
|
||||
recipes — alternative provisioning paths, less-common
|
||||
configurations, troubleshooting techniques. Once you have the
|
||||
shape of FIPS in your head from these tutorials, the how-tos are
|
||||
where you'll go to look up "how do I do X?" without being walked
|
||||
through the surrounding context.
|
||||
@@ -0,0 +1,435 @@
|
||||
# Advertise Your Node on Nostr
|
||||
|
||||
After
|
||||
[resolve-peers-via-nostr](resolve-peers-via-nostr.md) your
|
||||
daemon can look up a peer's current endpoint by npub. This
|
||||
tutorial flips it around: you publish a signed advert listing
|
||||
your own endpoint(s), so any other operator who knows your
|
||||
npub can dial you the same way you dialed `test-us01`.
|
||||
|
||||
The whole exercise should take about ten minutes if you have
|
||||
a public IP or a UDP listener that's reachable from outside.
|
||||
A short final section covers `udp:nat` best-effort hole-punching
|
||||
for the cases where direct UDP advertising isn't an option.
|
||||
|
||||
## What you'll build
|
||||
|
||||
```text
|
||||
┌───────────────────────┐
|
||||
│ your fips daemon │
|
||||
│ persistent npub │
|
||||
└──────────┬────────────┘
|
||||
│ signed advert (Kind 37195)
|
||||
│ { udp:<your-public-ip>:2121, ... }
|
||||
│ refreshes every 30 min
|
||||
▼
|
||||
┌──────────────────────────────────────────┐
|
||||
│ Nostr relays │
|
||||
│ relay.damus.io / nos.lol / offchain.pub│
|
||||
└──────────────────────┬───────────────────┘
|
||||
│
|
||||
│ "what's <your-npub>'s endpoint?"
|
||||
│
|
||||
┌───────────┴───────────┐
|
||||
│ another fips daemon │
|
||||
│ knows your npub, │
|
||||
│ via_nostr: true │
|
||||
└───────────────────────┘
|
||||
```
|
||||
|
||||
You will change two things in `/etc/fips/fips.yaml`:
|
||||
|
||||
- Flip `discovery.nostr.advertise` from `false` to `true`.
|
||||
- Add `advertise_on_nostr: true` and `public: true` under
|
||||
`transports.udp`.
|
||||
|
||||
After restart, your daemon publishes a Kind 37195 event tied
|
||||
to your npub, listing the UDP endpoint other peers should
|
||||
dial.
|
||||
|
||||
## How advertising works
|
||||
|
||||
> **Adverts are signed Nostr events.** Every advert is a Kind
|
||||
> 37195 event signed by your daemon's secret key. Anyone
|
||||
> reading it can verify the advert really came from the npub
|
||||
> claiming the endpoint. The advert is the `(npub → current
|
||||
> endpoints)` mapping, signed and published.
|
||||
|
||||
The advert lists transports the daemon is willing to expose,
|
||||
and only those:
|
||||
|
||||
> **Endpoints are opt-in per transport.** Only transports
|
||||
> with `advertise_on_nostr: true` are listed in your advert.
|
||||
> Transports without that flag stay private — they still
|
||||
> work for peers who reach you via static config, but they
|
||||
> won't appear in your published advert.
|
||||
|
||||
For UDP specifically, the daemon needs to know what IP and
|
||||
port to put in the advert:
|
||||
|
||||
> **Determining the advertised endpoint (wildcard-bound UDP).**
|
||||
> With UDP bound to a wildcard like `0.0.0.0:2121`, the daemon
|
||||
> doesn't know its own public IP at startup. You have two ways
|
||||
> to tell it what to put in the advert:
|
||||
>
|
||||
> - `public: true` — daemon does a one-shot STUN observation
|
||||
> against the configured STUN servers and uses the reflexive
|
||||
> IPv4 it learns. Right when your public IP is dynamic or
|
||||
> you'd rather not pin it in config. Note: STUN observes the
|
||||
> reflexive IP from an ephemeral socket, then pairs it with
|
||||
> the listener's bind port for the advert — the advert is
|
||||
> only useful if your listener really is reachable at that
|
||||
> public IP/port, which the daemon can't tell from STUN
|
||||
> alone. A manual probe from a second host is the only sure
|
||||
> check.
|
||||
> - `external_addr: "<ip>[:<port>]"` — explicit override.
|
||||
> Right when you already know your public IP — a static
|
||||
> residential IP, an Elastic IP behind 1:1 NAT, a cloud
|
||||
> instance whose advertised port differs from the bind
|
||||
> port — and you don't want to depend on STUN reachability.
|
||||
> Required for TCP on cloud setups where binding directly
|
||||
> to the public IP returns `EADDRNOTAVAIL`.
|
||||
>
|
||||
> If you bind UDP to a specific public IP rather than
|
||||
> `0.0.0.0`, neither STUN nor `external_addr` is needed — but
|
||||
> `advertise_on_nostr: true` and `public: true` are still both
|
||||
> required for the daemon to publish the endpoint.
|
||||
|
||||
Adverts don't sit on the relays forever:
|
||||
|
||||
> **TTL and refresh.** Adverts have a 1-hour expiration
|
||||
> (NIP-40 `expiration` tag) and the daemon re-publishes every
|
||||
> 30 minutes. If your daemon goes offline, your advert decays
|
||||
> from caches in roughly an hour and consumers stop trying.
|
||||
|
||||
## Step 1: Confirm your starting state
|
||||
|
||||
You should be coming out of
|
||||
[resolve-peers-via-nostr](resolve-peers-via-nostr.md) with:
|
||||
|
||||
- A persistent npub (`fipsctl show status | grep '"npub"'`).
|
||||
- Nostr discovery in consume-only mode
|
||||
(`discovery.nostr.enabled: true`,
|
||||
`discovery.nostr.advertise: false`).
|
||||
- A peer entry for `test-us01` with `via_nostr: true` and no
|
||||
static address. `fipsctl show peers` shows the link
|
||||
established.
|
||||
|
||||
If any of those isn't true, finish the previous tutorials
|
||||
first.
|
||||
|
||||
Capture your npub now — you'll need it for the verification
|
||||
step:
|
||||
|
||||
```sh
|
||||
sudo fipsctl show status | grep '"npub"'
|
||||
```
|
||||
|
||||
Copy the value.
|
||||
|
||||
## Step 2: Enable advertising in the config
|
||||
|
||||
Open `/etc/fips/fips.yaml` and change two things.
|
||||
|
||||
**Change 1: flip `advertise` to `true`.** Find the
|
||||
`discovery.nostr` block under `node:` and set:
|
||||
|
||||
```yaml
|
||||
node:
|
||||
identity:
|
||||
persistent: true
|
||||
discovery:
|
||||
nostr:
|
||||
enabled: true
|
||||
advertise: true
|
||||
```
|
||||
|
||||
(The previous tutorial set `advertise: false`; you're flipping
|
||||
that bit now.)
|
||||
|
||||
**Change 2: add the UDP advert flags.** Find the `udp:` block
|
||||
under `transports:`. The wildcard-bind default
|
||||
(`0.0.0.0:2121`) means the daemon needs help knowing what to
|
||||
advertise — pick one of the two approaches from the callout
|
||||
above.
|
||||
|
||||
If you want STUN auto-discovery (works for full-cone NATs and
|
||||
nodes with a directly-bound public IP):
|
||||
|
||||
```yaml
|
||||
transports:
|
||||
udp:
|
||||
bind_addr: "0.0.0.0:2121"
|
||||
advertise_on_nostr: true
|
||||
public: true
|
||||
```
|
||||
|
||||
If you already know your public IP (e.g., a static residential
|
||||
IP or a cloud Elastic IP behind 1:1 NAT) and want to skip the
|
||||
STUN dependency:
|
||||
|
||||
```yaml
|
||||
transports:
|
||||
udp:
|
||||
bind_addr: "0.0.0.0:2121"
|
||||
advertise_on_nostr: true
|
||||
public: true
|
||||
external_addr: "203.0.113.45:2121"
|
||||
```
|
||||
|
||||
Replace `203.0.113.45:2121` with your actual public IP and
|
||||
port. The bare-IP form `external_addr: "203.0.113.45"` is also
|
||||
accepted; the daemon combines it with the bind port. `public:
|
||||
true` is still required as the master switch that gates UDP
|
||||
advertisement; setting `external_addr` alongside it wins, and
|
||||
STUN auto-discovery is skipped entirely (no logging
|
||||
cross-check).
|
||||
|
||||
`advertise_on_nostr: true` is the bit that says "include this
|
||||
transport in my published advert" — common to both paths.
|
||||
|
||||
Save the file.
|
||||
|
||||
## Step 3: Restart the daemon
|
||||
|
||||
```sh
|
||||
sudo systemctl restart fips
|
||||
sudo systemctl status fips
|
||||
```
|
||||
|
||||
Status should show `active (running)`. Within a few seconds the
|
||||
daemon will:
|
||||
|
||||
1. Determine the address to advertise. If you set `external_addr`,
|
||||
the daemon uses it directly and skips STUN. If you set only
|
||||
`public: true`, the daemon runs a one-shot STUN observation
|
||||
against the default STUN servers and uses the reflexive IPv4 it
|
||||
learns.
|
||||
2. Build a Kind 37195 advert listing
|
||||
`udp:<public-ip>:2121` (and any other transports you have
|
||||
`advertise_on_nostr: true` on).
|
||||
3. Sign the advert with the daemon's nsec.
|
||||
4. Publish it to the three default advert relays.
|
||||
5. Schedule a refresh every 30 minutes.
|
||||
|
||||
If you took the `public: true` path and STUN fails (for example,
|
||||
the network blocks outbound UDP/3478), the daemon emits a WARN
|
||||
line in the journal and suppresses the UDP entry from the advert
|
||||
rather than publishing a wrong address. The link to `test-us01`
|
||||
from the previous tutorial keeps working regardless — only the
|
||||
publish side is gated on STUN, and only on the STUN path. The
|
||||
`external_addr` path doesn't depend on STUN reachability at all.
|
||||
|
||||
Quick sanity check on the journal:
|
||||
|
||||
```sh
|
||||
sudo journalctl -u fips -n 200 | grep -iE 'STUN|advert|warn' | head -20
|
||||
```
|
||||
|
||||
If you see `WARN` lines mentioning STUN or wildcard-bind
|
||||
fallthrough, jump to [Troubleshooting](#troubleshooting); the
|
||||
rest of the tutorial assumes the publish succeeded.
|
||||
|
||||
## Step 4: Verify your advert is on the network
|
||||
|
||||
The advert is a public Nostr event — anyone, including you,
|
||||
can fetch it. With the `nak` Nostr CLI installed, query the
|
||||
relays for adverts published by your npub:
|
||||
|
||||
```sh
|
||||
nak req -k 37195 -d "fips-overlay-v1" \
|
||||
-a $(nak decode <your-npub> | jq -r .pubkey) \
|
||||
--limit 1 wss://relay.damus.io
|
||||
```
|
||||
|
||||
Replace `<your-npub>` with the npub you copied in Step 1. The
|
||||
inner `nak decode` converts your bech32 npub to the hex pubkey
|
||||
the relay filter expects.
|
||||
|
||||
Expect one event back. The interesting fields:
|
||||
|
||||
- `pubkey` — your npub in hex form.
|
||||
- `tags` — includes `["d","fips-overlay-v1"]` (the namespace),
|
||||
`["protocol","fips-overlay-v1"]`, and an `["expiration", …]`
|
||||
tag set ~1 hour in the future.
|
||||
- `content` — JSON listing the `endpoints` array. You should
|
||||
see one entry like:
|
||||
|
||||
```json
|
||||
{"transport":"udp","addr":"<your-public-ip>:2121"}
|
||||
```
|
||||
|
||||
That `<your-public-ip>` is what STUN learned. Confirm it
|
||||
matches what you'd expect for your network — for a home node,
|
||||
it should be your residential IP, not a `192.168.x.x` LAN
|
||||
address.
|
||||
|
||||
## Step 5: Watch for inbound connections
|
||||
|
||||
Your advert is now consumable by any FIPS daemon running open
|
||||
discovery on the same `fips-overlay-v1` namespace. The public
|
||||
test mesh nodes do exactly this — they subscribe to all
|
||||
adverts in the namespace and try to dial new publishers.
|
||||
|
||||
Within a minute or two of restart, run:
|
||||
|
||||
```sh
|
||||
sudo fipsctl show peers
|
||||
```
|
||||
|
||||
In addition to your configured `test-us01` peer, you may see
|
||||
an entry for `test-us03` (the open-discovery test mesh node).
|
||||
It will have `connectivity` active and its own
|
||||
`transport_addr`. This peering appeared without you
|
||||
configuring anything — the test-mesh open-discovery node saw
|
||||
your advert, dialed the endpoint, and Noise IK established
|
||||
the link.
|
||||
|
||||
If no inbound peers appear, that's not necessarily a failure
|
||||
of advertising — it just means no one has consumed your advert
|
||||
*and* dialed back yet. The advert is on the relays regardless,
|
||||
verifiable in Step 4.
|
||||
|
||||
## What you've learned
|
||||
|
||||
- **Adverts are publish + sign.** Every running FIPS daemon
|
||||
with `advertise: true` publishes a signed advert; reading it
|
||||
is one Nostr event lookup.
|
||||
- **Endpoint inclusion is per-transport.** Only the transports
|
||||
you set `advertise_on_nostr: true` on appear in the advert.
|
||||
- **`public: true` invokes STUN.** Wildcard-bound UDP with
|
||||
`public: true` runs a one-shot STUN observation to learn
|
||||
its public IP.
|
||||
- **Refresh is automatic.** Adverts re-publish every 30
|
||||
minutes; consumers cache them with a 1-hour staleness
|
||||
bound.
|
||||
- **The publish side stands alone.** Once your advert is on
|
||||
the relays, peers can dial you whether you're advertising
|
||||
to them specifically or not. The test mesh's open-discovery
|
||||
nodes will pick you up automatically.
|
||||
|
||||
## If your direct UDP advert isn't reachable
|
||||
|
||||
`public: true` advertises the IP STUN observes paired with your
|
||||
listener's bind port. That advert is only useful if your listener
|
||||
really is reachable at that public IP/port — STUN can confirm the
|
||||
public IP but not that an unsolicited inbound packet to the bind
|
||||
port will make it through. The most common cause of the listener
|
||||
being unreachable is symmetric NAT (where the public port a peer
|
||||
sees varies per remote endpoint), but other configurations can
|
||||
have the same effect.
|
||||
|
||||
When direct UDP advertising can't be relied on, the alternative
|
||||
is `udp:nat` mode, which advertises a placeholder `udp:nat`
|
||||
endpoint along with the daemon's signaling-relay and STUN-server
|
||||
lists, and performs UDP hole-punching at dial time. Hole-punching
|
||||
is best-effort — it works reliably when both sides are full-cone
|
||||
or port-restricted, and symmetric NAT on either side typically
|
||||
defeats it. Both sides need matching configs.
|
||||
|
||||
The minimal config switch:
|
||||
|
||||
```yaml
|
||||
transports:
|
||||
udp:
|
||||
bind_addr: "0.0.0.0:2121"
|
||||
advertise_on_nostr: true
|
||||
public: false # ← was true; change to false
|
||||
```
|
||||
|
||||
And add the signaling/STUN block under `discovery.nostr`:
|
||||
|
||||
```yaml
|
||||
discovery:
|
||||
nostr:
|
||||
enabled: true
|
||||
advertise: true
|
||||
dm_relays:
|
||||
- "wss://relay.damus.io"
|
||||
- "wss://nos.lol"
|
||||
stun_servers:
|
||||
- "stun:stun.l.google.com:19302"
|
||||
- "stun:stun.cloudflare.com:3478"
|
||||
```
|
||||
|
||||
For the full setup including peer-side config and the punch-
|
||||
duration knob, see
|
||||
[../how-to/enable-nostr-discovery.md § When the node is behind NAT](../how-to/enable-nostr-discovery.md#when-the-node-is-behind-nat).
|
||||
|
||||
Separately from NAT considerations, FIPS supports running a
|
||||
node behind a Tor onion service as a deployment shape in its
|
||||
own right — chosen for the privacy, anonymity, and
|
||||
censorship-resistance properties it brings, not as a fallback
|
||||
when UDP or TCP fail. If those properties are an independent
|
||||
goal for your node, see
|
||||
[../how-to/enable-nostr-discovery.md § Tor onion node](../how-to/enable-nostr-discovery.md#tor-onion-node)
|
||||
and
|
||||
[../how-to/deploy-tor-onion.md](../how-to/deploy-tor-onion.md).
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
If your advert doesn't appear on the relays:
|
||||
|
||||
- **STUN failed.** Check the journal for WARN lines mentioning
|
||||
STUN or wildcard-bind. The most common causes are outbound
|
||||
UDP/3478 blocked or DNS for `stun.l.google.com` failing.
|
||||
Try: `dig stun.l.google.com` and
|
||||
`nc -uvz stun.l.google.com 19302` to verify reachability.
|
||||
- **Wrong public IP advertised.** If `nak` shows your advert
|
||||
with a non-public address (e.g., `10.x.x.x` or
|
||||
`192.168.x.x`), STUN didn't see your real public IP — likely
|
||||
you're behind a CGNAT that NATs your STUN traffic too, or a
|
||||
corporate firewall that proxies it. Two correct fixes:
|
||||
(a) keep `public: true` and add `external_addr: <your-IP>`
|
||||
(the explicit override wins and skips STUN); or (b) bind
|
||||
directly to your public interface
|
||||
(`bind_addr: <pub-ip>:2121`) and keep `advertise_on_nostr:
|
||||
true` and `public: true`. Don't drop those flags.
|
||||
|
||||
- **Relay reachability.** `nak req` against a relay you can
|
||||
reach but no events return — possibly the publish failed
|
||||
silently because the daemon couldn't connect to that
|
||||
specific relay. Try the other two:
|
||||
|
||||
```sh
|
||||
nak req ... wss://nos.lol
|
||||
nak req ... wss://offchain.pub
|
||||
```
|
||||
|
||||
- **`advertise_on_nostr` typo.** YAML is case-sensitive. The
|
||||
config parser rejects unknown keys via
|
||||
`serde(deny_unknown_fields)` on the per-section structs, so a
|
||||
misspelled field will refuse the daemon's start with a
|
||||
parse-error line in the journal naming the unknown field.
|
||||
If the daemon is running but `nak` returns no advert, the
|
||||
field was accepted but something else is wrong; double-check
|
||||
the spelling on the UDP block and that
|
||||
`discovery.nostr.advertise: true` is also set.
|
||||
|
||||
## What's next
|
||||
|
||||
- **Open discovery.**
|
||||
[open-discovery](open-discovery.md) flips the consume side
|
||||
symmetric — switch your daemon to `policy: open` and watch
|
||||
your peer list populate from the ambient
|
||||
`fips-overlay-v1` namespace, the same mechanism
|
||||
`test-us03` is using right now to find you.
|
||||
|
||||
- **Host a service of your own.**
|
||||
[host-a-service](host-a-service.md) walks through bringing up
|
||||
an HTTP server addressable as `<your-npub>.fips`, the same
|
||||
way the connecting node now reaches `test-us01`. The natural
|
||||
follow-on now that other operators can dial you by npub.
|
||||
|
||||
For the operator-style scenario reference covering all five
|
||||
shapes of Nostr discovery side-by-side (consume-only,
|
||||
publish-direct, publish-Tor, NAT traversal, open):
|
||||
|
||||
- [../how-to/enable-nostr-discovery.md](../how-to/enable-nostr-discovery.md)
|
||||
|
||||
For the wire-format and discovery design:
|
||||
|
||||
- [../reference/nostr-events.md](../reference/nostr-events.md)
|
||||
— Kind 37195 advert format, Kind 21059 traversal signaling.
|
||||
- [../design/fips-nostr-discovery.md](../design/fips-nostr-discovery.md)
|
||||
— discovery runtime design, security and threat model.
|
||||
@@ -0,0 +1,480 @@
|
||||
# Deploy a `fips-gateway` on an OpenWrt AP
|
||||
|
||||
In every other tutorial in this set you put FIPS *on* the host that
|
||||
needs to talk to the mesh. This one is the exception. Here you stand
|
||||
up a `fips-gateway` on an OpenWrt access point so the unmodified LAN
|
||||
behind it — phones, laptops, smart-home gear — can reach mesh
|
||||
destinations by `<npub>.fips` without any FIPS software of their own,
|
||||
and so a service running on a LAN box can be exposed to mesh peers
|
||||
through a port forward. Two halves, one binary, one config.
|
||||
|
||||
This is **advanced** material. It assumes you have already worked
|
||||
through [join-the-test-mesh](join-the-test-mesh.md) on some other
|
||||
machine, you understand what `<npub>.fips` means, and now you want to
|
||||
fold an existing LAN into the mesh from the edge router rather than
|
||||
installing FIPS on every device. The whole exercise should take about
|
||||
forty-five minutes.
|
||||
|
||||
## What you'll build
|
||||
|
||||
```text
|
||||
┌──────────────────────────────────────────────────────────────┐
|
||||
│ OpenWrt access point │
|
||||
│ │
|
||||
│ br-lan ┌──────────────┐ fips0 ┌──────────────┐ │
|
||||
│ (LAN side) ◀──│ fips-gateway │──────────▶│ fips daemon │ │
|
||||
│ │ (service) │ │ │──┼─▶ mesh
|
||||
│ │ │ fd97:.. │ fd97:.. │ │
|
||||
│ │ fd01::/112 │ │ │ │
|
||||
│ │ pool │ │ │ │
|
||||
│ └──────┬───────┘ └──────────────┘ │
|
||||
│ │ │
|
||||
│ ▼ │
|
||||
│ nftables NAT │
|
||||
│ (inet fips_gateway) │
|
||||
└─────────────────┬────────────────────────────────────────────┘
|
||||
│ br-lan
|
||||
▼
|
||||
┌──────────────────────────────────────────────┐
|
||||
│ LAN clients (phones, laptops, smart-home) │
|
||||
│ │
|
||||
│ no FIPS install — just IPv6 + DNS │
|
||||
│ ▲ │
|
||||
│ │ curl http://test-us01.fips/ │
|
||||
│ └────── DNS to dnsmasq ──▶ gateway DNS │
|
||||
│ │
|
||||
└──────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
By the end you will have:
|
||||
|
||||
- An OpenWrt AP whose LAN clients can fetch `http://test-us01.fips/`
|
||||
with no FIPS software installed on them.
|
||||
- An inbound port forward exposing one LAN service to the mesh as
|
||||
`<your-gateway-npub>.fips:<port>`.
|
||||
- An understanding of which LAN-side glue OpenWrt automates for you
|
||||
(DNS forwarding, RA route, IPv6 prefix) and which the operator owns
|
||||
(port forwards, mesh firewall).
|
||||
|
||||
## Why an OpenWrt AP
|
||||
|
||||
The gateway has very specific dependencies on the box it runs on. It
|
||||
needs to own DNS for the LAN, it needs to advertise an IPv6 route to
|
||||
the LAN, it needs a stable LAN-side interface, and it needs to be the
|
||||
default IPv6 router for the segment. An OpenWrt-based access point
|
||||
already does all of those things — it runs `dnsmasq`, it runs
|
||||
`odhcpd` for IPv6 RA, it owns `br-lan`, and clients are already using
|
||||
it as their gateway. The OpenWrt ipk leans into that: the `gateway:`
|
||||
block in `/etc/fips/fips.yaml` is pre-populated, and the
|
||||
`/etc/init.d/fips-gateway` init script wires up the LAN-side glue
|
||||
automatically when you start the service.
|
||||
|
||||
On a non-OpenWrt host the same integration is manual; that path is
|
||||
covered by [../how-to/deploy-gateway.md](../how-to/deploy-gateway.md).
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- An OpenWrt 22.03+ AP serving DHCP and DNS to a wired or wireless
|
||||
LAN segment.
|
||||
- The FIPS ipk installed and a working FIPS daemon on the AP. If you
|
||||
haven't done that yet, follow
|
||||
[../../packaging/openwrt-ipk/README.md](../../packaging/openwrt-ipk/README.md)
|
||||
for the install, then come back here.
|
||||
- The AP joined to the mesh — at least one healthy peer link. If it
|
||||
isn't, work through [join-the-test-mesh](join-the-test-mesh.md) on
|
||||
the AP first.
|
||||
- Root SSH to the AP. Every command in this tutorial runs on the AP.
|
||||
- A LAN client (phone, laptop) to test the outbound half from.
|
||||
|
||||
## Step 1: Verify the FIPS daemon is up
|
||||
|
||||
Confirm the daemon is running, the TUN is up, and at least one mesh
|
||||
peer is reachable:
|
||||
|
||||
```sh
|
||||
service fips status
|
||||
ip -6 addr show fips0
|
||||
fipsctl show peers
|
||||
```
|
||||
|
||||
You should see:
|
||||
|
||||
- `service fips status` reports `running`.
|
||||
- `fips0` exists and has one `inet6 fd97:...` address. That is the
|
||||
AP's mesh-side identity.
|
||||
- `fipsctl show peers` lists at least one peer with active
|
||||
connectivity (not `idle` / not zero bytes).
|
||||
|
||||
Confirm the AP can resolve a known mesh node by name:
|
||||
|
||||
```sh
|
||||
ping6 -c 2 test-us01.fips
|
||||
```
|
||||
|
||||
If any of these fail, fix the daemon side first — the gateway is a
|
||||
separate service that runs alongside a working daemon, not a
|
||||
substitute for one.
|
||||
|
||||
## Step 2: Inspect the pre-populated gateway config
|
||||
|
||||
The OpenWrt ipk ships `/etc/fips/fips.yaml` with the `gateway:` block
|
||||
already filled in. View the relevant section:
|
||||
|
||||
```sh
|
||||
sed -n '/^gateway:/,$p' /etc/fips/fips.yaml
|
||||
```
|
||||
|
||||
You will see roughly:
|
||||
|
||||
```yaml
|
||||
gateway:
|
||||
enabled: true
|
||||
pool: "fd01::/112" # virtual IP range (up to 65535 addresses)
|
||||
lan_interface: "br-lan" # LAN-facing interface for proxy NDP
|
||||
dns:
|
||||
upstream: "[::1]:5354" # FIPS daemon DNS resolver (matches daemon default)
|
||||
ttl: 60 # DNS TTL and mapping lifetime (seconds)
|
||||
pool_grace_period: 60 # seconds after last session before reclaiming
|
||||
```
|
||||
|
||||
Three things to notice:
|
||||
|
||||
- `pool: "fd01::/112"` — the virtual-IP CIDR the gateway hands out
|
||||
to LAN clients. 65 536 addresses, the gateway's hard cap. Pick a
|
||||
different `fdXX::/N` prefix if `fd01::/112` collides with anything
|
||||
on your network.
|
||||
- `lan_interface: "br-lan"` — the OpenWrt LAN bridge. The gateway
|
||||
installs proxy-NDP entries on this interface so LAN clients can
|
||||
ARP-equivalent for pool addresses.
|
||||
- No `dns.listen` line — the source default `[::1]:5353` is exactly
|
||||
what OpenWrt wants. The gateway listens on IPv6 loopback only;
|
||||
dnsmasq, which owns LAN port 53, forwards `.fips` queries to it.
|
||||
The init script wires up that forwarding; you don't bind to a LAN
|
||||
address yourself.
|
||||
|
||||
For the full reference, see
|
||||
[../reference/configuration.md § Gateway](../reference/configuration.md#gateway-gateway).
|
||||
|
||||
## Step 3: Enable and start fips-gateway
|
||||
|
||||
The service is shipped disabled — enable it once and start it:
|
||||
|
||||
```sh
|
||||
service fips-gateway enable
|
||||
service fips-gateway start
|
||||
```
|
||||
|
||||
Behind that single command, the init script
|
||||
(`/etc/init.d/fips-gateway`) does five things:
|
||||
|
||||
1. **Loads gateway sysctls.** `net.ipv6.conf.all.proxy_ndp=1` and
|
||||
`net.ipv6.conf.all.forwarding=1` from
|
||||
`/etc/sysctl.d/fips-gateway.conf`.
|
||||
2. **Reconfigures dnsmasq via UCI** so `.fips` queries arriving at
|
||||
the LAN's port 53 are forwarded to the gateway's loopback
|
||||
listener on port 5353 instead of going straight to the daemon's
|
||||
resolver on port 5354. (Dnsmasq still owns 53; the gateway sits
|
||||
in front of the daemon for `.fips` only.)
|
||||
3. **Adds a global-scope IPv6 prefix** to `br-lan`. Without a
|
||||
non-ULA address on the local interface, Android and Chrome
|
||||
suppress AAAA queries entirely — they assume the LAN has no
|
||||
real IPv6 and don't bother. The init script adds a small
|
||||
benchmarking-range prefix to convince them otherwise.
|
||||
4. **Adds an RA route for the virtual pool.** A UCI `route6` entry
|
||||
under `dhcp` tells `odhcpd` to advertise the pool CIDR via Router
|
||||
Advertisement (RFC 4191), so LAN clients learn how to reach pool
|
||||
addresses automatically. No per-client static routes needed.
|
||||
5. **Spawns `fips-gateway` under procd** with `--config
|
||||
/etc/fips/fips.yaml`, with crash-respawn.
|
||||
|
||||
Verify it is running:
|
||||
|
||||
```sh
|
||||
service fips-gateway status
|
||||
logread | grep fips-gateway | tail
|
||||
```
|
||||
|
||||
Expect a `running` status and a startup log line of the form
|
||||
`fips-gateway 0.x.y starting`, followed by entries for DNS bind, NAT
|
||||
table install, and pool initialisation.
|
||||
|
||||
> **What just changed on the LAN.** The AP is now offering two
|
||||
> things it wasn't offering a moment ago: AAAA records under `.fips`
|
||||
> that resolve to virtual IPs in `fd01::/112`, and a route to that
|
||||
> CIDR in its Router Advertisements. Existing LAN clients pick both
|
||||
> up the next time they re-resolve a name and the next time `odhcpd`
|
||||
> sends an RA, respectively. No reboot required on the client side.
|
||||
|
||||
## Step 4: Test the outbound half from a LAN client
|
||||
|
||||
Before bringing a LAN client into the picture, confirm from the AP
|
||||
itself that the mesh side is still healthy after the gateway start:
|
||||
|
||||
```sh
|
||||
ping6 -c 2 test-us01.fips
|
||||
```
|
||||
|
||||
This isolates the router-to-mesh path before involving the LAN
|
||||
segment. If this fails, the troubleshooting target is the daemon /
|
||||
mesh side, not the gateway-to-client side. If it succeeds and the
|
||||
LAN-client test below fails, the target is the LAN segment —
|
||||
`proxy_ndp`, the RA pool route, or DNS forwarding through dnsmasq.
|
||||
|
||||
Now from a phone or laptop on the AP's LAN — anything that does IPv6
|
||||
and DNS, with no FIPS software installed — try one of the public test
|
||||
mesh nodes:
|
||||
|
||||
```sh
|
||||
dig test-us01.fips AAAA
|
||||
ping6 -c 4 test-us01.fips
|
||||
curl -6 http://test-us01.fips/
|
||||
```
|
||||
|
||||
Expectations:
|
||||
|
||||
- `dig` returns an AAAA in `fd01::...`, **not** `fd97:...`. The
|
||||
`fd01:` address is the gateway's virtual-IP allocation; the LAN
|
||||
client never sees the raw mesh address.
|
||||
- `ping6` succeeds. ICMPv6 echo travels through the NAT pipeline and
|
||||
back.
|
||||
- `curl` fetches the page (whatever the test mesh is currently
|
||||
serving on `test-us01`).
|
||||
|
||||
> **What just happened end to end.** Your client asked dnsmasq for
|
||||
> `test-us01.fips`. Dnsmasq forwarded the query to the gateway's
|
||||
> loopback listener on port 5353. The gateway forwarded the query on
|
||||
> to the daemon's resolver on port 5354. The daemon answered with
|
||||
> `test-us01`'s mesh address (`fd97:...`). The gateway allocated a
|
||||
> virtual IP from `fd01::/112`, installed nftables DNAT/SNAT/
|
||||
> masquerade rules pinning that virtual IP to the mesh address,
|
||||
> installed a proxy-NDP entry on `br-lan` so the client could resolve
|
||||
> the virtual IP at the link layer, and returned the virtual IP in
|
||||
> the AAAA reply. Your client then routed traffic to the virtual IP
|
||||
> via the RA-advertised pool route, the AP's kernel rewrote the
|
||||
> destination to the mesh address, and the daemon's adapter carried
|
||||
> the packets across the mesh. Return traffic followed conntrack
|
||||
> back. The client never knew the mesh existed.
|
||||
|
||||
## Step 5 (optional): Inspect the gateway state
|
||||
|
||||
The gateway exposes its own control socket separate from the daemon's.
|
||||
Two useful queries:
|
||||
|
||||
```sh
|
||||
echo '{"command":"show_gateway"}' | nc -U /run/fips/gateway.sock
|
||||
echo '{"command":"show_mappings"}' | nc -U /run/fips/gateway.sock
|
||||
```
|
||||
|
||||
`show_gateway` reports pool utilisation, the DNS listen address,
|
||||
uptime, and the conntrack/NAT counters. `show_mappings` lists each
|
||||
allocated virtual IP, the mesh address it points at, the DNS name
|
||||
that triggered the allocation, and the mapping's lifecycle state
|
||||
(`Allocated`, `Active`, `Draining`).
|
||||
|
||||
For the full command catalog and JSON shapes, see
|
||||
[../reference/control-socket.md § Gateway Command Catalog](../reference/control-socket.md#gateway-command-catalog).
|
||||
The same data is rendered visually in the **Gateway** tab of
|
||||
[`fipstop`](../reference/cli-fipstop.md).
|
||||
|
||||
If you want to see the kernel rules the gateway installed:
|
||||
|
||||
```sh
|
||||
nft list table inet fips_gateway
|
||||
```
|
||||
|
||||
You will see DNAT, SNAT, and masquerade chains populated with one
|
||||
rule per active mapping.
|
||||
|
||||
## Step 6 (Optional): Add an inbound port-forward for a LAN service
|
||||
|
||||
The outbound half is the steady-state use of a gateway. The inbound
|
||||
half — exposing a LAN service to mesh peers — is a separate decision,
|
||||
configured per service under `gateway.port_forwards[]`.
|
||||
|
||||
For the worked example, run a one-page static web server on the AP
|
||||
itself, bound to its `br-lan` address, and expose it to the mesh
|
||||
through a port-forward. Anything would do — the point of the exercise
|
||||
is the port-forward, not the service. We use what is already on the
|
||||
AP: `busybox httpd`. In a real deployment the LAN-side target would
|
||||
typically be a separate host (a NAS, a home server, a dev box on the
|
||||
LAN); the rule shape is identical.
|
||||
|
||||
Find the AP's `br-lan` IPv6 address and save it for the rest of the
|
||||
step:
|
||||
|
||||
```sh
|
||||
BR_LAN_ADDR=$(ip -6 addr show br-lan \
|
||||
| awk '/inet6 fd|inet6 2/ && !/scope link/ {print $2}' \
|
||||
| head -1 | cut -d/ -f1)
|
||||
echo "$BR_LAN_ADDR"
|
||||
```
|
||||
|
||||
Pick from the global-scope benchmarking prefix the init script added
|
||||
in Step 3, or your own ULA if `br-lan` has one — anything except a
|
||||
link-local `fe80::/10` address.
|
||||
|
||||
Set up a one-file docroot and start a foreground `busybox httpd`
|
||||
bound to that LAN address on port 8000:
|
||||
|
||||
```sh
|
||||
mkdir -p /tmp/mesh-demo
|
||||
echo '<h1>Hello from the mesh-gateway demo</h1>' > /tmp/mesh-demo/index.html
|
||||
busybox httpd -f -p "[${BR_LAN_ADDR}]:8000" -h /tmp/mesh-demo
|
||||
```
|
||||
|
||||
Leave it running in this shell. Open a second SSH session on the AP
|
||||
to add the port-forward.
|
||||
|
||||
Edit `/etc/fips/fips.yaml`. Inside the existing `gateway:` block,
|
||||
add a `port_forwards:` list:
|
||||
|
||||
```yaml
|
||||
gateway:
|
||||
enabled: true
|
||||
pool: "fd01::/112"
|
||||
lan_interface: "br-lan"
|
||||
dns:
|
||||
upstream: "[::1]:5354"
|
||||
ttl: 60
|
||||
pool_grace_period: 60
|
||||
port_forwards:
|
||||
- listen_port: 8080
|
||||
proto: tcp
|
||||
target: "[<BR_LAN_ADDR>]:8000"
|
||||
```
|
||||
|
||||
Substitute the real address for `<BR_LAN_ADDR>`. The IPv6 form
|
||||
(`[addr]:port`) is required — IPv4 targets are rejected at config
|
||||
load.
|
||||
|
||||
Restart the gateway so it re-reads the config:
|
||||
|
||||
```sh
|
||||
service fips-gateway restart
|
||||
```
|
||||
|
||||
From any *other* mesh node, fetch the demo page through the gateway
|
||||
using the AP's npub:
|
||||
|
||||
```sh
|
||||
# on the AP, get the npub:
|
||||
NPUB=$(cat /etc/fips/fips.pub)
|
||||
echo "$NPUB"
|
||||
```
|
||||
|
||||
Then on the remote mesh node:
|
||||
|
||||
```sh
|
||||
curl -6 "http://${NPUB}.fips:8080/"
|
||||
```
|
||||
|
||||
Expect:
|
||||
|
||||
```text
|
||||
<h1>Hello from the mesh-gateway demo</h1>
|
||||
```
|
||||
|
||||
The connection landed on the gateway's `fips0` ingress on TCP/8080,
|
||||
nftables DNAT rewrote the destination to `[BR_LAN_ADDR]:8000`,
|
||||
LAN-side masquerade rewrote the source so `busybox httpd` saw a
|
||||
LAN-routable address, and the response retraced via conntrack.
|
||||
|
||||
> **What you exposed.** With the port-forward active, every mesh
|
||||
> peer that can route to your AP can hit `${NPUB}.fips:8080/` and
|
||||
> reach this service. That is exactly what the inbound half is
|
||||
> for — but if you want to scope visibility to a specific subset
|
||||
> of peers, the FIPS mesh firewall is the layer that does it; see
|
||||
> [../how-to/enable-mesh-firewall.md](../how-to/enable-mesh-firewall.md).
|
||||
> The port-forward rule and the firewall rule are independent: the
|
||||
> port-forward installs the rewrite; the firewall decides who is
|
||||
> allowed to reach the listen port.
|
||||
|
||||
## Step 7: Tidy up
|
||||
|
||||
In the first shell, stop `busybox httpd` with `Ctrl-C`. The demo
|
||||
docroot at `/tmp/mesh-demo` can stay — it is wiped on reboot — or
|
||||
remove it now (`rm -rf /tmp/mesh-demo`).
|
||||
|
||||
If you want to keep the outbound half but withdraw the inbound
|
||||
forward, remove the `port_forwards:` entry from `/etc/fips/fips.yaml`
|
||||
and `service fips-gateway restart`. The mesh-side listener disappears
|
||||
and so does the corresponding nftables rule.
|
||||
|
||||
To turn the gateway off entirely:
|
||||
|
||||
```sh
|
||||
service fips-gateway stop
|
||||
service fips-gateway disable
|
||||
```
|
||||
|
||||
The init script's `stop_service` handler reverses the LAN-side
|
||||
integration on the way out: dnsmasq's `.fips` forwarder is pointed
|
||||
back at the daemon's port 5354, the RA route for the pool is
|
||||
withdrawn from `odhcpd`, and the global-scope IPv6 prefix on
|
||||
`br-lan` is removed. The LAN reverts to the state it was in before
|
||||
you ran `service fips-gateway start` in Step 3.
|
||||
|
||||
The daemon and the rest of `/etc/fips/` are untouched. Existing mesh
|
||||
peering on the AP itself continues to work.
|
||||
|
||||
## What you've learned
|
||||
|
||||
- **The gateway is a niche feature for a niche box.** Most FIPS
|
||||
hosts run the daemon and reach the mesh directly. The gateway
|
||||
exists so an AP can fold an entire unmodified LAN behind it into
|
||||
the mesh in one place.
|
||||
- **Two halves of the same binary.** Outbound mode hands LAN clients
|
||||
virtual IPs and NATs them onto the mesh; inbound mode listens on
|
||||
`fips0` and forwards to LAN targets. They share one nftables
|
||||
table, one control socket, and one config block, but each half
|
||||
has its own use case.
|
||||
- **OpenWrt does the LAN-side glue for you.** The init script
|
||||
reconfigures dnsmasq, installs the RA route, adds the global IPv6
|
||||
prefix, and loads sysctls. On a non-OpenWrt host that integration
|
||||
is manual — see [../how-to/deploy-gateway.md](../how-to/deploy-gateway.md).
|
||||
- **Inbound forwards stay manual on every distro.** The
|
||||
`port_forwards[]` block is uniform across hosts, and on every
|
||||
distro you still own the decision of which LAN target to expose
|
||||
and on which mesh-side port.
|
||||
- **The mesh firewall is a separate decision.** Opening a port
|
||||
forward on the gateway side does not open it on the firewall
|
||||
side; if `fips-firewall.service` is enabled, you still need a
|
||||
drop-in that admits the listen port.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
If something doesn't work as described above, the operator-recipe
|
||||
guide [../how-to/troubleshoot-gateway.md](../how-to/troubleshoot-gateway.md)
|
||||
groups the common failures by symptom:
|
||||
|
||||
| Symptom | Where to look |
|
||||
| ------- | ------------- |
|
||||
| LAN client gets `fd97:...`, not `fd01:...` | DNS path: dnsmasq still pointing at port 5354. See "DNS queries fail". |
|
||||
| `dig` succeeds with a pool address but `ping6` times out | Pool route or proxy NDP. See "Virtual IP unreachable from client". |
|
||||
| `ping6` works but TCP times out | NAT pipeline or mesh-side firewall. See "Ping works but TCP does not". |
|
||||
| Gateway service won't start | "No gateway section in configuration" recipe. |
|
||||
| Inbound `curl` hits the listen port but never reaches the LAN target | Mesh-side firewall first, then the port-forward rule. |
|
||||
|
||||
The first thing the troubleshoot guide does in any of these cases is
|
||||
ask the gateway directly via `show_gateway` and `show_mappings`. If
|
||||
the mapping you expect is not there, the failure is on the DNS path;
|
||||
if it is there in `state: Active` but traffic still fails, the
|
||||
failure is downstream.
|
||||
|
||||
## What's next
|
||||
|
||||
- [../how-to/deploy-gateway.md](../how-to/deploy-gateway.md) —
|
||||
Manual deployment on a non-OpenWrt Linux host. Same gateway,
|
||||
same config, but you wire up dnsmasq/Unbound/etc. yourself,
|
||||
install a pool route per LAN client (or via your own RA daemon),
|
||||
and manage the systemd unit instead of the procd init script.
|
||||
- [../design/fips-gateway.md](../design/fips-gateway.md) — The
|
||||
design doc: NAT pipeline (DNAT, SNAT, masquerade, inbound DNAT),
|
||||
virtual-IP pool lifecycle (Allocated -> Active -> Draining ->
|
||||
reclaimed), DNS resolution flow, conntrack integration.
|
||||
- [../reference/configuration.md § Gateway](../reference/configuration.md#gateway-gateway)
|
||||
— Every field of the `gateway:` block, including the conntrack
|
||||
timeout overrides not used in this tutorial.
|
||||
- [../reference/cli-fips-gateway.md](../reference/cli-fips-gateway.md)
|
||||
— The `fips-gateway` binary's CLI options, exit codes, and
|
||||
environment variables.
|
||||
@@ -0,0 +1,518 @@
|
||||
# Build a Mesh from the Ground Up
|
||||
|
||||
The earlier tutorials in this progression rode existing IP — your
|
||||
daemon reached `test-us01` over the public internet through your
|
||||
ISP, your ISP's upstream, and however many hops separate you from
|
||||
the test node. That is the *overlay* deployment mode of FIPS:
|
||||
useful, but not the new ground.
|
||||
|
||||
This tutorial is about the other mode. Two devices, a wire (or a
|
||||
radio link) between them, no IP between them, and FIPS daemons on
|
||||
each end. The two daemons discover each other over the raw link,
|
||||
peer over Noise, and bring up an end-to-end mesh with addressing,
|
||||
naming, and reachability — all from layer 2 up. There is no DHCP,
|
||||
no router, no upstream. The mesh is the network.
|
||||
|
||||
This is the deployment mode FIPS was designed for. Overlay mode
|
||||
exists because riding existing IP is a useful convenience; the
|
||||
ground-up mode is what FIPS uniquely enables.
|
||||
|
||||
> **The two modes are not exclusive.** A node can carry overlay
|
||||
> peers and ground-up peers at the same time — different transports
|
||||
> on the same daemon. If you have already worked through
|
||||
> [join-the-test-mesh](join-the-test-mesh.md), the static peer to
|
||||
> `test-us01` you configured there can stay in place; the Ethernet
|
||||
> peer you add in this tutorial sits alongside it. Traffic flows
|
||||
> through whichever path is shortest by mesh metric, and a node on
|
||||
> one side can reach a node on the other through your machine
|
||||
> acting as a bridge between the two.
|
||||
|
||||
## What you'll build
|
||||
|
||||
```text
|
||||
┌──────────────────────┐ raw Ethernet frames ┌──────────────────────┐
|
||||
│ node A │ ─────────────────────── │ node B │
|
||||
│ npub1aaa… │ EtherType 0x2121 │ npub1bbb… │
|
||||
│ fips0 fd97:..:A │ no IP between them │ fips0 fd97:..:B │
|
||||
└──────────────────────┘ └──────────────────────┘
|
||||
│ │
|
||||
│ a single Ethernet cable │
|
||||
│ (or both NICs on the same │
|
||||
│ unmanaged switch — no DHCP, │
|
||||
│ no router, no IP at all) │
|
||||
└─────────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
Two machines, each running `fips`, joined by a physical Ethernet
|
||||
link. After the worked example:
|
||||
|
||||
- The two daemons have discovered each other via L2 beacons on
|
||||
the link, peered over Noise IK, and brought up an FMP link.
|
||||
- Each `fips0` adapter has a routable mesh address; each can
|
||||
ping the other by `<npub>.fips`.
|
||||
- Nothing between the two machines speaks IP. The link carries
|
||||
raw FIPS frames at EtherType `0x2121`.
|
||||
|
||||
The whole exercise should take about twenty minutes if you have
|
||||
the hardware ready.
|
||||
|
||||
## Why ground-up
|
||||
|
||||
Most networking tutorials assume IP is already there: an address
|
||||
arrived from DHCP, a default gateway routes you onward, DNS
|
||||
resolves names. FIPS does not need any of that. Two devices and
|
||||
a way to deliver bytes between them at layer 2 is enough — FIPS
|
||||
supplies the rest:
|
||||
|
||||
- **Identity**: each daemon has an npub (the same kind you saw
|
||||
in the overlay tutorials). Nothing in the ground-up case
|
||||
depends on a network identity from a router; the npub is the
|
||||
identity.
|
||||
- **Addressing**: the `fips0` adapter takes an `fd97:...` ULA
|
||||
derived from the npub. No DHCP. No SLAAC. The address is
|
||||
cryptographically tied to the identity.
|
||||
- **Discovery**: each daemon broadcasts a small beacon on the
|
||||
link advertising its npub; the other daemon's listener picks
|
||||
it up and dials in over the same link.
|
||||
- **Routing**: the FIPS mesh layer builds its own spanning tree
|
||||
across whatever links it has. Add a third node (peered to
|
||||
either A or B) and traffic reaches it transparently.
|
||||
|
||||
The point is not that ground-up replaces overlay. It's that
|
||||
overlay is one of two modes the same daemon supports, and
|
||||
ground-up is what unlocks the use cases overlay cannot —
|
||||
ad-hoc local meshes, partitioned networks, situations where
|
||||
no IP infrastructure exists or can be relied on.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
Two devices (call them **node A** and **node B**) and a way to
|
||||
join them at layer 2:
|
||||
|
||||
- Ethernet (the worked example): a direct cable between two
|
||||
modern NICs (auto-MDI/MDIX handles crossover for you), or
|
||||
both machines on a small unmanaged switch with no DHCP
|
||||
server. USB-Ethernet dongles work; a typical "USB-to-RJ45"
|
||||
adapter is fine on either end. The link does **not** need
|
||||
to be the machine's primary network interface — a second
|
||||
NIC dedicated to the mesh is the cleanest setup.
|
||||
- WiFi (a one-line variation, covered later): both machines
|
||||
associated to a common AP that has client (station)
|
||||
isolation **off**.
|
||||
- Bluetooth LE (a separate worked example via a how-to,
|
||||
covered later): two BLE-capable Linux hosts within roughly
|
||||
10 metres line of sight.
|
||||
|
||||
On both nodes:
|
||||
|
||||
- `fips` installed and running, per [getting-started](../getting-started.md).
|
||||
- A persistent identity from
|
||||
[persistent-identity](persistent-identity.md). Ephemeral
|
||||
identities work, but on each restart the npub regenerates
|
||||
and you'll have to re-check `fipsctl show peers` to see the
|
||||
new identity. Persistent makes the lesson stick.
|
||||
- The daemon running with `CAP_NET_RAW` (the shipped systemd
|
||||
unit runs as root and gets this for free; running
|
||||
interactively from a user account requires `setcap` —
|
||||
noted at the relevant step below).
|
||||
|
||||
You do **not** need:
|
||||
|
||||
- An IP address on the chosen interface. The Ethernet
|
||||
transport opens a raw socket directly; the kernel does not
|
||||
need to assign an IP to the NIC.
|
||||
- A default route. The mesh routes itself.
|
||||
- DNS resolution between the machines via any external
|
||||
service. The local `.fips` resolver supplies names from
|
||||
the npubs the daemons exchange.
|
||||
|
||||
## Step 1: Identify the link interface on each node
|
||||
|
||||
On each node, list the network interfaces and pick the one that
|
||||
sits on the link between the two machines. If it's a dedicated
|
||||
NIC for the mesh, that NIC has no other purpose; if it's a
|
||||
USB-Ethernet dongle, plug it in first so the kernel names it.
|
||||
|
||||
```sh
|
||||
ip link show
|
||||
```
|
||||
|
||||
Pick out the interface name. Common forms:
|
||||
|
||||
- `enp3s0`, `eno1` — built-in NICs under predictable naming.
|
||||
- `eth0` — older or container-style naming.
|
||||
- `enxAABBCCDDEEFF` — USB-Ethernet dongles often appear under
|
||||
this MAC-derived form.
|
||||
|
||||
Bring the interface up if it isn't:
|
||||
|
||||
```sh
|
||||
sudo ip link set dev <interface> up
|
||||
```
|
||||
|
||||
Confirm:
|
||||
|
||||
```sh
|
||||
ip -br link show <interface>
|
||||
```
|
||||
|
||||
You want `UP` and `LOWER_UP` in the flags. The interface does
|
||||
not need an IP address — `LOWER_UP` indicates the NIC sees
|
||||
carrier (cable plugged into something at the other end), and
|
||||
that is all the Ethernet transport needs.
|
||||
|
||||
For the rest of the tutorial we'll write the chosen interface
|
||||
as `<eth>`. Substitute the actual name on each node when you
|
||||
run the commands. Note that node A and node B may have
|
||||
different interface names — that is normal.
|
||||
|
||||
> **No IP needed.** If your chosen interface has an address
|
||||
> from a previous DHCP lease, leave it alone or remove it with
|
||||
> `sudo ip addr flush dev <eth>` — the FIPS Ethernet transport
|
||||
> uses raw `AF_PACKET` sockets that bypass the IP stack
|
||||
> entirely. The interface needs to be `up` and `LOWER_UP`,
|
||||
> nothing more.
|
||||
|
||||
## Step 2: Configure the Ethernet transport on each node
|
||||
|
||||
Edit `/etc/fips/fips.yaml` on **both** nodes. Under
|
||||
`transports:`, add an `ethernet:` block. The key settings are
|
||||
the four discovery flags — both nodes must opt in to all four,
|
||||
and they default to off:
|
||||
|
||||
```yaml
|
||||
transports:
|
||||
ethernet:
|
||||
interface: "<eth>" # the name from Step 1
|
||||
announce: true # broadcast our beacon on the link
|
||||
discovery: true # listen for beacons (default; shown for clarity)
|
||||
auto_connect: true # dial peers we discover
|
||||
accept_connections: true # accept dial-ins from peers we discover
|
||||
```
|
||||
|
||||
Each flag does one thing:
|
||||
|
||||
- `announce: true` — emit a small beacon every
|
||||
`beacon_interval_secs` (default 30s) carrying our npub.
|
||||
- `discovery: true` — listen for incoming beacons; populate a
|
||||
candidate-peer list keyed by source MAC and observed npub.
|
||||
- `auto_connect: true` — when we see a beacon from an npub
|
||||
we have not yet peered with, initiate the outbound Noise
|
||||
handshake.
|
||||
- `accept_connections: true` — when a remote npub initiates
|
||||
the handshake on this transport, complete it.
|
||||
|
||||
If only one node sets `announce`, the other won't see it; if
|
||||
only one side sets `auto_connect` or `accept_connections`, the
|
||||
roles are asymmetric and the link won't establish unless both
|
||||
are configured. The cleanest pattern for a ground-up tutorial
|
||||
is "all four flags on both ends."
|
||||
|
||||
> **Multiple Ethernet links.** If a node has more than one
|
||||
> physical interface that participates in the mesh, configure
|
||||
> each one as a *named instance* under `ethernet:`:
|
||||
>
|
||||
> ```yaml
|
||||
> transports:
|
||||
> ethernet:
|
||||
> lan:
|
||||
> interface: "eth0"
|
||||
> announce: true
|
||||
> discovery: true
|
||||
> auto_connect: true
|
||||
> accept_connections: true
|
||||
> dongle:
|
||||
> interface: "enx00aabbccddee"
|
||||
> announce: true
|
||||
> # ...
|
||||
> ```
|
||||
>
|
||||
> Each named instance runs its own socket and discovery state.
|
||||
> A single ground-up link only needs the flat form shown
|
||||
> first; named instances become useful when the same node
|
||||
> bridges multiple physical segments.
|
||||
|
||||
## Step 3: Grant the daemon permission to open raw sockets
|
||||
|
||||
The Ethernet transport opens an `AF_PACKET` `SOCK_DGRAM` socket
|
||||
bound to the chosen interface. That requires `CAP_NET_RAW`.
|
||||
|
||||
If you installed FIPS via the Debian package and run via the
|
||||
shipped systemd unit, the daemon runs as root and has
|
||||
`CAP_NET_RAW` already — there is nothing to do here. Skip to
|
||||
Step 4.
|
||||
|
||||
If you are running the daemon interactively as your user (a
|
||||
from-source / development setup), grant the capability once on
|
||||
the binary:
|
||||
|
||||
```sh
|
||||
sudo setcap CAP_NET_RAW,CAP_NET_ADMIN+ep "$(which fips)"
|
||||
```
|
||||
|
||||
`CAP_NET_ADMIN` is what the daemon needs for the `fips0` TUN
|
||||
adapter regardless; `CAP_NET_RAW` is the ground-up addition.
|
||||
The `setcap` invocation only needs to be repeated when the
|
||||
binary is replaced.
|
||||
|
||||
## Step 4: Restart the daemon on each node
|
||||
|
||||
```sh
|
||||
sudo systemctl restart fips
|
||||
```
|
||||
|
||||
Or, if running interactively, restart your `fips` invocation
|
||||
in whichever way you started it.
|
||||
|
||||
Watch the startup logs for the Ethernet transport coming up:
|
||||
|
||||
```sh
|
||||
sudo journalctl -u fips -f --since="1 minute ago"
|
||||
```
|
||||
|
||||
Look for landmarks like:
|
||||
|
||||
- A line indicating the Ethernet transport opened the chosen
|
||||
interface and started its receive loop.
|
||||
- Periodic outbound beacon messages (one per
|
||||
`beacon_interval_secs` window).
|
||||
- After the second beacon round on the *other* node, an
|
||||
inbound beacon parsed and a candidate-peer entry created.
|
||||
- Once each side dials, a Noise handshake completion log
|
||||
message naming the remote npub.
|
||||
|
||||
Beacon interval defaults to 30s, so the first peering can take
|
||||
up to a minute (one beacon window per side, plus handshake).
|
||||
Lower the interval for the tutorial if you want faster
|
||||
feedback:
|
||||
|
||||
```yaml
|
||||
transports:
|
||||
ethernet:
|
||||
# ...
|
||||
beacon_interval_secs: 10 # minimum allowed
|
||||
```
|
||||
|
||||
## Step 5: Verify the link
|
||||
|
||||
On either node:
|
||||
|
||||
```sh
|
||||
sudo fipsctl show peers
|
||||
```
|
||||
|
||||
Expect one entry whose `npub` matches the **other** node and
|
||||
whose `addresses` line shows `transport: ethernet`. Your
|
||||
existing overlay peers (if any from earlier tutorials) appear
|
||||
alongside it. Each peer has its own row, and the link status
|
||||
columns show whether the Noise session is up.
|
||||
|
||||
```sh
|
||||
sudo fipsctl show transports
|
||||
```
|
||||
|
||||
Confirms that the Ethernet transport is running and shows the
|
||||
beacon counters incrementing. Both `beacons_sent` and
|
||||
`beacons_received` should be non-zero if the link is healthy.
|
||||
|
||||
## Step 6: Reach the other node by name
|
||||
|
||||
On node A, ping node B by `.fips` name. Get node B's npub
|
||||
from its `fipsctl show status` output (it's the persistent
|
||||
identity you established earlier), then:
|
||||
|
||||
```sh
|
||||
ping6 npub1bbb…long-string….fips
|
||||
```
|
||||
|
||||
Expect ICMPv6 echo replies. The path is:
|
||||
|
||||
1. The local `.fips` resolver translates the npub-form name
|
||||
into an `fd97:...` mesh address (cryptographically derived
|
||||
from the npub on both ends — the resolver does the
|
||||
computation locally, with no network round trip).
|
||||
2. The kernel routes the packet via `fips0`.
|
||||
3. The FIPS daemon accepts it from the TUN, looks up the
|
||||
mesh route, and hands it to the FMP link to node B.
|
||||
4. The Ethernet transport on node A frames the FMP packet as
|
||||
a raw EtherType `0x2121` Ethernet frame addressed to node
|
||||
B's MAC, learned from B's beacons.
|
||||
5. Node B's daemon receives the frame, peels off the
|
||||
Ethernet/FIPS framing, and the packet emerges on node B's
|
||||
`fips0`.
|
||||
6. The kernel on node B sees an inbound ICMPv6 echo and
|
||||
replies, and the same path runs in reverse.
|
||||
|
||||
If you have a hosts file with shortnames configured (see
|
||||
[host-aliases](../how-to/host-aliases.md)), substitute the
|
||||
shortname for the full npub form.
|
||||
|
||||
## Step 7: Try a forward composition
|
||||
|
||||
If node A also has the `test-us01` overlay peer from
|
||||
[join-the-test-mesh](join-the-test-mesh.md), node B can
|
||||
reach `test-us01` *through* node A — even though node B has
|
||||
no direct internet path of its own:
|
||||
|
||||
On node B:
|
||||
|
||||
```sh
|
||||
ping6 npub1qmc3cvfz0yu2hx96nq3gp55zdan2qclealn7xshgr448d3nh6lks7zel98.fips
|
||||
```
|
||||
|
||||
The packet leaves B's `fips0`, traverses the Ethernet link to
|
||||
A, gets forwarded by A across the overlay UDP transport to
|
||||
`test-us01`, and the reply comes back the same way.
|
||||
|
||||
This is the composition the chapter intro flagged: the two
|
||||
deployment modes coexist on a single daemon. Node A is
|
||||
participating in the test mesh via the internet *and* in your
|
||||
local Ethernet mesh. From node B's perspective, the test mesh
|
||||
is reachable. From `test-us01`'s perspective, B is reachable.
|
||||
The mesh handles the rest.
|
||||
|
||||
## Variations
|
||||
|
||||
### WiFi (AP mode), same shape as Ethernet
|
||||
|
||||
Replace `<eth>` with the WiFi interface name (typically
|
||||
`wlan0` or `wlp3s0`) on each node. The WiFi NIC is presented
|
||||
as an Ethernet-class interface to the kernel by the
|
||||
`mac80211` abstraction; the FIPS Ethernet transport opens
|
||||
the same `AF_PACKET` socket on it. No FIPS-side configuration
|
||||
change beyond the interface name.
|
||||
|
||||
What you do need on the AP side:
|
||||
|
||||
- Both nodes associated to the same SSID.
|
||||
- **Client (station) isolation must be OFF** on the AP.
|
||||
Most consumer routers ship with it off; many guest
|
||||
networks and "secure" enterprise APs ship with it on.
|
||||
When client isolation is on, the AP refuses to forward
|
||||
station-to-station frames — the broadcast beacons never
|
||||
arrive at the other node, and discovery fails silently.
|
||||
If beacons aren't crossing, this is the first thing to
|
||||
check.
|
||||
|
||||
There is no FIPS-specific configuration for WiFi versus
|
||||
Ethernet on the daemon side; the choice is purely the
|
||||
adapter name.
|
||||
|
||||
### Bluetooth LE (experimental but works)
|
||||
|
||||
BLE is a separate transport (`transports.ble.*`) with its own
|
||||
discovery model — L2CAP advertisements rather than raw L2
|
||||
broadcasts. The shape of the tutorial is the same (advertise +
|
||||
scan + auto-connect + accept), but the prerequisites are
|
||||
different: BlueZ, `bluetoothd`, an HCI adapter, and the
|
||||
`bluetooth` group or capability set.
|
||||
|
||||
The full operator recipe is in
|
||||
[../how-to/set-up-bluetooth-peer.md](../how-to/set-up-bluetooth-peer.md).
|
||||
Mark this transport as experimental: it works in most
|
||||
configurations but the BLE stack has more variability than
|
||||
Ethernet — adapter quirks, BlueZ version differences, and the
|
||||
shorter range all matter.
|
||||
|
||||
The BLE transport is **Linux-only** at present; macOS and
|
||||
Windows builds skip it.
|
||||
|
||||
## What you've learned
|
||||
|
||||
- **Ground-up is the new ground.** FIPS does not need any IP
|
||||
infrastructure between two devices to mesh them. A wire (or
|
||||
a radio link), `CAP_NET_RAW`, and a few config flags on each
|
||||
end are sufficient. The mesh supplies its own identity,
|
||||
addressing, discovery, and routing.
|
||||
- **Discovery is a four-flag opt-in.** `announce`, `discovery`,
|
||||
`auto_connect`, and `accept_connections` each control one
|
||||
thing; both ends must agree before a link will form.
|
||||
- **The two modes coexist.** Overlay peers and ground-up peers
|
||||
ride the same daemon — same FMP link layer, same FSP session
|
||||
layer, same `fips0` adapter. A node can be a bridge between
|
||||
the two without any extra plumbing.
|
||||
- **No IP on the link.** The Ethernet transport bypasses the
|
||||
kernel IP stack via `AF_PACKET`. Whether the interface has
|
||||
an IP address is irrelevant; whether it has carrier is what
|
||||
matters.
|
||||
- **Names work the same way.** `<npub>.fips` resolves locally
|
||||
via the cryptographically-derived ULA. The resolver does
|
||||
not care whether the destination is reached over Ethernet,
|
||||
UDP overlay, or some hop chain combining both.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
- **No beacons received.** On either node, `sudo fipsctl show
|
||||
transports` should show `beacons_received` incrementing
|
||||
every `beacon_interval_secs` once the other node is also
|
||||
running. If it stays at zero:
|
||||
- Confirm the chosen interface is `LOWER_UP` (carrier
|
||||
present).
|
||||
- Confirm the other node is announcing (its `beacons_sent`
|
||||
should be non-zero).
|
||||
- On WiFi: confirm AP client isolation is off.
|
||||
- On a switch: confirm the switch is unmanaged or that
|
||||
EtherType `0x2121` is not being filtered. Most consumer
|
||||
switches forward all EtherTypes; managed switches
|
||||
sometimes don't.
|
||||
- **Beacons received but no peer entry.** The handshake is
|
||||
failing. Tail logs (`journalctl -u fips`) for Noise
|
||||
handshake errors. Common causes: peer ACL active and not
|
||||
including the remote npub (out of scope for this tutorial,
|
||||
but check `/etc/fips/peers.allow` if you have set one);
|
||||
daemon's clock drift large enough to fail freshness
|
||||
checks (rare).
|
||||
- **Daemon won't start with the Ethernet transport.** Likely
|
||||
a permissions error. Check `journalctl -u fips` for an
|
||||
`EPERM` or "operation not permitted" message; if running
|
||||
interactively, confirm the binary has `CAP_NET_RAW`
|
||||
(`getcap "$(which fips)"`).
|
||||
- **Beacons in both directions, peers entries on both sides,
|
||||
but ping6 times out.** The handshake completed but the FSP
|
||||
session is not flowing data. Check `fipsctl show peers`'s
|
||||
link status columns — if the FMP link is healthy but FSP
|
||||
is not, the mesh-layer side is fine and the issue is one
|
||||
layer up. The
|
||||
[reach-mesh-services § Troubleshooting](reach-mesh-services.md#troubleshooting)
|
||||
section covers symptoms at this level.
|
||||
- **`AF_PACKET` socket bind fails on a kernel-protected
|
||||
interface.** Some hardened kernels (`grsec`, certain
|
||||
containers, certain VMs) restrict raw-socket access even
|
||||
with `CAP_NET_RAW`. The daemon log will name the failing
|
||||
syscall. The fix is host-side: relax the restriction or
|
||||
pick a different interface.
|
||||
|
||||
## What's next
|
||||
|
||||
You now have the second deployment mode of FIPS in your
|
||||
hands. From here:
|
||||
|
||||
- **Add a third node.** Bring up a third machine on the same
|
||||
Ethernet segment, configure it identically, and watch all
|
||||
three nodes form a mesh. The FIPS spanning tree picks a
|
||||
root and routing converges within a few beacon intervals.
|
||||
- **Mix transports.** Add an overlay peer (per
|
||||
[join-the-test-mesh](join-the-test-mesh.md)) to one of
|
||||
your ground-up nodes; the local mesh now reaches the test
|
||||
mesh through that node, and vice versa.
|
||||
- **Host services.** Anything you do on `fips0` with overlay
|
||||
peers — bind an HTTP server (per
|
||||
[host-a-service](host-a-service.md)), reach a service via
|
||||
the daemon's IPv6 adapter (per
|
||||
[reach-mesh-services](reach-mesh-services.md)) — works
|
||||
identically on a ground-up mesh. The data plane is the
|
||||
same.
|
||||
|
||||
For more depth on the link-layer machinery:
|
||||
|
||||
- [../reference/transports.md § Ethernet](../reference/transports.md)
|
||||
— full Ethernet transport reference (counter inventory,
|
||||
per-instance configuration, MTU model).
|
||||
- [../reference/configuration.md § Ethernet](../reference/configuration.md#ethernet-transportsethernet)
|
||||
— every configuration key and its default.
|
||||
- [../how-to/set-up-bluetooth-peer.md](../how-to/set-up-bluetooth-peer.md)
|
||||
— operator recipe for the BLE variant.
|
||||
- [../design/fips-transport-layer.md](../design/fips-transport-layer.md)
|
||||
— the design doc that describes the per-link MTU model and
|
||||
why each transport is treated as link-layer rather than
|
||||
network-layer.
|
||||
@@ -0,0 +1,517 @@
|
||||
# Host a Service of Your Own
|
||||
|
||||
In [reach-mesh-services](reach-mesh-services.md) you used
|
||||
ordinary IPv6 tools to reach services on other mesh nodes. This
|
||||
tutorial flips the direction. You will bring up a small HTTP
|
||||
server on your machine, make a deliberate choice about which
|
||||
interface it binds to, and turn on the mesh firewall to keep
|
||||
the exposure to what you intended. By the end you will have
|
||||
hosted your first peer-reachable service and made an informed
|
||||
decision about who can reach it.
|
||||
|
||||
The whole exercise should take about twenty minutes. You should
|
||||
have already worked through
|
||||
[persistent-identity](persistent-identity.md) so that the npub
|
||||
your service is reachable at does not change between restarts.
|
||||
|
||||
## What you'll build
|
||||
|
||||
```text
|
||||
┌──────────────────────────────────────────┐
|
||||
│ your fips node │
|
||||
│ │
|
||||
│ python3 -m http.server │
|
||||
│ --bind <fips0-addr> 8080 │
|
||||
│ │ │
|
||||
│ ▼ │
|
||||
│ fips0 fd97:….:Y ◀─── port 8080 open │
|
||||
└────────────┬─────────────────────────────┘
|
||||
│
|
||||
│ reachable as
|
||||
│ http://<your-npub>.fips:8080/
|
||||
▼
|
||||
any mesh peer that can route to <your-npub>.fips
|
||||
```
|
||||
|
||||
You will have:
|
||||
|
||||
- A single-page HTTP server bound to `fips0` only — not the
|
||||
public internet, not your LAN, just the mesh.
|
||||
- The mesh firewall baseline active, default-deny on `fips0`
|
||||
inbound, with one explicit drop-in that allows TCP/8080.
|
||||
- A clear understanding of which interface choice corresponds
|
||||
to which audience.
|
||||
|
||||
## Why bind interface matters
|
||||
|
||||
The single most important decision when hosting a service is
|
||||
**which interface (and therefore which audience) the service is
|
||||
exposed to**. This is true for every IPv6 service, not just
|
||||
FIPS — but FIPS makes it stark because your machine often has
|
||||
several interfaces with very different exposure profiles.
|
||||
|
||||
> **Bind interface = exposure surface.** A typical FIPS host
|
||||
> has at least two distinct audiences:
|
||||
>
|
||||
> - `fips0` (`fd97:…`, reachable as `<your-npub>.fips`) —
|
||||
> reachable only from FIPS peers that have a working link to
|
||||
> your node. Bound by Noise authentication and (optionally)
|
||||
> the peer ACL.
|
||||
> - `eth0` / `wlan0` (your LAN address) — reachable from anyone
|
||||
> on your local network segment, with no FIPS auth in the
|
||||
> way.
|
||||
>
|
||||
> When you run a server, the `--bind` argument decides which of
|
||||
> these audiences sees the service. Binding to a *specific*
|
||||
> address opts in to one audience. Binding to `[::]` or
|
||||
> `0.0.0.0` opts in to **all** of them at once — including any
|
||||
> you forgot you had.
|
||||
>
|
||||
> **Audit what's already listening.** Bringing up `fips0` adds
|
||||
> a new audience to every service on this host that was already
|
||||
> bound to `0.0.0.0` or `[::]`. SSH, your web server, a database
|
||||
> — if any of them was listening on all interfaces before you
|
||||
> joined the mesh, they are now reachable from mesh peers too.
|
||||
> A quick `ss -tulnp` will show you everything currently
|
||||
> listening and on which addresses. The mesh firewall (Step 5
|
||||
> below) is one way to bring those exposures back under explicit
|
||||
> control; rebinding the affected services to a specific
|
||||
> non-mesh address is another.
|
||||
|
||||
There is nothing FIPS-specific about this rule; it applies to
|
||||
SSH, web servers, databases, anything. FIPS just gives you the
|
||||
option of "mesh peers only" as a distinct audience, which most
|
||||
hosts otherwise don't have.
|
||||
|
||||
## Step 1: Find your node's mesh address
|
||||
|
||||
You need the `fd97:...` address assigned to your `fips0`
|
||||
adapter. Two equivalent ways to get it:
|
||||
|
||||
```sh
|
||||
ip -6 addr show fips0
|
||||
```
|
||||
|
||||
Look for the `inet6 fd97:...` line. The address up to the `/`
|
||||
is what you want.
|
||||
|
||||
Or via the daemon:
|
||||
|
||||
```sh
|
||||
sudo fipsctl show identities
|
||||
```
|
||||
|
||||
The first JSON entry has `local: true` and a `ula` field — that
|
||||
is your address.
|
||||
|
||||
For the rest of this tutorial we will write the address as
|
||||
`<your-fips0-addr>`. Substitute the actual `fd97:...` value
|
||||
when you run the commands. Save it to a shell variable for
|
||||
convenience:
|
||||
|
||||
```sh
|
||||
FIPS0_ADDR=$(ip -6 addr show fips0 | awk '/inet6 fd97:/ {print $2}' | cut -d/ -f1)
|
||||
echo "$FIPS0_ADDR"
|
||||
```
|
||||
|
||||
You should also know your npub from
|
||||
[persistent-identity](persistent-identity.md):
|
||||
|
||||
```sh
|
||||
NPUB=$(sudo cat /etc/fips/fips.pub)
|
||||
echo "$NPUB"
|
||||
```
|
||||
|
||||
`<your-npub>.fips` and `<your-fips0-addr>` are two names for
|
||||
the same destination.
|
||||
|
||||
## Step 2: Bring up an HTTP server bound to fips0
|
||||
|
||||
Make a small directory with one file in it so the server has
|
||||
something to serve:
|
||||
|
||||
```sh
|
||||
mkdir -p /tmp/mesh-demo
|
||||
echo '<h1>Hello from the mesh</h1>' > /tmp/mesh-demo/index.html
|
||||
cd /tmp/mesh-demo
|
||||
```
|
||||
|
||||
Start a Python HTTP server bound to your `fips0` address:
|
||||
|
||||
```sh
|
||||
python3 -m http.server --bind "$FIPS0_ADDR" 8080
|
||||
```
|
||||
|
||||
Leave the server running. The terminal will show:
|
||||
|
||||
```text
|
||||
Serving HTTP on fd97:... port 8080 (http://[fd97:...]:8080/) ...
|
||||
```
|
||||
|
||||
Two things to notice:
|
||||
|
||||
- The "Serving HTTP on …" line names your `fd97:...` address
|
||||
explicitly. Python is binding only to that one address.
|
||||
- The default would have been `0.0.0.0` — which is **not**
|
||||
what you want here. Without `--bind`, the server would be
|
||||
reachable from your LAN and any public IP this host has,
|
||||
not just from the mesh.
|
||||
|
||||
## Step 3: Verify the service locally
|
||||
|
||||
Open a second terminal. From the same host, fetch the page:
|
||||
|
||||
```sh
|
||||
curl -6 "http://[$FIPS0_ADDR]:8080/"
|
||||
```
|
||||
|
||||
Expect:
|
||||
|
||||
```text
|
||||
<h1>Hello from the mesh</h1>
|
||||
```
|
||||
|
||||
Now fetch it by name. Both forms should work:
|
||||
|
||||
```sh
|
||||
curl -6 "http://${NPUB}.fips:8080/"
|
||||
```
|
||||
|
||||
The daemon's local DNS responder turned the npub-form name into
|
||||
the `fd97:...` address, and the request landed at your HTTP
|
||||
server.
|
||||
|
||||
> **What just happened.** The kernel routed your local request
|
||||
> via the loopback path because the destination address is
|
||||
> assigned to one of your own interfaces. You exercised the
|
||||
> client side (DNS, IPv6 socket open, HTTP request) and the
|
||||
> server side (HTTP listener, response). What you have not yet
|
||||
> verified is reachability *from another mesh node*. That is
|
||||
> the next concern.
|
||||
|
||||
If you have `fipstop` available, open it now in another terminal
|
||||
and switch to the **Node** tab. The right-half of the Traffic
|
||||
block — the **Listening on fips0** panel — should list a `tcp`
|
||||
row at port 8080 with a `python(<pid>)` Process column. The State
|
||||
column reads `OPEN` because the firewall has not been turned on
|
||||
yet; everything bound to `fips0` is currently mesh-reachable. The
|
||||
yellow banner above the panel says
|
||||
"fips-firewall.service inactive — all listeners exposed". Both
|
||||
signals will flip in the next two steps.
|
||||
|
||||
## Step 4: Reachability from a mesh node
|
||||
|
||||
Any mesh node — a direct peer, or a node several hops away —
|
||||
reaches your service the same way you reached `test-us01` in
|
||||
[reach-mesh-services](reach-mesh-services.md): it looks up
|
||||
`<your-npub>.fips`, gets back your `fd97:...` address, opens
|
||||
a TCP connection to it, and the FIPS data plane carries the
|
||||
packets across the mesh to you. From the remote host the curl
|
||||
looks identical to yours:
|
||||
|
||||
```sh
|
||||
curl -6 "http://${NPUB}.fips:8080/"
|
||||
```
|
||||
|
||||
If you have a second machine on the mesh — or you can ask
|
||||
another operator to try it — this is the moment to confirm.
|
||||
A node that can already `ping6 ${NPUB}.fips` should also be
|
||||
able to fetch your page. If it can ping but the curl times
|
||||
out, jump to [Troubleshooting](#troubleshooting) — but most
|
||||
likely the firewall step in this tutorial has not happened
|
||||
yet, so a remote attempt right now will succeed straight
|
||||
through to your HTTP server.
|
||||
|
||||
That is the problem. With no firewall in place, *any* mesh
|
||||
node that can route to you — your direct peers, and every node
|
||||
beyond them in the mesh — can reach port 8080. You have not
|
||||
yet made a deliberate decision about whether you want that.
|
||||
The rest of this tutorial replaces the implicit "every port
|
||||
with a listener is reachable" with an explicit "only the ports
|
||||
I have opened are reachable, optionally only from specific
|
||||
mesh nodes." That is the firewall's job, and it is the only
|
||||
mechanism in play in the rest of this tutorial.
|
||||
|
||||
(There is a separate, unrelated control called the *peer ACL*
|
||||
that decides which npubs may establish a peer connection with
|
||||
your node at the transport layer. It is not part of the
|
||||
firewall and does not affect what is described below; Step 7
|
||||
is a brief signpost to it.)
|
||||
|
||||
## Step 5: Activate the mesh firewall baseline
|
||||
|
||||
FIPS ships a default-deny nftables baseline at
|
||||
`/etc/fips/fips.nft` that restricts inbound traffic on `fips0`
|
||||
to ICMPv6 echo and conntrack replies. The baseline is **not**
|
||||
enabled by default — activation is an explicit step the
|
||||
operator has to take.
|
||||
|
||||
Activate it:
|
||||
|
||||
```sh
|
||||
sudo systemctl enable --now fips-firewall.service
|
||||
```
|
||||
|
||||
This loads the table immediately and arranges for it to load
|
||||
on every subsequent boot. Confirm:
|
||||
|
||||
```sh
|
||||
sudo nft list table inet fips
|
||||
```
|
||||
|
||||
You will see one chain named `inbound` hooked at `input`,
|
||||
roughly:
|
||||
|
||||
```text
|
||||
table inet fips {
|
||||
chain inbound {
|
||||
type filter hook input priority filter; policy accept;
|
||||
iifname != "fips0" return
|
||||
ct state established,related accept
|
||||
icmpv6 type echo-request accept
|
||||
counter packets 0 bytes 0 drop
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
The chain admits ICMPv6 echo (so `ping6` from any mesh node
|
||||
still works) and conntrack replies (so your *outbound*
|
||||
connections still get their replies back). Everything else
|
||||
inbound on `fips0` hits the final `counter ... drop`.
|
||||
|
||||
> **What this changed.** Your HTTP server is still running and
|
||||
> still reachable *from this same host* (same-host traffic to
|
||||
> `fd97:...` goes via the loopback path, which has
|
||||
> `iifname != "fips0"` and short-circuits at the first rule).
|
||||
> But any mesh node trying to reach `fd97:...:8080` now has its
|
||||
> TCP SYN dropped before it can reach your server. From the
|
||||
> remote end the connection times out.
|
||||
|
||||
The fipstop panel reflects the change immediately: the yellow
|
||||
"firewall inactive" banner disappears, the panel title becomes a
|
||||
plain "Listening on fips0", and your `tcp 8080 python(<pid>)` row
|
||||
flips to **DarkGray** with `filt` in the State column. Every
|
||||
other row also goes DarkGray — none of them have an explicit
|
||||
accept rule yet, and the chain falls through to `counter drop`.
|
||||
|
||||
So the firewall is in the right shape but in the wrong state
|
||||
for our purpose: we *want* mesh nodes to reach port 8080. The
|
||||
next step opens that one port.
|
||||
|
||||
## Step 6: Open port 8080 via a drop-in
|
||||
|
||||
Drop-ins live under `/etc/fips/fips.d/` with the `.nft`
|
||||
suffix. Each file is included into the `inbound` chain at the
|
||||
marked point and may contain any nftables rule lines valid in
|
||||
that context.
|
||||
|
||||
Create one for your HTTP service:
|
||||
|
||||
```sh
|
||||
sudo tee /etc/fips/fips.d/http-mesh-demo.nft >/dev/null <<'EOF'
|
||||
tcp dport 8080 accept
|
||||
EOF
|
||||
```
|
||||
|
||||
Reload the firewall:
|
||||
|
||||
```sh
|
||||
sudo systemctl reload-or-restart fips-firewall.service
|
||||
```
|
||||
|
||||
Confirm the rule is live:
|
||||
|
||||
```sh
|
||||
sudo nft list table inet fips
|
||||
```
|
||||
|
||||
The `inbound` chain now contains your `tcp dport 8080 accept`
|
||||
rule between the conntrack rule and the final `counter drop`.
|
||||
|
||||
A curl from any mesh node will now reach the HTTP server. The
|
||||
path is: remote node's mesh data plane → forwarded across the
|
||||
mesh → your direct peer's link to you → `fips0` ingress →
|
||||
`inbound` chain → matches `tcp dport 8080 accept` → delivered
|
||||
to the HTTP server.
|
||||
|
||||
In the fipstop panel, your `tcp 8080 python(<pid>)` row flips
|
||||
back to **default White** with `OPEN` in the State column on the
|
||||
next poll tick. No other row changes — they remain DarkGray
|
||||
`filt` because you have only opened this one port. The panel
|
||||
doubles as a security screen for the rest of the tutorial: any
|
||||
service whose row reads `OPEN` is mesh-reachable, anything
|
||||
DarkGray is filtered. If you later add a saddr-restricted
|
||||
drop-in (covered just below), the row will land at `filt?`
|
||||
rather than `OPEN`, signalling that the rule exists but is
|
||||
source-scoped — the panel deliberately does not classify
|
||||
restricted accepts as fully open.
|
||||
|
||||
If you only want to expose the service to a *specific* node
|
||||
or set of nodes, source-filter the rule. The address filter
|
||||
applies to the mesh-source address as it arrives on `fips0`,
|
||||
which is the originating node's address — not necessarily a
|
||||
direct peer. Replace the drop-in contents with something like:
|
||||
|
||||
```nft
|
||||
ip6 saddr fd97:1234:5678:9abc:def0:1234:5678:9abc tcp dport 8080 accept
|
||||
```
|
||||
|
||||
The source address is the node's mesh address, which it
|
||||
publishes in its `fips.pub` (and which you can resolve from
|
||||
its npub). For multiple nodes, use a set:
|
||||
|
||||
```nft
|
||||
ip6 saddr {
|
||||
fd97:1111:2222:3333:4444:5555:6666:7777,
|
||||
fd97:8888:9999:aaaa:bbbb:cccc:dddd:eeee
|
||||
} tcp dport 8080 accept
|
||||
```
|
||||
|
||||
For the worked example, leave the drop-in unfiltered — any
|
||||
mesh node that can route to you can fetch your page.
|
||||
|
||||
## Step 7: A note on the peer ACL
|
||||
|
||||
The firewall you just configured is the only control in scope
|
||||
for this tutorial. There is a separate, optional control
|
||||
called the *peer ACL* that you may run across in other docs;
|
||||
it is unrelated to the firewall and worth a sentence here only
|
||||
so you do not confuse the two.
|
||||
|
||||
The peer ACL decides which npubs may establish a peer
|
||||
connection with your node at the transport layer. It does not
|
||||
look at ports, drop-ins, or `fips0` traffic. You do not need
|
||||
it for this tutorial.
|
||||
|
||||
For when you do:
|
||||
|
||||
- [../how-to/enable-nostr-discovery.md](../how-to/enable-nostr-discovery.md)
|
||||
— operator recipe.
|
||||
- [../reference/security.md § Peer ACL](../reference/security.md#peer-acl)
|
||||
— file format, evaluation order, alias handling.
|
||||
|
||||
## Step 8: Stop the server and tidy up
|
||||
|
||||
When you are done, stop the HTTP server in the first terminal
|
||||
with `Ctrl-C`. The drop-in stays in place; remove it if you
|
||||
do not want port 8080 reachable after the demo:
|
||||
|
||||
```sh
|
||||
sudo rm /etc/fips/fips.d/http-mesh-demo.nft
|
||||
sudo systemctl reload-or-restart fips-firewall.service
|
||||
```
|
||||
|
||||
The `fips-firewall.service` itself can stay enabled —
|
||||
default-deny on `fips0` is a sensible posture even with no
|
||||
extra services running. To turn it back off:
|
||||
|
||||
```sh
|
||||
sudo systemctl disable --now fips-firewall.service
|
||||
```
|
||||
|
||||
## What you've learned
|
||||
|
||||
- **Bind interface = audience.** Binding to a specific address
|
||||
opts in to one audience; binding to wildcard
|
||||
(`0.0.0.0` / `[::]`) opts in to *all* of them, including
|
||||
ones you forgot you had. For mesh-only exposure, bind to
|
||||
your `fd97:...` address. The fipstop **Listening on fips0**
|
||||
panel marks wildcard binds with a trailing `*` after the
|
||||
process name as a reminder that the bind is not
|
||||
fips0-specific.
|
||||
- **Same-host loopback is misleading.** A local curl to your
|
||||
own `fd97:...` address goes via the loopback path, not
|
||||
through `fips0` ingress. To actually verify mesh-side
|
||||
reachability you need a second machine, or to read what
|
||||
the firewall is doing in `nft list table inet fips`.
|
||||
- **The mesh firewall is opt-in.** `fips-firewall.service` is
|
||||
not enabled by default. Once it is enabled, `fips0` is
|
||||
default-deny except for ICMPv6 echo and conntrack replies.
|
||||
- **Ports open via drop-ins.** Each file under
|
||||
`/etc/fips/fips.d/*.nft` adds rules into the `inbound`
|
||||
chain. Source-filter with `ip6 saddr` to scope a port to
|
||||
specific mesh nodes.
|
||||
- **Two independent controls at two different layers.** The
|
||||
firewall is a layer-3 filter on `fips0`: it controls which
|
||||
TCP/UDP ports are reachable and (optionally) which mesh
|
||||
source addresses may reach them. The peer ACL is a
|
||||
transport-layer admission filter on Noise handshakes: it
|
||||
controls which npubs may become direct peers of your node.
|
||||
They are unrelated — the ACL does not touch fips0 traffic,
|
||||
and the firewall does not look at npubs.
|
||||
|
||||
You now have the mental model for hosting any IPv6 service
|
||||
behind a deliberate exposure policy. The mechanics generalize:
|
||||
SSH on port 22, a database on port 5432, a custom protocol on
|
||||
its own port — same `--bind` rule, same drop-in shape.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
- **A remote mesh node cannot reach the service after the
|
||||
firewall reload.** Check the drop-in syntax with
|
||||
`sudo nft -c -f /etc/fips/fips.nft` before reloading; a
|
||||
syntax error in any drop-in causes the whole table to fail
|
||||
to load and the previous rules persist. Then
|
||||
`sudo nft list table inet fips` to confirm your
|
||||
`tcp dport 8080 accept` rule is present in the `inbound`
|
||||
chain.
|
||||
- **Local curl works, remote curl times out.** The packet is
|
||||
reaching `fips0` ingress and being dropped by the baseline.
|
||||
Either your drop-in did not load (see above) or it has a
|
||||
source filter that excludes the remote node's address.
|
||||
- **Local curl fails after binding to fips0.** Double-check
|
||||
that your `FIPS0_ADDR` matches the address shown in
|
||||
`ip -6 addr show fips0`. The Python server message also
|
||||
echoes the bound address — confirm it starts with `fd97:`,
|
||||
not `127.0.0.1` or `::`.
|
||||
- **`Address already in use` from Python.** Another process
|
||||
holds port 8080. Pick a different port (`8081`, `9000`, …)
|
||||
for both the `python3 -m http.server` invocation and the
|
||||
drop-in.
|
||||
- **Watch the firewall counter to confirm drops.** The
|
||||
`counter ... drop` line at the bottom of the chain
|
||||
increments on every dropped inbound packet. After a remote
|
||||
mesh node attempts to reach a port you have not opened,
|
||||
`sudo nft list table inet fips` will show the counter
|
||||
packet count rising.
|
||||
- **Use fipstop to spot-check listener and filter state.** The
|
||||
Listening on fips0 panel on the Node tab shows every
|
||||
fips0-reachable listener and its current filter state. A row
|
||||
staying `filt` after you expected `OPEN` usually means the
|
||||
drop-in failed to load (a syntax error in any file under
|
||||
`/etc/fips/fips.d/` aborts the whole reload, leaving the
|
||||
previous ruleset in place) or the drop-in carries a source
|
||||
filter and now reads `filt?` rather than `OPEN`.
|
||||
|
||||
## What's next
|
||||
|
||||
- [ground-up-mesh.md](ground-up-mesh.md) — Bring up two devices
|
||||
on a shared physical link — Ethernet, WiFi, or Bluetooth —
|
||||
with no pre-existing IP infrastructure between them. The
|
||||
second deployment mode of FIPS, where the mesh is the
|
||||
network rather than an overlay on top of one. Coexists with
|
||||
overlay peers; the same daemon carries both.
|
||||
|
||||
For more depth on the firewall and ACL surface:
|
||||
|
||||
- [../how-to/enable-mesh-firewall.md](../how-to/enable-mesh-firewall.md)
|
||||
— operator recipes for the baseline, drop-in patterns,
|
||||
and how to fold the baseline into an existing
|
||||
`nftables.conf`.
|
||||
- [../reference/security.md](../reference/security.md) —
|
||||
consolidated security reference: nftables baseline rules,
|
||||
drop-in format, peer ACL semantics, default exposures by
|
||||
transport, threat-resistance matrix.
|
||||
- [../design/fips-security.md](../design/fips-security.md) —
|
||||
threat model, why the baseline is opt-in, the metadata-
|
||||
privacy posture.
|
||||
|
||||
If you want to host a service that is *not* on a FIPS node — say,
|
||||
an existing HTTP server on a regular LAN box — and expose it to
|
||||
mesh peers through a `fips-gateway`, that's the inbound
|
||||
port-forward mode: the gateway runs a mesh-side listener on `fips0`
|
||||
and forwards to a LAN target. The operator recipe is at
|
||||
[../how-to/deploy-gateway.md#inbound-port-forwarding](../how-to/deploy-gateway.md#inbound-port-forwarding);
|
||||
a hand-held walk-through on an OpenWrt AP is at
|
||||
[deploy-fips-gateway.md](deploy-fips-gateway.md) under "Advanced"
|
||||
in [README.md](README.md).
|
||||