Files
fips/docs/design/fips-nostr-discovery.md
T
Johnathan Corgan 6a564e26ac Prepare the v0.5.0 release content
Everything the release needs except the version number, which stays at
0.5.0-dev until the tag.

The changelog entry covers only the work that is new on this line. The
point release's forty-six entries arrived under their own heading with the
forward merge and are left alone; the twenty that remained are regrouped by
topic and eight more added for changes no entry covered. Three of those
eight matter to someone upgrading. Five root modules and four re-exports
left the public library surface and Node::connections narrowed, none of it
recorded anywhere; the entry names what to use instead and distinguishes
the removed connection-phase enum from the Noise type of the same name,
which is a different type that still exists. Tracing targets moved, so an
existing RUST_LOG filter stops matching rather than erroring. And the
handshake resend interval key no longer governs the first resend, which is
now a constant, though it still governs later ones.

Seven more entries cover the work that landed after the first content pass
was written: the experimental native datagram API, the fipsctl probe
diagnostic, per-instance transport addressing, the app-owned UDP socket
seam, and the connect, disconnect and path-MTU fixes. The four bug fixes
among them all reach the deployed line, so the release notes no longer
claim this release carries exactly one fix for a shipped bug; it carries
four.

There is no security section, because after the split every security entry
belongs to the point release. The release notes say so plainly rather than
leaving a reader upgrading across both releases to conclude this one
carries no security work.

The notes are organized by audience, since the release spans OpenWrt
routers, embedders, FreeBSD, and the existing platforms, and a single list
serves none of them. The native datagram API is given a section of its own
rather than folded into the embedding seam: it is a client-facing API
rather than a way to host a node, and its one rule with no Berkeley-socket
counterpart, that the v1 wire carries no half-close, needs to be somewhere
a client author will read it. FreeBSD is advertised as supported on x86_64
only, stated wherever the platform appears. Android is advertised as an
embedding seam and not as a supported platform: a compile-gated library
surface with no artifact and no host application guide.

The configuration table rename is carried through every shipped file that
taught the old spelling: nine documentation files, the OpenWrt sample
config and a test generator, twenty-two sites in all. Guides written this
same cycle were among them, which is how the omission was found. The
documentation that arrived with the native API was checked for the same
omission and was already clean. The compatibility tests keep the old
spelling deliberately, since they exist to test the fold.

The changelog section is the fold of master's [Unreleased], not a snapshot
of it. An earlier version of this commit took a copy that then drifted, so
each section ended up holding a bullet the other did not and re-folding
them would have picked a winner silently. Both causes were fixed on master
instead — the NixOS module had never been recorded there, and the
pre-release batch of fixes was new — so [Unreleased] is a strict superset
and this is a copy rather than a merge. [0.5.0] carries all forty-six
bullets byte for byte, [Unreleased] is empty, and [0.4.2] is untouched,
checked by hashing it against master's copy.

The BLE work landed after the content pass and gets one summary entry in
the changelog and one section in the release notes rather than nine
bullets: the ble_available gate replacing target_os = "linux",
packet-boundary recovery for stream-oriented backends, peer recognition by
node identity instead of a rotating link address, the L2CAP PSM moving
into the backend seam and onto the advertisement, the embedder-supplied
Android radio, bounded probe retry, and inbound handshakes moved off the
accept loop.

The two release-notes copies no longer share their link paths. Relative
links resolve from one directory only, so the seven written for
docs/releases/ all 404ed from the root copy. The root copy now uses paths
from the repository root and the versioned copy keeps the ../ form; both
sets were resolved against the tree. The same two links are broken the
same way in the v0.4.0 through v0.4.2 notes, left as shipped history.

The contributor tallies are re-derived against maint..HEAD rather than
adjusted: twenty commits from outside the project and 171 from me, with
Arjen at fifteen and fr34aky at two. An earlier count of twelve and 138
was carried from a measurement taken three days before this content was
written, and the BLE branch widened the gap after it. Arjen's NixOS flake
module, the UDP sin6_scope_id fix and most of the BLE rework were
uncredited, as was fr34aky's L2CAP PSM seam. They want one last re-derive
at tag time if anything lands before the tag.

A sweep of all 99 tracked markdown files against the tree corrected
fifty-three of them. Four told the reader to run a build.sh that does not
exist; the only harness builder is testing/scripts/build.sh. The BLE build
prerequisites were described as optional on the strength of a probe that
build.rs does not perform, and bluez was named a build prerequisite when
libdbus-sys asks only for libdbus-1-dev and pkg-config and bluez is the
runtime daemon. Link cost is the primary sort key in next-hop ranking, not
reserved for future use; Ethernet runs on macOS as well as Linux; the BLE
MTU is the L2CAP CoC MTU rather than a negotiated ATT_MTU; effective
Ethernet MTU is 1497; the LAN discovery subsystem is src/mdns and eight
citations still named a src/discovery that never existed here. The
connectivity states in three tutorials were invented, and their jq filters
matched nothing including healthy peers. One command filtered on a literal
fd97: address prefix, which only the first byte of fixes, so it returned
empty for all but one reader in 256 and every later step using the
variable failed silently. transports.tor.advertise_on_nostr was
undocumented despite being validated against node.rendezvous.nostr.enabled.

The transport design document gains the BLE section it never had, written
from the source: the backend cascade and its compile_error tripwire, the
platform gate, the PSM advertisement wire layout and the byte budget that
forces a 16-bit service-data key, and the probe and admission bounds.

Three source files carried the same class of staleness and are corrected
with the documentation: the OpenWrt ipk usage line and Makefile error text
both named a packaging/openwrt that does not exist, and chaos.sh parsed
--subnet without listing it.

Folded in with the content commit, having been prepared alongside it:

The three GitHub Action pins that had gone stale. Every third-party
action is pinned to a commit SHA, nothing reports that a pin has aged,
and re-resolving all ten against their tags found dorny/test-reporter@v2,
taiki-e/install-action@v2 and vmactions/freebsd-vm@v1 had moved. The
three install-action@nextest references stay unpinned, since that action
reads the tool to install from the ref name. check-action-pins.sh passes
at 75 references and all nine workflow files parse.

The lockfile refresh, which is the mutating half of the dependency sweep.
Thirty-six packages move to their latest semver-compatible versions and
every one is transitive; nothing declared in Cargo.toml changes version.
No advisory forces any of them. It was taken before the validation
battery, because a gate run against a lockfile that later moves proves
nothing about what ships.

The sha2 0.10 to 0.11, hkdf 0.12 to 0.13 and bech32 0.11 to 0.12 majors,
three of the four deferred at v0.4.0 for change surface rather than
security. All three land with no source change. sha2 and hkdf must move
together, since both depend on digest 0.11, and neither changes an
algorithm. That matters because the chaining-key KDF in the Noise
handshake is built on Hkdf::<Sha256>, where an output change would be a
wire break rather than a compile error; no known-answer vectors exist for
that path, so the wire-compatibility gate is what covers it. secp256k1
0.31 is deliberately absent, since nostr's own requirement would leave
two copies of the ECC library in the tree.

The README support matrix, rebuilt as one feature table broken out by
Linux variety. A single Linux column hid that Debian, Ubuntu, Arch and
NixOS are one glibc build differing in packaging, that OpenWrt is musl
and drops BLE, and that Android is not a daemon platform. Transport rows
sort by how many platforms carry them. A Native API row reads its
platform set from the cfg gates. The installer row becomes a package
format row naming the artifact, and only the .deb is exercised per
release.

Four changelog and release-note gaps the BLE re-walk found: a Bluetooth
LE bullet stranded inside the released 0.4.2 section, a missing Fixed
entry for the scan and probe loop counting a pool-refused connection as
an established link, the unnamed embedder call that installs an
application-owned radio, and the fact that stopping the transport now
stops scanning as well as advertising.

Three release-document gaps found walking the unsurveyed commits: the UDP
reuse-flag fix stated in the direction opposite to the one it was made,
with the silent second-daemon bind it prevents left unsaid; the corrected
native-API socket paragraph carried into both release-note copies, which
still named SOCK_SEQPACKET on FreeBSD and two kernels where three are
handled; and the coordinate-cache hardening, which shipped with no text
anywhere despite adding four operator-visible status fields. That last
entry states plainly that the checks are mitigations and not a closure,
since the coordinate is still not authenticated.

Also folded in, the documentation pass that followed the content commit:

A stage-pipeline diagram for the probe, embedded in the fipsctl
reference under the five-stage list. It draws the five stages left to
right with each stage's failure reasons below it, and the bypass that
skips both lookup stages when the coordinates are cached or the target
is a direct peer. Its branches come from the probe state machine rather
than from the report, so the path stage is drawn as the one failure that
does not stop the probe.

A rewrite of the README's "What FIPS does" section. It now opens with
what a machine running FIPS gets, rather than with the two deployment
modes, and gives the self-organizing and permissionless property its own
paragraph since it holds for both modes.

A regrouping of the README's feature list into the mesh, getting traffic
onto it, and running a node, with a bullet added for the native datagram
API, which had none despite sitting in the support matrix. The Quick
start now leads with the released packages rather than a source build.
It also fixes a real defect: the package enables fips.service and
fips-dns.service and starts neither on a fresh install, so .fips name
resolution was silently dead until the next reboot and neither page said
to start the service.

A rewrite of the release notes. They opened with seven subsections of
upgrade caveats and reached the first feature two hundred lines in; they
now open with a summary of the release and elaborate below it in the
same order. Android is stated as supported through an embedded crate
rather than as a standalone daemon, consistently across all three
documents. The OpenWrt pair is corrected: it is 802.11s between routers
with FIPS supplying encryption, authentication and routing, plus a
convention of an open !FIPS SSID a client joins over WiFi, not meshing
over a router's own radios. The probe's path output is described as the
least-common-ancestor walk, which is the worst-case fallback route
rather than the route a packet takes. Detail that did not change what a
reader does was cut from the notes and kept in the changelog.
2026-08-30 10:42:59 +00:00

32 KiB
Raw Blame History

FIPS Discovery: Nostr-Mediated and LAN/mDNS

FIPS nodes have two discovery mechanisms beyond the static peers[] list. The bulk of this document describes Nostr-mediated discovery, which works across the internet using public Nostr relays as a signaling channel and can punch through UDP NAT. A second, much simpler mechanism — LAN/mDNS discovery — finds peers on the same local link with no relay, STUN, or NAT traversal at all; it is described in its own section near the end. The two are independent: a node can enable either, both, or neither.

Nostr-mediated discovery lets FIPS nodes find each other, and if necessary, punch through UDP NAT, using public Nostr relays as the signaling channel. A node publishes its reachable transport endpoints to a small set of relays under its own Nostr identity (which is also its FIPS identity), and peers resolve those endpoints at dial time by npub. For peers behind UDP NAT, the same relay channel carries an encrypted offer/answer exchange, and STUN supplies the reflexive address used for a coordinated hole-punch.

Nostr discovery is unconditionally compiled into the fips binary on every supported platform and ships in every published release artifact (.deb, AUR, systemd tarball, OpenWrt .ipk and .apk, FreeBSD .pkg, macOS .pkg, Windows .zip). It is runtime-opt-in: the YAML configuration defaults to disabled (node.rendezvous.nostr.enabled: false), so the discovery runtime stays dormant — and opens no relay connections — until an operator flips the flag. Default relay and STUN-server lists ship in the config; both are optional overrides. When disabled, nodes behave exactly as before: only the static peers[] addresses are used.

Role

The feature adds three capabilities on top of FIPS's static peer model:

  • Advertising. A node publishes the transport endpoints it wants peers to use (direct UDP, direct TCP, a Tor onion, or the special udp:nat rendezvous token) as a signed Nostr event. The advert is anchored to the node's FIPS identity key — a peer that knows the npub knows the advert is authentic.
  • Lookup. When dialing a configured peer marked via_nostr, or any peer in policy: open mode, the node fetches that peer's advert from the configured relays and appends the advertised endpoints to its dial list. Static addresses are always tried first.
  • UDP NAT hole-punch. When both sides of a connection have UDP NAT endpoints, the advert carries enough information to run a STUN-based offer/answer exchange over encrypted (NIP-59) Nostr events. Each side observes its reflexive address via STUN, exchanges candidate pairs through the relay, and both sides send UDP probes at a shared punch time. On the first successful probe, the punch socket is handed to FMP and becomes a normal UDP transport.

When to use it

  • You run a public node and want peers who know your npub to reach you without you distributing an address list out-of-band.
  • You want to reach a peer behind UDP NAT without deploying a relay or running Tor on both sides. The peer advertises udp:nat and you dial by npub.
  • You want zero-touch peer discovery within a known application namespace (policy: open), subject to an admission budget.
  • You want to advertise a Tor onion so peers don't need to know the .onion address out-of-band.

Skip the feature when every peer is already reachable through a stable static address (a LAN mesh, a pre-configured test bed, or a deployment where operators distribute peers[] blocks directly). The feature adds relay dependencies, STUN round-trips for NAT cases, and a small ambient background of relay traffic; none of that is useful when you already know where peers are.

Scenarios and configuration

For end-to-end operator recipes — each of the five activation scenarios (advertise a directly-reachable UDP node, advertise a Tor onion node, look up a configured peer by npub without advertising, NAT hole-punch between two configured peers, and open discovery within an app namespace) — see ../how-to/enable-nostr-discovery.md. The full configuration knob tables, per-transport keys, and startup validation rules live in ../reference/configuration.md under node.rendezvous.nostr.*. The Kind 37195 advert event format is in ../reference/nostr-events.md. The rest of this document covers the design of the discovery runtime itself.

Under the covers

The rest of this document describes how the feature works inside the node. For the generic protocol shape (event tags, NIP usage, on-the- wire offer/answer schema, failure-suppression machinery), see port-advertisement-and-nat-traversal.md.

Overview

The discovery runtime is a background task group started during node initialization when nostr.enabled is true. It maintains a single nostr-sdk client connected to the union of advert_relays and dm_relays, and runs four loops: advert publication, advert subscription (for open discovery and cache warming), DM subscription (for incoming offers and answers), and a periodic advert-cache prune. Discovery has no CLI surface; all operations are driven by the configuration and by connection attempts made by the rest of the node.

                    +-----------------------+
                    |   Discovery runtime   |
                    +-----------------------+
                       |       |       |
        advert publish |       | DM sub (offers, answers)
                       |       |
                       v       v
              +-------------------------+
              |   Nostr relay pool      |  (advert_relays ∪ dm_relays)
              +-------------------------+
                       ^       ^
    advert fetch/cache |       | encrypted signaling
                       |       |
   +----------------+  |       |  +--------------------+
   | connect_peer   |--+       +->|  offer / answer    |
   |  (node side)   |             |  handler           |
   +----------------+             +--------------------+
           |                                |
           v                                v
      +---------+                    +--------------+
      |  STUN   |<-- same socket --->|  UDP punch   |
      +---------+                    +--------------+
                                            |
                                            v
                                   adopt_established_traversal()
                                            |
                                            v
                                      FMP IK handshake
                                      on adopted socket

Phase 1 — Advertisement

Adverts are published as Nostr kind 37195 parameterized replaceable events (FIPS-specific, in the application-defined replaceable range 30000–39999; the digits visually spell FIPS — 7=F, 1=I, 9=P, 5=S). The d tag is hardcoded to the wire-format identifier fips-overlay-v1 (or fips-overlay-v1-next on the next branch), so each node has a single, in-place-updatable advert under its identity. The configurable app value populates a separate protocol tag, which scopes adverts within a relay set without splitting them across multiple d-tag streams. The event is signed with the node's FIPS identity key; there is no separate Nostr key. A NIP-40 expiration tag is set to now + advert_ttl_secs, and a version tag carries the protocol version. The advert content is a JSON document shaped as OverlayAdvert (see ../reference/nostr-events.md for the schema).

Publication happens on startup, again whenever the set of advertised endpoints changes (for example, when a Tor onion hostname first becomes available), and on a refresh timer every advert_refresh_secs. If the advertise flag is turned off, the previous advert event is deleted using a NIP-9 kind 5 delete event. Advert publication is fan-out: the same event is sent to every relay in advert_relays with no explicit failover — relay redundancy is implicit.

For a UDP or TCP transport with public: true, the address advertised follows a fixed precedence: an operator-supplied external_addr wins; otherwise a non-wildcard bound local_addr is used directly; otherwise — only for UDP — the runtime asks stun_servers for the reflexive address of the bound socket and advertises that. TCP has no STUN equivalent, so wildcard-bound TCP without external_addr produces a loud WARN and the endpoint is omitted from the advert.

Phase 2 — Lookup

When the node decides to dial a peer that is eligible for Nostr resolution (a via_nostr peer, or any peer under policy: open), it issues a Nostr REQ filtered by author = peer_pubkey, kind = 37195, #d = fips-overlay-v1. The fetch is time-bounded (~2 s) and runs against all configured advert_relays in parallel. The first valid advert wins; adverts whose protocol tag does not match the local app value are rejected at validation.

Results are kept in an in-memory cache keyed by author npub. Cache entries carry the advert's expiration time; a periodic prune drops expired entries, and an LRU-by-expiry eviction enforces advert_cache_max_entries. A parallel long-lived subscription on the advert relays populates the cache passively, so open-discovery candidates do not require per-dial fetches.

On cache hit, advert endpoints are appended to the peer's static address list with lower priority; the static list is tried first.

Phase 3 — Offer/Answer signaling

For any endpoint shaped as udp:nat, dialing triggers an offer/answer exchange before the first packet is sent. Signaling events are Nostr kind 21059 (ephemeral, not stored by conforming relays), gift-wrapped per NIP-59 and encrypted with NIP-44, so only the intended recipient can decrypt the payload.

The initiator performs STUN first (see Phase 4), then builds a TraversalOffer containing:

  • A unique sessionId and a random nonce (used to correlate the answer).
  • Its reflexive address (if STUN succeeded).
  • Its list of local (private) addresses for same-LAN paths.
  • The STUN server it used, for informational reporting only.
  • An expiresAt equal to now + signal_ttl_secs.

The offer is sealed to the recipient's npub and published to the peer's preferred signaling relays — the node first tries to resolve the peer's NIP-17 DM relay list (kind 10050), and falls back to dm_relays if the inbox-relays fetch fails. Each side also publishes its own inbox relay list on startup so dialers can discover it.

On the receiving side, admission is a pair of bounds taken together: a per-sender allowance of max_concurrent_offers_per_npub, keyed on the npub that signed the gift wrap, nested inside a global max_concurrent_incoming_offers. A sender over its own allowance is refused at debug, since by definition it is sending faster than the node wants and a record per rejection would turn the spam into log volume; the global bound being reached is the operator-visible warn, because that one says the node is genuinely saturated. Together they keep one identity from holding the whole pool. They do not make the pool inexhaustible: nostr identities are free to generate, so an attacker running ceil(max_concurrent_incoming_offers / max_concurrent_offers_per_npub) throwaway npubs still saturates it at the same total offer rate. Raising the attacker's cost beyond keypairs would mean pricing the offer itself. A sessionId replay cache (bounded by seen_sessions_max_entries, with entries valid for replay_window_secs) rejects duplicates.

The responder runs its own STUN query and replies with a TraversalAnswer carrying its reflexive and local addresses plus a PunchHint { startAtMs, intervalMs, durationMs } that tells both sides when to begin probing and how aggressively. If the responder has no usable addresses at all, it replies with accepted: false and a reason string.

Phase 4 — UDP hole-punch

Each side runs STUN (parsing XOR-MAPPED-ADDRESS from the response, all other attributes ignored) on the same UDP socket it will later use for punching and for the adopted FMP transport. This is critical: NAT state is per-socket, so the punch has to reuse the socket that taught the NAT about this binding.

Given its own reflexive + local addresses and the peer's, each side builds a candidate-pair plan that tries, in priority order:

  1. Reflexive ↔ reflexive. The classic STUN path. Tried first because it is the only candidate that's reliable across arbitrary network topologies — host candidates from one peer that happen to be reachable from the other (via a corporate VPN, a Tailscale subnet route, or overlapping private address space) will succeed at the socket layer in the punch but fail in the FMP handshake when the return path doesn't match.
  2. LAN ↔ LAN. If both sides share a /24 prefix, same-subnet private addresses are likely reachable directly. Only fires when both peers shared local host candidates (which requires share_local_candidates to be enabled — off by default).
  3. Mixed. Reflexive on one side, local on the other — catches hairpin and one-side-public scenarios.

At startAtMs both sides begin sending 24-byte probe packets on the candidate pair(s) at intervalMs cadence for up to durationMs. A probe carries a 4-byte magic (NPTC), a 4-byte sequence, and the first 16 bytes of SHA256(sessionId); both sides can compute the same session hash independently from the public sessionId, so no shared secret is needed on the punch path itself. On receiving a valid probe, a side replies with an NPTA ack. The first valid probe or ack seen from the far side records the working remote address and completes the attempt.

On timeout (attempt_timeout_secs as overall bound, punch_duration_ms as probe window), both sides issue NIP-9 deletes for their offer and answer events and report failure up to the discovery runtime's BootstrapEvent::Failed channel.

Phase 5 — Adoption

On success, the discovery runtime emits BootstrapEvent::Established carrying the session id, the punch socket, and the learned remote address. adopt_established_traversal() in the node lifecycle takes the socket, registers it with the UDP transport layer as a new transport instance, and calls initiate_connection() with the peer's FIPS identity as the expected remote. FMP's Noise IK handshake runs on the same socket — there is no "promote link" step between punch and handshake; the punch socket is the FMP socket.

From that moment on, the connection is a normal FMP link and is subject to the usual liveness (MMP heartbeats), rekey, and removal behavior. A link-dead event does not re-enter the discovery runtime automatically; reconnection relies on auto_reconnect and the same dial path that triggered the original punch.

Auto-connect semantics

Discovery does not itself initiate connections. It only supplies addresses. Dial attempts originate from the existing peer-connection machinery:

  • Configured peers (peers[] with connect_policy: auto_connect) are dialed on startup and on retry. When via_nostr is set, advert endpoints are appended to the dial list with lower priority than static entries.
  • Open discovery peers are assembled from the advert cache, fenced by the peer ACL, and enqueued into a bounded retry queue sized by open_discovery_max_pending. There is no event-driven "connect on every advert" — a peer re-enters the queue only when its prior attempt has drained.
  • Manual dials (fipsctl connect) can target any configured peer and use the same dial path, including Nostr resolution if configured.

Rate limits and safeguards

Mechanism Default What it prevents Behavior at limit
Offer semaphore (max_concurrent_incoming_offers) 16 CPU and memory exhaustion from offer spam on DM relays. Warn log, offer dropped.
Per-npub offer allowance (max_concurrent_offers_per_npub) 4 One sender identity holding every offer slot and denying traversal onboarding to everyone else. Does not prevent the same denial from several throwaway npubs. Debug log, offer dropped.
Advert cache (advert_cache_max_entries) 2048 Memory growth from ambient advert traffic under policy: open. LRU-by-expiry eviction.
Seen-sessions (seen_sessions_max_entries) 2048 Replay of stale sessionId values. Oldest entry evicted.
Signal TTL (signal_ttl_secs) 120 s Indefinite in-flight offers on relays. Expired offers rejected at validation.
Open discovery queue (open_discovery_max_pending) 64 Unbounded retry queue under ambient advert load. New candidates skipped until the queue drains.
Punch window (punch_duration_ms) 10 s Endless probe traffic after one side has given up. Attempt declared failed; sockets discarded.
Failure-streak threshold (failure_streak_threshold) 5 Repeated traversal attempts against a peer that keeps failing. Peer enters extended cooldown.
Extended cooldown (extended_cooldown_secs) 1800 s Tight retry loops after a failure streak. Per-peer suppression for the cooldown window.
WARN log throttle (warn_log_interval_secs) 300 s Log floods from a peer that fails on every attempt. One WARN per peer per interval; the rest demote to debug.
Failure-state cap (failure_state_max_entries) 4096 Memory growth from per-peer failure tracking. LRU eviction.

The load-shedding mechanisms (max_concurrent_incoming_offers and the failure-streak / extended-cooldown pair) are deliberately conservative so that a misbehaving relay cannot flood the node with offers and a chronically unreachable peer cannot keep the traversal pipeline saturated. The remaining rows are capacity bounds.

Adverts also undergo a stale-advert sweep: cached entries whose expiresAt has passed are evicted on the periodic prune tick. Inbound signaling tolerates ±60 s of clock skew between sender and receiver, and the runtime maintains an NTP-style skew estimate per remote so that consistently-skewed relays don't trip the freshness check.

Relay model

All configured relays (advert + DM) are opened on a single nostr-sdk::Client at startup. Publication is fan-out: the same event is sent to every relay in the target list, with no explicit retry or relay selection. Redundancy is implicit — a downed relay simply means its copy of the advert or signal is unavailable, while other relays still serve the same data.

For signaling specifically, the node prefers the recipient's NIP-17 DM relays when available (the recipient publishes its DM relay list as a kind 10050 event to its own DM relays on startup) and falls back to the local dm_relays list otherwise. This keeps the common case off the sender's DM relays when those are different from the recipient's, at the cost of one extra NIP-17 fetch per offer.

There is no per-relay rate limiting or health check. The relay model assumes that an operator chooses relays they trust to be best-effort available and that outright misbehavior is handled at the offer semaphore and replay-cache layers downstream.

Security and threat model

  • Relay operators can observe metadata. They see which npubs publish adverts, to whom offers are sent, and the timing of that traffic. The contents of offer and answer events are NIP-59/NIP-44 sealed — only the intended recipient decrypts them. Adverts are public by design.
  • STUN servers see the node's public IP and port. Only the STUN servers listed in the node's own stun_servers are ever contacted for reflexive discovery. Peer-advertised STUN values are informational; a malicious peer cannot steer this node to a chosen STUN target. See the doc comment on node.rendezvous.nostr.stun_servers.
  • The FIPS identity key signs adverts. Compromise of fips.key is compromise of the node's Nostr identity — an attacker can publish adverts on behalf of the node. The recovery path is the same as for any identity compromise: rotate the key and re-advertise. There is no separate Nostr keypair to rotate independently.
  • Tor advertising leaks timing via clearnet relays. When a Tor-only node advertises its onion address, the advert itself is published on clearnet WebSocket relays. Operators who want full unlinkability between the advertising identity and the node's IP must route relay traffic through Tor as well — for example by running fips inside a network namespace with a Tor SOCKS proxy as its only egress, or by pointing advert_relays and dm_relays at onion relay endpoints.
  • Open discovery accepts anyone publishing on the same app. Admission control is the peer ACL, not the discovery layer. Verify the ACL before enabling policy: open, and consider using a non-default app value to scope visibility.
  • Nothing about discovery bypasses FMP. A successful punch yields a UDP socket with a claimed remote identity. That identity is not trusted until FMP's Noise IK handshake completes. A peer whose advert says "I am npub X at 1.2.3.4:5678" but whose FMP handshake presents a different static key is rejected at the mesh layer.

LAN/mDNS discovery

LAN discovery is a separate, link-local discovery mechanism that finds peers on the same broadcast domain using mDNS / DNS-SD (RFC 6762 / RFC 6763). Unlike Nostr-mediated discovery, it contacts no relay, runs no STUN observation, and performs no NAT traversal: an endpoint learned from a LAN advert is by construction routable from the consumer's own link. The result is sub-second peer pairing on the same LAN.

It is unrelated to the "LAN candidate" terminology used in the NAT-traversal sections above (which refers to a host's own locally-bound address offered as a hole-punch candidate). LAN/mDNS discovery is a distinct subsystem under src/mdns/.

Role

LAN discovery adds two capabilities, both confined to the local link:

  • Advertising. The node publishes a _fips._udp.local. DNS-SD service advert carrying its npub, its protocol version, and (if configured) a discovery scope. The advert is multicast on the local link only; it does not leave the broadcast domain unless the operator's network bridges mDNS.
  • Browsing. The node concurrently browses for the same service type, learns the endpoints of other FIPS nodes on the link, and initiates a normal FMP link to each newly-seen peer.

The mDNS service type is _fips._udp.local. (src/mdns/mod.rs:45). Per RFC 6763 the _udp label denotes the IP transport used for the advert, not the FIPS upper protocol — both UDP and TCP FIPS endpoints announce under the same service type because the link-layer handshake travels over UDP either way. (In practice LAN discovery dials only over a UDP transport; see the handshake subsection.)

When to use it

  • You run several FIPS nodes on one LAN (a lab bench, an office segment, a home network) and want them to find each other without hand-maintaining peers[] blocks or standing up Nostr discovery.
  • You want the lowest-latency pairing path. Same-link pairing completes in well under a second with no relay round-trip.

Skip it when nodes are not on a shared broadcast domain (mDNS does not cross routed boundaries), or when you do not want the node to multicast its identity on the local link. LAN discovery is opt-in and disabled by default, so doing nothing leaves it off.

How it works

The LAN discovery runtime (src/mdns/mod.rs) is started during node initialization when node.rendezvous.lan.enabled is true. It is independent of Nostr discovery and runs even when Nostr is disabled (src/node/lifecycle/supervisor.rs:432-437). Startup requires an operational UDP transport: the node advertises the port of its lowest-TransportId operational, non-bootstrap UDP transport, chosen deterministically so the advertised port is stable across restarts (src/node/lifecycle/mod.rs:1598-1609). If no such port exists, the runtime returns NoAdvertisedPort and LAN discovery does not start (src/mdns/mod.rs:165-167).

The runtime does two things concurrently:

  1. Responder. It registers a DNS-SD service with instance name fips-<first-16-chars-of-npub> and a TXT record carrying the keys below. mdns-sd's address auto-detection appends every non-loopback interface address, with 127.0.0.1 seeded so same-host peers and integration tests can still resolve the advert (src/mdns/mod.rs:179-212).
  2. Browser. A background pump receives ServiceResolved events for the same service type. For each resolved advert it extracts the npub and scope TXT values, drops adverts that echo the node's own npub, drops cross-scope adverts (see scope filtering), drops records without an npub, and surfaces one LanDiscoveredPeer per routable interface address (src/mdns/mod.rs:230-297). IPv6 unicast link-local addresses without an interface scope id are skipped, since they cannot be dialed unambiguously (src/mdns/mod.rs:357-370).

The TXT record carries three keys (src/mdns/mod.rs:48-55):

TXT key Contents
npub bech32-encoded npub of the advertising node
scope the node's discovery scope, if one is configured (omitted otherwise)
v FIPS protocol version (the same PROTOCOL_VERSION used by the Nostr advert)

Once per node tick, the node drains browser events and acts on them in poll_lan_rendezvous() (src/node/lifecycle/mod.rs:1131, called from src/node/dataplane/rx_loop.rs:444). For each discovered peer it finds a UDP transport whose family matches the peer address, parses the npub into a PeerIdentity, skips peers it is already connected to or currently connecting to, and otherwise initiates a connection.

Handshake: Noise IK

LAN-discovered peers are dialed through the standard FMP outbound link path. poll_lan_rendezvous() calls initiate_connection() (src/node/lifecycle/mod.rs:448), which, for connectionless transports such as UDP, allocates a link and starts the Noise IK handshake (documented at src/node/lifecycle/mod.rs:438-442). This is the same link-layer handshake used by every other FMP connection — IK at the link layer per the FIPS architecture — not a different pattern for LAN peers.

The mDNS advert is unauthenticated: anyone on the link can multicast a TXT claiming any npub. Identity is proven end-to-end by the Noise IK handshake against the observed endpoint. A spoofed advert carrying another node's npub fails the handshake — the impostor does not hold the matching static key — and the half-open link is dropped. The mDNS advert is therefore a routing hint, never an identity assertion, exactly as a Nostr advert is treated (a successful contact is not trusted until FMP's Noise IK handshake completes).

Note: stale source doc-comments at src/mdns/mod.rs:14, 76, 153 describe this path as a "Noise XX" handshake. Those comments are inaccurate — the path uses Noise IK as described above. They are flagged for a separate source fix and do not reflect actual behavior.

Scope filtering

When a discovery scope is configured, the advert carries it in the scope TXT entry and the browser surfaces only peers whose advert carries a matching scope. Nodes on the same physical LAN but configured for different mesh networks therefore do not cross-feed each other.

The scope is resolved by lan_rendezvous_scope() (src/node/lifecycle/mod.rs:1104): the explicit node.rendezvous.lan.scope, if non-empty, is used directly. Otherwise the node falls back to deriving a scope from the Nostr discovery app tag (stripping the fips-overlay-v1: prefix when present). This lets an application keep its public, relay-visible Nostr app tag generic while still isolating LAN discovery per private network, or share one value across both. A node with no scope on either side surfaces all adverts it sees on the link.

Configuration

LAN discovery is configured under node.rendezvous.lan.* (src/config/node.rs:334, src/mdns/mod.rs:92-114):

Key Type Default Meaning
node.rendezvous.lan.enabled bool false Master switch. LAN discovery is opt-in; default-off avoids an unexpected per-link identity multicast on upgrade.
node.rendezvous.lan.service_type string _fips._udp.local. DNS-SD service type. Overridable mainly so integration tests can isolate multiple services on one loopback interface.
node.rendezvous.lan.scope string (optional) unset Application/network scope carried in the LAN-only scope TXT record. Kept deliberately separate from the public Nostr app tag. When unset, the scope falls back to the derived Nostr app value.

The identity surface published over mDNS (npub, version, optional scope) is a strict subset of what nostr.advertise already publishes publicly, so enabling LAN discovery adds no marginal privacy cost beyond making the node's presence observable on its own local link.

Relationship to Nostr discovery

The two mechanisms are complementary and independent:

Nostr-mediated LAN/mDNS
Reach Internet-wide, via relays Same broadcast domain only
Signaling channel Public Nostr relays mDNS multicast on the local link
NAT traversal STUN + UDP hole-punch for udp:nat peers None — endpoint is link-routable by construction
Identity carrier signed kind 37195 advert (authenticated at publish) unauthenticated mDNS TXT (routing hint only)
Identity proof FMP Noise IK on the connection FMP Noise IK on the connection
Default disabled (nostr.enabled: false) disabled (lan.enabled: false)
Scope key app tag (public) scope TXT (link-local), falls back to app

Both ultimately converge on the same trust boundary: discovery only supplies candidate endpoints, and no peer is trusted until FMP's Noise IK handshake confirms the claimed identity. A node may run both at once — for example, advertising globally over Nostr while also pairing instantly with same-LAN peers — with no interaction between the two beyond the shared scope fallback.

See also