feat(node): publish the DNS responder's bound address for embedders

An embedder that owns the TUN fd has no system DNS socket to point at the
built-in `.fips` responder. On Android specifically, `VpnService.Builder`
exposes `addDnsServer(address)` with no port — the OS resolver always uses
53, which an unprivileged app UID cannot bind — and it aims the resolver
*into* the tunnel, so `.fips` queries surface as IPv6/UDP packets on the
app's own fd rather than at any socket FIPS holds.

The app can still use the responder rather than reimplementing resolution:
lift the DNS payload out of the packet it read, send it to the responder
over an ordinary UDP socket of its own, and splice the answer back into a
reply packet. Nothing in the responder's start-up is desktop-specific —
`bind_dns_socket` is plain socket2, `lookup_mesh_ifindex` returns None with
no system TUN so the mesh filter self-disables, and `HostMapReloader` on an
absent hosts file settles at a no-op stat. What was missing is the address
to dial and whether anything is listening at it.

`dns_local_addr()` answers both, as a one-shot read taken after `start()`
returns and before the node is moved into a background task. That is the
only window in which an embedder running `run_rx_loop` holds a `&Node` to
call it on, and the value is settled by then: the responder is either up for
the rest of the node's life or it never came up. It reports the address read
back off the bound socket, so a `dns.port = 0` config yields the port the
kernel assigned rather than 0. Config alone cannot answer the second
question — `dns.enabled` with a failed bind leaves `bind_addr` naming a
plausible target nothing is listening on, and a bind failure only warns
rather than failing node start.

It is deliberately not a liveness feed. Watching a responder that dies later
needs a way to read live node state from a backgrounded `run_rx_loop`, which
is a general gap and not one an accessor should try to close.

Routing through the responder rather than resolving in the app is what keeps
route warming intact. Answering a `<npub>.fips` query is what puts that
peer's public key in the node's identity cache, and a FipsAddress is
SHA-256(pubkey) truncated twice: the key cannot be recovered from the IPv6
address. With no cache entry the first packet to a freshly-resolved name is
rejected with ICMPv6 "No route" — a failure that direct neighbours mask
entirely, since their identity arrives with the Noise handshake and never
needed resolving.

The address is retracted on `stop()`. `retract_child_publications` also
handles a responder that exits on its own at runtime, where the FSM's
`ChildExited` handling republishes node health but touches no per-child
handles. That consumer is dormant as written and documented as such:
`run_dns_responder` is an unconditional loop whose every failure arm
continues, so it never returns and the `Child::Dns` send after it is
unreachable. It lands here so a producer fix does not have to rediscover the
consuming side. A panicking responder is not covered either way, since the
unwind goes past the send rather than through it — true of every child
producer, not just this one.

The IPv6-adapter design doc gains an App-Owned DNS Path section beside the
App-Owned TUN one it mirrors, plus implementation-status rows for both.

Three tests. `dns_responder_serves_a_proxying_embedder` is the load-bearing
one: port-0 read-back, a proxied query answered with the right AAAA, and the
resolved identity arriving on the channel `run_rx_loop` drains into
`register_identity`. `dns_local_addr_stays_none_when_the_bind_fails` forces
`EADDRINUSE` against a socket the test holds open — `bind_dns_socket` sets
neither `SO_REUSEADDR` nor `SO_REUSEPORT`, so that is deterministic, where
naming a non-local address is not: `net.ipv4.ip_nonlocal_bind = 1` is
ordinary on hosts running keepalived or HAProxy and makes the bind succeed.
`retract_child_publications_clears_the_dns_address` is scoped and named for
the helper rather than the scenario, because deleting the `run_rx_loop` call
site leaves it green; that wiring is covered by nothing.

Full suite 1672 passed, 0 failed. fmt, `clippy --all-targets -D warnings`
and the Android `cargo ndk clippy --lib -D warnings` gate are clean.

The changelog entry was added at merge rather than in the pull request:
the app-owned TUN seam it mirrors gained an Unreleased entry in the
master-only sweep, so this one would otherwise recreate that debt.
This commit is contained in:
Arjen
2026-08-13 10:19:52 +00:00
committed by Johnathan Corgan
parent b3ecc1de4e
commit 672d828ef5
7 changed files with 345 additions and 1 deletions
+8
View File
@@ -51,6 +51,14 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
`fd00::/8`-destined packets and clamp TCP MSS on outbound SYNs. Desktop
builds are unchanged and no Cargo features are introduced.
- `Node::dns_local_addr()`, the DNS companion to the app-owned TUN seam above.
An embedder whose resolver is pointed into the tunnel has no system socket
aimed at the built-in `.fips` responder, so the accessor reports the address
read back off the bound socket: `dns.port = 0` therefore yields the
kernel-assigned port, and it returns `Some` only while the responder is up.
Read it once, after `start()` returns and before the node is moved into a
background task; it is not a liveness feed (#136).
- A bounded graceful-shutdown drain phase, controlled by the new
`node.drain_timeout_secs` (default 2s). On the shutdown signal the node
broadcasts Disconnect to all peers and then keeps serving for that window,
+43
View File
@@ -330,6 +330,47 @@ and the [TUN-Side TCP MSS Clamping](#tun-side-tcp-mss-clamping). The embedder is
therefore responsible for routing only `fd00::/8` to its TUN (so only mesh-bound
packets arrive) and for clamping TCP MSS on outbound SYNs.
### App-Owned DNS Path (embedded hosts)
An embedded host that owns the TUN fd generally has no system DNS socket to aim
at the responder either. On Android, `VpnService.Builder.addDnsServer()` takes an
address with no port — the resolver always uses 53, which an unprivileged app UID
cannot bind — and it points the resolver *into* the tunnel, so `.fips` queries
surface as IPv6/UDP packets on the app's own fd rather than at any socket FIPS
holds. The [DNS responder](#dns-integration) itself needs no changes for this:
its bind is a plain UDP socket, and with no system TUN the
[mesh-interface filter](#mesh-interface-query-filter) self-disables because the
interface name does not resolve.
`Node::dns_local_addr()` closes the gap. It reports the address read back off the
bound socket — so a `dns.port = 0` config yields the port the kernel assigned —
and is `None` when no responder came up. The embedder lifts the DNS payload out
of the packet it read, sends it to that address over an ordinary UDP socket of
its own, and splices the answer back into a reply packet.
It is a one-shot read taken after `start()` returns and before the node is moved
into a background task, because `run_rx_loop` then borrows the node exclusively
for its whole lifetime and no `&Node` remains to call it on. By that point the
value is settled: the responder is either up for the rest of the node's life or
it never came up.
Proxying to the responder rather than resolving in the app is what keeps the
[identity cache](#identity-cache) warm. Answering a `<npub>.fips` query is what
registers that peer's public key, and a FIPS address is a truncated hash of a
hash of the pubkey — the key cannot be recovered from the IPv6 address alone. An
app that resolves the AAAA itself leaves the cache empty, and the first packet to
the resolved name is rejected with ICMPv6 Destination Unreachable. Direct
neighbours mask the omission, since their identity arrives with the Noise
handshake and never needed resolving.
The address is retracted on `stop()`. A retraction hook for a responder that
exits on its own at runtime is wired on the consuming side, but is dormant:
`run_dns_responder` never returns, so nothing produces the `Child::Dns` exit
event it consumes. Watching a responder that dies mid-run is therefore not
something an embedder can do today, and would in any case need a way to read
live node state from a backgrounded `run_rx_loop` — a general gap rather than a
DNS-specific one.
## Implementation Status
| Feature | Status |
@@ -343,6 +384,8 @@ packets arrive) and for clamping TCP MSS on outbound SYNs.
| TCP MSS clamping (SYN + SYN-ACK) | **Implemented** |
| DNS service (.fips domain) | **Implemented** |
| DNS responder mesh-interface filter | **Implemented** |
| App-owned TUN (`Node::enable_app_owned_tun`) | **Implemented** |
| App-owned DNS path (`Node::dns_local_addr`) | **Implemented** |
| Port-based service multiplexing (port 256) | **Implemented** |
| IPv6 header compression (format 0x00) | **Implemented** |
| Per-destination route MTU (netlink) | Planned |
+5
View File
@@ -271,6 +271,11 @@ impl Node {
// `PublishState`; other variants are ignored defensively.
maybe_child = child_exit_rx.recv() => {
if let Some(child) = maybe_child {
// Drop anything the dead child published for embedders
// (e.g. the DNS responder's bound address) before
// republishing health, so nothing outside the node can
// observe an address the listener no longer answers on.
self.retract_child_publications(child);
let actions = self
.supervisor
.fsm
+40 -1
View File
@@ -1747,6 +1747,12 @@ impl Node {
let bind = std::net::SocketAddr::new(ip, self.config().dns.port());
match Self::bind_dns_socket(bind) {
Ok(socket) => {
// Read the bound address back off the socket
// rather than reusing `bind`: a port-0 config
// resolves to the kernel-assigned port here,
// and this is the address an embedder that
// proxies queries to us has to dial.
let local_addr = socket.local_addr().unwrap_or(bind);
let dns_channel_size = self.config().node.buffers.dns_channel;
let (identity_tx, identity_rx) =
tokio::sync::mpsc::channel(dns_channel_size);
@@ -1771,7 +1777,7 @@ impl Node {
let mesh_ifindex =
Self::lookup_mesh_ifindex(self.config().tun.name());
info!(
bind = %bind,
bind = %local_addr,
hosts = reloader.hosts().len(),
mesh_ifindex = ?mesh_ifindex,
"DNS responder started for .fips domain (auto-reload enabled)"
@@ -1797,6 +1803,7 @@ impl Node {
});
self.supervisor.dns_identity_rx = Some(identity_rx);
self.supervisor.dns_task = Some(handle);
self.supervisor.dns_local_addr = Some(local_addr);
Event::SubstrateUp { child }
}
Err(e) => {
@@ -2057,6 +2064,10 @@ impl Node {
handle.abort();
debug!("DNS responder stopped");
}
// Retract the published address in the same step that kills
// the listener, so an embedder polling `dns_local_addr()`
// never dials a socket that is already gone.
self.supervisor.dns_local_addr.take();
}
Child::Nostr => {
// Stop Nostr overlay discovery background work and withdraw
@@ -2153,6 +2164,34 @@ impl Node {
}
}
/// Retract anything a child published about itself, after it exited on its
/// own at runtime (as opposed to being torn down by [`Self::stop`]).
///
/// The FSM's `ChildExited` handling only republishes node health; it does
/// not touch per-child handles. That is fine for state nobody outside the
/// node reads, but not for an address an embedder dials: a stale
/// [`Node::dns_local_addr`] would have the app proxying `.fips` queries
/// into a socket that is gone, and the only symptom would be resolution
/// quietly timing out.
///
/// Deliberately narrow — it clears published facts, not handles.
/// `dns_task` is left alone because [`Self::reconstruct_supervised_up`]
/// reads it to rebuild the teardown set, and aborting an already-finished
/// handle there is harmless.
///
/// **Dormant for `Dns` as written.** `run_dns_responder` is an unconditional
/// loop whose every failure arm continues, so it never returns and the
/// `Child::Dns` send that follows it is unreachable — nothing produces the
/// event this consumes. The consumer side is correct and lands here so the
/// producer fix does not have to rediscover it. A responder that *panics* is
/// not covered either way, since the unwind goes past the send rather than
/// through it; that is true of every child producer, not just this one.
pub(in crate::node) fn retract_child_publications(&mut self, child: Child) {
if matches!(child, Child::Dns) {
self.supervisor.dns_local_addr.take();
}
}
/// Reconstruct the supervised up-set from observed runtime presence, so the
/// FSM authors the teardown order regardless of how the node reached
/// `Running`. Worker pools are deliberately excluded: today's teardown never
+6
View File
@@ -692,6 +692,11 @@ pub(crate) struct Supervisor {
pub(in crate::node) dns_identity_rx: Option<crate::upper::dns::DnsIdentityRx>,
/// DNS responder task handle.
pub(in crate::node) dns_task: Option<tokio::task::JoinHandle<()>>,
/// Address the DNS responder actually bound, read back from the socket
/// after `bind` so a port-0 config resolves to the assigned port. `Some`
/// only while the responder is up; published to embedders through
/// [`Node::dns_local_addr`](crate::Node::dns_local_addr).
pub(in crate::node) dns_local_addr: Option<std::net::SocketAddr>,
/// Node-side driver state for the Nostr overlay peer-rendezvous
/// subsystem: the engine handle, its startup timestamp, the one-shot
@@ -737,6 +742,7 @@ impl Supervisor {
tun_shutdown_fd: None,
dns_identity_rx: None,
dns_task: None,
dns_local_addr: None,
nostr_rendezvous: crate::nostr::RendezvousDriver::default(),
lan_rendezvous: None,
#[cfg(unix)]
+50
View File
@@ -2874,6 +2874,56 @@ impl Node {
(outbound_tx, tun_rx)
}
/// Address the built-in `.fips` DNS responder is listening on, or `None`
/// when it is not running (`dns.enabled = false`, the bind failed, or the
/// node is stopped).
///
/// This is the companion to [`Self::enable_app_owned_tun`] for embedders
/// that own the TUN fd. On a platform with no system DNS socket to point
/// at us — an Android `VpnService`, whose `addDnsServer()` takes an address
/// with no port and aims the OS resolver *into* the tunnel — `.fips`
/// queries arrive as IPv6/UDP packets on the app's own fd. The app can
/// forward the DNS payload here and splice the answer back into a reply
/// packet, rather than reimplementing resolution:
///
/// ```no_run
/// # async fn f(node: &fips::Node, query: &[u8]) -> std::io::Result<()> {
/// let Some(dns) = node.dns_local_addr() else { return Ok(()) };
/// let sock = tokio::net::UdpSocket::bind("[::1]:0").await?;
/// sock.send_to(query, dns).await?; // payload only, no IP/UDP header
/// let mut answer = [0u8; 512];
/// let (n, _) = sock.recv_from(&mut answer).await?;
/// # let _ = n; Ok(())
/// # }
/// ```
///
/// Going through the responder rather than resolving in the app is what
/// keeps route warming working: answering a `<npub>.fips` query is what
/// populates the node's identity cache with that peer's public key, and a
/// `FipsAddress` is a truncated hash — the pubkey cannot be recovered from
/// the IPv6 address alone. Without a cache entry the first packet to a
/// freshly-resolved name is rejected with ICMPv6 "No route". Direct
/// neighbours mask this, since their identity comes from the Noise
/// handshake and never needed resolving.
///
/// Read this **once, after [`Self::start`] returns and before the node is
/// moved into a background task** — that is the only window in which an
/// embedder running [`Self::run_rx_loop`] holds a `&Node` to call it on, and
/// the value is fixed by then: the responder is either up for the rest of
/// the node's life or it never came up. The address is read back off the
/// bound socket, so a `dns.port = 0` config reports the port the kernel
/// actually assigned.
///
/// This is a one-shot read, not a liveness feed. `None` distinguishes "no
/// responder" from "responder at this address" at that moment; it is not a
/// signal an embedder can watch for a responder that dies later, because
/// `run_rx_loop` borrows the node exclusively for its whole lifetime.
/// Reading live node state from a backgrounded loop is a general gap, not
/// one this accessor tries to close.
pub fn dns_local_addr(&self) -> Option<std::net::SocketAddr> {
self.supervisor.dns_local_addr
}
// === Sending ===
/// Encrypt and send a link-layer message to an authenticated peer.
+193
View File
@@ -2798,6 +2798,199 @@ async fn start_skips_system_tun_when_app_owned() {
node.stop().await.unwrap();
}
/// The embedder-facing DNS contract, end to end.
///
/// An embedder that owns the TUN fd (Android `VpnService`) has no system DNS
/// socket to point at us, so it proxies `.fips` query payloads it lifts out of
/// its own tunnel to the built-in responder. That requires three things to
/// hold, and this pins all three:
///
/// 1. `dns_local_addr()` publishes where to send — read back off the bound
/// socket, so a `port = 0` config reports the assigned port, not 0.
/// 2. The responder answers a proxied query with the right AAAA.
/// 3. The resolved identity reaches `dns_identity_rx` — the channel
/// `run_rx_loop` drains into `register_identity`. This is the leg that
/// populates the identity cache, without which the first packet to a
/// freshly-resolved `<npub>.fips` is rejected with ICMPv6 "No route".
#[tokio::test]
async fn dns_responder_serves_a_proxying_embedder() {
let mut config = crate::Config::new();
config.transports.udp = crate::config::TransportInstances::Single(crate::config::UdpConfig {
bind_addr: Some("127.0.0.1:0".to_string()),
..Default::default()
});
config.dns.enabled = true;
config.dns.bind_addr = Some("::1".to_string());
// Port 0: proves the address is read back off the socket rather than
// echoed from config — an embedder dialling 0 would reach nothing.
config.dns.port = Some(0);
// The TUN is app-owned, as it is on the platform this seam serves.
let mut node = make_node_with(config);
let (_outbound_tx, _tun_rx) = node.enable_app_owned_tun();
assert!(
node.dns_local_addr().is_none(),
"no responder before start()",
);
node.start().await.unwrap();
let dns_addr = node
.dns_local_addr()
.expect("responder is up, so its address is published");
assert_ne!(dns_addr.port(), 0, "must report the kernel-assigned port");
// Proxy a query the way the embedder would: payload only, no IP/UDP header
// (it strips those off the packet it read from its own TUN fd).
let peer = Identity::generate();
let query = {
use simple_dns::{CLASS, Name, Packet, QCLASS, QTYPE, Question, TYPE};
let mut packet = Packet::new_query(0x1234);
packet.questions.push(Question::new(
Name::new_unchecked(&format!("{}.fips", peer.npub())).into_owned(),
QTYPE::TYPE(TYPE::AAAA),
QCLASS::CLASS(CLASS::IN),
false,
));
packet.build_bytes_vec().unwrap()
};
let client = tokio::net::UdpSocket::bind("[::1]:0").await.unwrap();
client.send_to(&query, dns_addr).await.unwrap();
let mut buf = [0u8; 512];
let (len, _) = tokio::time::timeout(
std::time::Duration::from_secs(2),
client.recv_from(&mut buf),
)
.await
.expect("responder answered within the timeout")
.unwrap();
let answer = simple_dns::Packet::parse(&buf[..len]).expect("well-formed DNS response");
let rdata = &answer.answers.first().expect("one AAAA answer").rdata;
let simple_dns::rdata::RData::AAAA(aaaa) = rdata else {
panic!("expected an AAAA record, got {rdata:?}");
};
assert_eq!(
std::net::Ipv6Addr::from(aaaa.address),
peer.address().to_ipv6(),
"AAAA must be the peer's FipsAddress",
);
// The identity leg. `run_rx_loop` owns the node for its whole life, so the
// embedder cannot register identities itself — the responder publishes them
// on this channel instead. Drain and register exactly as the rx-loop arm in
// `dataplane/rx_loop.rs` does, then assert the cache is populated.
let identity = tokio::time::timeout(
std::time::Duration::from_secs(2),
node.supervisor
.dns_identity_rx
.as_mut()
.expect("responder installed the identity receiver")
.recv(),
)
.await
.expect("identity published within the timeout")
.expect("channel is open");
assert_eq!(identity.node_addr, *peer.node_addr());
node.register_identity(identity.node_addr, identity.pubkey);
assert!(
node.has_cached_identity(peer.node_addr()),
"resolving a name must warm the identity cache, or the first packet \
to that address is rejected with ICMPv6 \"No route\"",
);
node.stop().await.unwrap();
assert!(
node.dns_local_addr().is_none(),
"the published address must be retracted with the listener",
);
}
/// `retract_child_publications(Dns)` clears the published address.
///
/// Scoped to the helper deliberately, and named for that rather than for the
/// scenario: no responder dies here, and deleting the `run_rx_loop` call site
/// leaves this green. Driving a real exit through the loop needs the node moved
/// into a task, which puts `dns_local_addr()` out of reach — and the producer
/// side cannot deliver `Child::Dns` today regardless, since `run_dns_responder`
/// never returns.
///
/// What it does pin is the behavior the eventual wiring depends on: the FSM's
/// `ChildExited` handling only republishes node health, so without this
/// retraction `dns_local_addr()` would keep naming a socket nobody is listening
/// on, and a proxying embedder would see `.fips` queries silently time out
/// rather than any error it could act on.
#[tokio::test]
async fn retract_child_publications_clears_the_dns_address() {
let mut config = crate::Config::new();
config.transports.udp = crate::config::TransportInstances::Single(crate::config::UdpConfig {
bind_addr: Some("127.0.0.1:0".to_string()),
..Default::default()
});
config.dns.enabled = true;
config.dns.bind_addr = Some("::1".to_string());
config.dns.port = Some(0);
let mut node = make_node_with(config);
node.start().await.unwrap();
assert!(node.dns_local_addr().is_some(), "responder came up");
// What `run_rx_loop` does when the DNS task self-reports its exit.
node.retract_child_publications(crate::node::lifecycle::supervisor::Child::Dns);
assert!(
node.dns_local_addr().is_none(),
"a dead responder must not keep publishing an address to dial",
);
node.stop().await.unwrap();
}
/// `dns.enabled` with a bind that fails must report `None`, not an address.
///
/// This is the third state an embedder has to tell apart, and the one that
/// would otherwise be indistinguishable from a healthy responder by reading
/// config alone: DNS is switched on, so `config.dns.bind_addr()` names a
/// plausible target, but nothing is listening there. A bind failure is only
/// warned about and leaves the node running, so config is not evidence —
/// `dns_local_addr()` is.
///
/// The failure is forced with `EADDRINUSE` against a socket this test holds
/// open, rather than by naming an address the host has no interface for.
/// `bind_dns_socket` sets neither `SO_REUSEADDR` nor `SO_REUSEPORT`, so the
/// collision is deterministic on Linux and macOS. A non-local address is not:
/// `net.ipv4.ip_nonlocal_bind = 1` is ordinary on hosts running keepalived or
/// HAProxy and makes the bind succeed, which reds the test on a developer
/// machine while CI — at the default `0` — stays green.
#[tokio::test]
async fn dns_local_addr_stays_none_when_the_bind_fails() {
// Hold the port for the whole test so the responder's bind collides.
let squatter = tokio::net::UdpSocket::bind("[::1]:0").await.unwrap();
let taken = squatter.local_addr().unwrap();
let mut config = crate::Config::new();
config.transports.udp = crate::config::TransportInstances::Single(crate::config::UdpConfig {
bind_addr: Some("127.0.0.1:0".to_string()),
..Default::default()
});
config.dns.enabled = true;
config.dns.bind_addr = Some("::1".to_string());
config.dns.port = Some(taken.port());
let mut node = make_node_with(config);
node.start().await.unwrap();
assert!(
node.dns_local_addr().is_none(),
"an unbound responder must not publish an address",
);
node.stop().await.unwrap();
drop(squatter);
}
/// A connection whose handshake failed is retained with BOTH Noise handles
/// empty, and the stale-connection sweep depends on that: presence of the
/// pending connection — not presence of a handle — is what marks a machine as