`AndroidIo` captured the injected bridge `Arc` at construction, and the BLE
transport was only built if a bridge already existed when the node started.
Both facts are wrong for the platform this backend serves: the radio is owned
by an Android foreground service whose lifetime is independent of the node's —
it can start *after* the node, and it mints a fresh bridge every time it
starts. So the only way for a running node to adopt a radio was to stop and
rebuild it, which drops every peer, every session, and every route with it.
Turning Bluetooth on took the whole mesh down for as long as re-handshaking
took, which is not what enabling a transport should cost.
`AndroidIo` now resolves the process-wide bridge per operation, so a fresh
bridge is picked up in place, and the transport is constructed whether or not
one exists yet. Operations attempted with no radio present return a transport
error instead of being unreachable, and recover on their own once a radio
appears. `AndroidIo::new` is kept for tests that drive a specific bridge.
Live streams still hold the radio they were opened on rather than migrating to
a new one — a channel belongs to the socket that created it, and those die with
the radio that owned them.
The UDP transport opens one wildcard socket and selects the egress path per
destination address. That assumes the host routes by destination alone —
true on an ordinary Unix box, not true everywhere. On hosts that associate
each socket with one "network" and steer traffic by that association, a peer
reachable only over a secondary network is unreachable in a way FIPS cannot
see or fix: the address is well-formed, the send succeeds, the peer receives
our handshake and replies, and the reply is discarded by the host before it
reaches our socket. The link just retries msg1 forever.
Nothing in FIPS can correct that, because the correction is a socket option
on a socket the transport keeps private. So hand the embedder the fd:
let rx = node.enable_app_owned_udp_fd(); // before start()
node.start().await?;
let fd = rx.recv()?; // once the transport is up
This follows the existing `enable_app_owned_tun` / `enable_app_owned_dns`
contract — call after `Node::new` and before `start()`, get a channel back —
and stays deliberately narrow: FIPS keeps owning the socket, and the fd
carries no promise beyond "this is the transport's socket, now open". If no
UDP transport is configured, or it fails to start, nothing is ever sent.
Also plumbs `raw_fd()` through `TransportHandle` and `UdpTransport`. Unix
only, since `RawFd` is a unix concept; non-UDP transports return None.
A mesh-edge node (only direct BLE / Wi-Fi Aware / LAN peers, sparse bloom
filters) could not route to its own directly-connected neighbours: a
neighbour is never in another peer's bloom filter, so discovery was gated
off and find_next_hop had no coordinates. Traffic to such a peer failed
with "no route to destination" even with an established session.
- discovery: don't suppress a lookup on a bloom miss unless we have >2
peers (trustworthy blooms); and when no tree peer advertises the target,
flood the LookupRequest to every sendable peer — including the target
itself, which answers a lookup for its own address, so the querier learns
its coordinates and can route.
- session: when an established-session send hits no-route, trigger discovery
(previously only session *initiation* did), so a coordless session — e.g.
a platform-pushed peer whose Noise handshake came up before tree
discovery — self-warms and the app's retransmit succeeds.
- node: add Node::enable_app_owned_dns() — returns the DnsIdentityTx an
embedder (Android VpnService pump) uses to register identities it resolved
itself, giving the same identity-cache/route-warming the built-in DNS
responder provides.
When a peer is reachable over more than one transport at once, prefer the
faster one and hold it. roam_current_addr upgrades eagerly to a higher-
preference transport (BLE -> Wi-Fi Aware / UDP) but falls back to a lower
one only after the preferred link has been silent past a hysteresis
window (~2 heartbeat intervals), so a stray BLE keepalive can't drag an
active Aware session back to BLE. poll_transport_discovery additionally
stops re-probing an alternate path strictly worse than the one the peer
is already on, so BLE rediscovery no longer keeps grabbing a peer that
belongs on Aware. Equal preferences reduce to plain last-authenticated-
packet-wins roaming, so single-transport deployments are unaffected.
Add a transport-agnostic seam for an embedding platform (e.g. an Android
Wi-Fi Aware radio) to push "peer npub reachable at addr over transport T"
events into a running node — the generalization of the UDP-only LAN mDNS
drain. A process-global queue (fips::discovery::platform) is drained each
tick by poll_platform_discovery, which selects the transport family-aware
(an IPv6 target picks an IPv6 socket) and initiates a Noise IK handshake;
the pushed npub is only a routing hint, the handshake authenticates. An
event for an already-active peer starts an alternate-path handshake, and
a Lost event closes the pooled connection.
sockaddr_to_socket_addr dropped the IPv6 scope id when converting an
inbound source address, so a reply to a link-local peer (fe80::/10) had
no interface scope and could not be routed. This stalled the Noise
handshake reply (msg2) over a Wi-Fi Aware NDP interface, where every
address is link-local: msg1 reached the peer (sent to a scoped address)
but msg2 never routed back. Rebuild the SocketAddrV6 with sin6_scope_id;
non-link-local sources carry scope 0 and are unaffected.
BLE delivery was unreliable on every backend except BlueZ. The receive
path assumed one `recv()` returned exactly one whole FIPS packet, which
only holds for BlueZ's SOCK_SEQPACKET (boundary-preserving). Android's
BluetoothSocket input stream and macOS CoreBluetooth are byte-stream
oriented: a read can return a fragment of a packet or several packets
coalesced. Under the old loop a fragment was shipped up as a runt packet
(rejected by FMP/Noise) and a coalesced tail was silently truncated and
dropped — so packets were lost and the transport thrashed.
Recover packet boundaries from the byte stream instead of trusting the
OS to preserve them. FIPS packets are self-delimiting via the 4-byte FMP
common prefix, so this reuses the exact length-prefixed framer TCP already
uses (`tcp::stream::read_fmp_packet`) — kept transport-agnostic for this
reason. A new `BleStreamRead` adapter turns the datagram-shaped `BleStream`
into the `AsyncRead` that framer expects, buffering leftover bytes across
reads. The recv future owns its scratch buffer and returns an owned Vec so
it is `'static` and storable across `poll_read` calls.
One reader is threaded through the pubkey exchange and the receive loop per
connection, so bytes a peer coalesces after the 33-byte pubkey stay buffered
rather than being lost at the handoff. The pubkey exchange now reads via
`read_exact` (reassembles a fragmented pubkey) instead of a brittle single
`recv()` with an exact-length check.
On BlueZ this is a transparent pass-through (one recv already equals one
packet); on stream backends it reassembles. Either way the layer above sees
one whole packet per read, identically on every platform.
Also bump the Android outbound SEND_QUEUE_CAP from 8 to 32. Swept against
the peer speedtest: 8 starved the radio's connection events (~half
throughput), 64 bufferbloated TCP, 32 was best (~200/500 kbps up/down). The
sweep was noisy and non-monotonic — run-to-run BLE variance (RF, 2M PHY /
connection-priority grants) rivals the knob's effect — so 32 is the
best-observed value pending re-validation with PHY/interval instrumentation,
not a proven optimum.
Import ordering and multi-line wrapping rustfmt would enforce; android_io
is cfg(target_os = "android") so it never compiled on the host CI matrix
and the drift went unnoticed until the mobile build was wired up.
The per-stream outbound byte queue was 256 deep (~384 KB) on a link whose
bandwidth-delay product is ~1 packet. FSP/MMP filled it, RTT ballooned to several
seconds, and TCP above couldn't ramp. A shallow tail-drop queue is no better — it
sheds packets and collapses TCP throughput.
Give the outbound queue its own shallow cap (SEND_QUEUE_CAP=8) and make
BleStream::send backpressure — await a free slot rather than try_send-dropping —
so flow control propagates up through FSP/MMP to TCP. On-device this cut RTT from
~5s to ~1.2s with throughput preserved and 0% loss.
Per-peer PSM discovery keys the learned PSM by BLE address, but RPAs rotate
between the scan that learns a peer's PSM and the dial, so resolve() missed and
fell back to the legacy default 0x0085 — which no node listens on. Every outbound
L2CAP connect was rejected and the link ran inbound-only at a fraction of its rate.
A node advertises exactly one OS-assigned listener PSM, so on an exact-address
miss, dial the most-recently-learned PSM instead of the fixed default; a wrong
guess only costs a dial-retry.
Node::enable_app_owned_tun() lets an embedder that owns the TUN fd (e.g. an
Android VpnService) exchange IPv6 packet bytes with FIPS over channels instead of
FIPS creating a system TUN device. It returns (app_outbound_tx, app_inbound_rx):
the embedder pushes packets read from its fd into the outbound sender (app ->
mesh) and pulls packets destined for its fd from the inbound receiver (mesh ->
app). Called after Node::new and before start(), mirroring control_read_handle().
start() now gates TunDevice::create on tun_tx being unset, so when the app-owned
channels are pre-installed it skips system-TUN creation and does no system-TUN
ops. The inbound IPv6-shim delivery already writes to tun_tx, and run_rx_loop
already drains tun_outbound_rx into handle_tun_outbound, so both directions reuse
the existing wiring.
Packets entering via app_outbound_tx bypass the system-TUN reader's
handle_tun_packet, so the embedder must push only fd::/8-destined packets (FIPS no
longer filters the destination) and clamp TCP MSS on outbound SYNs; the rustdoc
and the IPv6-adapter design doc spell this out.
Tests: app_owned_tun_seam_wires_channels covers the channel round-trip and the
Active state; start_skips_system_tun_when_app_owned runs start() and asserts no
named system device is created (tun_name stays unset).
deliver_scan now carries RSSI as well, and records the latest advert per address
(PSM + RSSI) into a map exposed via advert_views() -> Vec<AdvertView>. Cleared
each scan cycle (addresses rotate with MAC privacy). This is the radio-level
view, distinct from the node's mesh-level peer table.
- next_send: clone the per-channel receiver out before blocking (no longer holds
the channels lock across recv_timeout), and treat a dropped sender
(Disconnected) as closed so the Kotlin writer thread exits when the stream is
gone.
- channel_open: also reports a closed (dropped-stream) channel, not just a
removed one — so writers stop promptly.
- set_android_ble_bridge: replaceable (Mutex<Option>) so a stop→start cycle
re-injects a fresh bridge for the rebuilt node.
Expose ControlReadHandle (was pub(crate)) + make Node::control_read_handle()
public, and add ControlReadHandle::peer_views() -> Vec<PeerView>
{ node_addr_hex, npub, connected }, read lock-free from the tick-published stats
snapshot (peer_meta). This lets an embedder run run_rx_loop on a background task
(which exclusively borrows &mut Node) and still poll live peer state from a
cloned handle — no &Node access needed. Used by the Myco app's developer UI.
The Android BLE radio lives in Kotlin, so AndroidIo/Stream/Acceptor/Scanner
implement BleIo by delegating to AndroidBleBridge — the channel machinery shared
with the JNI layer in the embedder. The AndroidRadio trait is the object-safe
command surface Kotlin implements (listen/connect/advertise/scan/close).
- Inbound bytes/events are pushed non-blocking into tokio channels
(deliver_recv/inbound/scan/connect_result); outbound bytes are pulled by a
per-channel Kotlin writer thread via next_send. BleStream::send never calls
JNI — the byte hot path is pure channel push.
- Per-peer PSM: connect() substitutes the learned PSM; advertising emits the
OS-assigned listener PSM (16-bit LE service-data).
- Wired as DefaultBleTransport on Android, plus a node construction arm that
reads the embedder-injected bridge (set_android_ble_bridge).
Pure Rust (no JNI here — that's in myco-core), so the channel logic unit-tests
on the host with a mock radio. Cross-compiles clean for arm64-android.
Foundation for "BLE v2" universal per-peer PSM discovery, shared by every
BleIo backend.
- ble/psm.rs: a 16-bit little-endian service-data PSM codec + a short-lived
BleAddr->PSM map (learn/lookup/resolve/clear), with unit tests. Every node
advertises its OS-assigned listener PSM and every dialer reads a peer's
advertised PSM before connect(), replacing the fixed DEFAULT_PSM (0x0085)
assumption that only BlueZ could satisfy (Android/macOS get OS-assigned
listener PSMs).
- Un-gate the BLE transport module from linux-only to a new `ble_available`
cfg (linux/macos/android). The pool/discovery/psm logic and the generic
BleTransport<I> are platform-agnostic; only the concrete BleIo backend is
platform-specific (BluerIo on linux-glibc, else the MockBleIo fallback).
The macOS BluestIo and Android AndroidIo backends, and the per-backend
advertise/scan/connect wiring of the PSM map, land in follow-ups.
A plain `cargo build` now compiles correctly for every target with no
flags: the compiler selects platform code by `target_os`, enabling what
each platform supports and disabling what it doesn't. Android no longer
needs `--no-default-features`.
- ethernet: gated `any(target_os = "linux", target_os = "macos")` (raw
AF_PACKET / BPF). Android is `target_os = "android"` — not "linux" — so
the raw-socket transport self-excludes there, as it already did on Windows.
- system-tun: real platform ops gated per `target_os = "linux"`/`"macos"`;
Windows keeps its own path. Android gets a no-op stub — the TUN is
app-owned by the embedder (e.g. an Android VpnService).
- dns: ipi6_ifindex is i32 on Android, u32 on macOS — cast.
- node: gate the unix-only ESTABLISHED_HEADER_SIZE import so the Windows
build is warning-clean.
No Cargo features are introduced; desktop builds are unchanged.
handle_peer_removal_tree_cleanup reparents or self-roots the node when
its parent link drops, but omitted the coordinate-cache invalidation that
every other position-change path performs. Cached entries for downstream
destinations kept the node's now-stale coordinate prefix, and because
find_next_hop refreshes the TTL on every routing access, an actively
routed stale entry never self-expired — corrected only by a fresh insert.
Mirror the loop-detection branch: inside the parent-loss changed block,
invalidate both classes — invalidate_via_node (reparent) and
invalidate_other_roots (self-root). Add regression tests for a parent-link
removal that reparents (via-node entry dropped, same-root sibling
preserved) and one that self-roots (both via-node and stale old-root
entries dropped).
Add six forwarding counters that partition transit-forwarded packets by their
tree relationship to the chosen next hop: tree-up (peer is our ancestor),
tree-down (peer is our descendant and the destination is within its subtree),
tree-down-cross (peer is our descendant but the destination is outside its
subtree), cross-link descend (lateral peer, destination within its subtree),
cross-link ascend (lateral peer, destination outside its subtree), and
direct-peer. The six classes sum to forwarded_packets, asserted by a unit
test. Classification is computed from tree coordinates at the transit
chokepoint, so the error-signal routing callers are excluded.
The two "outside the chosen peer's subtree" classes are both up-and-over
forwards but differ in what they depend on. Tree-down-cross is the
dive-to-tree-child cut-through: we forward down to our own child for a
destination not beneath it, which is only possible because the child
advertised cross-link reach upward to us, beyond its own subtree. Its count
measures how much forwarding depends on that upward advertisement, i.e. what
would change if cross-link advertisements were narrowed to subtree-entry only.
Cross-link ascend, by contrast, uses the node's own lateral cross-link learned
from a peer's split-horizon advertisement, so it does not depend on any upward
advertisement.
Surface the counters through the forwarding stats snapshot (control socket,
show_routing and show_status) and reorganize the fipstop routing tab so its
two columns separate own/endpoint traffic (received, delivered, originated)
from forwarded/transit traffic (the route-class breakdown and drop reasons),
with the tree-down-cross line visually flagged.
Self-addressed TCP/UDP connections to a node's own <npub>.fips address
half-opened and hung on macOS. macOS routes self-traffic as loopback (a
LOCAL route via lo0), which defers the transport TX checksum, but the
point-to-point utun then egresses the packet into the daemon with only
the pseudo-header partial checksum present. The hairpin path added in
9a9e90a re-injected these verbatim, so the local stack dropped every
segment whose checksum MSS clamping didn't happen to rewrite: the
SYN/SYN-ACK got through (clamping recomputes them) but the bare ACK,
data, and FIN were dropped for a bad checksum, leaving the listener
stuck in SYN_RCVD.
Recompute the TCP/UDP checksum for self-addressed packets on the hairpin
path before re-injection, completing the self-delivery 9a9e90a started
(which only covered ICMP and the TCP handshake). Linux is unaffected: it
loops self-traffic via lo before the TUN, so the hairpin branch never
fires and checksums are already valid.
Confirmed on macOS: a self-connect that previously timed out now
completes in ~9ms with payload delivered.
A packet destined for our own mesh address reached the TUN reader on
macOS (utun egresses self-traffic into the daemon despite the lo0 host
route) and was pushed onto the mesh outbound path, where it was dropped
for lack of a session/route to self. Hairpin self-addressed packets back
to the TUN writer instead, so ping6 and connections to our own
<npub>.fips address are delivered locally.
On Linux the kernel already loops self-traffic via `lo` before it reaches
the TUN, so the branch never fires there; the check is kept unconditional
as a daemon-level delivery invariant and to keep it covered by Linux CI.
poll_lan_discovery's comment said Noise XX, but LAN-discovered peers dial over
UDP through initiate_connection, which uses Noise IK (IK at FMP). compute_mesh_size's
header comment still described the obsolete sum-of-disjoint-subtrees estimate;
the function OR-unions every connected peer's inbound filter plus self and
estimates cardinality once (matching the body comment). Comment-only, no
behavior change.
The connect_refused stat counter (the Refused line in fipstop) was
defined but never incremented: every SOCKS5 connect failure recorded
socks5_errors instead, so the counter sat at zero and the operator-facing
gauge was permanently misleading. Both the synchronous connect path and
the background connect_async task now count a genuine SOCKS5 REP=0x05
refusal as connect_refused and every other failure as a socks5_error,
distinguishing them precisely via tokio_socks::Error::ConnectionRefused.
Extends the mock SOCKS5 server with a configurable reply code and adds
two tests covering the refused and general-failure paths.
Add an outbound-only Nym mixnet transport that tunnels FMP peer links
through a local nym-socks5-client SOCKS5 proxy into the Nym mixnet. It
structurally mirrors the Tor SOCKS5 transport (connection pool,
connect-on-send background promotion, FMP-v0 framing reused from TCP)
with the onion, inbound-listener, and control-port machinery removed.
Wires the transport through the full TransportHandle dispatch, NymConfig
(standard transport-instance pattern), and node instantiation, and
surfaces its counters in fipstop. Includes a mock SOCKS5 harness and unit
coverage for the address-parsing paths.
Also adds an isolated single-container example
(examples/sidecar-nostr-mixnet-relay/) demonstrating FIPS peering across
the mixnet end to end. No new crate dependencies: tokio_socks, socket2,
and futures are already pulled in by the Tor transport.
Two independent rendering glitches, both most visible over SSH and inside
tmux:
Startup: ratatui::try_init() enters the alternate screen but never clears
it, and the first terminal.draw() only emits cells that differ from an
assumed-blank internal buffer. On terminals that don't hand back a cleared
alternate buffer (notably tmux, and amplified by SSH latency) the prior
contents show through. Force a full repaint with terminal.clear() before
the first draw.
Quit: the input EventHandler spawned a detached thread that polled stdin in
a loop outliving the main loop, so at quit it kept reading after raw mode
was disabled and stray bytes (a keystroke or a terminal query response)
echoed onto the restored screen. Give the thread a stop flag and join it
before restoring the terminal; poll on a short fixed interval (decoupled
from the refresh tick) so quit stays responsive.
git log --oneline -1
When an FMP msg1 or FSP msg3 rekey retransmission budget is exhausted, the
cycle is abandoned and retried on the next timer. On lossy or high-latency
links this is an expected, self-limiting outcome: the existing session stays
valid and keeps carrying traffic, so the abort is not a failure that warrants
operator-level visibility. Demote both abandon-cycle messages from warn to
debug to cut steady-state log noise on nodes with many flapping peers.
A manual disconnect tore down only the local side and sent the peer nothing, so
the peer kept its session and never re-emitted its tree and filter
announcements; on reconnect it was never re-adopted as a child and its bloom
filter was never recorded. Send the disconnected peer a scoped Disconnect, the
same message graceful shutdown sends to all peers, so both sides tear down and
re-handshake cleanly on the next connection.
Reworks the fipstop TUI across its rendering, the control read surface it
draws from, and its interaction model, on a machine-verified base.
Test infrastructure:
- Add a ratatui TestBackend snapshot harness (testkit + snapshots
modules) that renders any ui::draw_* into an in-memory Buffer from
canned show_* JSON and asserts the text grid plus per-cell style.
Layout, columns, alignment, labels, grouping, and colour are now
checkable under cargo test; every render below ships a snapshot.
Control read surface (each new field emitted byte-identically on the
live and off-loop builders, published once from the tick, with schema
fixtures regenerated and the parity asserts holding):
- show_status: effective persistence (persistent || nsec.is_some());
root and is_root; and a per-configured-transport-type peer-count map
in which idle-but-configured types stay visible at zero.
- show_peers: per-peer effective_depth (depth + link_cost, the value
evaluate_parent ranks on), null when unmeasured or coordless so
fipstop never recomputes it.
- show_tree: root_npub, resolved once daemon-side (self when root, an
attested peer npub, or an identity-cache hit).
- show_bloom: the last-actually-sent uptree filter fill ratio and
subtree estimate, null for a root or before the first announce.
- show_mmp: session-layer srtt, loss, and etx trend labels.
Rendering:
- Display a 6-byte non-UTF-8 TransportAddr as a colon-separated MAC at
the type layer, so daemon logs, fipsctl, and JSON consumers all
benefit; non-6-byte payloads stay bare hex.
- Right-justify the Bloom Peer Filters numerics into aligned fixed-width
columns, render the Routing panes through a kv_lines helper that shares
one value column across a key-value group, and right-justify the Graphs
by-peer summary columns.
- Truncate an over-long peer name (the npub shown when no friendly name
exists) in the Tree, Bloom, and MMP peer lists so it no longer runs
into the next column.
- Group the Peers table by role (parent, then STP children, then other)
and render it as a full grouped view with styled group labels and
blank separators; the selection stays a peer index and the cursor only
ever lands on a peer row. Apply the same role grouping to the Tree and
Bloom peer lists, joining each peer's role from the peers view by node
address.
- Show min in the Graphs plot titles, rest a steady non-zero metric on
the baseline as a row of dots, render a genuine zero as an empty plot,
and keep a distinct no-data placeholder.
- Replace the metric-by-peer grid, which squeezed plots to nothing once
peers overflowed, with a master/detail Graphs view: a scrollable
per-peer summary list that expands (Enter) to a full-pane btop plot,
with up/down to flip peer, n/N to switch statistic, m to cycle mode,
and Esc to return.
- Put inline colored trend arrows on the Link and Session MMP values
(drawn only on a rising or falling trend, with a fixed blank slot when
stable so the value columns stay aligned), via a shared helper.
- Cycle column sorting on the Link MMP, Session MMP, and Graphs by-peer
tables (one key cycles the active column, another toggles direction),
with the active column marked in each table's header.
- Render the new daemon-surfaced fields: the dashboard root line (a
self-is-root marker, otherwise a truncated root hex), a
transports-by-type line, and an "approx. mesh estimate" line; an
effective_depth column and lines on the Peers, peer-detail, and Tree
sites from the single daemon derivation, showing a dash placeholder
when unmeasured rather than a misleading zero; the full Tree root hex
plus an Npub line; and the Bloom uptree fill and subtree-estimate lines.
Interaction model:
- Add a declarative keybinding registry keyed by (Tab, UiMode) that both
the context footer and the ? help overlay render from, so the two
cannot drift; a test asserts every registry key has a dispatch handler.
- Add a modal ? help overlay, and a context-aware footer that shows the
current state's actions first, drops global hints when the terminal is
narrow, and always keeps a Help affordance as the overflow path.
- Generalize per-pane focus and scroll state on App, wired across the
Tree, Filters, Routing, and MMP tabs (f cycles pane focus and the
focused pane scrolls instead of clipping its overflow); on the MMP tab
the column sort acts on the focused pane. Esc deselects the active row
when no detail is open (detail-close still takes priority).
- Add a Del-disconnect confirmation modal naming the peer, the only
state-mutating action, issuing the control-socket disconnect on confirm
and noting that the peer stays disconnected until manually reconnected.
Complete the control-plane read-isolation work: every pure-read show_*
query now renders in the control accept task from published read
snapshots, so none round-trips the data-plane receive loop. Only the
mutating connect/disconnect commands still reach that loop.
Three subsystem snapshots are published via ArcSwap and served through the
read handle's snapshot_dispatch:
- A routing read view (spanning tree, bloom filters, coordinate cache,
identity cache, and the discovery F-queue summary scalars), published
from the tick, serving show_tree/show_bloom/show_cache/show_routing/
show_identity_cache.
- A per-entity read view (peers, sessions, links, connections, transports,
and the MMP link/session views) as Vec<Arc<Row>> tables reconciled
against the prior snapshot so a republish reuses unchanged rows by
pointer and re-allocates only changed or new rows, keeping the per-tick
publish cost bounded as the peer/session count grows. Serves
show_peers/show_sessions/show_links/show_connections/show_transports/
show_mmp.
- The stats snapshot is extended with the peer-ACL status and a per-peer
metadata map (is_active, npub, display name), resolved at publish time,
serving show_acl and the two per-peer stats queries.
Display names and other cross-subsystem fields are resolved at publish
time; time-relative fields are derived at render time from captured
absolute timestamps, so rendered output is byte-identical to the prior
on-loop handlers, which are retained as the equality oracle.
With every read query served off-loop, the show_* branch is removed from
the rx_loop control handler and the now-dead on-loop dispatcher deleted.
The snapshot projections are forward-compatible with the later structural
extraction of the derived-state and session tables: they become thin
views over the extracted types without changing the read-handle interface.
Introduce a read-snapshot plane so pure-snapshot control queries render in
the control-socket task instead of round-tripping the rx_loop, removing the
head-of-line coupling that let a busy or slow rx_loop time out fipsctl and
fipstop observability.
- ControlReadHandle: a cloneable bundle the control accept loop holds, over
the node's already-shared NodeContext and MetricsRegistry plus an
ArcSwap-published StatsSnapshot. A snapshot_dispatch seam serves cut-over
commands off-loop and falls through to the rx_loop for the rest, keeping
the rx_loop's ownership of Node intact.
- StatsSnapshot is published from the tick (the natural and sole mutator of
stats_history), carrying the history rings plus the scalar gauges and
counts show_status reports. Readers serve the latest snapshot
unconditionally, with staleness bounded by the tick interval and no
IO_TIMEOUT-coupled fallback.
- Off-loop now: show_status, show_stats_history, show_stats_all_history,
show_listening_sockets, show_stats_list, and a new counter-only
show_metrics (exposed as fipsctl "stats metrics", the enabler for a
Prometheus scraper at no hot-path cost). Queries that need live per-entity
state (peers, links, sessions, routing, and the per-peer stats variants)
stay on the rx_loop path pending later phases.
Quartet green; forward-merge to next verified clean.
The pool's TTL clock (VirtualIpMapping.last_referenced) advanced only on
DNS re-query, never on traffic, and the mapping-TTL is wired equal to the
DNS TTL, so an in-use mapping was forced to drain at TTL and reclaimed at
the first zero-conntrack tick (a stale drain_start gave no grace effective
protection), breaking long-lived, bursty, or DNS-cached clients.
In tick(), refresh last_referenced whenever conntrack reports sessions > 0
so an actively used mapping never ages out, and recover a Draining mapping
to Active (clearing drain_start) when traffic resumes, so a later drain
gets a fresh grace window instead of a stale one. The Active arm now only
drains an idle mapping. The DNS-TTL / idle-reclaim-TTL wiring is unchanged.
Adds regression tests for continuous-traffic-survives-past-TTL, bursty
drain-then-recover, and fresh-grace-on-redrain.
The per-transport TCP inbound cap was hardwired to 256 and never read
node.limits.max_connections, so raising max_connections was a silent
no-op for inbound TCP. Resolve the effective cap with precedence:
explicit per-transport max_inbound_connections, then node-wide
max_connections, then the built-in default of 256. Established peers
remain bounded node-wide by add_connection, so deriving the per-transport
raw-accept ceiling from max_connections does not admit more real peers
across multiple transports.
Add effective_max_inbound on the TCP transport with a node_max_connections
setter wired from create_transports, plus a precedence unit test.
Estimate the OR-union cardinality over self plus every connected peer's
inbound filter, dropping the parent/child tree gating in
compute_mesh_size. Filter propagation is split-horizon, so cross-links
advertise near-complete mesh views; unioning all peers yields the same
set as the tree-only union in steady state (OR dedups overlap and no
filter can over-count) while damping the node-count flap on parent
switches, since dropping the parent no longer collapses the upward leg.
This also removes the estimate's dependence on tree-declaration cache
freshness.
Rename the debug-log child_count to contributor_count, adapt the two
membership-invariant tests to all-peers semantics, and add a test that
the estimate stays stable across a parent drop when a healthy cross-link
is present.
The inbound FilterAnnounce FPR cap rejects filters whose false-positive
rate (fill^k) exceeds the configured maximum. On the fixed 1 KB / k=5
filter, 0.05 corresponds to fill 0.549 (~1,300 reachable entries), and
the busiest nodes' aggregates were beginning to hit that ceiling as the
mesh grew. Raise the default to 0.10 (fill 0.631, ~1,630 entries) to
restore headroom toward the fixed-filter capacity limit. A saturated or
poisoned filter is ~100% FPR and remains rejected, so the antipoison
gate is not materially weakened.
Updates the config default, the config-reference and bloom-filter
design docs, and the changelog.
A node with a single tree peer has its periodic parent re-evaluation
disabled (it needs at least two peers for a meaningful comparison), so
it depends entirely on its peer pushing a TreeAnnounce for it to attach.
That push happens once at promotion time plus on the parent's slow
periodic no-change re-broadcast. If the one-shot attaching announce is
lost, the single-uplink node falls back to self-root and cannot recover
until the next periodic re-broadcast (reeval_interval_secs later),
stranding it out of the tree and unreachable end-to-end in the interim.
Make tree-position exchange self-healing on the receive path: when an
accepted TreeAnnounce advertises a root strictly worse (higher NodeAddr,
since election is smallest-wins) than our own, echo our current
declaration back to that peer. A stranded self-root node's announce now
provokes its better-rooted peer to re-push its real position immediately,
so the node re-attaches within a round-trip instead of waiting for the
periodic cadence.
Echo only in that one direction. If the peer's root is lower (better)
than ours, we are the stale side: the peer would ignore our worse root
anyway and we converge via the parent re-evaluation that follows, so
echoing back is pure waste and would double announce traffic in the
learning direction during a root change or partition merge. Equal roots
are already converged. The echo is bounded by the existing per-peer
500 ms tree-announce rate limiter and is a no-op once the peer adopts our
root, so it adds no traffic in a converged mesh.
Add a spanning-tree unit test that drives a converged child back to
self-root with its peer-ancestry view of the root cleared (modelling the
lost attaching announce), and asserts the root re-pushes on the
resulting root disagreement and the child re-attaches.
Ignore duplicate or counter-regressed ReceiverReports before updating
RTT, loss, goodput, or ETX, so a delayed or reordered report can no
longer poison link metrics. Compute the RTT-from-echo sample with
checked timestamp arithmetic and reject zero, negative, or out-of-range
results instead of risking wrap or underflow on untrusted wire values.
On the sender side, when receiver dwell time overflows the u16 wire
field, suppress the timestamp echo (send 0) and saturate dwell to
u16::MAX rather than truncating, so a bogus small RTT cannot be formed.
Adds duplicate, out-of-order, wrapped-add, and future-dated (checked_sub)
sample tests, asserts loss and goodput stay unchanged on a dropped
duplicate, and covers the dwell-overflow echo suppression. Documents the
behavior in the MMP design note and CHANGELOG.
Co-authored-by: Johnathan Corgan <johnathan@corganlabs.com>
Brings the maint-line OR-union mesh-size estimator fix and the
per-peer / capacity-cap log-level demotions forward to master.
Conflict resolution: compute_mesh_size resolved to master's config()
accessor form carrying maint's OR-union rewrite (drop the stale summing
`total`); the log-level demotions auto-merged into master's handler
versions. The equivalent master-only log "connected UDP socket installed"
(connected_udp.rs, a file that does not exist on maint) was demoted
info -> debug as part of this merge so the connected-UDP path matches the
rest of the per-peer lifecycle logging.
On a saturated public-mesh node the connection-lifecycle and capacity-cap
events fire continuously and drown out the genuinely notable INFO/WARN
lines. Demote them to debug and drop a redundant duplicate:
- FMP K-bit cutover promotion (encrypted): info -> debug
- "Connection promoted to active peer" (handshake): info -> debug, and
remove the duplicate "Inbound peer promoted to active" line that
shadowed it on the inbound path
- "Peer restart detected" (handshake): info -> debug
- "Peer removed and state cleaned up" (dispatch): info -> debug
- "Rejecting inbound TCP connection (max_inbound_connections reached)"
(tcp): warn -> debug
- "Congestion detected, CE flag set on forwarded packet" (forwarding):
warn -> debug
- "Removing peer: link dead timeout" (mmp): warn -> debug
These are expected, high-frequency conditions on a busy public node (new
and reconnecting peers, ECN CE marking, the inbound connection cap, and
link-dead churn), not operator-actionable signals.
The mesh-size estimator summed the per-filter cardinality of the parent
filter and each child filter, which assumes those filters are perfectly
disjoint. When they overlap -- a stale or oversized parent filter, or a
routing loop -- the sum over-counts and inflates the reported mesh size
to as much as several times the true size.
Estimate the cardinality of the OR-union of the contributing filters
(self + parent + children) once instead. OR is idempotent, so any
overlap is deduplicated: the result equals the old sum in the disjoint
case and stays correct under overlap. The union is seeded from a clone
of a contributing filter so it keeps that filter's size class, and a
filter whose size class does not match is skipped rather than panicking.
The refuse-to-estimate behavior on a saturated or above-cap filter is
preserved.
Add a regression test with overlapping parent and child filters where
the naive sum over-counts and the union estimate tracks the distinct
member count.
The transport layer used Mutex::lock().unwrap() at ten sites across the
UDP, BLE, and Ethernet code. A std mutex poisons if a thread panics
while holding it, after which every lock().unwrap() on that same mutex
also panics, turning one fault into a cascade. These critical sections
only perform short HashMap/Vec operations on locally constructed values
and are not reachable from peer input, but the idiom is fragile against
any future in-section panic. Replace each with
lock().unwrap_or_else(|e| e.into_inner()), which recovers the guarded
data and removes the cascade with no new dependency and no call-graph
change.
Also replace four self.local_addr.unwrap() calls in the UDP start and
adopt paths with a sentinel fallback. The value is provably set just
above each log line today, but the unwrap is brittle against a future
reordering; logging an unbound sentinel is harmless and cannot panic.