Files
fips/src/node/reloadable.rs
T
Johnathan Corgan d246f84da5 Rebuild alias name resolution and ACL alias entries when the peer list is replaced at runtime
Node::update_peers replaced the configured peer list but left every map
that reads peer aliases as it was at startup. Each hosts map (the display
map, the peer ACL's alias resolution and the .fips DNS responder's map)
kept the startup aliases as a fixed base and only re-merged the hosts file
over it. After an update, a new peer's alias did not resolve as a .fips
name or match an ACL entry naming it, a removed or moved alias kept
resolving to the old npub, and a deny entry written as an alias that moved
to another key kept denying the old key while admitting the new one. A
kept peer whose alias was removed also kept showing the old alias as its
display name.

update_peers now rebuilds the alias base from the new peer list and hands
it to all three maps: the display map directly, the peer ACL through a
forced rebuild that runs before any added peer is dialed, and the running
DNS responder through a watch channel it checks before each answer. Each
map keeps the hosts file as last read, so the file stays merged on top and
still wins, including edits picked up at runtime. The kept peer's display
name falls back to its short npub when its alias is removed.

run_dns_responder keeps its signature. This reaches only embedders that
call Node::update_peers.
2026-09-26 18:59:43 +00:00

401 lines
16 KiB
Rust

//! Lock-free reloadable configuration / resource snapshots.
//!
//! Several node-owned resources are loaded from disk at startup and may be
//! re-read when the backing file changes (for example the `/etc/fips/hosts`
//! map). Historically each one carried its own ad-hoc reloader with a
//! slightly different shape. The [`Reloadable`] trait normalizes them onto a
//! single contract built around an [`arc_swap::ArcSwap`] snapshot.
//!
//! # Canonical Arc-wrapper template
//!
//! These resources follow a single-writer / many-reader pattern: every write
//! goes through `&mut Node`, from the node tick task or from
//! `Node::update_peers`, so there is one writer at a time, while the hot path
//! reads the current value frequently and must never block.
//!
//! - The reader-facing immutable snapshot lives in an
//! [`arc_swap::ArcSwap<T>`]. Readers call [`Reloadable::load`], which yields
//! a lock-free [`arc_swap::Guard<Arc<T>>`] that derefs straight to the
//! snapshot — no mutex, no clone on the read path.
//! - The owning struct also holds the change-detection state (file mtime,
//! base data, source path). That state is touched only by
//! [`Reloadable::reload`] and, for the host map's peer-alias base, by
//! `HostMapReloadable::set_base`, both reached only through `&mut Node`.
//! - `reload` builds a brand-new `T` and then stores `Arc::new(new)` into the
//! `ArcSwap`, so a reader either sees the entire old snapshot or the entire
//! new one — never a partial update.
//! - Construction performs the initial synchronous load so the snapshot is
//! valid before the node starts serving reads.
//!
//! `reload` returns `bool` rather than a value or `Result`: the underlying
//! loaders already absorb I/O errors internally (a missing or unreadable
//! source degrades to an empty/base snapshot plus a warning), so callers have
//! nothing to handle. The return flag reports only whether the snapshot was
//! replaced, which is all the tick loop needs for logging.
//!
//! # What is and isn't `Reloadable`
//!
//! Two node-owned resources are deliberately *not* `Reloadable` because
//! neither has a "re-read from a backing source" concept:
//!
//! - `path_mtu_lookup` is an event-driven cache (`Arc<RwLock<HashMap>>`)
//! populated from observed path-MTU discovery traffic, not loaded from a
//! file. There is nothing to poll. Release is mostly event-driven for the
//! same reason: an entry is dropped when the path it describes is declared
//! invalid (a `PathBroken` report, session idle expiry, or handshake
//! timeout) and the locally derived link MTU is reseeded in its place. All
//! three of those events read session state, which leaves one carrier
//! uncovered: a lookup `LookupResponse` writes an entry for a destination
//! this node may never open a session with. Those entries, and only those,
//! carry a learn time and are expired at the coordinate cache's TTL by
//! `purge_expired_path_mtu` on the same tick, which then reseeds any direct
//! peer whose entry went. Locally derived link MTUs and values learned
//! inside a session carry no deadline. That sweep is expiry, not a reload.
//! (The read side could adopt the same lock-free `ArcSwap` shape in the
//! future, but that is an optimization, not a reload.)
//! - `nostr_rendezvous` is an async spawned subsystem, not a snapshot of disk
//! state.
//!
//! Both [`HostMapReloadable`] and the peer ACL reloader currently stat
//! `/etc/fips/hosts` independently each tick (the ACL reloader embeds its own
//! hosts reloader for alias resolution). A single small-file `stat` per tick
//! is cheap, so the duplicate is left in place; collapsing the two onto a
//! single shared mtime observation is a possible future cleanup.
use std::sync::Arc;
use crate::upper::hosts::{HostMap, file_mtime};
/// A resource backed by a lock-free [`arc_swap::ArcSwap`] snapshot that can be
/// re-read from its source on demand.
///
/// See the [module documentation](self) for the canonical Arc-wrapper
/// template that implementors follow.
pub trait Reloadable: Send {
/// The immutable snapshot type readers observe.
type Snapshot;
/// Re-read the backing source and replace the snapshot if it changed.
///
/// Returns `true` if a new snapshot was stored, `false` if nothing
/// changed. I/O errors are absorbed internally (degrading to an
/// empty/base snapshot with a warning) rather than surfaced.
///
/// Driven once per node tick from the rx loop, alongside the other
/// reloadable resources.
async fn reload(&mut self) -> bool;
/// Acquire a lock-free guard over the current snapshot.
///
/// This is the hot-path read: it performs no locking and no allocation.
fn load(&self) -> arc_swap::Guard<Arc<Self::Snapshot>>;
}
/// Reloadable hostname → npub map (base peer aliases merged with the operator
/// hosts file).
///
/// Holds the base map (from peer-config aliases, replaced when the peer list
/// is replaced at runtime) plus the change-detection state for the hosts
/// file. The effective map (base merged
/// with the hosts file) is published through an [`arc_swap::ArcSwap`] so the
/// display path can read it without locking.
pub struct HostMapReloadable {
/// Reader-facing effective snapshot (base merged with hosts file).
snapshot: arc_swap::ArcSwap<HostMap>,
/// Base map from peer-config aliases. Read by `reload` and replaced by
/// `set_base`, both reached only through `&mut Node`.
base: HostMap,
/// The hosts file as last read, kept so a new base can be merged under
/// it without reading the file again. Written by `new` and `reload`.
file: HostMap,
/// Path to the operator hosts file. Read only by `reload`.
path: std::path::PathBuf,
/// Last observed modification time of the hosts file (`None` if absent).
/// Read only by `reload`.
last_mtime: Option<std::time::SystemTime>,
}
impl HostMapReloadable {
/// Create a reloadable host map.
///
/// Performs the initial load of the hosts file and merges it over the
/// base map so the published snapshot is valid immediately.
pub fn new(base: HostMap, path: std::path::PathBuf) -> Self {
let last_mtime = file_mtime(&path);
let hosts_file = HostMap::load_hosts_file(&path);
let mut effective = base.clone();
effective.merge(hosts_file.clone());
Self {
snapshot: arc_swap::ArcSwap::from(Arc::new(effective)),
base,
file: hosts_file,
path,
last_mtime,
}
}
/// Replace the peer-alias base and publish it merged with the hosts file
/// as last read, which still wins on conflicts.
///
/// Reads no file and leaves the recorded mtime alone. Like `reload`, it is
/// reached only through `&mut Node`, which keeps one writer at a time;
/// readers see either the whole old snapshot or the whole new one.
pub(crate) fn set_base(&mut self, base: HostMap) {
let mut effective = base.clone();
effective.merge(self.file.clone());
self.base = base;
self.snapshot.store(Arc::new(effective));
}
}
impl Reloadable for HostMapReloadable {
type Snapshot = HostMap;
async fn reload(&mut self) -> bool {
let current_mtime = file_mtime(&self.path);
if current_mtime == self.last_mtime {
return false;
}
// File appeared, disappeared, or was modified.
self.last_mtime = current_mtime;
let hosts_file = HostMap::load_hosts_file(&self.path);
let mut new_effective = self.base.clone();
new_effective.merge(hosts_file.clone());
self.file = hosts_file;
let count = new_effective.len();
self.snapshot.store(Arc::new(new_effective));
tracing::info!(
path = %self.path.display(),
entries = count,
"Reloaded hosts file"
);
true
}
fn load(&self) -> arc_swap::Guard<Arc<HostMap>> {
self.snapshot.load()
}
}
#[cfg(test)]
mod tests {
use super::*;
use crate::Identity;
#[tokio::test]
async fn test_initial_load_base_and_file() {
let id_base = Identity::generate();
let id_file = Identity::generate();
let mut base = HostMap::new();
base.insert("core", &id_base.npub()).unwrap();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("hosts");
std::fs::write(&path, format!("gateway {}\n", id_file.npub())).unwrap();
let reloadable = HostMapReloadable::new(base, path);
let snapshot = reloadable.load();
assert_eq!(snapshot.len(), 2);
assert!(snapshot.lookup_npub("core").is_some());
assert!(snapshot.lookup_npub("gateway").is_some());
}
#[tokio::test]
async fn test_initial_load_no_file_base_only() {
let id = Identity::generate();
let mut base = HostMap::new();
base.insert("core", &id.npub()).unwrap();
let reloadable =
HostMapReloadable::new(base, std::path::PathBuf::from("/nonexistent/hosts"));
let snapshot = reloadable.load();
assert_eq!(snapshot.len(), 1);
assert!(snapshot.lookup_npub("core").is_some());
}
#[tokio::test]
async fn test_reload_detects_file_change() {
let id1 = Identity::generate();
let id2 = Identity::generate();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("hosts");
std::fs::write(&path, format!("gateway {}\n", id1.npub())).unwrap();
let mut reloadable = HostMapReloadable::new(HostMap::new(), path.clone());
assert_eq!(reloadable.load().len(), 1);
assert_eq!(
reloadable.load().lookup_npub("gateway"),
Some(id1.npub().as_str())
);
// No change yet.
assert!(!reloadable.reload().await);
// Bump mtime by rewriting; sleep for filesystem mtime granularity.
std::thread::sleep(std::time::Duration::from_millis(50));
std::fs::write(
&path,
format!("gateway {}\nnew-host {}\n", id1.npub(), id2.npub()),
)
.unwrap();
assert!(reloadable.reload().await);
let snapshot = reloadable.load();
assert_eq!(snapshot.len(), 2);
assert!(snapshot.lookup_npub("new-host").is_some());
}
#[tokio::test]
async fn test_reload_no_change_returns_false() {
let id = Identity::generate();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("hosts");
std::fs::write(&path, format!("gateway {}\n", id.npub())).unwrap();
let mut reloadable = HostMapReloadable::new(HostMap::new(), path);
assert!(!reloadable.reload().await);
assert!(!reloadable.reload().await);
}
#[tokio::test]
async fn test_reload_detects_file_deletion() {
let id = Identity::generate();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("hosts");
std::fs::write(&path, format!("gateway {}\n", id.npub())).unwrap();
let mut reloadable = HostMapReloadable::new(HostMap::new(), path.clone());
assert_eq!(reloadable.load().len(), 1);
std::fs::remove_file(&path).unwrap();
assert!(reloadable.reload().await);
assert!(reloadable.load().is_empty());
}
#[tokio::test]
async fn test_reload_detects_file_creation() {
let id = Identity::generate();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("hosts");
let mut reloadable = HostMapReloadable::new(HostMap::new(), path.clone());
assert!(reloadable.load().is_empty());
std::fs::write(&path, format!("gateway {}\n", id.npub())).unwrap();
assert!(reloadable.reload().await);
let snapshot = reloadable.load();
assert_eq!(snapshot.len(), 1);
assert!(snapshot.lookup_npub("gateway").is_some());
}
#[tokio::test]
async fn test_reload_preserves_base() {
let id_base = Identity::generate();
let id_file = Identity::generate();
let mut base = HostMap::new();
base.insert("core", &id_base.npub()).unwrap();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("hosts");
std::fs::write(&path, format!("gateway {}\n", id_file.npub())).unwrap();
let mut reloadable = HostMapReloadable::new(base, path.clone());
assert_eq!(reloadable.load().len(), 2);
std::fs::remove_file(&path).unwrap();
assert!(reloadable.reload().await);
let snapshot = reloadable.load();
assert_eq!(snapshot.len(), 1);
assert!(snapshot.lookup_npub("core").is_some());
assert!(snapshot.lookup_npub("gateway").is_none());
}
/// The initial published snapshot must be byte-for-byte equivalent to the
/// pre-migration `Arc<HostMap>` built by `base.clone()` + `merge(file)`.
#[tokio::test]
async fn test_initial_snapshot_matches_pre_migration_construction() {
let id_base = Identity::generate();
let id_file = Identity::generate();
let mut base = HostMap::new();
base.insert("core", &id_base.npub()).unwrap();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("hosts");
std::fs::write(&path, format!("gateway {}\n", id_file.npub())).unwrap();
// Pre-migration construction.
let mut expected = base.clone();
expected.merge(HostMap::load_hosts_file(&path));
// Post-migration construction.
let reloadable = HostMapReloadable::new(base, path);
let snapshot = reloadable.load();
assert_eq!(snapshot.len(), expected.len());
for key in ["core", "gateway"] {
assert_eq!(snapshot.lookup_npub(key), expected.lookup_npub(key));
}
}
/// Build a one-entry host map.
fn one_entry(name: &str, id: &Identity) -> HostMap {
let mut map = HostMap::new();
map.insert(name, &id.npub()).unwrap();
map
}
/// Replacing the base swaps the peer aliases while the hosts file, as it
/// was last re-read at runtime, stays merged on top and still wins.
#[tokio::test]
async fn set_base_replaces_peer_aliases_and_keeps_the_last_reloaded_hosts_file_on_top() {
let [x, y, z, v, w] = std::array::from_fn(|_| Identity::generate());
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("hosts");
std::fs::write(&path, format!("f {}\na {}\n", y.npub(), z.npub())).unwrap();
let mut reloadable = HostMapReloadable::new(one_entry("a", &x), path.clone());
let npub = |r: &HostMapReloadable, name: &str| r.load().lookup_npub(name).map(String::from);
assert_eq!(npub(&reloadable, "f"), Some(y.npub()), "startup file entry");
std::thread::sleep(std::time::Duration::from_millis(50));
std::fs::write(&path, format!("g {}\na {}\n", v.npub(), z.npub())).unwrap();
assert!(reloadable.reload().await, "reload sees the rewrite");
reloadable.set_base(one_entry("b", &w));
assert_eq!(npub(&reloadable, "b"), Some(w.npub()), "new base alias");
assert_eq!(
npub(&reloadable, "g"),
Some(v.npub()),
"file entry added at runtime survives set_base"
);
assert_eq!(npub(&reloadable, "a"), Some(z.npub()), "file still wins");
assert_eq!(
npub(&reloadable, "f"),
None,
"file entry removed at runtime stays removed"
);
let x_addr = *crate::PeerIdentity::from_npub(&x.npub())
.unwrap()
.node_addr();
assert_eq!(
reloadable.load().lookup_hostname(&x_addr),
None,
"old base npub no longer reverse-resolves"
);
}
}