Files
fips/src/peer/mod.rs
T
Arjen 8c78e1f4b9 feat(peer): a peer with a session is never dialled; a handshake creates no path state
Three ways an address for a peer we already hold a session with used
to reach the dialler — a beacon on a new transport, `update_peers` or
`fipsctl connect`, a configured address whose transport came up later
— and each was a second handshake, which the far side read as a rekey
and this side resolved as a cross-connection, the two not composing.
Two phones hearing the same Wi-Fi return at the same moment both
dialled at once; one side swapped to its outbound session and freed
the index it had just handed out in the rekey reply, the other kept
its inbound session and the pre-rekey index, every frame between them
was dropped, and the link-dead reap tore the peer down. About a
minute dark on every Wi-Fi return.

A peer that holds a session is never dialled now. An address on a
transport it has no path over becomes a path candidate under that
session; one on a transport whose path is not eligible re-points that
path (the active path included: it is not answering, that is why we
are here); one on a transport whose path is carrying acknowledged
traffic changes nothing. The heartbeat tick probes the candidate under
the existing session — one authenticated, replay-checked round trip —
and the mandatory switch takes it if the current path stops answering.
Nothing is lost against the dial: a session that is truly gone answers
no probe either, is reaped by the link-dead timeout, and is dialled
then; a peer that restarted dials us with a new epoch and wins
promotion outright, as before. Applies to the control API's connect,
to update_peers, to configured addresses (checked once a tick) and to
transport discovery alike.

The counterpart: a handshake creates no path state. A dial that does
reach a peer with a session — a startup that lists two addresses
dials both, a caller that still dials by hand, an older node dialling
us — is classified and resolved exactly as before this work: rekey,
duplicate or restart on the responder, whichever transport the msg1
arrived on; the cross-connection tie-break on the initiator. The
address it ran to is left as a candidate for the probe exchange. Two
reasons. Both ends must resolve a handshake on the same information,
and "is this a new transport to a live peer" was a fact only one end
could see. And the IK responder commits at msg1, which carries no
freshness beyond the startup epoch: a captured msg1 replayed from any
address would otherwise have planted a path, probed full-size for the
life of the peering and counting as a transport the peer is on for the
decrypt-failure gate. So that gate now counts garbage only on the
active path or one the peer has acknowledged. On a connection-oriented
transport the connection a dial opened is kept as the candidate's
socket rather than closed as the losing leg: the probe rides it, and
closing it would only have the first probe dial again — or, at the
responder, find an ephemeral port that cannot be dialled at all.
`api_disconnect` closes every path's connection, the standby's
included; loopback records the closes it is asked for so a test can
say so.

Path heartbeats are gated and bounded. A peer with one live path is
not path-heartbeated: selection has nothing to move to, the link
heartbeat keeps liveness, and five probes a second on every single-path
link was cost without a decision behind it. A standby the peer never
acknowledges is given up after eight discovery probes, Dead and pruned
after the grace; the active path is never given up. The active path's
first probe is small, the handshake having proved it and seeded its
MTU. And a Dead path is probed again when its transport returns:
nothing on our side ever re-probed one, so after a NIC replug traffic
stayed on the standby until the grace pruned the path and a beacon
found it with no history. The presence edge now revives every Dead
path on the transport as Probing, RTT window and ETX kept.

Smaller: `add_path_candidate` re-points a known transport's path at a
moved address (`refresh_path_addr`), for a Wi-Fi Aware data path that
re-forms with a new link-local; `api_disconnect` closes every path's
connection, not the active one alone; `path_show` is built from the
`show_peers` path projection plus the three now-relative fields;
`PathState` and `TransportRole` render through `as_str()`;
`node.path.switch_margin` is validated finite and at least 1.0;
`PathPolicy::PERMISSIVE` had no users; four doc comments an inserted
function had split are put back on the function they describe.

The dual-udp-flap scenario is config-driven: the dial owner lists
udp/main and udp/<veth>, both dial at startup, and the second is
proven as a path under the first's session by the probe exchange.
2026-09-23 11:48:33 -03:00

115 lines
3.3 KiB
Rust

//! Peer Management
//!
//! Two-phase peer lifecycle:
//! 1. **PeerMachine** with a handshake carrier attached - handshake phase,
//! before identity is verified
//! 2. **ActivePeer** - Authenticated phase, after successful Noise handshake
mod active;
pub(crate) mod machine;
pub use active::{
ActivePeer, ConnectivityState, HeartbeatPlan, HeartbeatSend, HeartbeatTiming,
MAX_DISCOVERY_PROBES, PathPolicy, PathState, PathSwitch, PathWithdrawal, PeerPath,
SwitchReason,
};
use crate::NodeAddr;
use crate::transport::LinkId;
use thiserror::Error;
// ============================================================================
// Errors
// ============================================================================
/// Errors related to peer operations.
#[derive(Debug, Error)]
pub enum PeerError {
#[error("peer not authenticated")]
NotAuthenticated,
#[error("peer not found: {0:?}")]
NotFound(NodeAddr),
#[error("connection not found: {0}")]
ConnectionNotFound(LinkId),
#[error("peer already exists: {0:?}")]
AlreadyExists(NodeAddr),
#[error("handshake failed: {0}")]
HandshakeFailed(String),
#[error("handshake timeout")]
HandshakeTimeout,
#[error("identity mismatch: expected {expected:?}, got {actual:?}")]
IdentityMismatch {
expected: NodeAddr,
actual: NodeAddr,
},
#[error("peer disconnected")]
Disconnected,
#[error("max connections exceeded: {max}")]
MaxConnectionsExceeded { max: usize },
#[error("max peers exceeded: {max}")]
MaxPeersExceeded { max: usize },
}
// ============================================================================
// Tests
// ============================================================================
#[cfg(test)]
mod tests {
use crate::proto::fmp::PromotionResult;
use crate::transport::LinkId;
use crate::{Identity, PeerIdentity};
fn make_peer_identity() -> PeerIdentity {
let identity = Identity::generate();
PeerIdentity::from_pubkey(identity.pubkey())
}
#[test]
fn test_promotion_result_promoted() {
let identity = make_peer_identity();
let node_addr = *identity.node_addr();
let result = PromotionResult::Promoted(node_addr);
assert!(result.node_addr().is_some());
assert_eq!(result.node_addr(), Some(node_addr));
assert!(!result.should_close_this_connection());
assert!(result.link_to_close().is_none());
}
#[test]
fn test_promotion_result_cross_lost() {
let result = PromotionResult::CrossConnectionLost {
winner_link_id: LinkId::new(1),
};
assert!(result.node_addr().is_none());
assert!(result.should_close_this_connection());
assert!(result.link_to_close().is_none()); // Caller closes their own
}
#[test]
fn test_promotion_result_cross_won() {
let identity = make_peer_identity();
let node_addr = *identity.node_addr();
let result = PromotionResult::CrossConnectionWon {
loser_link_id: LinkId::new(1),
node_addr,
};
assert!(result.node_addr().is_some());
assert_eq!(result.node_addr(), Some(node_addr));
assert!(!result.should_close_this_connection());
assert_eq!(result.link_to_close(), Some(LinkId::new(1)));
}
}