diff --git a/CHANGELOG.md b/CHANGELOG.md index 0c15c687..4e7b6ef6 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -542,6 +542,15 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 non-positive rate is rejected at config validation rather than silently refusing every session. +- `node.limits.max_sessions`, defaulting to 1024, which bounds the end-to-end + session table. Zero means unlimited, which restores the previous behaviour + exactly and is the way to back the change out on a running node. The default + is four times the adjacent `node.session.pending_max_destinations`. A + session entry measures 6608 bytes of inline state plus heap, so the table + holds to roughly 7 MB, and a test pins that per-entry figure so the + arithmetic behind the default fails loudly if an entry grows. Existing + configurations parse unchanged, the key being optional. + - `node.rate_limit.established_handshake_burst` and `node.rate_limit.established_handshake_rate`, the parameters of the new established-link msg1 token bucket, which meters link-layer msg1 rather @@ -952,6 +961,33 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 #### Transports & config +- The UDP transport's DNS cache is now bounded and actually evicts. The map + held one entry per distinct hostname string ever dialed, and the TTL was + applied only on the read, so a stale entry was overwritten on the next dial + of the same name and otherwise stayed for the life of the process. Under a + rendezvous policy that accepts advertised endpoints the keys are strings a + remote party chose, which made the growth theirs to drive. A store now + sweeps entries past their TTL and, if the map is still full, drops the + oldest, holding it to 256 hostnames. Refreshing a name already cached + evicts nothing. Eviction is by insertion time rather than last use, so a + rarely dialed name in a very large peer list may re-resolve more often; the + cost of a wrong eviction is one DNS lookup, not a failed dial. + +- macOS: stopping an Ethernet transport under load no longer hangs the + process. The BPF reader thread handed each frame to the async consumer with + `blocking_send`, which parks with no way to be woken. Stopping the transport + aborts the consumer first, so nothing drains the 1024-frame channel, and the + socket's `Drop` then joined a thread that could never return: on a busy + interface the daemon had to be killed. The socket now drops the receiver + before joining, which releases a parked send at once, and the reader thread + sends through a helper that watches the same shutdown pipe its `select()` + already honours, so a send waiting for room cannot outlive a shutdown + request. The helper yields before it sleeps, so the saturated-path handoff + rate is unchanged. **Not covered by CI**: the reader thread is macOS-only + and Linux CI compiles none of it. What the tests prove is that the helper + the thread now waits in is cancellable; that a real BPF thread exits under + load still needs a manual check on a Mac. + - A failed private-key write no longer leaves a node silently running an ephemeral identity. Six write results in the identity path were discarded, and the sharpest was in `persistent` mode: a failed write to `fips.key` fell @@ -1048,6 +1084,61 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 #### FMP/FSP session integrity +- A frame whose counter is `u64::MAX` is now refused by the replay window + instead of being accepted as a new high-water mark. Accepting it pinned + `highest` at the ceiling, after which every subsequent counter from that peer + fell more than a replay window below it and was rejected, wedging that peer's + own receive path until a rekey replaced the session. The send side already + refuses to emit that counter (`take_send_counter` and `advance_nonce` both + return a nonce-overflow error), so no conforming peer can produce it and the + refusal is invisible on the wire; the highest counter an honest peer can send, + `u64::MAX - 1`, is still accepted. Reaching this required an + already-authenticated peer running modified code, and the damage was confined + to that peer's own session. + +- The MMP gap tracker advances its expected-counter state with a saturating add, + so a received counter of `u64::MAX` no longer overflows it. The wrap silently + reset the expectation to zero in a release build and aborted the task under a + build with overflow checks on, such as the test harness. Behaviour is + unchanged for every counter an honest peer can emit. + +- An epoch-mismatch msg1 no longer tears down a peering that is still + carrying authenticated traffic, and a second epoch change for the same peer + identity inside 15 seconds is refused. The epoch travels inside the AEAD, so + such a msg1 is authentic, but it stays authentic after capture: replaying + one destroyed a working peering, and with it the FSP session state that + peering carried, from off the path. The peering's last authenticated inbound + frame is the evidence that it is still alive, and nothing an unauthenticated + sender emits can refresh it, so a peer that genuinely restarted clears the + gate by having stopped sending. The refusal is a silent drop: no msg2 is + returned, since the stored msg2 is bound to the original msg1's ephemeral + and answering a sender-chosen address is free amplification. The interval is + stamped only when an epoch change is accepted, so a sustained replay cannot + starve a genuinely restarting peer. Both thresholds come from one constant, + sized so a restarting peer's msg1 resends still land inside its own first + handshake window and below `link_dead_timeout_secs`, and nothing changes on + the wire. + +- Retention of a superseded FSP key epoch is now capped at an absolute + ceiling measured from the cutover, defaulting to 120 seconds against the + 10-second drain window. The drain deadline slides forward on every inbound + frame that authenticates against the `previous` slot, which is what keeps a + peer that lost msg3 from having the old epoch erased out from under it, but + it also meant the authenticated peer holding that key could keep the retired + key resident for as long as it kept sealing frames in the old epoch. The + sliding grace is unchanged; it now delays erasure by a bounded amount rather + than preventing it. The ceiling is set to clear the worst-case legitimate + recovery of a peer that lost msg3 (the msg3 resend ladder, then + `handshake_timeout_secs` before the responder abandons, then the rekey + dampening window before it may re-initiate, about 90 seconds at stock + settings), and it is raised automatically if the configured handshake timers + imply a longer budget, so shortening a timer cannot push the ceiling under + the recovery it has to leave room for. Nothing changes on the wire; each side + runs its own drain. A peer that has still not recovered when the ceiling + fires is left with undecryptable frames until its own rekey retry + re-converges the epochs, since nothing tears an established session down on + repeated decrypt failure. + - A session setup message naming an already-established peer no longer replaces that peer's session. The handler did this whenever `node.rekey.enabled` was false: it ran a fresh responder handshake and overwrote the entry, discarding @@ -1154,6 +1245,151 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 #### NAT traversal / Nostr discovery +- An advert or inbox-relay list returned by a relay is now checked against + the peer it claims to describe before anything else looks at it. The relay + pool verifies every event's signature but does not check a reply against + the request filter, and neither of the two options that would make it do so + is enabled, so a relay may answer a request for one author's advert with an + event it signed itself. The stale-advert refetch picked the newest + `created_at` across everything returned, with no author test, and then wrote + the result into the advert cache under the requested peer's npub, so a + single hostile or compromised advert relay could pin an endpoint set of its + own choosing for that peer. The author test now runs before the timestamp + contest rather than after, so a future-dated foreign event cannot even + suppress the genuine advert by winning it. The same filter now applies to + the inbox-relay lookup, where the omission let an attacker-authored relay + list steer this node's direct-message and traversal-signal traffic. A + refetch that comes back with events, none of them signed by the peer, now + leaves the cached entry alone: that is no evidence the advert was + withdrawn, and evicting on it would hand the same relay a way to clear the + cache. A refetch that genuinely comes back empty still evicts. + +- An advert's `created_at` is now clamped forward to the same 60s of clock + skew the traversal-signal path already tolerates. An unbounded future + timestamp bought a cache entry a proportionally distant validity horizon + and an unbeatable position in every replacement comparison, so a later + genuine advert could never displace it and the size-cap eviction collected + it last. The clamp applies to the stored timestamp as well as the validity + window, at all three points where an advert is cached, so ordering and + expiry now agree. Clamping rather than refusing the event is deliberate: a + node whose own clock runs slow reads every peer's honest advert as + future-dated, and refusing would silently withdraw Nostr-mediated dialing + for every peer at once. + +- Inbound traversal signals are now rate limited before they are decrypted. A + rendezvous-enabled node handed every kind-21059 event straight to the unwrap, + which is two NIP-44 decrypts and a signature verify, inline on the single + task that also routes traversal answers and maintains the advert cache. + Nothing bounded how fast an unauthenticated stranger could schedule that + work: the per-npub offer admission cannot, because it keys on the sender's + public key, which only exists once the first decrypt has already run, and + because it is a concurrency semaphore rather than a limit over time. A token + bucket now sits ahead of the unwrap, so a flood costs a node its inbound + offers instead of the whole notify loop. + + What the limit can and cannot key on is worth stating, because it decides the + shape of the fix. Before decryption there is no sender identity at all: the + outer event is signed by a key generated per event, so bucketing on its + author would hand an attacker a fresh allowance for free, and the timestamp + and recipient tag are equally attacker-chosen. The arrival relay is drawn + from our own configured set but is not an isolation boundary either, since an + attacker publishes to the same relays an honest peer does. The shared + allowance is therefore a single global bucket and is indiscriminate by + construction, which on its own would shed our own traversals along with the + attacker's, and since the attacker sets the rate every retry would land in + the same shed. A second, smaller allowance is held in reserve and drawn only + while this node has traversals of its own outstanding, so a flood denies a + node its inbound offers, which nothing receiver-side can prevent without a + pre-decrypt identity, rather than also denying it the answers to offers it + sent. Shed signals are counted and reported at debug level per event with a + warning each time the running total doubles, so a bucket sized below a busy + node's real need shows up in the log rather than as apparent relay flakiness. + + Two limits on what this buys. It bounds the crypto path only: the advert + branch runs earlier in the same loop and is not metered here, so a stranger + can still put JSON parsing and a cache insert on the task per event. And the + relay SDK verifies each event's outer signature on its own per-relay task + before this loop ever sees it, which no receiver-side change short of + dropping the subscription can avoid. + +- A STUN binding response is now accepted only from the address the binding + request was sent to. The client discarded the source address `recv_from` + returned and let the parser decide, and the parser checks only the message + type, the magic cookie and the 12-byte transaction id. An on-path attacker + who could read the outbound request could therefore inject a reply carrying + a transaction id copied from it, and its chosen address became the reflexive + candidate the node published in a traversal offer or answer, redirecting the + peer's hole-punch packets. Datagrams from any other source are counted and + discarded, and one debug record per STUN attempt reports the count and the + last unexpected source, so a rejection is diagnosable without giving a + flooder control of the log rate. A server that answers from an address other + than the one dialed, which RFC 5389 forbids, now times out and the next + configured server is tried. + +- The exemption that lets a peer's reflexive address skip the private-address + gate is now conditional on our own vantage point. That exemption exists for + the deployment whose STUN server sits inside the private network, so the + observed reflexive address is legitimately private; it was applied + unconditionally, so a node whose own STUN result was public still punched + whatever private address a peer named as its reflexive one. Any sender whose + offer or answer was accepted could therefore aim a burst of UDP packets, + carrying this node's source address, at a host inside the node's own private + network, which is the one place the candidate filter was written to keep it + out of. The gate now applies whenever our own reflexive address is public. + It stays lifted when our own reflexive address is itself private, which is + the LAN-STUN deployment the exemption was for, and also when we have no + reflexive address at all, so a failed STUN probe cannot cost a node its + same-LAN peering. Two consequences to state rather than discover: a peer + behind a private STUN server talking to a node with a public one loses its + reflexive candidate, which was never reachable from us in any case, and + because the /24 comparison is IPv4-only a unique-local IPv6 reflexive + address is refused unless our own reflexive address is unique-local too. + An off-subnet refusal of a peer's reflexive address is a shape an honest + deployment now produces, so it no longer raises the refusal record to + warning level on its own; the never-routable, port-0 and unparsable classes + still do. + +- A peer's candidate list is now bounded before it is walked rather than only + after. The eight-target cap ran after both planning loops had finished, so + it bounded what a node punched but not what it spent deciding: a signal + naming several thousand candidates had every one of them parsed and vetted, + and the deduplicating scan that follows is quadratic in the plan those + candidates feed. At most 32 candidates are now vetted, four times the + target cap and four times what the candidate generator produces on the + widest host, and the excess is discarded rather than failing the offer, so + an honest many-homed peer loses the tail of its list instead of its + traversal. The refusal record carries the discarded count as a new + `over_offered` field and treats a non-zero one as an attack shape, since + nothing honest reaches the bound. + +- A NAT-punch packet is now accepted only from an address this node planned to + probe. The punch packet's discriminator is a plain digest of the session id, + a value both peers already know, and it travels in the clear in every probe, + so acceptance proved only that the sender had seen one. The receive loop + broke on the first packet whose digest matched, whatever its source, and + returned that source as the peer address, so anyone who observed a probe, or + who could reach the node and guess the session id, could have an arbitrary + address adopted as the peer: the legitimate traversal was denied, the Noise + handshake and its retransmissions went to an address of the attacker's + choosing, and the pair was charged a failure against its backoff state. The + npub-pinned handshake still could not authenticate to the wrong host, so this + was a denial and a misdirection rather than an impersonation. The source + address is now ranked against the planned target list before anything else: + an unplanned source is dropped and, deliberately, is not acked either, since + acking it is a reflection the node controls. A source matching a planned + target exactly is adopted immediately, as before. A source matching a planned + target's IP on a different port is what a symmetric NAT's fresh mapping looks + like, and it is still adopted, because that is the main class of NAT pairing + punching exists to rescue; it is held as a candidate for 250 ms first, so an + exact match arriving inside that window supersedes it. The honest path's + latency is unchanged. Two consequences to state rather than discover: an + attacker that can source packets from a planned target's IP on any port is + still accepted, which is the residue only an authenticated probe can close; + and an attempt under a flood of spoofed matching packets now runs to its full + timeout instead of ending on the first one, so the refused sources are + counted and reported once when the attempt ends rather than logged per + packet. + - Traversal punch targets taken from a peer's offer or answer are now filtered and bounded. A rendezvous-enabled node previously punched every address a signed offer named, including loopback, link-local, multicast, @@ -1206,8 +1442,104 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 attributes the acceptance to clock skew, since a peer configured with a longer signalling TTL than ours now reaches it too. +#### DNS responder + +- The DNS responder's mesh-interface filter now works on macOS and FreeBSD, + where it had never run. The filter drops `.fips` queries that arrive over the + mesh TUN, which is what keeps a widened `dns.bind_addr` from exposing the + hosts file's alias space to every mesh peer. It was keyed on the interface + index resolved from the *configured* TUN name, but macOS and FreeBSD assign + the device a name of the kernel's choosing (`utunN`, `tunN`), so the lookup + found nothing, the index came back `None`, and `None` disables the filter. + The index is now resolved from the name of the device the node actually + created, which the TUN startup path already records, and a live device whose + index will not resolve is logged rather than passed off as "no mesh + interface". Linux is unaffected, since the configured name is the device's + name there. **Behaviour change on macOS and FreeBSD**: a node with a + non-loopback `dns.bind_addr` stops answering `.fips` queries that arrive over + the mesh interface. **What this does not close**: with an app-owned TUN the + node never learns a device name, so the filter stays off there. **Not + measured**: whether macOS and FreeBSD attribute a locally originated query + sent to the node's own mesh address to the TUN interface, as Linux does. If + they do, such a query is now dropped on those platforms; the shipped resolver + drop-in targets `[::1]` rather than the mesh address, so the packaged path is + not affected. + #### Data-plane / routing signals +- A transit node's induced routing errors are now bounded by the authenticated + link peer that induced them. The 100 ms suppression gate on + `CoordsRequired`, `PathBroken` and `MtuExceeded` was keyed on the failed + datagram's destination address, which is an envelope field the sender picks, + so a fresh random destination on every packet was always a first sighting and + every packet was admitted. Each admission also inserted a key and then walked + the whole map, so per-packet cost grew with the flood rate while the sender's + cost stayed flat, and the error itself is addressed to the datagram's source + address, which nothing binds to the sender either. A new per-link-peer token + bucket, 20 signals a second sustained with a burst of 50, is now consulted + first, keyed on the AEAD-authenticated peer the frame arrived over: the one + value at that point a sender cannot mint. The per-destination interval is + kept unchanged behind it, because it still does the aggregate suppression a + genuine outage needs, and no gate was added on the address the error is + returned to, which would have handed a sender a way to silence honest errors + toward a victim it names by keeping that victim's key hot. + + Two ordering choices in there rather than left to be discovered. The peer's + token is peeked and only spent once the destination gate has also admitted, + so a single unroutable destination behind a high-fanout peer cannot burn that + peer's whole budget on signals nothing sends and silence every other + destination behind it. And the destination map now carries a hard ceiling of + 4096 entries with its expiry sweep amortized to once per eviction interval + rather than run on every admission; when it is full it admits without + recording rather than refusing, because refusing would turn a full map into + node-wide silence exactly during partition healing, when many destinations + are legitimately unroutable at once. Emission stays bounded by the peer + budget in that state. Three counters, rendered on the fipstop Routing tab, + make each of the three outcomes visible instead of silent. + +- A transit-emitted `PathBroken` no longer carries the reporter's cached + coordinates for the unreachable destination. The signal is returned to the + datagram's source address, so anyone able to reach the node could name any + address and have the node's coordinate cache read back to them, one entry per + packet. The field is optional on the wire and no receiver reads it, so this + is an emission change only: an unmodified peer parses the frame exactly as + before. Which of the two signals is emitted still discloses whether the entry + exists. +- A reactive `MtuExceeded` is now believed only when this node has actually + sent a frame larger than the bottleneck it reports. The signal is + unauthenticated: the admission gate narrows which destination may be named + but cannot say who named it, so a value at the floor was a legal value from + anyone, and one datagram drove a bound session's path MTU to 256 and pinned + the address-keyed entry the SYN-time MSS clamp reads, with recovery costing + three consecutive higher notifications across two notification intervals. + Each session now carries the largest frame this node has put on the wire + toward it since the last accepted decrease, and a report is refused unless it + names something smaller. Honest path-MTU discovery satisfies that by + construction, because the report exists only because a frame we sent did not + fit; a forgery has to wait for us to emit something bigger than the value it + wants to claim, which bounds every accepted claim from below by our own + traffic. The evidence is cleared on each accepted decrease and on release, so + one large send early in a session cannot vouch for the rest of it. + + The guard sits ahead of both effects rather than between them, which is also + where the existing floor check moved to: the floor previously ran after the + session's own path MTU had already been changed and so governed only the + lookup table. The reactive carrier now names its own floor constant, held + equal to the actionable floor so no hop legitimately configured with a small + transport MTU loses its feedback; corroboration, not the floor's value, is + what stops a legal-but-forged claim. A separate counter, rendered on the + fipstop Routing tab, distinguishes an uncorroborated refusal from a + below-floor one. + +- The path-MTU release a `PathBroken` drives is now rate limited per + destination on a budget of its own. That signal is unauthenticated too, and + the release discards a bottleneck this node learned by having a packet + dropped, so repeating the claim discarded a genuine value as fast as it could + be relearned. The limiter is a separate instance rather than the one the + coordinate warmup send already uses: a budget another signal can spend is not + a bound. Deferring a release is the safe direction, since the value kept is + the tighter one. + - The influence a remote party has over path MTU is now bounded, and the per-destination path MTU cache has a way back. The `path_mtu` field is an unsigned per-hop transit annotation carried outside the signed proof, and the @@ -1282,8 +1614,139 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 claimed source and destination pairing no honest forwarder could produce. The drop log line now carries the signal type and the refusal class. +- A discovery lookup response is now acted on only when it answers a lookup + this node actually has outstanding. The originator path took any response + whose `request_id` was not in the transit dedup map, so an admitted peer + could harvest one genuine signed response for a target and re-inject it at + will: each injection cleared the victim's in-flight lookup, recorded a + reachability success for a target that might be unreachable, refreshed the + cached coordinates for a further full TTL, and flushed the victim's queued + packets onto a route at a moment the sender chose. It also reached the + signature verify before any check that the response was wanted, so the + verify was the first cost gate on the path. The node now records the + `request_id` of every lookup request it sends on that target's pending + entry, and a response is dropped unless it names a target with a lookup + outstanding and carries one of the ids issued for it. Because the id is + fresh 64-bit randomness drawn per attempt and the target signs over it, a + harvested response is bound to the request it answered and cannot be + redirected or replayed. The check runs before the identity-cache resolve + and before the signature verify, so a response nobody asked for costs + nothing. Replies to earlier attempts of a still-outstanding lookup are + still accepted, which is the common case on a link whose round trip + exceeds the first rung of the retry ladder. Drops are counted as + `resp_unsolicited`, visible through `show routing`, `show metrics` and the + fipstop routing pane; the counter has a nonzero floor in healthy operation, + because a request is flooded to every qualifying tree peer and the + duplicate replies land there once the first has been accepted. + +- A flooded discovery dedup cache no longer makes a node unresolvable. The + cache is both the duplicate filter and the reverse-path table for lookup + responses, and at its 4096-entry bound it dropped the arriving request. + That drop sat ahead of both the check for whether the request names this + node and the forwarding path, so one link peer emitting fresh request_ids + could stop the node answering lookups for itself and stop it carrying + anyone else's, for as long as it kept the cache full. The cache now makes + room instead of refusing: over a peer's own share it drops that peer's + oldest entry, and at global capacity it drops the oldest entry of whichever + peer holds the most, so a light peer's reverse path is never taken to admit + a heavy one and extra identities buy a flooder proportionally less. A + peer's share is the cache divided by the current link-peer count, with a + floor of 64. The loosening this accepts is that an evicted request_id + arriving again inside the window is forwarded a second time rather than + recognised as a duplicate, which the per-target forward limiter and TTL + already bound. Evictions are counted as `req_dedup_evicted`; the old + `req_dedup_cache_full` counter stays in place, frozen at zero, so a + dashboard carried across versions does not lose the series. + +- Answering a lookup for ourselves is now metered per link peer. The response + proof is signed over the requester's `request_id`, so every request + addressed to this node costs a fresh Schnorr signature that cannot be + cached or served twice, and until now the only thing bounding that rate was + the dedup cache filling up, which is the defect above. A token bucket per + link peer, 256 signatures of burst refilling at 32 per second, absorbs the + legitimate burst that follows a topology change, when many correspondents + re-look-up at once through the few links that lead here, while capping what + one neighbour can make the node sign. Refusals are counted as + `req_sign_rate_limited` and visible in `show routing`, `show metrics` and + the fipstop routing pane. A refused request keeps its dedup entry, and + retries carry fresh request_ids, so a refusal cannot suppress the retry. + #### Admission / peer caps +- The Ethernet transport's discovery buffer is now bounded and no longer costs + a linear scan per beacon. Beacons are unauthenticated broadcast frames, and + the buffer deduplicated by scanning a `Vec` for the source MAC and had no + cap, so anything on the segment could name a fresh MAC per frame and drive + both quadratic CPU in the receive loop and unbounded memory. It is drained + once per tick only while the transport is operational, so a transport that + is receiving but not operational was never drained at all. The buffer is now + a map keyed on source MAC, capped at 1024 distinct MACs between drains, with + the drain order still oldest sighting first so which neighbour gets dialed + under a connect budget does not depend on hash iteration order. A MAC already + buffered is always refreshed, so a flood of new MACs cannot crowd out a + neighbour already seen. Refused beacons are counted in the transport's stats + as `beacons_dropped` and reported in the log on the first drop and then on + each power-of-ten thereafter, so the flooder does not set the log rate. + **What this does not close**: a flood can still crowd out a neighbour not yet + seen in that tick, and anything able to flood raw frames on the segment can + already jam the beacon at L2 more cheaply. + +- A read failure on `peers.allow`, `peers.deny` or the `hosts` file no longer + turns the node into an open one. Every read error other than a steadily + absent file was logged and swallowed, leaving that file's entries empty, and + the reloader published the result unconditionally: an unreadable `peers.deny` + admitted the peers it named, and an unreadable `peers.allow` took a node from + admitting a named few to admitting everyone, with one warning line as the + only signal. Because the recorded modification times advanced before the + load, nothing retried until the file changed again, so a persistent + permission or I/O fault left the empty ACL in force indefinitely. The + reloader now keeps the last loaded ACL when any input is present but + unreadable, leaves the modification times alone, retries on the next tick + regardless of them, and logs the fault once on the transition rather than + once per tick. An absent file is still a policy and still loads as an empty + set; a `NotFound` that a successful stat contradicts is treated as a file + being rewritten under us and held. A reload whose inputs all read cleanly but + which empties an enforcing ACL while its files are still on disk is held for + one tick, which catches a read that caught a non-atomic in-place edit + mid-write, and released on the next so a deliberate blanking still takes + effect. `fipsctl` ACL status gains a `stale` flag reporting that the policy + in force is older than the files on disk. **What this does not close**: + there is no last-good snapshot at startup, so a node whose ACL file is + unreadable at boot still comes up with no entries, now logged as an error and + armed to retry on the first tick. Admission is also checked only at handshake + time, so a peer admitted during a window that has already happened keeps its + link. + +- The end-to-end session table now has a bound. It was the one remotely-grown + map with none: an inbound SessionSetup naming an address nobody had seen + inserted an entry, and the two existing limits did not reach it, the setup + limiter governing the arrival rate rather than the population and the idle + purge only reaching entries a peer stops using. One neighbour sending setups + at the permitted rate could hold roughly 1440 half-open entries at any + moment and grow the table without limit by keeping them warm. Setups that + would grow the table past `node.limits.max_sessions` are now refused, ahead + of the setup limiter, so a full table costs no token, no responder handshake + and no ack; a refused setup emits nothing at all, which is indistinguishable + from loss to the sender and is already covered by its own msg1 resend + schedule. The test is whether admitting would grow the table, not whether + the sender is a stranger, so a resent setup for an entry already present is + still served and an in-flight handshake is not broken. Unauthenticated + half-open entries are additionally held to half the table, so a handshake + flood cannot deny the whole of it to peers that complete; that share is sized + to leave a reconnect storm, where every peer initiates at once after a + restart or a healed partition, room to land. Locally originated sessions are + capped at the same ceiling, answered with ICMPv6 destination unreachable so + the application gets an immediate error rather than a silent drop. The cap + refuses rather than evicts: the setup that triggers the decision is + unauthenticated at that point, so evicting would hand a stranger a way to + tear down sessions it has nothing to do with. Refusals are counted as + `table_full` and `half_open_full` in the session reject family. What stays + open is per-neighbour fairness among established sessions: one hostile + neighbour that completes handshakes and keeps each session warm can occupy + the table and hold new session establishment closed for as long as it keeps + doing so, which is a denial of new sessions rather than the unbounded memory + growth it replaces. + - An accepted inbound TCP connection no longer holds a slot indefinitely without sending anything. The cap was tested at accept and the pool insert and counter bump followed with no read in between, while the frame reader's @@ -1322,6 +1785,30 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 NXDOMAIN. Connecting the socket also means a dead upstream surfaces ECONNREFUSED immediately instead of stalling for five seconds. +#### Control socket + +- The control socket and the directory holding it are now created with a + restrictive mode rather than created wide and narrowed afterwards. `bind(2)` + makes the socket inode `0777 & ~umask`, so under a permissive umask the + socket was world-accessible for the window between the bind and the `chmod` + to 0770 that followed it; the bind now runs under a umask that masks the + "other" bits, so the inode is 0770 from creation and the chmod and chown stay + the authority on its final mode. The parent directory was worse than a + window: it was created with `create_dir_all`, which is also `0777 & ~umask`, + and nothing ever set a mode on it, so under a permissive umask the directory + holding the socket stayed world-writable for the life of the host, and a + world-writable parent lets an unprivileged account plant an entry at the + socket path. Directories this code creates now come out 0750, which is what + the systemd unit (`RuntimeDirectoryMode=0750`) and the FreeBSD rc script + (`install -d -m 0750`) already apply, so no packaged deployment sees a + different mode and no `fipsctl` user loses access. Both the daemon and the + gateway control sockets are covered. **What this does not close**: the window + between the stale-socket probe and the bind is documented at the site rather + than removed. Reaching it needs write access to the socket's parent + directory, which the packaged layouts give to root alone, and an account + holding it can deny the daemon its socket more simply by squatting the path + first. + #### Key material and identity files - Private key writes no longer follow a symlink, and the key file's mode is diff --git a/src/bin/fipstop/ui/routing.rs b/src/bin/fipstop/ui/routing.rs index 4f016c5a..1961c8c5 100644 --- a/src/bin/fipstop/ui/routing.rs +++ b/src/bin/fipstop/ui/routing.rs @@ -192,6 +192,8 @@ fn draw_routing_stats( ("Bloom Miss", lookup("req_bloom_miss")), ("Backoff Suppressed", lookup("req_backoff_suppressed")), ("Fwd Rate Limited", lookup("req_forward_rate_limited")), + ("Sign Rate Limited", lookup("req_sign_rate_limited")), + ("Dedup Evicted", lookup("req_dedup_evicted")), ("TTL Exhausted", lookup("req_ttl_exhausted")), ("Decode Error", lookup("req_decode_error")), ], @@ -206,6 +208,7 @@ fn draw_routing_stats( ("Timed Out", lookup("resp_timed_out")), ("Identity Miss", lookup("resp_identity_miss")), ("Proof Failed", lookup("resp_proof_failed")), + ("Unsolicited", lookup("resp_unsolicited")), ("Decode Error", lookup("resp_decode_error")), ], )); @@ -298,6 +301,13 @@ fn draw_routing_stats( ("Path Broken Refused", err("unbound_broken")), ("MTU Exceeded Refused", err("unbound_mtu")), ("Forged Pairing", err("unbound_forged")), + ("Emit Over Peer Budget", err("emit_over_peer_budget")), + ("Emit Over Dest Interval", err("emit_over_dest_interval")), + ("Emit Limiter At Capacity", err("emit_limiter_at_capacity")), + ( + "MTU Exceeded Uncorroborated", + err("mtu_exceeded_uncorroborated"), + ), ], )); right.push(Line::from("")); diff --git a/src/bin/fipstop/ui/snapshots.rs b/src/bin/fipstop/ui/snapshots.rs index b108cc66..5aa93f72 100644 --- a/src/bin/fipstop/ui/snapshots.rs +++ b/src/bin/fipstop/ui/snapshots.rs @@ -1151,7 +1151,7 @@ fn routing_focused_pane_scrolls() { // column is the taller of the two, so scrolling fully to the bottom would // over-scroll the right column past Congestion; this offset lands the // Congestion region inside the short window instead. - app1.scroll_offsets.insert((Tab::Routing, 2), 28); + app1.scroll_offsets.insert((Tab::Routing, 2), 32); let buf1 = testkit::render(100, 20, |frame, area| { super::routing::draw(frame, &app1, area); }); diff --git a/src/config/node.rs b/src/config/node.rs index 578491b9..20969af4 100644 --- a/src/config/node.rs +++ b/src/config/node.rs @@ -28,6 +28,20 @@ pub struct LimitsConfig { /// Max pending inbound handshakes (`node.limits.max_pending_inbound`). #[serde(default = "LimitsConfig::default_max_pending_inbound")] pub max_pending_inbound: usize, + /// Max end-to-end sessions (`node.limits.max_sessions`), `0` = unlimited. + /// + /// The session table is the only remotely-grown map with no bound: an + /// inbound SessionSetup from an address nobody has seen inserts an + /// entry, and the idle purge only reaches entries a peer stops using. + /// The default of 1024 is four times the adjacent + /// `node.session.pending_max_destinations`. One entry measures 6608 + /// bytes of inline state plus heap, so the table holds to roughly 7 MB + /// and a test pins the per-entry figure the default rests on. Raising it + /// raises the memory an attacker can make this node hold; lowering it + /// refuses new sessions sooner on a node that legitimately talks + /// end-to-end to many others, such as a gateway. + #[serde(default = "LimitsConfig::default_max_sessions")] + pub max_sessions: usize, } impl Default for LimitsConfig { @@ -37,6 +51,7 @@ impl Default for LimitsConfig { max_peers: 128, max_links: 256, max_pending_inbound: 1000, + max_sessions: 1024, } } } @@ -54,6 +69,9 @@ impl LimitsConfig { fn default_max_pending_inbound() -> usize { 1000 } + fn default_max_sessions() -> usize { + 1024 + } } /// Rate limiting (`node.rate_limit.*`). diff --git a/src/control/snapshot.rs b/src/control/snapshot.rs index 4a7cd47b..bea9d8d7 100644 --- a/src/control/snapshot.rs +++ b/src/control/snapshot.rs @@ -134,6 +134,7 @@ fn empty_acl_status() -> PeerAclStatus { deny_file_entries: Vec::new(), allow_entries: Vec::new(), deny_entries: Vec::new(), + stale: false, } } diff --git a/src/control/snapshots/show_routing.json b/src/control/snapshots/show_routing.json index 68117f54..0e1dad20 100644 --- a/src/control/snapshots/show_routing.json +++ b/src/control/snapshots/show_routing.json @@ -12,6 +12,7 @@ "req_bloom_miss": 0, "req_decode_error": 0, "req_dedup_cache_full": 0, + "req_dedup_evicted": 0, "req_deduplicated": 0, "req_duplicate": 0, "req_fallback_forwarded": 0, @@ -20,6 +21,7 @@ "req_initiated": 0, "req_no_tree_peer": 0, "req_received": 0, + "req_sign_rate_limited": 0, "req_target_is_us": 0, "req_ttl_exhausted": 0, "resp_accepted": 0, @@ -29,13 +31,18 @@ "resp_no_route": 0, "resp_proof_failed": 0, "resp_received": 0, - "resp_timed_out": 0 + "resp_timed_out": 0, + "resp_unsolicited": 0 }, "error_signals": { "coords_required": 0, + "emit_limiter_at_capacity": 0, + "emit_over_dest_interval": 0, + "emit_over_peer_budget": 0, "lookup_resp_mtu_below_floor": 0, "mtu_exceeded": 0, "mtu_exceeded_below_floor": 0, + "mtu_exceeded_uncorroborated": 0, "path_broken": 0, "path_mtu_notif_below_floor": 0, "unbound_broken": 0, @@ -77,6 +84,7 @@ "req_bloom_miss": 0, "req_decode_error": 0, "req_dedup_cache_full": 0, + "req_dedup_evicted": 0, "req_deduplicated": 0, "req_duplicate": 0, "req_fallback_forwarded": 0, @@ -85,6 +93,7 @@ "req_initiated": 0, "req_no_tree_peer": 0, "req_received": 0, + "req_sign_rate_limited": 0, "req_target_is_us": 0, "req_ttl_exhausted": 0, "resp_accepted": 0, @@ -94,7 +103,8 @@ "resp_no_route": 0, "resp_proof_failed": 0, "resp_received": 0, - "resp_timed_out": 0 + "resp_timed_out": 0, + "resp_unsolicited": 0 }, "pending_lookups": [], "pending_tun_destinations": 0, diff --git a/src/node/acl.rs b/src/node/acl.rs index 6dafd94e..1c53647f 100644 --- a/src/node/acl.rs +++ b/src/node/acl.rs @@ -20,7 +20,7 @@ use std::fmt; use std::path::{Path, PathBuf}; use std::sync::Arc; use std::time::SystemTime; -use tracing::{debug, info, warn}; +use tracing::{debug, error, info, warn}; /// Default path for the peer allow list. /// @@ -106,6 +106,28 @@ pub enum PeerAclContext { OutboundHandshake, } +/// How many consecutive reloads the empty-snapshot guard may hold back. +/// +/// A reload whose files all read cleanly but which yields an empty ACL where +/// an enforcing one was in force is more likely a read that raced an in-place +/// rewrite than a policy change, so the previous snapshot is held. The guard +/// releases after this many holds, so an operator who deliberately blanks an +/// ACL file in place still converges, one tick late. Raising it widens the +/// window in which a genuine emptying is ignored; setting it to zero disables +/// the torn-read protection. +const EMPTY_ACL_HOLD_LIMIT: u32 = 1; + +/// A peer ACL input file exists but could not be read. +#[derive(Debug, thiserror::Error)] +#[error("failed to read {}: {source}", path.display())] +pub struct AclLoadError { + /// The file whose read failed. + pub path: PathBuf, + /// The underlying I/O failure. + #[source] + pub source: std::io::Error, +} + /// Snapshot of the currently loaded ACL state. #[derive(Debug, Clone, PartialEq, Eq, Serialize)] pub struct PeerAclStatus { @@ -120,6 +142,9 @@ pub struct PeerAclStatus { pub deny_file_entries: Vec, pub allow_entries: Vec, pub deny_entries: Vec, + /// Whether the ACL in force is older than the files on disk because a + /// reload input could not be read. + pub stale: bool, } impl fmt::Display for PeerAclContext { @@ -155,26 +180,38 @@ impl PeerAcl { #[cfg(test)] pub fn load_files(allow_path: &Path, deny_path: &Path) -> Self { let hosts = HostMap::new(); - Self::load_files_with_hosts(allow_path, deny_path, &hosts) + Self::try_load_files_with_hosts(allow_path, deny_path, &hosts).unwrap() } /// Load the allow/deny files into a new ACL using alias resolution. - pub fn load_files_with_hosts(allow_path: &Path, deny_path: &Path, hosts: &HostMap) -> Self { + /// + /// An absent file is a policy and contributes an empty set; a file that + /// is present and unreadable is a fault and is returned as an error, so + /// the caller can keep enforcing whatever it loaded last rather than + /// silently becoming an open node. + pub fn try_load_files_with_hosts( + allow_path: &Path, + deny_path: &Path, + hosts: &HostMap, + ) -> Result { let mut acl = Self::new(); - acl.load_file(allow_path, true, hosts); - acl.load_file(deny_path, false, hosts); + acl.load_file(allow_path, true, hosts)?; + acl.load_file(deny_path, false, hosts)?; + acl.log_loaded(); + Ok(acl) + } - if !acl.is_empty() { + /// Log the shape of a freshly loaded ACL, unless it has no entries. + fn log_loaded(&self) { + if !self.is_empty() { debug!( - allow_entries = acl.allow.len(), - deny_entries = acl.deny.len(), - allow_all = acl.allow_all, - deny_all = acl.deny_all, + allow_entries = self.allow.len(), + deny_entries = self.deny.len(), + allow_all = self.allow_all, + deny_all = self.deny_all, "Loaded peer ACL files" ); } - - acl } /// Evaluate whether a peer is allowed. @@ -245,16 +282,29 @@ impl PeerAcl { self.deny_file_entries.iter().cloned().collect() } - fn load_file(&mut self, path: &Path, is_allow: bool, hosts: &HostMap) { + /// Merge one ACL file into this ACL. + /// + /// An absent file is a policy and an unreadable one is a fault, and + /// `NotFound` alone does not say which: a stat that still finds the file + /// after the read missed it means the file is being rewritten under us, + /// which is transient and must not be published as an empty policy. + fn load_file( + &mut self, + path: &Path, + is_allow: bool, + hosts: &HostMap, + ) -> Result<(), AclLoadError> { let contents = match std::fs::read_to_string(path) { Ok(c) => c, - Err(e) if e.kind() == std::io::ErrorKind::NotFound => { + Err(e) if e.kind() == std::io::ErrorKind::NotFound && file_mtime(path).is_none() => { debug!(path = %path.display(), "No ACL file found, skipping"); - return; + return Ok(()); } Err(e) => { - warn!(path = %path.display(), error = %e, "Failed to read ACL file"); - return; + return Err(AclLoadError { + path: path.to_path_buf(), + source: e, + }); } }; @@ -310,6 +360,8 @@ impl PeerAcl { self.deny_npubs.insert(resolved_npub); } } + + Ok(()) } fn resolve_entry(entry: &str, hosts: &HostMap) -> Result<(PeerIdentity, String), String> { @@ -342,6 +394,13 @@ pub struct PeerAclReloader { deny_path: PathBuf, last_allow_mtime: Option, last_deny_mtime: Option, + /// Set while a reload input is unreadable. Forces the next reload + /// attempt regardless of mtimes, because the mtime comparison alone + /// cannot see a change the hosts reloader has already consumed, and + /// gates the fault log to the transition into the held state. + retry_pending: bool, + /// Consecutive reloads held back by the empty-snapshot guard. + empty_holds: u32, } impl PeerAclReloader { @@ -377,7 +436,23 @@ impl PeerAclReloader { let last_allow_mtime = file_mtime(&allow_path); let last_deny_mtime = file_mtime(&deny_path); let hosts = HostMapReloader::new(base_hosts, hosts_path); - let acl = PeerAcl::load_files_with_hosts(&allow_path, &deny_path, hosts.hosts()); + + // There is no last-good snapshot to hold at startup, so an + // unreadable file still comes up on an empty ACL, as it always has. + // It is logged as the fault it is and armed for retry, so the first + // tick after the file becomes readable enforces the real policy. + let (acl, retry_pending) = + match PeerAcl::try_load_files_with_hosts(&allow_path, &deny_path, hosts.hosts()) { + Ok(acl) => (acl, false), + Err(e) => { + error!( + path = %e.path.display(), + error = %e.source, + "Peer ACL file is present but unreadable; starting with no ACL entries" + ); + (PeerAcl::new(), true) + } + }; Self { acl: arc_swap::ArcSwap::from(Arc::new(acl)), @@ -386,9 +461,28 @@ impl PeerAclReloader { deny_path, last_allow_mtime, last_deny_mtime, + retry_pending, + empty_holds: 0, } } + /// Keep the published snapshot after a reload input failed to read. + /// + /// Leaves the recorded mtimes and the ACL in force untouched, arms the + /// retry so the next tick reloads regardless of mtimes, and logs the + /// fault once, on the transition into the held state, rather than once + /// per tick for as long as the fault lasts. + fn hold_snapshot(&mut self, path: &Path, error: &dyn fmt::Display) { + if !self.retry_pending { + error!( + path = %path.display(), + error = %error, + "Peer ACL input is unreadable; holding the last loaded ACL" + ); + } + self.retry_pending = true; + } + /// Acquire a lock-free guard over the current ACL snapshot. pub fn acl(&self) -> arc_swap::Guard> { self.load() @@ -409,6 +503,7 @@ impl PeerAclReloader { deny_file_entries: acl.deny_file_entries(), allow_entries: acl.allow_entries(), deny_entries: acl.deny_entries(), + stale: self.retry_pending, } } } @@ -419,19 +514,70 @@ impl Reloadable for PeerAclReloader { async fn reload(&mut self) -> bool { let allow_mtime = file_mtime(&self.allow_path); let deny_mtime = file_mtime(&self.deny_path); - let hosts_changed = self.hosts.check_reload(); + let hosts_changed = match self.hosts.try_check_reload() { + Ok(changed) => changed, + Err(e) => { + let path = self.hosts.path().to_path_buf(); + self.hold_snapshot(&path, &e); + return false; + } + }; if allow_mtime == self.last_allow_mtime && deny_mtime == self.last_deny_mtime && !hosts_changed + && !self.retry_pending { return false; } + let new_acl = match PeerAcl::try_load_files_with_hosts( + &self.allow_path, + &self.deny_path, + self.hosts.hosts(), + ) { + Ok(acl) => acl, + Err(e) => { + self.hold_snapshot(&e.path.clone(), &e.source); + return false; + } + }; + + // Every input read cleanly and the policy still evaporated. With the + // ACL files themselves freshly written and still on disk that is more + // likely a read that caught one mid-rewrite than an operator emptying + // both lists, so hold and look again next tick. Deleting a file, or + // dropping the aliases an entry resolved through, remains an + // unambiguous way to say "no policy" and is published immediately. + let acl_files_changed = + allow_mtime != self.last_allow_mtime || deny_mtime != self.last_deny_mtime; + if new_acl.is_empty() + && !self.acl.load().is_empty() + && acl_files_changed + && (allow_mtime.is_some() || deny_mtime.is_some()) + && self.empty_holds < EMPTY_ACL_HOLD_LIMIT + { + self.empty_holds += 1; + self.retry_pending = true; + warn!( + allow_file = %self.allow_path.display(), + deny_file = %self.deny_path.display(), + "Peer ACL reload emptied an enforcing ACL; holding the last loaded ACL" + ); + return false; + } + + if self.retry_pending { + info!( + allow_file = %self.allow_path.display(), + deny_file = %self.deny_path.display(), + "Peer ACL inputs read cleanly again; publishing the files on disk" + ); + } + self.retry_pending = false; + self.empty_holds = 0; self.last_allow_mtime = allow_mtime; self.last_deny_mtime = deny_mtime; - let new_acl = - PeerAcl::load_files_with_hosts(&self.allow_path, &self.deny_path, self.hosts.hosts()); info!( allow_file = %self.allow_path.display(), @@ -786,19 +932,15 @@ mod tests { } #[test] - fn test_acl_read_error_is_ignored() { + fn test_acl_read_error_is_reported_rather_than_yielding_an_empty_acl() { let dir = tempfile::tempdir().unwrap(); let allow = dir.path().join("peers.allow"); let deny = dir.path().join("peers.deny"); std::fs::create_dir(&allow).unwrap(); - let acl = PeerAcl::load_files(&allow, &deny); + let err = PeerAcl::try_load_files_with_hosts(&allow, &deny, &HostMap::new()).unwrap_err(); - assert!(acl.is_empty()); - assert_eq!( - acl.check(&test_peer(&test_npub())), - PeerAclDecision::DefaultAllow - ); + assert_eq!(err.path, allow); } #[test] @@ -812,7 +954,7 @@ mod tests { hosts.insert("node-a", &npub).unwrap(); write_file(&allow, "NODE-A\n"); - let acl = PeerAcl::load_files_with_hosts(&allow, &deny, &hosts); + let acl = PeerAcl::try_load_files_with_hosts(&allow, &deny, &hosts).unwrap(); assert_eq!(acl.allow_file_entries(), vec!["NODE-A".to_string()]); assert_eq!(acl.allow_entries(), vec![npub.clone()]); @@ -830,7 +972,7 @@ mod tests { hosts.insert("node-a", &npub).unwrap(); write_file(&allow, &format!("node-a\n{npub}\nnode-a\n")); - let acl = PeerAcl::load_files_with_hosts(&allow, &deny, &hosts); + let acl = PeerAcl::try_load_files_with_hosts(&allow, &deny, &hosts).unwrap(); assert_eq!( acl.allow_file_entries(), @@ -887,6 +1029,160 @@ mod tests { ); } + /// The three permission-fault tests below make a file unreadable through + /// the unix mode bits, which Windows has no equivalent for: a read-only + /// NTFS file is still readable, so the fault they need cannot be produced. + /// They are gated to unix rather than made to pass vacuously elsewhere. + /// + /// **Coverage gap**: on Windows nothing exercises the reloader's + /// unreadable-input path, so the fail-open defect this fix closes is + /// unverified there. + /// Make a file unreadable, returning false if the effective uid can read + /// it anyway. Root bypasses the mode bits, so the permission-fault tests + /// cannot run there and skip instead of passing vacuously; that leaves + /// the EACCES path unexercised in any root CI job. + #[cfg(unix)] + fn make_unreadable(path: &Path) -> bool { + use std::os::unix::fs::PermissionsExt; + std::fs::set_permissions(path, std::fs::Permissions::from_mode(0o000)).unwrap(); + std::fs::read_to_string(path).is_err() + } + + #[cfg(unix)] + fn make_readable(path: &Path) { + use std::os::unix::fs::PermissionsExt; + std::fs::set_permissions(path, std::fs::Permissions::from_mode(0o644)).unwrap(); + } + + #[cfg(unix)] + #[tokio::test] + async fn test_acl_reload_holds_last_good_snapshot_when_deny_file_unreadable() { + let dir = tempfile::tempdir().unwrap(); + let allow = dir.path().join("peers.allow"); + let deny = dir.path().join("peers.deny"); + let denied = test_npub(); + + write_file(&deny, &format!("{denied}\n")); + let mut reloader = PeerAclReloader::with_paths(allow, deny.clone()); + assert_eq!( + reloader.acl().check(&test_peer(&denied)), + PeerAclDecision::DenyList + ); + + // Rewrite before revoking access so the mtime change makes the + // reloader actually attempt the read that then fails. + std::thread::sleep(std::time::Duration::from_millis(5)); + write_file(&deny, &format!("{denied}\n")); + if !make_unreadable(&deny) { + return; + } + + assert!(!reloader.reload().await); + assert_eq!( + reloader.acl().check(&test_peer(&denied)), + PeerAclDecision::DenyList + ); + assert_eq!(reloader.acl().effective_mode(), "denylist"); + assert!(reloader.status().stale); + + make_readable(&deny); + } + + #[cfg(unix)] + #[tokio::test] + async fn test_acl_reload_retries_after_a_transient_read_error() { + let dir = tempfile::tempdir().unwrap(); + let allow = dir.path().join("peers.allow"); + let deny = dir.path().join("peers.deny"); + let denied = test_npub(); + + write_file(&deny, &format!("{denied}\n")); + let mut reloader = PeerAclReloader::with_paths(allow, deny.clone()); + std::thread::sleep(std::time::Duration::from_millis(5)); + write_file(&deny, &format!("{denied}\n")); + if !make_unreadable(&deny) { + return; + } + assert!(!reloader.reload().await); + + // No further mtime change: only the armed retry can pick this up. + make_readable(&deny); + assert!(reloader.reload().await); + assert_eq!( + reloader.acl().check(&test_peer(&denied)), + PeerAclDecision::DenyList + ); + assert!(!reloader.status().stale); + } + + #[cfg(unix)] + #[tokio::test] + async fn test_acl_reload_holds_last_good_when_the_hosts_file_becomes_unreadable() { + let dir = tempfile::tempdir().unwrap(); + let allow = dir.path().join("peers.allow"); + let deny = dir.path().join("peers.deny"); + let hosts = dir.path().join("hosts"); + let npub = test_npub(); + + write_file(&allow, "node-a\n"); + write_file(&hosts, &format!("node-a {npub}\n")); + + let mut reloader = + PeerAclReloader::with_alias_sources(allow, deny, HostMap::new(), hosts.clone()); + assert_eq!( + reloader.acl().check(&test_peer(&npub)), + PeerAclDecision::AllowList + ); + + std::thread::sleep(std::time::Duration::from_millis(5)); + write_file(&hosts, &format!("node-a {npub}\n")); + if !make_unreadable(&hosts) { + return; + } + + assert!(!reloader.reload().await); + assert_eq!( + reloader.acl().check(&test_peer(&npub)), + PeerAclDecision::AllowList + ); + assert_eq!(reloader.acl().default_decision(), "allow"); + assert_eq!( + reloader.acl().allow_file_entries(), + vec!["node-a".to_string()] + ); + + make_readable(&hosts); + } + + #[tokio::test] + async fn test_acl_reload_does_not_publish_an_empty_acl_over_an_enforcing_one() { + let dir = tempfile::tempdir().unwrap(); + let allow = dir.path().join("peers.allow"); + let deny = dir.path().join("peers.deny"); + let allowed = test_npub(); + + write_file(&allow, &format!("{allowed}\n")); + let mut reloader = PeerAclReloader::with_paths(allow.clone(), deny); + assert_eq!( + reloader.acl().check(&test_peer(&allowed)), + PeerAclDecision::AllowList + ); + + std::thread::sleep(std::time::Duration::from_millis(5)); + write_file(&allow, ""); + + assert!(!reloader.reload().await); + assert_eq!( + reloader.acl().check(&test_peer(&allowed)), + PeerAclDecision::AllowList + ); + + // The hold is bounded: a file the operator really did blank in place + // is published on the following tick. + assert!(reloader.reload().await); + assert!(reloader.acl().is_empty()); + } + #[test] fn test_acl_status_reports_effective_state_and_entries() { let dir = tempfile::tempdir().unwrap(); @@ -977,7 +1273,7 @@ mod tests { hosts.insert("node-a", &npub).unwrap(); std::fs::write(&allow, "node-a\n").unwrap(); - let acl = PeerAcl::load_files_with_hosts(&allow, &deny, &hosts); + let acl = PeerAcl::try_load_files_with_hosts(&allow, &deny, &hosts).unwrap(); let peer = PeerIdentity::from_npub(&npub).unwrap(); assert_eq!(acl.allow_file_entries(), vec!["node-a".to_string()]); diff --git a/src/node/dataplane/forwarding.rs b/src/node/dataplane/forwarding.rs index 7c514f86..725a8de5 100644 --- a/src/node/dataplane/forwarding.rs +++ b/src/node/dataplane/forwarding.rs @@ -16,7 +16,7 @@ use crate::proto::fsp::wire::{ }; use crate::proto::fsp::{SessionAck, SessionSetup}; use crate::proto::link::{SessionDatagram, SessionDatagramRef}; -use crate::proto::routing::{DropReason, NextHop, RouteAction, RouteOutcome}; +use crate::proto::routing::{DropReason, LimitVerdict, NextHop, RouteAction, RouteOutcome}; use std::time::{Duration, Instant}; use tracing::{debug, warn}; @@ -125,7 +125,7 @@ impl Node { bytes = payload.len(), "Dropping transit SessionDatagram: no route to destination" ); - self.send_routing_error(&original).await; + self.send_routing_error(from, &original).await; } RouteOutcome::Forward { next_hop, @@ -157,7 +157,7 @@ impl Node { self.metrics() .forwarding .record_reject_bytes(ForwardingReject::MtuExceeded, payload.len()); - self.send_mtu_exceeded_error(dest, datagram_ref.src_addr, mtu) + self.send_mtu_exceeded_error(from, dest, datagram_ref.src_addr, mtu) .await; } Err(e) => { @@ -321,6 +321,39 @@ impl Node { } } + /// Spend one peer-budget token, after a gate further down has admitted. + /// + /// The budget is keyed on the authenticated link peer the frame arrived + /// over, which is the one value at the emission point a sender cannot + /// mint: every field of the datagram itself is chosen by whoever sent it, + /// so a per-destination or per-source gate is escaped by varying the field + /// it keys on. + /// + /// Peek and commit are separate because the per-destination interval gate + /// lives inside `routing::synth_routing_error` and runs after this. + /// Charging a suppressed signal would let a single unroutable destination + /// behind a high-fanout peer spend that peer's whole budget on emissions + /// nothing sends, silencing every other destination behind it. + fn commit_error_emission(&mut self, from: &NodeAddr) { + self.peer_error_budget.commit(from, Instant::now()); + } + + /// Count what the core's per-destination gate decided about one candidate + /// error signal. + /// + /// The three verdicts are counted apart because they mean different + /// things to an operator: `Suppress` is the interval doing its job during + /// an outage, while `AdmitAtCapacity` says the destination map is full and + /// the interval is no longer suppressing anything for this destination, so + /// only the per-peer budget is still bounding emission. + fn record_error_verdict(&mut self, verdict: LimitVerdict) { + match verdict { + LimitVerdict::Suppress => self.metrics().errors.emit_over_dest_interval.inc(), + LimitVerdict::AdmitAtCapacity => self.metrics().errors.emit_limiter_at_capacity.inc(), + LimitVerdict::Admit => {} + } + } + /// Generate and send a routing error signal back to the datagram's source. /// /// If we have cached coords for the destination, send PathBroken (we know @@ -329,7 +362,20 @@ impl Node { /// /// If we can't route the error back to the source either, drop silently. /// No cascading errors. - async fn send_routing_error(&mut self, original: &SessionDatagram) { + /// `from` is the authenticated link peer the original datagram arrived + /// from, and is what the emission is charged against. It is the one value + /// at this point a sender cannot mint: every field of the datagram itself + /// is chosen by whoever sent it. + async fn send_routing_error(&mut self, from: &NodeAddr, original: &SessionDatagram) { + // Peeked, not spent. The destination gate inside the core may still + // suppress this signal, and charging a suppressed emission would let a + // single unroutable destination behind a high-fanout peer burn that + // peer's whole budget on signals nothing sends. + if !self.peer_error_budget.has_token(from, Instant::now()) { + self.metrics().errors.emit_over_peer_budget.inc(); + return; + } + let my_addr = *self.node_addr(); let now_ms = std::time::SystemTime::now() .duration_since(std::time::UNIX_EPOCH) @@ -356,12 +402,22 @@ impl Node { default_ttl, ) }; - let RouteAction::SendError { toward, bytes } = match action { + self.record_error_verdict(action.verdict); + let RouteAction::SendError { toward, bytes } = match action.action { Some(action) => action, // Rate limited: drop silently. No cascading errors. None => return, }; + // Both gates have admitted, so the token peeked above is now spent. + // Charged here rather than at the peek so a destination the core + // suppressed costs the link peer nothing; see + // `commit_error_emission`. A later failure to resolve the reverse hop + // still leaves the token spent, which is deliberate: the work the + // budget bounds is the synthesis this node was induced to perform, + // not whether a hop happened to exist for it. + self.commit_error_emission(from); + // Resolve the reverse link hop only now, after the gate passed, so // `find_next_hop`'s coord-cache touch keeps its pre-refactor scope. let next_hop_addr = match self.find_next_hop(&toward) { @@ -402,12 +458,28 @@ impl Node { /// /// `dest` is the failed datagram's destination (rate-limit key); `toward` /// is its source, where the signal is routed back. + /// + /// `from` is the authenticated link peer the original datagram arrived + /// from, and is what the emission is charged against. MtuExceeded shares + /// the link peer's budget with the routing errors rather than holding its + /// own: a separate bucket would insulate path-MTU discovery from + /// routing-error pressure, at the cost of a second knob and of letting one + /// peer induce twice the total emission. async fn send_mtu_exceeded_error( &mut self, + from: &NodeAddr, dest: NodeAddr, toward: NodeAddr, bottleneck_mtu: u16, ) { + // Peeked, not spent, for the same reason as in `send_routing_error`: + // the per-destination gate inside the core runs below and may still + // suppress this signal. + if !self.peer_error_budget.has_token(from, Instant::now()) { + self.metrics().errors.emit_over_peer_budget.inc(); + return; + } + let my_addr = *self.node_addr(); let now_ms = Self::now_ms(); let default_ttl = self.config().node.session.default_ttl; @@ -421,12 +493,16 @@ impl Node { now_ms, default_ttl, ); - let RouteAction::SendError { toward, bytes } = match action { + self.record_error_verdict(action.verdict); + let RouteAction::SendError { toward, bytes } = match action.action { Some(action) => action, // Rate limited: drop silently. No cascading errors. None => return, }; + // Both gates have admitted; spend the token peeked above. + self.commit_error_emission(from); + // Resolve the reverse link hop only now, after the gate passed, so // `find_next_hop`'s coord-cache touch keeps its pre-refactor scope. let next_hop_addr = match self.find_next_hop(&toward) { diff --git a/src/node/handlers/handshake.rs b/src/node/handlers/handshake.rs index f3cb8f52..24833634 100644 --- a/src/node/handlers/handshake.rs +++ b/src/node/handlers/handshake.rs @@ -19,9 +19,33 @@ use crate::proto::fmp::{ }; use crate::transport::{Link, LinkDirection, LinkId, ReceivedPacket}; use crate::utils::index::SessionIndex; -use std::time::Duration; +use std::time::{Duration, Instant}; use tracing::{debug, info, warn}; +/// Minimum interval between accepted epoch changes for one peer identity, +/// and the recency threshold at which the peering an epoch change would +/// destroy still counts as live. +/// +/// An epoch-mismatch msg1 is authentic but replayable: a captured one stays +/// valid indefinitely, and accepting it tears down a working peering. Both +/// conditions are receiver-local. The liveness half is the one that closes +/// the replay, since a peering under attack is by construction still +/// heartbeating; the interval half bounds the churn a peer can drive on its +/// own. +/// +/// Sized against the peer's own recovery rather than against a round number: +/// a genuinely restarting peer's msg1 resends fire at roughly t+1, t+3, t+7 +/// and t+15 seconds and its attempt is reaped at `handshake_timeout_secs` +/// (30), so 15 is the largest value at which a real restart still re-peers +/// inside its first handshake window with no reconnect backoff. It also sits +/// below `link_dead_timeout_secs` (30), so the liveness gate can never +/// outlive the reaper that would have removed the peering anyway. +/// +/// Raising it lengthens the outage an attacker's accepted replay causes, +/// because the genuine peer's recovery msg1 hits the same arm. Lowering it +/// weakens both halves and, below the resend ladder, buys nothing. +const EPOCH_RESTART_MIN_INTERVAL_SECS: u64 = 15; + /// Why an inbound msg1 got past the `accept_connections` gate, and against /// what identity the post-DH confirmation must check it. /// @@ -714,6 +738,58 @@ impl Node { // executor's `InvalidateSendState` // (`ambient.verified_identity.node_addr()`) targets the same addr // as the pre-refactor `remove_active_peer(&peer)`. + // The epoch travels inside the AEAD, so this msg1 is + // authentic — but it stays authentic after capture, and + // replaying one destroys a working peering and the FSP session + // state it carries, from off the path. Two receiver-local + // conditions gate the teardown. The peering's last + // authenticated inbound frame is the evidence it is still + // alive, and nothing an unauthenticated sender emits can + // refresh it, so a peer that genuinely restarted clears this by + // having stopped sending. The interval half bounds the churn + // one peer can drive on its own. + let now_ms = Self::now_ms(); + let peering_idle_ms = self + .peers + .get(&peer) + .map(|p| p.idle_time(now_ms)) + .unwrap_or(u64::MAX); + let dampened = self + .restart_dampener + .get(&peer) + .is_some_and(|t| t.elapsed().as_secs() < EPOCH_RESTART_MIN_INTERVAL_SECS); + if peering_idle_ms < EPOCH_RESTART_MIN_INTERVAL_SECS * 1000 || dampened { + debug!( + peer = %self.peer_display_name(&peer), + idle_ms = peering_idle_ms, + dampened, + "Epoch mismatch dampened, dropping msg1" + ); + // Silent drop: the stored msg2 is bound to the original + // msg1's ephemeral, and answering an address the sender + // chose is free amplification. + // + // No registry cleanup is needed here. On the pre-refactor + // layout this arm removed the pending connection and its + // link, because both were inserted before msg1 was + // classified. The classification now runs against a local + // `machine` that enters `peer_machines` only at the promote + // tails below, and `link_id` is a bare allocation until + // then, so dropping out of the arm is the whole cleanup. + // The fresh leg holds no session index either (it is parked + // at `Handshaking{ReceivedMsg1}` with `our_index == None`), + // so nothing is leaked by returning. + self.stats_mut() + .record_reject(RejectReason::Handshake(HandshakeReject::BadState)); + return; + } + // Stamped on acceptance only. A refusal that slid the window + // would let a sustained replay starve a genuinely restarting + // peer for as long as it kept sending. + let cutoff = Duration::from_secs(EPOCH_RESTART_MIN_INTERVAL_SECS); + self.restart_dampener.retain(|_, t| t.elapsed() < cutoff); + self.restart_dampener.insert(peer, Instant::now()); + debug!( peer = %self.peer_display_name(&peer), "Peer restart detected (epoch mismatch), removing stale session" diff --git a/src/node/handlers/lookup.rs b/src/node/handlers/lookup.rs index 2cf81ae8..665e64ad 100644 --- a/src/node/handlers/lookup.rs +++ b/src/node/handlers/lookup.rs @@ -104,7 +104,8 @@ impl Node { let recent_expiry_ms = self.config().node.lookup.recent_expiry_secs * 1000; let my_addr = *self.node_addr(); use crate::proto::lookup::RequestOutcome; - match crate::proto::lookup::classify_request( + let peer_count = self.peers.len(); + let classification = crate::proto::lookup::classify_request( &mut self.lookup, &request, from, @@ -112,7 +113,22 @@ impl Node { now_ms, recent_expiry_ms, MAX_RECENT_LOOKUP_REQUESTS, - ) { + peer_count, + ); + // A full cache evicts rather than refuses, and the core charges the + // eviction to the peer that filled the cache. Count and log it here: + // the core does no metrics and no logging of its own. + if let Some(evicted) = classification.evicted { + self.metrics().lookup.req_dedup_evicted.inc(); + debug!( + request_id = evicted.request_id, + evicted_from = %self.peer_display_name(&evicted.peer), + admitting = %self.peer_display_name(from), + share = evicted.share, + "Lookup dedup cache full, evicting the oldest entry to make room" + ); + } + match classification.outcome { RequestOutcome::Duplicate => { self.metrics() .lookup @@ -123,19 +139,25 @@ impl Node { "Duplicate LookupRequest, dropping" ); } - RequestOutcome::DedupCacheFull { len } => { - self.metrics() - .lookup - .record_reject(DiscoveryReject::ReqDedupCacheFull); - debug!( - request_id = request.request_id, - from = %self.peer_display_name(from), - recent_requests = len, - max_recent_requests = MAX_RECENT_LOOKUP_REQUESTS, - "Discovery request dedup cache full, dropping LookupRequest" - ); - } RequestOutcome::RespondAsTarget => { + // Answering costs a fresh Schnorr signature every time: the + // proof is bound to the requester's request_id, so it cannot + // be cached or served twice. Meter that per link peer, or a + // neighbour generating request_ids sets this node's signing + // rate. The dedup entry the core recorded stays regardless, + // so a refused request still occupies its id and a retry, + // which carries a fresh id, is unaffected. + if !self.discovery_sign_limiter.should_sign(from) { + self.metrics() + .lookup + .record_reject(DiscoveryReject::ReqSignRateLimited); + debug!( + request_id = request.request_id, + from = %self.peer_display_name(from), + "Lookup signing budget spent for this peer, not answering" + ); + return; + } self.metrics().lookup.req_target_is_us.inc(); debug!( request_id = request.request_id, @@ -197,7 +219,11 @@ impl Node { let now_ms = Self::now_ms(); // Check if we forwarded this request (transit node) or originated it - match crate::proto::lookup::classify_response(&mut self.lookup, response.request_id) { + match crate::proto::lookup::classify_response( + &mut self.lookup, + response.request_id, + &response.target, + ) { crate::proto::lookup::ResponseRoute::AlreadyForwarded => { // Already forwarded a response for this request — drop to // prevent response routing loops. @@ -231,6 +257,25 @@ impl Node { ); } } + crate::proto::lookup::ResponseRoute::Unsolicited => { + // Nothing outstanding matches this, so acting on it would let + // one harvested signed response be replayed at will: each + // injection cleared the pending lookup, recorded a + // reachability success, refreshed the cached coordinates for a + // further full TTL, and flushed queued packets onto a route at + // a moment the sender chose. Dropped here, before the identity + // resolve and before the verify, so an unsolicited response + // costs nothing. This counter has a nonzero floor in healthy + // operation: a request is flooded to every qualifying tree + // peer, so duplicate replies land here once the first has been + // accepted. + self.metrics().lookup.resp_unsolicited.inc(); + debug!( + request_id = response.request_id, + target = %self.peer_display_name(&response.target), + "LookupResponse does not match an outstanding request, dropping" + ); + } crate::proto::lookup::ResponseRoute::Originator => { // We originated this request — verify proof before caching let target = response.target; @@ -587,6 +632,15 @@ impl Node { }; let request = LookupRequest::new(request_id, *target, origin, origin_coords, ttl, 0); + // Recorded here rather than in the callers, so "if a request went out, + // its id is recorded" holds for every caller. The response path + // correlates against this set. + self.lookup + .pending_lookups + .entry(*target) + .or_insert_with(|| crate::proto::lookup::PendingLookup::new(Self::now_ms())) + .record(request_id); + // Tree-peer bloom-match selection + single encode live in the sans-IO // core. The core keeps the tree-only (no non-tree fallback) behavior; // the shell drives the sends and keeps all metrics/logging. @@ -740,6 +794,21 @@ impl Node { } } + /// Remove expired entries from the recent-request dedup cache. + /// + /// The ordinary request path purges lazily inside `classify_request`; + /// this is the explicit entry point for callers that need the purge + /// without an arriving request. Cache and per-peer index are purged + /// together, or the eviction policy reads a stale index. + /// + /// Only the dedup regression tests call it: the production path's purge + /// happens inside `classify_request`. + #[cfg(test)] + pub(in crate::node) fn purge_expired_requests(&mut self, current_time_ms: u64) { + let expiry_ms = self.config().node.lookup.recent_expiry_secs * 1000; + self.lookup.purge_recent(current_time_ms, expiry_ms); + } + /// Min-fold our outgoing-link MTU into a LookupResponse's `path_mtu`. /// /// Used at both transit-side reverse-path forward and at the target's diff --git a/src/node/handlers/mod.rs b/src/node/handlers/mod.rs index c0e9b64e..dea42f47 100644 --- a/src/node/handlers/mod.rs +++ b/src/node/handlers/mod.rs @@ -6,6 +6,9 @@ mod mmp; mod native; pub(in crate::node) use native::PendingNative; pub(crate) mod probe; -mod rekey; +// Widened from private by the rekey drain cap: `node::session` calls +// `rekey::drain_max_retention_ms` to bound how long a superseded epoch is +// retained. `rx_loop` is not declared here; master moved it out of `handlers`. +pub(in crate::node) mod rekey; pub(in crate::node) mod session; mod timeout; diff --git a/src/node/handlers/rekey.rs b/src/node/handlers/rekey.rs index ac1246ac..a9371e2f 100644 --- a/src/node/handlers/rekey.rs +++ b/src/node/handlers/rekey.rs @@ -29,6 +29,48 @@ const DRAIN_WINDOW_SECS: u64 = 10; /// a peer's rekey msg1. FMP-scoped copy for `check_rekey`. const REKEY_DAMPENING_SECS: u64 = 30; +/// Floor on the absolute ceiling for `previous`-slot retention after a +/// cutover, in seconds. +/// +/// The drain deadline is peer-progress-aware: it slides forward on every +/// inbound frame that authenticates against the old epoch, so a peer that +/// keeps sealing in that epoch holds the retired key for as long as it +/// likes. This bounds that. It has to stay longer than the worst-case +/// recovery of a legitimate peer that lost msg3, which at stock defaults +/// is the msg3 resend ladder (about 31 s) plus `handshake_timeout_secs` +/// (30 s) before the responder abandons plus `REKEY_DAMPENING_SECS` +/// (30 s) before it may re-initiate, so about 90 s. 120 s clears that +/// with margin and still bounds retention to roughly one +/// `node.rekey.after_secs` period. `drain_max_retention_ms` takes the +/// larger of this floor and the budget the running configuration +/// actually implies, so a shortened handshake timer cannot push the +/// ceiling under the recovery it has to clear. +/// +/// Lowering it below that budget cuts off legitimate slow peers: their +/// frames go silently undecryptable until their own rekey retry +/// re-converges the epochs, because nothing tears an established session +/// down on repeated decrypt failure. Raising it lengthens the window in +/// which a retired key stays resident. +const DRAIN_MAX_RETENTION_SECS: u64 = crate::proto::fsp::limits::DRAIN_WINDOW_SECS * 12; + +/// Effective ceiling on total `previous`-slot retention, in milliseconds. +/// +/// The larger of `DRAIN_MAX_RETENTION_SECS` and the msg3 recovery budget +/// the configured handshake timers imply, so the ceiling always clears +/// the recovery it is supposed to leave room for. +pub(in crate::node) fn drain_max_retention_ms(rate_limit: &crate::config::RateLimitConfig) -> u64 { + let mut ladder_ms: u64 = 0; + let mut interval = rate_limit.handshake_resend_interval_ms as f64; + for _ in 0..rate_limit.handshake_max_resends { + ladder_ms = ladder_ms.saturating_add(interval as u64); + interval *= rate_limit.handshake_resend_backoff; + } + let recovery_budget_ms = ladder_ms + .saturating_add(rate_limit.handshake_timeout_secs.saturating_mul(1000)) + .saturating_add(crate::proto::fsp::limits::REKEY_DAMPENING_SECS * 1000); + (DRAIN_MAX_RETENTION_SECS * 1000).max(recovery_budget_ms) +} + impl Node { /// Periodic rekey check. Called from the tick loop. /// @@ -640,6 +682,13 @@ impl Node { // is the only stamp that path writes. A *completed* rekey has no such // bound and must not acquire one: see `FspAction::AbandonHandshake`. let stale_handshake_ms = self.config().node.rate_limit.handshake_timeout_secs * 1000; + // Absolute ceiling on `previous`-slot retention, measured from the + // cutover. The sliding drain deadline is peer-progress-aware, so an + // authenticated peer that keeps sealing in the old epoch can hold the + // retired key indefinitely; this bounds that without shortening the + // grace a peer that lost msg3 legitimately needs. Resolved here rather + // than in the core, which reads no clock and no configuration. + let drain_max_ms = drain_max_retention_ms(&self.config().node.rate_limit); self.sessions .iter() .filter(|(_, entry)| entry.is_established()) @@ -650,7 +699,7 @@ impl Node { is_rekey_initiator: entry.is_rekey_initiator(), cutover_timer_elapsed: cutover_timer_elapsed(now_ms, entry.rekey_completed_ms()), is_draining: entry.is_draining(), - drain_expired: entry.drain_expired(now_ms, drain_ms), + drain_expired: entry.drain_expired(now_ms, drain_ms, drain_max_ms), has_rekey_msg3_payload: entry.rekey_msg3_payload().is_some(), is_dampened: entry.is_rekey_dampened(now_ms, dampening_ms), armed_handshake_expired: entry.last_peer_rekey_ms() != 0 diff --git a/src/node/handlers/session.rs b/src/node/handlers/session.rs index 58bc204f..4d03749e 100644 --- a/src/node/handlers/session.rs +++ b/src/node/handlers/session.rs @@ -46,6 +46,49 @@ use crate::upper::icmp::FIPS_OVERHEAD; use secp256k1::PublicKey; use tracing::{debug, info, trace, warn}; +/// Minimum interval between path-MTU releases driven by `PathBroken` for one +/// destination. +/// +/// `PathBroken` is unauthenticated, so a release is a remote party's claim +/// that the path a tightened MTU described is gone. Without an interval the +/// claim can be repeated at line rate, discarding a genuinely learned +/// bottleneck as fast as it is relearned. Raising it defers a legitimate +/// release after a second real break, which costs throughput on the new path +/// but never a blackhole, since the deferred value is the tighter one. +pub(in crate::node) const PATH_MTU_RELEASE_MIN_INTERVAL: std::time::Duration = + std::time::Duration::from_millis(1000); + +/// Bytes the link layer adds to an encoded `SessionDatagram` on its way to the +/// wire: the established FMP header, the 4-byte session-relative timestamp and +/// the AEAD tag. Mirrors the buffer `send_encrypted_link_message_with_ce` +/// builds. +/// +/// Spelled out in full rather than through the `crate::proto::fmp::wire` +/// import above, which is `#[cfg(unix)]`. This constant feeds `link_wire_len`, +/// whose caller `send_session_datagram` is compiled on every platform, so +/// taking the name from that import fails to build on Windows. +const LINK_FRAME_OVERHEAD: usize = + crate::proto::fmp::wire::ESTABLISHED_HEADER_SIZE + 4 + crate::noise::TAG_SIZE; + +/// Wire size of an encoded `SessionDatagram` of `encoded_len` bytes. +fn link_wire_len(encoded_len: usize) -> usize { + encoded_len + LINK_FRAME_OVERHEAD +} + +/// Divisor giving the share of the session table that unauthenticated +/// half-open entries may hold, as `max_sessions / DIVISOR`. +/// +/// Two means a reconnect storm, where every peer that had a session +/// initiates at once after a restart or a healed partition, still fits in +/// half the table; a tighter share bites four times sooner and is felt by +/// a hub before it is felt by an attacker. Half-open entries are reaped +/// after `handshake_timeout_secs` while established ones survive +/// `idle_timeout_secs`, so they turn over faster than the share suggests. +/// Lowering the divisor raises the share, which lets a handshake flood +/// crowd out peers that complete; raising it refuses legitimate initiators +/// sooner in a storm. +const HALF_OPEN_SHARE_DIVISOR: usize = 2; + /// Inputs to `try_send_session_data_pipelined` — the FSP+FMP pipelined /// fast path that hands both AEAD operations to the encrypt worker /// in a single dispatch. @@ -547,6 +590,19 @@ impl Node { // no limit at all. A setup naming an established peer cannot grow the // table and is metered separately, so that a stranger flood over a // shared link cannot stop that peer's rekey from arming. + // Population cap, ahead of the limiter so a full table costs no + // token, no responder handshake and no ack. The predicate is "would + // admitting this grow the table", not "is this a stranger": `class` + // is Stranger for an existing Initiating or AwaitingMsg3 entry too, + // and refusing those would break in-flight legitimate handshakes and + // the duplicate-ack resend. Same shape as the pending-destination cap + // in `queue_pending_packet`. Refuse rather than evict: msg1 is + // unauthenticated here, so evicting would hand a stranger a teardown + // primitive it does not have. + if !self.admit_new_session(src_addr) { + return; + } + let class = if self .sessions .get(src_addr) @@ -1876,8 +1932,20 @@ impl Node { } } // The path this destination's stored MTU described is gone, so release - // it rather than carrying it onto whatever path replaces it. - self.path_mtu_lookup_release(&msg.dest_addr); + // it rather than carrying it onto whatever path replaces it. Rate + // limited per destination on its own budget: PathBroken is + // unauthenticated, and an unlimited release discards a genuinely + // learned bottleneck as fast as it is relearned. The budget is not + // shared with any other signal, so nothing else can spend it. + if self + .path_mtu_release_limiter + .should_send(&msg.dest_addr, Self::now_ms()) + { + self.path_mtu_lookup_release(&msg.dest_addr); + } else { + trace!(dest = %msg.dest_addr, + "PathBroken path MTU release rate-limited, keeping the stored value"); + } if !has_cached_identity { debug!(dest = %msg.dest_addr, @@ -1946,6 +2014,55 @@ impl Node { "MtuExceeded: transit router reports oversized packet" ); + // Both effects below — the session's own path MTU and the + // FipsAddress-keyed lookup the TUN MSS clamp reads — are refused from + // here, so one return covers both. The guards sit ahead of the apply + // rather than between the two effects, which is what makes the floor + // govern `current_mtu` and not only the lookup table. + + // Refuse a bottleneck too small to describe a usable path; a stored + // value that low drives the SYN-time MSS clamp into single digits or + // zero. The reactive carrier is unauthenticated, so it has its own + // floor constant, currently equal to the actionable one. + if msg.mtu < crate::upper::icmp::MIN_REACTIVE_PATH_MTU { + warn!( + dest = %peer_name, + reporter = %msg.reporter, + bottleneck_mtu = msg.mtu, + floor = crate::upper::icmp::MIN_REACTIVE_PATH_MTU, + "MtuExceeded reports a path MTU below the actionable floor; ignoring" + ); + self.metrics().errors.mtu_exceeded_below_floor.inc(); + return; + } + + // Corroboration. The admission gate narrows which destination may be + // named; it cannot authenticate the reporter, so a legal value is a + // legal value from anyone and the floor alone only sets the outcome of + // a forgery rather than preventing it. An honest report exists only + // because a frame this node emitted did not fit some hop, so require + // that this node has actually sent something larger than the value + // being claimed since the last accepted decrease. Honest path-MTU + // discovery satisfies this by construction; a forgery has to wait for + // us to emit a frame bigger than the value it wants to claim, which + // bounds every accepted claim from below by our own traffic. + let sent_wire_len = self + .sessions + .get(&msg.dest_addr) + .map(|e| e.max_sent_wire_len()) + .unwrap_or(0); + if msg.mtu >= sent_wire_len { + debug!( + dest = %peer_name, + reporter = %msg.reporter, + bottleneck_mtu = msg.mtu, + max_sent_wire_len = sent_wire_len, + "MtuExceeded reports a bottleneck no smaller than anything this node has sent; ignoring" + ); + self.metrics().errors.mtu_exceeded_uncorroborated.inc(); + return; + } + // Apply to PathMtuState: immediate decrease via apply_notification() if let Some(entry) = self.sessions.get_mut(&msg.dest_addr) && let Some(mmp) = entry.mmp_mut() @@ -1966,21 +2083,12 @@ impl Node { } } - // The admission gate above restricts which addresses may be written, - // not which values. Any node at any distance may legitimately report a - // bottleneck for a destination this node has bound, so refuse to store - // one too small to describe a usable path; a stored value that low - // drives the SYN-time MSS clamp into single digits or zero. - if msg.mtu < crate::upper::icmp::MIN_ACTIONABLE_PATH_MTU { - warn!( - dest = %peer_name, - reporter = %msg.reporter, - bottleneck_mtu = msg.mtu, - floor = crate::upper::icmp::MIN_ACTIONABLE_PATH_MTU, - "MtuExceeded reports a path MTU below the actionable floor; ignoring" - ); - self.metrics().errors.mtu_exceeded_below_floor.inc(); - return; + // Spent: the evidence vouched for this decrease and does not vouch for + // the next one. An initiating session has no `mmp` and so reaches this + // with the apply above skipped; the reset belongs to the acceptance, + // not to the apply. + if let Some(entry) = self.sessions.get_mut(&msg.dest_addr) { + entry.clear_sent_wire_len(); } // Mirror the bottleneck into the FipsAddress-keyed lookup used by @@ -2049,6 +2157,59 @@ impl Node { /// Creates a Noise XK handshake as initiator, wraps msg1 in a /// SessionSetup, encapsulates in a SessionDatagram, and routes /// toward the destination. + /// Whether a session for `addr` may be created, given the table cap. + /// + /// Returns true when an entry already exists, since admitting it cannot + /// grow the table. Counts its own refusals, so the two reasons are + /// distinguishable without turning on debug logging. + pub(in crate::node) fn admit_new_session(&mut self, addr: &NodeAddr) -> bool { + let max_sessions = self.config().node.limits.max_sessions; + if max_sessions == 0 || self.sessions.contains_key(addr) { + return true; + } + + if self.sessions.len() >= max_sessions { + debug!( + src = %self.peer_display_name(addr), + sessions = self.sessions.len(), + max_sessions = max_sessions, + "Session table full, refusing to create a session" + ); + self.stats_mut() + .record_reject(RejectReason::Session(SessionReject::TableFull)); + return false; + } + + // Half-open entries are unauthenticated and are reaped after + // `handshake_timeout_secs`, so they are the cheap half of the table + // to fill. Holding them to a share keeps room for peers that + // complete. The outer length test makes the scan unreachable below + // the share, and the table is itself bounded by the cap above. + // At least one, or a table capped at one would admit no inbound + // session at all rather than one. + let half_open_share = (max_sessions / HALF_OPEN_SHARE_DIVISOR).max(1); + if self.sessions.len() >= half_open_share { + let half_open = self + .sessions + .values() + .filter(|e| e.is_awaiting_msg3()) + .count(); + if half_open >= half_open_share { + debug!( + src = %self.peer_display_name(addr), + half_open = half_open, + half_open_share = half_open_share, + "Half-open session share exhausted, refusing to create a session" + ); + self.stats_mut() + .record_reject(RejectReason::Session(SessionReject::HalfOpenFull)); + return false; + } + } + + true + } + pub(in crate::node) async fn initiate_session( &mut self, dest_addr: NodeAddr, @@ -2516,6 +2677,7 @@ impl Node { if let Some(entry) = self.sessions.get_mut(dest_addr) { entry.record_sent(send.payload.len()); + entry.record_sent_wire_len(wire_capacity); if let Some(mmp) = entry.mmp_mut() { mmp.sender.record_sent( fsp_counter, @@ -2794,6 +2956,13 @@ impl Node { self.send_encrypted_link_message(&next_hop_addr, &encoded) .await?; self.metrics().forwarding.record_originated(encoded.len()); + + // Evidence for the reactive path-MTU carrier. A transit hop + // re-encapsulates what it forwards, so the frame that overflows a + // downstream link is the size this frame is here. + if let Some(entry) = self.sessions.get_mut(&datagram.dest_addr) { + entry.record_sent_wire_len(link_wire_len(encoded.len())); + } Ok(()) } @@ -2888,6 +3057,17 @@ impl Node { return; } + // No session, so this one would grow the table. Answer the local + // application the way an unroutable destination is answered rather + // than returning an error from `initiate_session`: the caller reads + // an error as "no route" and responds with a discovery lookup and a + // queued packet, which is outbound traffic on a node already at its + // limit. + if !self.admit_new_session(&dest_addr) { + self.send_icmpv6_dest_unreachable(&ipv6_packet); + return; + } + // No session: initiate one and queue the packet. // If session initiation fails (no route), trigger discovery and // queue the packet for retry when discovery completes. @@ -3015,6 +3195,10 @@ impl Node { return; } + if !self.admit_new_session(&dest_addr) { + return; + } + match self.initiate_session(dest_addr, dest_pubkey).await { Ok(()) => { debug!(dest = %self.peer_display_name(&dest_addr), "Session initiated after discovery"); diff --git a/src/node/lifecycle/mod.rs b/src/node/lifecycle/mod.rs index 1588f136..4e5f7c20 100644 --- a/src/node/lifecycle/mod.rs +++ b/src/node/lifecycle/mod.rs @@ -1797,13 +1797,21 @@ impl Node { ); // Resolve the TUN ifindex so the responder can // drop queries arriving on the mesh interface - // (fips0). Without this, the `::` bind exposes - // /etc/fips/hosts alias probing to any mesh peer. - // When TUN isn't enabled or the name can't be - // resolved, `None` disables the filter (there - // is no mesh surface to defend anyway). - let mesh_ifindex = - Self::lookup_mesh_ifindex(self.config().tun.name()); + // Without this, the `::` bind exposes the + // hosts file's alias space to any mesh peer. + // The name comes from the device the TUN + // path actually created, not the configured + // one: macOS and FreeBSD assign utunN/tunN + // of their own choosing and the configured + // name resolves to nothing there, which left + // the filter permanently off. + let mesh_ifindex = self.mesh_ifindex(); + if self.tun_name.is_some() && mesh_ifindex.is_none() { + warn!( + device = ?self.tun_name, + "Mesh interface index unresolved; DNS mesh filter disabled" + ); + } info!( bind = %local_addr, hosts = reloader.hosts().len(), @@ -1992,12 +2000,21 @@ impl Node { Ok(()) } - /// Resolve the mesh TUN interface index by name. + /// Resolve the index of the mesh TUN device this node actually created. /// - /// Returns `None` if the interface does not exist (e.g. TUN disabled - /// or not yet created). A `None` result disables the DNS responder's - /// mesh-interface filter — safe, because if there is no fips0 there - /// is no mesh exposure to defend against. + /// Reads the device name recorded when the TUN was brought up, which is + /// the kernel's name rather than the configured one. Returns `None` when + /// no TUN is up, which disables the DNS responder's mesh-interface + /// filter: with no mesh interface there is no mesh exposure to defend. + /// An app-owned TUN also leaves the name unset, so the filter stays off + /// there even though a mesh interface exists. + pub(crate) fn mesh_ifindex(&self) -> Option { + self.tun_name.as_deref().and_then(Self::lookup_mesh_ifindex) + } + + /// Resolve an interface index by name. + /// + /// Returns `None` if the interface does not exist. fn lookup_mesh_ifindex(name: &str) -> Option { #[cfg(unix)] { diff --git a/src/node/metrics.rs b/src/node/metrics.rs index 30180f9a..231780c8 100644 --- a/src/node/metrics.rs +++ b/src/node/metrics.rs @@ -233,6 +233,8 @@ pub struct LookupMetrics { pub req_decode_error: Counter, pub req_duplicate: Counter, pub req_dedup_cache_full: Counter, + pub req_dedup_evicted: Counter, + pub req_sign_rate_limited: Counter, pub req_target_is_us: Counter, pub req_forwarded: Counter, pub req_ttl_exhausted: Counter, @@ -248,6 +250,7 @@ pub struct LookupMetrics { pub resp_forwarded: Counter, pub resp_identity_miss: Counter, pub resp_proof_failed: Counter, + pub resp_unsolicited: Counter, pub resp_no_route: Counter, pub resp_accepted: Counter, pub resp_timed_out: Counter, @@ -262,10 +265,12 @@ impl LookupMetrics { DiscoveryReject::ReqDecodeError => self.req_decode_error.inc(), DiscoveryReject::ReqDuplicate => self.req_duplicate.inc(), DiscoveryReject::ReqDedupCacheFull => self.req_dedup_cache_full.inc(), + DiscoveryReject::ReqSignRateLimited => self.req_sign_rate_limited.inc(), DiscoveryReject::ReqTtlExhausted => self.req_ttl_exhausted.inc(), DiscoveryReject::RespDecodeError => self.resp_decode_error.inc(), DiscoveryReject::RespIdentityMiss => self.resp_identity_miss.inc(), DiscoveryReject::RespProofFailed => self.resp_proof_failed.inc(), + DiscoveryReject::RespUnsolicited => self.resp_unsolicited.inc(), DiscoveryReject::RespNoRoute => self.resp_no_route.inc(), } } @@ -277,6 +282,8 @@ impl LookupMetrics { req_decode_error: self.req_decode_error.get(), req_duplicate: self.req_duplicate.get(), req_dedup_cache_full: self.req_dedup_cache_full.get(), + req_dedup_evicted: self.req_dedup_evicted.get(), + req_sign_rate_limited: self.req_sign_rate_limited.get(), req_target_is_us: self.req_target_is_us.get(), req_forwarded: self.req_forwarded.get(), req_ttl_exhausted: self.req_ttl_exhausted.get(), @@ -292,6 +299,7 @@ impl LookupMetrics { resp_forwarded: self.resp_forwarded.get(), resp_identity_miss: self.resp_identity_miss.get(), resp_proof_failed: self.resp_proof_failed.get(), + resp_unsolicited: self.resp_unsolicited.get(), resp_no_route: self.resp_no_route.get(), resp_accepted: self.resp_accepted.get(), resp_timed_out: self.resp_timed_out.get(), @@ -474,7 +482,25 @@ pub struct ErrorMetrics { /// count means a forwarder on the reverse path is mangling the unsigned /// annotation. pub lookup_resp_mtu_below_floor: Counter, + /// `MtuExceeded` signals ignored because this node has not sent a frame + /// larger than the bottleneck they report since the last accepted + /// decrease. An honest report cannot arise without such a frame, so a + /// rising count is a forged or stale reactive signal. + pub mtu_exceeded_uncorroborated: Counter, pub unbound: UnboundSignals, + /// Routing errors this node declined to emit because the authenticated + /// link peer that induced them had spent its budget. A rising count is + /// either a peer flooding unroutable traffic or a hub relaying more + /// simultaneously-broken destinations than the budget allows. + pub emit_over_peer_budget: Counter, + /// Routing errors this node declined to emit because one for the same + /// destination went out within the per-destination interval. This is the + /// aggregate suppression a real outage produces. + pub emit_over_dest_interval: Counter, + /// Routing errors emitted without recording their destination, because + /// the per-destination limiter's map was full. The signal was still sent; + /// what was lost is interval suppression for that destination. + pub emit_limiter_at_capacity: Counter, } impl ErrorMetrics { @@ -491,6 +517,10 @@ impl ErrorMetrics { unbound_broken: self.unbound.broken.get(), unbound_mtu: self.unbound.mtu.get(), unbound_forged: self.unbound.forged.get(), + emit_over_peer_budget: self.emit_over_peer_budget.get(), + emit_over_dest_interval: self.emit_over_dest_interval.get(), + emit_limiter_at_capacity: self.emit_limiter_at_capacity.get(), + mtu_exceeded_uncorroborated: self.mtu_exceeded_uncorroborated.get(), } } } diff --git a/src/node/mod.rs b/src/node/mod.rs index 851ae047..d1ecbd8d 100644 --- a/src/node/mod.rs +++ b/src/node/mod.rs @@ -17,6 +17,7 @@ pub(crate) mod encrypt_worker; mod handlers; mod lifecycle; pub(crate) mod metrics; +mod peer_error_budget; mod peering; mod rate_limit; pub(crate) mod reject; @@ -30,7 +31,8 @@ pub(crate) mod stats_history; mod tests; mod tree; -use self::rate_limit::{HandshakeRateLimiter, SessionSetupRateLimiter}; +use self::peer_error_budget::PeerErrorBudget; +use self::rate_limit::{HandshakeRateLimiter, LookupSignRateLimiter, SessionSetupRateLimiter}; use self::reloadable::Reloadable; /// Half-range of the symmetric jitter applied to the per-session rekey timer. @@ -473,6 +475,11 @@ pub struct Node { /// Discovery-subsystem state: recent-request dedup cache, in-flight /// lookups, originator-side backoff, and transit-side forward limiter. lookup: Lookup, + /// Signing budget for lookups we answer about ourselves (target-side), + /// keyed on the link peer the request arrived over. Held here rather than + /// inside `lookup` because it is an `Instant`-based limiter and the + /// `proto` tree is clockless. + discovery_sign_limiter: LookupSignRateLimiter, // === Diagnostics === /// In-flight `probe` jobs plus their per-target ownership claims. Driven @@ -538,6 +545,12 @@ pub struct Node { /// Pending outbound handshakes by our sender_idx. /// Tracks which LinkId corresponds to which session index. pending_outbound: HashMap<(TransportId, u32), LinkId>, + /// When each peer identity's last ACCEPTED epoch change tore down its + /// peering. Keyed on identity rather than address, and held here rather + /// than on `ActivePeer`, because the teardown being dampened destroys + /// the peer entry itself. Pruned on insert; see + /// `EPOCH_RESTART_MIN_INTERVAL_SECS`. + restart_dampener: HashMap, // === Rate Limiting === /// Rate limiter for msg1 processing (DoS protection). @@ -546,6 +559,12 @@ pub struct Node { setup_rate_limiter: SessionSetupRateLimiter, /// Rate limiter for ICMP Packet Too Big messages. icmp_rate_limiter: IcmpRateLimiter, + /// Budget bounding the routing errors one authenticated link peer can + /// induce this node to emit. Keyed on the link peer because that is the + /// only value at the emission point a sender cannot mint; the + /// per-destination interval inside `routing` is keyed on a field the + /// sender chooses and is an aggregate suppressor, not a bound. + peer_error_budget: PeerErrorBudget, /// Routing-subsystem state (routing error-signal rate limiter). routing: Router, /// FMP connection-lifecycle decision anchor (stateless; drives the @@ -559,6 +578,12 @@ pub struct Node { mmp: Mmp, /// Rate limiter for source-side CoordsRequired/PathBroken responses. coords_response_rate_limiter: RoutingErrorRateLimiter, + /// Rate limiter for PathBroken-driven path-MTU releases, per destination. + /// Deliberately its own instance rather than a share of + /// `coords_response_rate_limiter`: a budget another signal can spend is + /// not a bound on this one, and one PathBroken drives both responses, so + /// a shared limiter would let the coord-warmup arm pay for the release. + path_mtu_release_limiter: RoutingErrorRateLimiter, // === Peering Homeostasis === /// Owner of the peering-reconciler state relocated off `Node`: the sans-IO @@ -806,9 +831,11 @@ impl Node { index_allocator: IndexAllocator::new(), peers_by_index: HashMap::new(), pending_outbound: HashMap::new(), + restart_dampener: HashMap::new(), msg1_rate_limiter, setup_rate_limiter, icmp_rate_limiter: IcmpRateLimiter::new(), + peer_error_budget: PeerErrorBudget::new(), routing: Router::new(), fmp: Fmp::new(), fsp: Fsp::new(), @@ -816,11 +843,15 @@ impl Node { coords_response_rate_limiter: RoutingErrorRateLimiter::with_interval_ms( coords_response_interval_ms, ), + path_mtu_release_limiter: RoutingErrorRateLimiter::with_interval_ms( + handlers::session::PATH_MTU_RELEASE_MIN_INTERVAL.as_millis() as u64, + ), probes: handlers::probe::ProbeRegistry::new(), lookup: Lookup::new( LookupBackoff::with_params(backoff_base_secs, backoff_max_secs), LookupForwardRateLimiter::with_interval_ms(forward_min_interval_secs * 1000), ), + discovery_sign_limiter: LookupSignRateLimiter::new(), peering: peering::reconcile::Peering::new(), last_parent_reeval: None, last_congestion_log: None, @@ -963,9 +994,11 @@ impl Node { index_allocator: IndexAllocator::new(), peers_by_index: HashMap::new(), pending_outbound: HashMap::new(), + restart_dampener: HashMap::new(), msg1_rate_limiter, setup_rate_limiter, icmp_rate_limiter: IcmpRateLimiter::new(), + peer_error_budget: PeerErrorBudget::new(), routing: Router::new(), fmp: Fmp::new(), fsp: Fsp::new(), @@ -973,8 +1006,12 @@ impl Node { coords_response_rate_limiter: RoutingErrorRateLimiter::with_interval_ms( coords_response_interval_ms, ), + path_mtu_release_limiter: RoutingErrorRateLimiter::with_interval_ms( + handlers::session::PATH_MTU_RELEASE_MIN_INTERVAL.as_millis() as u64, + ), probes: handlers::probe::ProbeRegistry::new(), lookup: Lookup::new(LookupBackoff::new(), LookupForwardRateLimiter::new()), + discovery_sign_limiter: LookupSignRateLimiter::new(), peering: peering::reconcile::Peering::new(), last_parent_reeval: None, last_congestion_log: None, @@ -1761,6 +1798,17 @@ impl Node { /// the way the rx_loop's handler does without standing up a client task and /// a socket pair. A parity test over an empty registry proves nothing about /// a publisher that drops fields, which is why this exists. + /// The instant at which `peer`'s last ACCEPTED epoch change was stamped, + /// or `None` if it has none. Test-only, and it exists for one assertion: + /// that a REFUSED epoch-mismatch msg1 leaves this untouched. The refusal + /// must not slide the window, or a sustained replay starves a genuinely + /// restarting peer for as long as it keeps sending — which is the whole + /// point of stamping on acceptance rather than on every sighting. + #[cfg(test)] + pub(crate) fn restart_dampener_stamp(&self, peer: &NodeAddr) -> Option { + self.restart_dampener.get(peer).copied() + } + #[cfg(test)] pub(crate) fn native_registry_for_test(&mut self) -> &mut crate::native::registry::Registry { &mut self.native @@ -2754,6 +2802,12 @@ impl Node { // === End-to-End Sessions === /// Get a session by remote NodeAddr. + /// Set the per-link-peer lookup signing budget (for tests). + #[cfg(test)] + pub(crate) fn set_discovery_sign_budget(&mut self, burst: f64, rate: f64) { + self.discovery_sign_limiter = LookupSignRateLimiter::with_params(burst, rate); + } + /// Disable the discovery forward rate limiter (for tests). #[cfg(test)] pub(crate) fn disable_discovery_forward_rate_limit(&mut self) { @@ -2846,6 +2900,12 @@ impl Node { /// `FipsAddress`-keyed map the TCP MSS clamp reads, and the session's own /// source-side path MTU estimate. fn path_mtu_lookup_release(&mut self, addr: &NodeAddr) { + // The evidence that corroborates a reactive MtuExceeded described the + // path being released, so it does not vouch for whatever replaces it. + if let Some(entry) = self.sessions.get_mut(addr) { + entry.clear_sent_wire_len(); + } + // The session's own source-side estimate described the same dead path, // and the increase ladder is the only thing that would ever raise it // again. Reset it here so the two halves of "this path is gone" stay diff --git a/src/node/peer_error_budget.rs b/src/node/peer_error_budget.rs new file mode 100644 index 00000000..785337bc --- /dev/null +++ b/src/node/peer_error_budget.rs @@ -0,0 +1,239 @@ +//! Per-link-peer budget for induced routing-error emissions. +//! +//! A transit node synthesizes a routing error (CoordsRequired, PathBroken or +//! MtuExceeded) in response to a datagram it could not forward. Every field of +//! that datagram is chosen by whoever sent it, so a per-destination or +//! per-source gate can be escaped by varying the field it is keyed on. The one +//! value at the emission point an attacker cannot mint is the authenticated +//! link peer the frame arrived from, whose cardinality is bounded by the peer +//! table and by admission. This budget is keyed on it, and is consulted ahead +//! of any address-keyed structure so that those structures only grow at the +//! budget rate. + +use crate::NodeAddr; +use std::collections::HashMap; +use std::time::{Duration, Instant}; + +/// Sustained rate, in signals per second, at which one authenticated link peer +/// may induce this node to emit routing errors. +/// +/// Bounds the reflection an admitted peer can aim at a victim it names, and +/// bounds how fast that peer can grow the per-destination limiter's map. +/// Raising it costs proportionally more reflected traffic per peer; lowering it +/// silences a hub peer that relays many sources through a genuine outage +/// sooner, which costs those sources their CoordsRequired and their +/// path-MTU feedback. +pub const PEER_ERROR_RATE_PER_SEC: u32 = 20; + +/// Number of routing errors one link peer may induce back to back before the +/// sustained rate applies. +/// +/// Sized so that an ordinary burst of unroutable traffic behind one peer still +/// signals promptly. Raising it lets a peer front-load a larger reflection; +/// lowering it makes a legitimate convergence burst arrive as a trickle. +pub const PEER_ERROR_BURST: u32 = 50; + +/// Tokens are carried in thousandths so the refill of a sub-millisecond +/// interval is not rounded away. +const MILLI: u64 = 1000; + +/// A peer whose bucket has refilled to full carries no state worth keeping, so +/// entries are dropped once per this interval to bound the map across peer +/// churn. +const SWEEP_INTERVAL: Duration = Duration::from_secs(30); + +/// One peer's token bucket. +struct Bucket { + /// Remaining tokens, in thousandths of a signal. + milli_tokens: u64, + /// When `milli_tokens` was last brought up to date. + last_refill: Instant, +} + +/// Token-bucket budget for routing errors, keyed on the authenticated link +/// peer that induced them. +pub struct PeerErrorBudget { + buckets: HashMap, + /// Refill rate in thousandths of a token per millisecond. + milli_per_ms: u64, + /// Bucket ceiling, in thousandths of a token. + capacity: u64, + last_sweep: Instant, +} + +impl PeerErrorBudget { + /// Create a budget at the shipped rate and burst. + pub fn new() -> Self { + Self::with_rate(PEER_ERROR_RATE_PER_SEC, PEER_ERROR_BURST) + } + + /// Create a budget with an explicit sustained rate and burst. + pub fn with_rate(per_sec: u32, burst: u32) -> Self { + Self { + buckets: HashMap::new(), + milli_per_ms: u64::from(per_sec), + capacity: u64::from(burst) * MILLI, + last_sweep: Instant::now(), + } + } + + /// Whether `peer` has a token to spend, without spending it. + /// + /// Separate from [`Self::commit`] so a signal that a later gate suppresses + /// does not consume budget: an outage behind a high-fanout peer would + /// otherwise spend that peer's whole budget on emissions the + /// per-destination interval discards, silencing every other destination + /// behind it. + pub fn has_token(&mut self, peer: &NodeAddr, now: Instant) -> bool { + self.refill(peer, now); + self.buckets + .get(peer) + .is_some_and(|b| b.milli_tokens >= MILLI) + } + + /// Spend one token for `peer`. Call only on the path that actually emits. + pub fn commit(&mut self, peer: &NodeAddr, now: Instant) { + self.refill(peer, now); + if let Some(bucket) = self.buckets.get_mut(peer) { + bucket.milli_tokens = bucket.milli_tokens.saturating_sub(MILLI); + } + self.sweep(now); + } + + /// Bring `peer`'s bucket up to date, creating a full one on first sighting. + fn refill(&mut self, peer: &NodeAddr, now: Instant) { + let capacity = self.capacity; + let milli_per_ms = self.milli_per_ms; + let bucket = self.buckets.entry(*peer).or_insert(Bucket { + milli_tokens: capacity, + last_refill: now, + }); + let elapsed_ms = now + .saturating_duration_since(bucket.last_refill) + .as_millis() as u64; + if elapsed_ms > 0 { + bucket.milli_tokens = (bucket.milli_tokens + elapsed_ms * milli_per_ms).min(capacity); + bucket.last_refill = now; + } + } + + /// Drop full buckets, at most once per [`SWEEP_INTERVAL`]. + fn sweep(&mut self, now: Instant) { + if now.saturating_duration_since(self.last_sweep) < SWEEP_INTERVAL { + return; + } + self.last_sweep = now; + let capacity = self.capacity; + let milli_per_ms = self.milli_per_ms; + self.buckets.retain(|_, b| { + let elapsed_ms = now.saturating_duration_since(b.last_refill).as_millis() as u64; + b.milli_tokens + elapsed_ms * milli_per_ms < capacity + }); + } + + #[cfg(test)] + pub fn len(&self) -> usize { + self.buckets.len() + } +} + +impl Default for PeerErrorBudget { + fn default() -> Self { + Self::new() + } +} + +#[cfg(test)] +mod tests { + use super::*; + + fn addr(val: u8) -> NodeAddr { + let mut bytes = [0u8; 16]; + bytes[0] = val; + NodeAddr::from_bytes(bytes) + } + + /// Spend one token per admitted emission. + fn spend(budget: &mut PeerErrorBudget, peer: &NodeAddr, now: Instant) -> bool { + if !budget.has_token(peer, now) { + return false; + } + budget.commit(peer, now); + true + } + + #[test] + fn a_peer_may_emit_its_full_burst_then_is_suppressed() { + let mut budget = PeerErrorBudget::new(); + let now = Instant::now(); + let peer = addr(1); + + for i in 0..PEER_ERROR_BURST { + assert!(spend(&mut budget, &peer, now), "burst signal {i} refused"); + } + assert!(!spend(&mut budget, &peer, now)); + } + + #[test] + fn an_exhausted_budget_refills_at_the_sustained_rate() { + let mut budget = PeerErrorBudget::new(); + let start = Instant::now(); + let peer = addr(1); + + for _ in 0..PEER_ERROR_BURST { + assert!(spend(&mut budget, &peer, start)); + } + assert!(!spend(&mut budget, &peer, start)); + + // One second of refill buys exactly the sustained rate back. + let later = start + Duration::from_secs(1); + for i in 0..PEER_ERROR_RATE_PER_SEC { + assert!( + spend(&mut budget, &peer, later), + "refilled signal {i} refused" + ); + } + assert!(!spend(&mut budget, &peer, later)); + } + + #[test] + fn one_peer_exhausting_its_budget_does_not_silence_another() { + let mut budget = PeerErrorBudget::new(); + let now = Instant::now(); + let noisy = addr(1); + let quiet = addr(2); + + for _ in 0..PEER_ERROR_BURST { + assert!(spend(&mut budget, &noisy, now)); + } + assert!(!spend(&mut budget, &noisy, now)); + assert!(spend(&mut budget, &quiet, now)); + } + + #[test] + fn peeking_does_not_spend_a_token() { + let mut budget = PeerErrorBudget::with_rate(1, 1); + let now = Instant::now(); + let peer = addr(1); + + assert!(budget.has_token(&peer, now)); + assert!(budget.has_token(&peer, now)); + budget.commit(&peer, now); + assert!(!budget.has_token(&peer, now)); + } + + #[test] + fn full_buckets_are_dropped_by_the_sweep() { + let mut budget = PeerErrorBudget::new(); + let start = Instant::now(); + for i in 0..50u8 { + assert!(spend(&mut budget, &addr(i), start)); + } + assert_eq!(budget.len(), 50); + + // Long enough for every bucket to have refilled to full. + let later = start + SWEEP_INTERVAL + Duration::from_secs(1); + assert!(spend(&mut budget, &addr(200), later)); + assert_eq!(budget.len(), 1); + } +} diff --git a/src/node/rate_limit.rs b/src/node/rate_limit.rs index 0e577b7e..b456c002 100644 --- a/src/node/rate_limit.rs +++ b/src/node/rate_limit.rs @@ -874,3 +874,164 @@ mod tests { ); } } + +// ============================================================================ +// Target-side: Lookup Signing Budget +// +// Homed here rather than in `proto::lookup::limits` because it is an +// `Instant`-based shell limiter, and the `proto` tree is `alloc`-based and +// clockless. On `maint` it lived in `src/node/discovery_rate_limit.rs`, which +// `master` dissolved into `proto/lookup/limits.rs`. +// ============================================================================ + +/// Signatures one link peer may buy in a burst before the refill paces it. +/// +/// Sized for the case that actually produces a burst: a topology change +/// flushes correspondents' coordinate caches and they all look this node up +/// at once, through whichever few link peers lead here, each retrying on the +/// `node.discovery.attempt_timeouts_secs` ladder. Lowering this makes a +/// genuinely popular node intermittently unresolvable, which is the same +/// symptom as the flood it defends against; raising it raises the worst-case +/// signing burst one neighbour can force. +const DEFAULT_SIGN_BURST: f64 = 256.0; + +/// Sustained signatures per second per link peer. +/// +/// At the default eight or so link peers this caps the node near 256 +/// signatures per second in the sustained case. The real cost of one +/// `Identity::sign` on this codebase has not been measured, so this number +/// is a bound rather than a tuned value; it is the one line to change if a +/// measurement says otherwise. +const DEFAULT_SIGN_RATE: f64 = 32.0; + +/// Maximum age of an idle bucket before cleanup. +const SIGN_MAX_AGE: Duration = Duration::from_secs(300); + +/// Token bucket per link peer for lookups this node answers about itself. +/// +/// A min-interval limiter is the wrong shape here: a popular node receives +/// legitimate bursts of lookups for itself through the few link peers that +/// lead to it, and a min interval refuses all but the first of each burst. +/// A bucket absorbs the burst and paces the sustained rate. +pub struct LookupSignRateLimiter { + buckets: HashMap, + burst: f64, + rate: f64, +} + +struct SignBucket { + /// Tokens remaining, at most `burst`. + tokens: f64, + /// When `tokens` was last refilled. + updated: Instant, +} + +impl LookupSignRateLimiter { + /// Create with default burst and refill rate. + pub fn new() -> Self { + Self::with_params(DEFAULT_SIGN_BURST, DEFAULT_SIGN_RATE) + } + + /// Create with a custom burst and refill rate. + pub fn with_params(burst: f64, rate: f64) -> Self { + Self { + buckets: HashMap::new(), + burst, + rate, + } + } + + /// Spend one token for `from`, or report that its budget is exhausted. + /// + /// Returns true when the signature may be produced. A zero burst is + /// read as "unlimited" rather than "refuse everything", so a + /// misconfiguration cannot make this node unresolvable. + pub fn should_sign(&mut self, from: &NodeAddr) -> bool { + if self.burst <= 0.0 { + return true; + } + let now = Instant::now(); + let burst = self.burst; + let rate = self.rate; + let bucket = self.buckets.entry(*from).or_insert(SignBucket { + tokens: burst, + updated: now, + }); + let elapsed = now.duration_since(bucket.updated).as_secs_f64(); + bucket.tokens = (bucket.tokens + elapsed * rate).min(burst); + bucket.updated = now; + if bucket.tokens < 1.0 { + return false; + } + bucket.tokens -= 1.0; + self.cleanup(now); + true + } + + /// Drop buckets untouched for longer than [`SIGN_MAX_AGE`]; a full + /// bucket carries no state worth keeping. + fn cleanup(&mut self, now: Instant) { + self.buckets + .retain(|_, b| now.duration_since(b.updated) < SIGN_MAX_AGE); + } + + #[cfg(test)] + pub fn len(&self) -> usize { + self.buckets.len() + } +} + +impl Default for LookupSignRateLimiter { + fn default() -> Self { + Self::new() + } +} + +// ============================================================================ +// Tests +// ============================================================================ + +#[cfg(test)] +mod sign_limiter_tests { + use super::*; + + fn addr(val: u8) -> NodeAddr { + let mut bytes = [0u8; 16]; + bytes[0] = val; + NodeAddr::from_bytes(bytes) + } + + // The `DiscoveryBackoff` and `DiscoveryForwardRateLimiter` tests that + // accompanied this limiter in `src/node/discovery_rate_limit.rs` are not + // repeated here: both types moved into `crate::proto::lookup::limits` as + // `LookupBackoff` and `LookupForwardRateLimiter`, and all thirteen of those + // cases live there, name for name, in `src/proto/lookup/tests/limits.rs`, + // ported to the clockless millisecond API. Only the signing budget, which + // is `Instant`-based and so stays in the shell, is exercised below. + + #[test] + fn test_sign_budget_is_spent_per_peer_and_does_not_touch_another_peer() { + let mut limiter = LookupSignRateLimiter::with_params(4.0, 0.0); + for _ in 0..4 { + assert!(limiter.should_sign(&addr(1))); + } + assert!( + !limiter.should_sign(&addr(1)), + "the burst is the whole budget when nothing refills it" + ); + assert!( + limiter.should_sign(&addr(2)), + "one peer spending its budget must not spend another's" + ); + assert_eq!(limiter.len(), 2); + } + + #[test] + fn test_sign_budget_of_zero_burst_is_read_as_unlimited() { + let mut limiter = LookupSignRateLimiter::with_params(0.0, 0.0); + for _ in 0..1000 { + assert!(limiter.should_sign(&addr(1))); + } + assert_eq!(limiter.len(), 0, "unlimited keeps no per-peer state"); + } +} diff --git a/src/node/reject.rs b/src/node/reject.rs index 5dd06b2a..fbaaeab2 100644 --- a/src/node/reject.rs +++ b/src/node/reject.rs @@ -120,7 +120,19 @@ pub enum DiscoveryReject { /// Request dedup cache (`recent_requests`) is at capacity, so the /// `LookupRequest` is dropped without being forwarded. Tracked via /// [`DiscoveryStats::req_dedup_cache_full`](crate::node::stats::DiscoveryStats). + /// + /// Frozen at zero: a full cache now evicts its oldest entry and admits + /// the request, counted as + /// [`DiscoveryStats::req_dedup_evicted`](crate::node::stats::DiscoveryStats). + /// The variant and its counter stay so an operator reading a dashboard + /// across versions does not find the series missing. ReqDedupCacheFull, + /// This node is the lookup target, but the link peer the request + /// arrived from has spent its signing budget. Answering costs a fresh + /// Schnorr signature per request, so the budget bounds what one + /// neighbour can make this node sign. Tracked via + /// [`DiscoveryStats::req_sign_rate_limited`](crate::node::stats::DiscoveryStats). + ReqSignRateLimited, /// Request arrived with TTL=0 — no more forwarding hops allowed. /// Tracked via /// [`DiscoveryStats::req_ttl_exhausted`](crate::node::stats::DiscoveryStats). @@ -136,6 +148,14 @@ pub enum DiscoveryReject { /// Response proof signature failed verification. Tracked via /// [`DiscoveryStats::resp_proof_failed`](crate::node::stats::DiscoveryStats). RespProofFailed, + /// Response arrived on the originator path but carries no + /// `request_id` this node has outstanding for the named target, so + /// it answers no lookup of ours. Expected to be nonzero in healthy + /// operation: the request is flooded to every qualifying tree peer, + /// so duplicate replies land here after the first is accepted. + /// Tracked via + /// [`DiscoveryStats::resp_unsolicited`](crate::node::stats::DiscoveryStats). + RespUnsolicited, /// Response could not be routed toward the origin: no reverse-path /// entry for the `request_id` and no greedy tree route to the /// origin. Tracked via @@ -238,6 +258,18 @@ pub enum SessionReject { /// before any handshake state was created or any ack sent. Tracked via /// [`SessionStats::setup_rate_limited`](crate::node::stats::SessionStats). SetupRateLimited, + /// A session would have been created but the table is at + /// `node.limits.max_sessions`. Refused rather than evicted: the + /// deciding message is unauthenticated at this point, so evicting + /// would hand a stranger a way to tear down established sessions. + /// Tracked via + /// [`SessionStats::table_full`](crate::node::stats::SessionStats). + TableFull, + /// A session would have been created but unauthenticated half-open + /// entries already hold their share of the table. Bounds what a + /// handshake flood can deny an established peer. Tracked via + /// [`SessionStats::half_open_full`](crate::node::stats::SessionStats). + HalfOpenFull, } /// MMP rejection reasons. @@ -368,9 +400,11 @@ mod tests { DiscoveryReject::ReqDuplicate, DiscoveryReject::ReqDedupCacheFull, DiscoveryReject::ReqTtlExhausted, + DiscoveryReject::ReqSignRateLimited, DiscoveryReject::RespDecodeError, DiscoveryReject::RespIdentityMiss, DiscoveryReject::RespProofFailed, + DiscoveryReject::RespUnsolicited, ]; for v in variants { let r = RejectReason::Discovery(v); diff --git a/src/node/session/mod.rs b/src/node/session/mod.rs index a248d482..635c81f9 100644 --- a/src/node/session/mod.rs +++ b/src/node/session/mod.rs @@ -100,6 +100,15 @@ pub(crate) struct SessionEntry { /// Whether this node initiated the Noise handshake. /// Used for spin bit role assignment in session-layer MMP. is_initiator: bool, + /// Largest on-the-wire frame this node has sent toward the remote since + /// the last accepted path-MTU decrease or release, in bytes. + /// + /// Corroborates a reactive `MtuExceeded`, which is unauthenticated: an + /// honest report exists only because a frame this node emitted did not + /// fit some hop, so an honest report always names a value below this. + /// Reset on each accepted decrease and on release so one historical + /// large send cannot vouch for a session's whole lifetime. + max_sent_wire_len: u16, /// Session-layer MMP state. Initialized on Established transition. mmp: Option, @@ -203,6 +212,7 @@ impl SessionEntry { session_start_ms: 0, coords_warmup_remaining: 0, is_initiator, + max_sent_wire_len: 0, mmp: None, packets_sent: 0, packets_recv: 0, @@ -306,6 +316,27 @@ impl SessionEntry { self.coords_warmup_remaining = value; } + /// Largest wire frame sent toward the remote since the last accepted + /// path-MTU decrease or release. + pub(crate) fn max_sent_wire_len(&self) -> u16 { + self.max_sent_wire_len + } + + /// Note a frame of `wire_len` bytes sent toward the remote, keeping the + /// largest. Frames beyond `u16::MAX` saturate, which only ever makes the + /// corroboration more permissive and cannot exceed what a path MTU can + /// name. + pub(crate) fn record_sent_wire_len(&mut self, wire_len: usize) { + let wire_len = u16::try_from(wire_len).unwrap_or(u16::MAX); + self.max_sent_wire_len = self.max_sent_wire_len.max(wire_len); + } + + /// Forget what has been sent, so the next reactive report needs fresh + /// evidence of its own. + pub(crate) fn clear_sent_wire_len(&mut self) { + self.max_sent_wire_len = 0; + } + /// Mark the session as started (transition to Established). /// /// Records the current time as the session start for computing @@ -729,10 +760,24 @@ impl SessionEntry { /// permanent silent decrypt failure. A peer that never catches up /// is instead handled by the FSP session liveness path (fresh /// handshake / teardown of a genuinely dead link). - pub(crate) fn drain_expired(&self, now_ms: u64, drain_ms: u64) -> bool { + /// + /// `max_drain_ms` is an absolute ceiling measured from the cutover + /// alone, so the sliding deadline delays erasure by a bounded amount + /// rather than preventing it: the only party that can push the + /// deadline out is the authenticated peer holding the old key, and + /// without a ceiling it holds that key for as long as it keeps using + /// it. What the ceiling costs is that a peer which has still not + /// recovered by then is cut off deliberately, and its frames are + /// undecryptable until its own rekey retry re-converges the epochs. + /// It must therefore stay above the worst-case legitimate recovery; + /// see `DRAIN_MAX_RETENTION_SECS`. + pub(crate) fn drain_expired(&self, now_ms: u64, drain_ms: u64, max_drain_ms: u64) -> bool { if self.drain_started_ms == 0 { return false; } + if now_ms.saturating_sub(self.drain_started_ms) >= max_drain_ms { + return true; + } let deadline_anchor = self.drain_started_ms.max(self.previous_last_used_ms); now_ms.saturating_sub(deadline_anchor) >= drain_ms } @@ -1229,6 +1274,9 @@ mod overlapping_epoch_tests { #[test] fn drain_expiry_is_peer_progress_aware() { const DRAIN_MS: u64 = 10_000; + // Well clear of the shipped ceiling, so this case still exercises + // the sliding deadline and nothing else. + const MAX_MS: u64 = 120_000; let cutover_ms = 1_000; // Build the post-cutover state via the production cutover path: @@ -1257,7 +1305,7 @@ mod overlapping_epoch_tests { // Even though `now - drain_started_ms` exceeds DRAIN_MS, the // window is NOT expired: the peer just used `previous`. assert!( - !entry.drain_expired(t, DRAIN_MS), + !entry.drain_expired(t, DRAIN_MS, MAX_MS), "previous slot must not be retired while peer keeps using it (t={t})" ); assert!( @@ -1270,11 +1318,11 @@ mod overlapping_epoch_tests { // last `previous`-slot use was at t=25_000; the window now // elapses DRAIN_MS after that, NOT DRAIN_MS after the cutover. assert!( - !entry.drain_expired(34_999, DRAIN_MS), + !entry.drain_expired(34_999, DRAIN_MS, MAX_MS), "window must not expire before DRAIN_MS past the last previous use" ); assert!( - entry.drain_expired(35_000, DRAIN_MS), + entry.drain_expired(35_000, DRAIN_MS, MAX_MS), "window must expire DRAIN_MS after the last previous-slot decrypt" ); @@ -1292,6 +1340,7 @@ mod overlapping_epoch_tests { #[test] fn drain_expiry_unaffected_when_peer_off_old_epoch() { const DRAIN_MS: u64 = 10_000; + const MAX_MS: u64 = 120_000; let cutover_ms = 1_000; let (_old_send, old_recv) = xk_pair(1, 2); @@ -1303,141 +1352,106 @@ mod overlapping_epoch_tests { // No old-epoch frames ever arrive: `previous_last_used_ms` stays // 0, the deadline anchor is the cutover time. assert!( - !entry.drain_expired(cutover_ms + DRAIN_MS - 1, DRAIN_MS), + !entry.drain_expired(cutover_ms + DRAIN_MS - 1, DRAIN_MS, MAX_MS), "window must not expire early" ); assert!( - entry.drain_expired(cutover_ms + DRAIN_MS, DRAIN_MS), + entry.drain_expired(cutover_ms + DRAIN_MS, DRAIN_MS, MAX_MS), "window must expire on the plain wall-clock timer when peer is off the old epoch" ); } - // ======================================================================== - // Rekey-policy characterization (pins `check_session_rekey`'s decision - // boundaries before the `Fsp::poll_rekey` hoist — these thresholds have no - // other test module). - // ======================================================================== + // The retention ceiling the drain cap added. The five rekey-policy + // characterization tests that sat here on `maint` are not carried: they + // pinned `check_session_rekey`'s boundaries before the `Fsp::poll_rekey` + // hoist, and the hoisted predicate has its own tests in + // `src/proto/fsp/tests/core.rs`. - /// The initiator liveness-cutover delay used by `check_session_rekey` - /// (`FSP_CUTOVER_DELAY_MS`). Mirrored here as the characterization anchor. - const CUTOVER_DELAY_MS: u64 = 2000; + // 12. A peer that keeps exercising the old epoch delays erasure by a + // bounded amount rather than preventing it. The refreshes must + // continue past the ceiling: a case that stops refreshing at the + // boundary passes without the ceiling and proves nothing. + #[test] + fn drain_retention_is_capped_against_a_peer_pinning_the_old_epoch() { + const DRAIN_MS: u64 = 10_000; + const MAX_MS: u64 = 120_000; - /// Build an established entry that has completed a rekey as initiator and - /// holds a pending session awaiting the K-bit cutover. - fn entry_pending_cutover(rekey_completed_ms: u64) -> SessionEntry { - let (_cur_send, cur_recv) = xk_pair(1, 2); + let (_old_send, old_recv) = xk_pair(1, 2); let (_new_send, new_recv) = xk_pair(3, 4); - let mut entry = entry_with_current(cur_recv); - // Mark ourselves the rekey initiator, then land the completed session - // as pending (clears rekey_state, so has_rekey_in_progress() == false). - entry.set_rekey_state(HandshakeState::new_xk_responder(keypair(7)), true); + let mut entry = entry_with_current(old_recv); entry.set_pending_session(new_recv); - entry.set_rekey_completed_ms(rekey_completed_ms); - entry - } + assert!(entry.cutover_to_new_session(1)); - // The initiator-side cutover predicate: pending session present, no rekey - // in progress, we are the initiator, and the liveness timer has elapsed. - #[test] - fn rekey_cutover_predicate_boundary() { - let completed = 1_000u64; - let entry = entry_pending_cutover(completed); + // One old-epoch frame every half window, which is what a peer + // pinning the drain deadline actually does. + let mut t = 1u64; + while t < MAX_MS { + entry.refresh_previous_use(t); + assert!( + !entry.drain_expired(t, DRAIN_MS, MAX_MS), + "ceiling fired before the peer's grace ran out (t={t})" + ); + t += DRAIN_MS / 2; + } - assert!(entry.pending_new_session().is_some()); - assert!(!entry.has_rekey_in_progress()); - assert!(entry.is_rekey_initiator()); - - // Not yet eligible one ms before the delay elapses. - let just_before = completed + CUTOVER_DELAY_MS - 1; + // Still refreshing, so the sliding deadline is nowhere near due. + entry.refresh_previous_use(MAX_MS + 1); assert!( - just_before.saturating_sub(entry.rekey_completed_ms()) < CUTOVER_DELAY_MS, - "cutover must not fire before the liveness delay" - ); - // Eligible exactly at the delay. - let at = completed + CUTOVER_DELAY_MS; - assert!( - at.saturating_sub(entry.rekey_completed_ms()) >= CUTOVER_DELAY_MS, - "cutover fires once the liveness delay has elapsed" + entry.drain_expired(MAX_MS + 1, DRAIN_MS, MAX_MS), + "a peer pinning the old epoch retained the retired key past the ceiling" ); } - // The rekey trigger's own threshold arithmetic is tested against the real - // predicate in src/proto/fmp/tests/core.rs, which drives poll_rekey. A test - // here previously reproduced that OR predicate as a local closure and - // asserted against its own copy, so it could not fail for the reason it - // existed: deleting the counter arm from the trigger left it green. Its one - // assertion over real code, that a fresh entry's jitter lies within the - // symmetric bound, is covered over 100 samples by - // test_session_entry_rekey_jitter_in_range in src/node/tests/session.rs. - - // Dampening boundary: within `dampening_ms` of the peer's rekey msg1, local - // initiation is suppressed; at/after the window it is not. + // 13. The ceiling must not shorten the grace the sliding deadline + // exists to give a peer that lost msg3 and is still catching up. #[test] - fn rekey_dampening_boundary() { - let (_s, recv) = xk_pair(1, 2); - let mut entry = entry_with_current(recv); - const DAMP_MS: u64 = 30_000; + fn drain_retention_cap_does_not_shorten_the_ordinary_grace() { + const DRAIN_MS: u64 = 10_000; + const MAX_MS: u64 = 120_000; + const CUTOVER_MS: u64 = 1_000; - // No peer rekey recorded → never dampened. - assert!(!entry.is_rekey_dampened(50_000, DAMP_MS)); + let (_old_send, old_recv) = xk_pair(1, 2); + let (_new_send, new_recv) = xk_pair(3, 4); + let mut entry = entry_with_current(old_recv); + entry.set_pending_session(new_recv); + assert!(entry.cutover_to_new_session(CUTOVER_MS)); + assert!(entry.is_draining()); - entry.record_peer_rekey(10_000); - assert!( - entry.is_rekey_dampened(10_000 + DAMP_MS - 1, DAMP_MS), - "dampened within the window" - ); - assert!( - !entry.is_rekey_dampened(10_000 + DAMP_MS, DAMP_MS), - "not dampened once the window has elapsed" - ); + // A peer that lost msg3 and is still catching up sends one + // old-epoch frame part-way through the window. + let last_use = CUTOVER_MS + DRAIN_MS / 2; + entry.refresh_previous_use(last_use); + assert!(!entry.drain_expired(CUTOVER_MS + DRAIN_MS, DRAIN_MS, MAX_MS)); + assert!(!entry.drain_expired(last_use + DRAIN_MS - 1, DRAIN_MS, MAX_MS)); + assert!(entry.drain_expired(last_use + DRAIN_MS, DRAIN_MS, MAX_MS)); } - // Epoch-reaction: a frame authenticating against `pending` while a msg3 - // retransmission is retained confirms the peer on the new epoch (clears the - // msg3 payload) and then promotes. + // 14. The shipped ceiling has to clear the worst-case legitimate + // recovery of a peer that lost msg3: the msg3 resend ladder, the + // responder's handshake timeout, and the rekey dampening window + // before it may re-initiate. A later tightening of the ceiling + // reds this rather than silently amputating that recovery. #[test] - fn epoch_reaction_pending_confirms_then_promotes() { - let (mut p_send, p_recv) = xk_pair(3, 4); - let (_cur_send, cur_recv) = xk_pair(1, 2); - let mut entry = entry_with_current(cur_recv); - let k_before = entry.current_k_bit(); - entry.set_pending_session(p_recv); - entry.set_rekey_msg3_payload(vec![0xAB; 8], 5_000); - assert!(entry.rekey_msg3_payload().is_some()); + fn drain_retention_cap_clears_the_msg3_recovery_budget() { + const DRAIN_MS: u64 = 10_000; + // 31 s resend ladder + 30 s handshake timeout + 30 s dampening. + const RECOVERY_MS: u64 = 91_000; + const CUTOVER_MS: u64 = 1_000; + let max_ms = crate::node::handlers::rekey::drain_max_retention_ms( + &crate::config::RateLimitConfig::default(), + ); - let (ct, counter, hdr) = seal(&mut p_send, b"new-epoch", !k_before); - let (_pt, slot) = entry - .fsp_trial_decrypt(&ct, counter, &hdr, !k_before, 2_000) - .expect("pending frame decrypts"); - assert_eq!(slot, EpochSlot::Pending); + let (_old_send, old_recv) = xk_pair(1, 2); + let (_new_send, new_recv) = xk_pair(3, 4); + let mut entry = entry_with_current(old_recv); + entry.set_pending_session(new_recv); + assert!(entry.cutover_to_new_session(CUTOVER_MS)); + assert!(entry.is_draining()); - // Reaction order: confirm (while pending still held) then promote. - entry.confirm_peer_new_epoch(); - assert!(entry.rekey_msg3_payload().is_none()); - entry.handle_peer_kbit_flip(2_000); - assert!(entry.pending_new_session().is_none()); - assert_ne!(entry.current_k_bit(), k_before); - } - - // Epoch-reaction: as the initiator that already cut over on its own timer - // (msg3 retained, no pending), a frame authenticating against `current` - // confirms the responder reached the new epoch. - #[test] - fn epoch_reaction_current_confirms_responder() { - let (mut cur_send, cur_recv) = xk_pair(1, 2); - let mut entry = entry_with_current(cur_recv); - entry.set_rekey_msg3_payload(vec![0xCD; 8], 5_000); - assert!(entry.pending_new_session().is_none()); - assert!(entry.rekey_msg3_payload().is_some()); - - let (ct, counter, hdr) = seal(&mut cur_send, b"steady", false); - let (_pt, slot) = entry - .fsp_trial_decrypt(&ct, counter, &hdr, false, 2_000) - .expect("current frame decrypts"); - assert_eq!(slot, EpochSlot::Current); - - // The Current-with-retained-msg3-and-no-pending arm confirms. - entry.confirm_peer_new_epoch(); - assert!(entry.rekey_msg3_payload().is_none()); + entry.refresh_previous_use(CUTOVER_MS + RECOVERY_MS); + assert!( + !entry.drain_expired(CUTOVER_MS + RECOVERY_MS, DRAIN_MS, max_ms), + "ceiling fires inside the msg3 recovery budget" + ); } } diff --git a/src/node/stats.rs b/src/node/stats.rs index e9699085..74fe5ed3 100644 --- a/src/node/stats.rs +++ b/src/node/stats.rs @@ -79,6 +79,14 @@ pub struct SessionStats { /// A setup message was refused by the per-link-peer setup limiter, /// before any handshake state was created or any ack sent. pub setup_rate_limited: u64, + /// A session would have been created but the table is at + /// `node.limits.max_sessions`. A sustained rate means either the cap + /// is sized below what this node legitimately carries, or something is + /// holding the table full. + pub table_full: u64, + /// A session would have been created but unauthenticated half-open + /// entries already hold their share of the table. + pub half_open_full: u64, } impl SessionStats { @@ -96,6 +104,8 @@ impl SessionStats { pending_replaced: self.pending_replaced, ack_handshake_failed: self.ack_handshake_failed, setup_rate_limited: self.setup_rate_limited, + table_full: self.table_full, + half_open_full: self.half_open_full, } } @@ -110,6 +120,8 @@ impl SessionStats { SessionReject::RekeyPending => self.rekey_pending += 1, SessionReject::AckHandshakeFailed => self.ack_handshake_failed += 1, SessionReject::SetupRateLimited => self.setup_rate_limited += 1, + SessionReject::TableFull => self.table_full += 1, + SessionReject::HalfOpenFull => self.half_open_full += 1, } } } @@ -319,6 +331,8 @@ pub struct LookupStatsSnapshot { pub req_decode_error: u64, pub req_duplicate: u64, pub req_dedup_cache_full: u64, + pub req_dedup_evicted: u64, + pub req_sign_rate_limited: u64, pub req_target_is_us: u64, pub req_forwarded: u64, pub req_ttl_exhausted: u64, @@ -334,6 +348,7 @@ pub struct LookupStatsSnapshot { pub resp_forwarded: u64, pub resp_identity_miss: u64, pub resp_proof_failed: u64, + pub resp_unsolicited: u64, pub resp_no_route: u64, pub resp_accepted: u64, pub resp_timed_out: u64, @@ -389,6 +404,8 @@ pub struct SessionStatsSnapshot { pub pending_replaced: u64, pub ack_handshake_failed: u64, pub setup_rate_limited: u64, + pub table_full: u64, + pub half_open_full: u64, } #[derive(Clone, Debug, Default, Serialize)] @@ -421,6 +438,10 @@ pub struct ErrorSignalStatsSnapshot { pub unbound_broken: u64, pub unbound_mtu: u64, pub unbound_forged: u64, + pub emit_over_peer_budget: u64, + pub emit_over_dest_interval: u64, + pub emit_limiter_at_capacity: u64, + pub mtu_exceeded_uncorroborated: u64, } #[derive(Clone, Debug, Default, Serialize)] diff --git a/src/node/tests/discovery.rs b/src/node/tests/discovery.rs index 83656d1d..dda11ee1 100644 --- a/src/node/tests/discovery.rs +++ b/src/node/tests/discovery.rs @@ -83,6 +83,19 @@ async fn test_request_ttl_zero_not_forwarded() { // Unit Tests — LookupResponse Handler // ============================================================================ +/// Record `request_id` as outstanding for `target`, exactly as +/// `initiate_lookup` does when it puts a request on the wire. The response +/// handler correlates against this, so a unit test that hands the handler a +/// response without it is testing the correlation gate rather than whatever +/// it names. +fn seed_pending_lookup(node: &mut Node, target: crate::NodeAddr, request_id: u64) { + node.lookup + .pending_lookups + .entry(target) + .or_insert_with(|| crate::proto::lookup::PendingLookup::new(Node::now_ms())) + .record(request_id); +} + #[tokio::test] async fn test_response_decode_error() { let mut node = make_node(); @@ -106,6 +119,8 @@ async fn test_response_originator_caches_route() { // Register target identity in cache so verification can find it node.register_identity(target, target_identity.pubkey_full()); + seed_pending_lookup(&mut node, target, 555); + // Create a valid response with a real proof signature (includes coords) let proof_data = LookupResponse::proof_bytes(555, &target, &coords); let proof = target_identity.sign(&proof_data); @@ -183,6 +198,8 @@ async fn test_response_proof_verification_success() { // Register target in identity_cache node.register_identity(target, target_identity.pubkey_full()); + seed_pending_lookup(&mut node, target, 700); + // Sign with correct proof_bytes (including coords) let proof_data = LookupResponse::proof_bytes(700, &target, &coords); let proof = target_identity.sign(&proof_data); @@ -218,6 +235,8 @@ async fn test_response_proof_verification_failure() { node.register_identity(target, target_identity.pubkey_full()); // Sign with a DIFFERENT identity (wrong key) + seed_pending_lookup(&mut node, target, 701); + let wrong_identity = Identity::generate(); let proof_data = LookupResponse::proof_bytes(701, &target, &coords); let proof = wrong_identity.sign(&proof_data); @@ -251,6 +270,8 @@ async fn test_response_identity_cache_miss() { // Do NOT register target in identity_cache + seed_pending_lookup(&mut node, target, 702); + let proof_data = LookupResponse::proof_bytes(702, &target, &coords); let proof = target_identity.sign(&proof_data); @@ -285,6 +306,8 @@ async fn test_response_coord_substitution_detected() { // Register target in identity_cache node.register_identity(target, target_identity.pubkey_full()); + seed_pending_lookup(&mut node, target, 703); + // Sign proof with real coords let proof_data = LookupResponse::proof_bytes(703, &target, &real_coords); let proof = target_identity.sign(&proof_data); @@ -305,6 +328,267 @@ async fn test_response_coord_substitution_detected() { ); } +/// Build a signed LookupResponse body for `target_identity` over +/// `request_id`, ready to hand to `handle_lookup_response`. +fn signed_response_body( + target_identity: &Identity, + request_id: u64, + coords: &TreeCoordinate, +) -> Vec { + let target = *target_identity.node_addr(); + let proof_data = LookupResponse::proof_bytes(request_id, &target, coords); + let proof = target_identity.sign(&proof_data); + LookupResponse::new(request_id, target, coords.clone(), proof).encode()[1..].to_vec() +} + +/// Register `target_identity` and return its address and a plausible +/// coordinate for it, the shared preamble of the correlation tests. +fn register_lookup_target(node: &mut Node, target_identity: &Identity) -> TreeCoordinate { + let target = *target_identity.node_addr(); + node.register_identity(target, target_identity.pubkey_full()); + TreeCoordinate::from_addrs(vec![target, make_node_addr(0xF0)]).unwrap() +} + +fn wall_clock_ms() -> u64 { + std::time::SystemTime::now() + .duration_since(std::time::UNIX_EPOCH) + .map(|d| d.as_millis() as u64) + .unwrap_or(0) +} + +#[tokio::test] +async fn test_unsolicited_lookup_response_is_dropped_before_proof_verification() { + // Any admitted peer can hand us a correctly signed response for a target + // we never asked about. Accepting it lets that peer clear our pending + // state, refresh a cache entry's TTL and flush our queued packets at a + // moment it picks, so the response must not be acted on at all. + let mut node = make_node(); + let from = make_node_addr(0xAA); + + let target_identity = Identity::generate(); + let target = *target_identity.node_addr(); + let coords = register_lookup_target(&mut node, &target_identity); + let body = signed_response_body(&target_identity, 900, &coords); + + assert!( + node.lookup.pending_lookups.is_empty(), + "precondition: this node has no lookup outstanding for anything" + ); + + node.handle_lookup_response(&from, &body).await; + + assert!( + !node.coord_cache().contains(&target, wall_clock_ms()), + "a response answering no request of ours must not reach the coordinate cache" + ); + assert_eq!( + node.metrics().lookup.resp_accepted.get(), + 0, + "an unsolicited response must not count as accepted" + ); + assert_eq!( + node.metrics().lookup.resp_unsolicited.get(), + 1, + "the drop must be visible on a counter, not only in a log" + ); + assert_eq!( + node.metrics().lookup.resp_proof_failed.get(), + 0, + "the drop must happen before the signature verify, so the verify is not a cost gate" + ); +} + +#[tokio::test] +async fn test_lookup_response_with_a_request_id_we_never_issued_is_dropped() { + // Correlating on the target alone would leave the attack open: there is + // no inbound limiter on responses, so a peer can spray a harvested one + // and land inside any window in which we happen to be looking that + // target up. The id must match too. + let mut node = make_node(); + let from = make_node_addr(0xAA); + + let target_identity = Identity::generate(); + let target = *target_identity.node_addr(); + let coords = register_lookup_target(&mut node, &target_identity); + + node.initiate_lookup(&target, 5).await; + let issued = node + .lookup + .pending_lookups + .get(&target) + .unwrap() + .ids + .clone(); + assert_eq!(issued.len(), 1, "precondition: one attempt went out"); + + let body = signed_response_body(&target_identity, issued[0] ^ 1, &coords); + node.handle_lookup_response(&from, &body).await; + + assert!( + !node.coord_cache().contains(&target, wall_clock_ms()), + "a response bearing an id we never issued must not reach the coordinate cache" + ); + assert!( + node.lookup.pending_lookups.contains_key(&target), + "it must not cancel the lookup that is genuinely outstanding" + ); + assert_eq!(node.metrics().lookup.resp_unsolicited.get(), 1); +} + +#[tokio::test] +async fn test_a_response_matching_a_pending_attempt_is_accepted_and_clears_the_pending_lookup() { + // The healthy path. A fix that reds a legitimate lookup is no use, and + // this is the test that catches it. + let mut node = make_node(); + let from = make_node_addr(0xAA); + + let target_identity = Identity::generate(); + let target = *target_identity.node_addr(); + let coords = register_lookup_target(&mut node, &target_identity); + + node.initiate_lookup(&target, 5).await; + let issued = node.lookup.pending_lookups.get(&target).unwrap().ids[0]; + + let body = signed_response_body(&target_identity, issued, &coords); + node.handle_lookup_response(&from, &body).await; + + assert_eq!( + node.coord_cache().get(&target, wall_clock_ms()), + Some(&coords), + "a response to our own outstanding request must be cached" + ); + assert!( + !node.lookup.pending_lookups.contains_key(&target), + "accepting it must clear the pending lookup" + ); + assert_eq!(node.metrics().lookup.resp_accepted.get(), 1); + assert_eq!(node.metrics().lookup.resp_unsolicited.get(), 0); +} + +#[tokio::test] +async fn test_a_late_response_for_an_earlier_retry_attempt_is_still_accepted() { + // Each retry draws a fresh id, and on any link with more than a second + // of round trip the reply to an earlier attempt is the common case. A + // correlator that remembered only the newest id would drop it. + let mut node = make_node(); + let from = make_node_addr(0xAA); + + let target_identity = Identity::generate(); + let target = *target_identity.node_addr(); + let coords = register_lookup_target(&mut node, &target_identity); + + node.initiate_lookup(&target, 5).await; + node.initiate_lookup(&target, 5).await; + let issued = node + .lookup + .pending_lookups + .get(&target) + .unwrap() + .ids + .clone(); + assert_eq!(issued.len(), 2, "precondition: two attempts, two ids"); + + let body = signed_response_body(&target_identity, issued[0], &coords); + node.handle_lookup_response(&from, &body).await; + + assert_eq!( + node.coord_cache().get(&target, wall_clock_ms()), + Some(&coords), + "the first attempt's id is still ours and its answer must be accepted" + ); +} + +#[tokio::test] +async fn test_a_second_genuine_response_after_the_first_is_accepted_is_dropped() { + // The request is flooded to every qualifying tree peer, so duplicate + // replies are routine. They are dropped at the correlation gate, which + // gives the unsolicited counter a nonzero floor in healthy operation: + // it is not by itself a sign of attack traffic. + let mut node = make_node(); + let from = make_node_addr(0xAA); + + let target_identity = Identity::generate(); + let target = *target_identity.node_addr(); + let coords = register_lookup_target(&mut node, &target_identity); + + node.initiate_lookup(&target, 5).await; + let issued = node.lookup.pending_lookups.get(&target).unwrap().ids[0]; + let body = signed_response_body(&target_identity, issued, &coords); + + node.handle_lookup_response(&from, &body).await; + node.handle_lookup_response(&from, &body).await; + + assert_eq!( + node.coord_cache().get(&target, wall_clock_ms()), + Some(&coords), + "the value written by the first response must still be there" + ); + assert_eq!( + node.metrics().lookup.resp_accepted.get(), + 1, + "only the first of the two answers our request" + ); + assert_eq!( + node.metrics().lookup.resp_unsolicited.get(), + 1, + "the duplicate is counted, which is why the counter has a healthy floor" + ); +} + +#[tokio::test] +async fn test_a_validly_signed_response_for_a_retired_lookup_is_dropped() { + // The pending entry's lifetime is what bounds how stale an accepted + // coordinate can be. Once the retry ladder is exhausted and the entry + // goes, a transit node holding the genuine reply can no longer deliver + // it late. + let mut node = make_node(); + let from = make_node_addr(0xAA); + + let target_identity = Identity::generate(); + let target = *target_identity.node_addr(); + let coords = register_lookup_target(&mut node, &target_identity); + + node.initiate_lookup(&target, 5).await; + let issued = node.lookup.pending_lookups.get(&target).unwrap().ids[0]; + + // Drive the whole ladder: three retries, then the final timeout. + let mut now_ms = Node::now_ms(); + for _ in 0..4 { + now_ms += 100_000; + node.check_pending_lookups(now_ms).await; + } + assert!( + !node.lookup.pending_lookups.contains_key(&target), + "precondition: the ladder retired the lookup" + ); + + let body = signed_response_body(&target_identity, issued, &coords); + node.handle_lookup_response(&from, &body).await; + + assert!( + !node.coord_cache().contains(&target, wall_clock_ms()), + "a reply to a retired lookup must not install a coordinate" + ); + assert_eq!(node.metrics().lookup.resp_unsolicited.get(), 1); +} + +#[test] +fn pending_lookup_id_set_evicts_the_oldest_id_rather_than_refusing_the_newest() { + // The retry ladder is operator configuration and can be longer than the + // recorded-id cap. Refusing the newest id would discard the attempt most + // likely to be answered and fail a healthy lookup. + let mut pending = crate::proto::lookup::PendingLookup::new(0); + for id in 0..12u64 { + pending.record(id); + } + assert!( + pending.matches(11), + "the newest attempt's id must always be remembered" + ); + assert!(!pending.matches(0), "the oldest id is the one evicted"); + assert_eq!(pending.ids.len(), 8, "the set stays bounded"); +} + // ============================================================================ // Unit Tests — RecentRequest Expiry // ============================================================================ @@ -345,6 +629,212 @@ async fn test_recent_request_expiry() { assert!(node.lookup.recent_requests.contains_key(&789)); } +// ============================================================================ +// Unit Tests — dedup cache capacity policy +// ============================================================================ + +use crate::proto::lookup::{MAX_RECENT_LOOKUP_REQUESTS, MIN_RECENT_PER_PEER}; + +/// Encode a LookupRequest for `target` carrying `request_id`, ready for +/// `handle_lookup_request` (which is handed the payload without the +/// msg_type byte). +fn lookup_request_payload(request_id: u64, target: &crate::NodeAddr) -> Vec { + let origin = make_node_addr(0xCC); + let coords = TreeCoordinate::from_addrs(vec![origin, make_node_addr(0)]).unwrap(); + LookupRequest::new(request_id, *target, origin, coords, 5, 0).encode()[1..].to_vec() +} + +/// Deliver `count` distinct requests from `from`, ids starting at `first_id`. +async fn flood_requests(node: &mut Node, from: &crate::NodeAddr, first_id: u64, count: u64) { + let target = make_node_addr(0xBB); + for i in 0..count { + let payload = lookup_request_payload(first_id + i, &target); + node.handle_lookup_request(from, &payload).await; + } +} + +/// Register `count` peers so the per-peer share of the dedup cache is the +/// floor rather than the whole cache, and return their addresses. +fn register_peers(node: &mut Node, count: usize) -> Vec { + (0..count) + .map(|i| { + let identity = Identity::generate(); + let addr = *identity.node_addr(); + let peer_identity = crate::PeerIdentity::from_pubkey(identity.pubkey()); + node.peers.insert( + addr, + ActivePeer::new(peer_identity, LinkId::new(i as u64), 0), + ); + addr + }) + .collect() +} + +#[tokio::test] +async fn test_a_full_dedup_cache_admits_the_new_request_by_evicting_the_oldest() { + // A full cache used to drop the arriving request, which let one peer + // spend 4096 fresh request_ids and stop the node forwarding anyone + // else's lookups until the entries aged out. + let mut node = make_node(); + let from = make_node_addr(0xAA); + + flood_requests(&mut node, &from, 1, MAX_RECENT_LOOKUP_REQUESTS as u64).await; + assert_eq!( + node.lookup.recent_requests.len(), + MAX_RECENT_LOOKUP_REQUESTS, + "precondition: the cache is full, or the rest observes nothing" + ); + + let payload = lookup_request_payload(u64::MAX, &make_node_addr(0xBB)); + node.handle_lookup_request(&from, &payload).await; + + assert!( + node.lookup.recent_requests.contains_key(&u64::MAX), + "the arriving request must be recorded, so its response can be routed back" + ); + assert!( + !node.lookup.recent_requests.contains_key(&1), + "room is made by dropping the oldest entry" + ); + assert_eq!( + node.lookup.recent_requests.len(), + MAX_RECENT_LOOKUP_REQUESTS, + "the cache stays at its bound" + ); + assert_eq!(node.metrics().lookup.req_dedup_evicted.get(), 1); + assert_eq!( + node.metrics().lookup.req_dedup_cache_full.get(), + 0, + "the cache-full drop is gone, and its counter stays frozen at zero" + ); +} + +#[tokio::test] +async fn test_a_flooding_peer_evicts_only_its_own_dedup_entries() { + // The whole point of partitioning the cache by link peer: one peer + // filling its share must not cost another peer the reverse path its own + // lookup depends on. + let mut node = make_node(); + let peers = register_peers(&mut node, 64); + let flooder = peers[0]; + let light = peers[1]; + + let payload = lookup_request_payload(7, &make_node_addr(0xBB)); + node.handle_lookup_request(&light, &payload).await; + + // One over the share, so the flooder pays for its own admission. + flood_requests(&mut node, &flooder, 1000, MIN_RECENT_PER_PEER as u64 + 1).await; + + assert!( + node.lookup.recent_requests.contains_key(&7), + "a light peer's reverse-path entry must survive a neighbour's flood" + ); + assert!( + !node.lookup.recent_requests.contains_key(&1000), + "the flooder's own oldest entry is what pays for its newest" + ); + assert!( + node.lookup + .recent_requests + .contains_key(&(1000 + MIN_RECENT_PER_PEER as u64)), + "and its newest is admitted rather than dropped" + ); +} + +#[tokio::test] +async fn test_a_node_whose_dedup_cache_is_flooded_still_answers_a_lookup_for_itself() { + // The availability claim. Filling the cache used to make the node + // unresolvable, because the cache-full drop sat ahead of the check for + // whether the request names us. + let mut node = make_node(); + let flooder = make_node_addr(0xAA); + let other = make_node_addr(0xAB); + + flood_requests(&mut node, &flooder, 1, MAX_RECENT_LOOKUP_REQUESTS as u64).await; + + let my_addr = *node.node_addr(); + let payload = lookup_request_payload(u64::MAX, &my_addr); + node.handle_lookup_request(&other, &payload).await; + + assert_eq!( + node.metrics().lookup.req_target_is_us.get(), + 1, + "a flooded cache must not stop the node answering lookups for itself" + ); +} + +#[tokio::test] +async fn test_the_dedup_index_stays_level_with_the_cache_across_insert_duplicate_and_purge() { + // Two containers where there was one, so the desync is the maintenance + // risk. Everything the eviction policy decides reads the index, so an + // index that has drifted evicts the wrong entry or none at all. + let mut node = make_node(); + let peers = register_peers(&mut node, 64); + + flood_requests(&mut node, &peers[0], 1, 70).await; + flood_requests(&mut node, &peers[1], 500, 5).await; + // Duplicates, which must not be indexed twice. + flood_requests(&mut node, &peers[1], 500, 5).await; + + let indexed: usize = node + .lookup + .recent_by_peer + .values() + .map(|ids| ids.len()) + .sum(); + assert_eq!( + indexed, + node.lookup.recent_requests.len(), + "every cached request is indexed exactly once" + ); + + // Age everything out and purge through the ordinary request path. + let expiry_ms = node.config().node.lookup.recent_expiry_secs * 1000; + let future = Node::now_ms() + expiry_ms + 1; + node.purge_expired_requests(future); + + assert!( + node.lookup.recent_requests.is_empty(), + "precondition: the purge removed everything" + ); + assert!( + node.lookup.recent_by_peer.is_empty(), + "the index must not keep entries the cache no longer holds" + ); +} + +#[tokio::test] +async fn test_answering_lookups_for_ourselves_stops_at_the_per_peer_signing_budget() { + // Each answer costs a fresh Schnorr signature, because the proof is + // bound to the requester's request_id and cannot be reused. Without a + // budget, one neighbour sets this node's signing rate. + let mut node = make_node(); + node.set_discovery_sign_budget(3.0, 0.0); + let from = make_node_addr(0xAA); + let other = make_node_addr(0xAB); + let my_addr = *node.node_addr(); + + for id in 0..4u64 { + let payload = lookup_request_payload(id, &my_addr); + node.handle_lookup_request(&from, &payload).await; + } + + assert_eq!( + node.metrics().lookup.req_target_is_us.get(), + 3, + "the burst is answered and the fourth request is not" + ); + assert_eq!(node.metrics().lookup.req_sign_rate_limited.get(), 1); + + let payload = lookup_request_payload(100, &my_addr); + node.handle_lookup_request(&other, &payload).await; + assert_eq!( + node.metrics().lookup.req_target_is_us.get(), + 4, + "one peer spending its budget must not make the node unresolvable through another" + ); +} + // ============================================================================ // Integration Tests — Multi-Node Forwarding // ============================================================================ @@ -888,6 +1378,8 @@ async fn test_originator_stores_path_mtu_in_cache() { node.register_identity(target, target_identity.pubkey_full()); + seed_pending_lookup(&mut node, target, 800); + let proof_data = LookupResponse::proof_bytes(800, &target, &coords); let proof = target_identity.sign(&proof_data); @@ -932,6 +1424,8 @@ async fn test_originator_ignores_sub_floor_path_mtu_but_still_caches_coords() { node.register_identity(target, target_identity.pubkey_full()); + seed_pending_lookup(&mut node, target, 801); + let proof_data = LookupResponse::proof_bytes(801, &target, &coords); let proof = target_identity.sign(&proof_data); @@ -999,6 +1493,8 @@ async fn test_actionable_lookup_response_path_mtu_does_not_bump_below_floor_coun node.register_identity(target, target_identity.pubkey_full()); + seed_pending_lookup(&mut node, target, 802); + let proof_data = LookupResponse::proof_bytes(802, &target, &coords); let proof = target_identity.sign(&proof_data); @@ -1045,6 +1541,8 @@ async fn test_originator_lookup_response_keeps_tighter_path_mtu_lookup() { let target_fips = crate::FipsAddress::from_node_addr(&target); node.path_mtu_lookup_insert(target_fips, 1280); + seed_pending_lookup(&mut node, target, 800); + let proof_data = LookupResponse::proof_bytes(800, &target, &coords); let proof = target_identity.sign(&proof_data); @@ -1079,6 +1577,7 @@ fn make_verified_lookup_response( let coords = TreeCoordinate::from_addrs(vec![target, root]).unwrap(); node.register_identity(target, target_identity.pubkey_full()); + seed_pending_lookup(node, target, request_id); let proof_data = LookupResponse::proof_bytes(request_id, &target, &coords); let proof = target_identity.sign(&proof_data); @@ -1312,7 +1811,7 @@ async fn test_open_discovery_sweep_queues_eligible_skips_filtered() { /// verified indirectly: each `req_initiated` increment corresponds /// to one fresh `initiate_lookup` call. /// 3. **Final-timeout state transitions** — `pending_lookups` entry is -/// removed, `discovery.resp_timed_out` counter ticks, queued packet +/// removed, `lookup.resp_timed_out` counter ticks, queued packet /// is drained, and an ICMPv6 Destination Unreachable frame is /// emitted via the TUN sender. /// @@ -1483,7 +1982,7 @@ async fn test_check_pending_lookups_default_sequence_unreachable() { assert_eq!( node.metrics().lookup.resp_timed_out.get(), baseline_timed_out + 1, - "final timeout must increment discovery.resp_timed_out" + "final timeout must increment lookup.resp_timed_out" ); // No additional initiate_lookup on the timeout step. assert_eq!( diff --git a/src/node/tests/forwarding.rs b/src/node/tests/forwarding.rs index 0706c737..fe19f636 100644 --- a/src/node/tests/forwarding.rs +++ b/src/node/tests/forwarding.rs @@ -5,11 +5,13 @@ //! multi-hop forwarding through live node topologies. use super::*; +use crate::node::peer_error_budget::PEER_ERROR_BURST; use crate::proto::fsp::wire::{FSP_FLAG_CP, build_fsp_header}; use crate::proto::fsp::{SessionAck, SessionSetup}; use crate::proto::link::SessionDatagram; use crate::proto::stp::TreeCoordinate; use crate::proto::stp::encode_coords; + use spanning_tree::{ TestNode, cleanup_nodes, populate_all_coord_caches, process_available_packets, run_tree_test, verify_tree_convergence, @@ -1195,3 +1197,90 @@ async fn test_coord_cache_warming_short_inner_payload_is_dropped_not_panic() { "24 frames of 39..=46 and 47..=62 outer bytes sum to 1212" ); } + +// --- Emission bounds on induced routing errors --- + +/// A distinct destination per index, standing for the fresh `dest_addr` a +/// flooding sender puts on every datagram to escape the per-destination gate. +fn minted_dest(val: u32) -> NodeAddr { + let mut bytes = [0u8; 16]; + bytes[..4].copy_from_slice(&val.to_le_bytes()); + bytes[15] = 0xfe; + NodeAddr::from_bytes(bytes) +} + +/// Feed one transit datagram whose destination this node cannot route. +async fn inject_unroutable(node: &mut Node, from: &NodeAddr, src: NodeAddr, dest: NodeAddr) { + let dg = SessionDatagram::new(src, dest, vec![0x10, 0x00, 0x00, 0x00]).with_ttl(8); + let encoded = dg.encode(); + node.handle_session_datagram(from, &encoded[1..], false) + .await; +} + +#[tokio::test] +async fn one_link_peer_cannot_induce_unbounded_errors_by_varying_the_destination() { + let mut node = make_node(); + let attacker = make_node_addr(0xAA); + let overshoot = 10u32; + + for i in 0..PEER_ERROR_BURST + overshoot { + // Fresh destination and fresh spoofed source per packet: neither + // address-keyed gate sees a repeat. + inject_unroutable( + &mut node, + &attacker, + minted_dest(i + 1_000_000), + minted_dest(i), + ) + .await; + } + + let errors = &node.metrics().errors; + assert_eq!( + errors.emit_over_dest_interval.get(), + 0, + "the per-destination gate cannot bound a sender that varies the destination" + ); + assert_eq!( + errors.emit_over_peer_budget.get(), + u64::from(overshoot), + "everything past the link peer's burst must be refused" + ); +} + +#[tokio::test] +async fn a_destination_suppressed_error_does_not_spend_the_link_peer_budget() { + let mut node = make_node(); + let peer = make_node_addr(0xAA); + let src = make_node_addr(0x01); + let dest = make_node_addr(0x02); + let injected = PEER_ERROR_BURST * 4; + + for _ in 0..injected { + inject_unroutable(&mut node, &peer, src, dest).await; + } + + let errors = &node.metrics().errors; + assert_eq!( + errors.emit_over_peer_budget.get(), + 0, + "an outage on one destination must not spend the peer's budget for the others" + ); + assert_eq!( + errors.emit_over_dest_interval.get(), + u64::from(injected - 1), + "only the first error for a destination goes out within the interval" + ); +} + +#[tokio::test] +async fn a_single_unroutable_datagram_still_produces_its_error() { + let mut node = make_node(); + let peer = make_node_addr(0xAA); + + inject_unroutable(&mut node, &peer, make_node_addr(0x01), make_node_addr(0x02)).await; + + let errors = &node.metrics().errors; + assert_eq!(errors.emit_over_peer_budget.get(), 0); + assert_eq!(errors.emit_over_dest_interval.get(), 0); +} diff --git a/src/node/tests/handshake.rs b/src/node/tests/handshake.rs index 736d0e98..32c3b7bb 100644 --- a/src/node/tests/handshake.rs +++ b/src/node/tests/handshake.rs @@ -1940,3 +1940,278 @@ async fn a_stale_reverse_address_entry_does_not_hide_a_peer_reachable_by_address returns above the insert that would overwrite the stale entry" ); } + +// ===== Epoch-restart dampening ===== +// +// An epoch-mismatch msg1 is authentic, because the epoch travels inside the +// AEAD, but it is replayable: a captured one stays valid forever and +// accepting it destroys a working peering. Two receiver-local conditions +// gate the teardown, and each of the first two cases below breaks one of +// them. + +/// Install a peering for `initiator` that carries an epoch its genuine msg1 +/// does not, and that has gone `idle_secs` without authenticated inbound +/// traffic. Returns the link the peering is bound to. +fn install_peering_at_a_different_epoch( + node: &mut Node, + initiator: &Node, + transport_id: TransportId, + source_addr: &TransportAddr, + idle_secs: u64, +) -> LinkId { + use crate::peer::ActivePeer; + + let identity = PeerIdentity::from_pubkey_full(initiator.identity().pubkey_full()); + let node_addr = *identity.node_addr(); + let link_id = node.allocate_link_id(); + let authenticated_at = Node::now_ms().saturating_sub(idle_secs * 1000); + let mut peer = ActivePeer::new(identity, link_id, authenticated_at); + peer.set_current_addr(transport_id, source_addr.clone()); + // Anything but the epoch the initiator's msg1 carries, so the msg1 reads + // as a restart. + peer.set_remote_epoch(Some([0xAA; 8])); + node.peers.insert(node_addr, peer); + node.addr_to_link + .insert((transport_id, source_addr.clone()), link_id); + link_id +} + +/// A peering long enough past its last authenticated inbound frame that the +/// liveness gate does not hold the restart back. +const IDLE_SECS: u64 = 60; + +#[tokio::test] +async fn an_epoch_mismatch_msg1_against_a_live_peering_leaves_it_intact() { + let transport_id = TransportId::new(1); + let mut node = make_node(); + let initiator = make_node(); + let initiator_addr = node_addr_of(&initiator); + let source_addr = TransportAddr::from_string("127.0.0.1:41001"); + + // The peering is carrying authenticated traffic: it decrypted a frame a + // moment ago. Under replay that is always the case, because the genuine + // peer is heartbeating. + let peer_link = + install_peering_at_a_different_epoch(&mut node, &initiator, transport_id, &source_addr, 0); + + let bad_state_before = node.stats().handshake.bad_state; + node.handle_msg1(ReceivedPacket::with_timestamp( + transport_id, + source_addr.clone(), + genuine_msg1(&initiator, &node), + Node::now_ms(), + )) + .await; + + let peer = node + .get_peer(&initiator_addr) + .expect("a live peering must survive an epoch-mismatch msg1"); + assert_eq!( + peer.link_id(), + peer_link, + "the peering must be the one that was already established, not a \ + replacement promoted from the msg1" + ); + assert_eq!( + peer.remote_epoch(), + Some([0xAA; 8]), + "the stored epoch must not have moved to the one the msg1 carried" + ); + assert_eq!( + node.connection_count(), + 0, + "the dropped msg1 must leave no connection behind" + ); + assert_eq!( + node.stats().handshake.bad_state - bad_state_before, + 1, + "the drop must be counted" + ); +} + +#[tokio::test] +async fn a_second_epoch_change_inside_the_dampening_interval_leaves_the_peering_intact() { + let transport_id = TransportId::new(1); + let mut node = make_node(); + let initiator = make_node(); + let initiator_addr = node_addr_of(&initiator); + let source_addr = TransportAddr::from_string("127.0.0.1:41002"); + + // First epoch change: the peering is genuinely idle, so it is accepted + // and stamps the dampener. + let first_link = install_peering_at_a_different_epoch( + &mut node, + &initiator, + transport_id, + &source_addr, + IDLE_SECS, + ); + node.handle_msg1(ReceivedPacket::with_timestamp( + transport_id, + source_addr.clone(), + genuine_msg1(&initiator, &node), + Node::now_ms(), + )) + .await; + let promoted = node + .get_peer(&initiator_addr) + .expect("the first epoch change must be accepted"); + assert_ne!( + promoted.link_id(), + first_link, + "the first epoch change must have replaced the peering" + ); + + // The peer moves epoch again straight away. Nothing about the second + // msg1 is distinguishable from the first, which is why the interval, + // not the message, has to be what refuses it. + let second_link = install_peering_at_a_different_epoch( + &mut node, + &initiator, + transport_id, + &source_addr, + IDLE_SECS, + ); + + let bad_state_before = node.stats().handshake.bad_state; + node.handle_msg1(ReceivedPacket::with_timestamp( + transport_id, + source_addr.clone(), + genuine_msg1(&initiator, &node), + Node::now_ms(), + )) + .await; + + let peer = node + .get_peer(&initiator_addr) + .expect("a second epoch change inside the interval must not tear the peering down"); + assert_eq!( + peer.link_id(), + second_link, + "the peering must be the one that was already established" + ); + assert_eq!( + peer.remote_epoch(), + Some([0xAA; 8]), + "the stored epoch must not have moved to the one the msg1 carried" + ); + assert_eq!( + node.connection_count(), + 0, + "the dropped msg1 must leave no connection behind" + ); + assert_eq!( + node.stats().handshake.bad_state - bad_state_before, + 1, + "the drop must be counted" + ); +} + +/// A REFUSED epoch-mismatch msg1 must not slide the dampening window. +/// +/// The stamp is written on acceptance only. If it were written on every +/// sighting, a sender replaying a captured msg1 faster than the interval would +/// hold the window open indefinitely and a genuinely restarting peer could +/// never re-peer — the refusal would become the denial of service it exists to +/// prevent. The 15s interval is far longer than a test can wait, so this pins +/// the ordering directly: after an accepted change stamps the dampener, +/// repeated refused msg1s must leave that stamp byte-identical. +/// +/// Discriminator: moving the three stamp lines above the gate in +/// `InboundDecision::RestartThenPromote` reds this and nothing else in the +/// module. +#[tokio::test] +async fn a_refused_epoch_change_does_not_slide_the_dampening_window() { + let transport_id = TransportId::new(1); + let mut node = make_node(); + let initiator = make_node(); + let initiator_addr = node_addr_of(&initiator); + let source_addr = TransportAddr::from_string("127.0.0.1:41009"); + + // Accepted change: stamps the dampener. + install_peering_at_a_different_epoch( + &mut node, + &initiator, + transport_id, + &source_addr, + IDLE_SECS, + ); + node.handle_msg1(ReceivedPacket::with_timestamp( + transport_id, + source_addr.clone(), + genuine_msg1(&initiator, &node), + Node::now_ms(), + )) + .await; + let stamped = node + .restart_dampener_stamp(&initiator_addr) + .expect("an accepted epoch change must stamp the dampener"); + + // Sustained replay: every one of these is refused by the dampener. + for _ in 0..5 { + install_peering_at_a_different_epoch( + &mut node, + &initiator, + transport_id, + &source_addr, + IDLE_SECS, + ); + node.handle_msg1(ReceivedPacket::with_timestamp( + transport_id, + source_addr.clone(), + genuine_msg1(&initiator, &node), + Node::now_ms(), + )) + .await; + } + + assert_eq!( + node.restart_dampener_stamp(&initiator_addr), + Some(stamped), + "a refused epoch change must not restamp the dampener; if it does, a \ + sustained replay holds the window open and starves a real restart" + ); +} + +/// Healthy path, and NOT discriminating: this passes with or without the +/// gates. It is here so that tightening either one, or a bug that stamps the +/// dampener on a refusal, reds the suite instead of silently refusing every +/// genuine restart. +#[tokio::test] +async fn a_first_epoch_change_against_a_silent_peering_still_restarts_it() { + let transport_id = TransportId::new(1); + let mut node = make_node(); + let initiator = make_node(); + let initiator_addr = node_addr_of(&initiator); + let source_addr = TransportAddr::from_string("127.0.0.1:41003"); + + let stale_link = install_peering_at_a_different_epoch( + &mut node, + &initiator, + transport_id, + &source_addr, + IDLE_SECS, + ); + + node.handle_msg1(ReceivedPacket::with_timestamp( + transport_id, + source_addr.clone(), + genuine_msg1(&initiator, &node), + Node::now_ms(), + )) + .await; + + let peer = node + .get_peer(&initiator_addr) + .expect("a restart with no prior epoch change must be promoted"); + assert_ne!( + peer.link_id(), + stale_link, + "the stale peering must have been torn down and replaced" + ); + assert_eq!( + peer.remote_epoch(), + Some(initiator.startup_epoch()), + "the replacement must carry the epoch the msg1 announced" + ); +} diff --git a/src/node/tests/session.rs b/src/node/tests/session.rs index 217dfb61..8e76a541 100644 --- a/src/node/tests/session.rs +++ b/src/node/tests/session.rs @@ -2345,6 +2345,17 @@ fn install_halfopen(node: &mut Node, claimed: NodeAddr) { node.sessions.insert(claimed, entry); } +/// Record that this node put a frame of `wire_len` bytes on the wire toward +/// `dest`, which is what corroborates a reactive `MtuExceeded` reporting a +/// smaller bottleneck. Honest path-MTU discovery produces this by sending; +/// a handler test that installs a session without sending has to state it. +fn note_sent_wire_len(node: &mut Node, dest: &NodeAddr, wire_len: usize) { + node.sessions + .get_mut(dest) + .expect("session must exist to corroborate a report") + .record_sent_wire_len(wire_len); +} + /// Install the entry `initiate_session` creates: an address this node chose /// itself, with the handshake still in flight and MMP not yet initialized. fn install_initiating(node: &mut Node, remote: &Identity) { @@ -2380,6 +2391,7 @@ async fn test_handle_mtu_exceeded_writes_path_mtu_lookup_when_empty() { "lookup should start empty for this destination" ); + note_sent_wire_len(&mut tn.node, &dest, 1400); let inner = build_mtu_exceeded_inner(&dest, &reporter, 1280); tn.node.handle_mtu_exceeded(&reporter, &inner).await; @@ -2406,6 +2418,7 @@ async fn test_handle_mtu_exceeded_tightens_existing_path_mtu_lookup() { // response that didn't reflect the forward-path bottleneck). tn.node.path_mtu_lookup_insert(dest_fips, 1500); + note_sent_wire_len(&mut tn.node, &dest, 1400); let inner = build_mtu_exceeded_inner(&dest, &reporter, 1280); tn.node.handle_mtu_exceeded(&reporter, &inner).await; @@ -2537,8 +2550,9 @@ async fn test_handle_mtu_exceeded_at_the_floor_still_writes_path_mtu_lookup() { let dest = *remote.node_addr(); let reporter = NodeAddr::from_bytes([0xBB; 16]); let dest_fips = crate::FipsAddress::from_node_addr(&dest); - let floor = crate::upper::icmp::MIN_ACTIONABLE_PATH_MTU; + let floor = crate::upper::icmp::MIN_REACTIVE_PATH_MTU; + note_sent_wire_len(&mut tn.node, &dest, 1400); let inner = build_mtu_exceeded_inner(&dest, &reporter, floor); tn.node.handle_mtu_exceeded(&reporter, &inner).await; @@ -2986,6 +3000,7 @@ async fn test_mtu_exceeded_for_a_session_we_initiated_seeds_path_mtu_lookup_befo let reporter = NodeAddr::from_bytes([0xBB; 16]); let dest_fips = crate::FipsAddress::from_node_addr(&dest); + note_sent_wire_len(&mut node, &dest, 1400); let inner = build_mtu_exceeded_inner(&dest, &reporter, 1280); node.handle_mtu_exceeded(&reporter, &inner).await; @@ -3012,6 +3027,7 @@ async fn test_mtu_exceeded_from_a_third_party_forwarder_still_tightens_an_active let reporter = NodeAddr::from_bytes([0xBB; 16]); let dest_fips = crate::FipsAddress::from_node_addr(&dest); + note_sent_wire_len(&mut node, &dest, 1400); let inner = build_mtu_exceeded_inner(&dest, &reporter, 1280); node.handle_mtu_exceeded(&reporter, &inner).await; @@ -4321,6 +4337,258 @@ async fn test_a_drained_stranger_bucket_still_admits_a_setup_naming_an_establish cleanup_nodes(&mut nodes).await; } +// ============================================================================ +// Integration tests: the session-table population cap +// ============================================================================ + +/// Build a two-node routable mesh with the session table capped for a test +/// and the setup limiter opened wide, so the cap is the only thing refusing. +async fn make_session_capped_pair(max_sessions: usize) -> Vec { + let configs = (0..2) + .map(|_| { + let mut config = Config::new(); + config.node.rekey.enabled = false; + config.node.limits.max_sessions = max_sessions; + config.node.rate_limit.session_setup_burst = 10_000; + config.node.rate_limit.session_setup_rate = 10_000.0; + config + }) + .collect(); + let mut nodes = run_tree_test_with_configs(configs, &[(0, 1)]).await; + verify_tree_convergence(&nodes); + populate_all_coord_caches(&mut nodes); + nodes +} + +#[tokio::test] +async fn test_forged_setups_stop_growing_the_session_table_once_the_cap_is_reached() { + // The table was the one remotely-grown map with no bound: each setup from + // an address nobody has seen inserted an entry, and neither existing limit + // reached it, the setup limiter governing arrival rate rather than + // population and the idle purge only reaching entries a peer stops using. + const MAX: usize = 8; + let share = MAX / 2; + let mut nodes = make_session_capped_pair(MAX).await; + + for _ in 0..share { + deliver_forged_setup_over_link(&mut nodes).await; + } + assert_eq!( + nodes[1].node.sessions.len(), + share, + "the admissible entries must be admitted, or this test would pass for \ + the wrong reason" + ); + assert_eq!(nodes[1].node.stats().session.half_open_full, 0); + + // Every SessionAck goes out through `send_session_datagram`, the only + // thing bumping this counter on a node with no transit traffic. A refused + // setup must not move it. + let originated = nodes[1].node.metrics().forwarding.originated_packets.get(); + + for _ in 0..4 { + deliver_forged_setup_over_link(&mut nodes).await; + } + + assert_eq!( + nodes[1].node.sessions.len(), + share, + "a table at its bound must stop growing" + ); + assert_eq!( + nodes[1].node.stats().session.half_open_full, + 4, + "each refusal must be counted; the DEBUG line is invisible by default" + ); + assert_eq!( + nodes[1].node.metrics().forwarding.originated_packets.get(), + originated, + "a refused setup must emit nothing at all" + ); + + cleanup_nodes(&mut nodes).await; +} + +#[tokio::test] +async fn test_a_setup_that_would_grow_a_full_table_is_refused_and_counted() { + // The table-full arm specifically: one established entry against a cap of + // one, so the half-open share is not what refuses. + let mut nodes = make_session_capped_pair(1).await; + establish_pair_session(&mut nodes).await; + assert_eq!( + nodes[1].node.sessions.len(), + 1, + "precondition: the table is full with the established peer" + ); + + let originated = nodes[1].node.metrics().forwarding.originated_packets.get(); + deliver_forged_setup_over_link(&mut nodes).await; + + assert_eq!( + nodes[1].node.sessions.len(), + 1, + "a full table must not grow for a stranger" + ); + assert_eq!(nodes[1].node.stats().session.table_full, 1); + assert_eq!( + nodes[1].node.metrics().forwarding.originated_packets.get(), + originated, + "a refused setup must cost no ack" + ); + + cleanup_nodes(&mut nodes).await; +} + +#[tokio::test] +async fn test_a_full_session_table_still_serves_a_setup_naming_an_existing_entry() { + // The guard against writing the cap as "refuse strangers". A setup for an + // entry already present cannot grow the table, and refusing it would break + // the duplicate-ack resend an initiator depends on. + let mut nodes = make_session_capped_pair(1).await; + establish_pair_session(&mut nodes).await; + + let node0_addr = *nodes[0].node.node_addr(); + let node1_addr = *nodes[1].node.node_addr(); + let refused_before = nodes[1].node.stats().session.table_full; + let originated = nodes[1].node.metrics().forwarding.originated_packets.get(); + + // A setup naming the established peer: the shape an inbound rekey has. + let setup = forge_setup_from_stranger(&nodes); + let datagram = SessionDatagram::new(node0_addr, node1_addr, setup).with_ttl(64); + let encoded = datagram.encode(); + nodes[1] + .node + .handle_session_datagram(&node0_addr, &encoded[1..], false) + .await; + + assert_eq!( + nodes[1].node.stats().session.table_full, + refused_before, + "a setup that cannot grow the table must not be refused by the cap" + ); + assert!( + nodes[1].node.metrics().forwarding.originated_packets.get() > originated, + "and it must still be answered" + ); + + cleanup_nodes(&mut nodes).await; +} + +#[tokio::test] +async fn test_a_full_session_table_does_not_evict_an_established_session() { + // Pins refuse-not-evict. The setup that triggers the decision is + // unauthenticated at that point, so evicting would hand a stranger a way + // to tear down a session it has nothing to do with. + let mut nodes = make_session_capped_pair(1).await; + establish_pair_session(&mut nodes).await; + let node0_addr = *nodes[0].node.node_addr(); + + for _ in 0..4 { + deliver_forged_setup_over_link(&mut nodes).await; + } + + assert!( + nodes[1] + .node + .get_session(&node0_addr) + .expect("the established session must survive a flood at the cap") + .is_established(), + "a stranger's setup must never cost an established peer its session" + ); + + cleanup_nodes(&mut nodes).await; +} + +#[tokio::test] +async fn test_the_session_table_admits_again_after_the_handshake_reaper_drains_it() { + // The cap is a ceiling, not a latch: half-open entries are reaped after + // `handshake_timeout_secs` and the room they free must be usable. + const MAX: usize = 8; + let share = MAX / 2; + let mut nodes = make_session_capped_pair(MAX).await; + + for _ in 0..(share + 2) { + deliver_forged_setup_over_link(&mut nodes).await; + } + assert!( + nodes[1].node.stats().session.half_open_full > 0, + "precondition: the table is refusing before the reaper runs" + ); + + let timeout_ms = nodes[1] + .node + .config() + .node + .rate_limit + .handshake_timeout_secs + * 1000; + let now_ms = Node::now_ms(); + nodes[1] + .node + .resend_pending_session_handshakes(now_ms + timeout_ms + 1) + .await; + assert_eq!( + nodes[1].node.sessions.len(), + 0, + "precondition: the reaper freed the half-open entries" + ); + + deliver_forged_setup_over_link(&mut nodes).await; + assert_eq!( + nodes[1].node.sessions.len(), + 1, + "room freed by the reaper must be usable, or the cap is a latch" + ); + + cleanup_nodes(&mut nodes).await; +} + +#[tokio::test] +async fn test_half_open_setups_cannot_consume_more_than_their_share_of_the_table() { + // Half-open entries are unauthenticated and cheap to create, so they are + // held to a share of the table rather than being allowed to fill it and + // deny it to every peer that would complete a handshake. + const MAX: usize = 16; + let share = MAX / 2; + let mut nodes = make_session_capped_pair(MAX).await; + + for _ in 0..(share + 2) { + deliver_forged_setup_over_link(&mut nodes).await; + } + + assert_eq!( + nodes[1].node.sessions.len(), + share, + "half-open entries must stop at their share, well below the table cap" + ); + assert_eq!(nodes[1].node.stats().session.half_open_full, 2); + assert_eq!( + nodes[1].node.stats().session.table_full, + 0, + "the table itself is not full, so the refusals must be attributed to \ + the share rather than to the cap" + ); + + cleanup_nodes(&mut nodes).await; +} + +#[test] +fn test_session_entry_size_stays_within_the_budget_the_cap_is_derived_from() { + // The default `max_sessions` is derived from what one entry costs. + // Measured at 6608 bytes of inline state when the cap was written, plus + // heap for the MMP window and handshake payloads, so 1024 sessions is + // roughly 7 MB. This is what fires if a large field is added later and + // the arithmetic behind that default stops holding. + const BUDGET: usize = 8192; + assert!( + std::mem::size_of::() <= BUDGET, + "SessionEntry is {} bytes, over the {} the max_sessions default \ + assumes; re-derive the default or shrink the entry", + std::mem::size_of::(), + BUDGET + ); +} + // ============================================================================ // Integration tests: a forged SessionAck against an in-flight initiation // ============================================================================ @@ -5090,3 +5358,226 @@ async fn test_peer_restart_reestablishes_through_a_pending_session_that_waited_o cleanup_nodes(&mut nodes).await; } + +// --------------------------------------------------------------------------- +// Reactive MtuExceeded: corroboration against what this node actually sent +// --------------------------------------------------------------------------- + +#[tokio::test] +async fn a_reactive_mtu_exceeded_at_the_floor_no_longer_pins_a_session_this_node_has_not_overfilled() + { + // The defect itself. A report of exactly the floor is a legal value, and + // the admission gate cannot tell an honest forwarder from anyone else, so + // one packet drove a bound session's path MTU to the floor and pinned the + // FipsAddress-keyed entry the SYN-time MSS clamp reads. Nothing this node + // sent could have overflowed a hop at that size, so no honest report of it + // exists. + let mut node = make_node(); + + let remote = Identity::generate(); + install_established_session_with_mmp(&mut node, &remote); + let dest = *remote.node_addr(); + let reporter = NodeAddr::from_bytes([0xBB; 16]); + let dest_fips = crate::FipsAddress::from_node_addr(&dest); + + let before = node + .sessions + .get(&dest) + .and_then(|e| e.mmp()) + .map(|m| m.path_mtu.current_mtu()); + + let inner = + build_mtu_exceeded_inner(&dest, &reporter, crate::upper::icmp::MIN_REACTIVE_PATH_MTU); + node.handle_mtu_exceeded(&reporter, &inner).await; + + assert_eq!( + node.sessions + .get(&dest) + .and_then(|e| e.mmp()) + .map(|m| m.path_mtu.current_mtu()), + before, + "an uncorroborated report must leave the session path MTU alone" + ); + assert_eq!( + node.path_mtu_lookup_get(&dest_fips), + None, + "an uncorroborated report must leave no clamp entry behind" + ); + assert_eq!( + node.metrics().errors.mtu_exceeded_uncorroborated.get(), + 1, + "the refusal must be counted apart from the below-floor refusal" + ); + assert_eq!( + node.metrics().errors.mtu_exceeded_below_floor.get(), + 0, + "the floor is not what refused this; the value is exactly at it" + ); +} + +#[tokio::test] +async fn an_initiating_session_refuses_an_uncorroborated_report_and_accepts_a_corroborated_one() { + // The lookup write is the effect that survives on an initiating session, + // which has no MMP state at all, so this branch needs its own coverage: + // a guard placed on the apply rather than ahead of it would miss it. + let mut node = make_node(); + + let remote = Identity::generate(); + install_initiating(&mut node, &remote); + let dest = *remote.node_addr(); + let reporter = NodeAddr::from_bytes([0xBB; 16]); + let dest_fips = crate::FipsAddress::from_node_addr(&dest); + + let inner = build_mtu_exceeded_inner(&dest, &reporter, 800); + node.handle_mtu_exceeded(&reporter, &inner).await; + assert_eq!( + node.path_mtu_lookup_get(&dest_fips), + None, + "nothing this node sent could have overflowed a hop at 800 bytes" + ); + + // A SessionSetup can itself be the datagram that overflows a hop, so an + // initiating session must still be able to act on a real report. + note_sent_wire_len(&mut node, &dest, 1400); + node.handle_mtu_exceeded(&reporter, &inner).await; + assert_eq!( + node.path_mtu_lookup_get(&dest_fips), + Some(800), + "a report corroborated by an oversized send must still be applied" + ); +} + +#[tokio::test] +async fn a_second_reactive_decrease_needs_its_own_corroborating_send() { + // The evidence is spent on the decrease it vouched for. Otherwise one + // large send early in a session would vouch for every forged report for + // the rest of that session's life. + let mut node = make_node(); + + let remote = Identity::generate(); + install_established_session_with_mmp(&mut node, &remote); + let dest = *remote.node_addr(); + let reporter = NodeAddr::from_bytes([0xBB; 16]); + + note_sent_wire_len(&mut node, &dest, 1400); + let first = build_mtu_exceeded_inner(&dest, &reporter, 1200); + node.handle_mtu_exceeded(&reporter, &first).await; + assert_eq!( + node.sessions + .get(&dest) + .and_then(|e| e.mmp()) + .map(|m| m.path_mtu.current_mtu()), + Some(1200), + "the corroborated first decrease is accepted" + ); + + let second = build_mtu_exceeded_inner(&dest, &reporter, 600); + node.handle_mtu_exceeded(&reporter, &second).await; + assert_eq!( + node.sessions + .get(&dest) + .and_then(|e| e.mmp()) + .map(|m| m.path_mtu.current_mtu()), + Some(1200), + "a further decrease needs evidence of its own" + ); + + // A genuine re-route onto a smaller hop is preceded by a send that hop + // drops, so the honest sequence still converges. + note_sent_wire_len(&mut node, &dest, 900); + node.handle_mtu_exceeded(&reporter, &second).await; + assert_eq!( + node.sessions + .get(&dest) + .and_then(|e| e.mmp()) + .map(|m| m.path_mtu.current_mtu()), + Some(600), + "once this node has again sent something that does not fit, the report applies" + ); +} + +#[tokio::test] +async fn a_corroborated_report_below_the_reactive_floor_is_still_refused() { + // Corroboration and the floor are independent refusals. A hop that really + // is tiny still cannot drive the clamp into the band where the derived + // MSS degenerates. + let mut node = make_node(); + + let remote = Identity::generate(); + install_established_session_with_mmp(&mut node, &remote); + let dest = *remote.node_addr(); + let reporter = NodeAddr::from_bytes([0xBB; 16]); + let dest_fips = crate::FipsAddress::from_node_addr(&dest); + + note_sent_wire_len(&mut node, &dest, 1400); + let inner = build_mtu_exceeded_inner( + &dest, + &reporter, + crate::upper::icmp::MIN_REACTIVE_PATH_MTU - 1, + ); + node.handle_mtu_exceeded(&reporter, &inner).await; + + assert_eq!(node.path_mtu_lookup_get(&dest_fips), None); + assert_eq!(node.metrics().errors.mtu_exceeded_below_floor.get(), 1); + assert_eq!(node.metrics().errors.mtu_exceeded_uncorroborated.get(), 0); +} + +#[tokio::test] +async fn the_authenticated_path_mtu_notification_still_applies_at_the_actionable_floor() { + // The reactive guards must not leak onto the carrier that arrives inside + // an established session on the decrypted path, which is authenticated and + // needs no corroboration. + let mut node = make_node(); + + let remote = Identity::generate(); + install_established_session_with_mmp(&mut node, &remote); + let dest = *remote.node_addr(); + + let floor = crate::upper::icmp::MIN_ACTIONABLE_PATH_MTU; + let body = build_path_mtu_notification_body(floor); + node.handle_session_path_mtu_notification(&dest, &body); + + assert_eq!( + node.sessions + .get(&dest) + .and_then(|e| e.mmp()) + .map(|m| m.path_mtu.current_mtu()), + Some(floor), + "the authenticated carrier still applies a value at the actionable floor" + ); +} + +#[tokio::test] +async fn a_path_broken_flood_releases_the_stored_path_mtu_only_once_per_interval() { + use crate::proto::routing::PathBroken; + + // PathBroken is unauthenticated and its release discards a bottleneck this + // node learned the hard way. Unlimited, the claim can be repeated as fast + // as it can be sent, so a genuinely learned value never survives. + let mut node = make_node(); + + let remote = Identity::generate(); + install_initiating(&mut node, &remote); + let dest = *remote.node_addr(); + let reporter = NodeAddr::from_bytes([0xBB; 16]); + let dest_fips = crate::FipsAddress::from_node_addr(&dest); + + let encoded = PathBroken::new(dest, reporter).encode(); + let inner = &encoded[5..]; + + node.path_mtu_lookup_insert(dest_fips, 700); + node.handle_path_broken(&reporter, inner).await; + assert_eq!( + node.path_mtu_lookup_get(&dest_fips), + None, + "the first PathBroken still releases" + ); + + node.path_mtu_lookup_insert(dest_fips, 700); + node.handle_path_broken(&reporter, inner).await; + assert_eq!( + node.path_mtu_lookup_get(&dest_fips), + Some(700), + "a second release for the same destination inside the interval is refused" + ); +} diff --git a/src/node/tests/unit.rs b/src/node/tests/unit.rs index 8c340994..2c785b12 100644 --- a/src/node/tests/unit.rs +++ b/src/node/tests/unit.rs @@ -2264,6 +2264,48 @@ fn spawn_blackhole_relay() -> String { format!("ws://127.0.0.1:{port}") } +/// The author filter on advert selection must not swallow a genuine eviction. +/// +/// Dropping foreign-authored events narrows what the selection can return, and +/// the eviction arm is guarded on the relays having answered with nothing at +/// all. If that guard is written too broadly it also suppresses the real case +/// this function exists for: the peer withdrew its advert and the cached entry +/// has to go. Discriminator: a seeded cache entry plus relays that return +/// nothing must still come back `Evicted` with the entry gone. +#[tokio::test] +async fn refetch_still_evicts_a_cached_advert_when_the_relays_return_nothing() { + let peer_npub = Identity::generate().npub(); + let mut bootstrap = crate::nostr::NostrRendezvous::new_for_test(); + bootstrap + .set_advert_relays_for_test(vec![spawn_blackhole_relay()]) + .await; + + let endpoint = crate::nostr::OverlayEndpointAdvert { + transport: crate::nostr::OverlayTransportKind::Udp, + addr: "203.0.113.7:2121".to_string(), + }; + let advert = + crate::nostr::NostrRendezvous::cached_advert_for_test(peer_npub.clone(), endpoint, 1_000); + bootstrap + .insert_advert_for_test(peer_npub.clone(), advert) + .await; + + let outcome = bootstrap.refetch_advert_for_stale_check(&peer_npub).await; + + assert_eq!( + outcome, + crate::nostr::NostrRefetchOutcome::Evicted, + "an empty relay answer is still evidence the advert is gone" + ); + assert!( + bootstrap + .cached_created_at_for_test(&peer_npub) + .await + .is_none(), + "the stale entry should have been removed from the cache" + ); +} + /// The per-tick retry loop must not await the pre-dial advert refetch. /// /// `process_pending_retries` runs inline on the node's 1s rx-loop tick. Each @@ -3896,3 +3938,31 @@ fn test_peer_display_name_tracks_alias_change() { peer_identity.short_npub() ); } + +/// The DNS mesh-interface filter is keyed on the device the node actually +/// created, not on the configured name. macOS and FreeBSD hand out utunN and +/// tunN of the kernel's choosing, so a filter keyed on the configured name +/// resolved to nothing there and was permanently off. +#[cfg(unix)] +#[test] +fn mesh_filter_resolves_the_live_tun_device_rather_than_the_configured_name() { + let loopback = if cfg!(target_os = "macos") { + "lo0" + } else { + "lo" + }; + let c_name = std::ffi::CString::new(loopback).unwrap(); + let expected = unsafe { libc::if_nametoindex(c_name.as_ptr()) }; + if expected == 0 { + return; + } + + let mut config = Config::new(); + config.tun.name = Some("fips-absent-dev".to_string()); + let mut node = Node::new(config).unwrap(); + + assert_eq!(node.mesh_ifindex(), None); + + node.tun_name = Some(loopback.to_string()); + assert_eq!(node.mesh_ifindex(), Some(expected)); +} diff --git a/src/noise/replay.rs b/src/noise/replay.rs index 2e4e6443..99217e05 100644 --- a/src/noise/replay.rs +++ b/src/noise/replay.rs @@ -31,6 +31,14 @@ impl ReplayWindow { /// Returns true if the counter is acceptable, false if it should be rejected. /// Does NOT update the window - call `accept` after successful decryption. pub fn check(&self, counter: u64) -> bool { + // The send side refuses to emit u64::MAX (`CipherState::advance_nonce`, + // `NoiseSession::take_send_counter`), so no conforming peer produces it. + // Refusing it here keeps `accept` from pinning `highest` at the ceiling, + // which would wedge the window against every later counter. + if counter == u64::MAX { + return false; + } + if counter > self.highest { // New highest - always acceptable return true; diff --git a/src/noise/tests.rs b/src/noise/tests.rs index 82cd7a93..baabfc04 100644 --- a/src/noise/tests.rs +++ b/src/noise/tests.rs @@ -387,6 +387,41 @@ fn test_replay_window_reset() { assert!(window.check(100)); } +#[test] +fn test_replay_window_max_counter_does_not_wedge_the_window() { + let mut window = ReplayWindow::new(); + + window.accept(100); + // Mirror the real receive path: check, and accept only if the check passed. + if window.check(u64::MAX) { + window.accept(u64::MAX); + } + + // The honest peer's next frame must still be acceptable. + assert!(window.check(101), "ceiling frame wedged the window"); +} + +#[test] +fn test_replay_window_rejects_max_counter() { + let window = ReplayWindow::new(); + assert!(!window.check(u64::MAX)); +} + +#[test] +fn test_replay_window_accepts_highest_counter_an_honest_peer_can_send() { + // take_send_counter refuses u64::MAX, so u64::MAX - 1 is the highest + // counter a conforming peer emits. The ceiling guard must be exactly one + // value wide and leave that one alone. + let mut window = ReplayWindow::new(); + assert!(window.check(u64::MAX - 1)); + + window.accept(u64::MAX - 1); + assert!( + !window.check(u64::MAX - 1), + "replay should still be rejected" + ); +} + #[test] fn test_session_replay_protection() { let keypair1 = generate_keypair(); diff --git a/src/nostr/mod.rs b/src/nostr/mod.rs index b7f5a5d3..23e12314 100644 --- a/src/nostr/mod.rs +++ b/src/nostr/mod.rs @@ -5,6 +5,7 @@ mod handoff; mod offer_admission; mod runtime; mod signal; +mod signal_gate; mod stun; mod traversal; mod traversal_machine; diff --git a/src/nostr/runtime.rs b/src/nostr/runtime.rs index 76b505fa..7903c94f 100644 --- a/src/nostr/runtime.rs +++ b/src/nostr/runtime.rs @@ -1,7 +1,7 @@ use std::collections::{HashMap, HashSet}; use std::net::SocketAddr; use std::sync::Arc; -use std::sync::atomic::{AtomicBool, Ordering}; +use std::sync::atomic::{AtomicBool, AtomicU64, Ordering}; use std::time::{Duration, Instant}; use nostr::nips::nip17; @@ -22,10 +22,11 @@ use super::failure_state::FailureState; use super::handoff::EstablishedTraversal; use super::offer_admission::{AdmissionReject, OfferAdmission}; use super::signal::{ - FreshnessOutcome, SignalEnvelope, build_signal_event, create_traversal_answer, - create_traversal_offer, estimate_clock_skew, unwrap_signal_event, validate_offer_freshness, - validate_traversal_answer_for_offer, + FRESHNESS_SKEW_TOLERANCE_MS, FreshnessOutcome, SignalEnvelope, build_signal_event, + create_traversal_answer, create_traversal_offer, estimate_clock_skew, unwrap_signal_event, + validate_offer_freshness, validate_traversal_answer_for_offer, }; +use super::signal_gate::SignalGate; use super::stun::observe_traversal_addresses; use super::traversal::{ PunchTargetTally, is_doc_ip, is_never_punchable_ip, is_private_ip, nonce, now_ms, @@ -84,7 +85,7 @@ pub(super) fn short_id(id: &str) -> String { /// node collects by default and those refusals are the ones an operator needs /// to see; the routine off-subnet case stays at `debug`. fn log_refusals(tally: &PunchTargetTally, peer: &str, session: &str) { - if tally.offered <= tally.admitted && tally.capped == 0 { + if tally.offered <= tally.admitted && tally.capped == 0 && tally.over_offered == 0 { return; } let sample = tally.sample.as_deref().unwrap_or("-"); @@ -100,6 +101,7 @@ fn log_refusals(tally: &PunchTargetTally, peer: &str, session: &str) { unroutable = tally.unroutable, offsubnet = tally.offsubnet, capped = tally.capped, + over_offered = tally.over_offered, reflexive = %reflexive, sample = %sample, "traversal: punch candidates refused" @@ -115,6 +117,7 @@ fn log_refusals(tally: &PunchTargetTally, peer: &str, session: &str) { unroutable = tally.unroutable, offsubnet = tally.offsubnet, capped = tally.capped, + over_offered = tally.over_offered, reflexive = %reflexive, sample = %sample, "traversal: punch candidates refused" @@ -207,6 +210,9 @@ pub struct NostrRendezvous { traversal: TraversalMachine, pending_answers: Mutex>>>, admission: OfferAdmission, + signal_gate: SignalGate, + /// Inbound traversal signals shed before decryption, since process start. + shed_signals: AtomicU64, event_tx: mpsc::UnboundedSender, event_rx: Mutex>, connect_task: Mutex>>, @@ -346,6 +352,8 @@ impl NostrRendezvous { traversal, pending_answers: Mutex::new(HashMap::new()), admission, + signal_gate: SignalGate::new(Instant::now()), + shed_signals: AtomicU64::new(0), event_tx, event_rx: Mutex::new(event_rx), connect_task: Mutex::new(None), @@ -606,21 +614,18 @@ impl NostrRendezvous { Err(_) => return NostrRefetchOutcome::Skipped, }; - let mut newest: Option<(u64, &Event)> = None; - for ev in events.iter() { - let ts = ev.created_at.as_secs(); - match newest { - Some((cur, _)) if ts <= cur => {} - _ => newest = Some((ts, ev)), + let Some(ev) = Self::newest_event_by_author(events.iter(), target_pubkey) else { + if !events.is_empty() { + // The relays answered, but nothing they returned was signed by + // this peer. That is no evidence of absence, so keep the entry. + return NostrRefetchOutcome::Skipped; } - } - - let Some((relay_created_at, ev)) = newest else { // Absent on relays. Evict any stale cache entry. self.advert.remove(peer_npub); self.failure_state.reset_streak_after_refresh(peer_npub); return NostrRefetchOutcome::Evicted; }; + let relay_created_at = Self::effective_created_at_secs(ev.created_at.as_secs(), now_ms()); match cached_created_at { Some(cached) if relay_created_at <= cached => NostrRefetchOutcome::SameAdvert, @@ -751,7 +756,15 @@ impl NostrRendezvous { Self::parse_overlay_advert_event(&event, &self.config.app) { let endpoints = endpoint_summary(&advert.endpoints); - let created_at = event.created_at.as_secs(); + // Clamped forward to the signal path's skew + // tolerance: an unbounded future `created_at` would + // win every replacement comparison in + // `observe_advert` and buy a proportionally distant + // validity horizon. + let created_at = Self::effective_created_at_secs( + event.created_at.as_secs(), + now_ms(), + ); if self.advert.observe_advert( &author_npub, advert, @@ -774,6 +787,41 @@ impl NostrRendezvous { continue; } + // Ahead of the unwrap, which is two NIP-44 decrypts and a + // signature verify run inline on the single task that also + // routes answers and processes adverts. Nothing about the + // sender is known yet — the outer event is signed by a key + // generated per event — so the allowance is necessarily + // shared and indiscriminate, and the reserve is what keeps + // a flood from also shedding the answers to traversals + // this node started. + let awaiting_answers = match self.pending_answers.try_lock() { + Ok(pending) => !pending.is_empty(), + // Contended rather than known empty, so treat it as + // outstanding: the fail-open direction here spends the + // reserve, it does not shed. + Err(_) => true, + }; + if let Err(shed) = self.signal_gate.admit(awaiting_answers, Instant::now()) { + let total = self.shed_signals.fetch_add(1, Ordering::Relaxed) + 1; + // Debug, not warn, per event: the party that trips this + // is by definition sending faster than the node wants, + // so a record per drop turns the flood into log volume. + // The doubling summary below is the operator's signal. + debug!( + reason = ?shed, + total, + "shed inbound traversal signal before decrypt" + ); + if total.is_power_of_two() { + warn!( + shed = total, + "inbound traversal signals shed before decrypt" + ); + } + continue; + } + let unwrapped = match unwrap_signal_event(&self.keys, &event).await { Ok(unwrapped) => unwrapped, Err(err) => { @@ -1104,6 +1152,10 @@ impl NostrRendezvous { let base_socket = std::net::UdpSocket::bind(("0.0.0.0", 0))?; base_socket.set_nonblocking(true)?; + // This drains every datagram on the traversal socket until the STUN + // deadline, so it must complete before any punch can be in flight: a + // retry or re-observation once punching has started would swallow the + // peer's punch packets. let (reflexive_address, local_addresses, stun_server) = observe_traversal_addresses( &base_socket, &self.config.stun_servers, @@ -1373,6 +1425,10 @@ impl NostrRendezvous { let base_socket = std::net::UdpSocket::bind(("0.0.0.0", 0))?; base_socket.set_nonblocking(true)?; + // This drains every datagram on the traversal socket until the STUN + // deadline, so it must complete before any punch can be in flight: a + // retry or re-observation once punching has started would swallow the + // peer's punch packets. let (reflexive_address, local_addresses, stun_server) = observe_traversal_addresses( &base_socket, &self.config.stun_servers, @@ -1511,15 +1567,16 @@ impl NostrRendezvous { if author_npub != peer_npub { continue; } + let created_at = Self::effective_created_at_secs(event.created_at.as_secs(), now_ms()); let replace = best .as_ref() - .map(|current| event.created_at.as_secs() >= current.created_at) + .map(|current| created_at >= current.created_at) .unwrap_or(true); if replace { best = Some(CachedOverlayAdvert { author_npub, advert, - created_at: event.created_at.as_secs(), + created_at, valid_until_ms, }); } @@ -1589,7 +1646,7 @@ impl NostrRendezvous { return Ok(self.config.dm_relays.clone()); } }; - let newest = events.iter().max_by_key(|event| event.created_at.as_secs()); + let newest = Self::newest_event_by_author(events.iter(), target_pubkey); if let Some(event) = newest { let relays = nip17::extract_relay_list(event) .map(|relay| relay.to_string()) @@ -1693,6 +1750,42 @@ impl NostrRendezvous { self.advert.event_valid_until_ms(event, now_ms()) } + /// Newest event in `events` that was actually signed by `author`. + /// + /// The relay pool verifies each event's signature but does not check a + /// reply against the REQ filter unless `verify_subscriptions` or + /// `ban_relay_on_mismatch` is set, and neither is. A relay may therefore + /// answer an author-filtered request with an event it signed itself, so + /// the author test happens here, before the timestamp contest, not after: + /// a future-dated foreign event must not be able to suppress the genuine + /// one by winning `created_at`. + pub(super) fn newest_event_by_author<'a>( + events: impl Iterator, + author: PublicKey, + ) -> Option<&'a Event> { + events + .filter(|event| event.pubkey == author) + .max_by_key(|event| event.created_at.as_secs()) + } + + /// A peer's advert `created_at`, in seconds, clamped so it can never read + /// more than `FRESHNESS_SKEW_TOLERANCE_MS` ahead of `now_ms`. + /// + /// An unbounded future `created_at` buys a cache entry two things it + /// should not have: a proportionally distant validity horizon, and an + /// unbeatable position in every replacement comparison, so a later genuine + /// advert can never displace it. Clamping rather than rejecting is + /// deliberate: a node whose own clock runs slow reads every peer's honest + /// advert as future-dated, and rejecting would take out Nostr-mediated + /// dialing for every peer at once with nothing but a cache miss to show + /// for it. Raising the tolerance widens the window in which a future-dated + /// advert outranks an honest one; lowering it makes an ordinary clock + /// difference look hostile. + pub(super) fn effective_created_at_secs(created_at_secs: u64, now_ms: u64) -> u64 { + let ceiling_secs = now_ms.saturating_add(FRESHNESS_SKEW_TOLERANCE_MS) / 1000; + created_at_secs.min(ceiling_secs) + } + pub(super) fn compute_advert_valid_until_ms( event: &Event, advert_max_age_ms: u64, @@ -1702,7 +1795,8 @@ impl NostrRendezvous { return None; } - let created_ms = event.created_at.as_secs().saturating_mul(1000); + let created_ms = Self::effective_created_at_secs(event.created_at.as_secs(), now_ms) + .saturating_mul(1000); let created_window_until = created_ms.saturating_add(advert_max_age_ms); if created_window_until <= now_ms { return None; @@ -1852,6 +1946,8 @@ impl NostrRendezvous { traversal, pending_answers: Mutex::new(HashMap::new()), admission, + signal_gate: SignalGate::new(Instant::now()), + shed_signals: AtomicU64::new(0), event_tx, event_rx: Mutex::new(event_rx), connect_task: Mutex::new(None), @@ -1923,6 +2019,12 @@ impl NostrRendezvous { self.advert.insert_fetched(&npub, advert); } + /// The cached `created_at` for `npub`, or `None` when nothing is cached. + /// Lets a unit test observe whether a refetch evicted an entry. + pub(crate) async fn cached_created_at_for_test(&self, npub: &str) -> Option { + self.advert.cached_created_at(npub) + } + /// Queue a bootstrap event directly for lifecycle tests without live relays /// or a running traversal task. pub(crate) fn push_event_for_test(&self, event: BootstrapEvent) { diff --git a/src/nostr/signal_gate.rs b/src/nostr/signal_gate.rs new file mode 100644 index 00000000..95b1ac55 --- /dev/null +++ b/src/nostr/signal_gate.rs @@ -0,0 +1,252 @@ +//! Rate limiting for inbound traversal signals, ahead of any cryptography. +//! +//! The notify loop used to hand every kind-21059 event straight to +//! `unwrap_signal_event`, which is two NIP-44 decrypts and a signature verify, +//! on the single task that also routes answers and processes adverts. Nothing +//! bounded how fast an unauthenticated stranger could schedule that work, and +//! the per-npub offer admission cannot: it keys on the sender's public key, +//! which only exists once the first decrypt has already run. +//! +//! **What can be keyed on, and what cannot.** Before decryption there is no +//! sender identity at all. The outer event is signed by a key generated per +//! event, so bucketing on its author hands an attacker a fresh allowance for +//! free, and `created_at` and the p-tag are equally attacker-chosen. The relay +//! the event arrived over is drawn from our own configured set, but it is not +//! an isolation boundary either: an attacker publishes to the same relays the +//! honest peer does, and a duplicate event is attributed to whichever relay +//! won the delivery race. So the shared allowance here is deliberately a +//! single global bucket, and it is indiscriminate by construction. +//! +//! **What the reserve is for.** An indiscriminate limit sheds our own +//! traversals along with the attacker's, and since the attacker sets the rate, +//! every retry lands in the same shed. The second bucket is drawn only while +//! this node has traversals of its own outstanding, so a flood costs a node +//! its inbound offers, which is irreducible without a pre-decrypt identity, +//! rather than also costing it the answers to offers it sent. +//! +//! **Lock discipline.** `admit` takes `now` as a parameter rather than reading +//! the clock, so the type is testable without sleeping and holds no state that +//! has to be advanced by a timer. It does its whole decision under one +//! `std::sync::Mutex` and never awaits inside it, as `offer_admission` does; +//! moving anything awaited inside that lock would hold it across a decrypt. + +use std::sync::Mutex; +use std::time::Instant; + +/// Sustained inbound traversal signals per second admitted for decryption, +/// across all senders and relays. +/// +/// A traversal exchange is a handful of events (one offer and one answer per +/// attempt), and the signal subscription is opened with `limit(0)` so relays +/// replay no stored backlog, which means there is no legitimate burst larger +/// than the number of peers bootstrapping in the same second. Raising this +/// buys a larger rendezvous hub headroom at the cost of handing an attacker +/// the same multiple of decrypt work; lowering it starts shedding honest +/// signals on a busy node, which shows up in the log as the shed counter +/// rather than as silence. +const SIGNAL_RATE_PER_SEC: f64 = 5.0; + +/// Burst capacity of the shared allowance, in signals. +const SIGNAL_BURST: f64 = 20.0; + +/// Sustained rate of the reserve, drawn only while this node has traversals +/// of its own outstanding. +/// +/// It is sized like the shared allowance rather than smaller because the case +/// it exists for is onboarding fanout: a node that has just sent offers to +/// tens of peers receives their answers back in a burst, and shedding those +/// looks to an operator like relay flakiness. Lowering it re-exposes that +/// case; raising it lets a flood arriving while we happen to be mid-traversal +/// buy more decrypt work than the shared allowance alone would. +const ANSWER_RESERVE_RATE_PER_SEC: f64 = 5.0; + +/// Burst capacity of the reserve, in signals. +const ANSWER_RESERVE_BURST: f64 = 20.0; + +/// Which allowance refused an inbound signal. +/// +/// The two are different operator stories: the first says the node shed a +/// signal while it had nothing of its own in flight, the second says a flood +/// is now deep enough to reach traversals this node started. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub(super) enum SignalShed { + /// The shared allowance is exhausted. + Shared, + /// The shared allowance and the answer reserve are both exhausted. + Reserve, +} + +/// One token bucket, refilled from the caller's clock. +#[derive(Debug)] +struct Bucket { + tokens: f64, + burst: f64, + rate_per_sec: f64, + last: Instant, +} + +impl Bucket { + fn new(rate_per_sec: f64, burst: f64, now: Instant) -> Self { + Self { + tokens: burst, + burst, + rate_per_sec, + last: now, + } + } + + /// Advance the bucket to `now`, capped at its burst. + fn refill(&mut self, now: Instant) { + let elapsed = now.saturating_duration_since(self.last).as_secs_f64(); + if elapsed > 0.0 { + self.tokens = (self.tokens + elapsed * self.rate_per_sec).min(self.burst); + self.last = now; + } + } + + /// Whether a whole token is available, without taking it. + fn ready(&self) -> bool { + self.tokens >= 1.0 + } + + fn take(&mut self) { + self.tokens -= 1.0; + } +} + +/// The pre-decrypt allowance for inbound traversal signals. +pub(super) struct SignalGate { + inner: Mutex, +} + +#[derive(Debug)] +struct Inner { + shared: Bucket, + reserve: Bucket, +} + +impl SignalGate { + /// Build a gate whose buckets start full as of `now`. + pub(super) fn new(now: Instant) -> Self { + Self::with_limits( + SIGNAL_RATE_PER_SEC, + SIGNAL_BURST, + ANSWER_RESERVE_RATE_PER_SEC, + ANSWER_RESERVE_BURST, + now, + ) + } + + fn with_limits( + shared_rate: f64, + shared_burst: f64, + reserve_rate: f64, + reserve_burst: f64, + now: Instant, + ) -> Self { + Self { + inner: Mutex::new(Inner { + shared: Bucket::new(shared_rate, shared_burst, now), + reserve: Bucket::new(reserve_rate, reserve_burst, now), + }), + } + } + + /// Take one signal's worth of allowance, or say which bucket refused it. + /// + /// `awaiting_answers` says whether this node has traversals of its own + /// outstanding; only then may the reserve be drawn. The shared bucket is + /// always tried first, so the reserve is spent only on what the flood + /// would otherwise have shed. + pub(super) fn admit(&self, awaiting_answers: bool, now: Instant) -> Result<(), SignalShed> { + let mut inner = self.inner.lock().expect("signal-gate mutex poisoned"); + inner.shared.refill(now); + if inner.shared.ready() { + inner.shared.take(); + return Ok(()); + } + if !awaiting_answers { + return Err(SignalShed::Shared); + } + inner.reserve.refill(now); + if inner.reserve.ready() { + inner.reserve.take(); + return Ok(()); + } + Err(SignalShed::Reserve) + } +} + +#[cfg(test)] +mod tests { + use super::*; + use std::time::Duration; + + #[test] + fn a_flood_is_admitted_up_to_the_burst_and_shed_after_it() { + let start = Instant::now(); + let gate = SignalGate::new(start); + + let admitted = (0..1000) + .filter(|_| gate.admit(false, start).is_ok()) + .count(); + + assert_eq!(admitted, SIGNAL_BURST as usize); + assert_eq!(gate.admit(false, start), Err(SignalShed::Shared)); + } + + #[test] + fn tokens_refill_so_a_steady_legitimate_signal_rate_is_never_shed() { + let start = Instant::now(); + let gate = SignalGate::new(start); + + // One second per iteration at exactly the sustained rate, for long + // enough that an arithmetic error in the refill drains the bucket. + for second in 0..100u64 { + let now = start + Duration::from_secs(second); + for signal in 0..SIGNAL_RATE_PER_SEC as usize { + assert_eq!( + gate.admit(false, now), + Ok(()), + "signal {signal} of second {second} should be admitted" + ); + } + } + } + + #[test] + fn a_flood_cannot_shed_the_answers_to_traversals_this_node_started() { + let start = Instant::now(); + let gate = SignalGate::new(start); + + // The flood arrives while this node has nothing outstanding, so it + // cannot reach the reserve at all. + for _ in 0..1000 { + let _ = gate.admit(false, start); + } + assert_eq!(gate.admit(false, start), Err(SignalShed::Shared)); + + let admitted = (0..1000) + .filter(|_| gate.admit(true, start).is_ok()) + .count(); + assert_eq!(admitted, ANSWER_RESERVE_BURST as usize); + assert_eq!(gate.admit(true, start), Err(SignalShed::Reserve)); + } + + #[test] + fn the_reserve_is_untouched_while_the_shared_allowance_still_has_tokens() { + let start = Instant::now(); + let gate = SignalGate::new(start); + + // Every one of these is inside the shared burst, so none of them may + // spend the reserve even though the caller is entitled to it. + for _ in 0..SIGNAL_BURST as usize { + assert_eq!(gate.admit(true, start), Ok(())); + } + + let reserve_admitted = (0..1000) + .filter(|_| gate.admit(true, start).is_ok()) + .count(); + assert_eq!(reserve_admitted, ANSWER_RESERVE_BURST as usize); + } +} diff --git a/src/nostr/stun.rs b/src/nostr/stun.rs index 3aa7a14a..777a2ec1 100644 --- a/src/nostr/stun.rs +++ b/src/nostr/stun.rs @@ -80,6 +80,20 @@ pub(super) async fn observe_traversal_addresses( Ok((None, local_addresses, None)) } +/// Send a STUN binding request on `socket` and return the reflexive address +/// the server reports. +/// +/// Only a datagram whose source is exactly the resolved server address is +/// parsed; anything else is consumed and discarded, so an on-path attacker +/// who has seen the transaction id cannot substitute a mapped address by +/// injecting a reply. The comparison is exact because every socket handed to +/// this function is bound to the IPv4 wildcard, so a source can never arrive +/// in v4-mapped IPv6 form. A dual-stack bind would require normalising both +/// sides with `IpAddr::to_canonical()` before comparing. +/// +/// Caller requirement: this drains and discards every datagram arriving on +/// `socket` until the deadline, so it must not be entered on a socket that +/// may concurrently carry other traffic the caller cares about. async fn perform_stun( socket: &std::net::UdpSocket, stun_server: &str, @@ -95,15 +109,32 @@ async fn perform_stun( udp.send_to(&request, addr).await?; let mut buf = [0u8; 2048]; let deadline = tokio::time::Instant::now() + response_timeout; + let mut rejected = 0u64; + let mut last_unexpected = None; loop { let result = tokio::time::timeout_at(deadline, udp.recv_from(&mut buf)).await; - let Ok(Ok((len, _remote))) = result else { + let Ok(Ok((len, remote))) = result else { break; }; + if remote != addr { + rejected += 1; + last_unexpected = Some(remote); + continue; + } if let Some(mapped) = parse_stun_binding_success(&buf[..len], &txn_id) { return Ok(Some(mapped)); } } + // One line per call rather than per datagram: a flooder controls the rate. + if rejected > 0 { + debug!( + stun_server = %stun_server, + expected = %addr, + rejected, + last_unexpected = ?last_unexpected, + "discarded STUN datagrams from unexpected sources" + ); + } Err(BootstrapError::Stun(format!( "timed out waiting for {}", stun_server @@ -521,4 +552,96 @@ mod tests { "2001:db8::1".parse::().unwrap() ))); } + + /// Build a complete Binding Success carrying one XOR-MAPPED-ADDRESS. + fn build_binding_success(mapped: std::net::SocketAddrV4, txn_id: &[u8; 12]) -> Vec { + let mut packet = build_success_header(0, txn_id); + let cookie = STUN_MAGIC_COOKIE.to_be_bytes(); + let octets = mapped.ip().octets(); + let xport = mapped.port() ^ ((STUN_MAGIC_COOKIE >> 16) as u16); + packet.extend_from_slice(&0x0020u16.to_be_bytes()); // XOR-MAPPED-ADDRESS + packet.extend_from_slice(&8u16.to_be_bytes()); + packet.push(0x00); // reserved + packet.push(0x01); // family IPv4 + packet.extend_from_slice(&xport.to_be_bytes()); + for (index, octet) in octets.iter().enumerate() { + packet.push(octet ^ cookie[index]); + } + let body_len = (packet.len() - 20) as u16; + packet[2..4].copy_from_slice(&body_len.to_be_bytes()); + packet + } + + /// Bind a loopback socket suitable for handing to `perform_stun`. + /// + /// `set_nonblocking` is mandatory rather than tidiness: `perform_stun` + /// passes the socket to `tokio::net::UdpSocket::from_std`, which requires + /// a non-blocking socket and does not make one. A blocking socket parks + /// the runtime thread and the deadline never fires. + fn bind_stun_caller() -> std::net::UdpSocket { + let socket = std::net::UdpSocket::bind("127.0.0.1:0").unwrap(); + socket.set_nonblocking(true).unwrap(); + socket + } + + #[tokio::test] + async fn stun_binding_response_from_an_unexpected_source_is_refused() { + let caller = bind_stun_caller(); + let server = std::net::UdpSocket::bind("127.0.0.1:0").unwrap(); + let attacker = std::net::UdpSocket::bind("127.0.0.1:0").unwrap(); + let server_addr = server.local_addr().unwrap(); + + // Stands in for an on-path attacker: it learns the transaction id the + // way a real one would, by reading the request, and answers from its + // own address while the server stays silent. + let forger = std::thread::spawn(move || { + let mut buf = [0u8; 2048]; + let (len, from) = server.recv_from(&mut buf).unwrap(); + assert!(len >= 20); + let mut txn_id = [0u8; 12]; + txn_id.copy_from_slice(&buf[8..20]); + let forged = build_binding_success("203.0.113.7:1".parse().unwrap(), &txn_id); + attacker.send_to(&forged, from).unwrap(); + }); + + let result = super::perform_stun( + &caller, + &server_addr.to_string(), + std::time::Duration::from_millis(300), + ) + .await; + forger.join().unwrap(); + assert!( + result.is_err(), + "a binding success from a host other than the server must not be believed, got {:?}", + result + ); + } + + #[tokio::test] + async fn stun_binding_response_from_the_server_is_accepted() { + let caller = bind_stun_caller(); + let server = std::net::UdpSocket::bind("127.0.0.1:0").unwrap(); + let server_addr = server.local_addr().unwrap(); + + let responder = std::thread::spawn(move || { + let mut buf = [0u8; 2048]; + let (_len, from) = server.recv_from(&mut buf).unwrap(); + let mut txn_id = [0u8; 12]; + txn_id.copy_from_slice(&buf[8..20]); + let reply = build_binding_success("198.51.100.9:4242".parse().unwrap(), &txn_id); + server.send_to(&reply, from).unwrap(); + }); + + let mapped = super::perform_stun( + &caller, + &server_addr.to_string(), + std::time::Duration::from_secs(2), + ) + .await + .unwrap() + .unwrap(); + responder.join().unwrap(); + assert_eq!(mapped.to_string(), "198.51.100.9:4242"); + } } diff --git a/src/nostr/tests.rs b/src/nostr/tests.rs index c7bcdee4..45400636 100644 --- a/src/nostr/tests.rs +++ b/src/nostr/tests.rs @@ -1,6 +1,8 @@ use std::collections::HashSet; use std::net::{IpAddr, SocketAddr}; +use std::time::Duration; +use nostr::nips::nip17; use nostr::prelude::{EventBuilder, Kind, RelayUrl, Tag, Timestamp}; use super::runtime::{ @@ -13,8 +15,9 @@ use super::signal::{ }; use super::stun::{parse_stun_binding_success, parse_stun_url}; use super::traversal::{ - PunchStrategy, build_punch_packet, is_doc_ip, is_never_punchable_ip, is_private_ip, now_ms, - parse_punch_packet, plan_punch_targets, planned_remote_endpoints, session_hash, + PunchStrategy, SourceRank, build_punch_packet, is_doc_ip, is_never_punchable_ip, is_private_ip, + now_ms, parse_punch_packet, plan_punch_targets, planned_remote_endpoints, rank_punch_source, + run_punch_attempt, session_hash, }; use super::traversal_machine::suppress_responder_for_own_initiator; use super::types::BootstrapError; @@ -47,14 +50,29 @@ fn can_reach(local_nat: NatType, remote_nat: NatType) -> bool { } fn signed_overlay_advert_event(created_at_secs: u64, expiration_secs: Option) -> nostr::Event { - let keys = nostr::Keys::generate(); + signed_overlay_advert_event_from(&nostr::Keys::generate(), created_at_secs, expiration_secs) +} + +fn signed_overlay_advert_event_from( + keys: &nostr::Keys, + created_at_secs: u64, + expiration_secs: Option, +) -> nostr::Event { let content = r#"{"identifier":"fips-overlay-v1","version":1,"endpoints":[{"transport":"tcp","addr":"8.8.8.8:443"}]}"#; let mut builder = EventBuilder::new(Kind::Custom(ADVERT_KIND), content) .custom_created_at(Timestamp::from(created_at_secs)); if let Some(expiration_secs) = expiration_secs { builder = builder.tags([Tag::expiration(Timestamp::from(expiration_secs))]); } - builder.sign_with_keys(&keys).unwrap() + builder.sign_with_keys(keys).unwrap() +} + +fn signed_inbox_relay_event(keys: &nostr::Keys, created_at_secs: u64, relay: &str) -> nostr::Event { + EventBuilder::new(Kind::InboxRelays, "") + .tags([Tag::relay(RelayUrl::parse(relay).unwrap())]) + .custom_created_at(Timestamp::from(created_at_secs)) + .sign_with_keys(keys) + .unwrap() } #[test] @@ -197,6 +215,102 @@ fn advert_freshness_rejects_stale_created_at_without_expiration() { assert!(valid_until.is_none()); } +/// A hostile advert relay may answer an author-filtered request with an event +/// it signed itself. Selection has to drop those before the newest-`created_at` +/// contest, or a future-dated foreign advert suppresses the genuine one. +#[test] +fn advert_selection_ignores_events_not_signed_by_the_target_peer() { + let now_secs = Timestamp::now().as_secs(); + let peer_keys = nostr::Keys::generate(); + let hostile_keys = nostr::Keys::generate(); + + let hostile = signed_overlay_advert_event_from(&hostile_keys, now_secs + 3_600, None); + let genuine = signed_overlay_advert_event_from(&peer_keys, now_secs.saturating_sub(10), None); + let events = [hostile, genuine]; + + let selected = NostrRendezvous::newest_event_by_author(events.iter(), peer_keys.public_key()) + .expect("the peer's own advert should be selected"); + assert_eq!(selected.pubkey, peer_keys.public_key()); +} + +/// Nothing signed by the peer means nothing to select, even though the relays +/// did answer. The caller reads this as "no evidence", not "withdrawn". +#[test] +fn advert_selection_returns_nothing_when_every_event_is_foreign() { + let now_secs = Timestamp::now().as_secs(); + let peer_keys = nostr::Keys::generate(); + let hostile_keys = nostr::Keys::generate(); + + let events = [signed_overlay_advert_event_from( + &hostile_keys, + now_secs + 3_600, + None, + )]; + + assert!( + NostrRendezvous::newest_event_by_author(events.iter(), peer_keys.public_key()).is_none() + ); +} + +/// The same omission on the inbox-relay lookup steers this node's DM and +/// signal traffic onto relays an attacker chose, so it gets the same filter. +#[test] +fn inbox_relay_selection_ignores_relay_lists_not_signed_by_the_target() { + let now_secs = Timestamp::now().as_secs(); + let peer_keys = nostr::Keys::generate(); + let hostile_keys = nostr::Keys::generate(); + + let events = [ + signed_inbox_relay_event(&hostile_keys, now_secs + 3_600, "wss://hostile.example/"), + signed_inbox_relay_event( + &peer_keys, + now_secs.saturating_sub(10), + "wss://genuine.example/", + ), + ]; + + let selected = NostrRendezvous::newest_event_by_author(events.iter(), peer_keys.public_key()) + .expect("the peer's own relay list should be selected"); + let relays = nip17::extract_relay_list(selected) + .map(|relay| relay.to_string()) + .collect::>(); + assert_eq!(relays, vec!["wss://genuine.example/".to_string()]); +} + +/// A far-future `created_at` must not buy a proportionally distant validity +/// horizon. The window is computed from the clamped timestamp instead, so the +/// entry expires on our clock rather than the publisher's. +#[test] +fn advert_freshness_clamps_created_at_beyond_the_forward_skew_tolerance() { + let now_secs = Timestamp::now().as_secs(); + let event = signed_overlay_advert_event(now_secs + 3_600, None); + let valid_until = + NostrRendezvous::compute_advert_valid_until_ms(&event, 600_000, now_secs * 1000) + .expect("a future-dated advert is still usable, just not for as long"); + assert_eq!(valid_until, (now_secs + 60) * 1000 + 600_000); +} + +/// Pins the forward bound to `FRESHNESS_SKEW_TOLERANCE_MS` exactly, mirroring +/// the signal path: 60s ahead is taken as published, 61s ahead is clamped. +/// This is the healthy-path half; an ordinary clock difference must not cost a +/// legitimate peer anything. +#[test] +fn advert_freshness_at_the_forward_skew_limit_is_untouched_and_one_second_beyond_is_clamped() { + let now_secs = Timestamp::now().as_secs(); + + let at_limit = signed_overlay_advert_event(now_secs + 60, None); + let valid_until = + NostrRendezvous::compute_advert_valid_until_ms(&at_limit, 600_000, now_secs * 1000) + .expect("an advert exactly at the forward tolerance should be accepted as published"); + assert_eq!(valid_until, (now_secs + 60) * 1000 + 600_000); + + let past_limit = signed_overlay_advert_event(now_secs + 61, None); + let clamped = + NostrRendezvous::compute_advert_valid_until_ms(&past_limit, 600_000, now_secs * 1000) + .expect("an advert one second past the tolerance is clamped, not refused"); + assert_eq!(clamped, (now_secs + 60) * 1000 + 600_000); +} + #[test] fn advert_freshness_uses_earliest_expiration_bound() { let now_secs = Timestamp::now().as_secs(); @@ -588,7 +702,11 @@ fn planned_remote_endpoints_bound_an_oversized_list_of_unroutable_candidates() { endpoints, vec!["198.51.100.20:63000".parse::().unwrap()] ); - assert_eq!(tally.unroutable, 300); + // Only the first MAX_OFFERED_CANDIDATES are vetted at all now, so the + // unroutable count is the bound rather than the whole list; the rest are + // recorded as never having been looked at. + assert_eq!(tally.unroutable, 32); + assert_eq!(tally.over_offered, 268); assert_eq!(tally.admitted, 1); assert!(tally.suspicious()); } @@ -597,6 +715,10 @@ fn planned_remote_endpoints_bound_an_oversized_list_of_unroutable_candidates() { /// so the observed reflexive address is itself private. Applying the /24 gate /// to a peer's reflexive address would drop it and remove the only branch /// that works across arbitrary NATs; this test reds if anyone does that. +/// +/// The exemption is conditional on exactly the vantage point this test sets +/// up: our own reflexive address is private here, so it still applies. The +/// two tests below cover the public and absent cases. #[test] fn planned_remote_endpoints_keep_private_reflexive_when_stun_is_on_the_lan() { let (endpoints, _tally) = planned_remote_endpoints( @@ -610,6 +732,88 @@ fn planned_remote_endpoints_keep_private_reflexive_when_stun_is_on_the_lan() { assert!(endpoints.contains(&"192.168.1.20:63000".parse().unwrap())); } +/// A node whose own STUN result is public shares no LAN with a private +/// address, so a peer's private reflexive address is only ever an address of +/// the peer's choosing. Admitting it made the reflexive branch a way to have +/// this node punch inside its own private network; the /24 gate now applies. +#[test] +fn a_peers_private_reflexive_address_is_refused_when_our_own_stun_result_is_public() { + let (endpoints, tally) = planned_remote_endpoints( + &[], + Some(&addr("203.0.113.10", 62000)), + &[], + Some(&addr("192.168.1.20", 63000)), + ) + .expect("endpoint planning should succeed"); + + assert!(endpoints.is_empty()); + assert_eq!(tally.reflexive, Some("off-subnet")); +} + +/// The conditional gate keys on our own reflexive address being private, and +/// a node with no reflexive address at all has to keep behaving as it did: +/// a failed STUN probe must not cost same-LAN peering. +#[test] +fn a_peers_private_reflexive_address_is_kept_when_we_have_no_stun_result_at_all() { + let (endpoints, tally) = planned_remote_endpoints( + &[addr("192.168.1.10", 62000)], + None, + &[], + Some(&addr("192.168.1.20", 63000)), + ) + .expect("endpoint planning should succeed"); + + assert!(endpoints.contains(&"192.168.1.20:63000".parse().unwrap())); + assert_eq!(tally.reflexive, None); +} + +/// Refusing a peer's private reflexive address is now something an honest +/// asymmetric-STUN deployment produces, so it must not warn on its own. The +/// never-routable case above still does. +#[test] +fn an_off_subnet_reflexive_refusal_alone_is_not_suspicious() { + let (_planned, tally) = plan_punch_targets( + &[], + Some(&addr("203.0.113.10", 62000)), + &[addr("203.0.113.5", 63000)], + Some(&addr("192.168.1.20", 63000)), + ); + + assert_eq!(tally.reflexive, Some("off-subnet")); + assert!( + tally.admitted > 0, + "the host-candidate path should still plan" + ); + assert!(!tally.suspicious()); +} + +/// The eight-target cap runs after both planning loops, so it bounds the +/// output and not the work. The discriminating assertion is `unroutable`: +/// vetting every candidate would count all thousand, so a count of exactly +/// `MAX_OFFERED_CANDIDATES` is what proves the excess was never walked. +#[test] +fn an_oversized_candidate_list_is_bounded_before_vetting() { + let mut remotes = Vec::new(); + for index in 0..1000u32 { + remotes.push(addr( + &format!("127.0.0.{}", 1 + (index % 254)), + 63000 + (index % 1000) as u16, + )); + } + + let (_planned, tally) = plan_punch_targets( + &[], + Some(&addr("203.0.113.10", 62000)), + &remotes, + Some(&addr("198.51.100.20", 63000)), + ); + + assert_eq!(tally.offered, 1001); + assert_eq!(tally.unroutable, 32); + assert_eq!(tally.over_offered, 968); + assert!(tally.suspicious()); +} + /// The four refusal classes tell four different operational stories, so a /// change that collapses them into one counter, or that makes the warning /// fire on the benign dual-homed shape, has to red here. @@ -1172,6 +1376,238 @@ async fn signal_events_use_current_timestamps() { assert!(created_at <= after); } +/// These punch tests distinguish a spoofer from a planned target by source +/// **IP**, so each needs its own loopback address. Only Linux treats the whole +/// of 127/8 as local; macOS and Windows bind 127.0.0.1 alone unless an alias is +/// added, so the bind panics there. They are gated to Linux rather than +/// rewritten onto one address, because collapsing them onto 127.0.0.1 would +/// make every source rank `RemappedPort` and the tests would stop testing what +/// they are for. +/// +/// **Coverage gap**: on macOS and Windows nothing exercises `run_punch_attempt` +/// end to end. The ranking decision itself is covered on every platform by the +/// `rank_punch_source_*` unit tests above, which take no sockets. +/// A loopback socket bound on `host`, non-blocking as both production call +/// sites leave it, since `run_punch_attempt` hands it straight to +/// `UdpSocket::from_std`. +#[cfg(target_os = "linux")] +fn punch_socket(host: &str) -> std::net::UdpSocket { + let socket = std::net::UdpSocket::bind(format!("{host}:0")).expect("bind a loopback socket"); + socket + .set_nonblocking(true) + .expect("the punch socket must be non-blocking"); + socket +} + +/// A hint that starts punching immediately. `start_at_ms` is absolute wall +/// clock, so anything plausible-looking in the future would sleep out the test. +fn immediate_punch_hint(duration_ms: u64) -> PunchHint { + PunchHint { + start_at_ms: 0, + interval_ms: 20, + duration_ms, + } +} + +/// Send one well-formed probe carrying `session_id`'s hash from `from` to +/// `to`, which is what a replay of captured punch bytes looks like. +#[cfg(target_os = "linux")] +fn send_probe(from: &std::net::UdpSocket, to: SocketAddr, session_id: &str) { + let packet = build_punch_packet(PunchPacketKind::Probe, 1, session_id); + from.send_to(&packet, to).expect("probe should send"); +} + +/// Whether anything readable on `socket` is a punch ack. +#[cfg(target_os = "linux")] +fn received_an_ack(socket: &std::net::UdpSocket) -> bool { + let mut buf = [0u8; 2048]; + while let Ok((len, _)) = socket.recv_from(&mut buf) { + if parse_punch_packet(&buf[..len]) + .map(|packet| packet.kind == PunchPacketKind::Ack) + .unwrap_or(false) + { + return true; + } + } + false +} + +#[test] +fn rank_punch_source_accepts_a_planned_target() { + let target: SocketAddr = "198.51.100.20:63000".parse().unwrap(); + assert_eq!(rank_punch_source(target, &[target]), SourceRank::Planned); +} + +#[test] +fn rank_punch_source_reports_a_planned_targets_other_port_as_remapped() { + let target: SocketAddr = "198.51.100.20:63000".parse().unwrap(); + let remapped: SocketAddr = "198.51.100.20:41234".parse().unwrap(); + assert_eq!( + rank_punch_source(remapped, &[target]), + SourceRank::RemappedPort + ); +} + +#[test] +fn rank_punch_source_rejects_an_address_we_never_planned_to_probe() { + let target: SocketAddr = "198.51.100.20:63000".parse().unwrap(); + let stranger: SocketAddr = "203.0.113.9:63000".parse().unwrap(); + assert_eq!( + rank_punch_source(stranger, &[target]), + SourceRank::Unplanned + ); +} + +/// The regression test for the defect. The punch packet's discriminator is a +/// digest of a value both peers already know and it travels in the clear in +/// every probe, so anyone who has seen one can replay it. Acceptance is now +/// constrained to the targets this node planned; the spoofer is neither +/// adopted nor acked, and an ack would be a reflection we control. +#[cfg(target_os = "linux")] +#[tokio::test] +async fn a_matching_punch_packet_from_an_unplanned_source_is_neither_adopted_nor_acked() { + let victim = punch_socket("127.0.0.3"); + let peer = punch_socket("127.0.0.1"); + let spoofer = punch_socket("127.0.0.2"); + let victim_addr = victim.local_addr().expect("victim address"); + let targets = vec![peer.local_addr().expect("peer address")]; + + send_probe(&spoofer, victim_addr, "session-unplanned"); + let result = run_punch_attempt( + &victim, + "session-unplanned", + &targets, + immediate_punch_hint(400), + Duration::from_millis(700), + ) + .await; + + assert!( + matches!(result, Err(BootstrapError::PunchTimeout(_))), + "a spoofed source must not be adopted, got {result:?}" + ); + assert!( + !received_an_ack(&spoofer), + "an unplanned source must not be acked" + ); +} + +/// The spoofer wins the race on arrival order and still loses on address. +#[cfg(target_os = "linux")] +#[tokio::test] +async fn a_planned_source_is_adopted_even_when_a_spoofer_replies_first() { + let victim = punch_socket("127.0.0.3"); + let peer = punch_socket("127.0.0.1"); + let spoofer = punch_socket("127.0.0.2"); + let victim_addr = victim.local_addr().expect("victim address"); + let peer_addr = peer.local_addr().expect("peer address"); + + send_probe(&spoofer, victim_addr, "session-race"); + send_probe(&peer, victim_addr, "session-race"); + let result = run_punch_attempt( + &victim, + "session-race", + &[peer_addr], + immediate_punch_hint(400), + Duration::from_millis(700), + ) + .await; + + assert_eq!( + result.expect("the planned peer should be adopted"), + peer_addr + ); +} + +/// The healthy path, which is the check that the source constraint does not +/// red a legitimately clean run: one probe from the single planned target is +/// adopted immediately and acked. +#[cfg(target_os = "linux")] +#[tokio::test] +async fn the_ordinary_probe_from_a_planned_target_is_still_adopted_and_acked() { + let victim = punch_socket("127.0.0.3"); + let peer = punch_socket("127.0.0.1"); + let victim_addr = victim.local_addr().expect("victim address"); + let peer_addr = peer.local_addr().expect("peer address"); + + send_probe(&peer, victim_addr, "session-healthy"); + let result = run_punch_attempt( + &victim, + "session-healthy", + &[peer_addr], + immediate_punch_hint(400), + Duration::from_millis(700), + ) + .await; + + assert_eq!( + result.expect("the planned peer should be adopted"), + peer_addr + ); + assert!(received_an_ack(&peer), "a planned probe should be acked"); +} + +/// Peer-reflexive discovery: a symmetric NAT allocates a fresh port toward us, +/// so the peer's probe arrives from an address that is not in the plan but +/// shares a planned target's IP. Adopting it is the main class of NAT pairing +/// punching exists to rescue, and this test reds if the rule is ever tightened +/// to exact matching without that being reopened deliberately. +#[cfg(target_os = "linux")] +#[tokio::test] +async fn a_planned_targets_remapped_port_is_adopted_when_that_is_all_that_arrives() { + let victim = punch_socket("127.0.0.3"); + let peer = punch_socket("127.0.0.1"); + let victim_addr = victim.local_addr().expect("victim address"); + let peer_addr = peer.local_addr().expect("peer address"); + // The address the peer's own STUN observation named, before its NAT + // remapped the port: same host, a port nothing is bound to. + let stale_target = SocketAddr::new(peer_addr.ip(), peer_addr.port().wrapping_add(1).max(1)); + + send_probe(&peer, victim_addr, "session-remapped"); + let result = run_punch_attempt( + &victim, + "session-remapped", + &[stale_target], + immediate_punch_hint(400), + Duration::from_millis(2000), + ) + .await; + + assert_eq!( + result.expect("a remapped port on a planned target should be adopted"), + peer_addr + ); +} + +/// An exact match inside the settle window supersedes a remapped one that +/// arrived first, which is what the window is for. +#[cfg(target_os = "linux")] +#[tokio::test] +async fn an_exact_target_supersedes_a_remapped_port_inside_the_settle_window() { + let victim = punch_socket("127.0.0.3"); + let peer = punch_socket("127.0.0.1"); + let neighbour = punch_socket("127.0.0.1"); + let victim_addr = victim.local_addr().expect("victim address"); + let peer_addr = peer.local_addr().expect("peer address"); + + send_probe(&neighbour, victim_addr, "session-settle"); + send_probe(&peer, victim_addr, "session-settle"); + let result = run_punch_attempt( + &victim, + "session-settle", + &[peer_addr], + immediate_punch_hint(400), + Duration::from_millis(2000), + ) + .await; + + assert_eq!( + result.expect("the exact target should win"), + peer_addr, + "an exact match must supersede a source that only shares the IP" + ); +} + fn node_addr(first_byte: u8) -> NodeAddr { let mut bytes = [0u8; 16]; bytes[0] = first_byte; diff --git a/src/nostr/traversal.rs b/src/nostr/traversal.rs index b5309a97..cf1779f9 100644 --- a/src/nostr/traversal.rs +++ b/src/nostr/traversal.rs @@ -3,6 +3,7 @@ use std::sync::Arc; use std::time::{Duration, Instant, SystemTime, UNIX_EPOCH}; use tokio::net::UdpSocket; +use tracing::debug; use super::types::{ BootstrapError, PUNCH_ACK_MAGIC, PUNCH_MAGIC, PunchHint, PunchPacket, PunchPacketKind, @@ -32,6 +33,61 @@ pub(super) enum PunchStrategy { /// 400 packets, about 21 KB on the wire at 52 bytes each for IPv4. const MAX_PUNCH_TARGETS: usize = 8; +/// Upper bound on how many candidates one peer's signal may have vetted. +/// +/// Vetting is linear in this number and the `push_unique` scan that follows +/// is quadratic in the plan it feeds, so an unbounded candidate list lets one +/// signal buy an unbounded amount of our planning work regardless of the +/// eight-target cap, which only applies after both loops have run. Thirty-two +/// is four times `MAX_PUNCH_TARGETS` and four times what the candidate +/// generator produces on the widest host we have seen, so an honest peer +/// never reaches it. Raising it costs planning work per admitted signal; +/// lowering it costs an honest many-homed peer the tail of its candidate +/// list, which the tally records either way. +const MAX_OFFERED_CANDIDATES: usize = 32; + +/// How long the punch loop keeps listening for an exact target match once it +/// has already accepted a planned target's address on a different port. +/// +/// A source that matches a planned target's IP but not its port is what a +/// symmetric NAT's fresh mapping toward us looks like, and it is worth +/// adopting; a source that matches a target exactly is worth more, so the +/// first remapped source does not end the attempt outright. Raising this +/// delays adoption on the remapped path only, never past the attempt timeout; +/// lowering it toward zero makes the first remapped source win. +const PUNCH_SETTLE_MS: u64 = 250; + +/// How much the source address of a punch packet is worth as a peer address. +/// +/// The packet's own discriminator is a plain digest of a value both peers +/// already know, so it proves only that the sender has seen a probe. What the +/// source address is checked against is the target list this node planned, +/// which is the difference between adopting a peer we chose to probe and +/// adopting whoever replayed those bytes first. +#[derive(Debug, Clone, Copy, PartialEq, Eq, PartialOrd, Ord)] +pub(super) enum SourceRank { + /// Not an address we planned to probe, and so not adoptable. + Unplanned, + /// A planned target's address on a different port. + RemappedPort, + /// Exactly a target we planned to probe. + Planned, +} + +/// Rank one punch packet's source address against the targets we planned. +/// +/// `targets` holds at most `MAX_PUNCH_TARGETS` entries, so the scan is +/// bounded by construction. +pub(super) fn rank_punch_source(remote: SocketAddr, targets: &[SocketAddr]) -> SourceRank { + if targets.contains(&remote) { + SourceRank::Planned + } else if targets.iter().any(|target| target.ip() == remote.ip()) { + SourceRank::RemappedPort + } else { + SourceRank::Unplanned + } +} + #[derive(Debug, Clone, PartialEq, Eq)] pub(super) struct PlannedPunchTarget { pub(super) strategy: PunchStrategy, @@ -44,6 +100,12 @@ pub(super) struct PlannedPunchTarget { pub(super) remote_ip: IpAddr, } +/// Whether a candidate's address text parses as a private or unique-local +/// address. +fn is_private_address(candidate: &TraversalAddress) -> bool { + candidate.ip.parse::().is_ok_and(is_private_ip) +} + fn same_subnet_24(left: &TraversalAddress, right: &TraversalAddress) -> bool { let left_parts = left.ip.split('.').collect::>(); let right_parts = right.ip.split('.').collect::>(); @@ -151,6 +213,8 @@ pub(super) struct PunchTargetTally { pub(super) offsubnet: usize, /// Planned targets discarded by the target cap. pub(super) capped: usize, + /// Candidates past `MAX_OFFERED_CANDIDATES` that were never vetted. + pub(super) over_offered: usize, /// The class label that refused the peer's reflexive address, if it was /// refused. Held apart from the candidate counts because losing the /// reflexive branch removes every path that works across arbitrary NATs, @@ -170,12 +234,21 @@ impl PunchTargetTally { /// entirely refused offer is the reflector case itself. An off-subnet-only /// refusal is the ordinary dual-homed shape and is not suspicious. /// - /// A refused reflexive address always counts: the /24 gate does not apply - /// to it, so the only ways it can be refused are the attacker-shaped ones. + /// A refused reflexive address counts unless the class is `OffSubnet`. + /// The /24 gate now applies to a peer's reflexive address whenever our own + /// reflexive address is public, so an off-subnet refusal of it is what an + /// honest peer behind a LAN STUN server produces against a node with a + /// public one. The other three classes still have no honest producer. + /// + /// A candidate list longer than `MAX_OFFERED_CANDIDATES` counts too: the + /// generator tops out near eight, so nothing honest reaches the bound. pub(super) fn suspicious(&self) -> bool { self.unroutable + self.zeroport + self.unparsable + self.capped > 0 + || self.over_offered > 0 || (self.offered > 0 && self.admitted == 0) - || self.reflexive.is_some() + || self + .reflexive + .is_some_and(|label| label != RejectClass::OffSubnet.label()) } /// Record one refused candidate against its class, keeping the first @@ -212,10 +285,13 @@ impl PunchTargetTally { /// /// Returns the parsed address, or the class of the check that refused it. /// `lan_refs` are our own addresses that a private candidate must share a /24 -/// with. `apply_private_gate` is false for the peer's reflexive address: a -/// STUN server inside the private network legitimately reports a private -/// reflexive address, and dropping it would remove the only branch that works -/// across arbitrary NATs. +/// with. `apply_private_gate` is conditionally false for the peer's reflexive +/// address: a STUN server inside the private network legitimately reports a +/// private reflexive address, and dropping it would remove the only branch +/// that works across arbitrary NATs. That exemption applies only when our own +/// reflexive address is itself private, or absent; a node whose own STUN +/// result is public has no LAN in common with a private reflexive address and +/// would only be punching an address of the peer's choosing. fn admit_remote( candidate: &TraversalAddress, lan_refs: &[TraversalAddress], @@ -261,28 +337,42 @@ pub(super) fn plan_punch_targets( ..PunchTargetTally::default() }; + // Whether our own vantage point is a LAN one: either STUN reported a + // private address for us, or it reported nothing at all. The second case + // is deliberately treated as a LAN vantage point rather than a public one, + // so a node whose STUN probe failed, or that runs without STUN, keeps + // admitting a same-LAN peer's private reflexive address as it always has. + let local_reflexive_on_lan = local_reflexive_address.is_none_or(is_private_address); + // Our own addresses a peer's private candidate has to share a /24 with. // The local reflexive address joins the set when it is itself private, // which is what keeps a LAN-STUN deployment able to match while the // shipped `share_local_candidates=false` leaves the local list empty. let mut lan_refs = local_addresses.to_vec(); if let Some(reflexive) = local_reflexive_address - && reflexive.ip.parse::().is_ok_and(is_private_ip) + && is_private_address(reflexive) { lan_refs.push(reflexive.clone()); } + // A peer names its own candidate list, so bound it before anything walks + // it. The excess is recorded and discarded rather than failing the whole + // offer, which would cost an honest many-homed peer its traversal. + let considered = &remote_addresses[..remote_addresses.len().min(MAX_OFFERED_CANDIDATES)]; + tally.over_offered = remote_addresses.len() - considered.len(); + // Everything on the remote side is peer-supplied, so it is vetted once // here and the branches below only ever see admitted candidates. - let remote_reflexive = - remote_reflexive_address.and_then(|remote| match admit_remote(remote, &lan_refs, false) { + let remote_reflexive = remote_reflexive_address.and_then(|remote| { + match admit_remote(remote, &lan_refs, !local_reflexive_on_lan) { Ok(ip) => Some((remote, ip)), Err(class) => { tally.refuse_reflexive(class, remote); None } - }); - let remote_candidates = remote_addresses + } + }); + let remote_candidates = considered .iter() .filter_map(|remote| match admit_remote(remote, &lan_refs, true) { Ok(ip) => Some((remote, ip)), @@ -393,6 +483,27 @@ pub(super) fn planned_remote_endpoints( Ok((remotes, tally)) } +/// Hold a source that matched a planned target's IP on a different port. +/// +/// That is what a symmetric NAT's fresh mapping toward us looks like, and it +/// is the main class of pairing punching exists to rescue, so it is adopted +/// rather than dropped. It is held for `PUNCH_SETTLE_MS` first so an exact +/// match arriving inside that window supersedes it; the honest path's latency +/// is unchanged, because an exact match breaks the loop immediately. +fn hold_remapped( + remote: SocketAddr, + candidate: &mut Option, + settle_at: &mut Option, + superseded: &mut usize, +) { + if candidate.replace(remote).is_some() { + *superseded += 1; + } + settle_at.get_or_insert_with(|| { + tokio::time::Instant::now() + Duration::from_millis(PUNCH_SETTLE_MS) + }); +} + pub(super) async fn run_punch_attempt( socket: &std::net::UdpSocket, session_id: &str, @@ -427,22 +538,59 @@ pub(super) async fn run_punch_attempt( let expected_hash = session_hash(session_id); let mut buf = [0u8; 2048]; + // Counted rather than logged per packet: an attacker sets how many of + // these arrive, so a record each would trade the adoption this closes for + // log volume. One record at the end of the attempt instead. + let mut unplanned = 0usize; + let mut superseded = 0usize; + let mut candidate: Option = None; + let mut settle_at: Option = None; let result = loop { - let recv = tokio::time::timeout_at(finish_at, udp.recv_from(&mut buf)).await; + let deadline = settle_at.map_or(finish_at, |settle| settle.min(finish_at)); + let recv = tokio::time::timeout_at(deadline, udp.recv_from(&mut buf)).await; let Ok(Ok((len, remote))) = recv else { - break Err(BootstrapError::PunchTimeout(session_id.to_string())); + break match candidate { + Some(remote) => Ok(remote), + None => Err(BootstrapError::PunchTimeout(session_id.to_string())), + }; }; + // Ranked ahead of the ack, not only ahead of the adoption: acking a + // source we never planned to probe is a reflection this node controls, + // and there is no reason to emit it. The packet's own discriminator is + // a digest of a value both peers already know and travels in the clear + // in every probe, so it proves only that the sender saw one. + let rank = rank_punch_source(remote, targets); + if rank == SourceRank::Unplanned { + unplanned += 1; + continue; + } match classify_punch_packet(&buf[..len], expected_hash) { PunchAction::Ignore => continue, PunchAction::Ack { sequence } => { let ack = build_punch_packet(PunchPacketKind::Ack, sequence, session_id); let _ = udp.send_to(&ack, remote).await; - break Ok(remote); + if rank == SourceRank::Planned { + break Ok(remote); + } + hold_remapped(remote, &mut candidate, &mut settle_at, &mut superseded); + } + PunchAction::Matched => { + if rank == SourceRank::Planned { + break Ok(remote); + } + hold_remapped(remote, &mut candidate, &mut settle_at, &mut superseded); } - PunchAction::Matched => break Ok(remote), } }; send_handle.abort(); + if unplanned > 0 || superseded > 0 { + debug!( + session = %super::runtime::short_id(session_id), + unplanned, + superseded, + "traversal: punch packets refused on their source address" + ); + } result } diff --git a/src/proto/lookup/core.rs b/src/proto/lookup/core.rs index 8967707d..6aeedf84 100644 --- a/src/proto/lookup/core.rs +++ b/src/proto/lookup/core.rs @@ -7,7 +7,7 @@ use alloc::sync::Arc; -use super::state::{Lookup, PendingLookup, RecentRequest}; +use super::state::{Lookup, PendingLookup}; use super::wire::LookupRequest; use crate::NodeAddr; @@ -148,8 +148,6 @@ pub(crate) fn plan_initiate(request: &LookupRequest, rv: &impl RoutingView) -> V pub(crate) enum RequestOutcome { /// request_id already in the dedup cache — drop. Duplicate, - /// dedup cache at capacity — drop. `len` is the current cache size (for the log). - DedupCacheFull { len: usize }, /// We are the lookup target — the shell generates + sends the response. RespondAsTarget, /// Forward the request onward (the shell calls the forward planner). @@ -160,10 +158,84 @@ pub(crate) enum RequestOutcome { TtlExhausted, } +/// One dedup-cache entry dropped to make room for an arriving request. +/// +/// Returned to the shell so it can count and log the eviction; the core does +/// no metrics and no logging itself. +pub(crate) struct Eviction { + /// The `request_id` that was dropped. Its reverse path is gone: a + /// response still in flight for it will be treated as unsolicited. + pub request_id: u64, + /// The link peer charged for the eviction — the one whose oldest entry + /// this was, which is not necessarily the peer being admitted. + pub peer: NodeAddr, + /// The per-peer share in force at the time, for the log line. + pub share: usize, +} + +/// The result of classifying an inbound LookupRequest: the route decision, +/// plus any entry that was evicted to make room for it. +pub(crate) struct Classification { + /// What the shell should do with the request. + pub outcome: RequestOutcome, + /// The entry dropped to admit this request, if one was. + pub evicted: Option, +} + +/// Evict from the dedup cache if admitting one more request would put this +/// peer over its share, or the cache over its capacity. +/// +/// Who pays is the whole point. Over its own share a peer pays for itself, +/// and at global capacity the peer holding the most entries pays, so a light +/// peer's reverse path is never taken to admit a heavy one and extra +/// identities buy a flooder proportionally less. Nothing is evicted while +/// the peer is under its share and the cache is under capacity. +fn make_room( + lookup: &mut Lookup, + from: &NodeAddr, + max_recent: usize, + peer_count: usize, +) -> Option { + let share = Lookup::peer_share(max_recent, peer_count); + let victim = if lookup.peer_entries(from) >= share { + *from + } else if lookup.recent_requests.len() >= max_recent { + // Never charge the arriving peer when it is under its share: charge + // whoever is holding the most. + lookup.heaviest_peer()? + } else { + return None; + }; + let request_id = lookup.evict_oldest_from(&victim)?; + Some(Eviction { + request_id, + peer: victim, + share, + }) +} + /// Classify an inbound LookupRequest against the recent-request dedup cache and /// the transit forward rate limiter. Purges expired dedup entries, records the /// request for reverse-path forwarding on the non-drop paths, and decides the /// route. Pure over Lookup state + node addr + injected clock; no I/O, no view. +/// +/// A full cache evicts rather than refuses. Refusing put the capacity check +/// ahead of the check for whether the request names this node, so one link +/// peer emitting fresh `request_id`s could stop the node answering lookups +/// for itself and stop it carrying anyone else's for as long as it kept the +/// cache full — a denial of exactly the service the cache exists to protect. +/// [`make_room`] charges the eviction to the peer that filled the cache. +/// +/// The loosening this accepts: an evicted `request_id` arriving again inside +/// the dedup window is forwarded a second time rather than recognised as a +/// duplicate, and a response still in flight for it has lost its reverse +/// path. The per-target forward limiter and the request TTL already bound +/// what that second forward can cost, and the alternative — refusing the +/// arrival — is the availability defect above. +/// +/// `peer_count` is the current link-peer count, supplied by the caller: the +/// core is sans-IO and clockless and has no view of the live peer table. +#[allow(clippy::too_many_arguments)] pub(crate) fn classify_request( lookup: &mut Lookup, request: &LookupRequest, @@ -172,28 +244,26 @@ pub(crate) fn classify_request( now_ms: u64, recent_expiry_ms: u64, max_recent: usize, -) -> RequestOutcome { - // Purge expired dedup entries (was purge_expired_requests). - lookup - .recent_requests - .retain(|_, entry| !entry.is_expired(now_ms, recent_expiry_ms)); + peer_count: usize, +) -> Classification { + // Purge expired dedup entries (was purge_expired_requests). Cache and + // per-peer index are purged together, or the eviction policy below reads + // a stale index and charges the wrong peer. + lookup.purge_recent(now_ms, recent_expiry_ms); if lookup.recent_requests.contains_key(&request.request_id) { - return RequestOutcome::Duplicate; - } - if lookup.recent_requests.len() >= max_recent { - return RequestOutcome::DedupCacheFull { - len: lookup.recent_requests.len(), + return Classification { + outcome: RequestOutcome::Duplicate, + evicted: None, }; } - lookup - .recent_requests - .insert(request.request_id, RecentRequest::new(*from, now_ms)); - if request.target == *my_addr { - return RequestOutcome::RespondAsTarget; - } - if request.can_forward() { + let evicted = make_room(lookup, from, max_recent, peer_count); + lookup.record_recent(request.request_id, *from, now_ms); + + let outcome = if request.target == *my_addr { + RequestOutcome::RespondAsTarget + } else if request.can_forward() { if lookup .forward_limiter .should_forward(&request.target, now_ms) @@ -204,7 +274,8 @@ pub(crate) fn classify_request( } } else { RequestOutcome::TtlExhausted - } + }; + Classification { outcome, evicted } } /// How an inbound LookupResponse should be routed, decided from the @@ -217,13 +288,21 @@ pub(crate) enum ResponseRoute { Transit { from_peer: NodeAddr }, /// We originated this request — the shell verifies the proof and caches. Originator, + /// Nobody asked for this: it is neither a request we transited nor an + /// answer to a lookup we have outstanding for its target. Dropped before + /// the identity resolve and the signature verify, so it costs nothing. + Unsolicited, } /// Classify an inbound LookupResponse against the recent-request dedup cache. /// /// Pure decision over `Lookup` state: sets `response_forwarded` when this is /// the first response we transit for the request. No I/O, no view, no metrics. -pub(crate) fn classify_response(lookup: &mut Lookup, request_id: u64) -> ResponseRoute { +pub(crate) fn classify_response( + lookup: &mut Lookup, + request_id: u64, + target: &NodeAddr, +) -> ResponseRoute { match lookup.recent_requests.get_mut(&request_id) { Some(recent) => { if recent.response_forwarded { @@ -235,7 +314,25 @@ pub(crate) fn classify_response(lookup: &mut Lookup, request_id: u64) -> Respons } } } - None => ResponseRoute::Originator, + // Not a request we transited, so it claims to answer one of ours. + // Require that it names a target with a lookup outstanding and carries + // an id issued for it. The id is fresh 64-bit randomness drawn per + // attempt and the target signs over it, so a harvested response is + // bound to the request it answered and cannot be redirected or + // replayed. Replies to earlier attempts of a still-outstanding lookup + // still match, which is the common case on a link whose round trip + // exceeds the first rung of the retry ladder. + None => { + let solicited = lookup + .pending_lookups + .get(target) + .is_some_and(|pending| pending.matches(request_id)); + if solicited { + ResponseRoute::Originator + } else { + ResponseRoute::Unsolicited + } + } } } @@ -398,3 +495,233 @@ pub(crate) fn initiate_failed(lookup: &mut Lookup, dest: &NodeAddr, now_ms: u64) lookup.pending_lookups.remove(dest); lookup.backoff.record_failure(dest, now_ms); } + +#[cfg(test)] +mod dedup_eviction_tests { + //! Capacity policy for the dedup cache. + //! + //! These live beside the policy rather than in `lookup/tests/core.rs` + //! because they are the regression tests for a security finding and read + //! directly against `make_room`'s two branches. + + use super::super::limits::{LookupBackoff, LookupForwardRateLimiter}; + use super::super::state::MIN_RECENT_PER_PEER; + use super::*; + use crate::TreeCoordinate; + use crate::testutil::make_node_addr; + + /// The cache bound used by the behavioural tests. Smaller than the + /// production 4096 so a saturation test stays cheap; the policy is a + /// function of the bound, not of its value. + const CACHE: usize = 128; + + fn empty() -> Lookup { + Lookup::new( + LookupBackoff::default(), + LookupForwardRateLimiter::default(), + ) + } + + fn request(request_id: u64, target: NodeAddr) -> LookupRequest { + let origin = make_node_addr(0xCC); + LookupRequest::new( + request_id, + target, + origin, + TreeCoordinate::root(origin), + 5, + 0, + ) + } + + /// Deliver one transit request from `from`, discarding the route + /// decision: these tests are about which entries survive, and the + /// forward limiter's verdict does not affect what is cached. + fn deliver( + lookup: &mut Lookup, + request_id: u64, + from: &NodeAddr, + my_addr: &NodeAddr, + peer_count: usize, + ) -> Classification { + let target = make_node_addr(0xBB); + classify_request( + lookup, + &request(request_id, target), + from, + my_addr, + 1_000, + 60_000, + CACHE, + peer_count, + ) + } + + #[test] + fn a_peer_over_its_share_evicts_its_own_oldest_and_not_a_light_peers() { + let mut lookup = empty(); + let heavy = make_node_addr(0x01); + let light = make_node_addr(0x02); + let me = make_node_addr(0x99); + // 64 peers over a 128-entry cache: 128/64 = 2, floored to 64. + let peer_count = 64; + let share = Lookup::peer_share(CACHE, peer_count); + + deliver(&mut lookup, 7, &light, &me, peer_count); + + // One request past the share, so the heavy peer pays for its own + // admission rather than the cache paying for it. + for i in 0..=share as u64 { + deliver(&mut lookup, 1_000 + i, &heavy, &me, peer_count); + } + + assert!( + lookup.recent_requests.contains_key(&7), + "a light peer's reverse-path entry must survive a neighbour's flood" + ); + assert!( + !lookup.recent_requests.contains_key(&1_000), + "the flooder's own oldest entry is what pays for its newest" + ); + assert!( + lookup.recent_requests.contains_key(&(1_000 + share as u64)), + "and its newest is admitted rather than dropped" + ); + assert_eq!( + lookup.peer_entries(&heavy), + share, + "the flooder is held at its share" + ); + } + + #[test] + fn the_per_peer_share_never_falls_below_the_floor_however_many_peers() { + // A node with as many peers as cache entries would otherwise give + // each peer a share of one, which no genuine transit burst survives. + assert_eq!(Lookup::peer_share(CACHE, CACHE), MIN_RECENT_PER_PEER); + assert_eq!(Lookup::peer_share(CACHE, usize::MAX), MIN_RECENT_PER_PEER); + // Above the floor the share still tracks the peer count. + assert_eq!(Lookup::peer_share(4096, 8), 512); + // And a zero peer count (no links up yet) must not divide by zero. + assert_eq!(Lookup::peer_share(CACHE, 0), CACHE); + + // Behaviourally: at the floor, entry number 64 costs the peer + // nothing and entry number 65 costs it its oldest. + let mut lookup = empty(); + let peer = make_node_addr(0x01); + let me = make_node_addr(0x99); + let peer_count = usize::MAX; + for i in 0..MIN_RECENT_PER_PEER as u64 { + let evicted = deliver(&mut lookup, i, &peer, &me, peer_count).evicted; + assert!(evicted.is_none(), "nothing is evicted below the floor"); + } + let evicted = deliver(&mut lookup, 999, &peer, &me, peer_count) + .evicted + .expect("the entry past the floor must evict"); + assert_eq!(evicted.request_id, 0, "the peer's oldest entry pays"); + assert_eq!(evicted.peer, peer); + assert_eq!(evicted.share, MIN_RECENT_PER_PEER); + } + + #[test] + fn a_node_whose_dedup_cache_is_saturated_still_answers_a_lookup_for_itself() { + // The availability claim, and the actual finding: the capacity check + // used to sit ahead of the check for whether the request names us, + // so a peer holding the cache full made this node unresolvable. + let mut lookup = empty(); + let flooder = make_node_addr(0x01); + let other = make_node_addr(0x02); + let me = make_node_addr(0x99); + // A single link peer, so its share is the whole cache and it can + // saturate without evicting itself. + let peer_count = 1; + + for i in 0..CACHE as u64 { + deliver(&mut lookup, i, &flooder, &me, peer_count); + } + assert_eq!( + lookup.recent_requests.len(), + CACHE, + "precondition: the cache is full, or the rest observes nothing" + ); + + let classification = classify_request( + &mut lookup, + &request(u64::MAX, me), + &other, + &me, + 1_000, + 60_000, + CACHE, + peer_count, + ); + + assert!( + matches!(classification.outcome, RequestOutcome::RespondAsTarget), + "a saturated cache must not stop the node answering lookups for itself" + ); + let evicted = classification + .evicted + .expect("room must have been made at capacity"); + assert_eq!( + evicted.peer, flooder, + "the peer holding the most entries pays, not the arriving one" + ); + assert_eq!(evicted.request_id, 0, "and it pays with its oldest"); + assert!( + lookup.recent_requests.contains_key(&u64::MAX), + "the arriving request is recorded, so its response can be routed back" + ); + assert_eq!( + lookup.recent_requests.len(), + CACHE, + "the cache stays at its bound" + ); + } + + #[test] + fn a_duplicate_is_not_indexed_twice_and_evicts_nothing() { + // The index is a second container over the same entries, so the + // maintenance risk is drift: everything the policy decides reads it. + let mut lookup = empty(); + let peer = make_node_addr(0x01); + let me = make_node_addr(0x99); + + deliver(&mut lookup, 42, &peer, &me, 1); + let repeat = deliver(&mut lookup, 42, &peer, &me, 1); + + assert!(matches!(repeat.outcome, RequestOutcome::Duplicate)); + assert!(repeat.evicted.is_none(), "a duplicate makes no room"); + assert_eq!(lookup.peer_entries(&peer), 1); + assert_eq!(lookup.recent_requests.len(), 1); + } + + #[test] + fn purging_expired_entries_leaves_the_index_level_with_the_cache() { + let mut lookup = empty(); + let a = make_node_addr(0x01); + let b = make_node_addr(0x02); + let me = make_node_addr(0x99); + + for i in 0..5u64 { + deliver(&mut lookup, i, &a, &me, 2); + } + for i in 100..103u64 { + deliver(&mut lookup, i, &b, &me, 2); + } + let indexed: usize = lookup.recent_by_peer.values().map(|ids| ids.len()).sum(); + assert_eq!( + indexed, + lookup.recent_requests.len(), + "every cached request is indexed exactly once" + ); + + // Entries were stamped at 1_000 with a 60s window; age them out. + lookup.purge_recent(1_000 + 60_000 + 1, 60_000); + assert!(lookup.recent_requests.is_empty()); + assert!( + lookup.recent_by_peer.is_empty(), + "the index must not keep entries the cache no longer holds" + ); + } +} diff --git a/src/proto/lookup/limits.rs b/src/proto/lookup/limits.rs index 249df866..b6fa5edd 100644 --- a/src/proto/lookup/limits.rs +++ b/src/proto/lookup/limits.rs @@ -14,7 +14,7 @@ //! nodes generating fresh request_ids at high rate. use crate::NodeAddr; -use crate::proto::rate_limit::PerAddrRateLimiter; +use crate::proto::rate_limit::{PerAddrRateLimiter, RecordOutcome}; use alloc::collections::BTreeMap; // ============================================================================ @@ -180,7 +180,10 @@ impl LookupForwardRateLimiter { /// Returns true if enough time has passed since the last forward /// for this target. Updates internal state when returning true. pub fn should_forward(&mut self, target: &NodeAddr, now_ms: u64) -> bool { - self.0.check_and_record(target, now_ms) + // A full map admits rather than refuses; for forwarding that is the + // same fail-open direction the routing-error limiter takes, and the + // per-target interval is the only thing lost. + self.0.check_and_record(target, now_ms) != RecordOutcome::Suppress } /// Replace the minimum interval in milliseconds (e.g., set to zero to disable). diff --git a/src/proto/lookup/mod.rs b/src/proto/lookup/mod.rs index 2a5d9053..942e0ed9 100644 --- a/src/proto/lookup/mod.rs +++ b/src/proto/lookup/mod.rs @@ -30,4 +30,8 @@ pub(crate) use limits::{LookupBackoff, LookupForwardRateLimiter, MAX_RECENT_LOOK #[cfg(test)] pub(crate) use state::RecentRequest; pub(crate) use state::{Lookup, PendingLookup}; +// The eviction policy reads the share floor through `Lookup::peer_share`; +// only the node-level regression tests name the constant itself. +#[cfg(test)] +pub(crate) use state::MIN_RECENT_PER_PEER; pub use wire::{LookupRequest, LookupResponse}; diff --git a/src/proto/lookup/state.rs b/src/proto/lookup/state.rs index 45f1c26c..3eee5d55 100644 --- a/src/proto/lookup/state.rs +++ b/src/proto/lookup/state.rs @@ -6,7 +6,8 @@ //! evolve toward a sans-IO core without threading four fields through //! `Node`. -use alloc::collections::BTreeMap; +use alloc::collections::{BTreeMap, VecDeque}; +use alloc::vec::Vec; use super::limits::{LookupBackoff, LookupForwardRateLimiter}; use crate::NodeAddr; @@ -44,6 +45,18 @@ impl RecentRequest { } } +/// How many outstanding `request_id`s one pending lookup remembers. +/// +/// Bounds the per-target correlator at eight u64s. The retry ladder +/// (`node.discovery.attempt_timeouts_secs`) is operator configuration and can +/// be longer than this, so the recorder evicts the oldest id rather than +/// refusing the newest: dropping the newest would discard the id most likely +/// to be answered and fail a healthy lookup. Raising this costs eight bytes +/// per extra attempt on every pending target and widens the set of ids a late +/// response may still match; lowering it means a reply to an early attempt on +/// a long ladder is dropped as unsolicited. +const MAX_RECORDED_IDS: usize = 8; + /// Tracks a pending lookup with retry state. pub struct PendingLookup { /// When the lookup was first initiated. @@ -52,6 +65,11 @@ pub struct PendingLookup { pub last_sent_ms: u64, /// Current attempt number (1 = initial, 2 = first retry, ...). pub attempt: u8, + /// `request_id`s issued for this target, oldest first, capped at + /// [`MAX_RECORDED_IDS`]. A response is only acted on when it carries one + /// of these, which is what makes the accept path solicited. The entry + /// itself is dropped at ladder timeout, so this set needs no expiry. + pub ids: Vec, } impl PendingLookup { @@ -60,15 +78,55 @@ impl PendingLookup { initiated_ms: now_ms, last_sent_ms: now_ms, attempt: 1, + ids: Vec::new(), } } + + /// Remember a `request_id` just put on the wire for this target. + pub fn record(&mut self, request_id: u64) { + if self.ids.contains(&request_id) { + return; + } + if self.ids.len() >= MAX_RECORDED_IDS { + self.ids.remove(0); + } + self.ids.push(request_id); + } + + /// Whether `request_id` is one this node issued for this target. + pub fn matches(&self, request_id: u64) -> bool { + self.ids.contains(&request_id) + } } +/// Floor under one link peer's share of the dedup cache. +/// +/// A peer's share is the cache size divided by the current link-peer count, +/// and this is what stops that share collapsing to nothing on a node with +/// very many links. It is a cap and not a reservation: shares can sum past +/// the cache size, in which case the peer holding the most entries pays for +/// the next admission. Raising it lets one busy neighbour hold more of the +/// cache; lowering it clips a genuine transit burst. +pub(crate) const MIN_RECENT_PER_PEER: usize = 64; + /// Mesh lookup subsystem state. pub(crate) struct Lookup { /// Recent lookup requests (dedup + reverse-path forwarding). /// Maps request_id → RecentRequest. pub(crate) recent_requests: BTreeMap, + /// Arrival-order index over `recent_requests`, partitioned by the link + /// peer each request arrived from. The cache is full-then-evict rather + /// than full-then-refuse, and this is what lets an eviction be charged + /// to the peer that filled the cache instead of to whoever happens to be + /// oldest. `now_ms` is nondecreasing across inserts, so each deque is in + /// arrival order and its front is that peer's oldest entry. + /// + /// The index and the cache are two containers where there was one, so + /// they must be maintained together: every mutation of `recent_requests` + /// goes through [`Lookup::record_recent`], [`Lookup::evict_oldest_from`] + /// or [`Lookup::purge_recent`], which keep the two level. A drifted index + /// evicts the wrong entry, or none at all. + pub(crate) recent_by_peer: BTreeMap>, /// Tracks in-flight lookups. Maps target NodeAddr to the /// initiation timestamp (Unix ms). Prevents duplicate flood queries. pub(crate) pending_lookups: BTreeMap, @@ -87,12 +145,85 @@ impl Lookup { pub(crate) fn new(backoff: LookupBackoff, forward_limiter: LookupForwardRateLimiter) -> Self { Self { recent_requests: BTreeMap::new(), + recent_by_peer: BTreeMap::new(), pending_lookups: BTreeMap::new(), backoff, forward_limiter, } } + /// One link peer's share of the dedup cache at the current peer count. + /// + /// The share tracks the peer count rather than being pinned to a number + /// a many-peer node outgrows, with [`MIN_RECENT_PER_PEER`] as its floor. + /// `peer_count` is supplied by the caller: the core is sans-IO and has no + /// view of the live peer table. + pub(crate) fn peer_share(max_recent: usize, peer_count: usize) -> usize { + (max_recent / peer_count.max(1)).max(MIN_RECENT_PER_PEER) + } + + /// How many cached entries `peer` currently holds. + pub(crate) fn peer_entries(&self, peer: &NodeAddr) -> usize { + self.recent_by_peer.get(peer).map_or(0, VecDeque::len) + } + + /// The link peer holding the most cached entries, if any. + /// + /// Ties resolve to the highest `NodeAddr`, because `max_by_key` keeps the + /// last maximum and the index is ordered. Arbitrary but deterministic: + /// nothing about the policy depends on which of two equally heavy peers + /// pays, only that the choice does not vary run to run. + pub(crate) fn heaviest_peer(&self) -> Option { + self.recent_by_peer + .iter() + .max_by_key(|(_, ids)| ids.len()) + .map(|(peer, _)| *peer) + } + + /// Record a request for dedup and reverse-path forwarding, indexing it + /// under the link peer it arrived from. + /// + /// The caller has already established that `request_id` is not cached; a + /// duplicate must not reach here, or the index would hold it twice. + pub(crate) fn record_recent(&mut self, request_id: u64, from: NodeAddr, now_ms: u64) { + self.recent_requests + .insert(request_id, RecentRequest::new(from, now_ms)); + self.recent_by_peer + .entry(from) + .or_default() + .push_back(request_id); + } + + /// Drop `peer`'s oldest cached entry, returning the evicted `request_id`. + /// + /// Returns `None` when the peer holds nothing, which the eviction policy + /// treats as "no room could be made" rather than as an error. + pub(crate) fn evict_oldest_from(&mut self, peer: &NodeAddr) -> Option { + let ids = self.recent_by_peer.get_mut(peer)?; + let evicted = ids.pop_front()?; + if ids.is_empty() { + self.recent_by_peer.remove(peer); + } + self.recent_requests.remove(&evicted); + Some(evicted) + } + + /// Purge expired dedup entries, from the cache and the index together. + /// + /// The index is rebuilt from what survived rather than aged on its own + /// clock, so the two cannot drift apart: an id is indexed if and only if + /// the cache still holds it, and a peer disappears from the index when + /// its last entry does. + pub(crate) fn purge_recent(&mut self, now_ms: u64, expiry_ms: u64) { + self.recent_requests + .retain(|_, entry| !entry.is_expired(now_ms, expiry_ms)); + let recent = &self.recent_requests; + self.recent_by_peer.retain(|_, ids| { + ids.retain(|id| recent.contains_key(id)); + !ids.is_empty() + }); + } + /// Reset lookup backoff on topology changes. Returns the number of /// entries cleared (0 if already empty) so the shell can log the reset — /// observability stays out of the pure core. diff --git a/src/proto/lookup/tests/core.rs b/src/proto/lookup/tests/core.rs index 9bb2e19a..858f5953 100644 --- a/src/proto/lookup/tests/core.rs +++ b/src/proto/lookup/tests/core.rs @@ -152,7 +152,9 @@ fn classify_response_transit_on_fresh_forwarded_request() { .recent_requests .insert(42, RecentRequest::new(from_peer, 1000)); - match classify_response(&mut lookup, 42) { + // Transit is decided by the dedup record alone, so the target the response + // names plays no part here — pass one no pending lookup mentions. + match classify_response(&mut lookup, 42, &make_node_addr(0xF1)) { ResponseRoute::Transit { from_peer: peer } => assert_eq!(peer, from_peer), _ => panic!("expected Transit"), } @@ -168,21 +170,70 @@ fn classify_response_already_forwarded_on_second_call() { .recent_requests .insert(7, RecentRequest::new(from_peer, 1000)); + let target = make_node_addr(0xF2); assert!(matches!( - classify_response(&mut lookup, 7), + classify_response(&mut lookup, 7, &target), ResponseRoute::Transit { .. } )); assert!(matches!( - classify_response(&mut lookup, 7), + classify_response(&mut lookup, 7, &target), ResponseRoute::AlreadyForwarded )); } #[test] -fn classify_response_originator_when_request_absent() { +fn classify_response_originator_when_a_pending_lookup_issued_the_id() { + // No dedup record, so this is not a response we transit. It counts as ours + // only because the target has a lookup outstanding and that lookup issued + // the very request_id the response carries. + let target = make_node_addr(0x51); let mut lookup = empty_lookup(); + let mut pending = PendingLookup::new(1000); + pending.record(999); + lookup.pending_lookups.insert(target, pending); + assert!(matches!( - classify_response(&mut lookup, 999), + classify_response(&mut lookup, 999, &target), + ResponseRoute::Originator + )); +} + +#[test] +fn classify_response_unsolicited_when_nothing_correlates_the_id() { + // The three ways a response can fail to correlate, each of which must be + // dropped before the identity resolve and the signature verify rather than + // being treated as an answer to something we asked for. + let target = make_node_addr(0x52); + let other = make_node_addr(0x53); + let mut lookup = empty_lookup(); + + // 1. Nothing outstanding at all. + assert!(matches!( + classify_response(&mut lookup, 999, &target), + ResponseRoute::Unsolicited + )); + + // 2. A lookup is outstanding for the target, but it never issued this id: + // a harvested response cannot be replayed against a live lookup. + let mut pending = PendingLookup::new(1000); + pending.record(1); + lookup.pending_lookups.insert(target, pending); + assert!(matches!( + classify_response(&mut lookup, 999, &target), + ResponseRoute::Unsolicited + )); + + // 3. The id was issued, but for a different target: the id is bound to the + // exchange it was drawn for and cannot be redirected onto another name. + assert!(matches!( + classify_response(&mut lookup, 1, &other), + ResponseRoute::Unsolicited + )); + + // The matching pair still classifies as ours, so the checks above are + // discriminating rather than rejecting everything. + assert!(matches!( + classify_response(&mut lookup, 1, &target), ResponseRoute::Originator )); } @@ -365,7 +416,8 @@ fn classify_request_forwards_fresh_and_records_it() { let target = make_node_addr(0xAA); let request = make_request_id(1, target, 3); - let outcome = classify_request(&mut lookup, &request, &from, &my_addr, 1000, 5000, 4096); + let outcome = + classify_request(&mut lookup, &request, &from, &my_addr, 1000, 5000, 4096, 1).outcome; assert!(matches!(outcome, RequestOutcome::Forward)); // Recorded for reverse-path forwarding. assert!(lookup.recent_requests.contains_key(&1)); @@ -381,32 +433,39 @@ fn classify_request_duplicate_on_second_call() { let request = make_request_id(1, target, 3); assert!(matches!( - classify_request(&mut lookup, &request, &from, &my_addr, 1000, 5000, 4096), + classify_request(&mut lookup, &request, &from, &my_addr, 1000, 5000, 4096, 1).outcome, RequestOutcome::Forward )); assert!(matches!( - classify_request(&mut lookup, &request, &from, &my_addr, 1000, 5000, 4096), + classify_request(&mut lookup, &request, &from, &my_addr, 1000, 5000, 4096, 1).outcome, RequestOutcome::Duplicate )); } #[test] -fn classify_request_dedup_cache_full() { +fn classify_request_evicts_rather_than_refusing_a_full_dedup_cache() { + // Regression. The cache-full path used to drop the arriving request, so + // one peer emitting fresh request_ids + // could stop this node forwarding anyone's lookups and answering + // lookups for itself. A full cache now evicts instead, charged to the + // peer holding the most entries. let mut lookup = empty_lookup(); let from = make_node_addr(0x01); let my_addr = make_node_addr(0x99); let target = make_node_addr(0xAA); - // Fill the cache to max_recent with distinct request_ids. + // Fill the cache to max_recent with distinct request_ids, through + // `record_recent` so the per-peer index stays level with the cache. A + // direct `recent_requests.insert` would leave the index short and turn + // the eviction policy into a no-op, so the test would pass without + // exercising it. let max_recent = 3usize; for id in 100..(100 + max_recent as u64) { - lookup - .recent_requests - .insert(id, RecentRequest::new(from, 1000)); + lookup.record_recent(id, from, 1000); } assert_eq!(lookup.recent_requests.len(), max_recent); let request = make_request_id(1, target, 3); - match classify_request( + let classification = classify_request( &mut lookup, &request, &from, @@ -414,12 +473,25 @@ fn classify_request_dedup_cache_full() { 1000, 5000, max_recent, - ) { - RequestOutcome::DedupCacheFull { len } => assert_eq!(len, max_recent), - _ => panic!("expected DedupCacheFull"), - } - // The new request must not have been recorded on the drop path. - assert!(!lookup.recent_requests.contains_key(&1)); + 1, + ); + assert!( + matches!(classification.outcome, RequestOutcome::Forward), + "a full cache must not stop the node forwarding" + ); + let evicted = classification + .evicted + .expect("admitting into a full cache must evict something"); + assert_eq!(evicted.request_id, 100, "the oldest entry is what pays"); + assert_eq!( + evicted.peer, from, + "and the peer that filled the cache is what pays" + ); + // The arriving request is admitted, the evicted one is gone, and the + // cache has not grown past its bound. + assert!(lookup.recent_requests.contains_key(&1)); + assert!(!lookup.recent_requests.contains_key(&100)); + assert_eq!(lookup.recent_requests.len(), max_recent); } #[test] @@ -431,7 +503,7 @@ fn classify_request_respond_as_target() { let request = make_request_id(1, my_addr, 3); assert!(matches!( - classify_request(&mut lookup, &request, &from, &my_addr, 1000, 5000, 4096), + classify_request(&mut lookup, &request, &from, &my_addr, 1000, 5000, 4096, 1).outcome, RequestOutcome::RespondAsTarget )); // Recorded before the target decision. @@ -448,7 +520,7 @@ fn classify_request_ttl_exhausted_for_non_target() { let request = make_request_id(1, target, 0); assert!(matches!( - classify_request(&mut lookup, &request, &from, &my_addr, 1000, 5000, 4096), + classify_request(&mut lookup, &request, &from, &my_addr, 1000, 5000, 4096, 1).outcome, RequestOutcome::TtlExhausted )); } @@ -465,7 +537,7 @@ fn classify_request_forward_rate_limited() { let request = make_request_id(1, target, 3); assert!(matches!( - classify_request(&mut lookup, &request, &from, &my_addr, 1000, 5000, 4096), + classify_request(&mut lookup, &request, &from, &my_addr, 1000, 5000, 4096, 1).outcome, RequestOutcome::ForwardRateLimited )); } @@ -478,12 +550,20 @@ fn classify_request_purges_expired_entries() { let target = make_node_addr(0xAA); // Seed an entry that is expired at now_ms with the given expiry window. // is_expired: now - timestamp > expiry_ms → expired. - lookup - .recent_requests - .insert(55, RecentRequest::new(from, 1000)); + lookup.record_recent(55, from, 1000); // now_ms = 10_000, expiry_ms = 5000 → 9000 > 5000 → expired. let request = make_request_id(1, target, 3); - let outcome = classify_request(&mut lookup, &request, &from, &my_addr, 10_000, 5000, 4096); + let outcome = classify_request( + &mut lookup, + &request, + &from, + &my_addr, + 10_000, + 5000, + 4096, + 1, + ) + .outcome; assert!(matches!(outcome, RequestOutcome::Forward)); // The expired entry (55) must have been purged. assert!(!lookup.recent_requests.contains_key(&55)); diff --git a/src/proto/mmp/receiver.rs b/src/proto/mmp/receiver.rs index 15b9f44e..c4ce7533 100644 --- a/src/proto/mmp/receiver.rs +++ b/src/proto/mmp/receiver.rs @@ -61,7 +61,7 @@ impl GapTracker { fn observe(&mut self, counter: u64) -> u64 { let Some(expected) = self.expected_next else { // First frame: initialize - self.expected_next = Some(counter + 1); + self.expected_next = Some(counter.saturating_add(1)); return 0; }; @@ -90,7 +90,7 @@ impl GapTracker { // Update expected (always advance to counter+1 or keep expected if // this was a late/reordered frame) if counter >= expected { - self.expected_next = Some(counter + 1); + self.expected_next = Some(counter.saturating_add(1)); } lost diff --git a/src/proto/mmp/tests/state.rs b/src/proto/mmp/tests/state.rs index 5a36ed51..34b959e6 100644 --- a/src/proto/mmp/tests/state.rs +++ b/src/proto/mmp/tests/state.rs @@ -750,3 +750,50 @@ fn test_reverse_delivery_rekey_reset() { m.delivery_ratio_reverse ); } + +// The ceiling-counter behaviour the gap tracker gained with the counter +// overflow fix. Driven through `ReceiverState` like the burst tests above, +// because `GapTracker` is private to `receiver.rs`; the discriminator is +// unchanged either way, since pre-fix the unchecked `counter + 1` aborts the +// process under the dev profile's overflow checks rather than failing an +// assertion. The abort is the red, not a harness fault. + +#[test] +fn test_gap_tracker_saturates_on_a_first_frame_at_the_ceiling_counter() { + let mut r = ReceiverState::new(32); + r.record_recv(u64::MAX, 0, 100, false, 0); + // A saturated expectation stops the tracker advancing; a repeat of the + // same counter then takes the in-order branch and reports no burst. + r.record_recv(u64::MAX, 0, 100, false, 0); + let rr = r.build_report(0).unwrap(); + assert_eq!(rr.burst_loss_count, 0); + assert_eq!(rr.max_burst_loss, 0); +} + +#[test] +fn test_gap_tracker_saturates_when_advancing_onto_the_ceiling_counter() { + let mut r = ReceiverState::new(32); + // Prime the tracker so the advance branch is taken rather than the + // first-frame branch. + r.record_recv(1, 0, 100, false, 0); + // Reaching this line at all is the discriminator: pre-fix the unchecked + // `counter + 1` overflows here and aborts the process under the dev + // profile's overflow checks. The jump itself is a legitimate enormous gap + // and is correctly reported as one, so this does NOT assert an empty + // burst — an earlier version of this test did, and was simply wrong. + r.record_recv(u64::MAX, 0, 100, false, 0); + let jump = r.build_report(0).unwrap(); + assert!( + jump.burst_loss_count >= 1, + "the jump to the ceiling is a real gap and should be reported" + ); + + // Saturated: the expectation cannot advance past the ceiling, so a repeat + // of the same counter takes the in-order branch and opens no new burst. + r.record_recv(u64::MAX, 0, 100, false, 0); + let after = r.build_report(0).unwrap(); + assert_eq!( + after.burst_loss_count, 0, + "a saturated expectation must not keep opening bursts" + ); +} diff --git a/src/proto/rate_limit.rs b/src/proto/rate_limit.rs index 2eeb83a7..c9ed1298 100644 --- a/src/proto/rate_limit.rs +++ b/src/proto/rate_limit.rs @@ -5,12 +5,54 @@ use crate::NodeAddr; use alloc::collections::BTreeMap; +/// Maximum number of addresses one limiter remembers at once. +/// +/// A hard ceiling on the map a sender can grow by varying the address it keys +/// on, which for the routing-error limiter is a field the sender picks freely. +/// Raising it costs one `NodeAddr` plus one `u64` per entry and buys interval +/// suppression across more simultaneously-active addresses; lowering it makes +/// [`RecordOutcome::AdmitAtCapacity`] the common case sooner, which weakens the +/// interval gate but never the caller's own budget. +pub(crate) const MAX_ENTRIES: usize = 4096; + +/// Fraction of `max_age_ms` between amortized sweeps. +/// +/// The sweep is a full-map `retain`, so running it on every admit made +/// per-event cost linear in a map the sender sizes. Eight sweeps per entry +/// lifetime keeps expired entries from accumulating without putting the scan on +/// the per-event path. +const SWEEPS_PER_MAX_AGE: u64 = 8; + +/// What the limiter decided about one candidate event. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub(crate) enum RecordOutcome { + /// Admit it; the address was recorded. + Admit, + /// Admit it, but the map was full so the address was not recorded and the + /// interval will not suppress its successor. + /// + /// This gate fails open deliberately. Failing closed would turn a full map + /// into node-wide silence, and the map is fullest exactly during partition + /// healing, when many destinations are legitimately unroutable at once and + /// sources most need the signal. + AdmitAtCapacity, + /// Suppress it; an event for this address occurred within the interval. + Suppress, +} + /// Per-address minimum-interval rate limiter. Tracks the last event time per -/// address and enforces a minimum interval, evicting entries older than a max age. +/// address and enforces a minimum interval, evicting entries older than a max +/// age. The map is bounded and its sweep is amortized; see [`MAX_ENTRIES`] and +/// [`SWEEPS_PER_MAX_AGE`]. pub(crate) struct PerAddrRateLimiter { last: BTreeMap, min_interval_ms: u64, max_age_ms: u64, + last_sweep_ms: u64, + /// Sweeps run since construction. Read by the test that holds the + /// amortization property: a full-map scan per admit is the denial-of- + /// service multiplier this counter exists to catch coming back. + sweeps: u64, } impl PerAddrRateLimiter { @@ -19,20 +61,38 @@ impl PerAddrRateLimiter { last: BTreeMap::new(), min_interval_ms, max_age_ms, + last_sweep_ms: 0, + sweeps: 0, } } - /// Returns true (and records `now_ms`) if enough time has elapsed since the - /// last event for `addr`, or this is the first; false if within the interval. - pub(crate) fn check_and_record(&mut self, addr: &NodeAddr, now_ms: u64) -> bool { - if let Some(&last) = self.last.get(addr) + /// Decide one event for `addr`, recording `now_ms` when it is admitted and + /// there is room. See [`RecordOutcome`] for why a full map admits rather + /// than refuses. + pub(crate) fn check_and_record(&mut self, addr: &NodeAddr, now_ms: u64) -> RecordOutcome { + let known = self.last.get(addr).copied(); + if let Some(last) = known && now_ms.saturating_sub(last) < self.min_interval_ms { - return false; + return RecordOutcome::Suppress; + } + self.maybe_sweep(now_ms); + if known.is_none() && self.last.len() >= MAX_ENTRIES { + return RecordOutcome::AdmitAtCapacity; } self.last.insert(*addr, now_ms); + RecordOutcome::Admit + } + + /// Run the expiry sweep at most once per `max_age_ms / SWEEPS_PER_MAX_AGE`. + fn maybe_sweep(&mut self, now_ms: u64) { + let interval = (self.max_age_ms / SWEEPS_PER_MAX_AGE).max(1); + if now_ms.saturating_sub(self.last_sweep_ms) < interval && self.last_sweep_ms != 0 { + return; + } + self.last_sweep_ms = now_ms; + self.sweeps = self.sweeps.saturating_add(1); self.cleanup(now_ms); - true } pub(crate) fn cleanup(&mut self, now_ms: u64) { @@ -40,6 +100,11 @@ impl PerAddrRateLimiter { .retain(|_, &mut last| now_ms.saturating_sub(last) < self.max_age_ms); } + #[cfg(test)] + pub(crate) fn sweeps(&self) -> u64 { + self.sweeps + } + #[cfg(test)] pub(crate) fn set_interval_ms(&mut self, interval_ms: u64) { self.min_interval_ms = interval_ms; diff --git a/src/proto/routing/core.rs b/src/proto/routing/core.rs index a6ef7fea..f956715b 100644 --- a/src/proto/routing/core.rs +++ b/src/proto/routing/core.rs @@ -14,6 +14,7 @@ //! shell hands over only raw per-peer reads; all routing narrowing and decision //! logic lives here. +use super::limits::LimitVerdict; use super::state::Router; use super::wire::{CoordsRequired, MtuExceeded, PathBroken}; use crate::proto::link::{SessionDatagram, SessionDatagramRef}; @@ -171,8 +172,14 @@ impl Router { /// CoordsRequired. The chosen PDU is wrapped in a fresh SessionDatagram /// addressed back to `toward` (the failed datagram's source) and encoded. /// - /// Returns `None` when the rate-limit gate suppresses the signal (the shell - /// drops silently). On `Some`, the shell resolves the reverse link hop for + /// The returned [`ErrorSynth`] carries the gate's verdict alongside the + /// action, rather than collapsing it to a bool. The shell counts the three + /// verdicts separately: suppression and a full destination map say + /// different things about the node, and an operator cannot tell a genuine + /// outage from a limiter that has stopped limiting without the split. + /// + /// `action` is `None` exactly when the verdict is `Suppress` (the shell + /// drops silently). Otherwise the shell resolves the reverse link hop for /// `toward` and sends — resolving the hop only after this gate preserves /// the pre-refactor ordering (rate-limit before `find_next_hop`'s cache /// touch) and lets the shell distinguish suppression from no-reverse-route @@ -185,21 +192,34 @@ impl Router { rv: &impl RoutingView, now_ms: u64, default_ttl: u8, - ) -> Option { - if !self.error_limiter.should_send(dest, now_ms) { - return None; + ) -> ErrorSynth { + let verdict = self.error_limiter.check(dest, now_ms); + if verdict == LimitVerdict::Suppress { + return ErrorSynth { + verdict, + action: None, + }; } - let error_payload = match rv.cached_coords(dest, now_ms) { - Some(coords) => PathBroken::new(*dest, *my_addr) - .with_last_coords(coords) - .encode(), - None => CoordsRequired::new(*dest, *my_addr).encode(), + // Which of the two signals is emitted still discloses whether this + // node holds coords for the destination, but the coordinates + // themselves are not attached: the error is returned to the datagram's + // own src_addr, which nothing binds to the peer that sent it, so + // attaching them would answer a coordinate-cache read to whoever names + // an address, one entry per packet. The field is optional on the wire + // and no receiver reads it, so this is an emission change only. + let error_payload = if rv.cached_coords(dest, now_ms).is_some() { + PathBroken::new(*dest, *my_addr).encode() + } else { + CoordsRequired::new(*dest, *my_addr).encode() }; let error_dg = SessionDatagram::new(*my_addr, *toward, error_payload).with_ttl(default_ttl); - Some(RouteAction::SendError { - toward: *toward, - bytes: error_dg.encode(), - }) + ErrorSynth { + verdict, + action: Some(RouteAction::SendError { + toward: *toward, + bytes: error_dg.encode(), + }), + } } /// Synthesize an MtuExceeded error signal after a forward send failed with @@ -208,11 +228,12 @@ impl Router { /// carrying `bottleneck_mtu`, wraps it in a fresh SessionDatagram addressed /// back to `toward` (the failed datagram's source), and encodes it. /// - /// Returns `None` when the gate suppresses the signal. On `Some`, the shell - /// resolves the reverse link hop for `toward` and sends — resolving the hop - /// only after this gate preserves the pre-refactor ordering (rate-limit - /// before `find_next_hop`'s cache touch). No coordinate read is involved; - /// unlike routing errors, the PDU is unconditional once the gate passes. + /// The returned [`ErrorSynth`]'s `action` is `None` exactly when the + /// verdict is `Suppress`. Otherwise the shell resolves the reverse link hop + /// for `toward` and sends — resolving the hop only after this gate + /// preserves the pre-refactor ordering (rate-limit before + /// `find_next_hop`'s cache touch). No coordinate read is involved; unlike + /// routing errors, the PDU is unconditional once the gate passes. pub(crate) fn synth_mtu_exceeded( &mut self, dest: &NodeAddr, @@ -221,19 +242,38 @@ impl Router { bottleneck_mtu: u16, now_ms: u64, default_ttl: u8, - ) -> Option { - if !self.error_limiter.should_send(dest, now_ms) { - return None; + ) -> ErrorSynth { + let verdict = self.error_limiter.check(dest, now_ms); + if verdict == LimitVerdict::Suppress { + return ErrorSynth { + verdict, + action: None, + }; } let error_payload = MtuExceeded::new(*dest, *my_addr, bottleneck_mtu).encode(); let error_dg = SessionDatagram::new(*my_addr, *toward, error_payload).with_ttl(default_ttl); - Some(RouteAction::SendError { - toward: *toward, - bytes: error_dg.encode(), - }) + ErrorSynth { + verdict, + action: Some(RouteAction::SendError { + toward: *toward, + bytes: error_dg.encode(), + }), + } } } +/// One candidate error signal, as decided by the per-destination gate. +/// +/// The verdict travels with the action so the shell can count `Suppress` and +/// `AdmitAtCapacity` separately; a bool return collapsed the two admits +/// together and left both counters with no writer. +pub(crate) struct ErrorSynth { + /// What the per-destination interval gate decided. + pub verdict: LimitVerdict, + /// The signal to send. `None` exactly when `verdict` is `Suppress`. + pub action: Option, +} + /// Route class of a transit-forwarded packet, classified from tree /// coordinates at the forwarding decision point. The six variants /// partition `forwarded_packets` exactly. diff --git a/src/proto/routing/limits.rs b/src/proto/routing/limits.rs index 01762d1f..4ece2f52 100644 --- a/src/proto/routing/limits.rs +++ b/src/proto/routing/limits.rs @@ -9,7 +9,7 @@ //! portability and deterministic ordering. use crate::NodeAddr; -use crate::proto::rate_limit::PerAddrRateLimiter; +use crate::proto::rate_limit::{PerAddrRateLimiter, RecordOutcome}; /// Default minimum interval between error signals: 100 ms (max 10 errors/sec /// per destination). @@ -22,6 +22,19 @@ const MAX_AGE_MS: u64 = 10_000; /// /// Tracks the last time a routing error was sent for each destination /// address and enforces a minimum interval to prevent floods. +/// What the limiter decided about one candidate routing-error signal. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum LimitVerdict { + /// Send it; the destination was recorded. + Admit, + /// Send it, but the destination map was full so the destination was not + /// recorded and the interval will not suppress its successor. Emission + /// stays bounded by the per-peer error budget in this state. + AdmitAtCapacity, + /// Suppress it; an error for this destination went out within the interval. + Suppress, +} + pub struct RoutingErrorRateLimiter(PerAddrRateLimiter); impl RoutingErrorRateLimiter { @@ -44,7 +57,22 @@ impl RoutingErrorRateLimiter { /// this destination, or if this is the first error. Updates internal /// state when returning true. pub fn should_send(&mut self, dest_addr: &NodeAddr, now_ms: u64) -> bool { - self.0.check_and_record(dest_addr, now_ms) + self.check(dest_addr, now_ms) != LimitVerdict::Suppress + } + + /// Decide one candidate error signal, distinguishing an ordinary admit + /// from one made only because the destination map was full. + /// + /// The caller needs the difference because the two say different things + /// about the node: `AdmitAtCapacity` means the interval gate is no longer + /// suppressing anything for this destination, and only the per-peer budget + /// is bounding emission. + pub fn check(&mut self, dest_addr: &NodeAddr, now_ms: u64) -> LimitVerdict { + match self.0.check_and_record(dest_addr, now_ms) { + RecordOutcome::Admit => LimitVerdict::Admit, + RecordOutcome::AdmitAtCapacity => LimitVerdict::AdmitAtCapacity, + RecordOutcome::Suppress => LimitVerdict::Suppress, + } } /// Remove entries older than max_age. @@ -57,6 +85,13 @@ impl RoutingErrorRateLimiter { pub fn len(&self) -> usize { self.0.len() } + + /// Sweeps the inner map has run since construction. Read by the test + /// holding the amortization property. + #[cfg(test)] + pub(crate) fn sweeps(&self) -> u64 { + self.0.sweeps() + } } impl Default for RoutingErrorRateLimiter { diff --git a/src/proto/routing/mod.rs b/src/proto/routing/mod.rs index b9fa4bd8..10e8c025 100644 --- a/src/proto/routing/mod.rs +++ b/src/proto/routing/mod.rs @@ -29,7 +29,7 @@ pub(crate) use core::{ DropReason, NextHop, RouteAction, RouteClass, RouteOutcome, RoutingView, classify_forward, select_best_candidate, }; -pub(crate) use limits::RoutingErrorRateLimiter; +pub(crate) use limits::{LimitVerdict, RoutingErrorRateLimiter}; pub(crate) use state::Router; pub use wire::{ COORDS_REQUIRED_SIZE, CoordsRequired, MTU_EXCEEDED_SIZE, MtuExceeded, PathBroken, diff --git a/src/proto/routing/tests/core.rs b/src/proto/routing/tests/core.rs index c6be0dad..f1483f81 100644 --- a/src/proto/routing/tests/core.rs +++ b/src/proto/routing/tests/core.rs @@ -4,7 +4,7 @@ use super::util::{MockPeer, MockRoutingView, make_coords, make_datagram_ref, mak use crate::proto::link::SessionDatagramRef; use crate::proto::routing::RoutingSignalType; use crate::proto::routing::{ - DropReason, RouteAction, RouteOutcome, Router, RoutingView, select_best_candidate, + DropReason, LimitVerdict, RouteAction, RouteOutcome, Router, RoutingView, select_best_candidate, }; use crate::testutil::make_node_addr; use crate::{NodeAddr, TreeCoordinate}; @@ -396,6 +396,7 @@ fn synth_uses_pathbroken_when_coords_cached() { }; let action = router .synth_routing_error(&dest, &source, &my_addr, &rv, 0, 64) + .action .expect("gate passes on first call"); let RouteAction::SendError { toward, .. } = &action; assert_eq!( @@ -418,6 +419,7 @@ fn synth_uses_coords_required_when_not_cached() { let rv = MockRoutingView::new(false); // empty coord table let action = router .synth_routing_error(&dest, &source, &my_addr, &rv, 0, 64) + .action .expect("gate passes on first call"); assert_eq!( error_pdu_type(&action), @@ -434,26 +436,24 @@ fn synth_rate_limit_gate_suppresses_second_call() { let my_addr = make_node_addr(0x10); let rv = MockRoutingView::new(false); // First call for this destination passes the gate. - assert!( - router - .synth_routing_error(&dest, &source, &my_addr, &rv, 0, 64) - .is_some() - ); + let first = router.synth_routing_error(&dest, &source, &my_addr, &rv, 0, 64); + assert!(first.action.is_some()); + assert_eq!(first.verdict, LimitVerdict::Admit); // An immediate second call for the same destination is within the // rate-limit window and is suppressed (no sleeps needed — the two calls // are microseconds apart, well under the 100 ms interval). - assert!( - router - .synth_routing_error(&dest, &source, &my_addr, &rv, 0, 64) - .is_none() + let second = router.synth_routing_error(&dest, &source, &my_addr, &rv, 0, 64); + assert!(second.action.is_none()); + assert_eq!( + second.verdict, + LimitVerdict::Suppress, + "the shell counts emit_over_dest_interval off this verdict" ); // A different destination is independent and still allowed. let other = make_node_addr(0x22); - assert!( - router - .synth_routing_error(&other, &source, &my_addr, &rv, 0, 64) - .is_some() - ); + let third = router.synth_routing_error(&other, &source, &my_addr, &rv, 0, 64); + assert!(third.action.is_some()); + assert_eq!(third.verdict, LimitVerdict::Admit); } /// Extract the bottleneck MTU (trailing u16 LE) from an MtuExceeded action's @@ -474,6 +474,7 @@ fn synth_mtu_exceeded_carries_bottleneck_and_targets_source() { let my_addr = make_node_addr(0x10); let action = router .synth_mtu_exceeded(&dest, &source, &my_addr, 1280, 0, 64) + .action .expect("gate passes on first call"); let RouteAction::SendError { toward, .. } = &action; assert_eq!( @@ -499,22 +500,20 @@ fn synth_mtu_exceeded_rate_limit_gate_suppresses_second_call() { let source = make_node_addr(0x21); let my_addr = make_node_addr(0x10); // First call for this destination passes the gate. - assert!( - router - .synth_mtu_exceeded(&dest, &source, &my_addr, 1280, 0, 64) - .is_some() - ); + let first = router.synth_mtu_exceeded(&dest, &source, &my_addr, 1280, 0, 64); + assert!(first.action.is_some()); + assert_eq!(first.verdict, LimitVerdict::Admit); // Immediate second call for the same destination is suppressed. - assert!( - router - .synth_mtu_exceeded(&dest, &source, &my_addr, 1280, 0, 64) - .is_none() + let second = router.synth_mtu_exceeded(&dest, &source, &my_addr, 1280, 0, 64); + assert!(second.action.is_none()); + assert_eq!( + second.verdict, + LimitVerdict::Suppress, + "the shell counts emit_over_dest_interval off this verdict" ); // A different destination is independent and still allowed. let other = make_node_addr(0x22); - assert!( - router - .synth_mtu_exceeded(&other, &source, &my_addr, 1280, 0, 64) - .is_some() - ); + let third = router.synth_mtu_exceeded(&other, &source, &my_addr, 1280, 0, 64); + assert!(third.action.is_some()); + assert_eq!(third.verdict, LimitVerdict::Admit); } diff --git a/src/proto/routing/tests/limits.rs b/src/proto/routing/tests/limits.rs index d93f41b4..be7206b8 100644 --- a/src/proto/routing/tests/limits.rs +++ b/src/proto/routing/tests/limits.rs @@ -1,8 +1,19 @@ //! Tests for routing error-signal rate limiting. -use crate::proto::routing::RoutingErrorRateLimiter; +use crate::NodeAddr; +use crate::proto::rate_limit::MAX_ENTRIES; +use crate::proto::routing::{LimitVerdict, RoutingErrorRateLimiter}; use crate::testutil::make_node_addr as addr; +/// A distinct destination address per index, standing for the fresh +/// `dest_addr` a flooding sender puts on every datagram. +fn minted_addr(val: u32) -> NodeAddr { + let mut bytes = [0u8; 16]; + bytes[..4].copy_from_slice(&val.to_le_bytes()); + bytes[15] = 0xff; + NodeAddr::from_bytes(bytes) +} + #[test] fn test_first_send_allowed() { let mut limiter = RoutingErrorRateLimiter::new(); @@ -68,3 +79,72 @@ fn test_with_interval_custom_rate() { // Allowed at 500 ms. assert!(limiter.should_send(&addr(1), 500)); } + +#[test] +fn the_map_stays_bounded_when_a_sender_mints_distinct_destination_keys() { + // `dest_addr` is a field the sender picks, so an unbounded map is memory + // the sender sizes. + let mut limiter = RoutingErrorRateLimiter::new(); + + for i in 0..100_000u32 { + limiter.check(&minted_addr(i), 0); + } + + assert!( + limiter.len() <= MAX_ENTRIES, + "limiter held {} entries, above the {MAX_ENTRIES} ceiling", + limiter.len() + ); +} + +#[test] +fn an_admission_at_capacity_still_sends_rather_than_going_silent() { + // The ceiling fails open on purpose: the map is fullest during partition + // healing, when sources most need the signal. + let mut limiter = RoutingErrorRateLimiter::new(); + + for i in 0..MAX_ENTRIES as u32 { + assert_eq!(limiter.check(&minted_addr(i), 0), LimitVerdict::Admit); + } + + // The map is full and nothing in it is old enough to evict, so the next + // distinct destination cannot be recorded. It must still be sent. + assert_eq!( + limiter.check(&minted_addr(MAX_ENTRIES as u32), 0), + LimitVerdict::AdmitAtCapacity + ); +} + +#[test] +fn the_map_scan_does_not_run_once_per_admitted_destination() { + // The sweep is a full-map retain. Running it per admit would make the + // per-event cost linear in a map the sender sizes, which is the denial of + // service the ceiling above exists to prevent. + let mut limiter = RoutingErrorRateLimiter::new(); + limiter.check(&minted_addr(0), 1); + let before = limiter.sweeps(); + + for i in 1..1_000u32 { + limiter.check(&minted_addr(i), 1); + } + + assert_eq!( + limiter.sweeps() - before, + 0, + "the full-map scan ran inside a single sweep interval" + ); +} + +#[test] +fn the_map_scan_still_runs_once_a_sweep_interval_has_passed() { + let mut limiter = RoutingErrorRateLimiter::new(); + limiter.check(&minted_addr(0), 1); + let before = limiter.sweeps(); + + // 11 s later: past both the sweep interval and the 10 s max age. + limiter.check(&minted_addr(1), 11_000); + + assert_eq!(limiter.sweeps() - before, 1); + // The first destination aged out, so the sweep did its job. + assert_eq!(limiter.len(), 1); +} diff --git a/src/transport/ethernet/io.rs b/src/transport/ethernet/io.rs index 392a3b80..c60b37dd 100644 --- a/src/transport/ethernet/io.rs +++ b/src/transport/ethernet/io.rs @@ -21,6 +21,90 @@ mod platform; #[cfg(unix)] pub use platform::PacketSocket; +/// Outcome of `send_frame`. +#[cfg(unix)] +#[cfg_attr(not(target_os = "macos"), allow(dead_code))] +pub(crate) enum SendOutcome { + Sent, + Stop, +} + +/// Retry iterations spent yielding before the send loop starts sleeping. +/// +/// A transiently full channel drains in microseconds, so yielding keeps the +/// saturated-path handoff rate uncapped, which is the whole reason this +/// module has a dedicated reader thread. Raising it burns more CPU against a +/// genuinely stuck consumer; lowering it puts a sleep in the common case. +#[cfg(unix)] +#[cfg_attr(not(target_os = "macos"), allow(dead_code))] +const SEND_YIELD_SPINS: u32 = 64; + +/// Longest the send loop sleeps between attempts on a full channel. +/// +/// This bounds only how quickly a parked send notices a shutdown request that +/// closing the receiver has not already covered. Raising it delays that +/// notice; lowering it costs more wakeups under sustained backpressure. +#[cfg(unix)] +#[cfg_attr(not(target_os = "macos"), allow(dead_code))] +const SEND_RETRY_MAX: std::time::Duration = std::time::Duration::from_millis(1); + +/// Send one item, waiting out a full channel but waking on `shutdown_fd`. +/// +/// Returns `Stop` when the receiver is gone or shutdown has been requested, +/// which is the reader thread's cue to exit. Unlike `blocking_send` this +/// cannot park past a shutdown request, so the `join()` in `Drop` always +/// returns. The caller must keep the socket owning `shutdown_fd` alive across +/// the call; `poll` on a closed fd reports `POLLNVAL` rather than `POLLIN`, so +/// even a lifetime mistake degrades to waiting rather than to a false stop. +/// +/// Compiled on every unix so Linux CI exercises the tests below; only the +/// macOS reader thread calls it. +#[cfg(unix)] +#[cfg_attr(not(target_os = "macos"), allow(dead_code))] +pub(crate) fn send_frame( + tx: &tokio::sync::mpsc::Sender, + item: T, + shutdown_fd: std::os::unix::io::RawFd, +) -> SendOutcome { + use tokio::sync::mpsc::error::TrySendError; + + let mut item = item; + let mut spins = 0u32; + let mut backoff = std::time::Duration::from_micros(50); + loop { + match tx.try_send(item) { + Ok(()) => return SendOutcome::Sent, + Err(TrySendError::Closed(_)) => return SendOutcome::Stop, + Err(TrySendError::Full(returned)) => { + if fd_is_readable(shutdown_fd) { + return SendOutcome::Stop; + } + item = returned; + if spins < SEND_YIELD_SPINS { + spins += 1; + std::thread::yield_now(); + } else { + std::thread::sleep(backoff); + backoff = (backoff * 2).min(SEND_RETRY_MAX); + } + } + } + } +} + +/// True if `fd` has data ready, tested without blocking. +#[cfg(unix)] +#[cfg_attr(not(target_os = "macos"), allow(dead_code))] +pub(crate) fn fd_is_readable(fd: std::os::unix::io::RawFd) -> bool { + let mut pfd = libc::pollfd { + fd, + events: libc::POLLIN, + revents: 0, + }; + let ret = unsafe { libc::poll(&mut pfd, 1, 0) }; + ret > 0 && (pfd.revents & libc::POLLIN) != 0 +} + // ============================================================================= // Linux: AsyncFd-based async wrapper // ============================================================================= @@ -111,7 +195,9 @@ mod async_impl { pub struct AsyncPacketSocket { inner: Arc, - rx: tokio::sync::Mutex>, + /// `None` once shutdown has taken the receiver, which is what makes + /// a reader thread parked on a full channel return at once. + rx: tokio::sync::Mutex>>, reader_thread: Option>, } @@ -146,7 +232,13 @@ mod async_impl { match result { Ok((n, mac)) => { let data = read_buf[..n].to_vec(); - if tx.blocking_send((data, mac)).is_err() { + // Not blocking_send: a send parked on a + // full channel must still notice shutdown, + // or Drop's join() never returns. + if matches!( + super::send_frame(&tx, (data, mac), shutdown_fd), + super::SendOutcome::Stop + ) { return; } } @@ -207,7 +299,7 @@ mod async_impl { Ok(Self { inner, - rx: tokio::sync::Mutex::new(rx), + rx: tokio::sync::Mutex::new(Some(rx)), reader_thread: Some(reader_thread), }) } @@ -230,7 +322,10 @@ mod async_impl { } pub async fn recv_from(&self, buf: &mut [u8]) -> Result<(usize, [u8; 6]), TransportError> { - let mut rx = self.rx.lock().await; + let mut guard = self.rx.lock().await; + let Some(rx) = guard.as_mut() else { + return Err(TransportError::RecvFailed("reader thread stopped".into())); + }; match rx.recv().await { Some((data, mac)) => { let n = data.len().min(buf.len()); @@ -247,15 +342,25 @@ mod async_impl { /// Signal the reader thread to stop. /// - /// Sets the shutdown flag; the reader thread checks it after - /// each BPF read timeout (~250ms) and exits. + /// Drops the receiver where it can, which makes a send parked on a + /// full channel fail immediately, then writes the shutdown pipe that + /// the thread's `select()` and `send_frame` both watch. The receiver + /// is unavailable while a `recv_from` holds the lock; `Drop` takes it + /// unconditionally, so the pipe is what covers that window. pub fn shutdown(&self) { + if let Ok(mut guard) = self.rx.try_lock() { + guard.take(); + } self.inner.request_shutdown(); } } impl Drop for AsyncPacketSocket { fn drop(&mut self) { + // Drop the receiver before joining: a send parked on a full + // channel then returns at once, with no polling and no latency + // added to the steady-state path. + self.rx.get_mut().take(); self.inner.request_shutdown(); if let Some(handle) = self.reader_thread.take() { let _ = handle.join(); @@ -284,3 +389,133 @@ pub struct PacketSocket; #[cfg(windows)] pub struct AsyncPacketSocket; + +// ============================================================================= +// Tests +// ============================================================================= + +#[cfg(all(test, unix))] +mod tests { + use super::{SendOutcome, fd_is_readable, send_frame}; + use std::sync::mpsc; + use std::time::Duration; + + /// A pipe, as the shutdown signal, returned as (read fd, write fd). + /// + /// Leaked deliberately: these live for the length of one test and closing + /// them mid-poll is exactly the confusion the test is meant to avoid. + fn shutdown_pipe() -> (std::os::unix::io::RawFd, std::os::unix::io::RawFd) { + let mut fds = [0i32; 2]; + let ret = unsafe { libc::pipe(fds.as_mut_ptr()) }; + assert_eq!(ret, 0, "pipe() failed"); + (fds[0], fds[1]) + } + + fn signal(write_fd: std::os::unix::io::RawFd) { + let byte = [1u8]; + let ret = unsafe { libc::write(write_fd, byte.as_ptr() as *const libc::c_void, 1) }; + assert_eq!(ret, 1, "write() to shutdown pipe failed"); + } + + #[test] + fn fd_is_readable_is_false_for_an_unwritten_pipe_and_true_after_a_write() { + let (read_fd, write_fd) = shutdown_pipe(); + assert!(!fd_is_readable(read_fd)); + signal(write_fd); + assert!(fd_is_readable(read_fd)); + } + + #[test] + fn send_frame_delivers_when_the_channel_has_room() { + let (tx, mut rx) = tokio::sync::mpsc::channel::>(1); + let (read_fd, _write_fd) = shutdown_pipe(); + + assert!(matches!( + send_frame(&tx, vec![1u8, 2, 3], read_fd), + SendOutcome::Sent + )); + assert_eq!(rx.try_recv().unwrap(), vec![1u8, 2, 3]); + } + + #[test] + fn send_frame_returns_stop_when_the_receiver_is_gone() { + let (tx, rx) = tokio::sync::mpsc::channel::>(1); + let (read_fd, _write_fd) = shutdown_pipe(); + drop(rx); + + assert!(matches!( + send_frame(&tx, vec![0u8], read_fd), + SendOutcome::Stop + )); + } + + #[test] + fn send_frame_returns_stop_when_the_receiver_is_dropped_while_the_channel_is_full() { + // The mechanism `Drop` relies on: closing the channel releases a + // sender that is waiting for room. + let (tx, rx) = tokio::sync::mpsc::channel::>(1); + let (read_fd, _write_fd) = shutdown_pipe(); + tx.try_send(vec![0u8]).unwrap(); + + let (done_tx, done_rx) = mpsc::channel(); + let sender = std::thread::spawn(move || { + let outcome = send_frame(&tx, vec![1u8], read_fd); + done_tx.send(matches!(outcome, SendOutcome::Stop)).unwrap(); + }); + // The send is parked on a full channel; only the drop frees it. + assert!(done_rx.recv_timeout(Duration::from_millis(50)).is_err()); + drop(rx); + + let stopped = done_rx + .recv_timeout(Duration::from_secs(5)) + .expect("send_frame did not return after the receiver was dropped"); + sender.join().unwrap(); + assert!(stopped); + } + + #[test] + fn send_frame_returns_stop_when_shutdown_is_requested_and_the_channel_is_full() { + // The defect: `blocking_send` on a full channel nobody is draining + // parks forever, so the reader thread never sees shutdown and the + // `join()` in `Drop` never returns. See the ignored test below for + // the same fixture against `blocking_send`. + let (tx, _rx) = tokio::sync::mpsc::channel::>(1); + let (read_fd, write_fd) = shutdown_pipe(); + tx.try_send(vec![0u8]).unwrap(); + signal(write_fd); + + let (done_tx, done_rx) = mpsc::channel(); + let sender = std::thread::spawn(move || { + let outcome = send_frame(&tx, vec![1u8], read_fd); + done_tx.send(matches!(outcome, SendOutcome::Stop)).unwrap(); + }); + + let stopped = done_rx + .recv_timeout(Duration::from_secs(5)) + .expect("send_frame parked past a shutdown request"); + sender.join().unwrap(); + assert!(stopped); + } + + #[test] + #[ignore = "demonstrates the defect: blocking_send never returns, so this hangs"] + fn blocking_send_parks_past_a_shutdown_request_when_the_channel_is_full() { + // Run with `--ignored` to watch the old send site hang. Kept as the + // observed red-before for the test above, which cannot itself fail + // against the old code because `send_frame` did not exist then. + let (tx, _rx) = tokio::sync::mpsc::channel::>(1); + let (_read_fd, write_fd) = shutdown_pipe(); + tx.try_send(vec![0u8]).unwrap(); + signal(write_fd); + + let (done_tx, done_rx) = mpsc::channel(); + std::thread::spawn(move || { + let _ = tx.blocking_send(vec![1u8]); + done_tx.send(()).unwrap(); + }); + + done_rx + .recv_timeout(Duration::from_secs(5)) + .expect("blocking_send returned, so the send site was already cancellable"); + } +} diff --git a/src/transport/ethernet/mod.rs b/src/transport/ethernet/mod.rs index d8209ed9..f134205a 100644 --- a/src/transport/ethernet/mod.rs +++ b/src/transport/ethernet/mod.rs @@ -449,7 +449,13 @@ async fn ethernet_receive_loop( stats.record_beacon_recv(); if listen_enabled && let Some(pubkey) = parse_beacon(&buf[..len]) { - neighbor_buffer.add_peer(src_mac, pubkey); + // `add_peer` reports whether the buffer took the + // beacon. It refuses once the distinct-MAC cap is + // reached, which is the bound on an unauthenticated + // broadcast frame naming a fresh MAC every time. + if !neighbor_buffer.add_peer(src_mac, pubkey) { + stats.record_beacon_dropped(); + } trace!( transport_id = %transport_id, remote_mac = %format_mac(&src_mac), diff --git a/src/transport/ethernet/neighbor.rs b/src/transport/ethernet/neighbor.rs index a778d2c2..1c543bf0 100644 --- a/src/transport/ethernet/neighbor.rs +++ b/src/transport/ethernet/neighbor.rs @@ -7,7 +7,10 @@ use crate::transport::{DiscoveredPeer, TransportAddr, TransportId}; use secp256k1::XOnlyPublicKey; +use std::collections::HashMap; +use std::collections::hash_map::Entry; use std::sync::Mutex; +use tracing::warn; /// Beacon protocol version. pub const BEACON_VERSION: u8 = 0x01; @@ -46,10 +49,33 @@ pub fn parse_beacon(data: &[u8]) -> Option { XOnlyPublicKey::from_slice(&data[2..34]).ok() } +/// Maximum distinct source MACs held between drains. +/// +/// Beacons are unauthenticated broadcast frames, so anything on the segment +/// can name as many source MACs as it likes; without a bound the buffer grows +/// with the flood rate, and it is not drained at all while the transport is +/// not operational. This caps it at roughly a thousand small structs, tens of +/// kilobytes. Raising it costs that much more memory per transport; lowering +/// it risks truncating discovery on a very large segment. A thousand distinct +/// beaconing FIPS neighbors within one tick is far outside anything a real +/// deployment produces. +const MAX_BUFFERED_PEERS: usize = 1024; + /// Buffer for discovered peers, drained by `discover()`. pub struct NeighborBuffer { transport_id: TransportId, - peers: Mutex>, + peers: Mutex, +} + +/// Peers keyed by source MAC, plus the sighting order `take()` restores. +#[derive(Default)] +struct Buffered { + by_mac: HashMap<[u8; 6], (u64, DiscoveredPeer)>, + seq: u64, + /// Beacons refused for want of room, cumulative and never reset. + dropped: u64, + /// Cumulative drop count that earns the next log record. + warn_at: u64, } impl NeighborBuffer { @@ -57,25 +83,82 @@ impl NeighborBuffer { pub fn new(transport_id: TransportId) -> Self { Self { transport_id, - peers: Mutex::new(Vec::new()), + peers: Mutex::new(Buffered::default()), } } /// Add a discovered peer from a received beacon. - pub fn add_peer(&self, src_mac: [u8; 6], pubkey: XOnlyPublicKey) { - let addr = TransportAddr::from_bytes(&src_mac); - let peer = DiscoveredPeer::with_hint(self.transport_id, addr, pubkey); - let mut peers = self.peers.lock().unwrap_or_else(|e| e.into_inner()); - // Deduplicate by MAC address — keep the latest - peers.retain(|p| p.addr.as_bytes() != src_mac); - peers.push(peer); + /// + /// Returns false when the beacon was refused because the buffer is full. + /// A MAC already buffered is always refreshed, so a flood of new MACs + /// cannot stop a known neighbor from being seen again. + pub fn add_peer(&self, src_mac: [u8; 6], pubkey: XOnlyPublicKey) -> bool { + let mut buffered = self.peers.lock().unwrap_or_else(|e| e.into_inner()); + buffered.seq += 1; + let seq = buffered.seq; + let full = buffered.by_mac.len() >= MAX_BUFFERED_PEERS; + let stored = match buffered.by_mac.entry(src_mac) { + // Refreshing moves the MAC to the end, as retain-then-push did. + Entry::Occupied(mut slot) => { + slot.insert((seq, self.peer(src_mac, pubkey))); + true + } + Entry::Vacant(slot) if !full => { + slot.insert((seq, self.peer(src_mac, pubkey))); + true + } + Entry::Vacant(_) => false, + }; + if !stored { + buffered.dropped += 1; + } + stored } - /// Drain all discovered peers since the last call. + /// Drain all discovered peers since the last call, oldest sighting first. pub fn take(&self) -> Vec { - let mut peers = self.peers.lock().unwrap_or_else(|e| e.into_inner()); - std::mem::take(&mut *peers) + let mut buffered = self.peers.lock().unwrap_or_else(|e| e.into_inner()); + let mut ordered: Vec<(u64, DiscoveredPeer)> = + buffered.by_mac.drain().map(|(_, entry)| entry).collect(); + // The reconcile layer spends a finite connect budget in this order, so + // which neighbor gets dialed must not depend on hash iteration order. + ordered.sort_unstable_by_key(|(seq, _)| *seq); + // Rate-limited: the drop rate is whatever the flooder chooses, and one + // record per drain would hand it the log volume too. + if buffered.dropped >= buffered.warn_at.max(1) { + warn!( + transport_id = %self.transport_id, + dropped = buffered.dropped, + cap = MAX_BUFFERED_PEERS, + "discovery buffer full, beacons from unseen neighbors refused" + ); + buffered.warn_at = next_decade(buffered.dropped); + } + ordered.into_iter().map(|(_, peer)| peer).collect() } + + /// Beacons refused for want of room since this buffer was created. + pub fn dropped(&self) -> u64 { + self.peers.lock().unwrap_or_else(|e| e.into_inner()).dropped + } + + /// Build the buffered peer record for one beacon. + fn peer(&self, src_mac: [u8; 6], pubkey: XOnlyPublicKey) -> DiscoveredPeer { + let addr = TransportAddr::from_bytes(&src_mac); + DiscoveredPeer::with_hint(self.transport_id, addr, pubkey) + } +} + +/// Smallest power of ten strictly greater than `n`, saturating at `u64::MAX`. +fn next_decade(n: u64) -> u64 { + let mut threshold = 1u64; + while threshold <= n { + match threshold.checked_mul(10) { + Some(next) => threshold = next, + None => return u64::MAX, + } + } + threshold } // ============================================================================ @@ -163,4 +246,88 @@ mod tests { let peers = buffer.take(); assert_eq!(peers.len(), 1); } + + /// Distinct MAC number `n`, for filling the buffer. + fn nth_mac(n: usize) -> [u8; 6] { + let bytes = (n as u64).to_be_bytes(); + [0x02, bytes[3], bytes[4], bytes[5], bytes[6], bytes[7]] + } + + #[test] + fn discovery_buffer_stops_buffering_past_the_cap() { + // The defect: an unauthenticated flood of source MACs grew the buffer + // without bound. Fails against the uncapped Vec, which returns all of + // them. + let buffer = NeighborBuffer::new(TransportId::new(1)); + let pubkey = test_pubkey(); + for n in 0..MAX_BUFFERED_PEERS + 50 { + buffer.add_peer(nth_mac(n), pubkey); + } + + let peers = buffer.take(); + assert_eq!(peers.len(), MAX_BUFFERED_PEERS); + // Drop-new keeps the earliest sightings. + assert_eq!(peers[0].addr.as_bytes(), &nth_mac(0)); + } + + #[test] + fn discovery_buffer_counts_dropped_beacons() { + let buffer = NeighborBuffer::new(TransportId::new(1)); + let pubkey = test_pubkey(); + for n in 0..MAX_BUFFERED_PEERS { + assert!(buffer.add_peer(nth_mac(n), pubkey)); + } + for n in MAX_BUFFERED_PEERS..MAX_BUFFERED_PEERS + 7 { + assert!(!buffer.add_peer(nth_mac(n), pubkey)); + } + + assert_eq!(buffer.dropped(), 7); + } + + #[test] + fn discovery_buffer_repeat_beacon_from_a_full_buffer_still_refreshes() { + let buffer = NeighborBuffer::new(TransportId::new(1)); + let pubkey = test_pubkey(); + for n in 0..MAX_BUFFERED_PEERS + 50 { + buffer.add_peer(nth_mac(n), pubkey); + } + // A neighbor already buffered must not be refused by a full buffer. + assert!(buffer.add_peer(nth_mac(0), pubkey)); + + let peers = buffer.take(); + assert_eq!(peers.len(), MAX_BUFFERED_PEERS); + assert_eq!(peers[peers.len() - 1].addr.as_bytes(), &nth_mac(0)); + } + + #[test] + fn discovery_buffer_drain_preserves_last_seen_order() { + // A regression pin on the map rewrite rather than a test of the + // defect: retain-then-push already produced this order. + let buffer = NeighborBuffer::new(TransportId::new(1)); + let pubkey = test_pubkey(); + let a = [0xaa; 6]; + let b = [0xbb; 6]; + let c = [0xcc; 6]; + + buffer.add_peer(a, pubkey); + buffer.add_peer(b, pubkey); + buffer.add_peer(c, pubkey); + buffer.add_peer(a, pubkey); + + let macs: Vec<_> = buffer + .take() + .iter() + .map(|p| p.addr.as_bytes().to_vec()) + .collect(); + assert_eq!(macs, vec![b.to_vec(), c.to_vec(), a.to_vec()]); + } + + #[test] + fn next_decade_steps_by_powers_of_ten() { + assert_eq!(next_decade(0), 1); + assert_eq!(next_decade(1), 10); + assert_eq!(next_decade(9), 10); + assert_eq!(next_decade(10), 100); + assert_eq!(next_decade(u64::MAX), u64::MAX); + } } diff --git a/src/transport/ethernet/stats.rs b/src/transport/ethernet/stats.rs index 03462321..5e10659a 100644 --- a/src/transport/ethernet/stats.rs +++ b/src/transport/ethernet/stats.rs @@ -17,6 +17,7 @@ pub struct EthernetStats { pub recv_errors: AtomicU64, pub beacons_sent: AtomicU64, pub beacons_recv: AtomicU64, + pub beacons_dropped: AtomicU64, pub frames_too_short: AtomicU64, pub frames_too_long: AtomicU64, } @@ -33,6 +34,7 @@ impl EthernetStats { recv_errors: AtomicU64::new(0), beacons_sent: AtomicU64::new(0), beacons_recv: AtomicU64::new(0), + beacons_dropped: AtomicU64::new(0), frames_too_short: AtomicU64::new(0), frames_too_long: AtomicU64::new(0), } @@ -70,6 +72,11 @@ impl EthernetStats { self.beacons_recv.fetch_add(1, Ordering::Relaxed); } + /// Record a received beacon the discovery buffer had no room for. + pub fn record_beacon_dropped(&self) { + self.beacons_dropped.fetch_add(1, Ordering::Relaxed); + } + /// Take a snapshot of all counters. pub fn snapshot(&self) -> EthernetStatsSnapshot { EthernetStatsSnapshot { @@ -81,6 +88,7 @@ impl EthernetStats { recv_errors: self.recv_errors.load(Ordering::Relaxed), beacons_sent: self.beacons_sent.load(Ordering::Relaxed), beacons_recv: self.beacons_recv.load(Ordering::Relaxed), + beacons_dropped: self.beacons_dropped.load(Ordering::Relaxed), frames_too_short: self.frames_too_short.load(Ordering::Relaxed), frames_too_long: self.frames_too_long.load(Ordering::Relaxed), } @@ -104,6 +112,7 @@ pub struct EthernetStatsSnapshot { pub recv_errors: u64, pub beacons_sent: u64, pub beacons_recv: u64, + pub beacons_dropped: u64, pub frames_too_short: u64, pub frames_too_long: u64, } diff --git a/src/transport/udp/mod.rs b/src/transport/udp/mod.rs index 6f4ff868..9580ed47 100644 --- a/src/transport/udp/mod.rs +++ b/src/transport/udp/mod.rs @@ -25,6 +25,18 @@ use tracing::{debug, info, trace, warn}; /// DNS cache TTL for hostname resolution (60 seconds). const DNS_CACHE_TTL: Duration = Duration::from_secs(60); +/// Upper bound on the number of hostnames the DNS cache holds at once. +/// +/// The cache is keyed by the address string a dial was asked for, and under a +/// rendezvous policy that accepts advertised endpoints those strings come from +/// remote parties, so without a bound the map grows for the life of the +/// process. 256 sits about two orders of magnitude above the number of +/// distinct hostnames a configured peer list produces, so no ordinary +/// deployment reaches it. Lowering it starts to be reachable by a large peer +/// list, and the only cost of an eviction is one extra DNS lookup on the next +/// dial of that name; raising it buys nothing but resident memory. +const DNS_CACHE_MAX_ENTRIES: usize = 256; + /// UDP transport for FIPS. /// /// Provides connectionless, unreliable packet delivery over UDP/IP. @@ -155,10 +167,8 @@ impl UdpTransport { // Check cache { let cache = self.dns_cache.lock().unwrap_or_else(|e| e.into_inner()); - if let Some((resolved, cached_at)) = cache.get(addr) - && cached_at.elapsed() < DNS_CACHE_TTL - { - return Ok(*resolved); + if let Some(resolved) = cache_lookup(&cache, addr, Instant::now()) { + return Ok(resolved); } } @@ -168,7 +178,13 @@ impl UdpTransport { // Store in cache { let mut cache = self.dns_cache.lock().unwrap_or_else(|e| e.into_inner()); - cache.insert(addr.clone(), (resolved, Instant::now())); + cache_store( + &mut cache, + addr.clone(), + resolved, + Instant::now(), + DNS_CACHE_MAX_ENTRIES, + ); } Ok(resolved) @@ -611,6 +627,56 @@ async fn udp_receive_loop( } } +/// A cached resolution for `key`, if one is present and still inside +/// `DNS_CACHE_TTL` at `now`. +fn cache_lookup( + cache: &HashMap, + key: &TransportAddr, + now: Instant, +) -> Option { + cache + .get(key) + .filter(|(_, cached_at)| now.duration_since(*cached_at) < DNS_CACHE_TTL) + .map(|(resolved, _)| *resolved) +} + +/// Record a resolution, keeping the cache at or below `cap` entries. +/// +/// Refreshing a name already present never evicts anything. Otherwise every +/// entry past its TTL is dropped first, and only if that leaves the map full +/// is the oldest remaining entry evicted. Eviction is by insertion time rather +/// than by last use: the timestamp is already there as the TTL clock, and +/// tracking last use would mean writing to the map on the read path of every +/// dial. The sweep is linear in `cap` and runs only on a resolution miss, so +/// at most once per TTL per name. +fn cache_store( + cache: &mut HashMap, + key: TransportAddr, + resolved: SocketAddr, + now: Instant, + cap: usize, +) { + if let Some(entry) = cache.get_mut(&key) { + *entry = (resolved, now); + return; + } + + cache.retain(|_, (_, cached_at)| now.duration_since(*cached_at) < DNS_CACHE_TTL); + + while cache.len() >= cap { + let Some(oldest) = cache + .iter() + .min_by_key(|(_, (_, cached_at))| *cached_at) + .map(|(key, _)| key.clone()) + else { + break; + }; + cache.remove(&oldest); + } + + cache.insert(key, (resolved, now)); +} + // ============================================================================ // Tests // ============================================================================ @@ -621,6 +687,112 @@ mod tests { use crate::transport::packet_channel; use tokio::time::{Duration, timeout}; + /// A distinct hostname key, so each store is a fresh entry. + fn dns_key(n: usize) -> TransportAddr { + TransportAddr::from(format!("host{n}.example:2121")) + } + + fn dns_value() -> SocketAddr { + "198.51.100.1:2121".parse().unwrap() + } + + /// The cache is keyed by strings a remote party can choose, so its size + /// has to be bounded no matter how many distinct names are dialed. + #[test] + fn dns_cache_store_refuses_to_exceed_the_cap() { + const CAP: usize = 8; + let now = Instant::now(); + let mut cache = HashMap::new(); + + for n in 0..CAP + 5 { + cache_store(&mut cache, dns_key(n), dns_value(), now, CAP); + assert!( + cache.len() <= CAP, + "cache grew to {} entries past a cap of {CAP}", + cache.len() + ); + } + } + + /// A stale entry used to be overwritten on the next dial of the same name + /// and otherwise never removed, so a name dialed once sat there forever. + #[test] + fn dns_cache_store_evicts_entries_past_their_ttl() { + let now = Instant::now(); + let expired_at = now.checked_sub(DNS_CACHE_TTL * 2).expect("monotonic clock"); + let mut cache = HashMap::new(); + cache.insert(dns_key(0), (dns_value(), expired_at)); + + cache_store( + &mut cache, + dns_key(1), + dns_value(), + now, + DNS_CACHE_MAX_ENTRIES, + ); + + assert!( + !cache.contains_key(&dns_key(0)), + "an entry past its TTL should be swept, not left to accumulate" + ); + assert!(cache_lookup(&cache, &dns_key(0), now).is_none()); + assert!(cache_lookup(&cache, &dns_key(1), now).is_some()); + } + + /// With nothing expired, the cap is enforced by dropping the oldest entry. + /// The ages here are all well inside the TTL, so the expiry sweep cannot + /// be what makes room and the eviction branch is the one under test. + #[test] + fn dns_cache_store_evicts_the_oldest_entry_when_every_entry_is_fresh() { + const CAP: usize = 4; + let now = Instant::now(); + let mut cache = HashMap::new(); + for n in 0..CAP { + let age = Duration::from_secs((CAP - n) as u64); + assert!(age < DNS_CACHE_TTL, "fixture must stay inside the TTL"); + let cached_at = now.checked_sub(age).expect("monotonic clock"); + cache.insert(dns_key(n), (dns_value(), cached_at)); + } + assert_eq!(cache.len(), CAP, "no entry should be expired going in"); + + cache_store(&mut cache, dns_key(CAP), dns_value(), now, CAP); + + assert_eq!(cache.len(), CAP); + assert!( + !cache.contains_key(&dns_key(0)), + "the oldest entry should be the one evicted" + ); + for n in 1..=CAP { + assert!( + cache.contains_key(&dns_key(n)), + "entry {n} should have survived" + ); + } + } + + /// Re-resolving a name already cached is the common case on a live node. + /// It must not cost another entry its place. + #[test] + fn dns_cache_store_refreshing_an_existing_key_evicts_nothing() { + const CAP: usize = 4; + let now = Instant::now(); + let mut cache = HashMap::new(); + for n in 0..CAP { + let cached_at = now + .checked_sub(Duration::from_secs((CAP - n) as u64)) + .expect("monotonic clock"); + cache.insert(dns_key(n), (dns_value(), cached_at)); + } + + cache_store(&mut cache, dns_key(0), dns_value(), now, CAP); + + assert_eq!(cache.len(), CAP); + for n in 0..CAP { + assert!(cache.contains_key(&dns_key(n)), "entry {n} should remain"); + } + assert_eq!(cache_lookup(&cache, &dns_key(0), now), Some(dns_value())); + } + fn make_config(port: u16) -> UdpConfig { UdpConfig { bind_addr: Some(format!("127.0.0.1:{}", port)), diff --git a/src/upper/hosts.rs b/src/upper/hosts.rs index e68eb7e1..f9e9f3ff 100644 --- a/src/upper/hosts.rs +++ b/src/upper/hosts.rs @@ -134,17 +134,34 @@ impl HostMap { /// /// If the file does not exist, returns an empty map (not an error). /// Parse errors on individual lines are logged as warnings and skipped. + /// A read failure is logged and also yields an empty map; a caller that + /// must not mistake an unreadable file for an empty one uses + /// [`Self::try_load_hosts_file`] instead. pub fn load_hosts_file(path: &Path) -> Self { - let contents = match std::fs::read_to_string(path) { - Ok(c) => c, - Err(e) if e.kind() == std::io::ErrorKind::NotFound => { - debug!(path = %path.display(), "No hosts file found, skipping"); - return Self::new(); - } + match Self::try_load_hosts_file(path) { + Ok(map) => map, Err(e) => { warn!(path = %path.display(), error = %e, "Failed to read hosts file"); - return Self::new(); + Self::new() } + } + } + + /// Load a host map from a hosts file, reporting read failures. + /// + /// An absent file is a policy, not a fault: it resolves to an empty map + /// and `Ok`. Anything else — no read permission, an I/O error, non-UTF-8 + /// content, or a `NotFound` that contradicts a successful stat and so + /// means the file is being rewritten under us — is returned as an error + /// so the caller can keep whatever it loaded last. + pub fn try_load_hosts_file(path: &Path) -> Result { + let contents = match std::fs::read_to_string(path) { + Ok(c) => c, + Err(e) if e.kind() == std::io::ErrorKind::NotFound && file_mtime(path).is_none() => { + debug!(path = %path.display(), "No hosts file found, skipping"); + return Ok(Self::new()); + } + Err(e) => return Err(e), }; let mut map = Self::new(); @@ -183,7 +200,7 @@ impl HostMap { if !map.is_empty() { info!(path = %path.display(), count = map.len(), "Loaded hosts file"); } - map + Ok(map) } /// Merge another host map into this one. The other map wins on conflicts. @@ -224,8 +241,16 @@ impl HostMapReloader { /// /// Performs the initial load of the hosts file and merges with the base map. pub fn new(base: HostMap, path: std::path::PathBuf) -> Self { - let last_mtime = file_mtime(&path); - let hosts_file = HostMap::load_hosts_file(&path); + // A failed initial read records no mtime, so the next check sees a + // change and retries rather than treating the unread file as empty + // for the lifetime of the process. + let (last_mtime, hosts_file) = match HostMap::try_load_hosts_file(&path) { + Ok(map) => (file_mtime(&path), map), + Err(e) => { + warn!(path = %path.display(), error = %e, "Failed to read hosts file"); + (None, HostMap::new()) + } + }; let mut effective = base.clone(); effective.merge(hosts_file); @@ -242,6 +267,11 @@ impl HostMapReloader { &self.effective } + /// Path of the hosts file this reloader tracks. + pub fn path(&self) -> &Path { + &self.path + } + /// Check if the hosts file has been modified and reload if so. /// /// Returns `true` if the map was reloaded. @@ -254,7 +284,32 @@ impl HostMapReloader { // File appeared, disappeared, or was modified self.last_mtime = current_mtime; - let hosts_file = HostMap::load_hosts_file(&self.path); + self.apply(HostMap::load_hosts_file(&self.path)); + true + } + + /// Check if the hosts file has been modified and reload if so, reporting + /// read failures. + /// + /// On failure neither the recorded mtime nor the effective map is + /// touched, so the caller keeps its last-good state and the next call + /// retries. Returns `true` if the map was reloaded. + pub fn try_check_reload(&mut self) -> Result { + let current_mtime = file_mtime(&self.path); + + if current_mtime == self.last_mtime { + return Ok(false); + } + + let hosts_file = HostMap::try_load_hosts_file(&self.path)?; + self.last_mtime = current_mtime; + self.apply(hosts_file); + Ok(true) + } + + /// Replace the effective map with the base merged with a freshly read + /// hosts file. + fn apply(&mut self, hosts_file: HostMap) { let mut new_effective = self.base.clone(); new_effective.merge(hosts_file); @@ -266,7 +321,6 @@ impl HostMapReloader { entries = count, "Reloaded hosts file" ); - true } } diff --git a/src/upper/icmp.rs b/src/upper/icmp.rs index f6d87ab1..fecfe8e5 100644 --- a/src/upper/icmp.rs +++ b/src/upper/icmp.rs @@ -112,6 +112,22 @@ pub const FIPS_IPV6_OVERHEAD: u16 = 77; /// the `MtuExceeded` signal — reaches it through this module. pub use crate::proto::mmp::MIN_ACTIONABLE_PATH_MTU; +/// Smallest path MTU this node will act on when the claim arrives on the +/// unauthenticated reactive carrier, `MtuExceeded`. +/// +/// Held equal to [`MIN_ACTIONABLE_PATH_MTU`] so no hop legitimately configured +/// with a small transport MTU loses reactive feedback. It is a separate +/// constant because the two carriers differ in what they prove: the +/// authenticated `PathMtuNotification` and the proof-carrying discovery +/// response come from a party this node has verified, whereas this one comes +/// from whoever could route a datagram here. What keeps a legal-but-forged +/// claim from pinning a session is corroboration against what this node has +/// actually sent, not this floor. Raising it (576 is the value the original +/// path-MTU floor design proposed, and derives an inner IPv6 MTU of 499) +/// bounds the outcome of an uncorroborated claim further, at the cost of +/// ignoring an honest report from any hop configured between the two values. +pub const MIN_REACTIVE_PATH_MTU: u16 = MIN_ACTIONABLE_PATH_MTU; + /// Calculate the effective IPv6 MTU for FIPS-encapsulated traffic. /// /// Given a transport MTU (e.g., UDP payload size), returns the maximum diff --git a/src/utils/mod.rs b/src/utils/mod.rs index 9c18a16b..b961292e 100644 --- a/src/utils/mod.rs +++ b/src/utils/mod.rs @@ -1,8 +1,10 @@ //! Utility modules. //! //! Shared infrastructure that doesn't belong to a specific protocol layer: -//! session index allocation, Unix socket binding, and other cross-cutting -//! concerns. +//! session index allocation, Unix socket binding and its permission +//! primitives, and other cross-cutting concerns. pub mod index; pub mod sockbind; +#[cfg(unix)] +pub mod sockperm; diff --git a/src/utils/sockbind.rs b/src/utils/sockbind.rs index 9a32245c..85a97746 100644 --- a/src/utils/sockbind.rs +++ b/src/utils/sockbind.rs @@ -48,7 +48,12 @@ pub fn bind(path: &Path, what: &str) -> Result { remove_stale_socket(path, what)?; } - let listener = UnixListener::bind(path)?; + // Bound through `sockperm` rather than `UnixListener::bind` directly: + // bind(2) creates the socket inode as `0777 & !umask`, so under a + // permissive umask it is world-accessible for the window between the + // bind and the `set_socket_access` chmod below. That chmod stays the + // authority on the final mode; this only closes the window. + let listener = crate::utils::sockperm::bind(path)?; set_socket_access(path, managed_parent.as_deref(), chown_to_fips_group)?; @@ -65,11 +70,21 @@ pub fn bind(path: &Path, what: &str) -> Result { /// socket's private directory. #[cfg(unix)] fn ensure_socket_parent(parent: &Path) -> Result { + use std::os::unix::fs::DirBuilderExt; + if parent.as_os_str().is_empty() { return Ok(false); } - match std::fs::create_dir(parent) { + // Mode carried on creation rather than applied afterwards. A directory + // made by plain `create_dir` is `0777 & !umask`, and an intermediate + // ancestor is never chmodded by anything below, so under a permissive + // umask it would stay world-writable for the life of the host — and a + // world-writable parent lets an unprivileged account plant an entry at + // the socket path. 0750 is what `set_socket_access` applies to a managed + // parent anyway, and what the systemd unit and FreeBSD rc script already + // use. + match std::fs::DirBuilder::new().mode(0o750).create(parent) { Ok(()) => Ok(true), Err(error) if error.kind() == std::io::ErrorKind::AlreadyExists => { if parent.is_dir() { @@ -120,6 +135,14 @@ fn set_socket_access( /// If the file exists but no one is listening, remove it so we can bind. This /// handles unclean daemon exits. A live listener yields `AddrInUse` instead, so /// two daemons cannot silently take the same path. +/// +/// The gap between the connect probe and the bind that follows is accepted +/// rather than closed. Reaching it needs write access to the socket's parent +/// directory, which the packaged layouts give to root alone (0750 and +/// root-owned under both systemd and the FreeBSD rc script), and an account +/// holding it can deny the daemon its socket more simply by squatting the path +/// before the daemon starts. The removal itself unlinks a symlink rather than +/// its target, so it is not an arbitrary delete. #[cfg(unix)] fn remove_stale_socket(path: &Path, what: &str) -> Result<(), std::io::Error> { match std::os::unix::net::UnixStream::connect(path) { diff --git a/src/utils/sockperm.rs b/src/utils/sockperm.rs new file mode 100644 index 00000000..2e1d1d08 --- /dev/null +++ b/src/utils/sockperm.rs @@ -0,0 +1,157 @@ +//! Permission-safe creation of Unix domain sockets and the directories +//! holding them. +//! +//! The socket inode and its parent directory are created with a mode the +//! process umask can only tighten, rather than created wide and narrowed +//! afterwards. The caller's own chmod and chown stay where they are and +//! remain the authority on the socket's final mode; this closes the window +//! between creation and that fix-up, and the case of an intermediate +//! directory that nothing fixes up at all. + +use std::path::Path; +use tokio::net::UnixListener; + +/// Mode for a directory this module creates to hold a control socket. +/// +/// Matches what the packaging already applies (systemd's +/// `RuntimeDirectoryMode=0750`, `install -d -m 0750` in the FreeBSD rc +/// script), so no packaged deployment sees a different directory mode than +/// it does today. Widening it would expose the socket path to accounts that +/// cannot reach it now; the umask can still tighten it further. +const SOCKET_DIR_MODE: u32 = 0o750; + +/// umask held across the socket bind. +/// +/// `bind(2)` creates the socket inode with `0777 & !umask`, so under a +/// permissive umask the socket is world-accessible until the chmod that +/// follows it. Masking the "other" bits makes the inode 0770 at creation, +/// which is the mode the caller applies a moment later anyway. Changing +/// this changes the mode the socket is created with, not the mode it ends +/// up with. +const BIND_UMASK: libc::mode_t = 0o007; + +/// Restores the process umask when dropped. +struct UmaskGuard(libc::mode_t); + +impl UmaskGuard { + /// Install `mask` as the process umask, remembering the previous one. + fn tighten(mask: libc::mode_t) -> Self { + // SAFETY: umask(2) cannot fail and touches only process state. + Self(unsafe { libc::umask(mask) }) + } +} + +impl Drop for UmaskGuard { + fn drop(&mut self) { + // SAFETY: as above; restoring the mask this guard replaced. + unsafe { + libc::umask(self.0); + } + } +} + +/// Create the directory that will hold a socket, and any missing ancestors. +/// +/// Directories come out 0750 rather than `0777 & !umask`. Nothing chmods an +/// intermediate directory afterwards, so one created under a permissive +/// umask would stay world-writable for the life of the host, and a +/// world-writable parent lets an unprivileged account plant an entry at the +/// socket path. +pub fn make_parent(parent: &Path) -> Result<(), std::io::Error> { + use std::os::unix::fs::DirBuilderExt; + + std::fs::DirBuilder::new() + .recursive(true) + .mode(SOCKET_DIR_MODE) + .create(parent) +} + +/// Bind a Unix listener whose inode is never world-accessible. +/// +/// The umask is process-global, so it is held across the bind alone. It +/// only clears bits, so anything else created inside that window comes out +/// more restrictive, never less. +pub fn bind(path: &Path) -> Result { + let _umask = UmaskGuard::tighten(BIND_UMASK); + UnixListener::bind(path) +} + +#[cfg(test)] +mod tests { + use super::*; + use std::os::unix::fs::PermissionsExt; + use std::sync::Mutex; + + /// The umask is process-global, so the tests that set it run one at a + /// time. This does not serialize against the rest of the test binary; + /// the mask used is 0o022, the ordinary default, so a file another test + /// creates in the window is unaffected. + static UMASK_LOCK: Mutex<()> = Mutex::new(()); + + /// Take the umask lock, ignoring poisoning: a test that fails while + /// holding it must not turn its siblings red for an unrelated reason. + fn umask_lock() -> std::sync::MutexGuard<'static, ()> { + UMASK_LOCK.lock().unwrap_or_else(|e| e.into_inner()) + } + + /// Read the current umask, which is only observable by replacing it. + fn current_umask() -> libc::mode_t { + // SAFETY: umask(2) cannot fail; the value read is put straight back. + unsafe { + let old = libc::umask(0o022); + libc::umask(old); + old + } + } + + fn mode_of(path: &Path) -> u32 { + std::fs::symlink_metadata(path) + .unwrap() + .permissions() + .mode() + } + + #[tokio::test] + async fn socket_is_created_without_other_access_under_a_permissive_umask() { + let _lock = umask_lock(); + let dir = tempfile::tempdir().unwrap(); + let path = dir.path().join("control.sock"); + + let restore = UmaskGuard::tighten(0o022); + let listener = bind(&path).unwrap(); + drop(restore); + + assert_eq!(mode_of(&path) & 0o007, 0); + drop(listener); + } + + #[tokio::test] + async fn socket_bind_leaves_the_process_umask_as_it_found_it() { + let _lock = umask_lock(); + let dir = tempfile::tempdir().unwrap(); + let path = dir.path().join("control.sock"); + + let restore = UmaskGuard::tighten(0o022); + let listener = bind(&path).unwrap(); + let after = current_umask(); + drop(restore); + + assert_eq!(after, 0o022); + drop(listener); + } + + #[test] + fn socket_parent_and_its_ancestors_are_created_without_other_access() { + let _lock = umask_lock(); + let dir = tempfile::tempdir().unwrap(); + let intermediate = dir.path().join("run"); + let parent = intermediate.join("fips"); + + let restore = UmaskGuard::tighten(0o022); + make_parent(&parent).unwrap(); + drop(restore); + + assert_eq!(mode_of(&intermediate) & 0o007, 0); + assert_eq!(mode_of(&parent) & 0o007, 0); + } +}