Merge maint into master, carrying the twenty-one security fixes

Not a replay: master had independently reorganized or reimplemented nearly
every area the batch touches, so most of this is re-derivation. Four files
maint modified do not exist on master at all, and git marks those DU with no
conflict markers, so the lazy resolution would have dropped 499 lines while
the tree still built and still passed tests.

Two conflicts turned out to be the same defect already present on master in a
different place, and fixing it there reaches further than the original did.
Master had centralized socket binding into sockbind across three sockets, one
of which maint does not have, with the same create-wide-then-narrow window; and
master's shared PerAddrRateLimiter swept the whole map on every admit with no
entry ceiling, which is the routing-error limiter finding in a primitive the
lookup-forward limiter also uses.

Two others would have added dead code if replayed. Master had already deleted
the pre-refactor restart block the epoch dampener patched, so it went onto
InboundDecision::RestartThenPromote instead; and master already carried the
relocated reactive-MTU floor check, so the copy inside the conflict was the old
post-apply one that had been removed.

The lookup dedup eviction was re-derived into the hoisted lookup core, and its
target-side signing gate re-added to the request path, where the merge had
dropped it entirely and left the limiter inert.

Two regression tests were repaired rather than accepted. The gap-tracker test
adapted into master's ReceiverState style asserted an empty burst after a jump
to the ceiling counter, which is not what that code does; and the epoch
dampener's stamp-on-acceptance ordering was pinned by nothing, so a sustained
replay could have starved a restarting peer with every test still green. Both
now red when the fix they guard is reverted.

Gate: fmt, build, clippy -D warnings, 2273 tests passing, six repo guards on
the committed tree.
This commit is contained in:
Johnathan Corgan
2026-08-23 15:49:39 +01:00
57 changed files with 6444 additions and 408 deletions
+487
View File
@@ -542,6 +542,15 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
non-positive rate is rejected at config validation rather than silently
refusing every session.
- `node.limits.max_sessions`, defaulting to 1024, which bounds the end-to-end
session table. Zero means unlimited, which restores the previous behaviour
exactly and is the way to back the change out on a running node. The default
is four times the adjacent `node.session.pending_max_destinations`. A
session entry measures 6608 bytes of inline state plus heap, so the table
holds to roughly 7 MB, and a test pins that per-entry figure so the
arithmetic behind the default fails loudly if an entry grows. Existing
configurations parse unchanged, the key being optional.
- `node.rate_limit.established_handshake_burst` and
`node.rate_limit.established_handshake_rate`, the parameters of the new
established-link msg1 token bucket, which meters link-layer msg1 rather
@@ -952,6 +961,33 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
#### Transports & config
- The UDP transport's DNS cache is now bounded and actually evicts. The map
held one entry per distinct hostname string ever dialed, and the TTL was
applied only on the read, so a stale entry was overwritten on the next dial
of the same name and otherwise stayed for the life of the process. Under a
rendezvous policy that accepts advertised endpoints the keys are strings a
remote party chose, which made the growth theirs to drive. A store now
sweeps entries past their TTL and, if the map is still full, drops the
oldest, holding it to 256 hostnames. Refreshing a name already cached
evicts nothing. Eviction is by insertion time rather than last use, so a
rarely dialed name in a very large peer list may re-resolve more often; the
cost of a wrong eviction is one DNS lookup, not a failed dial.
- macOS: stopping an Ethernet transport under load no longer hangs the
process. The BPF reader thread handed each frame to the async consumer with
`blocking_send`, which parks with no way to be woken. Stopping the transport
aborts the consumer first, so nothing drains the 1024-frame channel, and the
socket's `Drop` then joined a thread that could never return: on a busy
interface the daemon had to be killed. The socket now drops the receiver
before joining, which releases a parked send at once, and the reader thread
sends through a helper that watches the same shutdown pipe its `select()`
already honours, so a send waiting for room cannot outlive a shutdown
request. The helper yields before it sleeps, so the saturated-path handoff
rate is unchanged. **Not covered by CI**: the reader thread is macOS-only
and Linux CI compiles none of it. What the tests prove is that the helper
the thread now waits in is cancellable; that a real BPF thread exits under
load still needs a manual check on a Mac.
- A failed private-key write no longer leaves a node silently running an
ephemeral identity. Six write results in the identity path were discarded,
and the sharpest was in `persistent` mode: a failed write to `fips.key` fell
@@ -1048,6 +1084,61 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
#### FMP/FSP session integrity
- A frame whose counter is `u64::MAX` is now refused by the replay window
instead of being accepted as a new high-water mark. Accepting it pinned
`highest` at the ceiling, after which every subsequent counter from that peer
fell more than a replay window below it and was rejected, wedging that peer's
own receive path until a rekey replaced the session. The send side already
refuses to emit that counter (`take_send_counter` and `advance_nonce` both
return a nonce-overflow error), so no conforming peer can produce it and the
refusal is invisible on the wire; the highest counter an honest peer can send,
`u64::MAX - 1`, is still accepted. Reaching this required an
already-authenticated peer running modified code, and the damage was confined
to that peer's own session.
- The MMP gap tracker advances its expected-counter state with a saturating add,
so a received counter of `u64::MAX` no longer overflows it. The wrap silently
reset the expectation to zero in a release build and aborted the task under a
build with overflow checks on, such as the test harness. Behaviour is
unchanged for every counter an honest peer can emit.
- An epoch-mismatch msg1 no longer tears down a peering that is still
carrying authenticated traffic, and a second epoch change for the same peer
identity inside 15 seconds is refused. The epoch travels inside the AEAD, so
such a msg1 is authentic, but it stays authentic after capture: replaying
one destroyed a working peering, and with it the FSP session state that
peering carried, from off the path. The peering's last authenticated inbound
frame is the evidence that it is still alive, and nothing an unauthenticated
sender emits can refresh it, so a peer that genuinely restarted clears the
gate by having stopped sending. The refusal is a silent drop: no msg2 is
returned, since the stored msg2 is bound to the original msg1's ephemeral
and answering a sender-chosen address is free amplification. The interval is
stamped only when an epoch change is accepted, so a sustained replay cannot
starve a genuinely restarting peer. Both thresholds come from one constant,
sized so a restarting peer's msg1 resends still land inside its own first
handshake window and below `link_dead_timeout_secs`, and nothing changes on
the wire.
- Retention of a superseded FSP key epoch is now capped at an absolute
ceiling measured from the cutover, defaulting to 120 seconds against the
10-second drain window. The drain deadline slides forward on every inbound
frame that authenticates against the `previous` slot, which is what keeps a
peer that lost msg3 from having the old epoch erased out from under it, but
it also meant the authenticated peer holding that key could keep the retired
key resident for as long as it kept sealing frames in the old epoch. The
sliding grace is unchanged; it now delays erasure by a bounded amount rather
than preventing it. The ceiling is set to clear the worst-case legitimate
recovery of a peer that lost msg3 (the msg3 resend ladder, then
`handshake_timeout_secs` before the responder abandons, then the rekey
dampening window before it may re-initiate, about 90 seconds at stock
settings), and it is raised automatically if the configured handshake timers
imply a longer budget, so shortening a timer cannot push the ceiling under
the recovery it has to leave room for. Nothing changes on the wire; each side
runs its own drain. A peer that has still not recovered when the ceiling
fires is left with undecryptable frames until its own rekey retry
re-converges the epochs, since nothing tears an established session down on
repeated decrypt failure.
- A session setup message naming an already-established peer no longer replaces
that peer's session. The handler did this whenever `node.rekey.enabled` was
false: it ran a fresh responder handshake and overwrote the entry, discarding
@@ -1154,6 +1245,151 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
#### NAT traversal / Nostr discovery
- An advert or inbox-relay list returned by a relay is now checked against
the peer it claims to describe before anything else looks at it. The relay
pool verifies every event's signature but does not check a reply against
the request filter, and neither of the two options that would make it do so
is enabled, so a relay may answer a request for one author's advert with an
event it signed itself. The stale-advert refetch picked the newest
`created_at` across everything returned, with no author test, and then wrote
the result into the advert cache under the requested peer's npub, so a
single hostile or compromised advert relay could pin an endpoint set of its
own choosing for that peer. The author test now runs before the timestamp
contest rather than after, so a future-dated foreign event cannot even
suppress the genuine advert by winning it. The same filter now applies to
the inbox-relay lookup, where the omission let an attacker-authored relay
list steer this node's direct-message and traversal-signal traffic. A
refetch that comes back with events, none of them signed by the peer, now
leaves the cached entry alone: that is no evidence the advert was
withdrawn, and evicting on it would hand the same relay a way to clear the
cache. A refetch that genuinely comes back empty still evicts.
- An advert's `created_at` is now clamped forward to the same 60s of clock
skew the traversal-signal path already tolerates. An unbounded future
timestamp bought a cache entry a proportionally distant validity horizon
and an unbeatable position in every replacement comparison, so a later
genuine advert could never displace it and the size-cap eviction collected
it last. The clamp applies to the stored timestamp as well as the validity
window, at all three points where an advert is cached, so ordering and
expiry now agree. Clamping rather than refusing the event is deliberate: a
node whose own clock runs slow reads every peer's honest advert as
future-dated, and refusing would silently withdraw Nostr-mediated dialing
for every peer at once.
- Inbound traversal signals are now rate limited before they are decrypted. A
rendezvous-enabled node handed every kind-21059 event straight to the unwrap,
which is two NIP-44 decrypts and a signature verify, inline on the single
task that also routes traversal answers and maintains the advert cache.
Nothing bounded how fast an unauthenticated stranger could schedule that
work: the per-npub offer admission cannot, because it keys on the sender's
public key, which only exists once the first decrypt has already run, and
because it is a concurrency semaphore rather than a limit over time. A token
bucket now sits ahead of the unwrap, so a flood costs a node its inbound
offers instead of the whole notify loop.
What the limit can and cannot key on is worth stating, because it decides the
shape of the fix. Before decryption there is no sender identity at all: the
outer event is signed by a key generated per event, so bucketing on its
author would hand an attacker a fresh allowance for free, and the timestamp
and recipient tag are equally attacker-chosen. The arrival relay is drawn
from our own configured set but is not an isolation boundary either, since an
attacker publishes to the same relays an honest peer does. The shared
allowance is therefore a single global bucket and is indiscriminate by
construction, which on its own would shed our own traversals along with the
attacker's, and since the attacker sets the rate every retry would land in
the same shed. A second, smaller allowance is held in reserve and drawn only
while this node has traversals of its own outstanding, so a flood denies a
node its inbound offers, which nothing receiver-side can prevent without a
pre-decrypt identity, rather than also denying it the answers to offers it
sent. Shed signals are counted and reported at debug level per event with a
warning each time the running total doubles, so a bucket sized below a busy
node's real need shows up in the log rather than as apparent relay flakiness.
Two limits on what this buys. It bounds the crypto path only: the advert
branch runs earlier in the same loop and is not metered here, so a stranger
can still put JSON parsing and a cache insert on the task per event. And the
relay SDK verifies each event's outer signature on its own per-relay task
before this loop ever sees it, which no receiver-side change short of
dropping the subscription can avoid.
- A STUN binding response is now accepted only from the address the binding
request was sent to. The client discarded the source address `recv_from`
returned and let the parser decide, and the parser checks only the message
type, the magic cookie and the 12-byte transaction id. An on-path attacker
who could read the outbound request could therefore inject a reply carrying
a transaction id copied from it, and its chosen address became the reflexive
candidate the node published in a traversal offer or answer, redirecting the
peer's hole-punch packets. Datagrams from any other source are counted and
discarded, and one debug record per STUN attempt reports the count and the
last unexpected source, so a rejection is diagnosable without giving a
flooder control of the log rate. A server that answers from an address other
than the one dialed, which RFC 5389 forbids, now times out and the next
configured server is tried.
- The exemption that lets a peer's reflexive address skip the private-address
gate is now conditional on our own vantage point. That exemption exists for
the deployment whose STUN server sits inside the private network, so the
observed reflexive address is legitimately private; it was applied
unconditionally, so a node whose own STUN result was public still punched
whatever private address a peer named as its reflexive one. Any sender whose
offer or answer was accepted could therefore aim a burst of UDP packets,
carrying this node's source address, at a host inside the node's own private
network, which is the one place the candidate filter was written to keep it
out of. The gate now applies whenever our own reflexive address is public.
It stays lifted when our own reflexive address is itself private, which is
the LAN-STUN deployment the exemption was for, and also when we have no
reflexive address at all, so a failed STUN probe cannot cost a node its
same-LAN peering. Two consequences to state rather than discover: a peer
behind a private STUN server talking to a node with a public one loses its
reflexive candidate, which was never reachable from us in any case, and
because the /24 comparison is IPv4-only a unique-local IPv6 reflexive
address is refused unless our own reflexive address is unique-local too.
An off-subnet refusal of a peer's reflexive address is a shape an honest
deployment now produces, so it no longer raises the refusal record to
warning level on its own; the never-routable, port-0 and unparsable classes
still do.
- A peer's candidate list is now bounded before it is walked rather than only
after. The eight-target cap ran after both planning loops had finished, so
it bounded what a node punched but not what it spent deciding: a signal
naming several thousand candidates had every one of them parsed and vetted,
and the deduplicating scan that follows is quadratic in the plan those
candidates feed. At most 32 candidates are now vetted, four times the
target cap and four times what the candidate generator produces on the
widest host, and the excess is discarded rather than failing the offer, so
an honest many-homed peer loses the tail of its list instead of its
traversal. The refusal record carries the discarded count as a new
`over_offered` field and treats a non-zero one as an attack shape, since
nothing honest reaches the bound.
- A NAT-punch packet is now accepted only from an address this node planned to
probe. The punch packet's discriminator is a plain digest of the session id,
a value both peers already know, and it travels in the clear in every probe,
so acceptance proved only that the sender had seen one. The receive loop
broke on the first packet whose digest matched, whatever its source, and
returned that source as the peer address, so anyone who observed a probe, or
who could reach the node and guess the session id, could have an arbitrary
address adopted as the peer: the legitimate traversal was denied, the Noise
handshake and its retransmissions went to an address of the attacker's
choosing, and the pair was charged a failure against its backoff state. The
npub-pinned handshake still could not authenticate to the wrong host, so this
was a denial and a misdirection rather than an impersonation. The source
address is now ranked against the planned target list before anything else:
an unplanned source is dropped and, deliberately, is not acked either, since
acking it is a reflection the node controls. A source matching a planned
target exactly is adopted immediately, as before. A source matching a planned
target's IP on a different port is what a symmetric NAT's fresh mapping looks
like, and it is still adopted, because that is the main class of NAT pairing
punching exists to rescue; it is held as a candidate for 250 ms first, so an
exact match arriving inside that window supersedes it. The honest path's
latency is unchanged. Two consequences to state rather than discover: an
attacker that can source packets from a planned target's IP on any port is
still accepted, which is the residue only an authenticated probe can close;
and an attempt under a flood of spoofed matching packets now runs to its full
timeout instead of ending on the first one, so the refused sources are
counted and reported once when the attempt ends rather than logged per
packet.
- Traversal punch targets taken from a peer's offer or answer are now
filtered and bounded. A rendezvous-enabled node previously punched every
address a signed offer named, including loopback, link-local, multicast,
@@ -1206,8 +1442,104 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
attributes the acceptance to clock skew, since a peer configured with a longer
signalling TTL than ours now reaches it too.
#### DNS responder
- The DNS responder's mesh-interface filter now works on macOS and FreeBSD,
where it had never run. The filter drops `.fips` queries that arrive over the
mesh TUN, which is what keeps a widened `dns.bind_addr` from exposing the
hosts file's alias space to every mesh peer. It was keyed on the interface
index resolved from the *configured* TUN name, but macOS and FreeBSD assign
the device a name of the kernel's choosing (`utunN`, `tunN`), so the lookup
found nothing, the index came back `None`, and `None` disables the filter.
The index is now resolved from the name of the device the node actually
created, which the TUN startup path already records, and a live device whose
index will not resolve is logged rather than passed off as "no mesh
interface". Linux is unaffected, since the configured name is the device's
name there. **Behaviour change on macOS and FreeBSD**: a node with a
non-loopback `dns.bind_addr` stops answering `.fips` queries that arrive over
the mesh interface. **What this does not close**: with an app-owned TUN the
node never learns a device name, so the filter stays off there. **Not
measured**: whether macOS and FreeBSD attribute a locally originated query
sent to the node's own mesh address to the TUN interface, as Linux does. If
they do, such a query is now dropped on those platforms; the shipped resolver
drop-in targets `[::1]` rather than the mesh address, so the packaged path is
not affected.
#### Data-plane / routing signals
- A transit node's induced routing errors are now bounded by the authenticated
link peer that induced them. The 100 ms suppression gate on
`CoordsRequired`, `PathBroken` and `MtuExceeded` was keyed on the failed
datagram's destination address, which is an envelope field the sender picks,
so a fresh random destination on every packet was always a first sighting and
every packet was admitted. Each admission also inserted a key and then walked
the whole map, so per-packet cost grew with the flood rate while the sender's
cost stayed flat, and the error itself is addressed to the datagram's source
address, which nothing binds to the sender either. A new per-link-peer token
bucket, 20 signals a second sustained with a burst of 50, is now consulted
first, keyed on the AEAD-authenticated peer the frame arrived over: the one
value at that point a sender cannot mint. The per-destination interval is
kept unchanged behind it, because it still does the aggregate suppression a
genuine outage needs, and no gate was added on the address the error is
returned to, which would have handed a sender a way to silence honest errors
toward a victim it names by keeping that victim's key hot.
Two ordering choices in there rather than left to be discovered. The peer's
token is peeked and only spent once the destination gate has also admitted,
so a single unroutable destination behind a high-fanout peer cannot burn that
peer's whole budget on signals nothing sends and silence every other
destination behind it. And the destination map now carries a hard ceiling of
4096 entries with its expiry sweep amortized to once per eviction interval
rather than run on every admission; when it is full it admits without
recording rather than refusing, because refusing would turn a full map into
node-wide silence exactly during partition healing, when many destinations
are legitimately unroutable at once. Emission stays bounded by the peer
budget in that state. Three counters, rendered on the fipstop Routing tab,
make each of the three outcomes visible instead of silent.
- A transit-emitted `PathBroken` no longer carries the reporter's cached
coordinates for the unreachable destination. The signal is returned to the
datagram's source address, so anyone able to reach the node could name any
address and have the node's coordinate cache read back to them, one entry per
packet. The field is optional on the wire and no receiver reads it, so this
is an emission change only: an unmodified peer parses the frame exactly as
before. Which of the two signals is emitted still discloses whether the entry
exists.
- A reactive `MtuExceeded` is now believed only when this node has actually
sent a frame larger than the bottleneck it reports. The signal is
unauthenticated: the admission gate narrows which destination may be named
but cannot say who named it, so a value at the floor was a legal value from
anyone, and one datagram drove a bound session's path MTU to 256 and pinned
the address-keyed entry the SYN-time MSS clamp reads, with recovery costing
three consecutive higher notifications across two notification intervals.
Each session now carries the largest frame this node has put on the wire
toward it since the last accepted decrease, and a report is refused unless it
names something smaller. Honest path-MTU discovery satisfies that by
construction, because the report exists only because a frame we sent did not
fit; a forgery has to wait for us to emit something bigger than the value it
wants to claim, which bounds every accepted claim from below by our own
traffic. The evidence is cleared on each accepted decrease and on release, so
one large send early in a session cannot vouch for the rest of it.
The guard sits ahead of both effects rather than between them, which is also
where the existing floor check moved to: the floor previously ran after the
session's own path MTU had already been changed and so governed only the
lookup table. The reactive carrier now names its own floor constant, held
equal to the actionable floor so no hop legitimately configured with a small
transport MTU loses its feedback; corroboration, not the floor's value, is
what stops a legal-but-forged claim. A separate counter, rendered on the
fipstop Routing tab, distinguishes an uncorroborated refusal from a
below-floor one.
- The path-MTU release a `PathBroken` drives is now rate limited per
destination on a budget of its own. That signal is unauthenticated too, and
the release discards a bottleneck this node learned by having a packet
dropped, so repeating the claim discarded a genuine value as fast as it could
be relearned. The limiter is a separate instance rather than the one the
coordinate warmup send already uses: a budget another signal can spend is not
a bound. Deferring a release is the safe direction, since the value kept is
the tighter one.
- The influence a remote party has over path MTU is now bounded, and the
per-destination path MTU cache has a way back. The `path_mtu` field is an
unsigned per-hop transit annotation carried outside the signed proof, and the
@@ -1282,8 +1614,139 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
claimed source and destination pairing no honest forwarder could produce.
The drop log line now carries the signal type and the refusal class.
- A discovery lookup response is now acted on only when it answers a lookup
this node actually has outstanding. The originator path took any response
whose `request_id` was not in the transit dedup map, so an admitted peer
could harvest one genuine signed response for a target and re-inject it at
will: each injection cleared the victim's in-flight lookup, recorded a
reachability success for a target that might be unreachable, refreshed the
cached coordinates for a further full TTL, and flushed the victim's queued
packets onto a route at a moment the sender chose. It also reached the
signature verify before any check that the response was wanted, so the
verify was the first cost gate on the path. The node now records the
`request_id` of every lookup request it sends on that target's pending
entry, and a response is dropped unless it names a target with a lookup
outstanding and carries one of the ids issued for it. Because the id is
fresh 64-bit randomness drawn per attempt and the target signs over it, a
harvested response is bound to the request it answered and cannot be
redirected or replayed. The check runs before the identity-cache resolve
and before the signature verify, so a response nobody asked for costs
nothing. Replies to earlier attempts of a still-outstanding lookup are
still accepted, which is the common case on a link whose round trip
exceeds the first rung of the retry ladder. Drops are counted as
`resp_unsolicited`, visible through `show routing`, `show metrics` and the
fipstop routing pane; the counter has a nonzero floor in healthy operation,
because a request is flooded to every qualifying tree peer and the
duplicate replies land there once the first has been accepted.
- A flooded discovery dedup cache no longer makes a node unresolvable. The
cache is both the duplicate filter and the reverse-path table for lookup
responses, and at its 4096-entry bound it dropped the arriving request.
That drop sat ahead of both the check for whether the request names this
node and the forwarding path, so one link peer emitting fresh request_ids
could stop the node answering lookups for itself and stop it carrying
anyone else's, for as long as it kept the cache full. The cache now makes
room instead of refusing: over a peer's own share it drops that peer's
oldest entry, and at global capacity it drops the oldest entry of whichever
peer holds the most, so a light peer's reverse path is never taken to admit
a heavy one and extra identities buy a flooder proportionally less. A
peer's share is the cache divided by the current link-peer count, with a
floor of 64. The loosening this accepts is that an evicted request_id
arriving again inside the window is forwarded a second time rather than
recognised as a duplicate, which the per-target forward limiter and TTL
already bound. Evictions are counted as `req_dedup_evicted`; the old
`req_dedup_cache_full` counter stays in place, frozen at zero, so a
dashboard carried across versions does not lose the series.
- Answering a lookup for ourselves is now metered per link peer. The response
proof is signed over the requester's `request_id`, so every request
addressed to this node costs a fresh Schnorr signature that cannot be
cached or served twice, and until now the only thing bounding that rate was
the dedup cache filling up, which is the defect above. A token bucket per
link peer, 256 signatures of burst refilling at 32 per second, absorbs the
legitimate burst that follows a topology change, when many correspondents
re-look-up at once through the few links that lead here, while capping what
one neighbour can make the node sign. Refusals are counted as
`req_sign_rate_limited` and visible in `show routing`, `show metrics` and
the fipstop routing pane. A refused request keeps its dedup entry, and
retries carry fresh request_ids, so a refusal cannot suppress the retry.
#### Admission / peer caps
- The Ethernet transport's discovery buffer is now bounded and no longer costs
a linear scan per beacon. Beacons are unauthenticated broadcast frames, and
the buffer deduplicated by scanning a `Vec` for the source MAC and had no
cap, so anything on the segment could name a fresh MAC per frame and drive
both quadratic CPU in the receive loop and unbounded memory. It is drained
once per tick only while the transport is operational, so a transport that
is receiving but not operational was never drained at all. The buffer is now
a map keyed on source MAC, capped at 1024 distinct MACs between drains, with
the drain order still oldest sighting first so which neighbour gets dialed
under a connect budget does not depend on hash iteration order. A MAC already
buffered is always refreshed, so a flood of new MACs cannot crowd out a
neighbour already seen. Refused beacons are counted in the transport's stats
as `beacons_dropped` and reported in the log on the first drop and then on
each power-of-ten thereafter, so the flooder does not set the log rate.
**What this does not close**: a flood can still crowd out a neighbour not yet
seen in that tick, and anything able to flood raw frames on the segment can
already jam the beacon at L2 more cheaply.
- A read failure on `peers.allow`, `peers.deny` or the `hosts` file no longer
turns the node into an open one. Every read error other than a steadily
absent file was logged and swallowed, leaving that file's entries empty, and
the reloader published the result unconditionally: an unreadable `peers.deny`
admitted the peers it named, and an unreadable `peers.allow` took a node from
admitting a named few to admitting everyone, with one warning line as the
only signal. Because the recorded modification times advanced before the
load, nothing retried until the file changed again, so a persistent
permission or I/O fault left the empty ACL in force indefinitely. The
reloader now keeps the last loaded ACL when any input is present but
unreadable, leaves the modification times alone, retries on the next tick
regardless of them, and logs the fault once on the transition rather than
once per tick. An absent file is still a policy and still loads as an empty
set; a `NotFound` that a successful stat contradicts is treated as a file
being rewritten under us and held. A reload whose inputs all read cleanly but
which empties an enforcing ACL while its files are still on disk is held for
one tick, which catches a read that caught a non-atomic in-place edit
mid-write, and released on the next so a deliberate blanking still takes
effect. `fipsctl` ACL status gains a `stale` flag reporting that the policy
in force is older than the files on disk. **What this does not close**:
there is no last-good snapshot at startup, so a node whose ACL file is
unreadable at boot still comes up with no entries, now logged as an error and
armed to retry on the first tick. Admission is also checked only at handshake
time, so a peer admitted during a window that has already happened keeps its
link.
- The end-to-end session table now has a bound. It was the one remotely-grown
map with none: an inbound SessionSetup naming an address nobody had seen
inserted an entry, and the two existing limits did not reach it, the setup
limiter governing the arrival rate rather than the population and the idle
purge only reaching entries a peer stops using. One neighbour sending setups
at the permitted rate could hold roughly 1440 half-open entries at any
moment and grow the table without limit by keeping them warm. Setups that
would grow the table past `node.limits.max_sessions` are now refused, ahead
of the setup limiter, so a full table costs no token, no responder handshake
and no ack; a refused setup emits nothing at all, which is indistinguishable
from loss to the sender and is already covered by its own msg1 resend
schedule. The test is whether admitting would grow the table, not whether
the sender is a stranger, so a resent setup for an entry already present is
still served and an in-flight handshake is not broken. Unauthenticated
half-open entries are additionally held to half the table, so a handshake
flood cannot deny the whole of it to peers that complete; that share is sized
to leave a reconnect storm, where every peer initiates at once after a
restart or a healed partition, room to land. Locally originated sessions are
capped at the same ceiling, answered with ICMPv6 destination unreachable so
the application gets an immediate error rather than a silent drop. The cap
refuses rather than evicts: the setup that triggers the decision is
unauthenticated at that point, so evicting would hand a stranger a way to
tear down sessions it has nothing to do with. Refusals are counted as
`table_full` and `half_open_full` in the session reject family. What stays
open is per-neighbour fairness among established sessions: one hostile
neighbour that completes handshakes and keeps each session warm can occupy
the table and hold new session establishment closed for as long as it keeps
doing so, which is a denial of new sessions rather than the unbounded memory
growth it replaces.
- An accepted inbound TCP connection no longer holds a slot indefinitely
without sending anything. The cap was tested at accept and the pool insert
and counter bump followed with no read in between, while the frame reader's
@@ -1322,6 +1785,30 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
NXDOMAIN. Connecting the socket also means a dead upstream surfaces
ECONNREFUSED immediately instead of stalling for five seconds.
#### Control socket
- The control socket and the directory holding it are now created with a
restrictive mode rather than created wide and narrowed afterwards. `bind(2)`
makes the socket inode `0777 & ~umask`, so under a permissive umask the
socket was world-accessible for the window between the bind and the `chmod`
to 0770 that followed it; the bind now runs under a umask that masks the
"other" bits, so the inode is 0770 from creation and the chmod and chown stay
the authority on its final mode. The parent directory was worse than a
window: it was created with `create_dir_all`, which is also `0777 & ~umask`,
and nothing ever set a mode on it, so under a permissive umask the directory
holding the socket stayed world-writable for the life of the host, and a
world-writable parent lets an unprivileged account plant an entry at the
socket path. Directories this code creates now come out 0750, which is what
the systemd unit (`RuntimeDirectoryMode=0750`) and the FreeBSD rc script
(`install -d -m 0750`) already apply, so no packaged deployment sees a
different mode and no `fipsctl` user loses access. Both the daemon and the
gateway control sockets are covered. **What this does not close**: the window
between the stale-socket probe and the bind is documented at the site rather
than removed. Reaching it needs write access to the socket's parent
directory, which the packaged layouts give to root alone, and an account
holding it can deny the daemon its socket more simply by squatting the path
first.
#### Key material and identity files
- Private key writes no longer follow a symlink, and the key file's mode is
+10
View File
@@ -192,6 +192,8 @@ fn draw_routing_stats(
("Bloom Miss", lookup("req_bloom_miss")),
("Backoff Suppressed", lookup("req_backoff_suppressed")),
("Fwd Rate Limited", lookup("req_forward_rate_limited")),
("Sign Rate Limited", lookup("req_sign_rate_limited")),
("Dedup Evicted", lookup("req_dedup_evicted")),
("TTL Exhausted", lookup("req_ttl_exhausted")),
("Decode Error", lookup("req_decode_error")),
],
@@ -206,6 +208,7 @@ fn draw_routing_stats(
("Timed Out", lookup("resp_timed_out")),
("Identity Miss", lookup("resp_identity_miss")),
("Proof Failed", lookup("resp_proof_failed")),
("Unsolicited", lookup("resp_unsolicited")),
("Decode Error", lookup("resp_decode_error")),
],
));
@@ -298,6 +301,13 @@ fn draw_routing_stats(
("Path Broken Refused", err("unbound_broken")),
("MTU Exceeded Refused", err("unbound_mtu")),
("Forged Pairing", err("unbound_forged")),
("Emit Over Peer Budget", err("emit_over_peer_budget")),
("Emit Over Dest Interval", err("emit_over_dest_interval")),
("Emit Limiter At Capacity", err("emit_limiter_at_capacity")),
(
"MTU Exceeded Uncorroborated",
err("mtu_exceeded_uncorroborated"),
),
],
));
right.push(Line::from(""));
+1 -1
View File
@@ -1151,7 +1151,7 @@ fn routing_focused_pane_scrolls() {
// column is the taller of the two, so scrolling fully to the bottom would
// over-scroll the right column past Congestion; this offset lands the
// Congestion region inside the short window instead.
app1.scroll_offsets.insert((Tab::Routing, 2), 28);
app1.scroll_offsets.insert((Tab::Routing, 2), 32);
let buf1 = testkit::render(100, 20, |frame, area| {
super::routing::draw(frame, &app1, area);
});
+18
View File
@@ -28,6 +28,20 @@ pub struct LimitsConfig {
/// Max pending inbound handshakes (`node.limits.max_pending_inbound`).
#[serde(default = "LimitsConfig::default_max_pending_inbound")]
pub max_pending_inbound: usize,
/// Max end-to-end sessions (`node.limits.max_sessions`), `0` = unlimited.
///
/// The session table is the only remotely-grown map with no bound: an
/// inbound SessionSetup from an address nobody has seen inserts an
/// entry, and the idle purge only reaches entries a peer stops using.
/// The default of 1024 is four times the adjacent
/// `node.session.pending_max_destinations`. One entry measures 6608
/// bytes of inline state plus heap, so the table holds to roughly 7 MB
/// and a test pins the per-entry figure the default rests on. Raising it
/// raises the memory an attacker can make this node hold; lowering it
/// refuses new sessions sooner on a node that legitimately talks
/// end-to-end to many others, such as a gateway.
#[serde(default = "LimitsConfig::default_max_sessions")]
pub max_sessions: usize,
}
impl Default for LimitsConfig {
@@ -37,6 +51,7 @@ impl Default for LimitsConfig {
max_peers: 128,
max_links: 256,
max_pending_inbound: 1000,
max_sessions: 1024,
}
}
}
@@ -54,6 +69,9 @@ impl LimitsConfig {
fn default_max_pending_inbound() -> usize {
1000
}
fn default_max_sessions() -> usize {
1024
}
}
/// Rate limiting (`node.rate_limit.*`).
+1
View File
@@ -134,6 +134,7 @@ fn empty_acl_status() -> PeerAclStatus {
deny_file_entries: Vec::new(),
allow_entries: Vec::new(),
deny_entries: Vec::new(),
stale: false,
}
}
+12 -2
View File
@@ -12,6 +12,7 @@
"req_bloom_miss": 0,
"req_decode_error": 0,
"req_dedup_cache_full": 0,
"req_dedup_evicted": 0,
"req_deduplicated": 0,
"req_duplicate": 0,
"req_fallback_forwarded": 0,
@@ -20,6 +21,7 @@
"req_initiated": 0,
"req_no_tree_peer": 0,
"req_received": 0,
"req_sign_rate_limited": 0,
"req_target_is_us": 0,
"req_ttl_exhausted": 0,
"resp_accepted": 0,
@@ -29,13 +31,18 @@
"resp_no_route": 0,
"resp_proof_failed": 0,
"resp_received": 0,
"resp_timed_out": 0
"resp_timed_out": 0,
"resp_unsolicited": 0
},
"error_signals": {
"coords_required": 0,
"emit_limiter_at_capacity": 0,
"emit_over_dest_interval": 0,
"emit_over_peer_budget": 0,
"lookup_resp_mtu_below_floor": 0,
"mtu_exceeded": 0,
"mtu_exceeded_below_floor": 0,
"mtu_exceeded_uncorroborated": 0,
"path_broken": 0,
"path_mtu_notif_below_floor": 0,
"unbound_broken": 0,
@@ -77,6 +84,7 @@
"req_bloom_miss": 0,
"req_decode_error": 0,
"req_dedup_cache_full": 0,
"req_dedup_evicted": 0,
"req_deduplicated": 0,
"req_duplicate": 0,
"req_fallback_forwarded": 0,
@@ -85,6 +93,7 @@
"req_initiated": 0,
"req_no_tree_peer": 0,
"req_received": 0,
"req_sign_rate_limited": 0,
"req_target_is_us": 0,
"req_ttl_exhausted": 0,
"resp_accepted": 0,
@@ -94,7 +103,8 @@
"resp_no_route": 0,
"resp_proof_failed": 0,
"resp_received": 0,
"resp_timed_out": 0
"resp_timed_out": 0,
"resp_unsolicited": 0
},
"pending_lookups": [],
"pending_tun_destinations": 0,
+327 -31
View File
@@ -20,7 +20,7 @@ use std::fmt;
use std::path::{Path, PathBuf};
use std::sync::Arc;
use std::time::SystemTime;
use tracing::{debug, info, warn};
use tracing::{debug, error, info, warn};
/// Default path for the peer allow list.
///
@@ -106,6 +106,28 @@ pub enum PeerAclContext {
OutboundHandshake,
}
/// How many consecutive reloads the empty-snapshot guard may hold back.
///
/// A reload whose files all read cleanly but which yields an empty ACL where
/// an enforcing one was in force is more likely a read that raced an in-place
/// rewrite than a policy change, so the previous snapshot is held. The guard
/// releases after this many holds, so an operator who deliberately blanks an
/// ACL file in place still converges, one tick late. Raising it widens the
/// window in which a genuine emptying is ignored; setting it to zero disables
/// the torn-read protection.
const EMPTY_ACL_HOLD_LIMIT: u32 = 1;
/// A peer ACL input file exists but could not be read.
#[derive(Debug, thiserror::Error)]
#[error("failed to read {}: {source}", path.display())]
pub struct AclLoadError {
/// The file whose read failed.
pub path: PathBuf,
/// The underlying I/O failure.
#[source]
pub source: std::io::Error,
}
/// Snapshot of the currently loaded ACL state.
#[derive(Debug, Clone, PartialEq, Eq, Serialize)]
pub struct PeerAclStatus {
@@ -120,6 +142,9 @@ pub struct PeerAclStatus {
pub deny_file_entries: Vec<String>,
pub allow_entries: Vec<String>,
pub deny_entries: Vec<String>,
/// Whether the ACL in force is older than the files on disk because a
/// reload input could not be read.
pub stale: bool,
}
impl fmt::Display for PeerAclContext {
@@ -155,26 +180,38 @@ impl PeerAcl {
#[cfg(test)]
pub fn load_files(allow_path: &Path, deny_path: &Path) -> Self {
let hosts = HostMap::new();
Self::load_files_with_hosts(allow_path, deny_path, &hosts)
Self::try_load_files_with_hosts(allow_path, deny_path, &hosts).unwrap()
}
/// Load the allow/deny files into a new ACL using alias resolution.
pub fn load_files_with_hosts(allow_path: &Path, deny_path: &Path, hosts: &HostMap) -> Self {
///
/// An absent file is a policy and contributes an empty set; a file that
/// is present and unreadable is a fault and is returned as an error, so
/// the caller can keep enforcing whatever it loaded last rather than
/// silently becoming an open node.
pub fn try_load_files_with_hosts(
allow_path: &Path,
deny_path: &Path,
hosts: &HostMap,
) -> Result<Self, AclLoadError> {
let mut acl = Self::new();
acl.load_file(allow_path, true, hosts);
acl.load_file(deny_path, false, hosts);
acl.load_file(allow_path, true, hosts)?;
acl.load_file(deny_path, false, hosts)?;
acl.log_loaded();
Ok(acl)
}
if !acl.is_empty() {
/// Log the shape of a freshly loaded ACL, unless it has no entries.
fn log_loaded(&self) {
if !self.is_empty() {
debug!(
allow_entries = acl.allow.len(),
deny_entries = acl.deny.len(),
allow_all = acl.allow_all,
deny_all = acl.deny_all,
allow_entries = self.allow.len(),
deny_entries = self.deny.len(),
allow_all = self.allow_all,
deny_all = self.deny_all,
"Loaded peer ACL files"
);
}
acl
}
/// Evaluate whether a peer is allowed.
@@ -245,16 +282,29 @@ impl PeerAcl {
self.deny_file_entries.iter().cloned().collect()
}
fn load_file(&mut self, path: &Path, is_allow: bool, hosts: &HostMap) {
/// Merge one ACL file into this ACL.
///
/// An absent file is a policy and an unreadable one is a fault, and
/// `NotFound` alone does not say which: a stat that still finds the file
/// after the read missed it means the file is being rewritten under us,
/// which is transient and must not be published as an empty policy.
fn load_file(
&mut self,
path: &Path,
is_allow: bool,
hosts: &HostMap,
) -> Result<(), AclLoadError> {
let contents = match std::fs::read_to_string(path) {
Ok(c) => c,
Err(e) if e.kind() == std::io::ErrorKind::NotFound => {
Err(e) if e.kind() == std::io::ErrorKind::NotFound && file_mtime(path).is_none() => {
debug!(path = %path.display(), "No ACL file found, skipping");
return;
return Ok(());
}
Err(e) => {
warn!(path = %path.display(), error = %e, "Failed to read ACL file");
return;
return Err(AclLoadError {
path: path.to_path_buf(),
source: e,
});
}
};
@@ -310,6 +360,8 @@ impl PeerAcl {
self.deny_npubs.insert(resolved_npub);
}
}
Ok(())
}
fn resolve_entry(entry: &str, hosts: &HostMap) -> Result<(PeerIdentity, String), String> {
@@ -342,6 +394,13 @@ pub struct PeerAclReloader {
deny_path: PathBuf,
last_allow_mtime: Option<SystemTime>,
last_deny_mtime: Option<SystemTime>,
/// Set while a reload input is unreadable. Forces the next reload
/// attempt regardless of mtimes, because the mtime comparison alone
/// cannot see a change the hosts reloader has already consumed, and
/// gates the fault log to the transition into the held state.
retry_pending: bool,
/// Consecutive reloads held back by the empty-snapshot guard.
empty_holds: u32,
}
impl PeerAclReloader {
@@ -377,7 +436,23 @@ impl PeerAclReloader {
let last_allow_mtime = file_mtime(&allow_path);
let last_deny_mtime = file_mtime(&deny_path);
let hosts = HostMapReloader::new(base_hosts, hosts_path);
let acl = PeerAcl::load_files_with_hosts(&allow_path, &deny_path, hosts.hosts());
// There is no last-good snapshot to hold at startup, so an
// unreadable file still comes up on an empty ACL, as it always has.
// It is logged as the fault it is and armed for retry, so the first
// tick after the file becomes readable enforces the real policy.
let (acl, retry_pending) =
match PeerAcl::try_load_files_with_hosts(&allow_path, &deny_path, hosts.hosts()) {
Ok(acl) => (acl, false),
Err(e) => {
error!(
path = %e.path.display(),
error = %e.source,
"Peer ACL file is present but unreadable; starting with no ACL entries"
);
(PeerAcl::new(), true)
}
};
Self {
acl: arc_swap::ArcSwap::from(Arc::new(acl)),
@@ -386,9 +461,28 @@ impl PeerAclReloader {
deny_path,
last_allow_mtime,
last_deny_mtime,
retry_pending,
empty_holds: 0,
}
}
/// Keep the published snapshot after a reload input failed to read.
///
/// Leaves the recorded mtimes and the ACL in force untouched, arms the
/// retry so the next tick reloads regardless of mtimes, and logs the
/// fault once, on the transition into the held state, rather than once
/// per tick for as long as the fault lasts.
fn hold_snapshot(&mut self, path: &Path, error: &dyn fmt::Display) {
if !self.retry_pending {
error!(
path = %path.display(),
error = %error,
"Peer ACL input is unreadable; holding the last loaded ACL"
);
}
self.retry_pending = true;
}
/// Acquire a lock-free guard over the current ACL snapshot.
pub fn acl(&self) -> arc_swap::Guard<Arc<PeerAcl>> {
self.load()
@@ -409,6 +503,7 @@ impl PeerAclReloader {
deny_file_entries: acl.deny_file_entries(),
allow_entries: acl.allow_entries(),
deny_entries: acl.deny_entries(),
stale: self.retry_pending,
}
}
}
@@ -419,19 +514,70 @@ impl Reloadable for PeerAclReloader {
async fn reload(&mut self) -> bool {
let allow_mtime = file_mtime(&self.allow_path);
let deny_mtime = file_mtime(&self.deny_path);
let hosts_changed = self.hosts.check_reload();
let hosts_changed = match self.hosts.try_check_reload() {
Ok(changed) => changed,
Err(e) => {
let path = self.hosts.path().to_path_buf();
self.hold_snapshot(&path, &e);
return false;
}
};
if allow_mtime == self.last_allow_mtime
&& deny_mtime == self.last_deny_mtime
&& !hosts_changed
&& !self.retry_pending
{
return false;
}
let new_acl = match PeerAcl::try_load_files_with_hosts(
&self.allow_path,
&self.deny_path,
self.hosts.hosts(),
) {
Ok(acl) => acl,
Err(e) => {
self.hold_snapshot(&e.path.clone(), &e.source);
return false;
}
};
// Every input read cleanly and the policy still evaporated. With the
// ACL files themselves freshly written and still on disk that is more
// likely a read that caught one mid-rewrite than an operator emptying
// both lists, so hold and look again next tick. Deleting a file, or
// dropping the aliases an entry resolved through, remains an
// unambiguous way to say "no policy" and is published immediately.
let acl_files_changed =
allow_mtime != self.last_allow_mtime || deny_mtime != self.last_deny_mtime;
if new_acl.is_empty()
&& !self.acl.load().is_empty()
&& acl_files_changed
&& (allow_mtime.is_some() || deny_mtime.is_some())
&& self.empty_holds < EMPTY_ACL_HOLD_LIMIT
{
self.empty_holds += 1;
self.retry_pending = true;
warn!(
allow_file = %self.allow_path.display(),
deny_file = %self.deny_path.display(),
"Peer ACL reload emptied an enforcing ACL; holding the last loaded ACL"
);
return false;
}
if self.retry_pending {
info!(
allow_file = %self.allow_path.display(),
deny_file = %self.deny_path.display(),
"Peer ACL inputs read cleanly again; publishing the files on disk"
);
}
self.retry_pending = false;
self.empty_holds = 0;
self.last_allow_mtime = allow_mtime;
self.last_deny_mtime = deny_mtime;
let new_acl =
PeerAcl::load_files_with_hosts(&self.allow_path, &self.deny_path, self.hosts.hosts());
info!(
allow_file = %self.allow_path.display(),
@@ -786,19 +932,15 @@ mod tests {
}
#[test]
fn test_acl_read_error_is_ignored() {
fn test_acl_read_error_is_reported_rather_than_yielding_an_empty_acl() {
let dir = tempfile::tempdir().unwrap();
let allow = dir.path().join("peers.allow");
let deny = dir.path().join("peers.deny");
std::fs::create_dir(&allow).unwrap();
let acl = PeerAcl::load_files(&allow, &deny);
let err = PeerAcl::try_load_files_with_hosts(&allow, &deny, &HostMap::new()).unwrap_err();
assert!(acl.is_empty());
assert_eq!(
acl.check(&test_peer(&test_npub())),
PeerAclDecision::DefaultAllow
);
assert_eq!(err.path, allow);
}
#[test]
@@ -812,7 +954,7 @@ mod tests {
hosts.insert("node-a", &npub).unwrap();
write_file(&allow, "NODE-A\n");
let acl = PeerAcl::load_files_with_hosts(&allow, &deny, &hosts);
let acl = PeerAcl::try_load_files_with_hosts(&allow, &deny, &hosts).unwrap();
assert_eq!(acl.allow_file_entries(), vec!["NODE-A".to_string()]);
assert_eq!(acl.allow_entries(), vec![npub.clone()]);
@@ -830,7 +972,7 @@ mod tests {
hosts.insert("node-a", &npub).unwrap();
write_file(&allow, &format!("node-a\n{npub}\nnode-a\n"));
let acl = PeerAcl::load_files_with_hosts(&allow, &deny, &hosts);
let acl = PeerAcl::try_load_files_with_hosts(&allow, &deny, &hosts).unwrap();
assert_eq!(
acl.allow_file_entries(),
@@ -887,6 +1029,160 @@ mod tests {
);
}
/// The three permission-fault tests below make a file unreadable through
/// the unix mode bits, which Windows has no equivalent for: a read-only
/// NTFS file is still readable, so the fault they need cannot be produced.
/// They are gated to unix rather than made to pass vacuously elsewhere.
///
/// **Coverage gap**: on Windows nothing exercises the reloader's
/// unreadable-input path, so the fail-open defect this fix closes is
/// unverified there.
/// Make a file unreadable, returning false if the effective uid can read
/// it anyway. Root bypasses the mode bits, so the permission-fault tests
/// cannot run there and skip instead of passing vacuously; that leaves
/// the EACCES path unexercised in any root CI job.
#[cfg(unix)]
fn make_unreadable(path: &Path) -> bool {
use std::os::unix::fs::PermissionsExt;
std::fs::set_permissions(path, std::fs::Permissions::from_mode(0o000)).unwrap();
std::fs::read_to_string(path).is_err()
}
#[cfg(unix)]
fn make_readable(path: &Path) {
use std::os::unix::fs::PermissionsExt;
std::fs::set_permissions(path, std::fs::Permissions::from_mode(0o644)).unwrap();
}
#[cfg(unix)]
#[tokio::test]
async fn test_acl_reload_holds_last_good_snapshot_when_deny_file_unreadable() {
let dir = tempfile::tempdir().unwrap();
let allow = dir.path().join("peers.allow");
let deny = dir.path().join("peers.deny");
let denied = test_npub();
write_file(&deny, &format!("{denied}\n"));
let mut reloader = PeerAclReloader::with_paths(allow, deny.clone());
assert_eq!(
reloader.acl().check(&test_peer(&denied)),
PeerAclDecision::DenyList
);
// Rewrite before revoking access so the mtime change makes the
// reloader actually attempt the read that then fails.
std::thread::sleep(std::time::Duration::from_millis(5));
write_file(&deny, &format!("{denied}\n"));
if !make_unreadable(&deny) {
return;
}
assert!(!reloader.reload().await);
assert_eq!(
reloader.acl().check(&test_peer(&denied)),
PeerAclDecision::DenyList
);
assert_eq!(reloader.acl().effective_mode(), "denylist");
assert!(reloader.status().stale);
make_readable(&deny);
}
#[cfg(unix)]
#[tokio::test]
async fn test_acl_reload_retries_after_a_transient_read_error() {
let dir = tempfile::tempdir().unwrap();
let allow = dir.path().join("peers.allow");
let deny = dir.path().join("peers.deny");
let denied = test_npub();
write_file(&deny, &format!("{denied}\n"));
let mut reloader = PeerAclReloader::with_paths(allow, deny.clone());
std::thread::sleep(std::time::Duration::from_millis(5));
write_file(&deny, &format!("{denied}\n"));
if !make_unreadable(&deny) {
return;
}
assert!(!reloader.reload().await);
// No further mtime change: only the armed retry can pick this up.
make_readable(&deny);
assert!(reloader.reload().await);
assert_eq!(
reloader.acl().check(&test_peer(&denied)),
PeerAclDecision::DenyList
);
assert!(!reloader.status().stale);
}
#[cfg(unix)]
#[tokio::test]
async fn test_acl_reload_holds_last_good_when_the_hosts_file_becomes_unreadable() {
let dir = tempfile::tempdir().unwrap();
let allow = dir.path().join("peers.allow");
let deny = dir.path().join("peers.deny");
let hosts = dir.path().join("hosts");
let npub = test_npub();
write_file(&allow, "node-a\n");
write_file(&hosts, &format!("node-a {npub}\n"));
let mut reloader =
PeerAclReloader::with_alias_sources(allow, deny, HostMap::new(), hosts.clone());
assert_eq!(
reloader.acl().check(&test_peer(&npub)),
PeerAclDecision::AllowList
);
std::thread::sleep(std::time::Duration::from_millis(5));
write_file(&hosts, &format!("node-a {npub}\n"));
if !make_unreadable(&hosts) {
return;
}
assert!(!reloader.reload().await);
assert_eq!(
reloader.acl().check(&test_peer(&npub)),
PeerAclDecision::AllowList
);
assert_eq!(reloader.acl().default_decision(), "allow");
assert_eq!(
reloader.acl().allow_file_entries(),
vec!["node-a".to_string()]
);
make_readable(&hosts);
}
#[tokio::test]
async fn test_acl_reload_does_not_publish_an_empty_acl_over_an_enforcing_one() {
let dir = tempfile::tempdir().unwrap();
let allow = dir.path().join("peers.allow");
let deny = dir.path().join("peers.deny");
let allowed = test_npub();
write_file(&allow, &format!("{allowed}\n"));
let mut reloader = PeerAclReloader::with_paths(allow.clone(), deny);
assert_eq!(
reloader.acl().check(&test_peer(&allowed)),
PeerAclDecision::AllowList
);
std::thread::sleep(std::time::Duration::from_millis(5));
write_file(&allow, "");
assert!(!reloader.reload().await);
assert_eq!(
reloader.acl().check(&test_peer(&allowed)),
PeerAclDecision::AllowList
);
// The hold is bounded: a file the operator really did blank in place
// is published on the following tick.
assert!(reloader.reload().await);
assert!(reloader.acl().is_empty());
}
#[test]
fn test_acl_status_reports_effective_state_and_entries() {
let dir = tempfile::tempdir().unwrap();
@@ -977,7 +1273,7 @@ mod tests {
hosts.insert("node-a", &npub).unwrap();
std::fs::write(&allow, "node-a\n").unwrap();
let acl = PeerAcl::load_files_with_hosts(&allow, &deny, &hosts);
let acl = PeerAcl::try_load_files_with_hosts(&allow, &deny, &hosts).unwrap();
let peer = PeerIdentity::from_npub(&npub).unwrap();
assert_eq!(acl.allow_file_entries(), vec!["node-a".to_string()]);
+82 -6
View File
@@ -16,7 +16,7 @@ use crate::proto::fsp::wire::{
};
use crate::proto::fsp::{SessionAck, SessionSetup};
use crate::proto::link::{SessionDatagram, SessionDatagramRef};
use crate::proto::routing::{DropReason, NextHop, RouteAction, RouteOutcome};
use crate::proto::routing::{DropReason, LimitVerdict, NextHop, RouteAction, RouteOutcome};
use std::time::{Duration, Instant};
use tracing::{debug, warn};
@@ -125,7 +125,7 @@ impl Node {
bytes = payload.len(),
"Dropping transit SessionDatagram: no route to destination"
);
self.send_routing_error(&original).await;
self.send_routing_error(from, &original).await;
}
RouteOutcome::Forward {
next_hop,
@@ -157,7 +157,7 @@ impl Node {
self.metrics()
.forwarding
.record_reject_bytes(ForwardingReject::MtuExceeded, payload.len());
self.send_mtu_exceeded_error(dest, datagram_ref.src_addr, mtu)
self.send_mtu_exceeded_error(from, dest, datagram_ref.src_addr, mtu)
.await;
}
Err(e) => {
@@ -321,6 +321,39 @@ impl Node {
}
}
/// Spend one peer-budget token, after a gate further down has admitted.
///
/// The budget is keyed on the authenticated link peer the frame arrived
/// over, which is the one value at the emission point a sender cannot
/// mint: every field of the datagram itself is chosen by whoever sent it,
/// so a per-destination or per-source gate is escaped by varying the field
/// it keys on.
///
/// Peek and commit are separate because the per-destination interval gate
/// lives inside `routing::synth_routing_error` and runs after this.
/// Charging a suppressed signal would let a single unroutable destination
/// behind a high-fanout peer spend that peer's whole budget on emissions
/// nothing sends, silencing every other destination behind it.
fn commit_error_emission(&mut self, from: &NodeAddr) {
self.peer_error_budget.commit(from, Instant::now());
}
/// Count what the core's per-destination gate decided about one candidate
/// error signal.
///
/// The three verdicts are counted apart because they mean different
/// things to an operator: `Suppress` is the interval doing its job during
/// an outage, while `AdmitAtCapacity` says the destination map is full and
/// the interval is no longer suppressing anything for this destination, so
/// only the per-peer budget is still bounding emission.
fn record_error_verdict(&mut self, verdict: LimitVerdict) {
match verdict {
LimitVerdict::Suppress => self.metrics().errors.emit_over_dest_interval.inc(),
LimitVerdict::AdmitAtCapacity => self.metrics().errors.emit_limiter_at_capacity.inc(),
LimitVerdict::Admit => {}
}
}
/// Generate and send a routing error signal back to the datagram's source.
///
/// If we have cached coords for the destination, send PathBroken (we know
@@ -329,7 +362,20 @@ impl Node {
///
/// If we can't route the error back to the source either, drop silently.
/// No cascading errors.
async fn send_routing_error(&mut self, original: &SessionDatagram) {
/// `from` is the authenticated link peer the original datagram arrived
/// from, and is what the emission is charged against. It is the one value
/// at this point a sender cannot mint: every field of the datagram itself
/// is chosen by whoever sent it.
async fn send_routing_error(&mut self, from: &NodeAddr, original: &SessionDatagram) {
// Peeked, not spent. The destination gate inside the core may still
// suppress this signal, and charging a suppressed emission would let a
// single unroutable destination behind a high-fanout peer burn that
// peer's whole budget on signals nothing sends.
if !self.peer_error_budget.has_token(from, Instant::now()) {
self.metrics().errors.emit_over_peer_budget.inc();
return;
}
let my_addr = *self.node_addr();
let now_ms = std::time::SystemTime::now()
.duration_since(std::time::UNIX_EPOCH)
@@ -356,12 +402,22 @@ impl Node {
default_ttl,
)
};
let RouteAction::SendError { toward, bytes } = match action {
self.record_error_verdict(action.verdict);
let RouteAction::SendError { toward, bytes } = match action.action {
Some(action) => action,
// Rate limited: drop silently. No cascading errors.
None => return,
};
// Both gates have admitted, so the token peeked above is now spent.
// Charged here rather than at the peek so a destination the core
// suppressed costs the link peer nothing; see
// `commit_error_emission`. A later failure to resolve the reverse hop
// still leaves the token spent, which is deliberate: the work the
// budget bounds is the synthesis this node was induced to perform,
// not whether a hop happened to exist for it.
self.commit_error_emission(from);
// Resolve the reverse link hop only now, after the gate passed, so
// `find_next_hop`'s coord-cache touch keeps its pre-refactor scope.
let next_hop_addr = match self.find_next_hop(&toward) {
@@ -402,12 +458,28 @@ impl Node {
///
/// `dest` is the failed datagram's destination (rate-limit key); `toward`
/// is its source, where the signal is routed back.
///
/// `from` is the authenticated link peer the original datagram arrived
/// from, and is what the emission is charged against. MtuExceeded shares
/// the link peer's budget with the routing errors rather than holding its
/// own: a separate bucket would insulate path-MTU discovery from
/// routing-error pressure, at the cost of a second knob and of letting one
/// peer induce twice the total emission.
async fn send_mtu_exceeded_error(
&mut self,
from: &NodeAddr,
dest: NodeAddr,
toward: NodeAddr,
bottleneck_mtu: u16,
) {
// Peeked, not spent, for the same reason as in `send_routing_error`:
// the per-destination gate inside the core runs below and may still
// suppress this signal.
if !self.peer_error_budget.has_token(from, Instant::now()) {
self.metrics().errors.emit_over_peer_budget.inc();
return;
}
let my_addr = *self.node_addr();
let now_ms = Self::now_ms();
let default_ttl = self.config().node.session.default_ttl;
@@ -421,12 +493,16 @@ impl Node {
now_ms,
default_ttl,
);
let RouteAction::SendError { toward, bytes } = match action {
self.record_error_verdict(action.verdict);
let RouteAction::SendError { toward, bytes } = match action.action {
Some(action) => action,
// Rate limited: drop silently. No cascading errors.
None => return,
};
// Both gates have admitted; spend the token peeked above.
self.commit_error_emission(from);
// Resolve the reverse link hop only now, after the gate passed, so
// `find_next_hop`'s coord-cache touch keeps its pre-refactor scope.
let next_hop_addr = match self.find_next_hop(&toward) {
+77 -1
View File
@@ -19,9 +19,33 @@ use crate::proto::fmp::{
};
use crate::transport::{Link, LinkDirection, LinkId, ReceivedPacket};
use crate::utils::index::SessionIndex;
use std::time::Duration;
use std::time::{Duration, Instant};
use tracing::{debug, info, warn};
/// Minimum interval between accepted epoch changes for one peer identity,
/// and the recency threshold at which the peering an epoch change would
/// destroy still counts as live.
///
/// An epoch-mismatch msg1 is authentic but replayable: a captured one stays
/// valid indefinitely, and accepting it tears down a working peering. Both
/// conditions are receiver-local. The liveness half is the one that closes
/// the replay, since a peering under attack is by construction still
/// heartbeating; the interval half bounds the churn a peer can drive on its
/// own.
///
/// Sized against the peer's own recovery rather than against a round number:
/// a genuinely restarting peer's msg1 resends fire at roughly t+1, t+3, t+7
/// and t+15 seconds and its attempt is reaped at `handshake_timeout_secs`
/// (30), so 15 is the largest value at which a real restart still re-peers
/// inside its first handshake window with no reconnect backoff. It also sits
/// below `link_dead_timeout_secs` (30), so the liveness gate can never
/// outlive the reaper that would have removed the peering anyway.
///
/// Raising it lengthens the outage an attacker's accepted replay causes,
/// because the genuine peer's recovery msg1 hits the same arm. Lowering it
/// weakens both halves and, below the resend ladder, buys nothing.
const EPOCH_RESTART_MIN_INTERVAL_SECS: u64 = 15;
/// Why an inbound msg1 got past the `accept_connections` gate, and against
/// what identity the post-DH confirmation must check it.
///
@@ -714,6 +738,58 @@ impl Node {
// executor's `InvalidateSendState`
// (`ambient.verified_identity.node_addr()`) targets the same addr
// as the pre-refactor `remove_active_peer(&peer)`.
// The epoch travels inside the AEAD, so this msg1 is
// authentic — but it stays authentic after capture, and
// replaying one destroys a working peering and the FSP session
// state it carries, from off the path. Two receiver-local
// conditions gate the teardown. The peering's last
// authenticated inbound frame is the evidence it is still
// alive, and nothing an unauthenticated sender emits can
// refresh it, so a peer that genuinely restarted clears this by
// having stopped sending. The interval half bounds the churn
// one peer can drive on its own.
let now_ms = Self::now_ms();
let peering_idle_ms = self
.peers
.get(&peer)
.map(|p| p.idle_time(now_ms))
.unwrap_or(u64::MAX);
let dampened = self
.restart_dampener
.get(&peer)
.is_some_and(|t| t.elapsed().as_secs() < EPOCH_RESTART_MIN_INTERVAL_SECS);
if peering_idle_ms < EPOCH_RESTART_MIN_INTERVAL_SECS * 1000 || dampened {
debug!(
peer = %self.peer_display_name(&peer),
idle_ms = peering_idle_ms,
dampened,
"Epoch mismatch dampened, dropping msg1"
);
// Silent drop: the stored msg2 is bound to the original
// msg1's ephemeral, and answering an address the sender
// chose is free amplification.
//
// No registry cleanup is needed here. On the pre-refactor
// layout this arm removed the pending connection and its
// link, because both were inserted before msg1 was
// classified. The classification now runs against a local
// `machine` that enters `peer_machines` only at the promote
// tails below, and `link_id` is a bare allocation until
// then, so dropping out of the arm is the whole cleanup.
// The fresh leg holds no session index either (it is parked
// at `Handshaking{ReceivedMsg1}` with `our_index == None`),
// so nothing is leaked by returning.
self.stats_mut()
.record_reject(RejectReason::Handshake(HandshakeReject::BadState));
return;
}
// Stamped on acceptance only. A refusal that slid the window
// would let a sustained replay starve a genuinely restarting
// peer for as long as it kept sending.
let cutoff = Duration::from_secs(EPOCH_RESTART_MIN_INTERVAL_SECS);
self.restart_dampener.retain(|_, t| t.elapsed() < cutoff);
self.restart_dampener.insert(peer, Instant::now());
debug!(
peer = %self.peer_display_name(&peer),
"Peer restart detected (epoch mismatch), removing stale session"
+84 -15
View File
@@ -104,7 +104,8 @@ impl Node {
let recent_expiry_ms = self.config().node.lookup.recent_expiry_secs * 1000;
let my_addr = *self.node_addr();
use crate::proto::lookup::RequestOutcome;
match crate::proto::lookup::classify_request(
let peer_count = self.peers.len();
let classification = crate::proto::lookup::classify_request(
&mut self.lookup,
&request,
from,
@@ -112,7 +113,22 @@ impl Node {
now_ms,
recent_expiry_ms,
MAX_RECENT_LOOKUP_REQUESTS,
) {
peer_count,
);
// A full cache evicts rather than refuses, and the core charges the
// eviction to the peer that filled the cache. Count and log it here:
// the core does no metrics and no logging of its own.
if let Some(evicted) = classification.evicted {
self.metrics().lookup.req_dedup_evicted.inc();
debug!(
request_id = evicted.request_id,
evicted_from = %self.peer_display_name(&evicted.peer),
admitting = %self.peer_display_name(from),
share = evicted.share,
"Lookup dedup cache full, evicting the oldest entry to make room"
);
}
match classification.outcome {
RequestOutcome::Duplicate => {
self.metrics()
.lookup
@@ -123,19 +139,25 @@ impl Node {
"Duplicate LookupRequest, dropping"
);
}
RequestOutcome::DedupCacheFull { len } => {
self.metrics()
.lookup
.record_reject(DiscoveryReject::ReqDedupCacheFull);
debug!(
request_id = request.request_id,
from = %self.peer_display_name(from),
recent_requests = len,
max_recent_requests = MAX_RECENT_LOOKUP_REQUESTS,
"Discovery request dedup cache full, dropping LookupRequest"
);
}
RequestOutcome::RespondAsTarget => {
// Answering costs a fresh Schnorr signature every time: the
// proof is bound to the requester's request_id, so it cannot
// be cached or served twice. Meter that per link peer, or a
// neighbour generating request_ids sets this node's signing
// rate. The dedup entry the core recorded stays regardless,
// so a refused request still occupies its id and a retry,
// which carries a fresh id, is unaffected.
if !self.discovery_sign_limiter.should_sign(from) {
self.metrics()
.lookup
.record_reject(DiscoveryReject::ReqSignRateLimited);
debug!(
request_id = request.request_id,
from = %self.peer_display_name(from),
"Lookup signing budget spent for this peer, not answering"
);
return;
}
self.metrics().lookup.req_target_is_us.inc();
debug!(
request_id = request.request_id,
@@ -197,7 +219,11 @@ impl Node {
let now_ms = Self::now_ms();
// Check if we forwarded this request (transit node) or originated it
match crate::proto::lookup::classify_response(&mut self.lookup, response.request_id) {
match crate::proto::lookup::classify_response(
&mut self.lookup,
response.request_id,
&response.target,
) {
crate::proto::lookup::ResponseRoute::AlreadyForwarded => {
// Already forwarded a response for this request — drop to
// prevent response routing loops.
@@ -231,6 +257,25 @@ impl Node {
);
}
}
crate::proto::lookup::ResponseRoute::Unsolicited => {
// Nothing outstanding matches this, so acting on it would let
// one harvested signed response be replayed at will: each
// injection cleared the pending lookup, recorded a
// reachability success, refreshed the cached coordinates for a
// further full TTL, and flushed queued packets onto a route at
// a moment the sender chose. Dropped here, before the identity
// resolve and before the verify, so an unsolicited response
// costs nothing. This counter has a nonzero floor in healthy
// operation: a request is flooded to every qualifying tree
// peer, so duplicate replies land here once the first has been
// accepted.
self.metrics().lookup.resp_unsolicited.inc();
debug!(
request_id = response.request_id,
target = %self.peer_display_name(&response.target),
"LookupResponse does not match an outstanding request, dropping"
);
}
crate::proto::lookup::ResponseRoute::Originator => {
// We originated this request — verify proof before caching
let target = response.target;
@@ -587,6 +632,15 @@ impl Node {
};
let request = LookupRequest::new(request_id, *target, origin, origin_coords, ttl, 0);
// Recorded here rather than in the callers, so "if a request went out,
// its id is recorded" holds for every caller. The response path
// correlates against this set.
self.lookup
.pending_lookups
.entry(*target)
.or_insert_with(|| crate::proto::lookup::PendingLookup::new(Self::now_ms()))
.record(request_id);
// Tree-peer bloom-match selection + single encode live in the sans-IO
// core. The core keeps the tree-only (no non-tree fallback) behavior;
// the shell drives the sends and keeps all metrics/logging.
@@ -740,6 +794,21 @@ impl Node {
}
}
/// Remove expired entries from the recent-request dedup cache.
///
/// The ordinary request path purges lazily inside `classify_request`;
/// this is the explicit entry point for callers that need the purge
/// without an arriving request. Cache and per-peer index are purged
/// together, or the eviction policy reads a stale index.
///
/// Only the dedup regression tests call it: the production path's purge
/// happens inside `classify_request`.
#[cfg(test)]
pub(in crate::node) fn purge_expired_requests(&mut self, current_time_ms: u64) {
let expiry_ms = self.config().node.lookup.recent_expiry_secs * 1000;
self.lookup.purge_recent(current_time_ms, expiry_ms);
}
/// Min-fold our outgoing-link MTU into a LookupResponse's `path_mtu`.
///
/// Used at both transit-side reverse-path forward and at the target's
+4 -1
View File
@@ -6,6 +6,9 @@ mod mmp;
mod native;
pub(in crate::node) use native::PendingNative;
pub(crate) mod probe;
mod rekey;
// Widened from private by the rekey drain cap: `node::session` calls
// `rekey::drain_max_retention_ms` to bound how long a superseded epoch is
// retained. `rx_loop` is not declared here; master moved it out of `handlers`.
pub(in crate::node) mod rekey;
pub(in crate::node) mod session;
mod timeout;
+50 -1
View File
@@ -29,6 +29,48 @@ const DRAIN_WINDOW_SECS: u64 = 10;
/// a peer's rekey msg1. FMP-scoped copy for `check_rekey`.
const REKEY_DAMPENING_SECS: u64 = 30;
/// Floor on the absolute ceiling for `previous`-slot retention after a
/// cutover, in seconds.
///
/// The drain deadline is peer-progress-aware: it slides forward on every
/// inbound frame that authenticates against the old epoch, so a peer that
/// keeps sealing in that epoch holds the retired key for as long as it
/// likes. This bounds that. It has to stay longer than the worst-case
/// recovery of a legitimate peer that lost msg3, which at stock defaults
/// is the msg3 resend ladder (about 31 s) plus `handshake_timeout_secs`
/// (30 s) before the responder abandons plus `REKEY_DAMPENING_SECS`
/// (30 s) before it may re-initiate, so about 90 s. 120 s clears that
/// with margin and still bounds retention to roughly one
/// `node.rekey.after_secs` period. `drain_max_retention_ms` takes the
/// larger of this floor and the budget the running configuration
/// actually implies, so a shortened handshake timer cannot push the
/// ceiling under the recovery it has to clear.
///
/// Lowering it below that budget cuts off legitimate slow peers: their
/// frames go silently undecryptable until their own rekey retry
/// re-converges the epochs, because nothing tears an established session
/// down on repeated decrypt failure. Raising it lengthens the window in
/// which a retired key stays resident.
const DRAIN_MAX_RETENTION_SECS: u64 = crate::proto::fsp::limits::DRAIN_WINDOW_SECS * 12;
/// Effective ceiling on total `previous`-slot retention, in milliseconds.
///
/// The larger of `DRAIN_MAX_RETENTION_SECS` and the msg3 recovery budget
/// the configured handshake timers imply, so the ceiling always clears
/// the recovery it is supposed to leave room for.
pub(in crate::node) fn drain_max_retention_ms(rate_limit: &crate::config::RateLimitConfig) -> u64 {
let mut ladder_ms: u64 = 0;
let mut interval = rate_limit.handshake_resend_interval_ms as f64;
for _ in 0..rate_limit.handshake_max_resends {
ladder_ms = ladder_ms.saturating_add(interval as u64);
interval *= rate_limit.handshake_resend_backoff;
}
let recovery_budget_ms = ladder_ms
.saturating_add(rate_limit.handshake_timeout_secs.saturating_mul(1000))
.saturating_add(crate::proto::fsp::limits::REKEY_DAMPENING_SECS * 1000);
(DRAIN_MAX_RETENTION_SECS * 1000).max(recovery_budget_ms)
}
impl Node {
/// Periodic rekey check. Called from the tick loop.
///
@@ -640,6 +682,13 @@ impl Node {
// is the only stamp that path writes. A *completed* rekey has no such
// bound and must not acquire one: see `FspAction::AbandonHandshake`.
let stale_handshake_ms = self.config().node.rate_limit.handshake_timeout_secs * 1000;
// Absolute ceiling on `previous`-slot retention, measured from the
// cutover. The sliding drain deadline is peer-progress-aware, so an
// authenticated peer that keeps sealing in the old epoch can hold the
// retired key indefinitely; this bounds that without shortening the
// grace a peer that lost msg3 legitimately needs. Resolved here rather
// than in the core, which reads no clock and no configuration.
let drain_max_ms = drain_max_retention_ms(&self.config().node.rate_limit);
self.sessions
.iter()
.filter(|(_, entry)| entry.is_established())
@@ -650,7 +699,7 @@ impl Node {
is_rekey_initiator: entry.is_rekey_initiator(),
cutover_timer_elapsed: cutover_timer_elapsed(now_ms, entry.rekey_completed_ms()),
is_draining: entry.is_draining(),
drain_expired: entry.drain_expired(now_ms, drain_ms),
drain_expired: entry.drain_expired(now_ms, drain_ms, drain_max_ms),
has_rekey_msg3_payload: entry.rekey_msg3_payload().is_some(),
is_dampened: entry.is_rekey_dampened(now_ms, dampening_ms),
armed_handshake_expired: entry.last_peer_rekey_ms() != 0
+201 -17
View File
@@ -46,6 +46,49 @@ use crate::upper::icmp::FIPS_OVERHEAD;
use secp256k1::PublicKey;
use tracing::{debug, info, trace, warn};
/// Minimum interval between path-MTU releases driven by `PathBroken` for one
/// destination.
///
/// `PathBroken` is unauthenticated, so a release is a remote party's claim
/// that the path a tightened MTU described is gone. Without an interval the
/// claim can be repeated at line rate, discarding a genuinely learned
/// bottleneck as fast as it is relearned. Raising it defers a legitimate
/// release after a second real break, which costs throughput on the new path
/// but never a blackhole, since the deferred value is the tighter one.
pub(in crate::node) const PATH_MTU_RELEASE_MIN_INTERVAL: std::time::Duration =
std::time::Duration::from_millis(1000);
/// Bytes the link layer adds to an encoded `SessionDatagram` on its way to the
/// wire: the established FMP header, the 4-byte session-relative timestamp and
/// the AEAD tag. Mirrors the buffer `send_encrypted_link_message_with_ce`
/// builds.
///
/// Spelled out in full rather than through the `crate::proto::fmp::wire`
/// import above, which is `#[cfg(unix)]`. This constant feeds `link_wire_len`,
/// whose caller `send_session_datagram` is compiled on every platform, so
/// taking the name from that import fails to build on Windows.
const LINK_FRAME_OVERHEAD: usize =
crate::proto::fmp::wire::ESTABLISHED_HEADER_SIZE + 4 + crate::noise::TAG_SIZE;
/// Wire size of an encoded `SessionDatagram` of `encoded_len` bytes.
fn link_wire_len(encoded_len: usize) -> usize {
encoded_len + LINK_FRAME_OVERHEAD
}
/// Divisor giving the share of the session table that unauthenticated
/// half-open entries may hold, as `max_sessions / DIVISOR`.
///
/// Two means a reconnect storm, where every peer that had a session
/// initiates at once after a restart or a healed partition, still fits in
/// half the table; a tighter share bites four times sooner and is felt by
/// a hub before it is felt by an attacker. Half-open entries are reaped
/// after `handshake_timeout_secs` while established ones survive
/// `idle_timeout_secs`, so they turn over faster than the share suggests.
/// Lowering the divisor raises the share, which lets a handshake flood
/// crowd out peers that complete; raising it refuses legitimate initiators
/// sooner in a storm.
const HALF_OPEN_SHARE_DIVISOR: usize = 2;
/// Inputs to `try_send_session_data_pipelined` — the FSP+FMP pipelined
/// fast path that hands both AEAD operations to the encrypt worker
/// in a single dispatch.
@@ -547,6 +590,19 @@ impl Node {
// no limit at all. A setup naming an established peer cannot grow the
// table and is metered separately, so that a stranger flood over a
// shared link cannot stop that peer's rekey from arming.
// Population cap, ahead of the limiter so a full table costs no
// token, no responder handshake and no ack. The predicate is "would
// admitting this grow the table", not "is this a stranger": `class`
// is Stranger for an existing Initiating or AwaitingMsg3 entry too,
// and refusing those would break in-flight legitimate handshakes and
// the duplicate-ack resend. Same shape as the pending-destination cap
// in `queue_pending_packet`. Refuse rather than evict: msg1 is
// unauthenticated here, so evicting would hand a stranger a teardown
// primitive it does not have.
if !self.admit_new_session(src_addr) {
return;
}
let class = if self
.sessions
.get(src_addr)
@@ -1876,8 +1932,20 @@ impl Node {
}
}
// The path this destination's stored MTU described is gone, so release
// it rather than carrying it onto whatever path replaces it.
self.path_mtu_lookup_release(&msg.dest_addr);
// it rather than carrying it onto whatever path replaces it. Rate
// limited per destination on its own budget: PathBroken is
// unauthenticated, and an unlimited release discards a genuinely
// learned bottleneck as fast as it is relearned. The budget is not
// shared with any other signal, so nothing else can spend it.
if self
.path_mtu_release_limiter
.should_send(&msg.dest_addr, Self::now_ms())
{
self.path_mtu_lookup_release(&msg.dest_addr);
} else {
trace!(dest = %msg.dest_addr,
"PathBroken path MTU release rate-limited, keeping the stored value");
}
if !has_cached_identity {
debug!(dest = %msg.dest_addr,
@@ -1946,6 +2014,55 @@ impl Node {
"MtuExceeded: transit router reports oversized packet"
);
// Both effects below — the session's own path MTU and the
// FipsAddress-keyed lookup the TUN MSS clamp reads — are refused from
// here, so one return covers both. The guards sit ahead of the apply
// rather than between the two effects, which is what makes the floor
// govern `current_mtu` and not only the lookup table.
// Refuse a bottleneck too small to describe a usable path; a stored
// value that low drives the SYN-time MSS clamp into single digits or
// zero. The reactive carrier is unauthenticated, so it has its own
// floor constant, currently equal to the actionable one.
if msg.mtu < crate::upper::icmp::MIN_REACTIVE_PATH_MTU {
warn!(
dest = %peer_name,
reporter = %msg.reporter,
bottleneck_mtu = msg.mtu,
floor = crate::upper::icmp::MIN_REACTIVE_PATH_MTU,
"MtuExceeded reports a path MTU below the actionable floor; ignoring"
);
self.metrics().errors.mtu_exceeded_below_floor.inc();
return;
}
// Corroboration. The admission gate narrows which destination may be
// named; it cannot authenticate the reporter, so a legal value is a
// legal value from anyone and the floor alone only sets the outcome of
// a forgery rather than preventing it. An honest report exists only
// because a frame this node emitted did not fit some hop, so require
// that this node has actually sent something larger than the value
// being claimed since the last accepted decrease. Honest path-MTU
// discovery satisfies this by construction; a forgery has to wait for
// us to emit a frame bigger than the value it wants to claim, which
// bounds every accepted claim from below by our own traffic.
let sent_wire_len = self
.sessions
.get(&msg.dest_addr)
.map(|e| e.max_sent_wire_len())
.unwrap_or(0);
if msg.mtu >= sent_wire_len {
debug!(
dest = %peer_name,
reporter = %msg.reporter,
bottleneck_mtu = msg.mtu,
max_sent_wire_len = sent_wire_len,
"MtuExceeded reports a bottleneck no smaller than anything this node has sent; ignoring"
);
self.metrics().errors.mtu_exceeded_uncorroborated.inc();
return;
}
// Apply to PathMtuState: immediate decrease via apply_notification()
if let Some(entry) = self.sessions.get_mut(&msg.dest_addr)
&& let Some(mmp) = entry.mmp_mut()
@@ -1966,21 +2083,12 @@ impl Node {
}
}
// The admission gate above restricts which addresses may be written,
// not which values. Any node at any distance may legitimately report a
// bottleneck for a destination this node has bound, so refuse to store
// one too small to describe a usable path; a stored value that low
// drives the SYN-time MSS clamp into single digits or zero.
if msg.mtu < crate::upper::icmp::MIN_ACTIONABLE_PATH_MTU {
warn!(
dest = %peer_name,
reporter = %msg.reporter,
bottleneck_mtu = msg.mtu,
floor = crate::upper::icmp::MIN_ACTIONABLE_PATH_MTU,
"MtuExceeded reports a path MTU below the actionable floor; ignoring"
);
self.metrics().errors.mtu_exceeded_below_floor.inc();
return;
// Spent: the evidence vouched for this decrease and does not vouch for
// the next one. An initiating session has no `mmp` and so reaches this
// with the apply above skipped; the reset belongs to the acceptance,
// not to the apply.
if let Some(entry) = self.sessions.get_mut(&msg.dest_addr) {
entry.clear_sent_wire_len();
}
// Mirror the bottleneck into the FipsAddress-keyed lookup used by
@@ -2049,6 +2157,59 @@ impl Node {
/// Creates a Noise XK handshake as initiator, wraps msg1 in a
/// SessionSetup, encapsulates in a SessionDatagram, and routes
/// toward the destination.
/// Whether a session for `addr` may be created, given the table cap.
///
/// Returns true when an entry already exists, since admitting it cannot
/// grow the table. Counts its own refusals, so the two reasons are
/// distinguishable without turning on debug logging.
pub(in crate::node) fn admit_new_session(&mut self, addr: &NodeAddr) -> bool {
let max_sessions = self.config().node.limits.max_sessions;
if max_sessions == 0 || self.sessions.contains_key(addr) {
return true;
}
if self.sessions.len() >= max_sessions {
debug!(
src = %self.peer_display_name(addr),
sessions = self.sessions.len(),
max_sessions = max_sessions,
"Session table full, refusing to create a session"
);
self.stats_mut()
.record_reject(RejectReason::Session(SessionReject::TableFull));
return false;
}
// Half-open entries are unauthenticated and are reaped after
// `handshake_timeout_secs`, so they are the cheap half of the table
// to fill. Holding them to a share keeps room for peers that
// complete. The outer length test makes the scan unreachable below
// the share, and the table is itself bounded by the cap above.
// At least one, or a table capped at one would admit no inbound
// session at all rather than one.
let half_open_share = (max_sessions / HALF_OPEN_SHARE_DIVISOR).max(1);
if self.sessions.len() >= half_open_share {
let half_open = self
.sessions
.values()
.filter(|e| e.is_awaiting_msg3())
.count();
if half_open >= half_open_share {
debug!(
src = %self.peer_display_name(addr),
half_open = half_open,
half_open_share = half_open_share,
"Half-open session share exhausted, refusing to create a session"
);
self.stats_mut()
.record_reject(RejectReason::Session(SessionReject::HalfOpenFull));
return false;
}
}
true
}
pub(in crate::node) async fn initiate_session(
&mut self,
dest_addr: NodeAddr,
@@ -2516,6 +2677,7 @@ impl Node {
if let Some(entry) = self.sessions.get_mut(dest_addr) {
entry.record_sent(send.payload.len());
entry.record_sent_wire_len(wire_capacity);
if let Some(mmp) = entry.mmp_mut() {
mmp.sender.record_sent(
fsp_counter,
@@ -2794,6 +2956,13 @@ impl Node {
self.send_encrypted_link_message(&next_hop_addr, &encoded)
.await?;
self.metrics().forwarding.record_originated(encoded.len());
// Evidence for the reactive path-MTU carrier. A transit hop
// re-encapsulates what it forwards, so the frame that overflows a
// downstream link is the size this frame is here.
if let Some(entry) = self.sessions.get_mut(&datagram.dest_addr) {
entry.record_sent_wire_len(link_wire_len(encoded.len()));
}
Ok(())
}
@@ -2888,6 +3057,17 @@ impl Node {
return;
}
// No session, so this one would grow the table. Answer the local
// application the way an unroutable destination is answered rather
// than returning an error from `initiate_session`: the caller reads
// an error as "no route" and responds with a discovery lookup and a
// queued packet, which is outbound traffic on a node already at its
// limit.
if !self.admit_new_session(&dest_addr) {
self.send_icmpv6_dest_unreachable(&ipv6_packet);
return;
}
// No session: initiate one and queue the packet.
// If session initiation fails (no route), trigger discovery and
// queue the packet for retry when discovery completes.
@@ -3015,6 +3195,10 @@ impl Node {
return;
}
if !self.admit_new_session(&dest_addr) {
return;
}
match self.initiate_session(dest_addr, dest_pubkey).await {
Ok(()) => {
debug!(dest = %self.peer_display_name(&dest_addr), "Session initiated after discovery");
+29 -12
View File
@@ -1797,13 +1797,21 @@ impl Node {
);
// Resolve the TUN ifindex so the responder can
// drop queries arriving on the mesh interface
// (fips0). Without this, the `::` bind exposes
// /etc/fips/hosts alias probing to any mesh peer.
// When TUN isn't enabled or the name can't be
// resolved, `None` disables the filter (there
// is no mesh surface to defend anyway).
let mesh_ifindex =
Self::lookup_mesh_ifindex(self.config().tun.name());
// Without this, the `::` bind exposes the
// hosts file's alias space to any mesh peer.
// The name comes from the device the TUN
// path actually created, not the configured
// one: macOS and FreeBSD assign utunN/tunN
// of their own choosing and the configured
// name resolves to nothing there, which left
// the filter permanently off.
let mesh_ifindex = self.mesh_ifindex();
if self.tun_name.is_some() && mesh_ifindex.is_none() {
warn!(
device = ?self.tun_name,
"Mesh interface index unresolved; DNS mesh filter disabled"
);
}
info!(
bind = %local_addr,
hosts = reloader.hosts().len(),
@@ -1992,12 +2000,21 @@ impl Node {
Ok(())
}
/// Resolve the mesh TUN interface index by name.
/// Resolve the index of the mesh TUN device this node actually created.
///
/// Returns `None` if the interface does not exist (e.g. TUN disabled
/// or not yet created). A `None` result disables the DNS responder's
/// mesh-interface filter — safe, because if there is no fips0 there
/// is no mesh exposure to defend against.
/// Reads the device name recorded when the TUN was brought up, which is
/// the kernel's name rather than the configured one. Returns `None` when
/// no TUN is up, which disables the DNS responder's mesh-interface
/// filter: with no mesh interface there is no mesh exposure to defend.
/// An app-owned TUN also leaves the name unset, so the filter stays off
/// there even though a mesh interface exists.
pub(crate) fn mesh_ifindex(&self) -> Option<u32> {
self.tun_name.as_deref().and_then(Self::lookup_mesh_ifindex)
}
/// Resolve an interface index by name.
///
/// Returns `None` if the interface does not exist.
fn lookup_mesh_ifindex(name: &str) -> Option<u32> {
#[cfg(unix)]
{
+30
View File
@@ -233,6 +233,8 @@ pub struct LookupMetrics {
pub req_decode_error: Counter,
pub req_duplicate: Counter,
pub req_dedup_cache_full: Counter,
pub req_dedup_evicted: Counter,
pub req_sign_rate_limited: Counter,
pub req_target_is_us: Counter,
pub req_forwarded: Counter,
pub req_ttl_exhausted: Counter,
@@ -248,6 +250,7 @@ pub struct LookupMetrics {
pub resp_forwarded: Counter,
pub resp_identity_miss: Counter,
pub resp_proof_failed: Counter,
pub resp_unsolicited: Counter,
pub resp_no_route: Counter,
pub resp_accepted: Counter,
pub resp_timed_out: Counter,
@@ -262,10 +265,12 @@ impl LookupMetrics {
DiscoveryReject::ReqDecodeError => self.req_decode_error.inc(),
DiscoveryReject::ReqDuplicate => self.req_duplicate.inc(),
DiscoveryReject::ReqDedupCacheFull => self.req_dedup_cache_full.inc(),
DiscoveryReject::ReqSignRateLimited => self.req_sign_rate_limited.inc(),
DiscoveryReject::ReqTtlExhausted => self.req_ttl_exhausted.inc(),
DiscoveryReject::RespDecodeError => self.resp_decode_error.inc(),
DiscoveryReject::RespIdentityMiss => self.resp_identity_miss.inc(),
DiscoveryReject::RespProofFailed => self.resp_proof_failed.inc(),
DiscoveryReject::RespUnsolicited => self.resp_unsolicited.inc(),
DiscoveryReject::RespNoRoute => self.resp_no_route.inc(),
}
}
@@ -277,6 +282,8 @@ impl LookupMetrics {
req_decode_error: self.req_decode_error.get(),
req_duplicate: self.req_duplicate.get(),
req_dedup_cache_full: self.req_dedup_cache_full.get(),
req_dedup_evicted: self.req_dedup_evicted.get(),
req_sign_rate_limited: self.req_sign_rate_limited.get(),
req_target_is_us: self.req_target_is_us.get(),
req_forwarded: self.req_forwarded.get(),
req_ttl_exhausted: self.req_ttl_exhausted.get(),
@@ -292,6 +299,7 @@ impl LookupMetrics {
resp_forwarded: self.resp_forwarded.get(),
resp_identity_miss: self.resp_identity_miss.get(),
resp_proof_failed: self.resp_proof_failed.get(),
resp_unsolicited: self.resp_unsolicited.get(),
resp_no_route: self.resp_no_route.get(),
resp_accepted: self.resp_accepted.get(),
resp_timed_out: self.resp_timed_out.get(),
@@ -474,7 +482,25 @@ pub struct ErrorMetrics {
/// count means a forwarder on the reverse path is mangling the unsigned
/// annotation.
pub lookup_resp_mtu_below_floor: Counter,
/// `MtuExceeded` signals ignored because this node has not sent a frame
/// larger than the bottleneck they report since the last accepted
/// decrease. An honest report cannot arise without such a frame, so a
/// rising count is a forged or stale reactive signal.
pub mtu_exceeded_uncorroborated: Counter,
pub unbound: UnboundSignals,
/// Routing errors this node declined to emit because the authenticated
/// link peer that induced them had spent its budget. A rising count is
/// either a peer flooding unroutable traffic or a hub relaying more
/// simultaneously-broken destinations than the budget allows.
pub emit_over_peer_budget: Counter,
/// Routing errors this node declined to emit because one for the same
/// destination went out within the per-destination interval. This is the
/// aggregate suppression a real outage produces.
pub emit_over_dest_interval: Counter,
/// Routing errors emitted without recording their destination, because
/// the per-destination limiter's map was full. The signal was still sent;
/// what was lost is interval suppression for that destination.
pub emit_limiter_at_capacity: Counter,
}
impl ErrorMetrics {
@@ -491,6 +517,10 @@ impl ErrorMetrics {
unbound_broken: self.unbound.broken.get(),
unbound_mtu: self.unbound.mtu.get(),
unbound_forged: self.unbound.forged.get(),
emit_over_peer_budget: self.emit_over_peer_budget.get(),
emit_over_dest_interval: self.emit_over_dest_interval.get(),
emit_limiter_at_capacity: self.emit_limiter_at_capacity.get(),
mtu_exceeded_uncorroborated: self.mtu_exceeded_uncorroborated.get(),
}
}
}
+61 -1
View File
@@ -17,6 +17,7 @@ pub(crate) mod encrypt_worker;
mod handlers;
mod lifecycle;
pub(crate) mod metrics;
mod peer_error_budget;
mod peering;
mod rate_limit;
pub(crate) mod reject;
@@ -30,7 +31,8 @@ pub(crate) mod stats_history;
mod tests;
mod tree;
use self::rate_limit::{HandshakeRateLimiter, SessionSetupRateLimiter};
use self::peer_error_budget::PeerErrorBudget;
use self::rate_limit::{HandshakeRateLimiter, LookupSignRateLimiter, SessionSetupRateLimiter};
use self::reloadable::Reloadable;
/// Half-range of the symmetric jitter applied to the per-session rekey timer.
@@ -473,6 +475,11 @@ pub struct Node {
/// Discovery-subsystem state: recent-request dedup cache, in-flight
/// lookups, originator-side backoff, and transit-side forward limiter.
lookup: Lookup,
/// Signing budget for lookups we answer about ourselves (target-side),
/// keyed on the link peer the request arrived over. Held here rather than
/// inside `lookup` because it is an `Instant`-based limiter and the
/// `proto` tree is clockless.
discovery_sign_limiter: LookupSignRateLimiter,
// === Diagnostics ===
/// In-flight `probe` jobs plus their per-target ownership claims. Driven
@@ -538,6 +545,12 @@ pub struct Node {
/// Pending outbound handshakes by our sender_idx.
/// Tracks which LinkId corresponds to which session index.
pending_outbound: HashMap<(TransportId, u32), LinkId>,
/// When each peer identity's last ACCEPTED epoch change tore down its
/// peering. Keyed on identity rather than address, and held here rather
/// than on `ActivePeer`, because the teardown being dampened destroys
/// the peer entry itself. Pruned on insert; see
/// `EPOCH_RESTART_MIN_INTERVAL_SECS`.
restart_dampener: HashMap<NodeAddr, std::time::Instant>,
// === Rate Limiting ===
/// Rate limiter for msg1 processing (DoS protection).
@@ -546,6 +559,12 @@ pub struct Node {
setup_rate_limiter: SessionSetupRateLimiter,
/// Rate limiter for ICMP Packet Too Big messages.
icmp_rate_limiter: IcmpRateLimiter,
/// Budget bounding the routing errors one authenticated link peer can
/// induce this node to emit. Keyed on the link peer because that is the
/// only value at the emission point a sender cannot mint; the
/// per-destination interval inside `routing` is keyed on a field the
/// sender chooses and is an aggregate suppressor, not a bound.
peer_error_budget: PeerErrorBudget,
/// Routing-subsystem state (routing error-signal rate limiter).
routing: Router,
/// FMP connection-lifecycle decision anchor (stateless; drives the
@@ -559,6 +578,12 @@ pub struct Node {
mmp: Mmp,
/// Rate limiter for source-side CoordsRequired/PathBroken responses.
coords_response_rate_limiter: RoutingErrorRateLimiter,
/// Rate limiter for PathBroken-driven path-MTU releases, per destination.
/// Deliberately its own instance rather than a share of
/// `coords_response_rate_limiter`: a budget another signal can spend is
/// not a bound on this one, and one PathBroken drives both responses, so
/// a shared limiter would let the coord-warmup arm pay for the release.
path_mtu_release_limiter: RoutingErrorRateLimiter,
// === Peering Homeostasis ===
/// Owner of the peering-reconciler state relocated off `Node`: the sans-IO
@@ -806,9 +831,11 @@ impl Node {
index_allocator: IndexAllocator::new(),
peers_by_index: HashMap::new(),
pending_outbound: HashMap::new(),
restart_dampener: HashMap::new(),
msg1_rate_limiter,
setup_rate_limiter,
icmp_rate_limiter: IcmpRateLimiter::new(),
peer_error_budget: PeerErrorBudget::new(),
routing: Router::new(),
fmp: Fmp::new(),
fsp: Fsp::new(),
@@ -816,11 +843,15 @@ impl Node {
coords_response_rate_limiter: RoutingErrorRateLimiter::with_interval_ms(
coords_response_interval_ms,
),
path_mtu_release_limiter: RoutingErrorRateLimiter::with_interval_ms(
handlers::session::PATH_MTU_RELEASE_MIN_INTERVAL.as_millis() as u64,
),
probes: handlers::probe::ProbeRegistry::new(),
lookup: Lookup::new(
LookupBackoff::with_params(backoff_base_secs, backoff_max_secs),
LookupForwardRateLimiter::with_interval_ms(forward_min_interval_secs * 1000),
),
discovery_sign_limiter: LookupSignRateLimiter::new(),
peering: peering::reconcile::Peering::new(),
last_parent_reeval: None,
last_congestion_log: None,
@@ -963,9 +994,11 @@ impl Node {
index_allocator: IndexAllocator::new(),
peers_by_index: HashMap::new(),
pending_outbound: HashMap::new(),
restart_dampener: HashMap::new(),
msg1_rate_limiter,
setup_rate_limiter,
icmp_rate_limiter: IcmpRateLimiter::new(),
peer_error_budget: PeerErrorBudget::new(),
routing: Router::new(),
fmp: Fmp::new(),
fsp: Fsp::new(),
@@ -973,8 +1006,12 @@ impl Node {
coords_response_rate_limiter: RoutingErrorRateLimiter::with_interval_ms(
coords_response_interval_ms,
),
path_mtu_release_limiter: RoutingErrorRateLimiter::with_interval_ms(
handlers::session::PATH_MTU_RELEASE_MIN_INTERVAL.as_millis() as u64,
),
probes: handlers::probe::ProbeRegistry::new(),
lookup: Lookup::new(LookupBackoff::new(), LookupForwardRateLimiter::new()),
discovery_sign_limiter: LookupSignRateLimiter::new(),
peering: peering::reconcile::Peering::new(),
last_parent_reeval: None,
last_congestion_log: None,
@@ -1761,6 +1798,17 @@ impl Node {
/// the way the rx_loop's handler does without standing up a client task and
/// a socket pair. A parity test over an empty registry proves nothing about
/// a publisher that drops fields, which is why this exists.
/// The instant at which `peer`'s last ACCEPTED epoch change was stamped,
/// or `None` if it has none. Test-only, and it exists for one assertion:
/// that a REFUSED epoch-mismatch msg1 leaves this untouched. The refusal
/// must not slide the window, or a sustained replay starves a genuinely
/// restarting peer for as long as it keeps sending — which is the whole
/// point of stamping on acceptance rather than on every sighting.
#[cfg(test)]
pub(crate) fn restart_dampener_stamp(&self, peer: &NodeAddr) -> Option<std::time::Instant> {
self.restart_dampener.get(peer).copied()
}
#[cfg(test)]
pub(crate) fn native_registry_for_test(&mut self) -> &mut crate::native::registry::Registry {
&mut self.native
@@ -2754,6 +2802,12 @@ impl Node {
// === End-to-End Sessions ===
/// Get a session by remote NodeAddr.
/// Set the per-link-peer lookup signing budget (for tests).
#[cfg(test)]
pub(crate) fn set_discovery_sign_budget(&mut self, burst: f64, rate: f64) {
self.discovery_sign_limiter = LookupSignRateLimiter::with_params(burst, rate);
}
/// Disable the discovery forward rate limiter (for tests).
#[cfg(test)]
pub(crate) fn disable_discovery_forward_rate_limit(&mut self) {
@@ -2846,6 +2900,12 @@ impl Node {
/// `FipsAddress`-keyed map the TCP MSS clamp reads, and the session's own
/// source-side path MTU estimate.
fn path_mtu_lookup_release(&mut self, addr: &NodeAddr) {
// The evidence that corroborates a reactive MtuExceeded described the
// path being released, so it does not vouch for whatever replaces it.
if let Some(entry) = self.sessions.get_mut(addr) {
entry.clear_sent_wire_len();
}
// The session's own source-side estimate described the same dead path,
// and the increase ladder is the only thing that would ever raise it
// again. Reset it here so the two halves of "this path is gone" stay
+239
View File
@@ -0,0 +1,239 @@
//! Per-link-peer budget for induced routing-error emissions.
//!
//! A transit node synthesizes a routing error (CoordsRequired, PathBroken or
//! MtuExceeded) in response to a datagram it could not forward. Every field of
//! that datagram is chosen by whoever sent it, so a per-destination or
//! per-source gate can be escaped by varying the field it is keyed on. The one
//! value at the emission point an attacker cannot mint is the authenticated
//! link peer the frame arrived from, whose cardinality is bounded by the peer
//! table and by admission. This budget is keyed on it, and is consulted ahead
//! of any address-keyed structure so that those structures only grow at the
//! budget rate.
use crate::NodeAddr;
use std::collections::HashMap;
use std::time::{Duration, Instant};
/// Sustained rate, in signals per second, at which one authenticated link peer
/// may induce this node to emit routing errors.
///
/// Bounds the reflection an admitted peer can aim at a victim it names, and
/// bounds how fast that peer can grow the per-destination limiter's map.
/// Raising it costs proportionally more reflected traffic per peer; lowering it
/// silences a hub peer that relays many sources through a genuine outage
/// sooner, which costs those sources their CoordsRequired and their
/// path-MTU feedback.
pub const PEER_ERROR_RATE_PER_SEC: u32 = 20;
/// Number of routing errors one link peer may induce back to back before the
/// sustained rate applies.
///
/// Sized so that an ordinary burst of unroutable traffic behind one peer still
/// signals promptly. Raising it lets a peer front-load a larger reflection;
/// lowering it makes a legitimate convergence burst arrive as a trickle.
pub const PEER_ERROR_BURST: u32 = 50;
/// Tokens are carried in thousandths so the refill of a sub-millisecond
/// interval is not rounded away.
const MILLI: u64 = 1000;
/// A peer whose bucket has refilled to full carries no state worth keeping, so
/// entries are dropped once per this interval to bound the map across peer
/// churn.
const SWEEP_INTERVAL: Duration = Duration::from_secs(30);
/// One peer's token bucket.
struct Bucket {
/// Remaining tokens, in thousandths of a signal.
milli_tokens: u64,
/// When `milli_tokens` was last brought up to date.
last_refill: Instant,
}
/// Token-bucket budget for routing errors, keyed on the authenticated link
/// peer that induced them.
pub struct PeerErrorBudget {
buckets: HashMap<NodeAddr, Bucket>,
/// Refill rate in thousandths of a token per millisecond.
milli_per_ms: u64,
/// Bucket ceiling, in thousandths of a token.
capacity: u64,
last_sweep: Instant,
}
impl PeerErrorBudget {
/// Create a budget at the shipped rate and burst.
pub fn new() -> Self {
Self::with_rate(PEER_ERROR_RATE_PER_SEC, PEER_ERROR_BURST)
}
/// Create a budget with an explicit sustained rate and burst.
pub fn with_rate(per_sec: u32, burst: u32) -> Self {
Self {
buckets: HashMap::new(),
milli_per_ms: u64::from(per_sec),
capacity: u64::from(burst) * MILLI,
last_sweep: Instant::now(),
}
}
/// Whether `peer` has a token to spend, without spending it.
///
/// Separate from [`Self::commit`] so a signal that a later gate suppresses
/// does not consume budget: an outage behind a high-fanout peer would
/// otherwise spend that peer's whole budget on emissions the
/// per-destination interval discards, silencing every other destination
/// behind it.
pub fn has_token(&mut self, peer: &NodeAddr, now: Instant) -> bool {
self.refill(peer, now);
self.buckets
.get(peer)
.is_some_and(|b| b.milli_tokens >= MILLI)
}
/// Spend one token for `peer`. Call only on the path that actually emits.
pub fn commit(&mut self, peer: &NodeAddr, now: Instant) {
self.refill(peer, now);
if let Some(bucket) = self.buckets.get_mut(peer) {
bucket.milli_tokens = bucket.milli_tokens.saturating_sub(MILLI);
}
self.sweep(now);
}
/// Bring `peer`'s bucket up to date, creating a full one on first sighting.
fn refill(&mut self, peer: &NodeAddr, now: Instant) {
let capacity = self.capacity;
let milli_per_ms = self.milli_per_ms;
let bucket = self.buckets.entry(*peer).or_insert(Bucket {
milli_tokens: capacity,
last_refill: now,
});
let elapsed_ms = now
.saturating_duration_since(bucket.last_refill)
.as_millis() as u64;
if elapsed_ms > 0 {
bucket.milli_tokens = (bucket.milli_tokens + elapsed_ms * milli_per_ms).min(capacity);
bucket.last_refill = now;
}
}
/// Drop full buckets, at most once per [`SWEEP_INTERVAL`].
fn sweep(&mut self, now: Instant) {
if now.saturating_duration_since(self.last_sweep) < SWEEP_INTERVAL {
return;
}
self.last_sweep = now;
let capacity = self.capacity;
let milli_per_ms = self.milli_per_ms;
self.buckets.retain(|_, b| {
let elapsed_ms = now.saturating_duration_since(b.last_refill).as_millis() as u64;
b.milli_tokens + elapsed_ms * milli_per_ms < capacity
});
}
#[cfg(test)]
pub fn len(&self) -> usize {
self.buckets.len()
}
}
impl Default for PeerErrorBudget {
fn default() -> Self {
Self::new()
}
}
#[cfg(test)]
mod tests {
use super::*;
fn addr(val: u8) -> NodeAddr {
let mut bytes = [0u8; 16];
bytes[0] = val;
NodeAddr::from_bytes(bytes)
}
/// Spend one token per admitted emission.
fn spend(budget: &mut PeerErrorBudget, peer: &NodeAddr, now: Instant) -> bool {
if !budget.has_token(peer, now) {
return false;
}
budget.commit(peer, now);
true
}
#[test]
fn a_peer_may_emit_its_full_burst_then_is_suppressed() {
let mut budget = PeerErrorBudget::new();
let now = Instant::now();
let peer = addr(1);
for i in 0..PEER_ERROR_BURST {
assert!(spend(&mut budget, &peer, now), "burst signal {i} refused");
}
assert!(!spend(&mut budget, &peer, now));
}
#[test]
fn an_exhausted_budget_refills_at_the_sustained_rate() {
let mut budget = PeerErrorBudget::new();
let start = Instant::now();
let peer = addr(1);
for _ in 0..PEER_ERROR_BURST {
assert!(spend(&mut budget, &peer, start));
}
assert!(!spend(&mut budget, &peer, start));
// One second of refill buys exactly the sustained rate back.
let later = start + Duration::from_secs(1);
for i in 0..PEER_ERROR_RATE_PER_SEC {
assert!(
spend(&mut budget, &peer, later),
"refilled signal {i} refused"
);
}
assert!(!spend(&mut budget, &peer, later));
}
#[test]
fn one_peer_exhausting_its_budget_does_not_silence_another() {
let mut budget = PeerErrorBudget::new();
let now = Instant::now();
let noisy = addr(1);
let quiet = addr(2);
for _ in 0..PEER_ERROR_BURST {
assert!(spend(&mut budget, &noisy, now));
}
assert!(!spend(&mut budget, &noisy, now));
assert!(spend(&mut budget, &quiet, now));
}
#[test]
fn peeking_does_not_spend_a_token() {
let mut budget = PeerErrorBudget::with_rate(1, 1);
let now = Instant::now();
let peer = addr(1);
assert!(budget.has_token(&peer, now));
assert!(budget.has_token(&peer, now));
budget.commit(&peer, now);
assert!(!budget.has_token(&peer, now));
}
#[test]
fn full_buckets_are_dropped_by_the_sweep() {
let mut budget = PeerErrorBudget::new();
let start = Instant::now();
for i in 0..50u8 {
assert!(spend(&mut budget, &addr(i), start));
}
assert_eq!(budget.len(), 50);
// Long enough for every bucket to have refilled to full.
let later = start + SWEEP_INTERVAL + Duration::from_secs(1);
assert!(spend(&mut budget, &addr(200), later));
assert_eq!(budget.len(), 1);
}
}
+161
View File
@@ -874,3 +874,164 @@ mod tests {
);
}
}
// ============================================================================
// Target-side: Lookup Signing Budget
//
// Homed here rather than in `proto::lookup::limits` because it is an
// `Instant`-based shell limiter, and the `proto` tree is `alloc`-based and
// clockless. On `maint` it lived in `src/node/discovery_rate_limit.rs`, which
// `master` dissolved into `proto/lookup/limits.rs`.
// ============================================================================
/// Signatures one link peer may buy in a burst before the refill paces it.
///
/// Sized for the case that actually produces a burst: a topology change
/// flushes correspondents' coordinate caches and they all look this node up
/// at once, through whichever few link peers lead here, each retrying on the
/// `node.discovery.attempt_timeouts_secs` ladder. Lowering this makes a
/// genuinely popular node intermittently unresolvable, which is the same
/// symptom as the flood it defends against; raising it raises the worst-case
/// signing burst one neighbour can force.
const DEFAULT_SIGN_BURST: f64 = 256.0;
/// Sustained signatures per second per link peer.
///
/// At the default eight or so link peers this caps the node near 256
/// signatures per second in the sustained case. The real cost of one
/// `Identity::sign` on this codebase has not been measured, so this number
/// is a bound rather than a tuned value; it is the one line to change if a
/// measurement says otherwise.
const DEFAULT_SIGN_RATE: f64 = 32.0;
/// Maximum age of an idle bucket before cleanup.
const SIGN_MAX_AGE: Duration = Duration::from_secs(300);
/// Token bucket per link peer for lookups this node answers about itself.
///
/// A min-interval limiter is the wrong shape here: a popular node receives
/// legitimate bursts of lookups for itself through the few link peers that
/// lead to it, and a min interval refuses all but the first of each burst.
/// A bucket absorbs the burst and paces the sustained rate.
pub struct LookupSignRateLimiter {
buckets: HashMap<NodeAddr, SignBucket>,
burst: f64,
rate: f64,
}
struct SignBucket {
/// Tokens remaining, at most `burst`.
tokens: f64,
/// When `tokens` was last refilled.
updated: Instant,
}
impl LookupSignRateLimiter {
/// Create with default burst and refill rate.
pub fn new() -> Self {
Self::with_params(DEFAULT_SIGN_BURST, DEFAULT_SIGN_RATE)
}
/// Create with a custom burst and refill rate.
pub fn with_params(burst: f64, rate: f64) -> Self {
Self {
buckets: HashMap::new(),
burst,
rate,
}
}
/// Spend one token for `from`, or report that its budget is exhausted.
///
/// Returns true when the signature may be produced. A zero burst is
/// read as "unlimited" rather than "refuse everything", so a
/// misconfiguration cannot make this node unresolvable.
pub fn should_sign(&mut self, from: &NodeAddr) -> bool {
if self.burst <= 0.0 {
return true;
}
let now = Instant::now();
let burst = self.burst;
let rate = self.rate;
let bucket = self.buckets.entry(*from).or_insert(SignBucket {
tokens: burst,
updated: now,
});
let elapsed = now.duration_since(bucket.updated).as_secs_f64();
bucket.tokens = (bucket.tokens + elapsed * rate).min(burst);
bucket.updated = now;
if bucket.tokens < 1.0 {
return false;
}
bucket.tokens -= 1.0;
self.cleanup(now);
true
}
/// Drop buckets untouched for longer than [`SIGN_MAX_AGE`]; a full
/// bucket carries no state worth keeping.
fn cleanup(&mut self, now: Instant) {
self.buckets
.retain(|_, b| now.duration_since(b.updated) < SIGN_MAX_AGE);
}
#[cfg(test)]
pub fn len(&self) -> usize {
self.buckets.len()
}
}
impl Default for LookupSignRateLimiter {
fn default() -> Self {
Self::new()
}
}
// ============================================================================
// Tests
// ============================================================================
#[cfg(test)]
mod sign_limiter_tests {
use super::*;
fn addr(val: u8) -> NodeAddr {
let mut bytes = [0u8; 16];
bytes[0] = val;
NodeAddr::from_bytes(bytes)
}
// The `DiscoveryBackoff` and `DiscoveryForwardRateLimiter` tests that
// accompanied this limiter in `src/node/discovery_rate_limit.rs` are not
// repeated here: both types moved into `crate::proto::lookup::limits` as
// `LookupBackoff` and `LookupForwardRateLimiter`, and all thirteen of those
// cases live there, name for name, in `src/proto/lookup/tests/limits.rs`,
// ported to the clockless millisecond API. Only the signing budget, which
// is `Instant`-based and so stays in the shell, is exercised below.
#[test]
fn test_sign_budget_is_spent_per_peer_and_does_not_touch_another_peer() {
let mut limiter = LookupSignRateLimiter::with_params(4.0, 0.0);
for _ in 0..4 {
assert!(limiter.should_sign(&addr(1)));
}
assert!(
!limiter.should_sign(&addr(1)),
"the burst is the whole budget when nothing refills it"
);
assert!(
limiter.should_sign(&addr(2)),
"one peer spending its budget must not spend another's"
);
assert_eq!(limiter.len(), 2);
}
#[test]
fn test_sign_budget_of_zero_burst_is_read_as_unlimited() {
let mut limiter = LookupSignRateLimiter::with_params(0.0, 0.0);
for _ in 0..1000 {
assert!(limiter.should_sign(&addr(1)));
}
assert_eq!(limiter.len(), 0, "unlimited keeps no per-peer state");
}
}
+34
View File
@@ -120,7 +120,19 @@ pub enum DiscoveryReject {
/// Request dedup cache (`recent_requests`) is at capacity, so the
/// `LookupRequest` is dropped without being forwarded. Tracked via
/// [`DiscoveryStats::req_dedup_cache_full`](crate::node::stats::DiscoveryStats).
///
/// Frozen at zero: a full cache now evicts its oldest entry and admits
/// the request, counted as
/// [`DiscoveryStats::req_dedup_evicted`](crate::node::stats::DiscoveryStats).
/// The variant and its counter stay so an operator reading a dashboard
/// across versions does not find the series missing.
ReqDedupCacheFull,
/// This node is the lookup target, but the link peer the request
/// arrived from has spent its signing budget. Answering costs a fresh
/// Schnorr signature per request, so the budget bounds what one
/// neighbour can make this node sign. Tracked via
/// [`DiscoveryStats::req_sign_rate_limited`](crate::node::stats::DiscoveryStats).
ReqSignRateLimited,
/// Request arrived with TTL=0 — no more forwarding hops allowed.
/// Tracked via
/// [`DiscoveryStats::req_ttl_exhausted`](crate::node::stats::DiscoveryStats).
@@ -136,6 +148,14 @@ pub enum DiscoveryReject {
/// Response proof signature failed verification. Tracked via
/// [`DiscoveryStats::resp_proof_failed`](crate::node::stats::DiscoveryStats).
RespProofFailed,
/// Response arrived on the originator path but carries no
/// `request_id` this node has outstanding for the named target, so
/// it answers no lookup of ours. Expected to be nonzero in healthy
/// operation: the request is flooded to every qualifying tree peer,
/// so duplicate replies land here after the first is accepted.
/// Tracked via
/// [`DiscoveryStats::resp_unsolicited`](crate::node::stats::DiscoveryStats).
RespUnsolicited,
/// Response could not be routed toward the origin: no reverse-path
/// entry for the `request_id` and no greedy tree route to the
/// origin. Tracked via
@@ -238,6 +258,18 @@ pub enum SessionReject {
/// before any handshake state was created or any ack sent. Tracked via
/// [`SessionStats::setup_rate_limited`](crate::node::stats::SessionStats).
SetupRateLimited,
/// A session would have been created but the table is at
/// `node.limits.max_sessions`. Refused rather than evicted: the
/// deciding message is unauthenticated at this point, so evicting
/// would hand a stranger a way to tear down established sessions.
/// Tracked via
/// [`SessionStats::table_full`](crate::node::stats::SessionStats).
TableFull,
/// A session would have been created but unauthenticated half-open
/// entries already hold their share of the table. Bounds what a
/// handshake flood can deny an established peer. Tracked via
/// [`SessionStats::half_open_full`](crate::node::stats::SessionStats).
HalfOpenFull,
}
/// MMP rejection reasons.
@@ -368,9 +400,11 @@ mod tests {
DiscoveryReject::ReqDuplicate,
DiscoveryReject::ReqDedupCacheFull,
DiscoveryReject::ReqTtlExhausted,
DiscoveryReject::ReqSignRateLimited,
DiscoveryReject::RespDecodeError,
DiscoveryReject::RespIdentityMiss,
DiscoveryReject::RespProofFailed,
DiscoveryReject::RespUnsolicited,
];
for v in variants {
let r = RejectReason::Discovery(v);
+129 -115
View File
@@ -100,6 +100,15 @@ pub(crate) struct SessionEntry {
/// Whether this node initiated the Noise handshake.
/// Used for spin bit role assignment in session-layer MMP.
is_initiator: bool,
/// Largest on-the-wire frame this node has sent toward the remote since
/// the last accepted path-MTU decrease or release, in bytes.
///
/// Corroborates a reactive `MtuExceeded`, which is unauthenticated: an
/// honest report exists only because a frame this node emitted did not
/// fit some hop, so an honest report always names a value below this.
/// Reset on each accepted decrease and on release so one historical
/// large send cannot vouch for a session's whole lifetime.
max_sent_wire_len: u16,
/// Session-layer MMP state. Initialized on Established transition.
mmp: Option<MmpSessionState>,
@@ -203,6 +212,7 @@ impl SessionEntry {
session_start_ms: 0,
coords_warmup_remaining: 0,
is_initiator,
max_sent_wire_len: 0,
mmp: None,
packets_sent: 0,
packets_recv: 0,
@@ -306,6 +316,27 @@ impl SessionEntry {
self.coords_warmup_remaining = value;
}
/// Largest wire frame sent toward the remote since the last accepted
/// path-MTU decrease or release.
pub(crate) fn max_sent_wire_len(&self) -> u16 {
self.max_sent_wire_len
}
/// Note a frame of `wire_len` bytes sent toward the remote, keeping the
/// largest. Frames beyond `u16::MAX` saturate, which only ever makes the
/// corroboration more permissive and cannot exceed what a path MTU can
/// name.
pub(crate) fn record_sent_wire_len(&mut self, wire_len: usize) {
let wire_len = u16::try_from(wire_len).unwrap_or(u16::MAX);
self.max_sent_wire_len = self.max_sent_wire_len.max(wire_len);
}
/// Forget what has been sent, so the next reactive report needs fresh
/// evidence of its own.
pub(crate) fn clear_sent_wire_len(&mut self) {
self.max_sent_wire_len = 0;
}
/// Mark the session as started (transition to Established).
///
/// Records the current time as the session start for computing
@@ -729,10 +760,24 @@ impl SessionEntry {
/// permanent silent decrypt failure. A peer that never catches up
/// is instead handled by the FSP session liveness path (fresh
/// handshake / teardown of a genuinely dead link).
pub(crate) fn drain_expired(&self, now_ms: u64, drain_ms: u64) -> bool {
///
/// `max_drain_ms` is an absolute ceiling measured from the cutover
/// alone, so the sliding deadline delays erasure by a bounded amount
/// rather than preventing it: the only party that can push the
/// deadline out is the authenticated peer holding the old key, and
/// without a ceiling it holds that key for as long as it keeps using
/// it. What the ceiling costs is that a peer which has still not
/// recovered by then is cut off deliberately, and its frames are
/// undecryptable until its own rekey retry re-converges the epochs.
/// It must therefore stay above the worst-case legitimate recovery;
/// see `DRAIN_MAX_RETENTION_SECS`.
pub(crate) fn drain_expired(&self, now_ms: u64, drain_ms: u64, max_drain_ms: u64) -> bool {
if self.drain_started_ms == 0 {
return false;
}
if now_ms.saturating_sub(self.drain_started_ms) >= max_drain_ms {
return true;
}
let deadline_anchor = self.drain_started_ms.max(self.previous_last_used_ms);
now_ms.saturating_sub(deadline_anchor) >= drain_ms
}
@@ -1229,6 +1274,9 @@ mod overlapping_epoch_tests {
#[test]
fn drain_expiry_is_peer_progress_aware() {
const DRAIN_MS: u64 = 10_000;
// Well clear of the shipped ceiling, so this case still exercises
// the sliding deadline and nothing else.
const MAX_MS: u64 = 120_000;
let cutover_ms = 1_000;
// Build the post-cutover state via the production cutover path:
@@ -1257,7 +1305,7 @@ mod overlapping_epoch_tests {
// Even though `now - drain_started_ms` exceeds DRAIN_MS, the
// window is NOT expired: the peer just used `previous`.
assert!(
!entry.drain_expired(t, DRAIN_MS),
!entry.drain_expired(t, DRAIN_MS, MAX_MS),
"previous slot must not be retired while peer keeps using it (t={t})"
);
assert!(
@@ -1270,11 +1318,11 @@ mod overlapping_epoch_tests {
// last `previous`-slot use was at t=25_000; the window now
// elapses DRAIN_MS after that, NOT DRAIN_MS after the cutover.
assert!(
!entry.drain_expired(34_999, DRAIN_MS),
!entry.drain_expired(34_999, DRAIN_MS, MAX_MS),
"window must not expire before DRAIN_MS past the last previous use"
);
assert!(
entry.drain_expired(35_000, DRAIN_MS),
entry.drain_expired(35_000, DRAIN_MS, MAX_MS),
"window must expire DRAIN_MS after the last previous-slot decrypt"
);
@@ -1292,6 +1340,7 @@ mod overlapping_epoch_tests {
#[test]
fn drain_expiry_unaffected_when_peer_off_old_epoch() {
const DRAIN_MS: u64 = 10_000;
const MAX_MS: u64 = 120_000;
let cutover_ms = 1_000;
let (_old_send, old_recv) = xk_pair(1, 2);
@@ -1303,141 +1352,106 @@ mod overlapping_epoch_tests {
// No old-epoch frames ever arrive: `previous_last_used_ms` stays
// 0, the deadline anchor is the cutover time.
assert!(
!entry.drain_expired(cutover_ms + DRAIN_MS - 1, DRAIN_MS),
!entry.drain_expired(cutover_ms + DRAIN_MS - 1, DRAIN_MS, MAX_MS),
"window must not expire early"
);
assert!(
entry.drain_expired(cutover_ms + DRAIN_MS, DRAIN_MS),
entry.drain_expired(cutover_ms + DRAIN_MS, DRAIN_MS, MAX_MS),
"window must expire on the plain wall-clock timer when peer is off the old epoch"
);
}
// ========================================================================
// Rekey-policy characterization (pins `check_session_rekey`'s decision
// boundaries before the `Fsp::poll_rekey` hoist — these thresholds have no
// other test module).
// ========================================================================
// The retention ceiling the drain cap added. The five rekey-policy
// characterization tests that sat here on `maint` are not carried: they
// pinned `check_session_rekey`'s boundaries before the `Fsp::poll_rekey`
// hoist, and the hoisted predicate has its own tests in
// `src/proto/fsp/tests/core.rs`.
/// The initiator liveness-cutover delay used by `check_session_rekey`
/// (`FSP_CUTOVER_DELAY_MS`). Mirrored here as the characterization anchor.
const CUTOVER_DELAY_MS: u64 = 2000;
// 12. A peer that keeps exercising the old epoch delays erasure by a
// bounded amount rather than preventing it. The refreshes must
// continue past the ceiling: a case that stops refreshing at the
// boundary passes without the ceiling and proves nothing.
#[test]
fn drain_retention_is_capped_against_a_peer_pinning_the_old_epoch() {
const DRAIN_MS: u64 = 10_000;
const MAX_MS: u64 = 120_000;
/// Build an established entry that has completed a rekey as initiator and
/// holds a pending session awaiting the K-bit cutover.
fn entry_pending_cutover(rekey_completed_ms: u64) -> SessionEntry {
let (_cur_send, cur_recv) = xk_pair(1, 2);
let (_old_send, old_recv) = xk_pair(1, 2);
let (_new_send, new_recv) = xk_pair(3, 4);
let mut entry = entry_with_current(cur_recv);
// Mark ourselves the rekey initiator, then land the completed session
// as pending (clears rekey_state, so has_rekey_in_progress() == false).
entry.set_rekey_state(HandshakeState::new_xk_responder(keypair(7)), true);
let mut entry = entry_with_current(old_recv);
entry.set_pending_session(new_recv);
entry.set_rekey_completed_ms(rekey_completed_ms);
entry
}
assert!(entry.cutover_to_new_session(1));
// The initiator-side cutover predicate: pending session present, no rekey
// in progress, we are the initiator, and the liveness timer has elapsed.
#[test]
fn rekey_cutover_predicate_boundary() {
let completed = 1_000u64;
let entry = entry_pending_cutover(completed);
// One old-epoch frame every half window, which is what a peer
// pinning the drain deadline actually does.
let mut t = 1u64;
while t < MAX_MS {
entry.refresh_previous_use(t);
assert!(
!entry.drain_expired(t, DRAIN_MS, MAX_MS),
"ceiling fired before the peer's grace ran out (t={t})"
);
t += DRAIN_MS / 2;
}
assert!(entry.pending_new_session().is_some());
assert!(!entry.has_rekey_in_progress());
assert!(entry.is_rekey_initiator());
// Not yet eligible one ms before the delay elapses.
let just_before = completed + CUTOVER_DELAY_MS - 1;
// Still refreshing, so the sliding deadline is nowhere near due.
entry.refresh_previous_use(MAX_MS + 1);
assert!(
just_before.saturating_sub(entry.rekey_completed_ms()) < CUTOVER_DELAY_MS,
"cutover must not fire before the liveness delay"
);
// Eligible exactly at the delay.
let at = completed + CUTOVER_DELAY_MS;
assert!(
at.saturating_sub(entry.rekey_completed_ms()) >= CUTOVER_DELAY_MS,
"cutover fires once the liveness delay has elapsed"
entry.drain_expired(MAX_MS + 1, DRAIN_MS, MAX_MS),
"a peer pinning the old epoch retained the retired key past the ceiling"
);
}
// The rekey trigger's own threshold arithmetic is tested against the real
// predicate in src/proto/fmp/tests/core.rs, which drives poll_rekey. A test
// here previously reproduced that OR predicate as a local closure and
// asserted against its own copy, so it could not fail for the reason it
// existed: deleting the counter arm from the trigger left it green. Its one
// assertion over real code, that a fresh entry's jitter lies within the
// symmetric bound, is covered over 100 samples by
// test_session_entry_rekey_jitter_in_range in src/node/tests/session.rs.
// Dampening boundary: within `dampening_ms` of the peer's rekey msg1, local
// initiation is suppressed; at/after the window it is not.
// 13. The ceiling must not shorten the grace the sliding deadline
// exists to give a peer that lost msg3 and is still catching up.
#[test]
fn rekey_dampening_boundary() {
let (_s, recv) = xk_pair(1, 2);
let mut entry = entry_with_current(recv);
const DAMP_MS: u64 = 30_000;
fn drain_retention_cap_does_not_shorten_the_ordinary_grace() {
const DRAIN_MS: u64 = 10_000;
const MAX_MS: u64 = 120_000;
const CUTOVER_MS: u64 = 1_000;
// No peer rekey recorded → never dampened.
assert!(!entry.is_rekey_dampened(50_000, DAMP_MS));
let (_old_send, old_recv) = xk_pair(1, 2);
let (_new_send, new_recv) = xk_pair(3, 4);
let mut entry = entry_with_current(old_recv);
entry.set_pending_session(new_recv);
assert!(entry.cutover_to_new_session(CUTOVER_MS));
assert!(entry.is_draining());
entry.record_peer_rekey(10_000);
assert!(
entry.is_rekey_dampened(10_000 + DAMP_MS - 1, DAMP_MS),
"dampened within the window"
);
assert!(
!entry.is_rekey_dampened(10_000 + DAMP_MS, DAMP_MS),
"not dampened once the window has elapsed"
);
// A peer that lost msg3 and is still catching up sends one
// old-epoch frame part-way through the window.
let last_use = CUTOVER_MS + DRAIN_MS / 2;
entry.refresh_previous_use(last_use);
assert!(!entry.drain_expired(CUTOVER_MS + DRAIN_MS, DRAIN_MS, MAX_MS));
assert!(!entry.drain_expired(last_use + DRAIN_MS - 1, DRAIN_MS, MAX_MS));
assert!(entry.drain_expired(last_use + DRAIN_MS, DRAIN_MS, MAX_MS));
}
// Epoch-reaction: a frame authenticating against `pending` while a msg3
// retransmission is retained confirms the peer on the new epoch (clears the
// msg3 payload) and then promotes.
// 14. The shipped ceiling has to clear the worst-case legitimate
// recovery of a peer that lost msg3: the msg3 resend ladder, the
// responder's handshake timeout, and the rekey dampening window
// before it may re-initiate. A later tightening of the ceiling
// reds this rather than silently amputating that recovery.
#[test]
fn epoch_reaction_pending_confirms_then_promotes() {
let (mut p_send, p_recv) = xk_pair(3, 4);
let (_cur_send, cur_recv) = xk_pair(1, 2);
let mut entry = entry_with_current(cur_recv);
let k_before = entry.current_k_bit();
entry.set_pending_session(p_recv);
entry.set_rekey_msg3_payload(vec![0xAB; 8], 5_000);
assert!(entry.rekey_msg3_payload().is_some());
fn drain_retention_cap_clears_the_msg3_recovery_budget() {
const DRAIN_MS: u64 = 10_000;
// 31 s resend ladder + 30 s handshake timeout + 30 s dampening.
const RECOVERY_MS: u64 = 91_000;
const CUTOVER_MS: u64 = 1_000;
let max_ms = crate::node::handlers::rekey::drain_max_retention_ms(
&crate::config::RateLimitConfig::default(),
);
let (ct, counter, hdr) = seal(&mut p_send, b"new-epoch", !k_before);
let (_pt, slot) = entry
.fsp_trial_decrypt(&ct, counter, &hdr, !k_before, 2_000)
.expect("pending frame decrypts");
assert_eq!(slot, EpochSlot::Pending);
let (_old_send, old_recv) = xk_pair(1, 2);
let (_new_send, new_recv) = xk_pair(3, 4);
let mut entry = entry_with_current(old_recv);
entry.set_pending_session(new_recv);
assert!(entry.cutover_to_new_session(CUTOVER_MS));
assert!(entry.is_draining());
// Reaction order: confirm (while pending still held) then promote.
entry.confirm_peer_new_epoch();
assert!(entry.rekey_msg3_payload().is_none());
entry.handle_peer_kbit_flip(2_000);
assert!(entry.pending_new_session().is_none());
assert_ne!(entry.current_k_bit(), k_before);
}
// Epoch-reaction: as the initiator that already cut over on its own timer
// (msg3 retained, no pending), a frame authenticating against `current`
// confirms the responder reached the new epoch.
#[test]
fn epoch_reaction_current_confirms_responder() {
let (mut cur_send, cur_recv) = xk_pair(1, 2);
let mut entry = entry_with_current(cur_recv);
entry.set_rekey_msg3_payload(vec![0xCD; 8], 5_000);
assert!(entry.pending_new_session().is_none());
assert!(entry.rekey_msg3_payload().is_some());
let (ct, counter, hdr) = seal(&mut cur_send, b"steady", false);
let (_pt, slot) = entry
.fsp_trial_decrypt(&ct, counter, &hdr, false, 2_000)
.expect("current frame decrypts");
assert_eq!(slot, EpochSlot::Current);
// The Current-with-retained-msg3-and-no-pending arm confirms.
entry.confirm_peer_new_epoch();
assert!(entry.rekey_msg3_payload().is_none());
entry.refresh_previous_use(CUTOVER_MS + RECOVERY_MS);
assert!(
!entry.drain_expired(CUTOVER_MS + RECOVERY_MS, DRAIN_MS, max_ms),
"ceiling fires inside the msg3 recovery budget"
);
}
}
+21
View File
@@ -79,6 +79,14 @@ pub struct SessionStats {
/// A setup message was refused by the per-link-peer setup limiter,
/// before any handshake state was created or any ack sent.
pub setup_rate_limited: u64,
/// A session would have been created but the table is at
/// `node.limits.max_sessions`. A sustained rate means either the cap
/// is sized below what this node legitimately carries, or something is
/// holding the table full.
pub table_full: u64,
/// A session would have been created but unauthenticated half-open
/// entries already hold their share of the table.
pub half_open_full: u64,
}
impl SessionStats {
@@ -96,6 +104,8 @@ impl SessionStats {
pending_replaced: self.pending_replaced,
ack_handshake_failed: self.ack_handshake_failed,
setup_rate_limited: self.setup_rate_limited,
table_full: self.table_full,
half_open_full: self.half_open_full,
}
}
@@ -110,6 +120,8 @@ impl SessionStats {
SessionReject::RekeyPending => self.rekey_pending += 1,
SessionReject::AckHandshakeFailed => self.ack_handshake_failed += 1,
SessionReject::SetupRateLimited => self.setup_rate_limited += 1,
SessionReject::TableFull => self.table_full += 1,
SessionReject::HalfOpenFull => self.half_open_full += 1,
}
}
}
@@ -319,6 +331,8 @@ pub struct LookupStatsSnapshot {
pub req_decode_error: u64,
pub req_duplicate: u64,
pub req_dedup_cache_full: u64,
pub req_dedup_evicted: u64,
pub req_sign_rate_limited: u64,
pub req_target_is_us: u64,
pub req_forwarded: u64,
pub req_ttl_exhausted: u64,
@@ -334,6 +348,7 @@ pub struct LookupStatsSnapshot {
pub resp_forwarded: u64,
pub resp_identity_miss: u64,
pub resp_proof_failed: u64,
pub resp_unsolicited: u64,
pub resp_no_route: u64,
pub resp_accepted: u64,
pub resp_timed_out: u64,
@@ -389,6 +404,8 @@ pub struct SessionStatsSnapshot {
pub pending_replaced: u64,
pub ack_handshake_failed: u64,
pub setup_rate_limited: u64,
pub table_full: u64,
pub half_open_full: u64,
}
#[derive(Clone, Debug, Default, Serialize)]
@@ -421,6 +438,10 @@ pub struct ErrorSignalStatsSnapshot {
pub unbound_broken: u64,
pub unbound_mtu: u64,
pub unbound_forged: u64,
pub emit_over_peer_budget: u64,
pub emit_over_dest_interval: u64,
pub emit_limiter_at_capacity: u64,
pub mtu_exceeded_uncorroborated: u64,
}
#[derive(Clone, Debug, Default, Serialize)]
+501 -2
View File
@@ -83,6 +83,19 @@ async fn test_request_ttl_zero_not_forwarded() {
// Unit Tests — LookupResponse Handler
// ============================================================================
/// Record `request_id` as outstanding for `target`, exactly as
/// `initiate_lookup` does when it puts a request on the wire. The response
/// handler correlates against this, so a unit test that hands the handler a
/// response without it is testing the correlation gate rather than whatever
/// it names.
fn seed_pending_lookup(node: &mut Node, target: crate::NodeAddr, request_id: u64) {
node.lookup
.pending_lookups
.entry(target)
.or_insert_with(|| crate::proto::lookup::PendingLookup::new(Node::now_ms()))
.record(request_id);
}
#[tokio::test]
async fn test_response_decode_error() {
let mut node = make_node();
@@ -106,6 +119,8 @@ async fn test_response_originator_caches_route() {
// Register target identity in cache so verification can find it
node.register_identity(target, target_identity.pubkey_full());
seed_pending_lookup(&mut node, target, 555);
// Create a valid response with a real proof signature (includes coords)
let proof_data = LookupResponse::proof_bytes(555, &target, &coords);
let proof = target_identity.sign(&proof_data);
@@ -183,6 +198,8 @@ async fn test_response_proof_verification_success() {
// Register target in identity_cache
node.register_identity(target, target_identity.pubkey_full());
seed_pending_lookup(&mut node, target, 700);
// Sign with correct proof_bytes (including coords)
let proof_data = LookupResponse::proof_bytes(700, &target, &coords);
let proof = target_identity.sign(&proof_data);
@@ -218,6 +235,8 @@ async fn test_response_proof_verification_failure() {
node.register_identity(target, target_identity.pubkey_full());
// Sign with a DIFFERENT identity (wrong key)
seed_pending_lookup(&mut node, target, 701);
let wrong_identity = Identity::generate();
let proof_data = LookupResponse::proof_bytes(701, &target, &coords);
let proof = wrong_identity.sign(&proof_data);
@@ -251,6 +270,8 @@ async fn test_response_identity_cache_miss() {
// Do NOT register target in identity_cache
seed_pending_lookup(&mut node, target, 702);
let proof_data = LookupResponse::proof_bytes(702, &target, &coords);
let proof = target_identity.sign(&proof_data);
@@ -285,6 +306,8 @@ async fn test_response_coord_substitution_detected() {
// Register target in identity_cache
node.register_identity(target, target_identity.pubkey_full());
seed_pending_lookup(&mut node, target, 703);
// Sign proof with real coords
let proof_data = LookupResponse::proof_bytes(703, &target, &real_coords);
let proof = target_identity.sign(&proof_data);
@@ -305,6 +328,267 @@ async fn test_response_coord_substitution_detected() {
);
}
/// Build a signed LookupResponse body for `target_identity` over
/// `request_id`, ready to hand to `handle_lookup_response`.
fn signed_response_body(
target_identity: &Identity,
request_id: u64,
coords: &TreeCoordinate,
) -> Vec<u8> {
let target = *target_identity.node_addr();
let proof_data = LookupResponse::proof_bytes(request_id, &target, coords);
let proof = target_identity.sign(&proof_data);
LookupResponse::new(request_id, target, coords.clone(), proof).encode()[1..].to_vec()
}
/// Register `target_identity` and return its address and a plausible
/// coordinate for it, the shared preamble of the correlation tests.
fn register_lookup_target(node: &mut Node, target_identity: &Identity) -> TreeCoordinate {
let target = *target_identity.node_addr();
node.register_identity(target, target_identity.pubkey_full());
TreeCoordinate::from_addrs(vec![target, make_node_addr(0xF0)]).unwrap()
}
fn wall_clock_ms() -> u64 {
std::time::SystemTime::now()
.duration_since(std::time::UNIX_EPOCH)
.map(|d| d.as_millis() as u64)
.unwrap_or(0)
}
#[tokio::test]
async fn test_unsolicited_lookup_response_is_dropped_before_proof_verification() {
// Any admitted peer can hand us a correctly signed response for a target
// we never asked about. Accepting it lets that peer clear our pending
// state, refresh a cache entry's TTL and flush our queued packets at a
// moment it picks, so the response must not be acted on at all.
let mut node = make_node();
let from = make_node_addr(0xAA);
let target_identity = Identity::generate();
let target = *target_identity.node_addr();
let coords = register_lookup_target(&mut node, &target_identity);
let body = signed_response_body(&target_identity, 900, &coords);
assert!(
node.lookup.pending_lookups.is_empty(),
"precondition: this node has no lookup outstanding for anything"
);
node.handle_lookup_response(&from, &body).await;
assert!(
!node.coord_cache().contains(&target, wall_clock_ms()),
"a response answering no request of ours must not reach the coordinate cache"
);
assert_eq!(
node.metrics().lookup.resp_accepted.get(),
0,
"an unsolicited response must not count as accepted"
);
assert_eq!(
node.metrics().lookup.resp_unsolicited.get(),
1,
"the drop must be visible on a counter, not only in a log"
);
assert_eq!(
node.metrics().lookup.resp_proof_failed.get(),
0,
"the drop must happen before the signature verify, so the verify is not a cost gate"
);
}
#[tokio::test]
async fn test_lookup_response_with_a_request_id_we_never_issued_is_dropped() {
// Correlating on the target alone would leave the attack open: there is
// no inbound limiter on responses, so a peer can spray a harvested one
// and land inside any window in which we happen to be looking that
// target up. The id must match too.
let mut node = make_node();
let from = make_node_addr(0xAA);
let target_identity = Identity::generate();
let target = *target_identity.node_addr();
let coords = register_lookup_target(&mut node, &target_identity);
node.initiate_lookup(&target, 5).await;
let issued = node
.lookup
.pending_lookups
.get(&target)
.unwrap()
.ids
.clone();
assert_eq!(issued.len(), 1, "precondition: one attempt went out");
let body = signed_response_body(&target_identity, issued[0] ^ 1, &coords);
node.handle_lookup_response(&from, &body).await;
assert!(
!node.coord_cache().contains(&target, wall_clock_ms()),
"a response bearing an id we never issued must not reach the coordinate cache"
);
assert!(
node.lookup.pending_lookups.contains_key(&target),
"it must not cancel the lookup that is genuinely outstanding"
);
assert_eq!(node.metrics().lookup.resp_unsolicited.get(), 1);
}
#[tokio::test]
async fn test_a_response_matching_a_pending_attempt_is_accepted_and_clears_the_pending_lookup() {
// The healthy path. A fix that reds a legitimate lookup is no use, and
// this is the test that catches it.
let mut node = make_node();
let from = make_node_addr(0xAA);
let target_identity = Identity::generate();
let target = *target_identity.node_addr();
let coords = register_lookup_target(&mut node, &target_identity);
node.initiate_lookup(&target, 5).await;
let issued = node.lookup.pending_lookups.get(&target).unwrap().ids[0];
let body = signed_response_body(&target_identity, issued, &coords);
node.handle_lookup_response(&from, &body).await;
assert_eq!(
node.coord_cache().get(&target, wall_clock_ms()),
Some(&coords),
"a response to our own outstanding request must be cached"
);
assert!(
!node.lookup.pending_lookups.contains_key(&target),
"accepting it must clear the pending lookup"
);
assert_eq!(node.metrics().lookup.resp_accepted.get(), 1);
assert_eq!(node.metrics().lookup.resp_unsolicited.get(), 0);
}
#[tokio::test]
async fn test_a_late_response_for_an_earlier_retry_attempt_is_still_accepted() {
// Each retry draws a fresh id, and on any link with more than a second
// of round trip the reply to an earlier attempt is the common case. A
// correlator that remembered only the newest id would drop it.
let mut node = make_node();
let from = make_node_addr(0xAA);
let target_identity = Identity::generate();
let target = *target_identity.node_addr();
let coords = register_lookup_target(&mut node, &target_identity);
node.initiate_lookup(&target, 5).await;
node.initiate_lookup(&target, 5).await;
let issued = node
.lookup
.pending_lookups
.get(&target)
.unwrap()
.ids
.clone();
assert_eq!(issued.len(), 2, "precondition: two attempts, two ids");
let body = signed_response_body(&target_identity, issued[0], &coords);
node.handle_lookup_response(&from, &body).await;
assert_eq!(
node.coord_cache().get(&target, wall_clock_ms()),
Some(&coords),
"the first attempt's id is still ours and its answer must be accepted"
);
}
#[tokio::test]
async fn test_a_second_genuine_response_after_the_first_is_accepted_is_dropped() {
// The request is flooded to every qualifying tree peer, so duplicate
// replies are routine. They are dropped at the correlation gate, which
// gives the unsolicited counter a nonzero floor in healthy operation:
// it is not by itself a sign of attack traffic.
let mut node = make_node();
let from = make_node_addr(0xAA);
let target_identity = Identity::generate();
let target = *target_identity.node_addr();
let coords = register_lookup_target(&mut node, &target_identity);
node.initiate_lookup(&target, 5).await;
let issued = node.lookup.pending_lookups.get(&target).unwrap().ids[0];
let body = signed_response_body(&target_identity, issued, &coords);
node.handle_lookup_response(&from, &body).await;
node.handle_lookup_response(&from, &body).await;
assert_eq!(
node.coord_cache().get(&target, wall_clock_ms()),
Some(&coords),
"the value written by the first response must still be there"
);
assert_eq!(
node.metrics().lookup.resp_accepted.get(),
1,
"only the first of the two answers our request"
);
assert_eq!(
node.metrics().lookup.resp_unsolicited.get(),
1,
"the duplicate is counted, which is why the counter has a healthy floor"
);
}
#[tokio::test]
async fn test_a_validly_signed_response_for_a_retired_lookup_is_dropped() {
// The pending entry's lifetime is what bounds how stale an accepted
// coordinate can be. Once the retry ladder is exhausted and the entry
// goes, a transit node holding the genuine reply can no longer deliver
// it late.
let mut node = make_node();
let from = make_node_addr(0xAA);
let target_identity = Identity::generate();
let target = *target_identity.node_addr();
let coords = register_lookup_target(&mut node, &target_identity);
node.initiate_lookup(&target, 5).await;
let issued = node.lookup.pending_lookups.get(&target).unwrap().ids[0];
// Drive the whole ladder: three retries, then the final timeout.
let mut now_ms = Node::now_ms();
for _ in 0..4 {
now_ms += 100_000;
node.check_pending_lookups(now_ms).await;
}
assert!(
!node.lookup.pending_lookups.contains_key(&target),
"precondition: the ladder retired the lookup"
);
let body = signed_response_body(&target_identity, issued, &coords);
node.handle_lookup_response(&from, &body).await;
assert!(
!node.coord_cache().contains(&target, wall_clock_ms()),
"a reply to a retired lookup must not install a coordinate"
);
assert_eq!(node.metrics().lookup.resp_unsolicited.get(), 1);
}
#[test]
fn pending_lookup_id_set_evicts_the_oldest_id_rather_than_refusing_the_newest() {
// The retry ladder is operator configuration and can be longer than the
// recorded-id cap. Refusing the newest id would discard the attempt most
// likely to be answered and fail a healthy lookup.
let mut pending = crate::proto::lookup::PendingLookup::new(0);
for id in 0..12u64 {
pending.record(id);
}
assert!(
pending.matches(11),
"the newest attempt's id must always be remembered"
);
assert!(!pending.matches(0), "the oldest id is the one evicted");
assert_eq!(pending.ids.len(), 8, "the set stays bounded");
}
// ============================================================================
// Unit Tests — RecentRequest Expiry
// ============================================================================
@@ -345,6 +629,212 @@ async fn test_recent_request_expiry() {
assert!(node.lookup.recent_requests.contains_key(&789));
}
// ============================================================================
// Unit Tests — dedup cache capacity policy
// ============================================================================
use crate::proto::lookup::{MAX_RECENT_LOOKUP_REQUESTS, MIN_RECENT_PER_PEER};
/// Encode a LookupRequest for `target` carrying `request_id`, ready for
/// `handle_lookup_request` (which is handed the payload without the
/// msg_type byte).
fn lookup_request_payload(request_id: u64, target: &crate::NodeAddr) -> Vec<u8> {
let origin = make_node_addr(0xCC);
let coords = TreeCoordinate::from_addrs(vec![origin, make_node_addr(0)]).unwrap();
LookupRequest::new(request_id, *target, origin, coords, 5, 0).encode()[1..].to_vec()
}
/// Deliver `count` distinct requests from `from`, ids starting at `first_id`.
async fn flood_requests(node: &mut Node, from: &crate::NodeAddr, first_id: u64, count: u64) {
let target = make_node_addr(0xBB);
for i in 0..count {
let payload = lookup_request_payload(first_id + i, &target);
node.handle_lookup_request(from, &payload).await;
}
}
/// Register `count` peers so the per-peer share of the dedup cache is the
/// floor rather than the whole cache, and return their addresses.
fn register_peers(node: &mut Node, count: usize) -> Vec<crate::NodeAddr> {
(0..count)
.map(|i| {
let identity = Identity::generate();
let addr = *identity.node_addr();
let peer_identity = crate::PeerIdentity::from_pubkey(identity.pubkey());
node.peers.insert(
addr,
ActivePeer::new(peer_identity, LinkId::new(i as u64), 0),
);
addr
})
.collect()
}
#[tokio::test]
async fn test_a_full_dedup_cache_admits_the_new_request_by_evicting_the_oldest() {
// A full cache used to drop the arriving request, which let one peer
// spend 4096 fresh request_ids and stop the node forwarding anyone
// else's lookups until the entries aged out.
let mut node = make_node();
let from = make_node_addr(0xAA);
flood_requests(&mut node, &from, 1, MAX_RECENT_LOOKUP_REQUESTS as u64).await;
assert_eq!(
node.lookup.recent_requests.len(),
MAX_RECENT_LOOKUP_REQUESTS,
"precondition: the cache is full, or the rest observes nothing"
);
let payload = lookup_request_payload(u64::MAX, &make_node_addr(0xBB));
node.handle_lookup_request(&from, &payload).await;
assert!(
node.lookup.recent_requests.contains_key(&u64::MAX),
"the arriving request must be recorded, so its response can be routed back"
);
assert!(
!node.lookup.recent_requests.contains_key(&1),
"room is made by dropping the oldest entry"
);
assert_eq!(
node.lookup.recent_requests.len(),
MAX_RECENT_LOOKUP_REQUESTS,
"the cache stays at its bound"
);
assert_eq!(node.metrics().lookup.req_dedup_evicted.get(), 1);
assert_eq!(
node.metrics().lookup.req_dedup_cache_full.get(),
0,
"the cache-full drop is gone, and its counter stays frozen at zero"
);
}
#[tokio::test]
async fn test_a_flooding_peer_evicts_only_its_own_dedup_entries() {
// The whole point of partitioning the cache by link peer: one peer
// filling its share must not cost another peer the reverse path its own
// lookup depends on.
let mut node = make_node();
let peers = register_peers(&mut node, 64);
let flooder = peers[0];
let light = peers[1];
let payload = lookup_request_payload(7, &make_node_addr(0xBB));
node.handle_lookup_request(&light, &payload).await;
// One over the share, so the flooder pays for its own admission.
flood_requests(&mut node, &flooder, 1000, MIN_RECENT_PER_PEER as u64 + 1).await;
assert!(
node.lookup.recent_requests.contains_key(&7),
"a light peer's reverse-path entry must survive a neighbour's flood"
);
assert!(
!node.lookup.recent_requests.contains_key(&1000),
"the flooder's own oldest entry is what pays for its newest"
);
assert!(
node.lookup
.recent_requests
.contains_key(&(1000 + MIN_RECENT_PER_PEER as u64)),
"and its newest is admitted rather than dropped"
);
}
#[tokio::test]
async fn test_a_node_whose_dedup_cache_is_flooded_still_answers_a_lookup_for_itself() {
// The availability claim. Filling the cache used to make the node
// unresolvable, because the cache-full drop sat ahead of the check for
// whether the request names us.
let mut node = make_node();
let flooder = make_node_addr(0xAA);
let other = make_node_addr(0xAB);
flood_requests(&mut node, &flooder, 1, MAX_RECENT_LOOKUP_REQUESTS as u64).await;
let my_addr = *node.node_addr();
let payload = lookup_request_payload(u64::MAX, &my_addr);
node.handle_lookup_request(&other, &payload).await;
assert_eq!(
node.metrics().lookup.req_target_is_us.get(),
1,
"a flooded cache must not stop the node answering lookups for itself"
);
}
#[tokio::test]
async fn test_the_dedup_index_stays_level_with_the_cache_across_insert_duplicate_and_purge() {
// Two containers where there was one, so the desync is the maintenance
// risk. Everything the eviction policy decides reads the index, so an
// index that has drifted evicts the wrong entry or none at all.
let mut node = make_node();
let peers = register_peers(&mut node, 64);
flood_requests(&mut node, &peers[0], 1, 70).await;
flood_requests(&mut node, &peers[1], 500, 5).await;
// Duplicates, which must not be indexed twice.
flood_requests(&mut node, &peers[1], 500, 5).await;
let indexed: usize = node
.lookup
.recent_by_peer
.values()
.map(|ids| ids.len())
.sum();
assert_eq!(
indexed,
node.lookup.recent_requests.len(),
"every cached request is indexed exactly once"
);
// Age everything out and purge through the ordinary request path.
let expiry_ms = node.config().node.lookup.recent_expiry_secs * 1000;
let future = Node::now_ms() + expiry_ms + 1;
node.purge_expired_requests(future);
assert!(
node.lookup.recent_requests.is_empty(),
"precondition: the purge removed everything"
);
assert!(
node.lookup.recent_by_peer.is_empty(),
"the index must not keep entries the cache no longer holds"
);
}
#[tokio::test]
async fn test_answering_lookups_for_ourselves_stops_at_the_per_peer_signing_budget() {
// Each answer costs a fresh Schnorr signature, because the proof is
// bound to the requester's request_id and cannot be reused. Without a
// budget, one neighbour sets this node's signing rate.
let mut node = make_node();
node.set_discovery_sign_budget(3.0, 0.0);
let from = make_node_addr(0xAA);
let other = make_node_addr(0xAB);
let my_addr = *node.node_addr();
for id in 0..4u64 {
let payload = lookup_request_payload(id, &my_addr);
node.handle_lookup_request(&from, &payload).await;
}
assert_eq!(
node.metrics().lookup.req_target_is_us.get(),
3,
"the burst is answered and the fourth request is not"
);
assert_eq!(node.metrics().lookup.req_sign_rate_limited.get(), 1);
let payload = lookup_request_payload(100, &my_addr);
node.handle_lookup_request(&other, &payload).await;
assert_eq!(
node.metrics().lookup.req_target_is_us.get(),
4,
"one peer spending its budget must not make the node unresolvable through another"
);
}
// ============================================================================
// Integration Tests — Multi-Node Forwarding
// ============================================================================
@@ -888,6 +1378,8 @@ async fn test_originator_stores_path_mtu_in_cache() {
node.register_identity(target, target_identity.pubkey_full());
seed_pending_lookup(&mut node, target, 800);
let proof_data = LookupResponse::proof_bytes(800, &target, &coords);
let proof = target_identity.sign(&proof_data);
@@ -932,6 +1424,8 @@ async fn test_originator_ignores_sub_floor_path_mtu_but_still_caches_coords() {
node.register_identity(target, target_identity.pubkey_full());
seed_pending_lookup(&mut node, target, 801);
let proof_data = LookupResponse::proof_bytes(801, &target, &coords);
let proof = target_identity.sign(&proof_data);
@@ -999,6 +1493,8 @@ async fn test_actionable_lookup_response_path_mtu_does_not_bump_below_floor_coun
node.register_identity(target, target_identity.pubkey_full());
seed_pending_lookup(&mut node, target, 802);
let proof_data = LookupResponse::proof_bytes(802, &target, &coords);
let proof = target_identity.sign(&proof_data);
@@ -1045,6 +1541,8 @@ async fn test_originator_lookup_response_keeps_tighter_path_mtu_lookup() {
let target_fips = crate::FipsAddress::from_node_addr(&target);
node.path_mtu_lookup_insert(target_fips, 1280);
seed_pending_lookup(&mut node, target, 800);
let proof_data = LookupResponse::proof_bytes(800, &target, &coords);
let proof = target_identity.sign(&proof_data);
@@ -1079,6 +1577,7 @@ fn make_verified_lookup_response(
let coords = TreeCoordinate::from_addrs(vec![target, root]).unwrap();
node.register_identity(target, target_identity.pubkey_full());
seed_pending_lookup(node, target, request_id);
let proof_data = LookupResponse::proof_bytes(request_id, &target, &coords);
let proof = target_identity.sign(&proof_data);
@@ -1312,7 +1811,7 @@ async fn test_open_discovery_sweep_queues_eligible_skips_filtered() {
/// verified indirectly: each `req_initiated` increment corresponds
/// to one fresh `initiate_lookup` call.
/// 3. **Final-timeout state transitions** — `pending_lookups` entry is
/// removed, `discovery.resp_timed_out` counter ticks, queued packet
/// removed, `lookup.resp_timed_out` counter ticks, queued packet
/// is drained, and an ICMPv6 Destination Unreachable frame is
/// emitted via the TUN sender.
///
@@ -1483,7 +1982,7 @@ async fn test_check_pending_lookups_default_sequence_unreachable() {
assert_eq!(
node.metrics().lookup.resp_timed_out.get(),
baseline_timed_out + 1,
"final timeout must increment discovery.resp_timed_out"
"final timeout must increment lookup.resp_timed_out"
);
// No additional initiate_lookup on the timeout step.
assert_eq!(
+89
View File
@@ -5,11 +5,13 @@
//! multi-hop forwarding through live node topologies.
use super::*;
use crate::node::peer_error_budget::PEER_ERROR_BURST;
use crate::proto::fsp::wire::{FSP_FLAG_CP, build_fsp_header};
use crate::proto::fsp::{SessionAck, SessionSetup};
use crate::proto::link::SessionDatagram;
use crate::proto::stp::TreeCoordinate;
use crate::proto::stp::encode_coords;
use spanning_tree::{
TestNode, cleanup_nodes, populate_all_coord_caches, process_available_packets, run_tree_test,
verify_tree_convergence,
@@ -1195,3 +1197,90 @@ async fn test_coord_cache_warming_short_inner_payload_is_dropped_not_panic() {
"24 frames of 39..=46 and 47..=62 outer bytes sum to 1212"
);
}
// --- Emission bounds on induced routing errors ---
/// A distinct destination per index, standing for the fresh `dest_addr` a
/// flooding sender puts on every datagram to escape the per-destination gate.
fn minted_dest(val: u32) -> NodeAddr {
let mut bytes = [0u8; 16];
bytes[..4].copy_from_slice(&val.to_le_bytes());
bytes[15] = 0xfe;
NodeAddr::from_bytes(bytes)
}
/// Feed one transit datagram whose destination this node cannot route.
async fn inject_unroutable(node: &mut Node, from: &NodeAddr, src: NodeAddr, dest: NodeAddr) {
let dg = SessionDatagram::new(src, dest, vec![0x10, 0x00, 0x00, 0x00]).with_ttl(8);
let encoded = dg.encode();
node.handle_session_datagram(from, &encoded[1..], false)
.await;
}
#[tokio::test]
async fn one_link_peer_cannot_induce_unbounded_errors_by_varying_the_destination() {
let mut node = make_node();
let attacker = make_node_addr(0xAA);
let overshoot = 10u32;
for i in 0..PEER_ERROR_BURST + overshoot {
// Fresh destination and fresh spoofed source per packet: neither
// address-keyed gate sees a repeat.
inject_unroutable(
&mut node,
&attacker,
minted_dest(i + 1_000_000),
minted_dest(i),
)
.await;
}
let errors = &node.metrics().errors;
assert_eq!(
errors.emit_over_dest_interval.get(),
0,
"the per-destination gate cannot bound a sender that varies the destination"
);
assert_eq!(
errors.emit_over_peer_budget.get(),
u64::from(overshoot),
"everything past the link peer's burst must be refused"
);
}
#[tokio::test]
async fn a_destination_suppressed_error_does_not_spend_the_link_peer_budget() {
let mut node = make_node();
let peer = make_node_addr(0xAA);
let src = make_node_addr(0x01);
let dest = make_node_addr(0x02);
let injected = PEER_ERROR_BURST * 4;
for _ in 0..injected {
inject_unroutable(&mut node, &peer, src, dest).await;
}
let errors = &node.metrics().errors;
assert_eq!(
errors.emit_over_peer_budget.get(),
0,
"an outage on one destination must not spend the peer's budget for the others"
);
assert_eq!(
errors.emit_over_dest_interval.get(),
u64::from(injected - 1),
"only the first error for a destination goes out within the interval"
);
}
#[tokio::test]
async fn a_single_unroutable_datagram_still_produces_its_error() {
let mut node = make_node();
let peer = make_node_addr(0xAA);
inject_unroutable(&mut node, &peer, make_node_addr(0x01), make_node_addr(0x02)).await;
let errors = &node.metrics().errors;
assert_eq!(errors.emit_over_peer_budget.get(), 0);
assert_eq!(errors.emit_over_dest_interval.get(), 0);
}
+275
View File
@@ -1940,3 +1940,278 @@ async fn a_stale_reverse_address_entry_does_not_hide_a_peer_reachable_by_address
returns above the insert that would overwrite the stale entry"
);
}
// ===== Epoch-restart dampening =====
//
// An epoch-mismatch msg1 is authentic, because the epoch travels inside the
// AEAD, but it is replayable: a captured one stays valid forever and
// accepting it destroys a working peering. Two receiver-local conditions
// gate the teardown, and each of the first two cases below breaks one of
// them.
/// Install a peering for `initiator` that carries an epoch its genuine msg1
/// does not, and that has gone `idle_secs` without authenticated inbound
/// traffic. Returns the link the peering is bound to.
fn install_peering_at_a_different_epoch(
node: &mut Node,
initiator: &Node,
transport_id: TransportId,
source_addr: &TransportAddr,
idle_secs: u64,
) -> LinkId {
use crate::peer::ActivePeer;
let identity = PeerIdentity::from_pubkey_full(initiator.identity().pubkey_full());
let node_addr = *identity.node_addr();
let link_id = node.allocate_link_id();
let authenticated_at = Node::now_ms().saturating_sub(idle_secs * 1000);
let mut peer = ActivePeer::new(identity, link_id, authenticated_at);
peer.set_current_addr(transport_id, source_addr.clone());
// Anything but the epoch the initiator's msg1 carries, so the msg1 reads
// as a restart.
peer.set_remote_epoch(Some([0xAA; 8]));
node.peers.insert(node_addr, peer);
node.addr_to_link
.insert((transport_id, source_addr.clone()), link_id);
link_id
}
/// A peering long enough past its last authenticated inbound frame that the
/// liveness gate does not hold the restart back.
const IDLE_SECS: u64 = 60;
#[tokio::test]
async fn an_epoch_mismatch_msg1_against_a_live_peering_leaves_it_intact() {
let transport_id = TransportId::new(1);
let mut node = make_node();
let initiator = make_node();
let initiator_addr = node_addr_of(&initiator);
let source_addr = TransportAddr::from_string("127.0.0.1:41001");
// The peering is carrying authenticated traffic: it decrypted a frame a
// moment ago. Under replay that is always the case, because the genuine
// peer is heartbeating.
let peer_link =
install_peering_at_a_different_epoch(&mut node, &initiator, transport_id, &source_addr, 0);
let bad_state_before = node.stats().handshake.bad_state;
node.handle_msg1(ReceivedPacket::with_timestamp(
transport_id,
source_addr.clone(),
genuine_msg1(&initiator, &node),
Node::now_ms(),
))
.await;
let peer = node
.get_peer(&initiator_addr)
.expect("a live peering must survive an epoch-mismatch msg1");
assert_eq!(
peer.link_id(),
peer_link,
"the peering must be the one that was already established, not a \
replacement promoted from the msg1"
);
assert_eq!(
peer.remote_epoch(),
Some([0xAA; 8]),
"the stored epoch must not have moved to the one the msg1 carried"
);
assert_eq!(
node.connection_count(),
0,
"the dropped msg1 must leave no connection behind"
);
assert_eq!(
node.stats().handshake.bad_state - bad_state_before,
1,
"the drop must be counted"
);
}
#[tokio::test]
async fn a_second_epoch_change_inside_the_dampening_interval_leaves_the_peering_intact() {
let transport_id = TransportId::new(1);
let mut node = make_node();
let initiator = make_node();
let initiator_addr = node_addr_of(&initiator);
let source_addr = TransportAddr::from_string("127.0.0.1:41002");
// First epoch change: the peering is genuinely idle, so it is accepted
// and stamps the dampener.
let first_link = install_peering_at_a_different_epoch(
&mut node,
&initiator,
transport_id,
&source_addr,
IDLE_SECS,
);
node.handle_msg1(ReceivedPacket::with_timestamp(
transport_id,
source_addr.clone(),
genuine_msg1(&initiator, &node),
Node::now_ms(),
))
.await;
let promoted = node
.get_peer(&initiator_addr)
.expect("the first epoch change must be accepted");
assert_ne!(
promoted.link_id(),
first_link,
"the first epoch change must have replaced the peering"
);
// The peer moves epoch again straight away. Nothing about the second
// msg1 is distinguishable from the first, which is why the interval,
// not the message, has to be what refuses it.
let second_link = install_peering_at_a_different_epoch(
&mut node,
&initiator,
transport_id,
&source_addr,
IDLE_SECS,
);
let bad_state_before = node.stats().handshake.bad_state;
node.handle_msg1(ReceivedPacket::with_timestamp(
transport_id,
source_addr.clone(),
genuine_msg1(&initiator, &node),
Node::now_ms(),
))
.await;
let peer = node
.get_peer(&initiator_addr)
.expect("a second epoch change inside the interval must not tear the peering down");
assert_eq!(
peer.link_id(),
second_link,
"the peering must be the one that was already established"
);
assert_eq!(
peer.remote_epoch(),
Some([0xAA; 8]),
"the stored epoch must not have moved to the one the msg1 carried"
);
assert_eq!(
node.connection_count(),
0,
"the dropped msg1 must leave no connection behind"
);
assert_eq!(
node.stats().handshake.bad_state - bad_state_before,
1,
"the drop must be counted"
);
}
/// A REFUSED epoch-mismatch msg1 must not slide the dampening window.
///
/// The stamp is written on acceptance only. If it were written on every
/// sighting, a sender replaying a captured msg1 faster than the interval would
/// hold the window open indefinitely and a genuinely restarting peer could
/// never re-peer — the refusal would become the denial of service it exists to
/// prevent. The 15s interval is far longer than a test can wait, so this pins
/// the ordering directly: after an accepted change stamps the dampener,
/// repeated refused msg1s must leave that stamp byte-identical.
///
/// Discriminator: moving the three stamp lines above the gate in
/// `InboundDecision::RestartThenPromote` reds this and nothing else in the
/// module.
#[tokio::test]
async fn a_refused_epoch_change_does_not_slide_the_dampening_window() {
let transport_id = TransportId::new(1);
let mut node = make_node();
let initiator = make_node();
let initiator_addr = node_addr_of(&initiator);
let source_addr = TransportAddr::from_string("127.0.0.1:41009");
// Accepted change: stamps the dampener.
install_peering_at_a_different_epoch(
&mut node,
&initiator,
transport_id,
&source_addr,
IDLE_SECS,
);
node.handle_msg1(ReceivedPacket::with_timestamp(
transport_id,
source_addr.clone(),
genuine_msg1(&initiator, &node),
Node::now_ms(),
))
.await;
let stamped = node
.restart_dampener_stamp(&initiator_addr)
.expect("an accepted epoch change must stamp the dampener");
// Sustained replay: every one of these is refused by the dampener.
for _ in 0..5 {
install_peering_at_a_different_epoch(
&mut node,
&initiator,
transport_id,
&source_addr,
IDLE_SECS,
);
node.handle_msg1(ReceivedPacket::with_timestamp(
transport_id,
source_addr.clone(),
genuine_msg1(&initiator, &node),
Node::now_ms(),
))
.await;
}
assert_eq!(
node.restart_dampener_stamp(&initiator_addr),
Some(stamped),
"a refused epoch change must not restamp the dampener; if it does, a \
sustained replay holds the window open and starves a real restart"
);
}
/// Healthy path, and NOT discriminating: this passes with or without the
/// gates. It is here so that tightening either one, or a bug that stamps the
/// dampener on a refusal, reds the suite instead of silently refusing every
/// genuine restart.
#[tokio::test]
async fn a_first_epoch_change_against_a_silent_peering_still_restarts_it() {
let transport_id = TransportId::new(1);
let mut node = make_node();
let initiator = make_node();
let initiator_addr = node_addr_of(&initiator);
let source_addr = TransportAddr::from_string("127.0.0.1:41003");
let stale_link = install_peering_at_a_different_epoch(
&mut node,
&initiator,
transport_id,
&source_addr,
IDLE_SECS,
);
node.handle_msg1(ReceivedPacket::with_timestamp(
transport_id,
source_addr.clone(),
genuine_msg1(&initiator, &node),
Node::now_ms(),
))
.await;
let peer = node
.get_peer(&initiator_addr)
.expect("a restart with no prior epoch change must be promoted");
assert_ne!(
peer.link_id(),
stale_link,
"the stale peering must have been torn down and replaced"
);
assert_eq!(
peer.remote_epoch(),
Some(initiator.startup_epoch()),
"the replacement must carry the epoch the msg1 announced"
);
}
+492 -1
View File
@@ -2345,6 +2345,17 @@ fn install_halfopen(node: &mut Node, claimed: NodeAddr) {
node.sessions.insert(claimed, entry);
}
/// Record that this node put a frame of `wire_len` bytes on the wire toward
/// `dest`, which is what corroborates a reactive `MtuExceeded` reporting a
/// smaller bottleneck. Honest path-MTU discovery produces this by sending;
/// a handler test that installs a session without sending has to state it.
fn note_sent_wire_len(node: &mut Node, dest: &NodeAddr, wire_len: usize) {
node.sessions
.get_mut(dest)
.expect("session must exist to corroborate a report")
.record_sent_wire_len(wire_len);
}
/// Install the entry `initiate_session` creates: an address this node chose
/// itself, with the handshake still in flight and MMP not yet initialized.
fn install_initiating(node: &mut Node, remote: &Identity) {
@@ -2380,6 +2391,7 @@ async fn test_handle_mtu_exceeded_writes_path_mtu_lookup_when_empty() {
"lookup should start empty for this destination"
);
note_sent_wire_len(&mut tn.node, &dest, 1400);
let inner = build_mtu_exceeded_inner(&dest, &reporter, 1280);
tn.node.handle_mtu_exceeded(&reporter, &inner).await;
@@ -2406,6 +2418,7 @@ async fn test_handle_mtu_exceeded_tightens_existing_path_mtu_lookup() {
// response that didn't reflect the forward-path bottleneck).
tn.node.path_mtu_lookup_insert(dest_fips, 1500);
note_sent_wire_len(&mut tn.node, &dest, 1400);
let inner = build_mtu_exceeded_inner(&dest, &reporter, 1280);
tn.node.handle_mtu_exceeded(&reporter, &inner).await;
@@ -2537,8 +2550,9 @@ async fn test_handle_mtu_exceeded_at_the_floor_still_writes_path_mtu_lookup() {
let dest = *remote.node_addr();
let reporter = NodeAddr::from_bytes([0xBB; 16]);
let dest_fips = crate::FipsAddress::from_node_addr(&dest);
let floor = crate::upper::icmp::MIN_ACTIONABLE_PATH_MTU;
let floor = crate::upper::icmp::MIN_REACTIVE_PATH_MTU;
note_sent_wire_len(&mut tn.node, &dest, 1400);
let inner = build_mtu_exceeded_inner(&dest, &reporter, floor);
tn.node.handle_mtu_exceeded(&reporter, &inner).await;
@@ -2986,6 +3000,7 @@ async fn test_mtu_exceeded_for_a_session_we_initiated_seeds_path_mtu_lookup_befo
let reporter = NodeAddr::from_bytes([0xBB; 16]);
let dest_fips = crate::FipsAddress::from_node_addr(&dest);
note_sent_wire_len(&mut node, &dest, 1400);
let inner = build_mtu_exceeded_inner(&dest, &reporter, 1280);
node.handle_mtu_exceeded(&reporter, &inner).await;
@@ -3012,6 +3027,7 @@ async fn test_mtu_exceeded_from_a_third_party_forwarder_still_tightens_an_active
let reporter = NodeAddr::from_bytes([0xBB; 16]);
let dest_fips = crate::FipsAddress::from_node_addr(&dest);
note_sent_wire_len(&mut node, &dest, 1400);
let inner = build_mtu_exceeded_inner(&dest, &reporter, 1280);
node.handle_mtu_exceeded(&reporter, &inner).await;
@@ -4321,6 +4337,258 @@ async fn test_a_drained_stranger_bucket_still_admits_a_setup_naming_an_establish
cleanup_nodes(&mut nodes).await;
}
// ============================================================================
// Integration tests: the session-table population cap
// ============================================================================
/// Build a two-node routable mesh with the session table capped for a test
/// and the setup limiter opened wide, so the cap is the only thing refusing.
async fn make_session_capped_pair(max_sessions: usize) -> Vec<TestNode> {
let configs = (0..2)
.map(|_| {
let mut config = Config::new();
config.node.rekey.enabled = false;
config.node.limits.max_sessions = max_sessions;
config.node.rate_limit.session_setup_burst = 10_000;
config.node.rate_limit.session_setup_rate = 10_000.0;
config
})
.collect();
let mut nodes = run_tree_test_with_configs(configs, &[(0, 1)]).await;
verify_tree_convergence(&nodes);
populate_all_coord_caches(&mut nodes);
nodes
}
#[tokio::test]
async fn test_forged_setups_stop_growing_the_session_table_once_the_cap_is_reached() {
// The table was the one remotely-grown map with no bound: each setup from
// an address nobody has seen inserted an entry, and neither existing limit
// reached it, the setup limiter governing arrival rate rather than
// population and the idle purge only reaching entries a peer stops using.
const MAX: usize = 8;
let share = MAX / 2;
let mut nodes = make_session_capped_pair(MAX).await;
for _ in 0..share {
deliver_forged_setup_over_link(&mut nodes).await;
}
assert_eq!(
nodes[1].node.sessions.len(),
share,
"the admissible entries must be admitted, or this test would pass for \
the wrong reason"
);
assert_eq!(nodes[1].node.stats().session.half_open_full, 0);
// Every SessionAck goes out through `send_session_datagram`, the only
// thing bumping this counter on a node with no transit traffic. A refused
// setup must not move it.
let originated = nodes[1].node.metrics().forwarding.originated_packets.get();
for _ in 0..4 {
deliver_forged_setup_over_link(&mut nodes).await;
}
assert_eq!(
nodes[1].node.sessions.len(),
share,
"a table at its bound must stop growing"
);
assert_eq!(
nodes[1].node.stats().session.half_open_full,
4,
"each refusal must be counted; the DEBUG line is invisible by default"
);
assert_eq!(
nodes[1].node.metrics().forwarding.originated_packets.get(),
originated,
"a refused setup must emit nothing at all"
);
cleanup_nodes(&mut nodes).await;
}
#[tokio::test]
async fn test_a_setup_that_would_grow_a_full_table_is_refused_and_counted() {
// The table-full arm specifically: one established entry against a cap of
// one, so the half-open share is not what refuses.
let mut nodes = make_session_capped_pair(1).await;
establish_pair_session(&mut nodes).await;
assert_eq!(
nodes[1].node.sessions.len(),
1,
"precondition: the table is full with the established peer"
);
let originated = nodes[1].node.metrics().forwarding.originated_packets.get();
deliver_forged_setup_over_link(&mut nodes).await;
assert_eq!(
nodes[1].node.sessions.len(),
1,
"a full table must not grow for a stranger"
);
assert_eq!(nodes[1].node.stats().session.table_full, 1);
assert_eq!(
nodes[1].node.metrics().forwarding.originated_packets.get(),
originated,
"a refused setup must cost no ack"
);
cleanup_nodes(&mut nodes).await;
}
#[tokio::test]
async fn test_a_full_session_table_still_serves_a_setup_naming_an_existing_entry() {
// The guard against writing the cap as "refuse strangers". A setup for an
// entry already present cannot grow the table, and refusing it would break
// the duplicate-ack resend an initiator depends on.
let mut nodes = make_session_capped_pair(1).await;
establish_pair_session(&mut nodes).await;
let node0_addr = *nodes[0].node.node_addr();
let node1_addr = *nodes[1].node.node_addr();
let refused_before = nodes[1].node.stats().session.table_full;
let originated = nodes[1].node.metrics().forwarding.originated_packets.get();
// A setup naming the established peer: the shape an inbound rekey has.
let setup = forge_setup_from_stranger(&nodes);
let datagram = SessionDatagram::new(node0_addr, node1_addr, setup).with_ttl(64);
let encoded = datagram.encode();
nodes[1]
.node
.handle_session_datagram(&node0_addr, &encoded[1..], false)
.await;
assert_eq!(
nodes[1].node.stats().session.table_full,
refused_before,
"a setup that cannot grow the table must not be refused by the cap"
);
assert!(
nodes[1].node.metrics().forwarding.originated_packets.get() > originated,
"and it must still be answered"
);
cleanup_nodes(&mut nodes).await;
}
#[tokio::test]
async fn test_a_full_session_table_does_not_evict_an_established_session() {
// Pins refuse-not-evict. The setup that triggers the decision is
// unauthenticated at that point, so evicting would hand a stranger a way
// to tear down a session it has nothing to do with.
let mut nodes = make_session_capped_pair(1).await;
establish_pair_session(&mut nodes).await;
let node0_addr = *nodes[0].node.node_addr();
for _ in 0..4 {
deliver_forged_setup_over_link(&mut nodes).await;
}
assert!(
nodes[1]
.node
.get_session(&node0_addr)
.expect("the established session must survive a flood at the cap")
.is_established(),
"a stranger's setup must never cost an established peer its session"
);
cleanup_nodes(&mut nodes).await;
}
#[tokio::test]
async fn test_the_session_table_admits_again_after_the_handshake_reaper_drains_it() {
// The cap is a ceiling, not a latch: half-open entries are reaped after
// `handshake_timeout_secs` and the room they free must be usable.
const MAX: usize = 8;
let share = MAX / 2;
let mut nodes = make_session_capped_pair(MAX).await;
for _ in 0..(share + 2) {
deliver_forged_setup_over_link(&mut nodes).await;
}
assert!(
nodes[1].node.stats().session.half_open_full > 0,
"precondition: the table is refusing before the reaper runs"
);
let timeout_ms = nodes[1]
.node
.config()
.node
.rate_limit
.handshake_timeout_secs
* 1000;
let now_ms = Node::now_ms();
nodes[1]
.node
.resend_pending_session_handshakes(now_ms + timeout_ms + 1)
.await;
assert_eq!(
nodes[1].node.sessions.len(),
0,
"precondition: the reaper freed the half-open entries"
);
deliver_forged_setup_over_link(&mut nodes).await;
assert_eq!(
nodes[1].node.sessions.len(),
1,
"room freed by the reaper must be usable, or the cap is a latch"
);
cleanup_nodes(&mut nodes).await;
}
#[tokio::test]
async fn test_half_open_setups_cannot_consume_more_than_their_share_of_the_table() {
// Half-open entries are unauthenticated and cheap to create, so they are
// held to a share of the table rather than being allowed to fill it and
// deny it to every peer that would complete a handshake.
const MAX: usize = 16;
let share = MAX / 2;
let mut nodes = make_session_capped_pair(MAX).await;
for _ in 0..(share + 2) {
deliver_forged_setup_over_link(&mut nodes).await;
}
assert_eq!(
nodes[1].node.sessions.len(),
share,
"half-open entries must stop at their share, well below the table cap"
);
assert_eq!(nodes[1].node.stats().session.half_open_full, 2);
assert_eq!(
nodes[1].node.stats().session.table_full,
0,
"the table itself is not full, so the refusals must be attributed to \
the share rather than to the cap"
);
cleanup_nodes(&mut nodes).await;
}
#[test]
fn test_session_entry_size_stays_within_the_budget_the_cap_is_derived_from() {
// The default `max_sessions` is derived from what one entry costs.
// Measured at 6608 bytes of inline state when the cap was written, plus
// heap for the MMP window and handshake payloads, so 1024 sessions is
// roughly 7 MB. This is what fires if a large field is added later and
// the arithmetic behind that default stops holding.
const BUDGET: usize = 8192;
assert!(
std::mem::size_of::<SessionEntry>() <= BUDGET,
"SessionEntry is {} bytes, over the {} the max_sessions default \
assumes; re-derive the default or shrink the entry",
std::mem::size_of::<SessionEntry>(),
BUDGET
);
}
// ============================================================================
// Integration tests: a forged SessionAck against an in-flight initiation
// ============================================================================
@@ -5090,3 +5358,226 @@ async fn test_peer_restart_reestablishes_through_a_pending_session_that_waited_o
cleanup_nodes(&mut nodes).await;
}
// ---------------------------------------------------------------------------
// Reactive MtuExceeded: corroboration against what this node actually sent
// ---------------------------------------------------------------------------
#[tokio::test]
async fn a_reactive_mtu_exceeded_at_the_floor_no_longer_pins_a_session_this_node_has_not_overfilled()
{
// The defect itself. A report of exactly the floor is a legal value, and
// the admission gate cannot tell an honest forwarder from anyone else, so
// one packet drove a bound session's path MTU to the floor and pinned the
// FipsAddress-keyed entry the SYN-time MSS clamp reads. Nothing this node
// sent could have overflowed a hop at that size, so no honest report of it
// exists.
let mut node = make_node();
let remote = Identity::generate();
install_established_session_with_mmp(&mut node, &remote);
let dest = *remote.node_addr();
let reporter = NodeAddr::from_bytes([0xBB; 16]);
let dest_fips = crate::FipsAddress::from_node_addr(&dest);
let before = node
.sessions
.get(&dest)
.and_then(|e| e.mmp())
.map(|m| m.path_mtu.current_mtu());
let inner =
build_mtu_exceeded_inner(&dest, &reporter, crate::upper::icmp::MIN_REACTIVE_PATH_MTU);
node.handle_mtu_exceeded(&reporter, &inner).await;
assert_eq!(
node.sessions
.get(&dest)
.and_then(|e| e.mmp())
.map(|m| m.path_mtu.current_mtu()),
before,
"an uncorroborated report must leave the session path MTU alone"
);
assert_eq!(
node.path_mtu_lookup_get(&dest_fips),
None,
"an uncorroborated report must leave no clamp entry behind"
);
assert_eq!(
node.metrics().errors.mtu_exceeded_uncorroborated.get(),
1,
"the refusal must be counted apart from the below-floor refusal"
);
assert_eq!(
node.metrics().errors.mtu_exceeded_below_floor.get(),
0,
"the floor is not what refused this; the value is exactly at it"
);
}
#[tokio::test]
async fn an_initiating_session_refuses_an_uncorroborated_report_and_accepts_a_corroborated_one() {
// The lookup write is the effect that survives on an initiating session,
// which has no MMP state at all, so this branch needs its own coverage:
// a guard placed on the apply rather than ahead of it would miss it.
let mut node = make_node();
let remote = Identity::generate();
install_initiating(&mut node, &remote);
let dest = *remote.node_addr();
let reporter = NodeAddr::from_bytes([0xBB; 16]);
let dest_fips = crate::FipsAddress::from_node_addr(&dest);
let inner = build_mtu_exceeded_inner(&dest, &reporter, 800);
node.handle_mtu_exceeded(&reporter, &inner).await;
assert_eq!(
node.path_mtu_lookup_get(&dest_fips),
None,
"nothing this node sent could have overflowed a hop at 800 bytes"
);
// A SessionSetup can itself be the datagram that overflows a hop, so an
// initiating session must still be able to act on a real report.
note_sent_wire_len(&mut node, &dest, 1400);
node.handle_mtu_exceeded(&reporter, &inner).await;
assert_eq!(
node.path_mtu_lookup_get(&dest_fips),
Some(800),
"a report corroborated by an oversized send must still be applied"
);
}
#[tokio::test]
async fn a_second_reactive_decrease_needs_its_own_corroborating_send() {
// The evidence is spent on the decrease it vouched for. Otherwise one
// large send early in a session would vouch for every forged report for
// the rest of that session's life.
let mut node = make_node();
let remote = Identity::generate();
install_established_session_with_mmp(&mut node, &remote);
let dest = *remote.node_addr();
let reporter = NodeAddr::from_bytes([0xBB; 16]);
note_sent_wire_len(&mut node, &dest, 1400);
let first = build_mtu_exceeded_inner(&dest, &reporter, 1200);
node.handle_mtu_exceeded(&reporter, &first).await;
assert_eq!(
node.sessions
.get(&dest)
.and_then(|e| e.mmp())
.map(|m| m.path_mtu.current_mtu()),
Some(1200),
"the corroborated first decrease is accepted"
);
let second = build_mtu_exceeded_inner(&dest, &reporter, 600);
node.handle_mtu_exceeded(&reporter, &second).await;
assert_eq!(
node.sessions
.get(&dest)
.and_then(|e| e.mmp())
.map(|m| m.path_mtu.current_mtu()),
Some(1200),
"a further decrease needs evidence of its own"
);
// A genuine re-route onto a smaller hop is preceded by a send that hop
// drops, so the honest sequence still converges.
note_sent_wire_len(&mut node, &dest, 900);
node.handle_mtu_exceeded(&reporter, &second).await;
assert_eq!(
node.sessions
.get(&dest)
.and_then(|e| e.mmp())
.map(|m| m.path_mtu.current_mtu()),
Some(600),
"once this node has again sent something that does not fit, the report applies"
);
}
#[tokio::test]
async fn a_corroborated_report_below_the_reactive_floor_is_still_refused() {
// Corroboration and the floor are independent refusals. A hop that really
// is tiny still cannot drive the clamp into the band where the derived
// MSS degenerates.
let mut node = make_node();
let remote = Identity::generate();
install_established_session_with_mmp(&mut node, &remote);
let dest = *remote.node_addr();
let reporter = NodeAddr::from_bytes([0xBB; 16]);
let dest_fips = crate::FipsAddress::from_node_addr(&dest);
note_sent_wire_len(&mut node, &dest, 1400);
let inner = build_mtu_exceeded_inner(
&dest,
&reporter,
crate::upper::icmp::MIN_REACTIVE_PATH_MTU - 1,
);
node.handle_mtu_exceeded(&reporter, &inner).await;
assert_eq!(node.path_mtu_lookup_get(&dest_fips), None);
assert_eq!(node.metrics().errors.mtu_exceeded_below_floor.get(), 1);
assert_eq!(node.metrics().errors.mtu_exceeded_uncorroborated.get(), 0);
}
#[tokio::test]
async fn the_authenticated_path_mtu_notification_still_applies_at_the_actionable_floor() {
// The reactive guards must not leak onto the carrier that arrives inside
// an established session on the decrypted path, which is authenticated and
// needs no corroboration.
let mut node = make_node();
let remote = Identity::generate();
install_established_session_with_mmp(&mut node, &remote);
let dest = *remote.node_addr();
let floor = crate::upper::icmp::MIN_ACTIONABLE_PATH_MTU;
let body = build_path_mtu_notification_body(floor);
node.handle_session_path_mtu_notification(&dest, &body);
assert_eq!(
node.sessions
.get(&dest)
.and_then(|e| e.mmp())
.map(|m| m.path_mtu.current_mtu()),
Some(floor),
"the authenticated carrier still applies a value at the actionable floor"
);
}
#[tokio::test]
async fn a_path_broken_flood_releases_the_stored_path_mtu_only_once_per_interval() {
use crate::proto::routing::PathBroken;
// PathBroken is unauthenticated and its release discards a bottleneck this
// node learned the hard way. Unlimited, the claim can be repeated as fast
// as it can be sent, so a genuinely learned value never survives.
let mut node = make_node();
let remote = Identity::generate();
install_initiating(&mut node, &remote);
let dest = *remote.node_addr();
let reporter = NodeAddr::from_bytes([0xBB; 16]);
let dest_fips = crate::FipsAddress::from_node_addr(&dest);
let encoded = PathBroken::new(dest, reporter).encode();
let inner = &encoded[5..];
node.path_mtu_lookup_insert(dest_fips, 700);
node.handle_path_broken(&reporter, inner).await;
assert_eq!(
node.path_mtu_lookup_get(&dest_fips),
None,
"the first PathBroken still releases"
);
node.path_mtu_lookup_insert(dest_fips, 700);
node.handle_path_broken(&reporter, inner).await;
assert_eq!(
node.path_mtu_lookup_get(&dest_fips),
Some(700),
"a second release for the same destination inside the interval is refused"
);
}
+70
View File
@@ -2264,6 +2264,48 @@ fn spawn_blackhole_relay() -> String {
format!("ws://127.0.0.1:{port}")
}
/// The author filter on advert selection must not swallow a genuine eviction.
///
/// Dropping foreign-authored events narrows what the selection can return, and
/// the eviction arm is guarded on the relays having answered with nothing at
/// all. If that guard is written too broadly it also suppresses the real case
/// this function exists for: the peer withdrew its advert and the cached entry
/// has to go. Discriminator: a seeded cache entry plus relays that return
/// nothing must still come back `Evicted` with the entry gone.
#[tokio::test]
async fn refetch_still_evicts_a_cached_advert_when_the_relays_return_nothing() {
let peer_npub = Identity::generate().npub();
let mut bootstrap = crate::nostr::NostrRendezvous::new_for_test();
bootstrap
.set_advert_relays_for_test(vec![spawn_blackhole_relay()])
.await;
let endpoint = crate::nostr::OverlayEndpointAdvert {
transport: crate::nostr::OverlayTransportKind::Udp,
addr: "203.0.113.7:2121".to_string(),
};
let advert =
crate::nostr::NostrRendezvous::cached_advert_for_test(peer_npub.clone(), endpoint, 1_000);
bootstrap
.insert_advert_for_test(peer_npub.clone(), advert)
.await;
let outcome = bootstrap.refetch_advert_for_stale_check(&peer_npub).await;
assert_eq!(
outcome,
crate::nostr::NostrRefetchOutcome::Evicted,
"an empty relay answer is still evidence the advert is gone"
);
assert!(
bootstrap
.cached_created_at_for_test(&peer_npub)
.await
.is_none(),
"the stale entry should have been removed from the cache"
);
}
/// The per-tick retry loop must not await the pre-dial advert refetch.
///
/// `process_pending_retries` runs inline on the node's 1s rx-loop tick. Each
@@ -3896,3 +3938,31 @@ fn test_peer_display_name_tracks_alias_change() {
peer_identity.short_npub()
);
}
/// The DNS mesh-interface filter is keyed on the device the node actually
/// created, not on the configured name. macOS and FreeBSD hand out utunN and
/// tunN of the kernel's choosing, so a filter keyed on the configured name
/// resolved to nothing there and was permanently off.
#[cfg(unix)]
#[test]
fn mesh_filter_resolves_the_live_tun_device_rather_than_the_configured_name() {
let loopback = if cfg!(target_os = "macos") {
"lo0"
} else {
"lo"
};
let c_name = std::ffi::CString::new(loopback).unwrap();
let expected = unsafe { libc::if_nametoindex(c_name.as_ptr()) };
if expected == 0 {
return;
}
let mut config = Config::new();
config.tun.name = Some("fips-absent-dev".to_string());
let mut node = Node::new(config).unwrap();
assert_eq!(node.mesh_ifindex(), None);
node.tun_name = Some(loopback.to_string());
assert_eq!(node.mesh_ifindex(), Some(expected));
}
+8
View File
@@ -31,6 +31,14 @@ impl ReplayWindow {
/// Returns true if the counter is acceptable, false if it should be rejected.
/// Does NOT update the window - call `accept` after successful decryption.
pub fn check(&self, counter: u64) -> bool {
// The send side refuses to emit u64::MAX (`CipherState::advance_nonce`,
// `NoiseSession::take_send_counter`), so no conforming peer produces it.
// Refusing it here keeps `accept` from pinning `highest` at the ceiling,
// which would wedge the window against every later counter.
if counter == u64::MAX {
return false;
}
if counter > self.highest {
// New highest - always acceptable
return true;
+35
View File
@@ -387,6 +387,41 @@ fn test_replay_window_reset() {
assert!(window.check(100));
}
#[test]
fn test_replay_window_max_counter_does_not_wedge_the_window() {
let mut window = ReplayWindow::new();
window.accept(100);
// Mirror the real receive path: check, and accept only if the check passed.
if window.check(u64::MAX) {
window.accept(u64::MAX);
}
// The honest peer's next frame must still be acceptable.
assert!(window.check(101), "ceiling frame wedged the window");
}
#[test]
fn test_replay_window_rejects_max_counter() {
let window = ReplayWindow::new();
assert!(!window.check(u64::MAX));
}
#[test]
fn test_replay_window_accepts_highest_counter_an_honest_peer_can_send() {
// take_send_counter refuses u64::MAX, so u64::MAX - 1 is the highest
// counter a conforming peer emits. The ceiling guard must be exactly one
// value wide and leave that one alone.
let mut window = ReplayWindow::new();
assert!(window.check(u64::MAX - 1));
window.accept(u64::MAX - 1);
assert!(
!window.check(u64::MAX - 1),
"replay should still be rejected"
);
}
#[test]
fn test_session_replay_protection() {
let keypair1 = generate_keypair();
+1
View File
@@ -5,6 +5,7 @@ mod handoff;
mod offer_admission;
mod runtime;
mod signal;
mod signal_gate;
mod stun;
mod traversal;
mod traversal_machine;
+121 -19
View File
@@ -1,7 +1,7 @@
use std::collections::{HashMap, HashSet};
use std::net::SocketAddr;
use std::sync::Arc;
use std::sync::atomic::{AtomicBool, Ordering};
use std::sync::atomic::{AtomicBool, AtomicU64, Ordering};
use std::time::{Duration, Instant};
use nostr::nips::nip17;
@@ -22,10 +22,11 @@ use super::failure_state::FailureState;
use super::handoff::EstablishedTraversal;
use super::offer_admission::{AdmissionReject, OfferAdmission};
use super::signal::{
FreshnessOutcome, SignalEnvelope, build_signal_event, create_traversal_answer,
create_traversal_offer, estimate_clock_skew, unwrap_signal_event, validate_offer_freshness,
validate_traversal_answer_for_offer,
FRESHNESS_SKEW_TOLERANCE_MS, FreshnessOutcome, SignalEnvelope, build_signal_event,
create_traversal_answer, create_traversal_offer, estimate_clock_skew, unwrap_signal_event,
validate_offer_freshness, validate_traversal_answer_for_offer,
};
use super::signal_gate::SignalGate;
use super::stun::observe_traversal_addresses;
use super::traversal::{
PunchTargetTally, is_doc_ip, is_never_punchable_ip, is_private_ip, nonce, now_ms,
@@ -84,7 +85,7 @@ pub(super) fn short_id(id: &str) -> String {
/// node collects by default and those refusals are the ones an operator needs
/// to see; the routine off-subnet case stays at `debug`.
fn log_refusals(tally: &PunchTargetTally, peer: &str, session: &str) {
if tally.offered <= tally.admitted && tally.capped == 0 {
if tally.offered <= tally.admitted && tally.capped == 0 && tally.over_offered == 0 {
return;
}
let sample = tally.sample.as_deref().unwrap_or("-");
@@ -100,6 +101,7 @@ fn log_refusals(tally: &PunchTargetTally, peer: &str, session: &str) {
unroutable = tally.unroutable,
offsubnet = tally.offsubnet,
capped = tally.capped,
over_offered = tally.over_offered,
reflexive = %reflexive,
sample = %sample,
"traversal: punch candidates refused"
@@ -115,6 +117,7 @@ fn log_refusals(tally: &PunchTargetTally, peer: &str, session: &str) {
unroutable = tally.unroutable,
offsubnet = tally.offsubnet,
capped = tally.capped,
over_offered = tally.over_offered,
reflexive = %reflexive,
sample = %sample,
"traversal: punch candidates refused"
@@ -207,6 +210,9 @@ pub struct NostrRendezvous {
traversal: TraversalMachine,
pending_answers: Mutex<HashMap<String, oneshot::Sender<SignalEnvelope<TraversalAnswer>>>>,
admission: OfferAdmission,
signal_gate: SignalGate,
/// Inbound traversal signals shed before decryption, since process start.
shed_signals: AtomicU64,
event_tx: mpsc::UnboundedSender<BootstrapEvent>,
event_rx: Mutex<mpsc::UnboundedReceiver<BootstrapEvent>>,
connect_task: Mutex<Option<JoinHandle<()>>>,
@@ -346,6 +352,8 @@ impl NostrRendezvous {
traversal,
pending_answers: Mutex::new(HashMap::new()),
admission,
signal_gate: SignalGate::new(Instant::now()),
shed_signals: AtomicU64::new(0),
event_tx,
event_rx: Mutex::new(event_rx),
connect_task: Mutex::new(None),
@@ -606,21 +614,18 @@ impl NostrRendezvous {
Err(_) => return NostrRefetchOutcome::Skipped,
};
let mut newest: Option<(u64, &Event)> = None;
for ev in events.iter() {
let ts = ev.created_at.as_secs();
match newest {
Some((cur, _)) if ts <= cur => {}
_ => newest = Some((ts, ev)),
let Some(ev) = Self::newest_event_by_author(events.iter(), target_pubkey) else {
if !events.is_empty() {
// The relays answered, but nothing they returned was signed by
// this peer. That is no evidence of absence, so keep the entry.
return NostrRefetchOutcome::Skipped;
}
}
let Some((relay_created_at, ev)) = newest else {
// Absent on relays. Evict any stale cache entry.
self.advert.remove(peer_npub);
self.failure_state.reset_streak_after_refresh(peer_npub);
return NostrRefetchOutcome::Evicted;
};
let relay_created_at = Self::effective_created_at_secs(ev.created_at.as_secs(), now_ms());
match cached_created_at {
Some(cached) if relay_created_at <= cached => NostrRefetchOutcome::SameAdvert,
@@ -751,7 +756,15 @@ impl NostrRendezvous {
Self::parse_overlay_advert_event(&event, &self.config.app)
{
let endpoints = endpoint_summary(&advert.endpoints);
let created_at = event.created_at.as_secs();
// Clamped forward to the signal path's skew
// tolerance: an unbounded future `created_at` would
// win every replacement comparison in
// `observe_advert` and buy a proportionally distant
// validity horizon.
let created_at = Self::effective_created_at_secs(
event.created_at.as_secs(),
now_ms(),
);
if self.advert.observe_advert(
&author_npub,
advert,
@@ -774,6 +787,41 @@ impl NostrRendezvous {
continue;
}
// Ahead of the unwrap, which is two NIP-44 decrypts and a
// signature verify run inline on the single task that also
// routes answers and processes adverts. Nothing about the
// sender is known yet — the outer event is signed by a key
// generated per event — so the allowance is necessarily
// shared and indiscriminate, and the reserve is what keeps
// a flood from also shedding the answers to traversals
// this node started.
let awaiting_answers = match self.pending_answers.try_lock() {
Ok(pending) => !pending.is_empty(),
// Contended rather than known empty, so treat it as
// outstanding: the fail-open direction here spends the
// reserve, it does not shed.
Err(_) => true,
};
if let Err(shed) = self.signal_gate.admit(awaiting_answers, Instant::now()) {
let total = self.shed_signals.fetch_add(1, Ordering::Relaxed) + 1;
// Debug, not warn, per event: the party that trips this
// is by definition sending faster than the node wants,
// so a record per drop turns the flood into log volume.
// The doubling summary below is the operator's signal.
debug!(
reason = ?shed,
total,
"shed inbound traversal signal before decrypt"
);
if total.is_power_of_two() {
warn!(
shed = total,
"inbound traversal signals shed before decrypt"
);
}
continue;
}
let unwrapped = match unwrap_signal_event(&self.keys, &event).await {
Ok(unwrapped) => unwrapped,
Err(err) => {
@@ -1104,6 +1152,10 @@ impl NostrRendezvous {
let base_socket = std::net::UdpSocket::bind(("0.0.0.0", 0))?;
base_socket.set_nonblocking(true)?;
// This drains every datagram on the traversal socket until the STUN
// deadline, so it must complete before any punch can be in flight: a
// retry or re-observation once punching has started would swallow the
// peer's punch packets.
let (reflexive_address, local_addresses, stun_server) = observe_traversal_addresses(
&base_socket,
&self.config.stun_servers,
@@ -1373,6 +1425,10 @@ impl NostrRendezvous {
let base_socket = std::net::UdpSocket::bind(("0.0.0.0", 0))?;
base_socket.set_nonblocking(true)?;
// This drains every datagram on the traversal socket until the STUN
// deadline, so it must complete before any punch can be in flight: a
// retry or re-observation once punching has started would swallow the
// peer's punch packets.
let (reflexive_address, local_addresses, stun_server) = observe_traversal_addresses(
&base_socket,
&self.config.stun_servers,
@@ -1511,15 +1567,16 @@ impl NostrRendezvous {
if author_npub != peer_npub {
continue;
}
let created_at = Self::effective_created_at_secs(event.created_at.as_secs(), now_ms());
let replace = best
.as_ref()
.map(|current| event.created_at.as_secs() >= current.created_at)
.map(|current| created_at >= current.created_at)
.unwrap_or(true);
if replace {
best = Some(CachedOverlayAdvert {
author_npub,
advert,
created_at: event.created_at.as_secs(),
created_at,
valid_until_ms,
});
}
@@ -1589,7 +1646,7 @@ impl NostrRendezvous {
return Ok(self.config.dm_relays.clone());
}
};
let newest = events.iter().max_by_key(|event| event.created_at.as_secs());
let newest = Self::newest_event_by_author(events.iter(), target_pubkey);
if let Some(event) = newest {
let relays = nip17::extract_relay_list(event)
.map(|relay| relay.to_string())
@@ -1693,6 +1750,42 @@ impl NostrRendezvous {
self.advert.event_valid_until_ms(event, now_ms())
}
/// Newest event in `events` that was actually signed by `author`.
///
/// The relay pool verifies each event's signature but does not check a
/// reply against the REQ filter unless `verify_subscriptions` or
/// `ban_relay_on_mismatch` is set, and neither is. A relay may therefore
/// answer an author-filtered request with an event it signed itself, so
/// the author test happens here, before the timestamp contest, not after:
/// a future-dated foreign event must not be able to suppress the genuine
/// one by winning `created_at`.
pub(super) fn newest_event_by_author<'a>(
events: impl Iterator<Item = &'a Event>,
author: PublicKey,
) -> Option<&'a Event> {
events
.filter(|event| event.pubkey == author)
.max_by_key(|event| event.created_at.as_secs())
}
/// A peer's advert `created_at`, in seconds, clamped so it can never read
/// more than `FRESHNESS_SKEW_TOLERANCE_MS` ahead of `now_ms`.
///
/// An unbounded future `created_at` buys a cache entry two things it
/// should not have: a proportionally distant validity horizon, and an
/// unbeatable position in every replacement comparison, so a later genuine
/// advert can never displace it. Clamping rather than rejecting is
/// deliberate: a node whose own clock runs slow reads every peer's honest
/// advert as future-dated, and rejecting would take out Nostr-mediated
/// dialing for every peer at once with nothing but a cache miss to show
/// for it. Raising the tolerance widens the window in which a future-dated
/// advert outranks an honest one; lowering it makes an ordinary clock
/// difference look hostile.
pub(super) fn effective_created_at_secs(created_at_secs: u64, now_ms: u64) -> u64 {
let ceiling_secs = now_ms.saturating_add(FRESHNESS_SKEW_TOLERANCE_MS) / 1000;
created_at_secs.min(ceiling_secs)
}
pub(super) fn compute_advert_valid_until_ms(
event: &Event,
advert_max_age_ms: u64,
@@ -1702,7 +1795,8 @@ impl NostrRendezvous {
return None;
}
let created_ms = event.created_at.as_secs().saturating_mul(1000);
let created_ms = Self::effective_created_at_secs(event.created_at.as_secs(), now_ms)
.saturating_mul(1000);
let created_window_until = created_ms.saturating_add(advert_max_age_ms);
if created_window_until <= now_ms {
return None;
@@ -1852,6 +1946,8 @@ impl NostrRendezvous {
traversal,
pending_answers: Mutex::new(HashMap::new()),
admission,
signal_gate: SignalGate::new(Instant::now()),
shed_signals: AtomicU64::new(0),
event_tx,
event_rx: Mutex::new(event_rx),
connect_task: Mutex::new(None),
@@ -1923,6 +2019,12 @@ impl NostrRendezvous {
self.advert.insert_fetched(&npub, advert);
}
/// The cached `created_at` for `npub`, or `None` when nothing is cached.
/// Lets a unit test observe whether a refetch evicted an entry.
pub(crate) async fn cached_created_at_for_test(&self, npub: &str) -> Option<u64> {
self.advert.cached_created_at(npub)
}
/// Queue a bootstrap event directly for lifecycle tests without live relays
/// or a running traversal task.
pub(crate) fn push_event_for_test(&self, event: BootstrapEvent) {
+252
View File
@@ -0,0 +1,252 @@
//! Rate limiting for inbound traversal signals, ahead of any cryptography.
//!
//! The notify loop used to hand every kind-21059 event straight to
//! `unwrap_signal_event`, which is two NIP-44 decrypts and a signature verify,
//! on the single task that also routes answers and processes adverts. Nothing
//! bounded how fast an unauthenticated stranger could schedule that work, and
//! the per-npub offer admission cannot: it keys on the sender's public key,
//! which only exists once the first decrypt has already run.
//!
//! **What can be keyed on, and what cannot.** Before decryption there is no
//! sender identity at all. The outer event is signed by a key generated per
//! event, so bucketing on its author hands an attacker a fresh allowance for
//! free, and `created_at` and the p-tag are equally attacker-chosen. The relay
//! the event arrived over is drawn from our own configured set, but it is not
//! an isolation boundary either: an attacker publishes to the same relays the
//! honest peer does, and a duplicate event is attributed to whichever relay
//! won the delivery race. So the shared allowance here is deliberately a
//! single global bucket, and it is indiscriminate by construction.
//!
//! **What the reserve is for.** An indiscriminate limit sheds our own
//! traversals along with the attacker's, and since the attacker sets the rate,
//! every retry lands in the same shed. The second bucket is drawn only while
//! this node has traversals of its own outstanding, so a flood costs a node
//! its inbound offers, which is irreducible without a pre-decrypt identity,
//! rather than also costing it the answers to offers it sent.
//!
//! **Lock discipline.** `admit` takes `now` as a parameter rather than reading
//! the clock, so the type is testable without sleeping and holds no state that
//! has to be advanced by a timer. It does its whole decision under one
//! `std::sync::Mutex` and never awaits inside it, as `offer_admission` does;
//! moving anything awaited inside that lock would hold it across a decrypt.
use std::sync::Mutex;
use std::time::Instant;
/// Sustained inbound traversal signals per second admitted for decryption,
/// across all senders and relays.
///
/// A traversal exchange is a handful of events (one offer and one answer per
/// attempt), and the signal subscription is opened with `limit(0)` so relays
/// replay no stored backlog, which means there is no legitimate burst larger
/// than the number of peers bootstrapping in the same second. Raising this
/// buys a larger rendezvous hub headroom at the cost of handing an attacker
/// the same multiple of decrypt work; lowering it starts shedding honest
/// signals on a busy node, which shows up in the log as the shed counter
/// rather than as silence.
const SIGNAL_RATE_PER_SEC: f64 = 5.0;
/// Burst capacity of the shared allowance, in signals.
const SIGNAL_BURST: f64 = 20.0;
/// Sustained rate of the reserve, drawn only while this node has traversals
/// of its own outstanding.
///
/// It is sized like the shared allowance rather than smaller because the case
/// it exists for is onboarding fanout: a node that has just sent offers to
/// tens of peers receives their answers back in a burst, and shedding those
/// looks to an operator like relay flakiness. Lowering it re-exposes that
/// case; raising it lets a flood arriving while we happen to be mid-traversal
/// buy more decrypt work than the shared allowance alone would.
const ANSWER_RESERVE_RATE_PER_SEC: f64 = 5.0;
/// Burst capacity of the reserve, in signals.
const ANSWER_RESERVE_BURST: f64 = 20.0;
/// Which allowance refused an inbound signal.
///
/// The two are different operator stories: the first says the node shed a
/// signal while it had nothing of its own in flight, the second says a flood
/// is now deep enough to reach traversals this node started.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub(super) enum SignalShed {
/// The shared allowance is exhausted.
Shared,
/// The shared allowance and the answer reserve are both exhausted.
Reserve,
}
/// One token bucket, refilled from the caller's clock.
#[derive(Debug)]
struct Bucket {
tokens: f64,
burst: f64,
rate_per_sec: f64,
last: Instant,
}
impl Bucket {
fn new(rate_per_sec: f64, burst: f64, now: Instant) -> Self {
Self {
tokens: burst,
burst,
rate_per_sec,
last: now,
}
}
/// Advance the bucket to `now`, capped at its burst.
fn refill(&mut self, now: Instant) {
let elapsed = now.saturating_duration_since(self.last).as_secs_f64();
if elapsed > 0.0 {
self.tokens = (self.tokens + elapsed * self.rate_per_sec).min(self.burst);
self.last = now;
}
}
/// Whether a whole token is available, without taking it.
fn ready(&self) -> bool {
self.tokens >= 1.0
}
fn take(&mut self) {
self.tokens -= 1.0;
}
}
/// The pre-decrypt allowance for inbound traversal signals.
pub(super) struct SignalGate {
inner: Mutex<Inner>,
}
#[derive(Debug)]
struct Inner {
shared: Bucket,
reserve: Bucket,
}
impl SignalGate {
/// Build a gate whose buckets start full as of `now`.
pub(super) fn new(now: Instant) -> Self {
Self::with_limits(
SIGNAL_RATE_PER_SEC,
SIGNAL_BURST,
ANSWER_RESERVE_RATE_PER_SEC,
ANSWER_RESERVE_BURST,
now,
)
}
fn with_limits(
shared_rate: f64,
shared_burst: f64,
reserve_rate: f64,
reserve_burst: f64,
now: Instant,
) -> Self {
Self {
inner: Mutex::new(Inner {
shared: Bucket::new(shared_rate, shared_burst, now),
reserve: Bucket::new(reserve_rate, reserve_burst, now),
}),
}
}
/// Take one signal's worth of allowance, or say which bucket refused it.
///
/// `awaiting_answers` says whether this node has traversals of its own
/// outstanding; only then may the reserve be drawn. The shared bucket is
/// always tried first, so the reserve is spent only on what the flood
/// would otherwise have shed.
pub(super) fn admit(&self, awaiting_answers: bool, now: Instant) -> Result<(), SignalShed> {
let mut inner = self.inner.lock().expect("signal-gate mutex poisoned");
inner.shared.refill(now);
if inner.shared.ready() {
inner.shared.take();
return Ok(());
}
if !awaiting_answers {
return Err(SignalShed::Shared);
}
inner.reserve.refill(now);
if inner.reserve.ready() {
inner.reserve.take();
return Ok(());
}
Err(SignalShed::Reserve)
}
}
#[cfg(test)]
mod tests {
use super::*;
use std::time::Duration;
#[test]
fn a_flood_is_admitted_up_to_the_burst_and_shed_after_it() {
let start = Instant::now();
let gate = SignalGate::new(start);
let admitted = (0..1000)
.filter(|_| gate.admit(false, start).is_ok())
.count();
assert_eq!(admitted, SIGNAL_BURST as usize);
assert_eq!(gate.admit(false, start), Err(SignalShed::Shared));
}
#[test]
fn tokens_refill_so_a_steady_legitimate_signal_rate_is_never_shed() {
let start = Instant::now();
let gate = SignalGate::new(start);
// One second per iteration at exactly the sustained rate, for long
// enough that an arithmetic error in the refill drains the bucket.
for second in 0..100u64 {
let now = start + Duration::from_secs(second);
for signal in 0..SIGNAL_RATE_PER_SEC as usize {
assert_eq!(
gate.admit(false, now),
Ok(()),
"signal {signal} of second {second} should be admitted"
);
}
}
}
#[test]
fn a_flood_cannot_shed_the_answers_to_traversals_this_node_started() {
let start = Instant::now();
let gate = SignalGate::new(start);
// The flood arrives while this node has nothing outstanding, so it
// cannot reach the reserve at all.
for _ in 0..1000 {
let _ = gate.admit(false, start);
}
assert_eq!(gate.admit(false, start), Err(SignalShed::Shared));
let admitted = (0..1000)
.filter(|_| gate.admit(true, start).is_ok())
.count();
assert_eq!(admitted, ANSWER_RESERVE_BURST as usize);
assert_eq!(gate.admit(true, start), Err(SignalShed::Reserve));
}
#[test]
fn the_reserve_is_untouched_while_the_shared_allowance_still_has_tokens() {
let start = Instant::now();
let gate = SignalGate::new(start);
// Every one of these is inside the shared burst, so none of them may
// spend the reserve even though the caller is entitled to it.
for _ in 0..SIGNAL_BURST as usize {
assert_eq!(gate.admit(true, start), Ok(()));
}
let reserve_admitted = (0..1000)
.filter(|_| gate.admit(true, start).is_ok())
.count();
assert_eq!(reserve_admitted, ANSWER_RESERVE_BURST as usize);
}
}
+124 -1
View File
@@ -80,6 +80,20 @@ pub(super) async fn observe_traversal_addresses(
Ok((None, local_addresses, None))
}
/// Send a STUN binding request on `socket` and return the reflexive address
/// the server reports.
///
/// Only a datagram whose source is exactly the resolved server address is
/// parsed; anything else is consumed and discarded, so an on-path attacker
/// who has seen the transaction id cannot substitute a mapped address by
/// injecting a reply. The comparison is exact because every socket handed to
/// this function is bound to the IPv4 wildcard, so a source can never arrive
/// in v4-mapped IPv6 form. A dual-stack bind would require normalising both
/// sides with `IpAddr::to_canonical()` before comparing.
///
/// Caller requirement: this drains and discards every datagram arriving on
/// `socket` until the deadline, so it must not be entered on a socket that
/// may concurrently carry other traffic the caller cares about.
async fn perform_stun(
socket: &std::net::UdpSocket,
stun_server: &str,
@@ -95,15 +109,32 @@ async fn perform_stun(
udp.send_to(&request, addr).await?;
let mut buf = [0u8; 2048];
let deadline = tokio::time::Instant::now() + response_timeout;
let mut rejected = 0u64;
let mut last_unexpected = None;
loop {
let result = tokio::time::timeout_at(deadline, udp.recv_from(&mut buf)).await;
let Ok(Ok((len, _remote))) = result else {
let Ok(Ok((len, remote))) = result else {
break;
};
if remote != addr {
rejected += 1;
last_unexpected = Some(remote);
continue;
}
if let Some(mapped) = parse_stun_binding_success(&buf[..len], &txn_id) {
return Ok(Some(mapped));
}
}
// One line per call rather than per datagram: a flooder controls the rate.
if rejected > 0 {
debug!(
stun_server = %stun_server,
expected = %addr,
rejected,
last_unexpected = ?last_unexpected,
"discarded STUN datagrams from unexpected sources"
);
}
Err(BootstrapError::Stun(format!(
"timed out waiting for {}",
stun_server
@@ -521,4 +552,96 @@ mod tests {
"2001:db8::1".parse::<Ipv6Addr>().unwrap()
)));
}
/// Build a complete Binding Success carrying one XOR-MAPPED-ADDRESS.
fn build_binding_success(mapped: std::net::SocketAddrV4, txn_id: &[u8; 12]) -> Vec<u8> {
let mut packet = build_success_header(0, txn_id);
let cookie = STUN_MAGIC_COOKIE.to_be_bytes();
let octets = mapped.ip().octets();
let xport = mapped.port() ^ ((STUN_MAGIC_COOKIE >> 16) as u16);
packet.extend_from_slice(&0x0020u16.to_be_bytes()); // XOR-MAPPED-ADDRESS
packet.extend_from_slice(&8u16.to_be_bytes());
packet.push(0x00); // reserved
packet.push(0x01); // family IPv4
packet.extend_from_slice(&xport.to_be_bytes());
for (index, octet) in octets.iter().enumerate() {
packet.push(octet ^ cookie[index]);
}
let body_len = (packet.len() - 20) as u16;
packet[2..4].copy_from_slice(&body_len.to_be_bytes());
packet
}
/// Bind a loopback socket suitable for handing to `perform_stun`.
///
/// `set_nonblocking` is mandatory rather than tidiness: `perform_stun`
/// passes the socket to `tokio::net::UdpSocket::from_std`, which requires
/// a non-blocking socket and does not make one. A blocking socket parks
/// the runtime thread and the deadline never fires.
fn bind_stun_caller() -> std::net::UdpSocket {
let socket = std::net::UdpSocket::bind("127.0.0.1:0").unwrap();
socket.set_nonblocking(true).unwrap();
socket
}
#[tokio::test]
async fn stun_binding_response_from_an_unexpected_source_is_refused() {
let caller = bind_stun_caller();
let server = std::net::UdpSocket::bind("127.0.0.1:0").unwrap();
let attacker = std::net::UdpSocket::bind("127.0.0.1:0").unwrap();
let server_addr = server.local_addr().unwrap();
// Stands in for an on-path attacker: it learns the transaction id the
// way a real one would, by reading the request, and answers from its
// own address while the server stays silent.
let forger = std::thread::spawn(move || {
let mut buf = [0u8; 2048];
let (len, from) = server.recv_from(&mut buf).unwrap();
assert!(len >= 20);
let mut txn_id = [0u8; 12];
txn_id.copy_from_slice(&buf[8..20]);
let forged = build_binding_success("203.0.113.7:1".parse().unwrap(), &txn_id);
attacker.send_to(&forged, from).unwrap();
});
let result = super::perform_stun(
&caller,
&server_addr.to_string(),
std::time::Duration::from_millis(300),
)
.await;
forger.join().unwrap();
assert!(
result.is_err(),
"a binding success from a host other than the server must not be believed, got {:?}",
result
);
}
#[tokio::test]
async fn stun_binding_response_from_the_server_is_accepted() {
let caller = bind_stun_caller();
let server = std::net::UdpSocket::bind("127.0.0.1:0").unwrap();
let server_addr = server.local_addr().unwrap();
let responder = std::thread::spawn(move || {
let mut buf = [0u8; 2048];
let (_len, from) = server.recv_from(&mut buf).unwrap();
let mut txn_id = [0u8; 12];
txn_id.copy_from_slice(&buf[8..20]);
let reply = build_binding_success("198.51.100.9:4242".parse().unwrap(), &txn_id);
server.send_to(&reply, from).unwrap();
});
let mapped = super::perform_stun(
&caller,
&server_addr.to_string(),
std::time::Duration::from_secs(2),
)
.await
.unwrap()
.unwrap();
responder.join().unwrap();
assert_eq!(mapped.to_string(), "198.51.100.9:4242");
}
}
+441 -5
View File
@@ -1,6 +1,8 @@
use std::collections::HashSet;
use std::net::{IpAddr, SocketAddr};
use std::time::Duration;
use nostr::nips::nip17;
use nostr::prelude::{EventBuilder, Kind, RelayUrl, Tag, Timestamp};
use super::runtime::{
@@ -13,8 +15,9 @@ use super::signal::{
};
use super::stun::{parse_stun_binding_success, parse_stun_url};
use super::traversal::{
PunchStrategy, build_punch_packet, is_doc_ip, is_never_punchable_ip, is_private_ip, now_ms,
parse_punch_packet, plan_punch_targets, planned_remote_endpoints, session_hash,
PunchStrategy, SourceRank, build_punch_packet, is_doc_ip, is_never_punchable_ip, is_private_ip,
now_ms, parse_punch_packet, plan_punch_targets, planned_remote_endpoints, rank_punch_source,
run_punch_attempt, session_hash,
};
use super::traversal_machine::suppress_responder_for_own_initiator;
use super::types::BootstrapError;
@@ -47,14 +50,29 @@ fn can_reach(local_nat: NatType, remote_nat: NatType) -> bool {
}
fn signed_overlay_advert_event(created_at_secs: u64, expiration_secs: Option<u64>) -> nostr::Event {
let keys = nostr::Keys::generate();
signed_overlay_advert_event_from(&nostr::Keys::generate(), created_at_secs, expiration_secs)
}
fn signed_overlay_advert_event_from(
keys: &nostr::Keys,
created_at_secs: u64,
expiration_secs: Option<u64>,
) -> nostr::Event {
let content = r#"{"identifier":"fips-overlay-v1","version":1,"endpoints":[{"transport":"tcp","addr":"8.8.8.8:443"}]}"#;
let mut builder = EventBuilder::new(Kind::Custom(ADVERT_KIND), content)
.custom_created_at(Timestamp::from(created_at_secs));
if let Some(expiration_secs) = expiration_secs {
builder = builder.tags([Tag::expiration(Timestamp::from(expiration_secs))]);
}
builder.sign_with_keys(&keys).unwrap()
builder.sign_with_keys(keys).unwrap()
}
fn signed_inbox_relay_event(keys: &nostr::Keys, created_at_secs: u64, relay: &str) -> nostr::Event {
EventBuilder::new(Kind::InboxRelays, "")
.tags([Tag::relay(RelayUrl::parse(relay).unwrap())])
.custom_created_at(Timestamp::from(created_at_secs))
.sign_with_keys(keys)
.unwrap()
}
#[test]
@@ -197,6 +215,102 @@ fn advert_freshness_rejects_stale_created_at_without_expiration() {
assert!(valid_until.is_none());
}
/// A hostile advert relay may answer an author-filtered request with an event
/// it signed itself. Selection has to drop those before the newest-`created_at`
/// contest, or a future-dated foreign advert suppresses the genuine one.
#[test]
fn advert_selection_ignores_events_not_signed_by_the_target_peer() {
let now_secs = Timestamp::now().as_secs();
let peer_keys = nostr::Keys::generate();
let hostile_keys = nostr::Keys::generate();
let hostile = signed_overlay_advert_event_from(&hostile_keys, now_secs + 3_600, None);
let genuine = signed_overlay_advert_event_from(&peer_keys, now_secs.saturating_sub(10), None);
let events = [hostile, genuine];
let selected = NostrRendezvous::newest_event_by_author(events.iter(), peer_keys.public_key())
.expect("the peer's own advert should be selected");
assert_eq!(selected.pubkey, peer_keys.public_key());
}
/// Nothing signed by the peer means nothing to select, even though the relays
/// did answer. The caller reads this as "no evidence", not "withdrawn".
#[test]
fn advert_selection_returns_nothing_when_every_event_is_foreign() {
let now_secs = Timestamp::now().as_secs();
let peer_keys = nostr::Keys::generate();
let hostile_keys = nostr::Keys::generate();
let events = [signed_overlay_advert_event_from(
&hostile_keys,
now_secs + 3_600,
None,
)];
assert!(
NostrRendezvous::newest_event_by_author(events.iter(), peer_keys.public_key()).is_none()
);
}
/// The same omission on the inbox-relay lookup steers this node's DM and
/// signal traffic onto relays an attacker chose, so it gets the same filter.
#[test]
fn inbox_relay_selection_ignores_relay_lists_not_signed_by_the_target() {
let now_secs = Timestamp::now().as_secs();
let peer_keys = nostr::Keys::generate();
let hostile_keys = nostr::Keys::generate();
let events = [
signed_inbox_relay_event(&hostile_keys, now_secs + 3_600, "wss://hostile.example/"),
signed_inbox_relay_event(
&peer_keys,
now_secs.saturating_sub(10),
"wss://genuine.example/",
),
];
let selected = NostrRendezvous::newest_event_by_author(events.iter(), peer_keys.public_key())
.expect("the peer's own relay list should be selected");
let relays = nip17::extract_relay_list(selected)
.map(|relay| relay.to_string())
.collect::<Vec<_>>();
assert_eq!(relays, vec!["wss://genuine.example/".to_string()]);
}
/// A far-future `created_at` must not buy a proportionally distant validity
/// horizon. The window is computed from the clamped timestamp instead, so the
/// entry expires on our clock rather than the publisher's.
#[test]
fn advert_freshness_clamps_created_at_beyond_the_forward_skew_tolerance() {
let now_secs = Timestamp::now().as_secs();
let event = signed_overlay_advert_event(now_secs + 3_600, None);
let valid_until =
NostrRendezvous::compute_advert_valid_until_ms(&event, 600_000, now_secs * 1000)
.expect("a future-dated advert is still usable, just not for as long");
assert_eq!(valid_until, (now_secs + 60) * 1000 + 600_000);
}
/// Pins the forward bound to `FRESHNESS_SKEW_TOLERANCE_MS` exactly, mirroring
/// the signal path: 60s ahead is taken as published, 61s ahead is clamped.
/// This is the healthy-path half; an ordinary clock difference must not cost a
/// legitimate peer anything.
#[test]
fn advert_freshness_at_the_forward_skew_limit_is_untouched_and_one_second_beyond_is_clamped() {
let now_secs = Timestamp::now().as_secs();
let at_limit = signed_overlay_advert_event(now_secs + 60, None);
let valid_until =
NostrRendezvous::compute_advert_valid_until_ms(&at_limit, 600_000, now_secs * 1000)
.expect("an advert exactly at the forward tolerance should be accepted as published");
assert_eq!(valid_until, (now_secs + 60) * 1000 + 600_000);
let past_limit = signed_overlay_advert_event(now_secs + 61, None);
let clamped =
NostrRendezvous::compute_advert_valid_until_ms(&past_limit, 600_000, now_secs * 1000)
.expect("an advert one second past the tolerance is clamped, not refused");
assert_eq!(clamped, (now_secs + 60) * 1000 + 600_000);
}
#[test]
fn advert_freshness_uses_earliest_expiration_bound() {
let now_secs = Timestamp::now().as_secs();
@@ -588,7 +702,11 @@ fn planned_remote_endpoints_bound_an_oversized_list_of_unroutable_candidates() {
endpoints,
vec!["198.51.100.20:63000".parse::<SocketAddr>().unwrap()]
);
assert_eq!(tally.unroutable, 300);
// Only the first MAX_OFFERED_CANDIDATES are vetted at all now, so the
// unroutable count is the bound rather than the whole list; the rest are
// recorded as never having been looked at.
assert_eq!(tally.unroutable, 32);
assert_eq!(tally.over_offered, 268);
assert_eq!(tally.admitted, 1);
assert!(tally.suspicious());
}
@@ -597,6 +715,10 @@ fn planned_remote_endpoints_bound_an_oversized_list_of_unroutable_candidates() {
/// so the observed reflexive address is itself private. Applying the /24 gate
/// to a peer's reflexive address would drop it and remove the only branch
/// that works across arbitrary NATs; this test reds if anyone does that.
///
/// The exemption is conditional on exactly the vantage point this test sets
/// up: our own reflexive address is private here, so it still applies. The
/// two tests below cover the public and absent cases.
#[test]
fn planned_remote_endpoints_keep_private_reflexive_when_stun_is_on_the_lan() {
let (endpoints, _tally) = planned_remote_endpoints(
@@ -610,6 +732,88 @@ fn planned_remote_endpoints_keep_private_reflexive_when_stun_is_on_the_lan() {
assert!(endpoints.contains(&"192.168.1.20:63000".parse().unwrap()));
}
/// A node whose own STUN result is public shares no LAN with a private
/// address, so a peer's private reflexive address is only ever an address of
/// the peer's choosing. Admitting it made the reflexive branch a way to have
/// this node punch inside its own private network; the /24 gate now applies.
#[test]
fn a_peers_private_reflexive_address_is_refused_when_our_own_stun_result_is_public() {
let (endpoints, tally) = planned_remote_endpoints(
&[],
Some(&addr("203.0.113.10", 62000)),
&[],
Some(&addr("192.168.1.20", 63000)),
)
.expect("endpoint planning should succeed");
assert!(endpoints.is_empty());
assert_eq!(tally.reflexive, Some("off-subnet"));
}
/// The conditional gate keys on our own reflexive address being private, and
/// a node with no reflexive address at all has to keep behaving as it did:
/// a failed STUN probe must not cost same-LAN peering.
#[test]
fn a_peers_private_reflexive_address_is_kept_when_we_have_no_stun_result_at_all() {
let (endpoints, tally) = planned_remote_endpoints(
&[addr("192.168.1.10", 62000)],
None,
&[],
Some(&addr("192.168.1.20", 63000)),
)
.expect("endpoint planning should succeed");
assert!(endpoints.contains(&"192.168.1.20:63000".parse().unwrap()));
assert_eq!(tally.reflexive, None);
}
/// Refusing a peer's private reflexive address is now something an honest
/// asymmetric-STUN deployment produces, so it must not warn on its own. The
/// never-routable case above still does.
#[test]
fn an_off_subnet_reflexive_refusal_alone_is_not_suspicious() {
let (_planned, tally) = plan_punch_targets(
&[],
Some(&addr("203.0.113.10", 62000)),
&[addr("203.0.113.5", 63000)],
Some(&addr("192.168.1.20", 63000)),
);
assert_eq!(tally.reflexive, Some("off-subnet"));
assert!(
tally.admitted > 0,
"the host-candidate path should still plan"
);
assert!(!tally.suspicious());
}
/// The eight-target cap runs after both planning loops, so it bounds the
/// output and not the work. The discriminating assertion is `unroutable`:
/// vetting every candidate would count all thousand, so a count of exactly
/// `MAX_OFFERED_CANDIDATES` is what proves the excess was never walked.
#[test]
fn an_oversized_candidate_list_is_bounded_before_vetting() {
let mut remotes = Vec::new();
for index in 0..1000u32 {
remotes.push(addr(
&format!("127.0.0.{}", 1 + (index % 254)),
63000 + (index % 1000) as u16,
));
}
let (_planned, tally) = plan_punch_targets(
&[],
Some(&addr("203.0.113.10", 62000)),
&remotes,
Some(&addr("198.51.100.20", 63000)),
);
assert_eq!(tally.offered, 1001);
assert_eq!(tally.unroutable, 32);
assert_eq!(tally.over_offered, 968);
assert!(tally.suspicious());
}
/// The four refusal classes tell four different operational stories, so a
/// change that collapses them into one counter, or that makes the warning
/// fire on the benign dual-homed shape, has to red here.
@@ -1172,6 +1376,238 @@ async fn signal_events_use_current_timestamps() {
assert!(created_at <= after);
}
/// These punch tests distinguish a spoofer from a planned target by source
/// **IP**, so each needs its own loopback address. Only Linux treats the whole
/// of 127/8 as local; macOS and Windows bind 127.0.0.1 alone unless an alias is
/// added, so the bind panics there. They are gated to Linux rather than
/// rewritten onto one address, because collapsing them onto 127.0.0.1 would
/// make every source rank `RemappedPort` and the tests would stop testing what
/// they are for.
///
/// **Coverage gap**: on macOS and Windows nothing exercises `run_punch_attempt`
/// end to end. The ranking decision itself is covered on every platform by the
/// `rank_punch_source_*` unit tests above, which take no sockets.
/// A loopback socket bound on `host`, non-blocking as both production call
/// sites leave it, since `run_punch_attempt` hands it straight to
/// `UdpSocket::from_std`.
#[cfg(target_os = "linux")]
fn punch_socket(host: &str) -> std::net::UdpSocket {
let socket = std::net::UdpSocket::bind(format!("{host}:0")).expect("bind a loopback socket");
socket
.set_nonblocking(true)
.expect("the punch socket must be non-blocking");
socket
}
/// A hint that starts punching immediately. `start_at_ms` is absolute wall
/// clock, so anything plausible-looking in the future would sleep out the test.
fn immediate_punch_hint(duration_ms: u64) -> PunchHint {
PunchHint {
start_at_ms: 0,
interval_ms: 20,
duration_ms,
}
}
/// Send one well-formed probe carrying `session_id`'s hash from `from` to
/// `to`, which is what a replay of captured punch bytes looks like.
#[cfg(target_os = "linux")]
fn send_probe(from: &std::net::UdpSocket, to: SocketAddr, session_id: &str) {
let packet = build_punch_packet(PunchPacketKind::Probe, 1, session_id);
from.send_to(&packet, to).expect("probe should send");
}
/// Whether anything readable on `socket` is a punch ack.
#[cfg(target_os = "linux")]
fn received_an_ack(socket: &std::net::UdpSocket) -> bool {
let mut buf = [0u8; 2048];
while let Ok((len, _)) = socket.recv_from(&mut buf) {
if parse_punch_packet(&buf[..len])
.map(|packet| packet.kind == PunchPacketKind::Ack)
.unwrap_or(false)
{
return true;
}
}
false
}
#[test]
fn rank_punch_source_accepts_a_planned_target() {
let target: SocketAddr = "198.51.100.20:63000".parse().unwrap();
assert_eq!(rank_punch_source(target, &[target]), SourceRank::Planned);
}
#[test]
fn rank_punch_source_reports_a_planned_targets_other_port_as_remapped() {
let target: SocketAddr = "198.51.100.20:63000".parse().unwrap();
let remapped: SocketAddr = "198.51.100.20:41234".parse().unwrap();
assert_eq!(
rank_punch_source(remapped, &[target]),
SourceRank::RemappedPort
);
}
#[test]
fn rank_punch_source_rejects_an_address_we_never_planned_to_probe() {
let target: SocketAddr = "198.51.100.20:63000".parse().unwrap();
let stranger: SocketAddr = "203.0.113.9:63000".parse().unwrap();
assert_eq!(
rank_punch_source(stranger, &[target]),
SourceRank::Unplanned
);
}
/// The regression test for the defect. The punch packet's discriminator is a
/// digest of a value both peers already know and it travels in the clear in
/// every probe, so anyone who has seen one can replay it. Acceptance is now
/// constrained to the targets this node planned; the spoofer is neither
/// adopted nor acked, and an ack would be a reflection we control.
#[cfg(target_os = "linux")]
#[tokio::test]
async fn a_matching_punch_packet_from_an_unplanned_source_is_neither_adopted_nor_acked() {
let victim = punch_socket("127.0.0.3");
let peer = punch_socket("127.0.0.1");
let spoofer = punch_socket("127.0.0.2");
let victim_addr = victim.local_addr().expect("victim address");
let targets = vec![peer.local_addr().expect("peer address")];
send_probe(&spoofer, victim_addr, "session-unplanned");
let result = run_punch_attempt(
&victim,
"session-unplanned",
&targets,
immediate_punch_hint(400),
Duration::from_millis(700),
)
.await;
assert!(
matches!(result, Err(BootstrapError::PunchTimeout(_))),
"a spoofed source must not be adopted, got {result:?}"
);
assert!(
!received_an_ack(&spoofer),
"an unplanned source must not be acked"
);
}
/// The spoofer wins the race on arrival order and still loses on address.
#[cfg(target_os = "linux")]
#[tokio::test]
async fn a_planned_source_is_adopted_even_when_a_spoofer_replies_first() {
let victim = punch_socket("127.0.0.3");
let peer = punch_socket("127.0.0.1");
let spoofer = punch_socket("127.0.0.2");
let victim_addr = victim.local_addr().expect("victim address");
let peer_addr = peer.local_addr().expect("peer address");
send_probe(&spoofer, victim_addr, "session-race");
send_probe(&peer, victim_addr, "session-race");
let result = run_punch_attempt(
&victim,
"session-race",
&[peer_addr],
immediate_punch_hint(400),
Duration::from_millis(700),
)
.await;
assert_eq!(
result.expect("the planned peer should be adopted"),
peer_addr
);
}
/// The healthy path, which is the check that the source constraint does not
/// red a legitimately clean run: one probe from the single planned target is
/// adopted immediately and acked.
#[cfg(target_os = "linux")]
#[tokio::test]
async fn the_ordinary_probe_from_a_planned_target_is_still_adopted_and_acked() {
let victim = punch_socket("127.0.0.3");
let peer = punch_socket("127.0.0.1");
let victim_addr = victim.local_addr().expect("victim address");
let peer_addr = peer.local_addr().expect("peer address");
send_probe(&peer, victim_addr, "session-healthy");
let result = run_punch_attempt(
&victim,
"session-healthy",
&[peer_addr],
immediate_punch_hint(400),
Duration::from_millis(700),
)
.await;
assert_eq!(
result.expect("the planned peer should be adopted"),
peer_addr
);
assert!(received_an_ack(&peer), "a planned probe should be acked");
}
/// Peer-reflexive discovery: a symmetric NAT allocates a fresh port toward us,
/// so the peer's probe arrives from an address that is not in the plan but
/// shares a planned target's IP. Adopting it is the main class of NAT pairing
/// punching exists to rescue, and this test reds if the rule is ever tightened
/// to exact matching without that being reopened deliberately.
#[cfg(target_os = "linux")]
#[tokio::test]
async fn a_planned_targets_remapped_port_is_adopted_when_that_is_all_that_arrives() {
let victim = punch_socket("127.0.0.3");
let peer = punch_socket("127.0.0.1");
let victim_addr = victim.local_addr().expect("victim address");
let peer_addr = peer.local_addr().expect("peer address");
// The address the peer's own STUN observation named, before its NAT
// remapped the port: same host, a port nothing is bound to.
let stale_target = SocketAddr::new(peer_addr.ip(), peer_addr.port().wrapping_add(1).max(1));
send_probe(&peer, victim_addr, "session-remapped");
let result = run_punch_attempt(
&victim,
"session-remapped",
&[stale_target],
immediate_punch_hint(400),
Duration::from_millis(2000),
)
.await;
assert_eq!(
result.expect("a remapped port on a planned target should be adopted"),
peer_addr
);
}
/// An exact match inside the settle window supersedes a remapped one that
/// arrived first, which is what the window is for.
#[cfg(target_os = "linux")]
#[tokio::test]
async fn an_exact_target_supersedes_a_remapped_port_inside_the_settle_window() {
let victim = punch_socket("127.0.0.3");
let peer = punch_socket("127.0.0.1");
let neighbour = punch_socket("127.0.0.1");
let victim_addr = victim.local_addr().expect("victim address");
let peer_addr = peer.local_addr().expect("peer address");
send_probe(&neighbour, victim_addr, "session-settle");
send_probe(&peer, victim_addr, "session-settle");
let result = run_punch_attempt(
&victim,
"session-settle",
&[peer_addr],
immediate_punch_hint(400),
Duration::from_millis(2000),
)
.await;
assert_eq!(
result.expect("the exact target should win"),
peer_addr,
"an exact match must supersede a source that only shares the IP"
);
}
fn node_addr(first_byte: u8) -> NodeAddr {
let mut bytes = [0u8; 16];
bytes[0] = first_byte;
+164 -16
View File
@@ -3,6 +3,7 @@ use std::sync::Arc;
use std::time::{Duration, Instant, SystemTime, UNIX_EPOCH};
use tokio::net::UdpSocket;
use tracing::debug;
use super::types::{
BootstrapError, PUNCH_ACK_MAGIC, PUNCH_MAGIC, PunchHint, PunchPacket, PunchPacketKind,
@@ -32,6 +33,61 @@ pub(super) enum PunchStrategy {
/// 400 packets, about 21 KB on the wire at 52 bytes each for IPv4.
const MAX_PUNCH_TARGETS: usize = 8;
/// Upper bound on how many candidates one peer's signal may have vetted.
///
/// Vetting is linear in this number and the `push_unique` scan that follows
/// is quadratic in the plan it feeds, so an unbounded candidate list lets one
/// signal buy an unbounded amount of our planning work regardless of the
/// eight-target cap, which only applies after both loops have run. Thirty-two
/// is four times `MAX_PUNCH_TARGETS` and four times what the candidate
/// generator produces on the widest host we have seen, so an honest peer
/// never reaches it. Raising it costs planning work per admitted signal;
/// lowering it costs an honest many-homed peer the tail of its candidate
/// list, which the tally records either way.
const MAX_OFFERED_CANDIDATES: usize = 32;
/// How long the punch loop keeps listening for an exact target match once it
/// has already accepted a planned target's address on a different port.
///
/// A source that matches a planned target's IP but not its port is what a
/// symmetric NAT's fresh mapping toward us looks like, and it is worth
/// adopting; a source that matches a target exactly is worth more, so the
/// first remapped source does not end the attempt outright. Raising this
/// delays adoption on the remapped path only, never past the attempt timeout;
/// lowering it toward zero makes the first remapped source win.
const PUNCH_SETTLE_MS: u64 = 250;
/// How much the source address of a punch packet is worth as a peer address.
///
/// The packet's own discriminator is a plain digest of a value both peers
/// already know, so it proves only that the sender has seen a probe. What the
/// source address is checked against is the target list this node planned,
/// which is the difference between adopting a peer we chose to probe and
/// adopting whoever replayed those bytes first.
#[derive(Debug, Clone, Copy, PartialEq, Eq, PartialOrd, Ord)]
pub(super) enum SourceRank {
/// Not an address we planned to probe, and so not adoptable.
Unplanned,
/// A planned target's address on a different port.
RemappedPort,
/// Exactly a target we planned to probe.
Planned,
}
/// Rank one punch packet's source address against the targets we planned.
///
/// `targets` holds at most `MAX_PUNCH_TARGETS` entries, so the scan is
/// bounded by construction.
pub(super) fn rank_punch_source(remote: SocketAddr, targets: &[SocketAddr]) -> SourceRank {
if targets.contains(&remote) {
SourceRank::Planned
} else if targets.iter().any(|target| target.ip() == remote.ip()) {
SourceRank::RemappedPort
} else {
SourceRank::Unplanned
}
}
#[derive(Debug, Clone, PartialEq, Eq)]
pub(super) struct PlannedPunchTarget {
pub(super) strategy: PunchStrategy,
@@ -44,6 +100,12 @@ pub(super) struct PlannedPunchTarget {
pub(super) remote_ip: IpAddr,
}
/// Whether a candidate's address text parses as a private or unique-local
/// address.
fn is_private_address(candidate: &TraversalAddress) -> bool {
candidate.ip.parse::<IpAddr>().is_ok_and(is_private_ip)
}
fn same_subnet_24(left: &TraversalAddress, right: &TraversalAddress) -> bool {
let left_parts = left.ip.split('.').collect::<Vec<_>>();
let right_parts = right.ip.split('.').collect::<Vec<_>>();
@@ -151,6 +213,8 @@ pub(super) struct PunchTargetTally {
pub(super) offsubnet: usize,
/// Planned targets discarded by the target cap.
pub(super) capped: usize,
/// Candidates past `MAX_OFFERED_CANDIDATES` that were never vetted.
pub(super) over_offered: usize,
/// The class label that refused the peer's reflexive address, if it was
/// refused. Held apart from the candidate counts because losing the
/// reflexive branch removes every path that works across arbitrary NATs,
@@ -170,12 +234,21 @@ impl PunchTargetTally {
/// entirely refused offer is the reflector case itself. An off-subnet-only
/// refusal is the ordinary dual-homed shape and is not suspicious.
///
/// A refused reflexive address always counts: the /24 gate does not apply
/// to it, so the only ways it can be refused are the attacker-shaped ones.
/// A refused reflexive address counts unless the class is `OffSubnet`.
/// The /24 gate now applies to a peer's reflexive address whenever our own
/// reflexive address is public, so an off-subnet refusal of it is what an
/// honest peer behind a LAN STUN server produces against a node with a
/// public one. The other three classes still have no honest producer.
///
/// A candidate list longer than `MAX_OFFERED_CANDIDATES` counts too: the
/// generator tops out near eight, so nothing honest reaches the bound.
pub(super) fn suspicious(&self) -> bool {
self.unroutable + self.zeroport + self.unparsable + self.capped > 0
|| self.over_offered > 0
|| (self.offered > 0 && self.admitted == 0)
|| self.reflexive.is_some()
|| self
.reflexive
.is_some_and(|label| label != RejectClass::OffSubnet.label())
}
/// Record one refused candidate against its class, keeping the first
@@ -212,10 +285,13 @@ impl PunchTargetTally {
///
/// Returns the parsed address, or the class of the check that refused it.
/// `lan_refs` are our own addresses that a private candidate must share a /24
/// with. `apply_private_gate` is false for the peer's reflexive address: a
/// STUN server inside the private network legitimately reports a private
/// reflexive address, and dropping it would remove the only branch that works
/// across arbitrary NATs.
/// with. `apply_private_gate` is conditionally false for the peer's reflexive
/// address: a STUN server inside the private network legitimately reports a
/// private reflexive address, and dropping it would remove the only branch
/// that works across arbitrary NATs. That exemption applies only when our own
/// reflexive address is itself private, or absent; a node whose own STUN
/// result is public has no LAN in common with a private reflexive address and
/// would only be punching an address of the peer's choosing.
fn admit_remote(
candidate: &TraversalAddress,
lan_refs: &[TraversalAddress],
@@ -261,28 +337,42 @@ pub(super) fn plan_punch_targets(
..PunchTargetTally::default()
};
// Whether our own vantage point is a LAN one: either STUN reported a
// private address for us, or it reported nothing at all. The second case
// is deliberately treated as a LAN vantage point rather than a public one,
// so a node whose STUN probe failed, or that runs without STUN, keeps
// admitting a same-LAN peer's private reflexive address as it always has.
let local_reflexive_on_lan = local_reflexive_address.is_none_or(is_private_address);
// Our own addresses a peer's private candidate has to share a /24 with.
// The local reflexive address joins the set when it is itself private,
// which is what keeps a LAN-STUN deployment able to match while the
// shipped `share_local_candidates=false` leaves the local list empty.
let mut lan_refs = local_addresses.to_vec();
if let Some(reflexive) = local_reflexive_address
&& reflexive.ip.parse::<IpAddr>().is_ok_and(is_private_ip)
&& is_private_address(reflexive)
{
lan_refs.push(reflexive.clone());
}
// A peer names its own candidate list, so bound it before anything walks
// it. The excess is recorded and discarded rather than failing the whole
// offer, which would cost an honest many-homed peer its traversal.
let considered = &remote_addresses[..remote_addresses.len().min(MAX_OFFERED_CANDIDATES)];
tally.over_offered = remote_addresses.len() - considered.len();
// Everything on the remote side is peer-supplied, so it is vetted once
// here and the branches below only ever see admitted candidates.
let remote_reflexive =
remote_reflexive_address.and_then(|remote| match admit_remote(remote, &lan_refs, false) {
let remote_reflexive = remote_reflexive_address.and_then(|remote| {
match admit_remote(remote, &lan_refs, !local_reflexive_on_lan) {
Ok(ip) => Some((remote, ip)),
Err(class) => {
tally.refuse_reflexive(class, remote);
None
}
});
let remote_candidates = remote_addresses
}
});
let remote_candidates = considered
.iter()
.filter_map(|remote| match admit_remote(remote, &lan_refs, true) {
Ok(ip) => Some((remote, ip)),
@@ -393,6 +483,27 @@ pub(super) fn planned_remote_endpoints(
Ok((remotes, tally))
}
/// Hold a source that matched a planned target's IP on a different port.
///
/// That is what a symmetric NAT's fresh mapping toward us looks like, and it
/// is the main class of pairing punching exists to rescue, so it is adopted
/// rather than dropped. It is held for `PUNCH_SETTLE_MS` first so an exact
/// match arriving inside that window supersedes it; the honest path's latency
/// is unchanged, because an exact match breaks the loop immediately.
fn hold_remapped(
remote: SocketAddr,
candidate: &mut Option<SocketAddr>,
settle_at: &mut Option<tokio::time::Instant>,
superseded: &mut usize,
) {
if candidate.replace(remote).is_some() {
*superseded += 1;
}
settle_at.get_or_insert_with(|| {
tokio::time::Instant::now() + Duration::from_millis(PUNCH_SETTLE_MS)
});
}
pub(super) async fn run_punch_attempt(
socket: &std::net::UdpSocket,
session_id: &str,
@@ -427,22 +538,59 @@ pub(super) async fn run_punch_attempt(
let expected_hash = session_hash(session_id);
let mut buf = [0u8; 2048];
// Counted rather than logged per packet: an attacker sets how many of
// these arrive, so a record each would trade the adoption this closes for
// log volume. One record at the end of the attempt instead.
let mut unplanned = 0usize;
let mut superseded = 0usize;
let mut candidate: Option<SocketAddr> = None;
let mut settle_at: Option<tokio::time::Instant> = None;
let result = loop {
let recv = tokio::time::timeout_at(finish_at, udp.recv_from(&mut buf)).await;
let deadline = settle_at.map_or(finish_at, |settle| settle.min(finish_at));
let recv = tokio::time::timeout_at(deadline, udp.recv_from(&mut buf)).await;
let Ok(Ok((len, remote))) = recv else {
break Err(BootstrapError::PunchTimeout(session_id.to_string()));
break match candidate {
Some(remote) => Ok(remote),
None => Err(BootstrapError::PunchTimeout(session_id.to_string())),
};
};
// Ranked ahead of the ack, not only ahead of the adoption: acking a
// source we never planned to probe is a reflection this node controls,
// and there is no reason to emit it. The packet's own discriminator is
// a digest of a value both peers already know and travels in the clear
// in every probe, so it proves only that the sender saw one.
let rank = rank_punch_source(remote, targets);
if rank == SourceRank::Unplanned {
unplanned += 1;
continue;
}
match classify_punch_packet(&buf[..len], expected_hash) {
PunchAction::Ignore => continue,
PunchAction::Ack { sequence } => {
let ack = build_punch_packet(PunchPacketKind::Ack, sequence, session_id);
let _ = udp.send_to(&ack, remote).await;
break Ok(remote);
if rank == SourceRank::Planned {
break Ok(remote);
}
hold_remapped(remote, &mut candidate, &mut settle_at, &mut superseded);
}
PunchAction::Matched => {
if rank == SourceRank::Planned {
break Ok(remote);
}
hold_remapped(remote, &mut candidate, &mut settle_at, &mut superseded);
}
PunchAction::Matched => break Ok(remote),
}
};
send_handle.abort();
if unplanned > 0 || superseded > 0 {
debug!(
session = %super::runtime::short_id(session_id),
unplanned,
superseded,
"traversal: punch packets refused on their source address"
);
}
result
}
+350 -23
View File
@@ -7,7 +7,7 @@
use alloc::sync::Arc;
use super::state::{Lookup, PendingLookup, RecentRequest};
use super::state::{Lookup, PendingLookup};
use super::wire::LookupRequest;
use crate::NodeAddr;
@@ -148,8 +148,6 @@ pub(crate) fn plan_initiate(request: &LookupRequest, rv: &impl RoutingView) -> V
pub(crate) enum RequestOutcome {
/// request_id already in the dedup cache — drop.
Duplicate,
/// dedup cache at capacity — drop. `len` is the current cache size (for the log).
DedupCacheFull { len: usize },
/// We are the lookup target — the shell generates + sends the response.
RespondAsTarget,
/// Forward the request onward (the shell calls the forward planner).
@@ -160,10 +158,84 @@ pub(crate) enum RequestOutcome {
TtlExhausted,
}
/// One dedup-cache entry dropped to make room for an arriving request.
///
/// Returned to the shell so it can count and log the eviction; the core does
/// no metrics and no logging itself.
pub(crate) struct Eviction {
/// The `request_id` that was dropped. Its reverse path is gone: a
/// response still in flight for it will be treated as unsolicited.
pub request_id: u64,
/// The link peer charged for the eviction — the one whose oldest entry
/// this was, which is not necessarily the peer being admitted.
pub peer: NodeAddr,
/// The per-peer share in force at the time, for the log line.
pub share: usize,
}
/// The result of classifying an inbound LookupRequest: the route decision,
/// plus any entry that was evicted to make room for it.
pub(crate) struct Classification {
/// What the shell should do with the request.
pub outcome: RequestOutcome,
/// The entry dropped to admit this request, if one was.
pub evicted: Option<Eviction>,
}
/// Evict from the dedup cache if admitting one more request would put this
/// peer over its share, or the cache over its capacity.
///
/// Who pays is the whole point. Over its own share a peer pays for itself,
/// and at global capacity the peer holding the most entries pays, so a light
/// peer's reverse path is never taken to admit a heavy one and extra
/// identities buy a flooder proportionally less. Nothing is evicted while
/// the peer is under its share and the cache is under capacity.
fn make_room(
lookup: &mut Lookup,
from: &NodeAddr,
max_recent: usize,
peer_count: usize,
) -> Option<Eviction> {
let share = Lookup::peer_share(max_recent, peer_count);
let victim = if lookup.peer_entries(from) >= share {
*from
} else if lookup.recent_requests.len() >= max_recent {
// Never charge the arriving peer when it is under its share: charge
// whoever is holding the most.
lookup.heaviest_peer()?
} else {
return None;
};
let request_id = lookup.evict_oldest_from(&victim)?;
Some(Eviction {
request_id,
peer: victim,
share,
})
}
/// Classify an inbound LookupRequest against the recent-request dedup cache and
/// the transit forward rate limiter. Purges expired dedup entries, records the
/// request for reverse-path forwarding on the non-drop paths, and decides the
/// route. Pure over Lookup state + node addr + injected clock; no I/O, no view.
///
/// A full cache evicts rather than refuses. Refusing put the capacity check
/// ahead of the check for whether the request names this node, so one link
/// peer emitting fresh `request_id`s could stop the node answering lookups
/// for itself and stop it carrying anyone else's for as long as it kept the
/// cache full — a denial of exactly the service the cache exists to protect.
/// [`make_room`] charges the eviction to the peer that filled the cache.
///
/// The loosening this accepts: an evicted `request_id` arriving again inside
/// the dedup window is forwarded a second time rather than recognised as a
/// duplicate, and a response still in flight for it has lost its reverse
/// path. The per-target forward limiter and the request TTL already bound
/// what that second forward can cost, and the alternative — refusing the
/// arrival — is the availability defect above.
///
/// `peer_count` is the current link-peer count, supplied by the caller: the
/// core is sans-IO and clockless and has no view of the live peer table.
#[allow(clippy::too_many_arguments)]
pub(crate) fn classify_request(
lookup: &mut Lookup,
request: &LookupRequest,
@@ -172,28 +244,26 @@ pub(crate) fn classify_request(
now_ms: u64,
recent_expiry_ms: u64,
max_recent: usize,
) -> RequestOutcome {
// Purge expired dedup entries (was purge_expired_requests).
lookup
.recent_requests
.retain(|_, entry| !entry.is_expired(now_ms, recent_expiry_ms));
peer_count: usize,
) -> Classification {
// Purge expired dedup entries (was purge_expired_requests). Cache and
// per-peer index are purged together, or the eviction policy below reads
// a stale index and charges the wrong peer.
lookup.purge_recent(now_ms, recent_expiry_ms);
if lookup.recent_requests.contains_key(&request.request_id) {
return RequestOutcome::Duplicate;
}
if lookup.recent_requests.len() >= max_recent {
return RequestOutcome::DedupCacheFull {
len: lookup.recent_requests.len(),
return Classification {
outcome: RequestOutcome::Duplicate,
evicted: None,
};
}
lookup
.recent_requests
.insert(request.request_id, RecentRequest::new(*from, now_ms));
if request.target == *my_addr {
return RequestOutcome::RespondAsTarget;
}
if request.can_forward() {
let evicted = make_room(lookup, from, max_recent, peer_count);
lookup.record_recent(request.request_id, *from, now_ms);
let outcome = if request.target == *my_addr {
RequestOutcome::RespondAsTarget
} else if request.can_forward() {
if lookup
.forward_limiter
.should_forward(&request.target, now_ms)
@@ -204,7 +274,8 @@ pub(crate) fn classify_request(
}
} else {
RequestOutcome::TtlExhausted
}
};
Classification { outcome, evicted }
}
/// How an inbound LookupResponse should be routed, decided from the
@@ -217,13 +288,21 @@ pub(crate) enum ResponseRoute {
Transit { from_peer: NodeAddr },
/// We originated this request — the shell verifies the proof and caches.
Originator,
/// Nobody asked for this: it is neither a request we transited nor an
/// answer to a lookup we have outstanding for its target. Dropped before
/// the identity resolve and the signature verify, so it costs nothing.
Unsolicited,
}
/// Classify an inbound LookupResponse against the recent-request dedup cache.
///
/// Pure decision over `Lookup` state: sets `response_forwarded` when this is
/// the first response we transit for the request. No I/O, no view, no metrics.
pub(crate) fn classify_response(lookup: &mut Lookup, request_id: u64) -> ResponseRoute {
pub(crate) fn classify_response(
lookup: &mut Lookup,
request_id: u64,
target: &NodeAddr,
) -> ResponseRoute {
match lookup.recent_requests.get_mut(&request_id) {
Some(recent) => {
if recent.response_forwarded {
@@ -235,7 +314,25 @@ pub(crate) fn classify_response(lookup: &mut Lookup, request_id: u64) -> Respons
}
}
}
None => ResponseRoute::Originator,
// Not a request we transited, so it claims to answer one of ours.
// Require that it names a target with a lookup outstanding and carries
// an id issued for it. The id is fresh 64-bit randomness drawn per
// attempt and the target signs over it, so a harvested response is
// bound to the request it answered and cannot be redirected or
// replayed. Replies to earlier attempts of a still-outstanding lookup
// still match, which is the common case on a link whose round trip
// exceeds the first rung of the retry ladder.
None => {
let solicited = lookup
.pending_lookups
.get(target)
.is_some_and(|pending| pending.matches(request_id));
if solicited {
ResponseRoute::Originator
} else {
ResponseRoute::Unsolicited
}
}
}
}
@@ -398,3 +495,233 @@ pub(crate) fn initiate_failed(lookup: &mut Lookup, dest: &NodeAddr, now_ms: u64)
lookup.pending_lookups.remove(dest);
lookup.backoff.record_failure(dest, now_ms);
}
#[cfg(test)]
mod dedup_eviction_tests {
//! Capacity policy for the dedup cache.
//!
//! These live beside the policy rather than in `lookup/tests/core.rs`
//! because they are the regression tests for a security finding and read
//! directly against `make_room`'s two branches.
use super::super::limits::{LookupBackoff, LookupForwardRateLimiter};
use super::super::state::MIN_RECENT_PER_PEER;
use super::*;
use crate::TreeCoordinate;
use crate::testutil::make_node_addr;
/// The cache bound used by the behavioural tests. Smaller than the
/// production 4096 so a saturation test stays cheap; the policy is a
/// function of the bound, not of its value.
const CACHE: usize = 128;
fn empty() -> Lookup {
Lookup::new(
LookupBackoff::default(),
LookupForwardRateLimiter::default(),
)
}
fn request(request_id: u64, target: NodeAddr) -> LookupRequest {
let origin = make_node_addr(0xCC);
LookupRequest::new(
request_id,
target,
origin,
TreeCoordinate::root(origin),
5,
0,
)
}
/// Deliver one transit request from `from`, discarding the route
/// decision: these tests are about which entries survive, and the
/// forward limiter's verdict does not affect what is cached.
fn deliver(
lookup: &mut Lookup,
request_id: u64,
from: &NodeAddr,
my_addr: &NodeAddr,
peer_count: usize,
) -> Classification {
let target = make_node_addr(0xBB);
classify_request(
lookup,
&request(request_id, target),
from,
my_addr,
1_000,
60_000,
CACHE,
peer_count,
)
}
#[test]
fn a_peer_over_its_share_evicts_its_own_oldest_and_not_a_light_peers() {
let mut lookup = empty();
let heavy = make_node_addr(0x01);
let light = make_node_addr(0x02);
let me = make_node_addr(0x99);
// 64 peers over a 128-entry cache: 128/64 = 2, floored to 64.
let peer_count = 64;
let share = Lookup::peer_share(CACHE, peer_count);
deliver(&mut lookup, 7, &light, &me, peer_count);
// One request past the share, so the heavy peer pays for its own
// admission rather than the cache paying for it.
for i in 0..=share as u64 {
deliver(&mut lookup, 1_000 + i, &heavy, &me, peer_count);
}
assert!(
lookup.recent_requests.contains_key(&7),
"a light peer's reverse-path entry must survive a neighbour's flood"
);
assert!(
!lookup.recent_requests.contains_key(&1_000),
"the flooder's own oldest entry is what pays for its newest"
);
assert!(
lookup.recent_requests.contains_key(&(1_000 + share as u64)),
"and its newest is admitted rather than dropped"
);
assert_eq!(
lookup.peer_entries(&heavy),
share,
"the flooder is held at its share"
);
}
#[test]
fn the_per_peer_share_never_falls_below_the_floor_however_many_peers() {
// A node with as many peers as cache entries would otherwise give
// each peer a share of one, which no genuine transit burst survives.
assert_eq!(Lookup::peer_share(CACHE, CACHE), MIN_RECENT_PER_PEER);
assert_eq!(Lookup::peer_share(CACHE, usize::MAX), MIN_RECENT_PER_PEER);
// Above the floor the share still tracks the peer count.
assert_eq!(Lookup::peer_share(4096, 8), 512);
// And a zero peer count (no links up yet) must not divide by zero.
assert_eq!(Lookup::peer_share(CACHE, 0), CACHE);
// Behaviourally: at the floor, entry number 64 costs the peer
// nothing and entry number 65 costs it its oldest.
let mut lookup = empty();
let peer = make_node_addr(0x01);
let me = make_node_addr(0x99);
let peer_count = usize::MAX;
for i in 0..MIN_RECENT_PER_PEER as u64 {
let evicted = deliver(&mut lookup, i, &peer, &me, peer_count).evicted;
assert!(evicted.is_none(), "nothing is evicted below the floor");
}
let evicted = deliver(&mut lookup, 999, &peer, &me, peer_count)
.evicted
.expect("the entry past the floor must evict");
assert_eq!(evicted.request_id, 0, "the peer's oldest entry pays");
assert_eq!(evicted.peer, peer);
assert_eq!(evicted.share, MIN_RECENT_PER_PEER);
}
#[test]
fn a_node_whose_dedup_cache_is_saturated_still_answers_a_lookup_for_itself() {
// The availability claim, and the actual finding: the capacity check
// used to sit ahead of the check for whether the request names us,
// so a peer holding the cache full made this node unresolvable.
let mut lookup = empty();
let flooder = make_node_addr(0x01);
let other = make_node_addr(0x02);
let me = make_node_addr(0x99);
// A single link peer, so its share is the whole cache and it can
// saturate without evicting itself.
let peer_count = 1;
for i in 0..CACHE as u64 {
deliver(&mut lookup, i, &flooder, &me, peer_count);
}
assert_eq!(
lookup.recent_requests.len(),
CACHE,
"precondition: the cache is full, or the rest observes nothing"
);
let classification = classify_request(
&mut lookup,
&request(u64::MAX, me),
&other,
&me,
1_000,
60_000,
CACHE,
peer_count,
);
assert!(
matches!(classification.outcome, RequestOutcome::RespondAsTarget),
"a saturated cache must not stop the node answering lookups for itself"
);
let evicted = classification
.evicted
.expect("room must have been made at capacity");
assert_eq!(
evicted.peer, flooder,
"the peer holding the most entries pays, not the arriving one"
);
assert_eq!(evicted.request_id, 0, "and it pays with its oldest");
assert!(
lookup.recent_requests.contains_key(&u64::MAX),
"the arriving request is recorded, so its response can be routed back"
);
assert_eq!(
lookup.recent_requests.len(),
CACHE,
"the cache stays at its bound"
);
}
#[test]
fn a_duplicate_is_not_indexed_twice_and_evicts_nothing() {
// The index is a second container over the same entries, so the
// maintenance risk is drift: everything the policy decides reads it.
let mut lookup = empty();
let peer = make_node_addr(0x01);
let me = make_node_addr(0x99);
deliver(&mut lookup, 42, &peer, &me, 1);
let repeat = deliver(&mut lookup, 42, &peer, &me, 1);
assert!(matches!(repeat.outcome, RequestOutcome::Duplicate));
assert!(repeat.evicted.is_none(), "a duplicate makes no room");
assert_eq!(lookup.peer_entries(&peer), 1);
assert_eq!(lookup.recent_requests.len(), 1);
}
#[test]
fn purging_expired_entries_leaves_the_index_level_with_the_cache() {
let mut lookup = empty();
let a = make_node_addr(0x01);
let b = make_node_addr(0x02);
let me = make_node_addr(0x99);
for i in 0..5u64 {
deliver(&mut lookup, i, &a, &me, 2);
}
for i in 100..103u64 {
deliver(&mut lookup, i, &b, &me, 2);
}
let indexed: usize = lookup.recent_by_peer.values().map(|ids| ids.len()).sum();
assert_eq!(
indexed,
lookup.recent_requests.len(),
"every cached request is indexed exactly once"
);
// Entries were stamped at 1_000 with a 60s window; age them out.
lookup.purge_recent(1_000 + 60_000 + 1, 60_000);
assert!(lookup.recent_requests.is_empty());
assert!(
lookup.recent_by_peer.is_empty(),
"the index must not keep entries the cache no longer holds"
);
}
}
+5 -2
View File
@@ -14,7 +14,7 @@
//! nodes generating fresh request_ids at high rate.
use crate::NodeAddr;
use crate::proto::rate_limit::PerAddrRateLimiter;
use crate::proto::rate_limit::{PerAddrRateLimiter, RecordOutcome};
use alloc::collections::BTreeMap;
// ============================================================================
@@ -180,7 +180,10 @@ impl LookupForwardRateLimiter {
/// Returns true if enough time has passed since the last forward
/// for this target. Updates internal state when returning true.
pub fn should_forward(&mut self, target: &NodeAddr, now_ms: u64) -> bool {
self.0.check_and_record(target, now_ms)
// A full map admits rather than refuses; for forwarding that is the
// same fail-open direction the routing-error limiter takes, and the
// per-target interval is the only thing lost.
self.0.check_and_record(target, now_ms) != RecordOutcome::Suppress
}
/// Replace the minimum interval in milliseconds (e.g., set to zero to disable).
+4
View File
@@ -30,4 +30,8 @@ pub(crate) use limits::{LookupBackoff, LookupForwardRateLimiter, MAX_RECENT_LOOK
#[cfg(test)]
pub(crate) use state::RecentRequest;
pub(crate) use state::{Lookup, PendingLookup};
// The eviction policy reads the share floor through `Lookup::peer_share`;
// only the node-level regression tests name the constant itself.
#[cfg(test)]
pub(crate) use state::MIN_RECENT_PER_PEER;
pub use wire::{LookupRequest, LookupResponse};
+132 -1
View File
@@ -6,7 +6,8 @@
//! evolve toward a sans-IO core without threading four fields through
//! `Node`.
use alloc::collections::BTreeMap;
use alloc::collections::{BTreeMap, VecDeque};
use alloc::vec::Vec;
use super::limits::{LookupBackoff, LookupForwardRateLimiter};
use crate::NodeAddr;
@@ -44,6 +45,18 @@ impl RecentRequest {
}
}
/// How many outstanding `request_id`s one pending lookup remembers.
///
/// Bounds the per-target correlator at eight u64s. The retry ladder
/// (`node.discovery.attempt_timeouts_secs`) is operator configuration and can
/// be longer than this, so the recorder evicts the oldest id rather than
/// refusing the newest: dropping the newest would discard the id most likely
/// to be answered and fail a healthy lookup. Raising this costs eight bytes
/// per extra attempt on every pending target and widens the set of ids a late
/// response may still match; lowering it means a reply to an early attempt on
/// a long ladder is dropped as unsolicited.
const MAX_RECORDED_IDS: usize = 8;
/// Tracks a pending lookup with retry state.
pub struct PendingLookup {
/// When the lookup was first initiated.
@@ -52,6 +65,11 @@ pub struct PendingLookup {
pub last_sent_ms: u64,
/// Current attempt number (1 = initial, 2 = first retry, ...).
pub attempt: u8,
/// `request_id`s issued for this target, oldest first, capped at
/// [`MAX_RECORDED_IDS`]. A response is only acted on when it carries one
/// of these, which is what makes the accept path solicited. The entry
/// itself is dropped at ladder timeout, so this set needs no expiry.
pub ids: Vec<u64>,
}
impl PendingLookup {
@@ -60,15 +78,55 @@ impl PendingLookup {
initiated_ms: now_ms,
last_sent_ms: now_ms,
attempt: 1,
ids: Vec::new(),
}
}
/// Remember a `request_id` just put on the wire for this target.
pub fn record(&mut self, request_id: u64) {
if self.ids.contains(&request_id) {
return;
}
if self.ids.len() >= MAX_RECORDED_IDS {
self.ids.remove(0);
}
self.ids.push(request_id);
}
/// Whether `request_id` is one this node issued for this target.
pub fn matches(&self, request_id: u64) -> bool {
self.ids.contains(&request_id)
}
}
/// Floor under one link peer's share of the dedup cache.
///
/// A peer's share is the cache size divided by the current link-peer count,
/// and this is what stops that share collapsing to nothing on a node with
/// very many links. It is a cap and not a reservation: shares can sum past
/// the cache size, in which case the peer holding the most entries pays for
/// the next admission. Raising it lets one busy neighbour hold more of the
/// cache; lowering it clips a genuine transit burst.
pub(crate) const MIN_RECENT_PER_PEER: usize = 64;
/// Mesh lookup subsystem state.
pub(crate) struct Lookup {
/// Recent lookup requests (dedup + reverse-path forwarding).
/// Maps request_id → RecentRequest.
pub(crate) recent_requests: BTreeMap<u64, RecentRequest>,
/// Arrival-order index over `recent_requests`, partitioned by the link
/// peer each request arrived from. The cache is full-then-evict rather
/// than full-then-refuse, and this is what lets an eviction be charged
/// to the peer that filled the cache instead of to whoever happens to be
/// oldest. `now_ms` is nondecreasing across inserts, so each deque is in
/// arrival order and its front is that peer's oldest entry.
///
/// The index and the cache are two containers where there was one, so
/// they must be maintained together: every mutation of `recent_requests`
/// goes through [`Lookup::record_recent`], [`Lookup::evict_oldest_from`]
/// or [`Lookup::purge_recent`], which keep the two level. A drifted index
/// evicts the wrong entry, or none at all.
pub(crate) recent_by_peer: BTreeMap<NodeAddr, VecDeque<u64>>,
/// Tracks in-flight lookups. Maps target NodeAddr to the
/// initiation timestamp (Unix ms). Prevents duplicate flood queries.
pub(crate) pending_lookups: BTreeMap<NodeAddr, PendingLookup>,
@@ -87,12 +145,85 @@ impl Lookup {
pub(crate) fn new(backoff: LookupBackoff, forward_limiter: LookupForwardRateLimiter) -> Self {
Self {
recent_requests: BTreeMap::new(),
recent_by_peer: BTreeMap::new(),
pending_lookups: BTreeMap::new(),
backoff,
forward_limiter,
}
}
/// One link peer's share of the dedup cache at the current peer count.
///
/// The share tracks the peer count rather than being pinned to a number
/// a many-peer node outgrows, with [`MIN_RECENT_PER_PEER`] as its floor.
/// `peer_count` is supplied by the caller: the core is sans-IO and has no
/// view of the live peer table.
pub(crate) fn peer_share(max_recent: usize, peer_count: usize) -> usize {
(max_recent / peer_count.max(1)).max(MIN_RECENT_PER_PEER)
}
/// How many cached entries `peer` currently holds.
pub(crate) fn peer_entries(&self, peer: &NodeAddr) -> usize {
self.recent_by_peer.get(peer).map_or(0, VecDeque::len)
}
/// The link peer holding the most cached entries, if any.
///
/// Ties resolve to the highest `NodeAddr`, because `max_by_key` keeps the
/// last maximum and the index is ordered. Arbitrary but deterministic:
/// nothing about the policy depends on which of two equally heavy peers
/// pays, only that the choice does not vary run to run.
pub(crate) fn heaviest_peer(&self) -> Option<NodeAddr> {
self.recent_by_peer
.iter()
.max_by_key(|(_, ids)| ids.len())
.map(|(peer, _)| *peer)
}
/// Record a request for dedup and reverse-path forwarding, indexing it
/// under the link peer it arrived from.
///
/// The caller has already established that `request_id` is not cached; a
/// duplicate must not reach here, or the index would hold it twice.
pub(crate) fn record_recent(&mut self, request_id: u64, from: NodeAddr, now_ms: u64) {
self.recent_requests
.insert(request_id, RecentRequest::new(from, now_ms));
self.recent_by_peer
.entry(from)
.or_default()
.push_back(request_id);
}
/// Drop `peer`'s oldest cached entry, returning the evicted `request_id`.
///
/// Returns `None` when the peer holds nothing, which the eviction policy
/// treats as "no room could be made" rather than as an error.
pub(crate) fn evict_oldest_from(&mut self, peer: &NodeAddr) -> Option<u64> {
let ids = self.recent_by_peer.get_mut(peer)?;
let evicted = ids.pop_front()?;
if ids.is_empty() {
self.recent_by_peer.remove(peer);
}
self.recent_requests.remove(&evicted);
Some(evicted)
}
/// Purge expired dedup entries, from the cache and the index together.
///
/// The index is rebuilt from what survived rather than aged on its own
/// clock, so the two cannot drift apart: an id is indexed if and only if
/// the cache still holds it, and a peer disappears from the index when
/// its last entry does.
pub(crate) fn purge_recent(&mut self, now_ms: u64, expiry_ms: u64) {
self.recent_requests
.retain(|_, entry| !entry.is_expired(now_ms, expiry_ms));
let recent = &self.recent_requests;
self.recent_by_peer.retain(|_, ids| {
ids.retain(|id| recent.contains_key(id));
!ids.is_empty()
});
}
/// Reset lookup backoff on topology changes. Returns the number of
/// entries cleared (0 if already empty) so the shell can log the reset —
/// observability stays out of the pure core.
+107 -27
View File
@@ -152,7 +152,9 @@ fn classify_response_transit_on_fresh_forwarded_request() {
.recent_requests
.insert(42, RecentRequest::new(from_peer, 1000));
match classify_response(&mut lookup, 42) {
// Transit is decided by the dedup record alone, so the target the response
// names plays no part here — pass one no pending lookup mentions.
match classify_response(&mut lookup, 42, &make_node_addr(0xF1)) {
ResponseRoute::Transit { from_peer: peer } => assert_eq!(peer, from_peer),
_ => panic!("expected Transit"),
}
@@ -168,21 +170,70 @@ fn classify_response_already_forwarded_on_second_call() {
.recent_requests
.insert(7, RecentRequest::new(from_peer, 1000));
let target = make_node_addr(0xF2);
assert!(matches!(
classify_response(&mut lookup, 7),
classify_response(&mut lookup, 7, &target),
ResponseRoute::Transit { .. }
));
assert!(matches!(
classify_response(&mut lookup, 7),
classify_response(&mut lookup, 7, &target),
ResponseRoute::AlreadyForwarded
));
}
#[test]
fn classify_response_originator_when_request_absent() {
fn classify_response_originator_when_a_pending_lookup_issued_the_id() {
// No dedup record, so this is not a response we transit. It counts as ours
// only because the target has a lookup outstanding and that lookup issued
// the very request_id the response carries.
let target = make_node_addr(0x51);
let mut lookup = empty_lookup();
let mut pending = PendingLookup::new(1000);
pending.record(999);
lookup.pending_lookups.insert(target, pending);
assert!(matches!(
classify_response(&mut lookup, 999),
classify_response(&mut lookup, 999, &target),
ResponseRoute::Originator
));
}
#[test]
fn classify_response_unsolicited_when_nothing_correlates_the_id() {
// The three ways a response can fail to correlate, each of which must be
// dropped before the identity resolve and the signature verify rather than
// being treated as an answer to something we asked for.
let target = make_node_addr(0x52);
let other = make_node_addr(0x53);
let mut lookup = empty_lookup();
// 1. Nothing outstanding at all.
assert!(matches!(
classify_response(&mut lookup, 999, &target),
ResponseRoute::Unsolicited
));
// 2. A lookup is outstanding for the target, but it never issued this id:
// a harvested response cannot be replayed against a live lookup.
let mut pending = PendingLookup::new(1000);
pending.record(1);
lookup.pending_lookups.insert(target, pending);
assert!(matches!(
classify_response(&mut lookup, 999, &target),
ResponseRoute::Unsolicited
));
// 3. The id was issued, but for a different target: the id is bound to the
// exchange it was drawn for and cannot be redirected onto another name.
assert!(matches!(
classify_response(&mut lookup, 1, &other),
ResponseRoute::Unsolicited
));
// The matching pair still classifies as ours, so the checks above are
// discriminating rather than rejecting everything.
assert!(matches!(
classify_response(&mut lookup, 1, &target),
ResponseRoute::Originator
));
}
@@ -365,7 +416,8 @@ fn classify_request_forwards_fresh_and_records_it() {
let target = make_node_addr(0xAA);
let request = make_request_id(1, target, 3);
let outcome = classify_request(&mut lookup, &request, &from, &my_addr, 1000, 5000, 4096);
let outcome =
classify_request(&mut lookup, &request, &from, &my_addr, 1000, 5000, 4096, 1).outcome;
assert!(matches!(outcome, RequestOutcome::Forward));
// Recorded for reverse-path forwarding.
assert!(lookup.recent_requests.contains_key(&1));
@@ -381,32 +433,39 @@ fn classify_request_duplicate_on_second_call() {
let request = make_request_id(1, target, 3);
assert!(matches!(
classify_request(&mut lookup, &request, &from, &my_addr, 1000, 5000, 4096),
classify_request(&mut lookup, &request, &from, &my_addr, 1000, 5000, 4096, 1).outcome,
RequestOutcome::Forward
));
assert!(matches!(
classify_request(&mut lookup, &request, &from, &my_addr, 1000, 5000, 4096),
classify_request(&mut lookup, &request, &from, &my_addr, 1000, 5000, 4096, 1).outcome,
RequestOutcome::Duplicate
));
}
#[test]
fn classify_request_dedup_cache_full() {
fn classify_request_evicts_rather_than_refusing_a_full_dedup_cache() {
// Regression. The cache-full path used to drop the arriving request, so
// one peer emitting fresh request_ids
// could stop this node forwarding anyone's lookups and answering
// lookups for itself. A full cache now evicts instead, charged to the
// peer holding the most entries.
let mut lookup = empty_lookup();
let from = make_node_addr(0x01);
let my_addr = make_node_addr(0x99);
let target = make_node_addr(0xAA);
// Fill the cache to max_recent with distinct request_ids.
// Fill the cache to max_recent with distinct request_ids, through
// `record_recent` so the per-peer index stays level with the cache. A
// direct `recent_requests.insert` would leave the index short and turn
// the eviction policy into a no-op, so the test would pass without
// exercising it.
let max_recent = 3usize;
for id in 100..(100 + max_recent as u64) {
lookup
.recent_requests
.insert(id, RecentRequest::new(from, 1000));
lookup.record_recent(id, from, 1000);
}
assert_eq!(lookup.recent_requests.len(), max_recent);
let request = make_request_id(1, target, 3);
match classify_request(
let classification = classify_request(
&mut lookup,
&request,
&from,
@@ -414,12 +473,25 @@ fn classify_request_dedup_cache_full() {
1000,
5000,
max_recent,
) {
RequestOutcome::DedupCacheFull { len } => assert_eq!(len, max_recent),
_ => panic!("expected DedupCacheFull"),
}
// The new request must not have been recorded on the drop path.
assert!(!lookup.recent_requests.contains_key(&1));
1,
);
assert!(
matches!(classification.outcome, RequestOutcome::Forward),
"a full cache must not stop the node forwarding"
);
let evicted = classification
.evicted
.expect("admitting into a full cache must evict something");
assert_eq!(evicted.request_id, 100, "the oldest entry is what pays");
assert_eq!(
evicted.peer, from,
"and the peer that filled the cache is what pays"
);
// The arriving request is admitted, the evicted one is gone, and the
// cache has not grown past its bound.
assert!(lookup.recent_requests.contains_key(&1));
assert!(!lookup.recent_requests.contains_key(&100));
assert_eq!(lookup.recent_requests.len(), max_recent);
}
#[test]
@@ -431,7 +503,7 @@ fn classify_request_respond_as_target() {
let request = make_request_id(1, my_addr, 3);
assert!(matches!(
classify_request(&mut lookup, &request, &from, &my_addr, 1000, 5000, 4096),
classify_request(&mut lookup, &request, &from, &my_addr, 1000, 5000, 4096, 1).outcome,
RequestOutcome::RespondAsTarget
));
// Recorded before the target decision.
@@ -448,7 +520,7 @@ fn classify_request_ttl_exhausted_for_non_target() {
let request = make_request_id(1, target, 0);
assert!(matches!(
classify_request(&mut lookup, &request, &from, &my_addr, 1000, 5000, 4096),
classify_request(&mut lookup, &request, &from, &my_addr, 1000, 5000, 4096, 1).outcome,
RequestOutcome::TtlExhausted
));
}
@@ -465,7 +537,7 @@ fn classify_request_forward_rate_limited() {
let request = make_request_id(1, target, 3);
assert!(matches!(
classify_request(&mut lookup, &request, &from, &my_addr, 1000, 5000, 4096),
classify_request(&mut lookup, &request, &from, &my_addr, 1000, 5000, 4096, 1).outcome,
RequestOutcome::ForwardRateLimited
));
}
@@ -478,12 +550,20 @@ fn classify_request_purges_expired_entries() {
let target = make_node_addr(0xAA);
// Seed an entry that is expired at now_ms with the given expiry window.
// is_expired: now - timestamp > expiry_ms → expired.
lookup
.recent_requests
.insert(55, RecentRequest::new(from, 1000));
lookup.record_recent(55, from, 1000);
// now_ms = 10_000, expiry_ms = 5000 → 9000 > 5000 → expired.
let request = make_request_id(1, target, 3);
let outcome = classify_request(&mut lookup, &request, &from, &my_addr, 10_000, 5000, 4096);
let outcome = classify_request(
&mut lookup,
&request,
&from,
&my_addr,
10_000,
5000,
4096,
1,
)
.outcome;
assert!(matches!(outcome, RequestOutcome::Forward));
// The expired entry (55) must have been purged.
assert!(!lookup.recent_requests.contains_key(&55));
+2 -2
View File
@@ -61,7 +61,7 @@ impl GapTracker {
fn observe(&mut self, counter: u64) -> u64 {
let Some(expected) = self.expected_next else {
// First frame: initialize
self.expected_next = Some(counter + 1);
self.expected_next = Some(counter.saturating_add(1));
return 0;
};
@@ -90,7 +90,7 @@ impl GapTracker {
// Update expected (always advance to counter+1 or keep expected if
// this was a late/reordered frame)
if counter >= expected {
self.expected_next = Some(counter + 1);
self.expected_next = Some(counter.saturating_add(1));
}
lost
+47
View File
@@ -750,3 +750,50 @@ fn test_reverse_delivery_rekey_reset() {
m.delivery_ratio_reverse
);
}
// The ceiling-counter behaviour the gap tracker gained with the counter
// overflow fix. Driven through `ReceiverState` like the burst tests above,
// because `GapTracker` is private to `receiver.rs`; the discriminator is
// unchanged either way, since pre-fix the unchecked `counter + 1` aborts the
// process under the dev profile's overflow checks rather than failing an
// assertion. The abort is the red, not a harness fault.
#[test]
fn test_gap_tracker_saturates_on_a_first_frame_at_the_ceiling_counter() {
let mut r = ReceiverState::new(32);
r.record_recv(u64::MAX, 0, 100, false, 0);
// A saturated expectation stops the tracker advancing; a repeat of the
// same counter then takes the in-order branch and reports no burst.
r.record_recv(u64::MAX, 0, 100, false, 0);
let rr = r.build_report(0).unwrap();
assert_eq!(rr.burst_loss_count, 0);
assert_eq!(rr.max_burst_loss, 0);
}
#[test]
fn test_gap_tracker_saturates_when_advancing_onto_the_ceiling_counter() {
let mut r = ReceiverState::new(32);
// Prime the tracker so the advance branch is taken rather than the
// first-frame branch.
r.record_recv(1, 0, 100, false, 0);
// Reaching this line at all is the discriminator: pre-fix the unchecked
// `counter + 1` overflows here and aborts the process under the dev
// profile's overflow checks. The jump itself is a legitimate enormous gap
// and is correctly reported as one, so this does NOT assert an empty
// burst — an earlier version of this test did, and was simply wrong.
r.record_recv(u64::MAX, 0, 100, false, 0);
let jump = r.build_report(0).unwrap();
assert!(
jump.burst_loss_count >= 1,
"the jump to the ceiling is a real gap and should be reported"
);
// Saturated: the expectation cannot advance past the ceiling, so a repeat
// of the same counter takes the in-order branch and opens no new burst.
r.record_recv(u64::MAX, 0, 100, false, 0);
let after = r.build_report(0).unwrap();
assert_eq!(
after.burst_loss_count, 0,
"a saturated expectation must not keep opening bursts"
);
}
+72 -7
View File
@@ -5,12 +5,54 @@
use crate::NodeAddr;
use alloc::collections::BTreeMap;
/// Maximum number of addresses one limiter remembers at once.
///
/// A hard ceiling on the map a sender can grow by varying the address it keys
/// on, which for the routing-error limiter is a field the sender picks freely.
/// Raising it costs one `NodeAddr` plus one `u64` per entry and buys interval
/// suppression across more simultaneously-active addresses; lowering it makes
/// [`RecordOutcome::AdmitAtCapacity`] the common case sooner, which weakens the
/// interval gate but never the caller's own budget.
pub(crate) const MAX_ENTRIES: usize = 4096;
/// Fraction of `max_age_ms` between amortized sweeps.
///
/// The sweep is a full-map `retain`, so running it on every admit made
/// per-event cost linear in a map the sender sizes. Eight sweeps per entry
/// lifetime keeps expired entries from accumulating without putting the scan on
/// the per-event path.
const SWEEPS_PER_MAX_AGE: u64 = 8;
/// What the limiter decided about one candidate event.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub(crate) enum RecordOutcome {
/// Admit it; the address was recorded.
Admit,
/// Admit it, but the map was full so the address was not recorded and the
/// interval will not suppress its successor.
///
/// This gate fails open deliberately. Failing closed would turn a full map
/// into node-wide silence, and the map is fullest exactly during partition
/// healing, when many destinations are legitimately unroutable at once and
/// sources most need the signal.
AdmitAtCapacity,
/// Suppress it; an event for this address occurred within the interval.
Suppress,
}
/// Per-address minimum-interval rate limiter. Tracks the last event time per
/// address and enforces a minimum interval, evicting entries older than a max age.
/// address and enforces a minimum interval, evicting entries older than a max
/// age. The map is bounded and its sweep is amortized; see [`MAX_ENTRIES`] and
/// [`SWEEPS_PER_MAX_AGE`].
pub(crate) struct PerAddrRateLimiter {
last: BTreeMap<NodeAddr, u64>,
min_interval_ms: u64,
max_age_ms: u64,
last_sweep_ms: u64,
/// Sweeps run since construction. Read by the test that holds the
/// amortization property: a full-map scan per admit is the denial-of-
/// service multiplier this counter exists to catch coming back.
sweeps: u64,
}
impl PerAddrRateLimiter {
@@ -19,20 +61,38 @@ impl PerAddrRateLimiter {
last: BTreeMap::new(),
min_interval_ms,
max_age_ms,
last_sweep_ms: 0,
sweeps: 0,
}
}
/// Returns true (and records `now_ms`) if enough time has elapsed since the
/// last event for `addr`, or this is the first; false if within the interval.
pub(crate) fn check_and_record(&mut self, addr: &NodeAddr, now_ms: u64) -> bool {
if let Some(&last) = self.last.get(addr)
/// Decide one event for `addr`, recording `now_ms` when it is admitted and
/// there is room. See [`RecordOutcome`] for why a full map admits rather
/// than refuses.
pub(crate) fn check_and_record(&mut self, addr: &NodeAddr, now_ms: u64) -> RecordOutcome {
let known = self.last.get(addr).copied();
if let Some(last) = known
&& now_ms.saturating_sub(last) < self.min_interval_ms
{
return false;
return RecordOutcome::Suppress;
}
self.maybe_sweep(now_ms);
if known.is_none() && self.last.len() >= MAX_ENTRIES {
return RecordOutcome::AdmitAtCapacity;
}
self.last.insert(*addr, now_ms);
RecordOutcome::Admit
}
/// Run the expiry sweep at most once per `max_age_ms / SWEEPS_PER_MAX_AGE`.
fn maybe_sweep(&mut self, now_ms: u64) {
let interval = (self.max_age_ms / SWEEPS_PER_MAX_AGE).max(1);
if now_ms.saturating_sub(self.last_sweep_ms) < interval && self.last_sweep_ms != 0 {
return;
}
self.last_sweep_ms = now_ms;
self.sweeps = self.sweeps.saturating_add(1);
self.cleanup(now_ms);
true
}
pub(crate) fn cleanup(&mut self, now_ms: u64) {
@@ -40,6 +100,11 @@ impl PerAddrRateLimiter {
.retain(|_, &mut last| now_ms.saturating_sub(last) < self.max_age_ms);
}
#[cfg(test)]
pub(crate) fn sweeps(&self) -> u64 {
self.sweeps
}
#[cfg(test)]
pub(crate) fn set_interval_ms(&mut self, interval_ms: u64) {
self.min_interval_ms = interval_ms;
+66 -26
View File
@@ -14,6 +14,7 @@
//! shell hands over only raw per-peer reads; all routing narrowing and decision
//! logic lives here.
use super::limits::LimitVerdict;
use super::state::Router;
use super::wire::{CoordsRequired, MtuExceeded, PathBroken};
use crate::proto::link::{SessionDatagram, SessionDatagramRef};
@@ -171,8 +172,14 @@ impl Router {
/// CoordsRequired. The chosen PDU is wrapped in a fresh SessionDatagram
/// addressed back to `toward` (the failed datagram's source) and encoded.
///
/// Returns `None` when the rate-limit gate suppresses the signal (the shell
/// drops silently). On `Some`, the shell resolves the reverse link hop for
/// The returned [`ErrorSynth`] carries the gate's verdict alongside the
/// action, rather than collapsing it to a bool. The shell counts the three
/// verdicts separately: suppression and a full destination map say
/// different things about the node, and an operator cannot tell a genuine
/// outage from a limiter that has stopped limiting without the split.
///
/// `action` is `None` exactly when the verdict is `Suppress` (the shell
/// drops silently). Otherwise the shell resolves the reverse link hop for
/// `toward` and sends — resolving the hop only after this gate preserves
/// the pre-refactor ordering (rate-limit before `find_next_hop`'s cache
/// touch) and lets the shell distinguish suppression from no-reverse-route
@@ -185,21 +192,34 @@ impl Router {
rv: &impl RoutingView,
now_ms: u64,
default_ttl: u8,
) -> Option<RouteAction> {
if !self.error_limiter.should_send(dest, now_ms) {
return None;
) -> ErrorSynth {
let verdict = self.error_limiter.check(dest, now_ms);
if verdict == LimitVerdict::Suppress {
return ErrorSynth {
verdict,
action: None,
};
}
let error_payload = match rv.cached_coords(dest, now_ms) {
Some(coords) => PathBroken::new(*dest, *my_addr)
.with_last_coords(coords)
.encode(),
None => CoordsRequired::new(*dest, *my_addr).encode(),
// Which of the two signals is emitted still discloses whether this
// node holds coords for the destination, but the coordinates
// themselves are not attached: the error is returned to the datagram's
// own src_addr, which nothing binds to the peer that sent it, so
// attaching them would answer a coordinate-cache read to whoever names
// an address, one entry per packet. The field is optional on the wire
// and no receiver reads it, so this is an emission change only.
let error_payload = if rv.cached_coords(dest, now_ms).is_some() {
PathBroken::new(*dest, *my_addr).encode()
} else {
CoordsRequired::new(*dest, *my_addr).encode()
};
let error_dg = SessionDatagram::new(*my_addr, *toward, error_payload).with_ttl(default_ttl);
Some(RouteAction::SendError {
toward: *toward,
bytes: error_dg.encode(),
})
ErrorSynth {
verdict,
action: Some(RouteAction::SendError {
toward: *toward,
bytes: error_dg.encode(),
}),
}
}
/// Synthesize an MtuExceeded error signal after a forward send failed with
@@ -208,11 +228,12 @@ impl Router {
/// carrying `bottleneck_mtu`, wraps it in a fresh SessionDatagram addressed
/// back to `toward` (the failed datagram's source), and encodes it.
///
/// Returns `None` when the gate suppresses the signal. On `Some`, the shell
/// resolves the reverse link hop for `toward` and sends — resolving the hop
/// only after this gate preserves the pre-refactor ordering (rate-limit
/// before `find_next_hop`'s cache touch). No coordinate read is involved;
/// unlike routing errors, the PDU is unconditional once the gate passes.
/// The returned [`ErrorSynth`]'s `action` is `None` exactly when the
/// verdict is `Suppress`. Otherwise the shell resolves the reverse link hop
/// for `toward` and sends — resolving the hop only after this gate
/// preserves the pre-refactor ordering (rate-limit before
/// `find_next_hop`'s cache touch). No coordinate read is involved; unlike
/// routing errors, the PDU is unconditional once the gate passes.
pub(crate) fn synth_mtu_exceeded(
&mut self,
dest: &NodeAddr,
@@ -221,19 +242,38 @@ impl Router {
bottleneck_mtu: u16,
now_ms: u64,
default_ttl: u8,
) -> Option<RouteAction> {
if !self.error_limiter.should_send(dest, now_ms) {
return None;
) -> ErrorSynth {
let verdict = self.error_limiter.check(dest, now_ms);
if verdict == LimitVerdict::Suppress {
return ErrorSynth {
verdict,
action: None,
};
}
let error_payload = MtuExceeded::new(*dest, *my_addr, bottleneck_mtu).encode();
let error_dg = SessionDatagram::new(*my_addr, *toward, error_payload).with_ttl(default_ttl);
Some(RouteAction::SendError {
toward: *toward,
bytes: error_dg.encode(),
})
ErrorSynth {
verdict,
action: Some(RouteAction::SendError {
toward: *toward,
bytes: error_dg.encode(),
}),
}
}
}
/// One candidate error signal, as decided by the per-destination gate.
///
/// The verdict travels with the action so the shell can count `Suppress` and
/// `AdmitAtCapacity` separately; a bool return collapsed the two admits
/// together and left both counters with no writer.
pub(crate) struct ErrorSynth {
/// What the per-destination interval gate decided.
pub verdict: LimitVerdict,
/// The signal to send. `None` exactly when `verdict` is `Suppress`.
pub action: Option<RouteAction>,
}
/// Route class of a transit-forwarded packet, classified from tree
/// coordinates at the forwarding decision point. The six variants
/// partition `forwarded_packets` exactly.
+37 -2
View File
@@ -9,7 +9,7 @@
//! portability and deterministic ordering.
use crate::NodeAddr;
use crate::proto::rate_limit::PerAddrRateLimiter;
use crate::proto::rate_limit::{PerAddrRateLimiter, RecordOutcome};
/// Default minimum interval between error signals: 100 ms (max 10 errors/sec
/// per destination).
@@ -22,6 +22,19 @@ const MAX_AGE_MS: u64 = 10_000;
///
/// Tracks the last time a routing error was sent for each destination
/// address and enforces a minimum interval to prevent floods.
/// What the limiter decided about one candidate routing-error signal.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum LimitVerdict {
/// Send it; the destination was recorded.
Admit,
/// Send it, but the destination map was full so the destination was not
/// recorded and the interval will not suppress its successor. Emission
/// stays bounded by the per-peer error budget in this state.
AdmitAtCapacity,
/// Suppress it; an error for this destination went out within the interval.
Suppress,
}
pub struct RoutingErrorRateLimiter(PerAddrRateLimiter);
impl RoutingErrorRateLimiter {
@@ -44,7 +57,22 @@ impl RoutingErrorRateLimiter {
/// this destination, or if this is the first error. Updates internal
/// state when returning true.
pub fn should_send(&mut self, dest_addr: &NodeAddr, now_ms: u64) -> bool {
self.0.check_and_record(dest_addr, now_ms)
self.check(dest_addr, now_ms) != LimitVerdict::Suppress
}
/// Decide one candidate error signal, distinguishing an ordinary admit
/// from one made only because the destination map was full.
///
/// The caller needs the difference because the two say different things
/// about the node: `AdmitAtCapacity` means the interval gate is no longer
/// suppressing anything for this destination, and only the per-peer budget
/// is bounding emission.
pub fn check(&mut self, dest_addr: &NodeAddr, now_ms: u64) -> LimitVerdict {
match self.0.check_and_record(dest_addr, now_ms) {
RecordOutcome::Admit => LimitVerdict::Admit,
RecordOutcome::AdmitAtCapacity => LimitVerdict::AdmitAtCapacity,
RecordOutcome::Suppress => LimitVerdict::Suppress,
}
}
/// Remove entries older than max_age.
@@ -57,6 +85,13 @@ impl RoutingErrorRateLimiter {
pub fn len(&self) -> usize {
self.0.len()
}
/// Sweeps the inner map has run since construction. Read by the test
/// holding the amortization property.
#[cfg(test)]
pub(crate) fn sweeps(&self) -> u64 {
self.0.sweeps()
}
}
impl Default for RoutingErrorRateLimiter {
+1 -1
View File
@@ -29,7 +29,7 @@ pub(crate) use core::{
DropReason, NextHop, RouteAction, RouteClass, RouteOutcome, RoutingView, classify_forward,
select_best_candidate,
};
pub(crate) use limits::RoutingErrorRateLimiter;
pub(crate) use limits::{LimitVerdict, RoutingErrorRateLimiter};
pub(crate) use state::Router;
pub use wire::{
COORDS_REQUIRED_SIZE, CoordsRequired, MTU_EXCEEDED_SIZE, MtuExceeded, PathBroken,
+28 -29
View File
@@ -4,7 +4,7 @@ use super::util::{MockPeer, MockRoutingView, make_coords, make_datagram_ref, mak
use crate::proto::link::SessionDatagramRef;
use crate::proto::routing::RoutingSignalType;
use crate::proto::routing::{
DropReason, RouteAction, RouteOutcome, Router, RoutingView, select_best_candidate,
DropReason, LimitVerdict, RouteAction, RouteOutcome, Router, RoutingView, select_best_candidate,
};
use crate::testutil::make_node_addr;
use crate::{NodeAddr, TreeCoordinate};
@@ -396,6 +396,7 @@ fn synth_uses_pathbroken_when_coords_cached() {
};
let action = router
.synth_routing_error(&dest, &source, &my_addr, &rv, 0, 64)
.action
.expect("gate passes on first call");
let RouteAction::SendError { toward, .. } = &action;
assert_eq!(
@@ -418,6 +419,7 @@ fn synth_uses_coords_required_when_not_cached() {
let rv = MockRoutingView::new(false); // empty coord table
let action = router
.synth_routing_error(&dest, &source, &my_addr, &rv, 0, 64)
.action
.expect("gate passes on first call");
assert_eq!(
error_pdu_type(&action),
@@ -434,26 +436,24 @@ fn synth_rate_limit_gate_suppresses_second_call() {
let my_addr = make_node_addr(0x10);
let rv = MockRoutingView::new(false);
// First call for this destination passes the gate.
assert!(
router
.synth_routing_error(&dest, &source, &my_addr, &rv, 0, 64)
.is_some()
);
let first = router.synth_routing_error(&dest, &source, &my_addr, &rv, 0, 64);
assert!(first.action.is_some());
assert_eq!(first.verdict, LimitVerdict::Admit);
// An immediate second call for the same destination is within the
// rate-limit window and is suppressed (no sleeps needed — the two calls
// are microseconds apart, well under the 100 ms interval).
assert!(
router
.synth_routing_error(&dest, &source, &my_addr, &rv, 0, 64)
.is_none()
let second = router.synth_routing_error(&dest, &source, &my_addr, &rv, 0, 64);
assert!(second.action.is_none());
assert_eq!(
second.verdict,
LimitVerdict::Suppress,
"the shell counts emit_over_dest_interval off this verdict"
);
// A different destination is independent and still allowed.
let other = make_node_addr(0x22);
assert!(
router
.synth_routing_error(&other, &source, &my_addr, &rv, 0, 64)
.is_some()
);
let third = router.synth_routing_error(&other, &source, &my_addr, &rv, 0, 64);
assert!(third.action.is_some());
assert_eq!(third.verdict, LimitVerdict::Admit);
}
/// Extract the bottleneck MTU (trailing u16 LE) from an MtuExceeded action's
@@ -474,6 +474,7 @@ fn synth_mtu_exceeded_carries_bottleneck_and_targets_source() {
let my_addr = make_node_addr(0x10);
let action = router
.synth_mtu_exceeded(&dest, &source, &my_addr, 1280, 0, 64)
.action
.expect("gate passes on first call");
let RouteAction::SendError { toward, .. } = &action;
assert_eq!(
@@ -499,22 +500,20 @@ fn synth_mtu_exceeded_rate_limit_gate_suppresses_second_call() {
let source = make_node_addr(0x21);
let my_addr = make_node_addr(0x10);
// First call for this destination passes the gate.
assert!(
router
.synth_mtu_exceeded(&dest, &source, &my_addr, 1280, 0, 64)
.is_some()
);
let first = router.synth_mtu_exceeded(&dest, &source, &my_addr, 1280, 0, 64);
assert!(first.action.is_some());
assert_eq!(first.verdict, LimitVerdict::Admit);
// Immediate second call for the same destination is suppressed.
assert!(
router
.synth_mtu_exceeded(&dest, &source, &my_addr, 1280, 0, 64)
.is_none()
let second = router.synth_mtu_exceeded(&dest, &source, &my_addr, 1280, 0, 64);
assert!(second.action.is_none());
assert_eq!(
second.verdict,
LimitVerdict::Suppress,
"the shell counts emit_over_dest_interval off this verdict"
);
// A different destination is independent and still allowed.
let other = make_node_addr(0x22);
assert!(
router
.synth_mtu_exceeded(&other, &source, &my_addr, 1280, 0, 64)
.is_some()
);
let third = router.synth_mtu_exceeded(&other, &source, &my_addr, 1280, 0, 64);
assert!(third.action.is_some());
assert_eq!(third.verdict, LimitVerdict::Admit);
}
+81 -1
View File
@@ -1,8 +1,19 @@
//! Tests for routing error-signal rate limiting.
use crate::proto::routing::RoutingErrorRateLimiter;
use crate::NodeAddr;
use crate::proto::rate_limit::MAX_ENTRIES;
use crate::proto::routing::{LimitVerdict, RoutingErrorRateLimiter};
use crate::testutil::make_node_addr as addr;
/// A distinct destination address per index, standing for the fresh
/// `dest_addr` a flooding sender puts on every datagram.
fn minted_addr(val: u32) -> NodeAddr {
let mut bytes = [0u8; 16];
bytes[..4].copy_from_slice(&val.to_le_bytes());
bytes[15] = 0xff;
NodeAddr::from_bytes(bytes)
}
#[test]
fn test_first_send_allowed() {
let mut limiter = RoutingErrorRateLimiter::new();
@@ -68,3 +79,72 @@ fn test_with_interval_custom_rate() {
// Allowed at 500 ms.
assert!(limiter.should_send(&addr(1), 500));
}
#[test]
fn the_map_stays_bounded_when_a_sender_mints_distinct_destination_keys() {
// `dest_addr` is a field the sender picks, so an unbounded map is memory
// the sender sizes.
let mut limiter = RoutingErrorRateLimiter::new();
for i in 0..100_000u32 {
limiter.check(&minted_addr(i), 0);
}
assert!(
limiter.len() <= MAX_ENTRIES,
"limiter held {} entries, above the {MAX_ENTRIES} ceiling",
limiter.len()
);
}
#[test]
fn an_admission_at_capacity_still_sends_rather_than_going_silent() {
// The ceiling fails open on purpose: the map is fullest during partition
// healing, when sources most need the signal.
let mut limiter = RoutingErrorRateLimiter::new();
for i in 0..MAX_ENTRIES as u32 {
assert_eq!(limiter.check(&minted_addr(i), 0), LimitVerdict::Admit);
}
// The map is full and nothing in it is old enough to evict, so the next
// distinct destination cannot be recorded. It must still be sent.
assert_eq!(
limiter.check(&minted_addr(MAX_ENTRIES as u32), 0),
LimitVerdict::AdmitAtCapacity
);
}
#[test]
fn the_map_scan_does_not_run_once_per_admitted_destination() {
// The sweep is a full-map retain. Running it per admit would make the
// per-event cost linear in a map the sender sizes, which is the denial of
// service the ceiling above exists to prevent.
let mut limiter = RoutingErrorRateLimiter::new();
limiter.check(&minted_addr(0), 1);
let before = limiter.sweeps();
for i in 1..1_000u32 {
limiter.check(&minted_addr(i), 1);
}
assert_eq!(
limiter.sweeps() - before,
0,
"the full-map scan ran inside a single sweep interval"
);
}
#[test]
fn the_map_scan_still_runs_once_a_sweep_interval_has_passed() {
let mut limiter = RoutingErrorRateLimiter::new();
limiter.check(&minted_addr(0), 1);
let before = limiter.sweeps();
// 11 s later: past both the sweep interval and the 10 s max age.
limiter.check(&minted_addr(1), 11_000);
assert_eq!(limiter.sweeps() - before, 1);
// The first destination aged out, so the sweep did its job.
assert_eq!(limiter.len(), 1);
}
+241 -6
View File
@@ -21,6 +21,90 @@ mod platform;
#[cfg(unix)]
pub use platform::PacketSocket;
/// Outcome of `send_frame`.
#[cfg(unix)]
#[cfg_attr(not(target_os = "macos"), allow(dead_code))]
pub(crate) enum SendOutcome {
Sent,
Stop,
}
/// Retry iterations spent yielding before the send loop starts sleeping.
///
/// A transiently full channel drains in microseconds, so yielding keeps the
/// saturated-path handoff rate uncapped, which is the whole reason this
/// module has a dedicated reader thread. Raising it burns more CPU against a
/// genuinely stuck consumer; lowering it puts a sleep in the common case.
#[cfg(unix)]
#[cfg_attr(not(target_os = "macos"), allow(dead_code))]
const SEND_YIELD_SPINS: u32 = 64;
/// Longest the send loop sleeps between attempts on a full channel.
///
/// This bounds only how quickly a parked send notices a shutdown request that
/// closing the receiver has not already covered. Raising it delays that
/// notice; lowering it costs more wakeups under sustained backpressure.
#[cfg(unix)]
#[cfg_attr(not(target_os = "macos"), allow(dead_code))]
const SEND_RETRY_MAX: std::time::Duration = std::time::Duration::from_millis(1);
/// Send one item, waiting out a full channel but waking on `shutdown_fd`.
///
/// Returns `Stop` when the receiver is gone or shutdown has been requested,
/// which is the reader thread's cue to exit. Unlike `blocking_send` this
/// cannot park past a shutdown request, so the `join()` in `Drop` always
/// returns. The caller must keep the socket owning `shutdown_fd` alive across
/// the call; `poll` on a closed fd reports `POLLNVAL` rather than `POLLIN`, so
/// even a lifetime mistake degrades to waiting rather than to a false stop.
///
/// Compiled on every unix so Linux CI exercises the tests below; only the
/// macOS reader thread calls it.
#[cfg(unix)]
#[cfg_attr(not(target_os = "macos"), allow(dead_code))]
pub(crate) fn send_frame<T>(
tx: &tokio::sync::mpsc::Sender<T>,
item: T,
shutdown_fd: std::os::unix::io::RawFd,
) -> SendOutcome {
use tokio::sync::mpsc::error::TrySendError;
let mut item = item;
let mut spins = 0u32;
let mut backoff = std::time::Duration::from_micros(50);
loop {
match tx.try_send(item) {
Ok(()) => return SendOutcome::Sent,
Err(TrySendError::Closed(_)) => return SendOutcome::Stop,
Err(TrySendError::Full(returned)) => {
if fd_is_readable(shutdown_fd) {
return SendOutcome::Stop;
}
item = returned;
if spins < SEND_YIELD_SPINS {
spins += 1;
std::thread::yield_now();
} else {
std::thread::sleep(backoff);
backoff = (backoff * 2).min(SEND_RETRY_MAX);
}
}
}
}
}
/// True if `fd` has data ready, tested without blocking.
#[cfg(unix)]
#[cfg_attr(not(target_os = "macos"), allow(dead_code))]
pub(crate) fn fd_is_readable(fd: std::os::unix::io::RawFd) -> bool {
let mut pfd = libc::pollfd {
fd,
events: libc::POLLIN,
revents: 0,
};
let ret = unsafe { libc::poll(&mut pfd, 1, 0) };
ret > 0 && (pfd.revents & libc::POLLIN) != 0
}
// =============================================================================
// Linux: AsyncFd-based async wrapper
// =============================================================================
@@ -111,7 +195,9 @@ mod async_impl {
pub struct AsyncPacketSocket {
inner: Arc<PacketSocket>,
rx: tokio::sync::Mutex<tokio::sync::mpsc::Receiver<Frame>>,
/// `None` once shutdown has taken the receiver, which is what makes
/// a reader thread parked on a full channel return at once.
rx: tokio::sync::Mutex<Option<tokio::sync::mpsc::Receiver<Frame>>>,
reader_thread: Option<std::thread::JoinHandle<()>>,
}
@@ -146,7 +232,13 @@ mod async_impl {
match result {
Ok((n, mac)) => {
let data = read_buf[..n].to_vec();
if tx.blocking_send((data, mac)).is_err() {
// Not blocking_send: a send parked on a
// full channel must still notice shutdown,
// or Drop's join() never returns.
if matches!(
super::send_frame(&tx, (data, mac), shutdown_fd),
super::SendOutcome::Stop
) {
return;
}
}
@@ -207,7 +299,7 @@ mod async_impl {
Ok(Self {
inner,
rx: tokio::sync::Mutex::new(rx),
rx: tokio::sync::Mutex::new(Some(rx)),
reader_thread: Some(reader_thread),
})
}
@@ -230,7 +322,10 @@ mod async_impl {
}
pub async fn recv_from(&self, buf: &mut [u8]) -> Result<(usize, [u8; 6]), TransportError> {
let mut rx = self.rx.lock().await;
let mut guard = self.rx.lock().await;
let Some(rx) = guard.as_mut() else {
return Err(TransportError::RecvFailed("reader thread stopped".into()));
};
match rx.recv().await {
Some((data, mac)) => {
let n = data.len().min(buf.len());
@@ -247,15 +342,25 @@ mod async_impl {
/// Signal the reader thread to stop.
///
/// Sets the shutdown flag; the reader thread checks it after
/// each BPF read timeout (~250ms) and exits.
/// Drops the receiver where it can, which makes a send parked on a
/// full channel fail immediately, then writes the shutdown pipe that
/// the thread's `select()` and `send_frame` both watch. The receiver
/// is unavailable while a `recv_from` holds the lock; `Drop` takes it
/// unconditionally, so the pipe is what covers that window.
pub fn shutdown(&self) {
if let Ok(mut guard) = self.rx.try_lock() {
guard.take();
}
self.inner.request_shutdown();
}
}
impl Drop for AsyncPacketSocket {
fn drop(&mut self) {
// Drop the receiver before joining: a send parked on a full
// channel then returns at once, with no polling and no latency
// added to the steady-state path.
self.rx.get_mut().take();
self.inner.request_shutdown();
if let Some(handle) = self.reader_thread.take() {
let _ = handle.join();
@@ -284,3 +389,133 @@ pub struct PacketSocket;
#[cfg(windows)]
pub struct AsyncPacketSocket;
// =============================================================================
// Tests
// =============================================================================
#[cfg(all(test, unix))]
mod tests {
use super::{SendOutcome, fd_is_readable, send_frame};
use std::sync::mpsc;
use std::time::Duration;
/// A pipe, as the shutdown signal, returned as (read fd, write fd).
///
/// Leaked deliberately: these live for the length of one test and closing
/// them mid-poll is exactly the confusion the test is meant to avoid.
fn shutdown_pipe() -> (std::os::unix::io::RawFd, std::os::unix::io::RawFd) {
let mut fds = [0i32; 2];
let ret = unsafe { libc::pipe(fds.as_mut_ptr()) };
assert_eq!(ret, 0, "pipe() failed");
(fds[0], fds[1])
}
fn signal(write_fd: std::os::unix::io::RawFd) {
let byte = [1u8];
let ret = unsafe { libc::write(write_fd, byte.as_ptr() as *const libc::c_void, 1) };
assert_eq!(ret, 1, "write() to shutdown pipe failed");
}
#[test]
fn fd_is_readable_is_false_for_an_unwritten_pipe_and_true_after_a_write() {
let (read_fd, write_fd) = shutdown_pipe();
assert!(!fd_is_readable(read_fd));
signal(write_fd);
assert!(fd_is_readable(read_fd));
}
#[test]
fn send_frame_delivers_when_the_channel_has_room() {
let (tx, mut rx) = tokio::sync::mpsc::channel::<Vec<u8>>(1);
let (read_fd, _write_fd) = shutdown_pipe();
assert!(matches!(
send_frame(&tx, vec![1u8, 2, 3], read_fd),
SendOutcome::Sent
));
assert_eq!(rx.try_recv().unwrap(), vec![1u8, 2, 3]);
}
#[test]
fn send_frame_returns_stop_when_the_receiver_is_gone() {
let (tx, rx) = tokio::sync::mpsc::channel::<Vec<u8>>(1);
let (read_fd, _write_fd) = shutdown_pipe();
drop(rx);
assert!(matches!(
send_frame(&tx, vec![0u8], read_fd),
SendOutcome::Stop
));
}
#[test]
fn send_frame_returns_stop_when_the_receiver_is_dropped_while_the_channel_is_full() {
// The mechanism `Drop` relies on: closing the channel releases a
// sender that is waiting for room.
let (tx, rx) = tokio::sync::mpsc::channel::<Vec<u8>>(1);
let (read_fd, _write_fd) = shutdown_pipe();
tx.try_send(vec![0u8]).unwrap();
let (done_tx, done_rx) = mpsc::channel();
let sender = std::thread::spawn(move || {
let outcome = send_frame(&tx, vec![1u8], read_fd);
done_tx.send(matches!(outcome, SendOutcome::Stop)).unwrap();
});
// The send is parked on a full channel; only the drop frees it.
assert!(done_rx.recv_timeout(Duration::from_millis(50)).is_err());
drop(rx);
let stopped = done_rx
.recv_timeout(Duration::from_secs(5))
.expect("send_frame did not return after the receiver was dropped");
sender.join().unwrap();
assert!(stopped);
}
#[test]
fn send_frame_returns_stop_when_shutdown_is_requested_and_the_channel_is_full() {
// The defect: `blocking_send` on a full channel nobody is draining
// parks forever, so the reader thread never sees shutdown and the
// `join()` in `Drop` never returns. See the ignored test below for
// the same fixture against `blocking_send`.
let (tx, _rx) = tokio::sync::mpsc::channel::<Vec<u8>>(1);
let (read_fd, write_fd) = shutdown_pipe();
tx.try_send(vec![0u8]).unwrap();
signal(write_fd);
let (done_tx, done_rx) = mpsc::channel();
let sender = std::thread::spawn(move || {
let outcome = send_frame(&tx, vec![1u8], read_fd);
done_tx.send(matches!(outcome, SendOutcome::Stop)).unwrap();
});
let stopped = done_rx
.recv_timeout(Duration::from_secs(5))
.expect("send_frame parked past a shutdown request");
sender.join().unwrap();
assert!(stopped);
}
#[test]
#[ignore = "demonstrates the defect: blocking_send never returns, so this hangs"]
fn blocking_send_parks_past_a_shutdown_request_when_the_channel_is_full() {
// Run with `--ignored` to watch the old send site hang. Kept as the
// observed red-before for the test above, which cannot itself fail
// against the old code because `send_frame` did not exist then.
let (tx, _rx) = tokio::sync::mpsc::channel::<Vec<u8>>(1);
let (_read_fd, write_fd) = shutdown_pipe();
tx.try_send(vec![0u8]).unwrap();
signal(write_fd);
let (done_tx, done_rx) = mpsc::channel();
std::thread::spawn(move || {
let _ = tx.blocking_send(vec![1u8]);
done_tx.send(()).unwrap();
});
done_rx
.recv_timeout(Duration::from_secs(5))
.expect("blocking_send returned, so the send site was already cancellable");
}
}
+7 -1
View File
@@ -449,7 +449,13 @@ async fn ethernet_receive_loop(
stats.record_beacon_recv();
if listen_enabled && let Some(pubkey) = parse_beacon(&buf[..len]) {
neighbor_buffer.add_peer(src_mac, pubkey);
// `add_peer` reports whether the buffer took the
// beacon. It refuses once the distinct-MAC cap is
// reached, which is the bound on an unauthenticated
// broadcast frame naming a fresh MAC every time.
if !neighbor_buffer.add_peer(src_mac, pubkey) {
stats.record_beacon_dropped();
}
trace!(
transport_id = %transport_id,
remote_mac = %format_mac(&src_mac),
+179 -12
View File
@@ -7,7 +7,10 @@
use crate::transport::{DiscoveredPeer, TransportAddr, TransportId};
use secp256k1::XOnlyPublicKey;
use std::collections::HashMap;
use std::collections::hash_map::Entry;
use std::sync::Mutex;
use tracing::warn;
/// Beacon protocol version.
pub const BEACON_VERSION: u8 = 0x01;
@@ -46,10 +49,33 @@ pub fn parse_beacon(data: &[u8]) -> Option<XOnlyPublicKey> {
XOnlyPublicKey::from_slice(&data[2..34]).ok()
}
/// Maximum distinct source MACs held between drains.
///
/// Beacons are unauthenticated broadcast frames, so anything on the segment
/// can name as many source MACs as it likes; without a bound the buffer grows
/// with the flood rate, and it is not drained at all while the transport is
/// not operational. This caps it at roughly a thousand small structs, tens of
/// kilobytes. Raising it costs that much more memory per transport; lowering
/// it risks truncating discovery on a very large segment. A thousand distinct
/// beaconing FIPS neighbors within one tick is far outside anything a real
/// deployment produces.
const MAX_BUFFERED_PEERS: usize = 1024;
/// Buffer for discovered peers, drained by `discover()`.
pub struct NeighborBuffer {
transport_id: TransportId,
peers: Mutex<Vec<DiscoveredPeer>>,
peers: Mutex<Buffered>,
}
/// Peers keyed by source MAC, plus the sighting order `take()` restores.
#[derive(Default)]
struct Buffered {
by_mac: HashMap<[u8; 6], (u64, DiscoveredPeer)>,
seq: u64,
/// Beacons refused for want of room, cumulative and never reset.
dropped: u64,
/// Cumulative drop count that earns the next log record.
warn_at: u64,
}
impl NeighborBuffer {
@@ -57,25 +83,82 @@ impl NeighborBuffer {
pub fn new(transport_id: TransportId) -> Self {
Self {
transport_id,
peers: Mutex::new(Vec::new()),
peers: Mutex::new(Buffered::default()),
}
}
/// Add a discovered peer from a received beacon.
pub fn add_peer(&self, src_mac: [u8; 6], pubkey: XOnlyPublicKey) {
let addr = TransportAddr::from_bytes(&src_mac);
let peer = DiscoveredPeer::with_hint(self.transport_id, addr, pubkey);
let mut peers = self.peers.lock().unwrap_or_else(|e| e.into_inner());
// Deduplicate by MAC address — keep the latest
peers.retain(|p| p.addr.as_bytes() != src_mac);
peers.push(peer);
///
/// Returns false when the beacon was refused because the buffer is full.
/// A MAC already buffered is always refreshed, so a flood of new MACs
/// cannot stop a known neighbor from being seen again.
pub fn add_peer(&self, src_mac: [u8; 6], pubkey: XOnlyPublicKey) -> bool {
let mut buffered = self.peers.lock().unwrap_or_else(|e| e.into_inner());
buffered.seq += 1;
let seq = buffered.seq;
let full = buffered.by_mac.len() >= MAX_BUFFERED_PEERS;
let stored = match buffered.by_mac.entry(src_mac) {
// Refreshing moves the MAC to the end, as retain-then-push did.
Entry::Occupied(mut slot) => {
slot.insert((seq, self.peer(src_mac, pubkey)));
true
}
Entry::Vacant(slot) if !full => {
slot.insert((seq, self.peer(src_mac, pubkey)));
true
}
Entry::Vacant(_) => false,
};
if !stored {
buffered.dropped += 1;
}
stored
}
/// Drain all discovered peers since the last call.
/// Drain all discovered peers since the last call, oldest sighting first.
pub fn take(&self) -> Vec<DiscoveredPeer> {
let mut peers = self.peers.lock().unwrap_or_else(|e| e.into_inner());
std::mem::take(&mut *peers)
let mut buffered = self.peers.lock().unwrap_or_else(|e| e.into_inner());
let mut ordered: Vec<(u64, DiscoveredPeer)> =
buffered.by_mac.drain().map(|(_, entry)| entry).collect();
// The reconcile layer spends a finite connect budget in this order, so
// which neighbor gets dialed must not depend on hash iteration order.
ordered.sort_unstable_by_key(|(seq, _)| *seq);
// Rate-limited: the drop rate is whatever the flooder chooses, and one
// record per drain would hand it the log volume too.
if buffered.dropped >= buffered.warn_at.max(1) {
warn!(
transport_id = %self.transport_id,
dropped = buffered.dropped,
cap = MAX_BUFFERED_PEERS,
"discovery buffer full, beacons from unseen neighbors refused"
);
buffered.warn_at = next_decade(buffered.dropped);
}
ordered.into_iter().map(|(_, peer)| peer).collect()
}
/// Beacons refused for want of room since this buffer was created.
pub fn dropped(&self) -> u64 {
self.peers.lock().unwrap_or_else(|e| e.into_inner()).dropped
}
/// Build the buffered peer record for one beacon.
fn peer(&self, src_mac: [u8; 6], pubkey: XOnlyPublicKey) -> DiscoveredPeer {
let addr = TransportAddr::from_bytes(&src_mac);
DiscoveredPeer::with_hint(self.transport_id, addr, pubkey)
}
}
/// Smallest power of ten strictly greater than `n`, saturating at `u64::MAX`.
fn next_decade(n: u64) -> u64 {
let mut threshold = 1u64;
while threshold <= n {
match threshold.checked_mul(10) {
Some(next) => threshold = next,
None => return u64::MAX,
}
}
threshold
}
// ============================================================================
@@ -163,4 +246,88 @@ mod tests {
let peers = buffer.take();
assert_eq!(peers.len(), 1);
}
/// Distinct MAC number `n`, for filling the buffer.
fn nth_mac(n: usize) -> [u8; 6] {
let bytes = (n as u64).to_be_bytes();
[0x02, bytes[3], bytes[4], bytes[5], bytes[6], bytes[7]]
}
#[test]
fn discovery_buffer_stops_buffering_past_the_cap() {
// The defect: an unauthenticated flood of source MACs grew the buffer
// without bound. Fails against the uncapped Vec, which returns all of
// them.
let buffer = NeighborBuffer::new(TransportId::new(1));
let pubkey = test_pubkey();
for n in 0..MAX_BUFFERED_PEERS + 50 {
buffer.add_peer(nth_mac(n), pubkey);
}
let peers = buffer.take();
assert_eq!(peers.len(), MAX_BUFFERED_PEERS);
// Drop-new keeps the earliest sightings.
assert_eq!(peers[0].addr.as_bytes(), &nth_mac(0));
}
#[test]
fn discovery_buffer_counts_dropped_beacons() {
let buffer = NeighborBuffer::new(TransportId::new(1));
let pubkey = test_pubkey();
for n in 0..MAX_BUFFERED_PEERS {
assert!(buffer.add_peer(nth_mac(n), pubkey));
}
for n in MAX_BUFFERED_PEERS..MAX_BUFFERED_PEERS + 7 {
assert!(!buffer.add_peer(nth_mac(n), pubkey));
}
assert_eq!(buffer.dropped(), 7);
}
#[test]
fn discovery_buffer_repeat_beacon_from_a_full_buffer_still_refreshes() {
let buffer = NeighborBuffer::new(TransportId::new(1));
let pubkey = test_pubkey();
for n in 0..MAX_BUFFERED_PEERS + 50 {
buffer.add_peer(nth_mac(n), pubkey);
}
// A neighbor already buffered must not be refused by a full buffer.
assert!(buffer.add_peer(nth_mac(0), pubkey));
let peers = buffer.take();
assert_eq!(peers.len(), MAX_BUFFERED_PEERS);
assert_eq!(peers[peers.len() - 1].addr.as_bytes(), &nth_mac(0));
}
#[test]
fn discovery_buffer_drain_preserves_last_seen_order() {
// A regression pin on the map rewrite rather than a test of the
// defect: retain-then-push already produced this order.
let buffer = NeighborBuffer::new(TransportId::new(1));
let pubkey = test_pubkey();
let a = [0xaa; 6];
let b = [0xbb; 6];
let c = [0xcc; 6];
buffer.add_peer(a, pubkey);
buffer.add_peer(b, pubkey);
buffer.add_peer(c, pubkey);
buffer.add_peer(a, pubkey);
let macs: Vec<_> = buffer
.take()
.iter()
.map(|p| p.addr.as_bytes().to_vec())
.collect();
assert_eq!(macs, vec![b.to_vec(), c.to_vec(), a.to_vec()]);
}
#[test]
fn next_decade_steps_by_powers_of_ten() {
assert_eq!(next_decade(0), 1);
assert_eq!(next_decade(1), 10);
assert_eq!(next_decade(9), 10);
assert_eq!(next_decade(10), 100);
assert_eq!(next_decade(u64::MAX), u64::MAX);
}
}
+9
View File
@@ -17,6 +17,7 @@ pub struct EthernetStats {
pub recv_errors: AtomicU64,
pub beacons_sent: AtomicU64,
pub beacons_recv: AtomicU64,
pub beacons_dropped: AtomicU64,
pub frames_too_short: AtomicU64,
pub frames_too_long: AtomicU64,
}
@@ -33,6 +34,7 @@ impl EthernetStats {
recv_errors: AtomicU64::new(0),
beacons_sent: AtomicU64::new(0),
beacons_recv: AtomicU64::new(0),
beacons_dropped: AtomicU64::new(0),
frames_too_short: AtomicU64::new(0),
frames_too_long: AtomicU64::new(0),
}
@@ -70,6 +72,11 @@ impl EthernetStats {
self.beacons_recv.fetch_add(1, Ordering::Relaxed);
}
/// Record a received beacon the discovery buffer had no room for.
pub fn record_beacon_dropped(&self) {
self.beacons_dropped.fetch_add(1, Ordering::Relaxed);
}
/// Take a snapshot of all counters.
pub fn snapshot(&self) -> EthernetStatsSnapshot {
EthernetStatsSnapshot {
@@ -81,6 +88,7 @@ impl EthernetStats {
recv_errors: self.recv_errors.load(Ordering::Relaxed),
beacons_sent: self.beacons_sent.load(Ordering::Relaxed),
beacons_recv: self.beacons_recv.load(Ordering::Relaxed),
beacons_dropped: self.beacons_dropped.load(Ordering::Relaxed),
frames_too_short: self.frames_too_short.load(Ordering::Relaxed),
frames_too_long: self.frames_too_long.load(Ordering::Relaxed),
}
@@ -104,6 +112,7 @@ pub struct EthernetStatsSnapshot {
pub recv_errors: u64,
pub beacons_sent: u64,
pub beacons_recv: u64,
pub beacons_dropped: u64,
pub frames_too_short: u64,
pub frames_too_long: u64,
}
+177 -5
View File
@@ -25,6 +25,18 @@ use tracing::{debug, info, trace, warn};
/// DNS cache TTL for hostname resolution (60 seconds).
const DNS_CACHE_TTL: Duration = Duration::from_secs(60);
/// Upper bound on the number of hostnames the DNS cache holds at once.
///
/// The cache is keyed by the address string a dial was asked for, and under a
/// rendezvous policy that accepts advertised endpoints those strings come from
/// remote parties, so without a bound the map grows for the life of the
/// process. 256 sits about two orders of magnitude above the number of
/// distinct hostnames a configured peer list produces, so no ordinary
/// deployment reaches it. Lowering it starts to be reachable by a large peer
/// list, and the only cost of an eviction is one extra DNS lookup on the next
/// dial of that name; raising it buys nothing but resident memory.
const DNS_CACHE_MAX_ENTRIES: usize = 256;
/// UDP transport for FIPS.
///
/// Provides connectionless, unreliable packet delivery over UDP/IP.
@@ -155,10 +167,8 @@ impl UdpTransport {
// Check cache
{
let cache = self.dns_cache.lock().unwrap_or_else(|e| e.into_inner());
if let Some((resolved, cached_at)) = cache.get(addr)
&& cached_at.elapsed() < DNS_CACHE_TTL
{
return Ok(*resolved);
if let Some(resolved) = cache_lookup(&cache, addr, Instant::now()) {
return Ok(resolved);
}
}
@@ -168,7 +178,13 @@ impl UdpTransport {
// Store in cache
{
let mut cache = self.dns_cache.lock().unwrap_or_else(|e| e.into_inner());
cache.insert(addr.clone(), (resolved, Instant::now()));
cache_store(
&mut cache,
addr.clone(),
resolved,
Instant::now(),
DNS_CACHE_MAX_ENTRIES,
);
}
Ok(resolved)
@@ -611,6 +627,56 @@ async fn udp_receive_loop(
}
}
/// A cached resolution for `key`, if one is present and still inside
/// `DNS_CACHE_TTL` at `now`.
fn cache_lookup(
cache: &HashMap<TransportAddr, (SocketAddr, Instant)>,
key: &TransportAddr,
now: Instant,
) -> Option<SocketAddr> {
cache
.get(key)
.filter(|(_, cached_at)| now.duration_since(*cached_at) < DNS_CACHE_TTL)
.map(|(resolved, _)| *resolved)
}
/// Record a resolution, keeping the cache at or below `cap` entries.
///
/// Refreshing a name already present never evicts anything. Otherwise every
/// entry past its TTL is dropped first, and only if that leaves the map full
/// is the oldest remaining entry evicted. Eviction is by insertion time rather
/// than by last use: the timestamp is already there as the TTL clock, and
/// tracking last use would mean writing to the map on the read path of every
/// dial. The sweep is linear in `cap` and runs only on a resolution miss, so
/// at most once per TTL per name.
fn cache_store(
cache: &mut HashMap<TransportAddr, (SocketAddr, Instant)>,
key: TransportAddr,
resolved: SocketAddr,
now: Instant,
cap: usize,
) {
if let Some(entry) = cache.get_mut(&key) {
*entry = (resolved, now);
return;
}
cache.retain(|_, (_, cached_at)| now.duration_since(*cached_at) < DNS_CACHE_TTL);
while cache.len() >= cap {
let Some(oldest) = cache
.iter()
.min_by_key(|(_, (_, cached_at))| *cached_at)
.map(|(key, _)| key.clone())
else {
break;
};
cache.remove(&oldest);
}
cache.insert(key, (resolved, now));
}
// ============================================================================
// Tests
// ============================================================================
@@ -621,6 +687,112 @@ mod tests {
use crate::transport::packet_channel;
use tokio::time::{Duration, timeout};
/// A distinct hostname key, so each store is a fresh entry.
fn dns_key(n: usize) -> TransportAddr {
TransportAddr::from(format!("host{n}.example:2121"))
}
fn dns_value() -> SocketAddr {
"198.51.100.1:2121".parse().unwrap()
}
/// The cache is keyed by strings a remote party can choose, so its size
/// has to be bounded no matter how many distinct names are dialed.
#[test]
fn dns_cache_store_refuses_to_exceed_the_cap() {
const CAP: usize = 8;
let now = Instant::now();
let mut cache = HashMap::new();
for n in 0..CAP + 5 {
cache_store(&mut cache, dns_key(n), dns_value(), now, CAP);
assert!(
cache.len() <= CAP,
"cache grew to {} entries past a cap of {CAP}",
cache.len()
);
}
}
/// A stale entry used to be overwritten on the next dial of the same name
/// and otherwise never removed, so a name dialed once sat there forever.
#[test]
fn dns_cache_store_evicts_entries_past_their_ttl() {
let now = Instant::now();
let expired_at = now.checked_sub(DNS_CACHE_TTL * 2).expect("monotonic clock");
let mut cache = HashMap::new();
cache.insert(dns_key(0), (dns_value(), expired_at));
cache_store(
&mut cache,
dns_key(1),
dns_value(),
now,
DNS_CACHE_MAX_ENTRIES,
);
assert!(
!cache.contains_key(&dns_key(0)),
"an entry past its TTL should be swept, not left to accumulate"
);
assert!(cache_lookup(&cache, &dns_key(0), now).is_none());
assert!(cache_lookup(&cache, &dns_key(1), now).is_some());
}
/// With nothing expired, the cap is enforced by dropping the oldest entry.
/// The ages here are all well inside the TTL, so the expiry sweep cannot
/// be what makes room and the eviction branch is the one under test.
#[test]
fn dns_cache_store_evicts_the_oldest_entry_when_every_entry_is_fresh() {
const CAP: usize = 4;
let now = Instant::now();
let mut cache = HashMap::new();
for n in 0..CAP {
let age = Duration::from_secs((CAP - n) as u64);
assert!(age < DNS_CACHE_TTL, "fixture must stay inside the TTL");
let cached_at = now.checked_sub(age).expect("monotonic clock");
cache.insert(dns_key(n), (dns_value(), cached_at));
}
assert_eq!(cache.len(), CAP, "no entry should be expired going in");
cache_store(&mut cache, dns_key(CAP), dns_value(), now, CAP);
assert_eq!(cache.len(), CAP);
assert!(
!cache.contains_key(&dns_key(0)),
"the oldest entry should be the one evicted"
);
for n in 1..=CAP {
assert!(
cache.contains_key(&dns_key(n)),
"entry {n} should have survived"
);
}
}
/// Re-resolving a name already cached is the common case on a live node.
/// It must not cost another entry its place.
#[test]
fn dns_cache_store_refreshing_an_existing_key_evicts_nothing() {
const CAP: usize = 4;
let now = Instant::now();
let mut cache = HashMap::new();
for n in 0..CAP {
let cached_at = now
.checked_sub(Duration::from_secs((CAP - n) as u64))
.expect("monotonic clock");
cache.insert(dns_key(n), (dns_value(), cached_at));
}
cache_store(&mut cache, dns_key(0), dns_value(), now, CAP);
assert_eq!(cache.len(), CAP);
for n in 0..CAP {
assert!(cache.contains_key(&dns_key(n)), "entry {n} should remain");
}
assert_eq!(cache_lookup(&cache, &dns_key(0), now), Some(dns_value()));
}
fn make_config(port: u16) -> UdpConfig {
UdpConfig {
bind_addr: Some(format!("127.0.0.1:{}", port)),
+66 -12
View File
@@ -134,17 +134,34 @@ impl HostMap {
///
/// If the file does not exist, returns an empty map (not an error).
/// Parse errors on individual lines are logged as warnings and skipped.
/// A read failure is logged and also yields an empty map; a caller that
/// must not mistake an unreadable file for an empty one uses
/// [`Self::try_load_hosts_file`] instead.
pub fn load_hosts_file(path: &Path) -> Self {
let contents = match std::fs::read_to_string(path) {
Ok(c) => c,
Err(e) if e.kind() == std::io::ErrorKind::NotFound => {
debug!(path = %path.display(), "No hosts file found, skipping");
return Self::new();
}
match Self::try_load_hosts_file(path) {
Ok(map) => map,
Err(e) => {
warn!(path = %path.display(), error = %e, "Failed to read hosts file");
return Self::new();
Self::new()
}
}
}
/// Load a host map from a hosts file, reporting read failures.
///
/// An absent file is a policy, not a fault: it resolves to an empty map
/// and `Ok`. Anything else — no read permission, an I/O error, non-UTF-8
/// content, or a `NotFound` that contradicts a successful stat and so
/// means the file is being rewritten under us — is returned as an error
/// so the caller can keep whatever it loaded last.
pub fn try_load_hosts_file(path: &Path) -> Result<Self, std::io::Error> {
let contents = match std::fs::read_to_string(path) {
Ok(c) => c,
Err(e) if e.kind() == std::io::ErrorKind::NotFound && file_mtime(path).is_none() => {
debug!(path = %path.display(), "No hosts file found, skipping");
return Ok(Self::new());
}
Err(e) => return Err(e),
};
let mut map = Self::new();
@@ -183,7 +200,7 @@ impl HostMap {
if !map.is_empty() {
info!(path = %path.display(), count = map.len(), "Loaded hosts file");
}
map
Ok(map)
}
/// Merge another host map into this one. The other map wins on conflicts.
@@ -224,8 +241,16 @@ impl HostMapReloader {
///
/// Performs the initial load of the hosts file and merges with the base map.
pub fn new(base: HostMap, path: std::path::PathBuf) -> Self {
let last_mtime = file_mtime(&path);
let hosts_file = HostMap::load_hosts_file(&path);
// A failed initial read records no mtime, so the next check sees a
// change and retries rather than treating the unread file as empty
// for the lifetime of the process.
let (last_mtime, hosts_file) = match HostMap::try_load_hosts_file(&path) {
Ok(map) => (file_mtime(&path), map),
Err(e) => {
warn!(path = %path.display(), error = %e, "Failed to read hosts file");
(None, HostMap::new())
}
};
let mut effective = base.clone();
effective.merge(hosts_file);
@@ -242,6 +267,11 @@ impl HostMapReloader {
&self.effective
}
/// Path of the hosts file this reloader tracks.
pub fn path(&self) -> &Path {
&self.path
}
/// Check if the hosts file has been modified and reload if so.
///
/// Returns `true` if the map was reloaded.
@@ -254,7 +284,32 @@ impl HostMapReloader {
// File appeared, disappeared, or was modified
self.last_mtime = current_mtime;
let hosts_file = HostMap::load_hosts_file(&self.path);
self.apply(HostMap::load_hosts_file(&self.path));
true
}
/// Check if the hosts file has been modified and reload if so, reporting
/// read failures.
///
/// On failure neither the recorded mtime nor the effective map is
/// touched, so the caller keeps its last-good state and the next call
/// retries. Returns `true` if the map was reloaded.
pub fn try_check_reload(&mut self) -> Result<bool, std::io::Error> {
let current_mtime = file_mtime(&self.path);
if current_mtime == self.last_mtime {
return Ok(false);
}
let hosts_file = HostMap::try_load_hosts_file(&self.path)?;
self.last_mtime = current_mtime;
self.apply(hosts_file);
Ok(true)
}
/// Replace the effective map with the base merged with a freshly read
/// hosts file.
fn apply(&mut self, hosts_file: HostMap) {
let mut new_effective = self.base.clone();
new_effective.merge(hosts_file);
@@ -266,7 +321,6 @@ impl HostMapReloader {
entries = count,
"Reloaded hosts file"
);
true
}
}
+16
View File
@@ -112,6 +112,22 @@ pub const FIPS_IPV6_OVERHEAD: u16 = 77;
/// the `MtuExceeded` signal — reaches it through this module.
pub use crate::proto::mmp::MIN_ACTIONABLE_PATH_MTU;
/// Smallest path MTU this node will act on when the claim arrives on the
/// unauthenticated reactive carrier, `MtuExceeded`.
///
/// Held equal to [`MIN_ACTIONABLE_PATH_MTU`] so no hop legitimately configured
/// with a small transport MTU loses reactive feedback. It is a separate
/// constant because the two carriers differ in what they prove: the
/// authenticated `PathMtuNotification` and the proof-carrying discovery
/// response come from a party this node has verified, whereas this one comes
/// from whoever could route a datagram here. What keeps a legal-but-forged
/// claim from pinning a session is corroboration against what this node has
/// actually sent, not this floor. Raising it (576 is the value the original
/// path-MTU floor design proposed, and derives an inner IPv6 MTU of 499)
/// bounds the outcome of an uncorroborated claim further, at the cost of
/// ignoring an honest report from any hop configured between the two values.
pub const MIN_REACTIVE_PATH_MTU: u16 = MIN_ACTIONABLE_PATH_MTU;
/// Calculate the effective IPv6 MTU for FIPS-encapsulated traffic.
///
/// Given a transport MTU (e.g., UDP payload size), returns the maximum
+4 -2
View File
@@ -1,8 +1,10 @@
//! Utility modules.
//!
//! Shared infrastructure that doesn't belong to a specific protocol layer:
//! session index allocation, Unix socket binding, and other cross-cutting
//! concerns.
//! session index allocation, Unix socket binding and its permission
//! primitives, and other cross-cutting concerns.
pub mod index;
pub mod sockbind;
#[cfg(unix)]
pub mod sockperm;
+25 -2
View File
@@ -48,7 +48,12 @@ pub fn bind(path: &Path, what: &str) -> Result<UnixListener, std::io::Error> {
remove_stale_socket(path, what)?;
}
let listener = UnixListener::bind(path)?;
// Bound through `sockperm` rather than `UnixListener::bind` directly:
// bind(2) creates the socket inode as `0777 & !umask`, so under a
// permissive umask it is world-accessible for the window between the
// bind and the `set_socket_access` chmod below. That chmod stays the
// authority on the final mode; this only closes the window.
let listener = crate::utils::sockperm::bind(path)?;
set_socket_access(path, managed_parent.as_deref(), chown_to_fips_group)?;
@@ -65,11 +70,21 @@ pub fn bind(path: &Path, what: &str) -> Result<UnixListener, std::io::Error> {
/// socket's private directory.
#[cfg(unix)]
fn ensure_socket_parent(parent: &Path) -> Result<bool, std::io::Error> {
use std::os::unix::fs::DirBuilderExt;
if parent.as_os_str().is_empty() {
return Ok(false);
}
match std::fs::create_dir(parent) {
// Mode carried on creation rather than applied afterwards. A directory
// made by plain `create_dir` is `0777 & !umask`, and an intermediate
// ancestor is never chmodded by anything below, so under a permissive
// umask it would stay world-writable for the life of the host — and a
// world-writable parent lets an unprivileged account plant an entry at
// the socket path. 0750 is what `set_socket_access` applies to a managed
// parent anyway, and what the systemd unit and FreeBSD rc script already
// use.
match std::fs::DirBuilder::new().mode(0o750).create(parent) {
Ok(()) => Ok(true),
Err(error) if error.kind() == std::io::ErrorKind::AlreadyExists => {
if parent.is_dir() {
@@ -120,6 +135,14 @@ fn set_socket_access(
/// If the file exists but no one is listening, remove it so we can bind. This
/// handles unclean daemon exits. A live listener yields `AddrInUse` instead, so
/// two daemons cannot silently take the same path.
///
/// The gap between the connect probe and the bind that follows is accepted
/// rather than closed. Reaching it needs write access to the socket's parent
/// directory, which the packaged layouts give to root alone (0750 and
/// root-owned under both systemd and the FreeBSD rc script), and an account
/// holding it can deny the daemon its socket more simply by squatting the path
/// before the daemon starts. The removal itself unlinks a symlink rather than
/// its target, so it is not an arbitrary delete.
#[cfg(unix)]
fn remove_stale_socket(path: &Path, what: &str) -> Result<(), std::io::Error> {
match std::os::unix::net::UnixStream::connect(path) {
+157
View File
@@ -0,0 +1,157 @@
//! Permission-safe creation of Unix domain sockets and the directories
//! holding them.
//!
//! The socket inode and its parent directory are created with a mode the
//! process umask can only tighten, rather than created wide and narrowed
//! afterwards. The caller's own chmod and chown stay where they are and
//! remain the authority on the socket's final mode; this closes the window
//! between creation and that fix-up, and the case of an intermediate
//! directory that nothing fixes up at all.
use std::path::Path;
use tokio::net::UnixListener;
/// Mode for a directory this module creates to hold a control socket.
///
/// Matches what the packaging already applies (systemd's
/// `RuntimeDirectoryMode=0750`, `install -d -m 0750` in the FreeBSD rc
/// script), so no packaged deployment sees a different directory mode than
/// it does today. Widening it would expose the socket path to accounts that
/// cannot reach it now; the umask can still tighten it further.
const SOCKET_DIR_MODE: u32 = 0o750;
/// umask held across the socket bind.
///
/// `bind(2)` creates the socket inode with `0777 & !umask`, so under a
/// permissive umask the socket is world-accessible until the chmod that
/// follows it. Masking the "other" bits makes the inode 0770 at creation,
/// which is the mode the caller applies a moment later anyway. Changing
/// this changes the mode the socket is created with, not the mode it ends
/// up with.
const BIND_UMASK: libc::mode_t = 0o007;
/// Restores the process umask when dropped.
struct UmaskGuard(libc::mode_t);
impl UmaskGuard {
/// Install `mask` as the process umask, remembering the previous one.
fn tighten(mask: libc::mode_t) -> Self {
// SAFETY: umask(2) cannot fail and touches only process state.
Self(unsafe { libc::umask(mask) })
}
}
impl Drop for UmaskGuard {
fn drop(&mut self) {
// SAFETY: as above; restoring the mask this guard replaced.
unsafe {
libc::umask(self.0);
}
}
}
/// Create the directory that will hold a socket, and any missing ancestors.
///
/// Directories come out 0750 rather than `0777 & !umask`. Nothing chmods an
/// intermediate directory afterwards, so one created under a permissive
/// umask would stay world-writable for the life of the host, and a
/// world-writable parent lets an unprivileged account plant an entry at the
/// socket path.
pub fn make_parent(parent: &Path) -> Result<(), std::io::Error> {
use std::os::unix::fs::DirBuilderExt;
std::fs::DirBuilder::new()
.recursive(true)
.mode(SOCKET_DIR_MODE)
.create(parent)
}
/// Bind a Unix listener whose inode is never world-accessible.
///
/// The umask is process-global, so it is held across the bind alone. It
/// only clears bits, so anything else created inside that window comes out
/// more restrictive, never less.
pub fn bind(path: &Path) -> Result<UnixListener, std::io::Error> {
let _umask = UmaskGuard::tighten(BIND_UMASK);
UnixListener::bind(path)
}
#[cfg(test)]
mod tests {
use super::*;
use std::os::unix::fs::PermissionsExt;
use std::sync::Mutex;
/// The umask is process-global, so the tests that set it run one at a
/// time. This does not serialize against the rest of the test binary;
/// the mask used is 0o022, the ordinary default, so a file another test
/// creates in the window is unaffected.
static UMASK_LOCK: Mutex<()> = Mutex::new(());
/// Take the umask lock, ignoring poisoning: a test that fails while
/// holding it must not turn its siblings red for an unrelated reason.
fn umask_lock() -> std::sync::MutexGuard<'static, ()> {
UMASK_LOCK.lock().unwrap_or_else(|e| e.into_inner())
}
/// Read the current umask, which is only observable by replacing it.
fn current_umask() -> libc::mode_t {
// SAFETY: umask(2) cannot fail; the value read is put straight back.
unsafe {
let old = libc::umask(0o022);
libc::umask(old);
old
}
}
fn mode_of(path: &Path) -> u32 {
std::fs::symlink_metadata(path)
.unwrap()
.permissions()
.mode()
}
#[tokio::test]
async fn socket_is_created_without_other_access_under_a_permissive_umask() {
let _lock = umask_lock();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("control.sock");
let restore = UmaskGuard::tighten(0o022);
let listener = bind(&path).unwrap();
drop(restore);
assert_eq!(mode_of(&path) & 0o007, 0);
drop(listener);
}
#[tokio::test]
async fn socket_bind_leaves_the_process_umask_as_it_found_it() {
let _lock = umask_lock();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("control.sock");
let restore = UmaskGuard::tighten(0o022);
let listener = bind(&path).unwrap();
let after = current_umask();
drop(restore);
assert_eq!(after, 0o022);
drop(listener);
}
#[test]
fn socket_parent_and_its_ancestors_are_created_without_other_access() {
let _lock = umask_lock();
let dir = tempfile::tempdir().unwrap();
let intermediate = dir.path().join("run");
let parent = intermediate.join("fips");
let restore = UmaskGuard::tighten(0o022);
make_parent(&parent).unwrap();
drop(restore);
assert_eq!(mode_of(&intermediate) & 0o007, 0);
assert_eq!(mode_of(&parent) & 0o007, 0);
}
}