mirror of
https://github.com/jmcorgan/fips.git
synced 2026-07-30 19:46:15 +00:00
Tune overly aggressive discovery rate limiting
The default 30s post-failure backoff (300s cap, doubling per consecutive failure) was set to bound traffic from chatty apps looking up unreachable targets, but in practice it dominates cold-start mesh convergence: a single timed-out lookup during initial bloom-filter propagation suppresses any retry for 30s, and the existing reset triggers (parent change, new peer, first RTT, reconnection) don't fire on a stable post-handshake topology. The suppression window winds up dictating the protocol's effective time-to-converge instead of bounding repeat traffic. Replaces the single-lookup-with-internal-retry model (`timeout_secs`/`retry_interval_secs`/`max_attempts`) with a per-attempt timeout sequence in `node.discovery.attempt_timeouts_secs`, defaulting to `[1, 2, 4, 8]`. Each attempt sends a fresh LookupRequest with a new random request_id so successive attempts can take different forwarding paths as the bloom and tree state evolve. The destination is declared unreachable only after the sequence is exhausted (15s total at the default). Disables post-failure suppression by default (`backoff_base_secs`/ `backoff_max_secs` now `0`/`0`). The `DiscoveryBackoff` machinery stays in tree (inert at zero base/cap); operators with chatty apps generating repeat lookups against unreachable destinations can opt back in. `PendingLookup` field shape unchanged so the control-socket `show_routing` JSON (`pending_lookups[].attempt`/`initiated_ms`/ `last_sent_ms`) keeps the same schema for fipstop and external consumers; `last_sent_ms` now means "current-attempt start" under the new state machine.
This commit is contained in:
@@ -165,12 +165,10 @@ Controls bloom-guided node discovery (LookupRequest/LookupResponse).
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `node.discovery.ttl` | u8 | `64` | Hop limit for LookupRequest forwarding |
|
||||
| `node.discovery.timeout_secs` | u64 | `10` | Lookup completion timeout |
|
||||
| `node.discovery.attempt_timeouts_secs` | array<u64> | `[1, 2, 4, 8]` | Per-attempt timeouts. Each entry is the deadline for one `LookupRequest` before sending the next attempt with a fresh `request_id`. Length determines total attempt count; default gives 4 attempts and a 15s total budget |
|
||||
| `node.discovery.recent_expiry_secs` | u64 | `10` | Dedup cache expiry for recent request IDs |
|
||||
| `node.discovery.retry_interval_secs` | u64 | `5` | Retry interval within the timeout window; after this interval without a response, resend the lookup |
|
||||
| `node.discovery.max_attempts` | u8 | `2` | Max attempts per lookup (1 = no retry, 2 = one retry) |
|
||||
| `node.discovery.backoff_base_secs` | u64 | `30` | Base for exponential backoff after lookup failure; doubles per consecutive failure |
|
||||
| `node.discovery.backoff_max_secs` | u64 | `300` | Cap on exponential backoff (5 minutes) |
|
||||
| `node.discovery.backoff_base_secs` | u64 | `0` | Optional post-failure suppression base in seconds; doubles per consecutive failure. `0` disables (default) — the per-attempt sequence is the only retry pacing |
|
||||
| `node.discovery.backoff_max_secs` | u64 | `0` | Cap on optional post-failure backoff |
|
||||
| `node.discovery.forward_min_interval_secs` | u64 | `2` | Transit-side rate limiting: minimum interval between forwarded lookups for the same target |
|
||||
|
||||
### Spanning Tree (`node.tree.*`)
|
||||
@@ -707,12 +705,10 @@ node:
|
||||
identity_size: 10000
|
||||
discovery:
|
||||
ttl: 64
|
||||
timeout_secs: 10
|
||||
attempt_timeouts_secs: [1, 2, 4, 8]
|
||||
recent_expiry_secs: 10
|
||||
retry_interval_secs: 5
|
||||
max_attempts: 2
|
||||
backoff_base_secs: 30
|
||||
backoff_max_secs: 300
|
||||
backoff_base_secs: 0
|
||||
backoff_max_secs: 0
|
||||
forward_min_interval_secs: 2
|
||||
tree:
|
||||
announce_min_interval_ms: 500
|
||||
|
||||
@@ -385,19 +385,31 @@ The default `max_attempts` is 2 (initial + one retry). Each retry generates
|
||||
a fresh `request_id` and re-evaluates bloom filter matches, so it can take
|
||||
a different path if the tree has restructured.
|
||||
|
||||
### Originator Backoff
|
||||
### Per-Attempt Timeouts
|
||||
|
||||
After a lookup times out or no peer's bloom filter contains the target, the
|
||||
originator enters **exponential backoff** before re-attempting discovery for
|
||||
the same target:
|
||||
Each discovery is a sequence of attempts with growing per-attempt timeouts.
|
||||
Default sequence is `[1s, 2s, 4s, 8s]` (configurable via
|
||||
`node.discovery.attempt_timeouts_secs`). When the current attempt's deadline
|
||||
elapses without a `LookupResponse`, the originator sends another
|
||||
`LookupRequest` with a **fresh `request_id`** and the next entry in the
|
||||
sequence as its deadline. Fresh request_ids let each attempt take a
|
||||
different forwarding path as the bloom and tree state evolve, which is
|
||||
particularly useful during cold-start convergence. The destination is
|
||||
declared unreachable only after the full sequence is exhausted (15s total
|
||||
with the default).
|
||||
|
||||
- **Base delay**: 30s (configurable via `backoff_base_secs`)
|
||||
- **Multiplier**: 2x per consecutive failure
|
||||
- **Cap**: 300s (configurable via `backoff_max_secs`)
|
||||
### Originator Backoff (optional, off by default)
|
||||
|
||||
Backoff is **reset on topology changes** that might make previously
|
||||
unreachable targets reachable: parent switch, new peer connection, first
|
||||
RTT measurement from MMP, or peer reconnection.
|
||||
After the per-attempt sequence is exhausted, the originator can additionally
|
||||
suppress further fresh lookups for the same target with exponential
|
||||
post-failure backoff. This is **disabled by default** (`backoff_base_secs:
|
||||
0`); the per-attempt sequence is the only retry pacing in the standard
|
||||
configuration. Operators may opt in via `node.discovery.backoff_base_secs`
|
||||
and `node.discovery.backoff_max_secs` if their deployment has chatty apps
|
||||
generating repeated lookups for genuinely unreachable destinations. When
|
||||
enabled, backoff is **reset on topology changes** that might make
|
||||
previously unreachable targets reachable: parent switch, new peer
|
||||
connection, first RTT measurement from MMP, or peer reconnection.
|
||||
|
||||
### Bloom Filter Pre-Check
|
||||
|
||||
|
||||
Reference in New Issue
Block a user