mirror of
https://github.com/jmcorgan/fips.git
synced 2026-08-09 08:14:42 +00:00
The Phase 1, Phase 3, and Phase 5 strict asserts each fire a single ping per directed pair. Under low-level packet loss (e.g. 1% i.i.d. per-direction loss from a CI runner under pressure), a single-shot round-trip fails at ~2% per pair, so a 20-pair strict assert misses with probability 1 - (0.98)^20 = ~33% per phase from ICMP noise alone, well above the routing-state signal the asserts are meant to catch. ping_one gains a max_attempts parameter (default 1, preserving existing call sites). On failure it retries up to MAX_PING_ATTEMPTS-1 additional times with PING_RETRY_DELAY seconds between attempts. Per-pair worst case under the defaults (4 attempts, 1 s spacing, 5 s ping6 -W timeout) is 4*5 + 3 = 23 s; per-rep worst case scales with the failing-pair count. Successful retries log "OK (RTT, attempt N)"; exhausted retries log "FAIL (after N attempts)". The retry budget is wired into all three strict asserts: - Phase 1 final ping_all (after wait_for_full_baseline converges) - Phase 3 ping_all (post-first-rekey) - Phase 5 ping_all (post-second-rekey) The wait_for_full_baseline convergence loop itself stays single-shot. Its job is to detect when the mesh first sees a fully clean 20-pair batch, and retries inside the loop would conflate transient ping loss with still-converging routing state. No daemon code changes.