mirror of
https://github.com/jmcorgan/fips.git
synced 2026-07-30 19:46:15 +00:00
node: periodically re-broadcast TreeAnnounce on no-change in check_periodic_parent_reeval
Closes the eventually-consistent gap in spanning-tree state distribution. Every existing send_tree_announce_to_all call site gates on a local state-change event (parent switch, self-root promotion, ancestry change, peer promotion, parent loss). Once a partition latches — for example a parent-switch announce stranded in the brief cross-init handshake swap window, where the announce arrives on a session-index whose decrypt-worker entry has been unregistered — neither side's state changes again, so neither side ever re-broadcasts. The existing 60 s check_periodic_parent_reeval was a re-evaluation, not a re-broadcast: it short-circuited silently on no-change. Production-side healing depended on incidental link churn; lab harnesses with stable docker-bridge links had no equivalent path. Add a final else branch that fires send_tree_announce_to_all unconditionally on the no-change path, alongside the existing switch and self-promote arms. Receivers coalesce by sequence comparison (ParentDeclaration::is_fresher_than) and short-circuit at the `if !updated` gate in handle_tree_announce; same-sequence repeats drop silently with no cascade. The per-peer 500 ms rate-limiter is well below this 60 s cadence and does not suppress the heartbeat broadcast. The fix is a general protocol-robustness improvement: it addresses any in-flight TreeAnnounce loss class, not only the specific cross-init swap-window drop site. testing/static/scripts/rekey-test.sh BASELINE_CONVERGENCE_TIMEOUT 60 -> 65 so a partition healed by the periodic broadcast at T+60 lands inside the convergence window. wait_for_full_baseline early-exits on PASS, so successful reps see no extra wall-clock.
This commit is contained in:
@@ -145,8 +145,14 @@ fi
|
||||
# ── Full test ─────────────────────────────────────────────────────────
|
||||
trap 'echo ""; echo "Test interrupted"; exit 130' INT
|
||||
|
||||
# Wait times derived from rekey timer
|
||||
BASELINE_CONVERGENCE_TIMEOUT=60
|
||||
# Wait times derived from rekey timer.
|
||||
# BASELINE_CONVERGENCE_TIMEOUT must cover one full daemon
|
||||
# node.tree.reeval_interval_secs (default 60) plus a small margin
|
||||
# so any partition that only heals via the periodic TreeAnnounce
|
||||
# re-broadcast lands inside the convergence window. wait_for_full_baseline
|
||||
# early-exits on PASS, so successful reps are unaffected by the
|
||||
# extra headroom.
|
||||
BASELINE_CONVERGENCE_TIMEOUT=65
|
||||
REKEY_SETTLE=12 # > DRAIN_WINDOW_SECS (10) so post-rekey samples are off the old session
|
||||
# First FMP rekey should follow shortly after the 35s interval once the mesh is
|
||||
# fully converged. Keep this bounded to preserve a meaningful scheduling check
|
||||
|
||||
Reference in New Issue
Block a user