fips-link-layer.md:
- Rewrite Liveness Detection: explicit Heartbeat (0x51) with 10s interval
and 30s dead timeout replaces vague gossip-as-heartbeat description
- Add Auto-Reconnect section: MMP dead timeout triggers retry with
unlimited backoff for auto_reconnect peers
- Add Handshake Message Retry section: link + session layer resend with
exponential backoff within timeout window
- Add Heartbeat to Link Message Types table
- Update Implementation Status with three new implemented features
fips-configuration.md:
- Add handshake_resend_interval_ms, handshake_resend_backoff,
handshake_max_resends to rate_limit table
- Add heartbeat_interval_secs, link_dead_timeout_secs to general table
- Add peers[].auto_reconnect to peers table
- Note auto-reconnect bypasses max_retries in retry section
- Update complete reference YAML with all new parameters
fips-wire-formats.md:
- Rename 0x51 from reserved Keepalive to implemented Heartbeat
- Update Disconnect reason 0x07 to Heartbeat liveness timeout
testing/chaos/README.md:
- Add runner.log to output files
- Add Directed Outbound Configs subsection
Auto-reconnect:
- Add per-peer auto_reconnect config (default true) to PeerConfig
- schedule_reconnect() feeds removed peers back into retry system with
unlimited retries and exponential backoff after MMP dead timeout
- RetryState gains reconnect flag to distinguish startup retries
(max_retries-limited) from auto-reconnect (unlimited)
Retry re-fire fix:
- process_pending_retries() now pushes retry_after_ms past the handshake
timeout window after successful initiate_peer_connection(), preventing
retries from firing every tick with no backoff
Chaos sim improvements:
- Directed outbound configs: BFS spanning tree + lower-ID-first assignment
eliminates dual-connect race conditions in simulation
- Save runner log (runner.log) alongside per-node logs for event correlation
- Increase churn-20 traffic aggressiveness and node churn (max_down_nodes
3→5, traffic interval min 0s, duration max 90s, concurrent flows 5→10)