Files
ngit-grasp/tests
DanConwayDev bccbb54b62 fix(sync): require stable recovery before resetting backoff
Production showed nostr-pub.wellorder.net repeatedly completing a WebSocket handshake, resetting under normal REQ load 30-42 seconds later, and returning to the five-second base reconnect delay. The health tracker cleared its failure streak at handshake time even though its state model and documentation already required five minutes of stable recovery.

Preserve connection-failure history across recovered handshakes, promote a connected relay only after the existing five-minute stability period, and remove the redundant quick-reconnect success record. A relay continues normal live and historic work while degraded; only its next reconnect delay escalates. Active rate-limit cooldown accounting remains independent.

A WebSocket/REQ fixture reproduces the short-lived-success lifecycle and verifies the recovered session retains its failure count. Unit tests cover escalated backoff, stable promotion, and the rule that disconnected relays cannot age into recovery.

Correctness assumes an uninterrupted five-minute session under normal workload is sufficient evidence to forgive the streak. Intermittent sessions remain one instability episode, including for the existing 24-hour dead-relay threshold. Changing workload pace, stability duration, backoff configuration, or relay subscription policy is deliberately excluded.

Validated with cargo test --lib (650 passed), cargo test --test sync (88 passed, 1 ignored), the focused reconnect scenario, and nix build .#ngit-grasp.
2026-08-07 16:25:21 +00:00
..