test(marmot): let the disband test converge instead of snapshotting

Test 29 asserted that wn reaches amy's post-disband epoch within one fixed
window. That is stricter than the protocol promises, and it failed on a
full-suite run for a case `group-lifecycle-v1.md` explicitly allows.

MDK rotates its own leaf shortly after joining, so it can commit between the
epoch-agreement gate and amy's disband — and then the two have forked. What the
spec guarantees from there is not "the disband lands first time" but that the
REQUEST survives, is regenerated against whichever branch was selected, and
lands eventually. The old assertion could only pass in the race-free case, and
a busier machine widens the race.

The loop now re-reads AMY's epoch each round, because regeneration advances it,
and drives amy's own sync, which is what carries a pass to settlement and
re-issues a disband that lost. It also adds a check the old version lacked: once
the two agree, amy must still read the group as disbanded. "Epochs agree but the
group is live" is the outcome actually worth catching, and counting epochs alone
never would have.

Evidence this is a race and not a regression from the audit fixes: the two
commits in the failing window carry different `h` tags, so they are commits in
two different groups rather than a fork, and the test passes in isolation
(`epoch 1 -> 2, wn at 2`). A full run with this change is the confirmation and
is not in yet.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016kCuA6tc4JQzHPCDd39GHq
This commit is contained in:
Claude
2026-09-11 03:11:13 +00:00
parent 025116acc8
commit 3a1947e733
+23 -4
View File
@@ -678,17 +678,36 @@ test_29_disband_amy_to_wn() {
record_result "$id" fail "amy's epoch did not advance past $before_epoch"; return
fi
local deadline=$(( $(date +%s) + 150 )) accepted=0 saw=""
# Converge, don't snapshot. MDK rotates its own leaf shortly after joining,
# so it can commit between the epoch gate above and the disband below — and
# then the two have forked. The protocol's answer is not "the disband lands
# first time": it is that the REQUEST survives, is regenerated against
# whichever branch was selected, and lands eventually. Asserting the
# race-free happy path made this test fail on a busy machine for a case the
# spec explicitly allows, so the loop re-reads amy's epoch each round —
# regeneration moves it — and drives amy's own sync, which is what carries a
# pass to settlement and re-issues a disband that lost.
local deadline=$(( $(date +%s) + 240 )) accepted=0 saw="" amy_now="$after_epoch"
while [[ $(date +%s) -lt $deadline ]]; do
amy_now=$(amy_json marmot group show "$gid" 2>/dev/null | jq -r '.epoch // empty')
saw=$(wn_b_json groups show "$mls_gid" 2>/dev/null | jq -r '.result.mls.epoch // empty')
if [[ -n "$saw" && "$saw" == "$after_epoch" ]]; then accepted=1; break; fi
if [[ -n "$saw" && -n "$amy_now" && "$saw" == "$amy_now" ]]; then accepted=1; break; fi
wn_b sync >/dev/null 2>&1 || true
sleep 5
done
printf 'disband29 epoch %s -> %s, wn at %s\n' "$before_epoch" "$after_epoch" "${saw:-<none>}" >>"$LOG_FILE"
printf 'disband29 epoch %s -> %s (amy now %s), wn at %s\n' \
"$before_epoch" "$after_epoch" "${amy_now:-?}" "${saw:-<none>}" >>"$LOG_FILE"
if [[ "$accepted" -ne 1 ]]; then
record_result "$id" fail "wn stayed at epoch ${saw:-<none>} instead of amy's $after_epoch — it rejected the disband commit"
record_result "$id" fail "wn stayed at epoch ${saw:-<none>} while amy is at ${amy_now:-?} — the disband never converged"
return
fi
# Converged — and amy must still read the group as ended. A disband that
# lost its branch and was never regenerated would agree on an epoch here
# while leaving the group live, which is the failure worth catching.
if [[ "$(amy_json marmot group show "$gid" 2>/dev/null | jq -r '.disbanded // false')" != "true" ]]; then
record_result "$id" fail "amy and wn agree on epoch ${saw} but the group is not disbanded"
return
fi
record_result "$id" pass