mirror of
https://github.com/Routstr/routstrd.git
synced 2026-10-05 20:38:22 +00:00
Problem
-------
0ce4c07 pruned expired pending mint quotes purely locally at startup:
any quote past its bolt11 expiry with no recorded PAID/ISSUED
observation was failed without contacting its mint. That invariant is
only forward-looking: a quote can be paid before expiry while the
daemon is down, leaving no local observation behind. Failing such a
quote strands the paid funds at the mint: failed operations are skipped
by recoverPendingMintOperations(), so the claimable proofs are never
claimed.
Concrete case: receiveBolt11 invoice created, daemon stops, user pays
within expiry, daemon restarts after expiry: the old prune failed the
op without ever asking the mint.
Change
------
Replace failExpiredMintsLocally() with settleExpiredMintQuotes(),
which adds one bounded observation round before any local fail:
1. Select expired, unobserved pending mint quotes as before
(selectCleanupOperations, minAgeMs 0).
2. For each candidate, ask its mint for the quote state via
MintOperationService.observePendingOperation() (the same check the
mint sweep uses, reached through the existing structural cast)
under a shared 15s wall-clock budget
(EXPIRED_MINT_OBSERVATION_DEADLINE_MS).
3. Act on the answer from the mint:
- UNPAID ("waiting"): the expired quote can never be issued, so
failing it locally cannot strand funds; failPendingOperation().
- PAID/ISSUED ("ready"/"completed"): leave pending; the mint
recovery sweep (or the processor, via the emitted
mint-op:quote-state-changed event) finalizes it and claims the
proofs.
- unreachable/slow mint or unknown quote: leave pending so a later
startup can still recover it. Nothing is failed without a mint
confirmation.
Why a deadline
--------------
coco-core issues mint requests via bare fetch() with no timeout, so a
hung mint could otherwise stall this phase (and with it the recovery
promise that gates value-moving operations) for minutes. The shared
budget caps the whole round at 15s; the unobserved remainder stays
pending and is handled by the normal sweep (background, per-op
contained).
Why not keep the blind local fail
---------------------------------
The mint sweep treats UNPAID as "waiting" and never fails expired
quotes itself, so some form of pruning is still required to keep
recovery quick on wallets with many dead quotes. The observation round
keeps that property: confirmed-unpaid quotes are failed before the
sweep and never contacted again, while the unsafe case (paid before
expiry, never observed) now goes through normal recovery.
Side effects
------------
- Asking the mint also closes the narrower race from the old flow
(watcher records PAID between selection and fail): quotes are now
failed only when the mint currently reports UNPAID past expiry.
- Recovery phase strings are now "Settling/Settled expired mint
quotes"; settlement counts are logged to the startup stream.
- The explicit wallet cleanup command keeps its local-only semantics:
it is user-invoked, supports dry-run, and defaults to a 7-day
minimum age, giving ample observation opportunity beforehand.
Testing
-------
- New settleExpiredMintQuotes unit tests (6): mint-confirmed unpaid is
failed locally; PAID/ISSUED is left for recovery; unreachable mint
is left pending; hung mint is bounded by the shared deadline;
unexpired/observed quotes untouched.
- bun run lint (tsc --noEmit) passes.
- bun run build passes.
- Wallet/cleanup tests pass (56/56).
- Full bun test shows one pre-existing, unrelated failure
(mergeHermesConfig) that also fails on the parent commit.