Two tests entered patch("routstr.auth.calculate_cost", ...) inside
concurrently gathered tasks. Interleaved patch exits restore in the
wrong order, leaving the mock permanently installed for every later
test in the session. Hoist the patch around the gather.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The USD-cost path (and the litellm pricing fallback) resolved the
provider fee via get_provider_for_model(model_id)[0] — the best-ranked
provider for the alias, not the one that served. Settlement callers in
the upstream handlers now pass their own provider_fee through
adjust_payment_for_tokens / get_x_cashu_cost into calculate_cost; the
string-derived fallback remains for callers without a serving provider.
Configured model pricing is unaffected (the fee is already baked into
cached pricing).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
X-Cashu handlers do not rewrite the upstream's echoed model string, so
get_x_cashu_cost previously priced whatever wire name the upstream
reported — the most collapse-prone alias lookup of all. The routed Model
is now threaded from forward_x_cashu_request through the chat and
Responses handler chains (and the litellm messages path) into
get_x_cashu_cost, so cost and refund are computed from the model that
actually served.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Settlement previously re-derived pricing from the response's model string
through the alias map, which resolves to the best-ranked candidate for
that alias — not necessarily the provider/model that actually served the
request. adjust_payment_for_tokens and calculate_cost now accept the
routed Model and bill its pricing directly; the string lookup remains as
a warning fallback for callers without routed identity (e.g. the generic
streaming finalizer).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
PPQ.AI (BYOK) requests were billed at ~5% of their true cost because
_resolve_usd_cost fell through to usage.cost (a small BYOK routing fee)
instead of using cost_details.upstream_inference_cost (the real inference
cost). The proxy operator absorbed the inference cost.
The fix adds a BYOK-specific branch in _resolve_usd_cost: when is_byok is
true and cost_details.upstream_inference_cost is present, bill
upstream_inference_cost + byok_fee — what PPQ actually deducts from the
balance. Non-BYOK providers (e.g. OpenRouter) are unaffected because their
usage.cost already equals upstream_inference_cost.
Regression tests mirror the live glm-5.2-fast request from GitHub issue #615,
asserting the corrected billing (940,274 msats vs the old 45,202 msats — a
20.8× undercharge).
Closes#615
- Case-insensitive comparison of served model vs forwarded_model_id
(matches get_model_instance's lowercasing semantics)
- Suppress spurious mismatch when get_model_instance resolves back to
the same model (e.g. date-versioned glm-5-2-20260415 -> glm-5-2)
- Preserve unknown-model warning only when lookup returns None
- Tests: case-insensitive match + date-versioned alias resolution
Tinfoil PR #385 added model=<name> to the X-Tinfoil-Usage-Metrics
header/trailer. This commit uses that field for accurate billing.
Changes in routstr/upstream/ehbp.py:
- parse_tinfoil_usage_metrics(): extract the model= field as a string
alongside the existing token counts (previously silently discarded
because int() failed on it).
- _build_cost_info(): accept optional actual_model parameter propagated
through to callers when a real mismatch is detected.
- _compute_ehbp_actual_cost(): compare the served model against
model_obj.forwarded_model_id (the expected upstream ID) rather than
model_obj.id (the client-facing alias). This prevents spurious
mismatches when a node runner maps e.g. tinfoil-glm-5-2 -> glm-5-2
and the header correctly reports glm-5-2. On a genuine mismatch
(failover to a different upstream model), look up the actual model's
pricing via get_model_instance() (forwarded_model_id values are
registered as routable aliases in the global model map).
- forward_ehbp_request() / forward_ehbp_x_cashu_request(): use the
actual served model for payment finalization and logging when a
mismatch is detected.
Tests: 6 new scenarios (alias match, real mismatch with alias, unknown
model fallback, old-format compat, cache token details + model), plus
forwarded_model_id set on all existing mock model objects to keep them
passing. All 49 Tinfoil/EHBP unit tests pass.