mirror of
https://github.com/Routstr/routstr-core.git
synced 2026-10-05 12:28:22 +00:00
The cache discount still did not apply after 3c813da because the EHBP
finalization path never handed the serving ``Model`` to ``calculate_cost``:
it passed only the model *string* (``{"model": pricing_model_id, ...}``),
so ``_get_pricing_rates`` fell into its "settling without routed model
identity" branch and re-derived pricing through the global alias map.
That lookup resolves the string to the *best-ranked* candidate for the
id, not the serving one. Tinfoil's catalog id (``deepseek-v4-1-flash``)
is also a bare cross-provider alias, and the cheaper cross-provider
candidate carries ``input_cache_read=0``. ``_get_pricing_rates`` treats a
zero cache rate as "missing" and falls back to the full input price, so
every request billed at the undiscounted rate (and at the cross-provider
price, ~28% low): production showed ``Applied model-specific pricing:
input=494.55, cache_read=494.55`` with ``source=configured`` even though
the enclave reported ``cached_prompt_tokens=12800``.
The sibling path is correct: ``BaseUpstreamProvider.get_x_cashu_cost``
already passes ``model_obj`` to ``calculate_cost``. The EHBP path is the
only caller that omitted it.
Fix: thread the pricing model object through ``_compute_ehbp_actual_cost``
— the routed ``model_obj`` normally, or the resolved served model on a
genuine mismatch — and pass it to ``calculate_cost``. The requested model
string stays in ``response_data["model"]`` for logging. 3c813da's
namespace resolution is still what decides *which* object to bill; this
commit makes that object actually reach pricing.
Verified numerically in the new regression test: 12800 cached tokens at
the Tinfoil cache rate (~112 msat/1k) bill ~1434 msats instead of ~8800
msats at the full input rate.
Tests: 2 new in tests/unit/test_tinfoil_integration.py; full suite 2001
passed, 14 skipped. ruff check + mypy clean.