Files
redshift 6c91eac02c fix(ehbp): bill the routed model object, not its bare id via the alias map
The cache discount still did not apply after 3c813da because the EHBP
finalization path never handed the serving ``Model`` to ``calculate_cost``:
it passed only the model *string* (``{"model": pricing_model_id, ...}``),
so ``_get_pricing_rates`` fell into its "settling without routed model
identity" branch and re-derived pricing through the global alias map.

That lookup resolves the string to the *best-ranked* candidate for the
id, not the serving one. Tinfoil's catalog id (``deepseek-v4-1-flash``)
is also a bare cross-provider alias, and the cheaper cross-provider
candidate carries ``input_cache_read=0``. ``_get_pricing_rates`` treats a
zero cache rate as "missing" and falls back to the full input price, so
every request billed at the undiscounted rate (and at the cross-provider
price, ~28% low): production showed ``Applied model-specific pricing:
input=494.55, cache_read=494.55`` with ``source=configured`` even though
the enclave reported ``cached_prompt_tokens=12800``.

The sibling path is correct: ``BaseUpstreamProvider.get_x_cashu_cost``
already passes ``model_obj`` to ``calculate_cost``. The EHBP path is the
only caller that omitted it.

Fix: thread the pricing model object through ``_compute_ehbp_actual_cost``
— the routed ``model_obj`` normally, or the resolved served model on a
genuine mismatch — and pass it to ``calculate_cost``. The requested model
string stays in ``response_data["model"]`` for logging. 3c813da's
namespace resolution is still what decides *which* object to bill; this
commit makes that object actually reach pricing.

Verified numerically in the new regression test: 12800 cached tokens at
the Tinfoil cache rate (~112 msat/1k) bill ~1434 msats instead of ~8800
msats at the full input rate.

Tests: 2 new in tests/unit/test_tinfoil_integration.py; full suite 2001
passed, 14 skipped. ruff check + mypy clean.
2026-09-17 22:50:23 +02:00
..
2025-08-06 20:31:55 -03:00