Files
routstr-core/routstr
redshift 3c813daedc fix(ehbp): resolve the served model within the tinfoil namespace before pricing
The cache discount never applied in production even though everything
downstream of it worked: the enclave reported cache splits
(cached_prompt_tokens=160512 of 161652) and the parse put them into
cache_read_input_tokens, but the applied cache-read rate was the full input
rate.

Root cause: the SDK strips the routstr "tinfoil-" namespace prefix for the
encrypted body (getTinfoilUpstreamModelId), so the enclave always reports the
bare upstream id ("deepseek-v4-1-flash") in X-Tinfoil-Usage-Metrics, while
the catalog registers the model as "tinfoil-deepseek-v4-1-flash" and
forwarded_model_id carries the prefix too. _normalize_upstream_model_id only
lowercases, so every request took the "served model differs" path and
re-derived pricing via get_model_instance on the bare id — which resolves
globally to a cheaper cross-provider model whose pricing has no cache rate
(input_cache_read=0), falling back to the full input price.

Billing was therefore doubly wrong: no cache discount, and E2EE requests
undercharged at the cross-provider rate instead of the Tinfoil rate.

Fix: when the requested model is namespaced "tinfoil-" and the served id is
not, resolve the served id within the same namespace first (with a bare-id
fallback). A same-model report then maps back onto the requested Tinfoil
model (keeping its pricing), and a genuine failover lands on the
actually-served Tinfoil model — while the bare-id lookup that previously
hijacked pricing is only used as a last resort.
2026-09-17 22:20:25 +02:00
..
2026-09-07 00:30:05 +02:00