The session was needed when model pricing lived in the DB (73d3613) and has
been dead since pricing moved to the in-memory model map (0da08fb), yet every
caller was still obliged to supply one. get_x_cashu_cost even opened a DB
session per x-cashu request solely to feed it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Cached prompt tokens were billed at the full input rate whenever a vendor's
usage dialect or cache pricing was unknown, overcharging DeepSeek topups
~5-10x on agentic workloads (hits are 10x cheaper upstream) and silently
mispricing OpenAI cached reads and Anthropic cache writes the same way.
Two root causes, two fixes:
- Usage dialects: DeepSeek reports prompt_cache_hit_tokens /
prompt_cache_miss_tokens, which billing never parsed. Usage normalization
now lives in payment/usage.py as a union parser over the known,
non-colliding dialects (OpenAI prompt_tokens_details, Anthropic additive
cache fields, DeepSeek hit/miss), producing one canonical NormalizedUsage.
Providers expose it as an overridable BaseUpstreamProvider.normalize_usage
hook — the escape hatch for future vendors whose fields genuinely
conflict — and every settlement call site passes the provider's result
through, so calculate_cost holds no vendor knowledge of its own.
- Cache rates: the OpenRouter model feed omits input_cache_read/-write for
most DeepSeek models (and e.g. openai/gpt-4o), so billing fell back to the
full input rate. Missing rates are now backfilled from litellm's bundled
cost map before the provider fee is applied; the input-rate fallback
remains only as the documented last resort.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>