mirror of
https://github.com/Routstr/routstr-core.git
synced 2026-08-10 19:16:31 +00:00
Cached prompt tokens were billed at the full input rate whenever a vendor's usage dialect or cache pricing was unknown, overcharging DeepSeek topups ~5-10x on agentic workloads (hits are 10x cheaper upstream) and silently mispricing OpenAI cached reads and Anthropic cache writes the same way. Two root causes, two fixes: - Usage dialects: DeepSeek reports prompt_cache_hit_tokens / prompt_cache_miss_tokens, which billing never parsed. Usage normalization now lives in payment/usage.py as a union parser over the known, non-colliding dialects (OpenAI prompt_tokens_details, Anthropic additive cache fields, DeepSeek hit/miss), producing one canonical NormalizedUsage. Providers expose it as an overridable BaseUpstreamProvider.normalize_usage hook — the escape hatch for future vendors whose fields genuinely conflict — and every settlement call site passes the provider's result through, so calculate_cost holds no vendor knowledge of its own. - Cache rates: the OpenRouter model feed omits input_cache_read/-write for most DeepSeek models (and e.g. openai/gpt-4o), so billing fell back to the full input rate. Missing rates are now backfilled from litellm's bundled cost map before the provider fee is applied; the input-rate fallback remains only as the documented last resort. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>