mirror of
https://github.com/Routstr/routstr-core.git
synced 2026-10-05 20:28:23 +00:00
_fold_cache_into_input_tokens rolled cache_read/cache_creation tokens into the visible prompt_tokens for every dialect. In the OpenAI family (Venice, OpenAI, DeepSeek, OpenRouter, litellm) prompt_tokens already includes the cached portion, so the client-visible prompt count double-counted cache hits (Venice: 14075-token prompt shown as 27997 after a 13922-token cache read). Billing was unaffected — normalize_usage already subtracts cache exactly once and the fold runs after cost calculation. The fold now mirrors normalize_usage: only Anthropic-native input_tokens (which excludes cache) gets the roll-up; prompt_tokens is left untouched. Adds dict-based regression tests for the Venice/OpenAI, Anthropic-native and litellm-mirror shapes (existing Mock tests never exercised the logic).