Files
redshift bc8172783d fix(usage): don't fold cache tokens into inclusive prompt_tokens
_fold_cache_into_input_tokens rolled cache_read/cache_creation tokens
into the visible prompt_tokens for every dialect. In the OpenAI family
(Venice, OpenAI, DeepSeek, OpenRouter, litellm) prompt_tokens already
includes the cached portion, so the client-visible prompt count
double-counted cache hits (Venice: 14075-token prompt shown as 27997
after a 13922-token cache read). Billing was unaffected — normalize_usage
already subtracts cache exactly once and the fold runs after cost
calculation.

The fold now mirrors normalize_usage: only Anthropic-native input_tokens
(which excludes cache) gets the roll-up; prompt_tokens is left untouched.

Adds dict-based regression tests for the Venice/OpenAI, Anthropic-native
and litellm-mirror shapes (existing Mock tests never exercised the logic).
2026-09-21 16:06:48 +03:00
..
2025-08-06 20:31:55 -03:00