mirror of
https://github.com/Routstr/routstr-core.git
synced 2026-10-05 12:28:22 +00:00
The Responses API reports usage in its own dialect: input_tokens: 9434 (inclusive grand total) input_tokens_details.cached_tokens: 8704 (cached subset) normalize_usage only knew prompt_tokens_details (chat completions), Anthropic top-level fields, and DeepSeek hit/miss. For /v1/responses payloads it found no cache fields and, seeing no prompt_tokens key, fell into the Anthropic-native branch that treats input_tokens as excluding cache — so cached tokens were billed at the full input rate and cache_read_input_tokens was recorded as 0. Two changes: * _extract_cache_tokens also reads input_tokens_details.cached_tokens and input_tokens_details.cache_write_tokens. * normalize_usage treats input_tokens as an inclusive grand total when input_tokens_details is present (Anthropic native never sends that object, so it safely disambiguates the field-name collision). Fixes both streaming and non-streaming /v1/responses billing, which share calculate_cost -> normalize_usage.