Files
redshift bc5b41b3bd fix(generic): carry Venice's native cache_input price into input_cache_read
The generic provider's _native_pricing parses Venice's bespoke model_spec
schema but read only pricing.input/output.usd, dropping cache_input.usd
(e.g. deepseek-v4-1-flash: $0.0075/1M cache reads vs $0.375/1M input).
With input_cache_read left at 0, billing's fallback priced cache reads at
the FULL input rate — a ~50x overcharge on cache hits compared to what the
upstream charges.

cache_input.usd is now coerced with the same rules as the token rates;
a malformed/negative cache rate is treated as absent (never carried),
mirroring the OpenRouter rung's drop-don't-carry behaviour.

Adds regression tests: cache rate carried through, and malformed cache
rates (-, Infinity, non-numeric) dropped while the model still resolves.
2026-09-21 16:26:01 +03:00
..
2026-08-26 22:51:09 +02:00
2026-06-13 23:40:39 +02:00
2026-06-27 16:01:45 +02:00