Files
routstr-core/routstr/payment/usage.py
T
9qeklajcand9qeklajc 7f44389b59 fix(certification): harden against adversarial inputs found by tester agents
Three independent tester subagents found 12 defects (1 critical, 2 high,
9 medium/low). This commit fixes all of them and adds 56 regression tests.

Critical (CLI dead on arrival):
- certify_upstream_url() called sats_usd_price() which raises ValueError
  in any fresh process (the module global is only set by the app lifespan
  task). Now resolves via _resolve_sats_usd_price(): module global → BTC
  global → exchange feed → None (warn row, not a crash). Adds
  --sats-usd-price CLI flag for explicit override.

High (non-finite tokens crash the billing path):
- parse_token_count() crashed on Infinity/NaN (json.loads accepts both).
  Fixed to reject non-finite values → 0. This was a shared-code bug in
  routstr/payment/usage.py, reachable from the main billing path too.
- usage_capture_row and cost_prompt_completion_row now guard normalize_usage
  in try/except via safe_row(), so a raising check becomes a fail row
  instead of a 500.

Medium:
- certification_row now coerces non-dict evidence to {} (was stored verbatim)
- endpoint_validity_row checks .hostname not .netloc (http://:8080 rejected)
- models_payload_row rejects empty-string ids (agrees with CLI discovery)
- models_payload_row / usage_capture_row guard against non-dict payloads
- _reported_usd_cost uses coerce_rate for parity with the engine
- _expected_token_msats raises ValueError on non-finite rates (not OverflowError)
- get_candidates() call in admin endpoint wrapped in try/except
- Admin timeout clamped to [1, 60] seconds
- Explicit --prompt-price validated via coerce_rate (negatives rejected)
- CLI --json-out flag writes strictly parseable JSON to a file
- CLI logs routed to stderr so stdout is the report's channel

Test plan: 1749 passed, 1 skipped (no regressions); ruff clean; mypy clean.
2026-09-21 22:53:33 +02:00

165 lines
7.2 KiB
Python

"""Vendor-agnostic normalization of upstream usage objects.
Upstream providers report token usage in vendor dialects that differ in field
names and in whether cached tokens are included in the input count:
* OpenAI / Azure / xAI / Groq / Moonshot / Qwen / Gemini-compat: cache reads in
``prompt_tokens_details.cached_tokens``, included in ``prompt_tokens``.
* OpenRouter: same as OpenAI plus cache *writes* in
``prompt_tokens_details.cache_write_tokens``, also included in
``prompt_tokens``.
* litellm-normalized: same nesting, but names the write field
``prompt_tokens_details.cache_creation_tokens`` (and additionally mirrors the
Anthropic top-level fields), with ``prompt_tokens`` as the grand total.
* Anthropic native: ``cache_read_input_tokens`` / ``cache_creation_input_tokens``
top-level, additive to (not included in) ``input_tokens``.
* DeepSeek: ``prompt_cache_hit_tokens`` / ``prompt_cache_miss_tokens``, with
``prompt_tokens = hit + miss``.
* OpenAI Responses API: ``input_tokens_details.cached_tokens`` (and
``cache_write_tokens``), included in ``input_tokens`` — same inclusive
semantics as ``prompt_tokens``, but under the Responses API field names.
What decides whether cached tokens must be subtracted out of the input count is
**which prompt field the vendor uses**, not which cache field appears:
* ``prompt_tokens`` present -> cached + cache-write tokens are *included* in it
(OpenAI family, DeepSeek, OpenRouter, litellm); subtract both so
``input_tokens`` holds only the regular-rate portion.
* ``input_tokens_details`` present -> OpenAI Responses API; ``input_tokens``
*includes* cached + cache-write tokens, so subtract both. This disambiguates
the field-name collision with Anthropic native, which never sends
``input_tokens_details``.
* only ``input_tokens`` (Anthropic native) -> cached tokens are *additive*;
leave ``input_tokens`` untouched.
``normalize_usage`` maps all of them onto one canonical ``NormalizedUsage``
shape so billing code needs no vendor knowledge. The known dialects' field
names do not collide, so a single union parser is safe; a vendor whose fields
would genuinely conflict needs a dedicated branch here.
"""
import math
from pydantic.v1 import BaseModel
class NormalizedUsage(BaseModel):
"""Canonical token usage: input_tokens never includes cached tokens."""
input_tokens: int = 0
output_tokens: int = 0
cache_read_tokens: int = 0
cache_write_tokens: int = 0
def parse_token_count(value: object) -> int:
"""Parse a token count from various formats (int, float, str, bool).
A non-finite count is not a count. ``json.loads`` accepts the bare
``Infinity``/``NaN`` literals and overflows ``1e999`` to ``inf``, so an
upstream — or an attacker who controls one — can put them on the wire.
``int(inf)`` raises ``OverflowError`` and ``int(nan)`` raises
``ValueError``; either would turn a billing path into a 500. Same rule as
``is_usable_rate``: reject the value, do not crash on it.
"""
if isinstance(value, bool):
return 0
if isinstance(value, int):
return max(0, value)
if isinstance(value, float):
return max(0, int(value)) if math.isfinite(value) else 0
if isinstance(value, str):
try:
parsed = float(value)
except (ValueError, OverflowError):
return 0
return max(0, int(parsed)) if math.isfinite(parsed) else 0
return 0
def _first_token_count(usage_data: dict, *fields: str) -> int:
"""Return the first positive token count among the given fields."""
for field in fields:
value = parse_token_count(usage_data.get(field, 0))
if value > 0:
return value
return 0
def _extract_cache_tokens(usage_data: dict) -> tuple[int, int]:
"""Pull (cache_read, cache_write) across all known dialects.
Precedence (highest first), independent for reads and writes:
* Anthropic top-level: ``cache_read_input_tokens`` /
``cache_creation_input_tokens``.
* Nested ``prompt_tokens_details``: ``cached_tokens`` for reads;
``cache_creation_tokens`` (litellm) or ``cache_write_tokens``
(OpenRouter) for writes.
* Nested ``input_tokens_details`` (OpenAI Responses API): ``cached_tokens``
for reads, ``cache_write_tokens`` for writes.
* DeepSeek: ``prompt_cache_hit_tokens`` for reads (no write concept).
"""
cache_read = parse_token_count(usage_data.get("cache_read_input_tokens", 0))
cache_write = parse_token_count(usage_data.get("cache_creation_input_tokens", 0))
prompt_details = usage_data.get("prompt_tokens_details")
if isinstance(prompt_details, dict):
if not cache_read:
cache_read = parse_token_count(prompt_details.get("cached_tokens", 0))
if not cache_write:
cache_write = _first_token_count(
prompt_details, "cache_creation_tokens", "cache_write_tokens"
)
input_details = usage_data.get("input_tokens_details")
if isinstance(input_details, dict):
if not cache_read:
cache_read = parse_token_count(input_details.get("cached_tokens", 0))
if not cache_write:
cache_write = parse_token_count(input_details.get("cache_write_tokens", 0))
if not cache_read:
# DeepSeek: prompt_tokens = prompt_cache_hit_tokens + prompt_cache_miss_tokens
cache_read = parse_token_count(usage_data.get("prompt_cache_hit_tokens", 0))
return cache_read, cache_write
def normalize_usage(usage_data: object) -> NormalizedUsage | None:
"""Map a vendor usage dict onto the canonical shape, or None if absent.
Cached reads and writes are subtracted from the input count exactly once,
only for dialects whose input grand total already includes them: the
``prompt_tokens`` family (OpenAI chat completions, DeepSeek, OpenRouter,
litellm) and the OpenAI Responses API (``input_tokens`` inclusive,
identified by the presence of ``input_tokens_details``). Anthropic native
reports them additively under ``input_tokens`` and is left untouched.
"""
if not isinstance(usage_data, dict):
return None
output_tokens = _first_token_count(usage_data, "completion_tokens", "output_tokens")
cache_read, cache_write = _extract_cache_tokens(usage_data)
# ``prompt_tokens`` is the inclusive grand total; ``input_tokens`` (Anthropic
# native) excludes cached tokens. The field chosen decides whether to subtract.
if "prompt_tokens" in usage_data:
input_tokens = parse_token_count(usage_data.get("prompt_tokens", 0))
input_tokens = max(0, input_tokens - cache_read - cache_write)
elif isinstance(usage_data.get("input_tokens_details"), dict):
# OpenAI Responses API: ``input_tokens`` is also an inclusive grand
# total (cached tokens are a subset of it), signalled by the nested
# ``input_tokens_details`` object Anthropic native never sends.
input_tokens = parse_token_count(usage_data.get("input_tokens", 0))
input_tokens = max(0, input_tokens - cache_read - cache_write)
else:
input_tokens = parse_token_count(usage_data.get("input_tokens", 0))
return NormalizedUsage(
input_tokens=input_tokens,
output_tokens=output_tokens,
cache_read_tokens=cache_read,
cache_write_tokens=cache_write,
)