redshift 6c91eac02c fix(ehbp): bill the routed model object, not its bare id via the alias map
The cache discount still did not apply after 3c813da because the EHBP
finalization path never handed the serving ``Model`` to ``calculate_cost``:
it passed only the model *string* (``{"model": pricing_model_id, ...}``),
so ``_get_pricing_rates`` fell into its "settling without routed model
identity" branch and re-derived pricing through the global alias map.

That lookup resolves the string to the *best-ranked* candidate for the
id, not the serving one. Tinfoil's catalog id (``deepseek-v4-1-flash``)
is also a bare cross-provider alias, and the cheaper cross-provider
candidate carries ``input_cache_read=0``. ``_get_pricing_rates`` treats a
zero cache rate as "missing" and falls back to the full input price, so
every request billed at the undiscounted rate (and at the cross-provider
price, ~28% low): production showed ``Applied model-specific pricing:
input=494.55, cache_read=494.55`` with ``source=configured`` even though
the enclave reported ``cached_prompt_tokens=12800``.

The sibling path is correct: ``BaseUpstreamProvider.get_x_cashu_cost``
already passes ``model_obj`` to ``calculate_cost``. The EHBP path is the
only caller that omitted it.

Fix: thread the pricing model object through ``_compute_ehbp_actual_cost``
— the routed ``model_obj`` normally, or the resolved served model on a
genuine mismatch — and pass it to ``calculate_cost``. The requested model
string stays in ``response_data["model"]`` for logging. 3c813da's
namespace resolution is still what decides *which* object to bill; this
commit makes that object actually reach pricing.

Verified numerically in the new regression test: 12800 cached tokens at
the Tinfoil cache rate (~112 msat/1k) bill ~1434 msats instead of ~8800
msats at the full input rate.

Tests: 2 new in tests/unit/test_tinfoil_integration.py; full suite 2001
passed, 14 skipped. ruff check + mypy clean.
2026-09-17 22:50:23 +02:00
2025-10-22 12:18:00 +08:00
2026-09-09 23:23:11 +02:00
2026-08-06 22:58:44 +02:00
2026-08-26 22:51:09 +02:00
2025-06-09 23:46:19 +02:00
2025-08-12 15:02:22 -03:00
2025-04-08 21:42:32 +08:00
2025-08-01 22:31:18 -03:00

Routstr Payment Proxy

License Stars Issues Release

Routstr is a decentralized protocol for permissionless, private, and censorship-resistant AI inference. It combines Nostr for discovery and Cashu for private Bitcoin micropayments.

This repo contains Routstr Core: a FastAPI-based reverse proxy that sits in front of OpenAI-compatible APIs and handles pay-per-request billing.

Start Here

Basic Usage

If you are a user/developer, you just point an OpenAI-compatible SDK at a Routstr node and pay with a Cashu token.

OpenAI SDK

from openai import OpenAI

client = OpenAI(
    base_url="https://api.routstr.com/v1",
    api_key="cashuBo2FteCJodHRwczovL21...",
)

response = client.chat.completions.create(
    model="gpt-5-nano",
    messages=[{"role": "user", "content": "hello"}],
)

print(response.choices[0].message.content)

cURL

curl https://api.routstr.com/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "x-cashu: cashuBo2FteCJodHRwczovL21..." \
  -d '{
    "model": "gpt-5-nano",
    "messages": [{"role": "user", "content": "hello"}]
  }'

Quick Start (Docker)

If you are a node runner, start a Routstr Core instance using Docker Compose:

  1. Prepare your .env:

    # Optional: encrypts node secrets at rest. If unset, the node generates a key
    # on first start, writes it to routstr_secret.key, and prints it once — back
    # up that file. Set it explicitly to manage the key yourself (recommended in
    # production).
    ROUTSTR_SECRET_KEY=<generated-key>
    NAME="My AI Node"
    DESCRIPTION="Fast access to models"
    RECEIVE_LN_ADDRESS=yourname@wallet.com
    

    Your Nostr identity (nsec) is not set in .env — configure it from the admin UI after first start, where it's stored encrypted in the database. (NSEC in .env is still read once as a legacy seed for existing deployments.)

    If you don't set one, a key is generated and printed on first start — save it somewhere safe (losing it makes previously encrypted secrets unreadable). To supply your own, generate it once and keep it stable:

    uv run python -c "from cryptography.fernet import Fernet; print(Fernet.generate_key().decode())"
    
  2. Start the services:

    docker compose up -d
    
  3. Get your admin password: On first start the node generates an admin password and logs it once with the /admin URL. Read it from the logs:

    docker compose logs routstr | grep -i admin
    

    (Lost it? Reset with docker compose exec routstr /.venv/bin/python scripts/reset_admin_password.py --regenerate.)

  4. Configure: Open http://localhost:8000/admin/ to connect your AI providers and set pricing.

For full instructions, see the Provider Quick Start Guide.

Development

make setup
cp .env.example .env
fastapi run routstr
S
Description
Routstr is a decentralized protocol for permissionless, private, and censorship-resistant AI inference.
Readme GPL-3.0
33 MiB
Languages
Python 79.8%
TypeScript 18.3%
HTML 1.2%
JavaScript 0.2%
Makefile 0.2%
Other 0.3%