Add TinfoilUpstreamProvider that uses inference.tinfoil.sh as a direct
EHBP upstream. Routstr acts as a blind relay: it forwards the opaque
encrypted body to the Tinfoil enclave without ever seeing plaintext,
and bills from the X-Tinfoil-Usage-Metrics response header.
- New routstr/upstream/tinfoil.py: fetches models from public
GET /v1/models, parses Tinfoil pricing into standard Model/Pricing
schema, supports_ehbp=True, proxies /attestation to atc.tinfoil.sh
- Updated routstr/upstream/ehbp.py:
- parse_tinfoil_usage_metrics() parses prompt=N,completion=N header
- _resolve_ehbp_target_url() honors X-Tinfoil-Enclave-Url from SDK
- _strip_proxy_headers() removes proxy-only headers before forwarding
- _compute_ehbp_actual_cost() converts usage to msats via calculate_cost
- forward_ehbp_request() finalizes with exact token cost when usage
header is present (non-streaming), falls back to max-cost otherwise
- forward_ehbp_x_cashu_request() computes refund from actual cost
when usage is available
- Updated routstr/proxy.py: forward /attestation and /.well-known/ paths
without model/cost/auth lookups
- Updated routstr/upstream/__init__.py and helpers.py: register and
auto-seed Tinfoil provider from TINFOIL_API_KEY env var
- Updated .env.example with TINFOIL_API_KEY
- 22 new unit tests in tests/unit/test_tinfoil_integration.py
- Updated docs/tinfoil-direct-integration.md and docs/ehbp-proxy-support.md
with implementation status and billing behavior table
7.8 KiB
EHBP Proxy Support for Tinfoil Models
Problem
The SDK's SecureClient.fetch encrypts request bodies with HPKE (EHBP protocol)
and sends them to the Routstr provider. The Routstr proxy had no EHBP handling:
- It tried to
json.loads()the binary HPKE-sealed body → failed with a 400 - The upstream PPQ.AI public endpoint (
/v1/chat/completions) doesn't speak EHBP, so the response had noEhbp-Response-Nonceheader SecureClientthrewMissing Ehbp-Response-Nonce headerbecause it expects every response from an EHBP-configuredbaseURLto carry that header
Root cause
PPQ.AI exposes EHBP-aware inference at /private/, separate from the public
/v1/ endpoint. The Routstr proxy was forwarding to /v1/ (the public
endpoint) instead of /private/ (the enclave endpoint). The public endpoint
can't decrypt the body, returns a normal HTTP response, and the SDK can't
decrypt it because there's no nonce header.
The PPQ private-mode proxy (ppq-private-mode-proxy/lib/proxy.ts) shows the
correct pattern: SecureClient talks to api.ppq.ai/private/v1/chat/completions,
which decrypts inside the attested enclave and returns an EHBP-encrypted
response with the Ehbp-Response-Nonce header.
What was changed
routstr/proxy.py
Detects EHBP requests by checking for the Ehbp-Encapsulated-Key header (set
by the EHBP transport on every encrypted request). For EHBP requests:
- Skips JSON body parsing (the body is binary ciphertext, not JSON)
- Reads the model ID from the
X-Routstr-Modelheader (set by the SDK) instead of frombody.model - Routes through new
forward_ehbp_request(bearer auth) andforward_ehbp_x_cashu_request(x-cashu auth) methods - Skips reactive 400 param correction (can't parse encrypted response body)
- Still charges the user via
pay_for_request(usesmax_cost_for_modelfrom the model registry, not the body)
routstr/upstream/base.py
Keeps EHBP as an explicit opt-in provider capability instead of making every upstream provider appear EHBP-capable:
supports_ehbp = Falseby defaultget_ehbp_forwarding_target(path, model_obj)raisesNotImplementedErrorunless a provider opts in and returns a provider-specific EHBP target
The actual EHBP forwarding logic does not live in base.py.
routstr/upstream/ehbp.py
Contains the shared opaque EHBP transport helpers:
EHBPForwardingTarget— provider-specific target URL plus extra headersforward_ehbp_request()— forwards the raw encrypted body to an EHBP-capable provider, streams the encrypted response back untouched, and finalizes bearer billing at max cost because usage is encryptedforward_ehbp_x_cashu_request()— redeems the Cashu token, forwards raw, refunds the full token on upstream failure, and refunds any value abovemax_cost_for_modelon success
routstr/upstream/ppqai.py
- Sets
supports_ehbp = True. - Implements
get_ehbp_forwarding_target()to forward tohttps://api.ppq.ai/private/v1/...— the PPQ.AI enclave endpoint that understands EHBP and returns theEhbp-Response-Nonceheader. - Adds
X-Private-Modelwith the model'sforwarded_model_id(e.g.private/kimi-k2-6). PPQ.AI's billing layer needs this since it can't decrypt the body.
Why it's done this way
The proxy is a blind relay for EHBP requests. It cannot decrypt the body (only the attested enclave can), so it must:
- Get the model ID from a header, not the body
- Forward the raw bytes without parsing or transformation
- Stream the response back without SSE/cost parsing
- Pass through EHBP protocol headers (
Ehbp-Encapsulated-Keyon request,Ehbp-Response-Nonceon response)
Cost tracking happens at the proxy level using max_cost_for_model from the
model registry. Because EHBP responses are encrypted, Routstr cannot reconcile
against token usage. Bearer requests reserve and then finalize max-cost billing;
X-Cashu requests redeem the token and refund any amount above max cost.
End-to-end flow
SDK Routstr Proxy PPQ.AI /private/
│ │ │
│── X-Routstr-Model: tinfoil-kimi-k2-6 ─│ │
│── Ehbp-Encapsulated-Key: <hex> ───────│ │
│── Authorization: Bearer <cashu> ──────│ │
│── body = HPKE-encrypted(kimi-k2-6) ───│ │
│ │ │
│ detects Ehbp-Encapsulated-Key │
│ reads model from X-Routstr-Model │
│ does billing/routing │
│ │ │
│ adds X-Private-Model: private/kimi-k2-6
│ forwards raw body to /private/v1/... │
│ │──────────────────────────────▶│
│ │ enclave decrypts
│ │ runs inference
│ │◀── Ehbp-Response-Nonce ──────│
│ │◀── encrypted response ────────│
│ │ │
│ streams response back untouched │
│◀── encrypted response ────────────────│ │
│ │ │
SecureClient reads nonce, decrypts │ │
SDK SSE processing sees plaintext │ │
Model ID mapping
Three parties see three different model IDs:
| Party | Header/Body | Value | Source |
|---|---|---|---|
| Routstr proxy | X-Routstr-Model header |
tinfoil-kimi-k2-6 |
SDK sends full caller-facing id |
| PPQ.AI billing | X-Private-Model header |
private/kimi-k2-6 |
Proxy sends forwarded_model_id |
| Tinfoil enclave | body.model (encrypted) |
kimi-k2-6 |
SDK strips tinfoil- prefix before encryption |
Implementation status
A dedicated TinfoilUpstreamProvider (routstr/upstream/tinfoil.py) now
implements the direct blind-upstream pattern described above. The shared EHBP
helpers in routstr/upstream/ehbp.py were extended to:
- Request usage metrics via
X-Tinfoil-Request-Usage-Metrics: true. - Parse
X-Tinfoil-Usage-Metricsfrom the response header (non-streaming). - Override the forwarding URL with
X-Tinfoil-Enclave-Urlwhen the SDK sends it. - Finalize bearer billing with actual token cost via
adjust_payment_for_tokens. - Compute X-Cashu refunds from actual cost instead of max cost.
See docs/tinfoil-direct-integration.md for the full implementation notes.
Not yet tested
These changes were written without integration testing due to the complexity
of the full stack (SDK + proxy + PPQ.AI enclave + Cashu mint). Needs end-to-end
verification with a real tinfoil-* model request.
Important assumptions to verify:
- PPQ.AI accepts
/private/v1/...withX-Private-Model. - PPQ.AI enforces consistency between
X-Private-Modeland the encryptedbody.model, otherwise a malicious client could understateX-Routstr-Modelfor billing. - SDK behavior on non-2xx proxy-generated errors that do not carry
Ehbp-Response-Nonce.