The doc asserted that usage-metrics headers are never returned even when `X-Tinfoil-Request-Usage-Metrics: true` is sent. That observation was made against PPQ's `/private/` endpoint on 2026-06-21 and contradicts the "Usage metrics header format" section added later (2026-07-08), which documents the header as present. Scope the older claim to the PPQ endpoint and cross-reference the verified section, so the doc no longer contradicts itself. While verifying against the direct enclave, two further gaps surfaced: - The `/v1/responses` open question is answered: it does return `X-Tinfoil-Usage-Metrics` as a plaintext response header with the same field set as `/v1/chat/completions`, including `cost_usd`. - Streaming trailers are emitted twice in a single comma-joined field. Note this so a future trailer reader is tolerant of the duplicate.
20 KiB
Tinfoil / PPQ Private-Mode Integration Notes
This document summarizes the current options for integrating Tinfoil/PPQ private models with Routstr, based on the EHBP work in this branch and local testing against ppq-private-mode-proxy.
Background
PPQ private models run behind a Tinfoil/EHBP flow:
- Request bodies are HPKE-encrypted by a Tinfoil client.
- The PPQ
/private/endpoint routes ciphertext to the attested enclave. - The enclave decrypts, runs inference, and returns an encrypted response.
- The caller's Tinfoil client decrypts the response locally.
The PPQ private-mode proxy (~/projects/ppq-private-mode-proxy) uses this pattern with the JavaScript tinfoil SDK:
import { SecureClient } from "tinfoil";
const apiBase = "https://api.ppq.ai";
const client = new SecureClient({
baseURL: `${apiBase}/private/`,
attestationBundleURL: `${apiBase}/private`,
transport: "ehbp",
});
await client.ready();
const response = await client.fetch(`${apiBase}/private/v1/chat/completions`, {
method: "POST",
headers: {
"Content-Type": "application/json",
Authorization: `Bearer ${process.env.PPQ_API_KEY}`,
"X-Private-Model": "private/gpt-oss-120b",
"x-query-source": "api",
},
body: JSON.stringify({
model: "gpt-oss-120b", // enclave-internal model id
messages: [{ role: "user", content: "Hello" }],
}),
});
const json = await response.json();
console.log(json.usage);
X-Private-Model carries the PPQ-facing private model id, while the encrypted JSON body uses the enclave-internal model id without the private/ prefix.
PPQ private model pricing
PPQ exposes private model pricing through:
GET https://api.ppq.ai/v1/models?type=all
Filter models whose IDs start with private/.
Example private pricing observed:
| Model | Input USD / 1M tokens | Output USD / 1M tokens |
|---|---|---|
private/gpt-oss-120b |
0.79125 |
1.31875 |
private/llama3-3-70b |
1.84625 |
2.90125 |
private/qwen3-vl-30b |
1.31875 |
4.22 |
private/glm-5-2 |
1.5825 |
5.53875 |
private/gemma4-31b |
0.47475 |
1.055 |
private/kimi-k2-6 |
1.5825 |
5.53875 |
The ppq-private-mode-proxy commit ba984214793d3bca0f7d046b6955d42abc1c6843 changed OpenClaw display metadata to align with a 5% API margin. The live PPQ model endpoint returns the more precise rates above, which already include that margin.
Actual PPQ private billing is:
price_usd =
input_tokens * input_per_1M_tokens / 1_000_000
+ output_tokens * output_per_1M_tokens / 1_000_000
A test request to private/gpt-oss-120b produced:
{
"input_count": 74,
"output_count": 3,
"price_in_usd": 0.00006250875
}
which exactly matches:
74 * 0.79125 / 1_000_000 + 3 * 1.31875 / 1_000_000
= 0.00006250875 USD
Integration architectures
There are three materially different ways Routstr could integrate Tinfoil/PPQ private inference.
Option A: User integrates Tinfoil directly
User app / SDK
-> Tinfoil SecureClient / TinfoilAI
-> EHBP-encrypted request
-> Tinfoil/PPQ private enclave
Properties:
- Best privacy for the user.
- Routstr is not in the request path.
- User's client encrypts requests and decrypts responses.
- Usage is visible to the user's app after decryption.
- Billing is handled directly by Tinfoil/PPQ.
Example with Tinfoil's OpenAI-compatible client:
import { TinfoilAI } from "tinfoil";
const client = new TinfoilAI({
apiKey: process.env.TINFOIL_API_KEY,
transport: "ehbp",
});
const res = await client.chat.completions.create({
model: "llama3-3-70b",
messages: [{ role: "user", content: "Hello" }],
});
console.log(res.choices[0].message.content);
console.log(res.usage);
This is not a Routstr marketplace flow unless Routstr only acts as discovery/UI around direct Tinfoil/PPQ usage.
Option B: Routstr integrates Tinfoil as an upstream client
User -> Routstr plaintext request
-> Routstr Tinfoil SecureClient encrypts to PPQ/Tinfoil
-> PPQ private enclave
-> Routstr receives decrypted response
-> Routstr bills from decrypted usage
-> Routstr returns plaintext response to user
Properties:
- Easier exact billing.
- Routstr can read the decrypted OpenAI response and
usageobject. - Routstr can charge exact PPQ token pricing.
- Privacy is different: the user sends plaintext to Routstr, and Routstr sees prompts/responses.
- End-to-end encryption is only Routstr-to-enclave, not user-to-enclave.
This should be considered a separate product/provider mode, not the same as an end-to-end private relay.
Practical implementation options:
- Run a Node sidecar that uses the
tinfoilnpm package and expose it as a local HTTP upstream to Routstr. - Port EHBP client behavior to Python.
- Reuse or adapt
ppq-private-mode-proxyas a local upstream.
A Node sidecar is probably the quickest implementation path because ppq-private-mode-proxy already demonstrates the full flow.
Option C: Routstr uses Tinfoil as a direct blind upstream
This is the current branch's design intent and is likely the best fit if Tinfoil/PPQ exposes usage metadata headers:
User Tinfoil SecureClient
-> encrypted request body
-> Routstr proxy
-> Tinfoil/PPQ private API
with X-Tinfoil-Request-Usage-Metrics: true
-> PPQ private enclave
<- encrypted response
plus X-Tinfoil-Usage-Metrics / cost headers
<- Routstr proxy
-> user decrypts response
In this mode Routstr is still a normal upstream proxy from the user's point of view, but the upstream is Tinfoil/PPQ private inference and the body remains opaque to Routstr.
Properties:
- Strongest privacy with Routstr in the path.
- Routstr never sees plaintext prompt or plaintext response.
- Routstr can authenticate and route based on plaintext headers.
- Routstr should request Tinfoil usage metadata by adding
X-Tinfoil-Request-Usage-Metrics: trueto the upstream request. - Tinfoil documents
X-Tinfoil-Usage-Metricsas an upstream response header for non-streaming requests. - For streaming requests, Tinfoil documents
X-Tinfoil-Usage-Metricsas an HTTP trailer available only after the response body completes. - If Tinfoil/PPQ returns
X-Tinfoil-Usage-Metricsor a cost header on the encrypted response, Routstr can bill exactly without decrypting the body. - If usage is only present inside the encrypted response body, Routstr still cannot read it and exact billing is not possible without a separate metadata path.
This is the only architecture that preserves end-to-end encryption from the user to the PPQ/Tinfoil enclave while still letting Routstr mediate payment. The key requirement is that usage/cost metadata must be returned outside the encrypted body, ideally as a response header available before body streaming begins.
Original Routstr problem
The original EHBP implementation charged successful EHBP requests at max_cost_for_model because Routstr could not decrypt the response body:
successful EHBP request -> charge full reserved max cost
That is incorrect for PPQ private models. Max cost should be only a reservation/solvency ceiling. Final charge should use actual PPQ private token pricing.
Desired behavior:
reserve max cost
forward encrypted request
obtain actual usage/cost metadata
finalize actual cost
refund/release the difference
Usage/cost metadata requirement
For blind-relay exact billing, PPQ should return one of the following outside the encrypted response body:
X-PPQ-Cost-USD: 0.00006250875
or:
X-Private-Usage-Metrics: input=74,output=3
or:
X-Tinfoil-Usage-Metrics: prompt=74,completion=3,total=77
Tinfoil proxy documentation references a usage-metrics flow where a proxy can request usage via:
X-Tinfoil-Request-Usage-Metrics: true
and read usage from:
X-Tinfoil-Usage-Metrics
Docs/example references:
- https://docs.tinfoil.sh/guides/proxy-server
- https://github.com/tinfoilsh/encrypted-request-proxy-example
During local PPQ testing, PPQ responses included this CORS exposure header:
Access-Control-Expose-Headers: Ehbp-Response-Nonce, X-Private-Usage-Metrics, X-Encrypted-Usage-Metrics, X-Tinfoil-Usage-Metrics
However, when this was tested against PPQ's /private/ endpoint
(private/gpt-oss-120b) on 2026-06-21, the non-streaming response did not
include any of these usage headers, even when
X-Tinfoil-Request-Usage-Metrics: true was sent. That observation does not
hold for the direct Tinfoil enclave upstream that Routstr ships: see
Usage metrics header format below, where the
response header and the streaming trailer are both verified present.
The decrypted body did include normal OpenAI usage, but only the decrypting Tinfoil client can see that body.
Query-history fallback
PPQ's query history endpoint exposes actual usage and cost:
GET https://api.ppq.ai/queries/history?page=1&page_count=...
A record includes:
{
"timestamp": "...",
"model": "private/gpt-oss-120b",
"input_count": 74,
"output_count": 3,
"price_in_usd": 0.00006250875,
"query_type": "chat_completion",
"query_source": "api"
}
This could be used as a fallback, but it is less robust than response headers/trailers because matching a request to a history row can be race-prone under concurrency. It would need a reliable request identifier or metadata field that PPQ stores in history.
Recommended Routstr direction
Use Tinfoil/PPQ as a direct blind upstream and have Routstr explicitly request usage metadata:
User encrypts body
Routstr reserves max cost
Routstr forwards encrypted body to Tinfoil/PPQ /private/
Routstr includes X-Tinfoil-Request-Usage-Metrics: true
Tinfoil/PPQ returns X-Tinfoil-Usage-Metrics as a response header for non-streaming,
or as an HTTP trailer after the body completes for streaming
Routstr finalizes exact charge
User decrypts encrypted response
This preserves both:
- privacy: Routstr cannot read prompts/responses;
- exact billing: Routstr can charge actual PPQ private model cost, assuming Tinfoil/PPQ returns usage/cost metadata outside the encrypted body.
Implementation steps:
-
Update PPQ model fetching to include private models:
GET https://api.ppq.ai/v1/models?type=all -
Register
private/*models and any Routstr-facing aliases with correctforwarded_model_id. -
Keep max-cost reservation for bearer keys and X-Cashu solvency checks.
-
Add parsing support for possible usage/cost headers:
X-PPQ-Cost-USD X-Private-Usage-Metrics X-Tinfoil-Usage-Metrics X-Encrypted-Usage-Metrics -
Finalize by actual cost instead of max cost:
actual_msats = ceil((actual_usd / sats_usd_price()) * 1000) actual_msats = max(actual_msats, settings.min_request_msat) actual_msats = min(actual_msats, reserved_msats)or, if only token counts are available:
actual_usd = input_tokens * input_per_1M_tokens / 1_000_000 + output_tokens * output_per_1M_tokens / 1_000_000 -
If usage/cost metadata is missing, choose an explicit policy:
- fail closed and refund/revert;
- query PPQ history as a fallback;
- fallback to max-cost billing only if explicitly configured and clearly disclosed.
Silent max-cost billing should not be the default for PPQ private requests.
X-Cashu consideration
For bearer-auth requests, finalization can happen after the response stream completes if usage is delivered as a trailer.
For X-Cashu, Routstr needs to return the refund token in the response headers. If usage/cost is only available after consuming the encrypted response stream, Routstr may need to buffer EHBP responses before sending them to the client so it can compute the refund amount first.
Possible approaches:
- Prefer a non-trailer response header with actual cost, available before streaming body starts.
- Buffer EHBP X-Cashu responses and then return
X-Cashurefund. - Introduce a later/refund-claim mechanism, which would be a larger protocol change.
Summary
- PPQ private models are billed per actual input/output tokens.
- Private model rates are available from
GET /v1/models?type=all. - Max-cost EHBP fallback is wrong for PPQ private models; current code releases/refunds when trusted usage metadata is absent.
- Direct Tinfoil integration inside Routstr would enable exact usage billing but would make Routstr see plaintext.
- A blind EHBP relay preserves privacy but requires PPQ/Tinfoil to expose usage/cost in plaintext headers/trailers.
- The preferred solution is to keep Routstr blind and have PPQ return billing metadata outside the encrypted body.
Implementation status
Direct Tinfoil upstream integration is implemented in routstr/upstream/tinfoil.py
and routstr/upstream/ehbp.py.
What was built
-
TinfoilUpstreamProvider(provider_type = "tinfoil"):- Base URL:
https://inference.tinfoil.sh - Fetches models from the public
GET /v1/modelsendpoint (no auth needed). - Parses Tinfoil's pricing (
inputTokenPricePer1M,outputTokenPricePer1M,cachedInputTokenPricePer1M,requestPrice) into the standardModel/Pricingschema. Cached reads use the cached rate when present, otherwise the full input rate; cache writes always use the full input rate. supports_ehbp = True— acts as a blind EHBP relay.get_ehbp_forwarding_target()returns a target that includesX-Tinfoil-Request-Usage-Metrics: true.forward_get_request()proxies/attestationtohttps://atc.tinfoil.sh/attestationso the SDK can fetch attestation bundles through Routstr.- Registered in
routstr/upstream/__init__.pyand seeded fromTINFOIL_API_KEYenv var.
- Base URL:
-
routstr/upstream/ehbp.py:parse_tinfoil_usage_metrics()parsesprompt=N,completion=N[,total=N][,cached_prompt_tokens=N, uncached_prompt_tokens=N][,model=<name>][,cost_usd=<usd>]into an OpenAI-style usage dict. Cache reads map tocache_read_input_tokenssocalculate_costcan bill them at the cached rate. Themodelfield (added in tinfoilsh/confidential-model-router PR #385) is extracted as a string;cost_usdis parsed for logging only._resolve_ehbp_target_url()overrides the forwarding URL withX-Tinfoil-Enclave-Urlwhen the SDK sends it._strip_proxy_headers()removesX-Routstr-Model,X-Tinfoil-Enclave-Url, andX-Tinfoil-Request-Usage-Metricsbefore forwarding to the enclave._compute_ehbp_actual_cost()converts the usage header into msats viacalculate_cost(), clamped to[min_request_msat, max_cost_for_model]. When the header'smodel=<name>differs from the requested model, the actual served model's pricing is used for cost calculation.forward_ehbp_request()(bearer auth): ifX-Tinfoil-Usage-Metricsis present in the response header, finalizes withadjust_payment_for_tokens()for exact billing; otherwise releases the reservation. The encrypted body cannot be estimated locally, and the authorization ceiling is not billed. Billing uses the actual served model when it differs from the requested one.forward_ehbp_x_cashu_request(): if usage is available, computes the refund from actual cost instead of max cost, using the actual served model's pricing when applicable.
-
routstr/proxy.py:/attestationand/tee/attestationpaths are forwarded to Tinfoil upstreams without model/cost/auth lookups.
Billing behavior
| Request shape | Usage source | Billing |
|---|---|---|
| Bearer, non-streaming | X-Tinfoil-Usage-Metrics response header |
Exact token cost via adjust_payment_for_tokens |
| Bearer, streaming | X-Tinfoil-Usage-Metrics HTTP trailer |
Exact token cost (h11 captures trailers) |
| Bearer, no usage header/trailer | N/A | Release reservation; zero charge |
| X-Cashu, non-streaming | X-Tinfoil-Usage-Metrics response header |
Refund = redeemed - actual_cost |
| X-Cashu, streaming | X-Tinfoil-Usage-Metrics HTTP trailer |
Refund = redeemed - actual_cost (h11 captures trailers) |
| X-Cashu, no usage header/trailer | N/A | Full refund |
Cost response headers
Since EHBP response bodies are opaque encrypted blobs, per-request cost cannot be injected into the JSON body (as done in the normal proxy flow). Instead, Routstr returns cost info as response headers:
| Header | Auth | Description |
|---|---|---|
X-Routstr-Cost-Msats |
Bearer, X-Cashu | Settled msats debited for this request |
X-Routstr-Computed-Cost-Msats |
Bearer, X-Cashu | Computed usage cost, emitted when it differs from the settled debit |
X-Routstr-Cost-Usd |
Bearer | USD equivalent of the computed usage |
X-Routstr-Input-Cost-Msats |
Bearer, X-Cashu | msats attributed to input tokens |
X-Routstr-Output-Cost-Msats |
Bearer, X-Cashu | msats attributed to output tokens |
X-Routstr-Cache-Read-Msats |
Bearer, X-Cashu | msats attributed to cached input (cache reads) |
X-Routstr-Cache-Creation-Msats |
Bearer, X-Cashu | msats attributed to cache creation (0 for Tinfoil today) |
The client/Tinfoil SDK can read these headers from the HTTP response without
needing to decrypt the body. A duplicate or rejected finalization can therefore
report a zero settled debit while preserving the non-zero computed usage cost.
Normal JSON responses use the same contract: total_msats and charged_msats
are settled values, while computed_msats retains the usage calculation when
it differs.
Setup
TINFOIL_API_KEY=your-tinfoil-api-key
The provider is auto-seeded on first startup.
Usage metrics header format
Tinfoil returns usage metrics in the X-Tinfoil-Usage-Metrics response header
(non-streaming) or HTTP trailer (streaming) when X-Tinfoil-Request-Usage-Metrics: true is sent. As of tinfoilsh/confidential-model-router PR #385, the format is:
prompt=<prompt_tokens>,completion=<completion_tokens>,total=<total_tokens>[,cached_prompt_tokens=<n>,uncached_prompt_tokens=<n>][,model=<served_model>][,cost_usd=<usd>]
prompt is the inclusive prompt total; cached_prompt_tokens is the portion
already in Tinfoil's prefix cache and is billed at the model's
cachedInputTokenPricePer1M rate (or the full input rate when the model has
no cached rate). cost_usd is Tinfoil's own computed request cost and is
currently parsed for observability only — Routstr bills from token counts.
Note that the header/trailer value is not always a single occurrence: for
streaming responses the trailer is emitted twice, so a client that reads the
trailer directly may see the same prompt=...,completion=...,... string twice
in one field, comma-joined. Parsers must be tolerant of the duplicate rather
than assuming exactly one occurrence. parse_tinfoil_usage_metrics() is
unaffected: it assigns each key=value part as it walks the comma-separated
value, and both occurrences carry identical numbers.
The model field carries the actual model name served by the enclave.
Routstr uses this to:
- Verify the served model matches the expected upstream model. The comparison
uses
model_obj.forwarded_model_id(the actual upstream ID, e.g.glm-5-2) rather thanmodel_obj.id(the client-facing alias, e.g.tinfoil-glm-5-2), so aliased models don't trigger a spurious mismatch. - When they genuinely differ (Tinfoil served a different upstream model than
expected), look up the actual served model's pricing and use it for billing.
The reverse lookup uses
get_model_instance, which resolvesforwarded_model_idvalues registered as routable aliases. - Log the discrepancy for observability.
If the actual model is not found in Routstr's model registry, billing falls back to the requested model's pricing.
What still needs verification
End-to-end test with a real Tinfoil SDK client against a Routstr node withVerified: both non-streaming (header) and streaming (trailer) responses includeTINFOIL_API_KEYset.model=<name>.- Streaming trailer capture is implemented by buffering the encrypted response
in
forward_with_trailer()and then using the dedicated EHBP payment finalizers for bearer and X-Cashu requests. This provides actual-cost billing today, at the cost of full time-to-last-byte latency for streaming responses. - Whether Tinfoil's
/v1/responsesendpoint also returns usage metrics headers or trailers. Verified: yes —/v1/responsesreturnsX-Tinfoil-Usage-Metricsas a plaintext response header, with the same field set as/v1/chat/completions(includingcost_usd).