Files
routstr-core/docs/api/errors.md
T
Paperclip Deployment Engineerandthefux 9d0cd41e5a fix(upstream): report upstream 5xx as 424 + UPSTREAM_UNAVAILABLE, not node-down
CORE-UPSTREAM-5XX-NOT-NODE-DOWN

An upstream-attributable failure (provider 5xx, EHBP timeout, transport error
after a Cashu token was redeemed) was forwarded to the caller as the
provider's own 5xx. Callers read that as *this node* being down: they marked
the node unhealthy, dropped it from rotation, or refused to retry an upstream
blip that the node had already retried across candidates.

The node is healthy in those cases — it accepted the request, authenticated
it, reserved payment, tried every candidate, and reverted the reservation.
That is now visible from the response alone:

  status            424 (UPSTREAM_ERROR_STATUS)
  error.code        UPSTREAM_UNAVAILABLE
  header            X-Routstr-Error-Scope: upstream
  error.upstream_status / error.details.upstream_status
                    the provider's own status, preserved

Applied in routstr/upstream/base.py (forward_upstream_error_response and the
post-redemption x-cashu paths for both chat-completions and responses),
routstr/payment/helpers.py (create_error_response / create_upstream_error_response),
routstr/proxy.py (424 is retryable across candidates on the bearer, EHBP and
unauthenticated-GET loops), routstr/upstream/ehbp.py and
routstr/upstream/tinfoil.py (attestation host). New module
routstr/core/error_scope.py holds the contract constants and the mapping.

Deliberately unchanged:
  * rate limits keep 429 + UPSTREAM_RATE_LIMIT (including a rate limit wrapped
    in a provider 5xx envelope: status 429, code unchanged) — the retry hint is
    worth more than the status class;
  * provider-side 4xx passes through unchanged;
  * genuine node faults stay 500 and carry no scope header (UpstreamError gained
    scope=, set to "node" on internal-exception paths) so "node broken" is still
    distinguishable from "upstream broken".

The new status is exported through CORS (x-routstr-error-scope) so browser
clients can read the attribution.

Tests: 15 stale assertions of the old 5xx contract updated (renames keep their
intent: refund still happens, bodies still redacted, pinned requests still do
not fall back), plus new acceptance coverage for 424 + scope header +
upstream_status on the provider, bearer, x-cashu, /v1/messages, EHBP and
unauthenticated-GET paths, node faults staying 500 with no header, and failover
past a 424 to a healthy candidate returning 200.

Docs: docs/api/errors.md (status table, new "Upstream attribution" section,
upstream error examples, retry list now includes 424), docs/api/overview.md
and docs/api/endpoints.md.

Full unit suite: 1649 passed, 1 skipped.
2026-09-23 15:16:44 +00:00

22 KiB

Error Handling

This guide covers error responses, codes, and handling strategies for the Routstr API.

Error Response Format

Versioned endpoints use a structured JSON error object:

{
  "error": {
    "type": "error_type",
    "message": "Human-readable error message",
    "code": "error_code",
    "details": {
      "additional": "context-specific information"
    }
  }
}

Lightning invoice errors

The /v2/lightning/* endpoints use the structured error object above. In the HTTP response, error is available at the top level and mirrored under detail.error. Clients should branch on error.code.

Endpoint case Status type code
Top-up without a credential 401 invalid_request_error topup_authorization_required
Top-up with a non-sk- credential 400 invalid_request_error topup_invalid_api_key_format
Top-up target key not found 404 invalid_request_error topup_api_key_not_found
Invoice not found during status or recovery 404 invalid_request_error invoice_not_found
Cashu mint rate-limited 503 mint_rate_limited lightning_mint_rate_limited
Cashu mint unreachable 503 mint_unreachable lightning_mint_unreachable
Unexpected invoice creation failure 500 api_error invoice_creation_failed

Request validation failures, including non-positive or excessive amounts, use FastAPI's standard 422 validation response. Only the 503 mint failures are retryable. Use backoff and honor the Retry-After header when present. The compatibility endpoints /lightning/* and /v1/balance/lightning/* retain their original string detail errors and legacy status behavior.

HTTP Status Codes

Status Meaning Common Causes
400 Bad Request Invalid parameters, malformed JSON
401 Unauthorized Invalid or missing API key
402 Payment Required Insufficient balance
403 Forbidden Access denied to resource
404 Not Found Endpoint or resource doesn't exist
422 Unprocessable Entity Validation errors
424 Failed Dependency An upstream inference provider failed. This node is healthy — see Upstream attribution
429 Too Many Requests Rate limit exceeded (this node or an upstream provider)
500 Internal Server Error Server-side error on this node
502 Bad Gateway Gateway-level failure
503 Service Unavailable Temporary outage

Upstream attribution (424 Failed Dependency)

Routstr fronts third-party inference providers. When one of them fails, this node is still healthy: it accepted the request, authenticated it, reserved payment, and reverted the reservation once the last candidate failed. Such failures are reported deliberately as a non-5xx status so a client does not mark the node down, drop it from rotation, or refuse to retry an upstream blip.

An upstream-attributable failure answers:

  • Status: 424
  • error.code: UPSTREAM_UNAVAILABLE
  • Header: X-Routstr-Error-Scope: upstream
  • error.upstream_status: the provider's own status (e.g. 503). Failures built by the payment helpers carry it in error.details.upstream_status instead
HTTP/1.1 424 Failed Dependency
X-Routstr-Error-Scope: upstream
Content-Type: application/json

{
  "error": {
    "type": "upstream_error",
    "message": "Service Unavailable",
    "code": "UPSTREAM_UNAVAILABLE",
    "upstream_status": 503
  }
}

Two exceptions keep their own status:

  • Rate limits answer 429 with error.code = UPSTREAM_RATE_LIMIT — the retry hint is worth more than the status class. This includes a rate limit a provider wrapped in a 5xx envelope: the status becomes 429 and the code stays UPSTREAM_RATE_LIMIT.
  • Provider-side 4xx (400/401/403/404/422) passes through unchanged: that is the provider's verdict on the request, not a health signal.

Genuine node faults are deliberately untouched: an unreachable mint, a database failure, or an internal exception while talking to a provider still answers 500 and carries no X-Routstr-Error-Scope header. That is what makes "node healthy, upstream failed" distinguishable from "node broken" from the response alone.

Error Types

Authentication Errors

Invalid API Key

{
  "error": {
    "type": "authentication_failed",
    "message": "Invalid API key provided",
    "code": "invalid_api_key"
  }
}

Status: 401
Resolution: Check API key format and validity

Expired API Key

{
  "error": {
    "type": "authentication_failed",
    "message": "API key has expired",
    "code": "key_expired",
    "details": {
      "expired_at": "2024-01-01T00:00:00Z",
      "refund_available": true
    }
  }
}

Status: 401
Resolution: Create new API key or contact admin for refund

Missing Authorization

{
  "error": {
    "type": "authentication_failed",
    "message": "Authorization header required",
    "code": "missing_auth"
  }
}

Status: 401
Resolution: Include Authorization: Bearer {api_key} header

Payment Errors

Insufficient Balance

{
  "error": {
    "type": "insufficient_balance",
    "message": "Insufficient balance for request",
    "code": "payment_required",
    "details": {
      "balance": 100,
      "required": 154,
      "shortfall": 54,
      "estimated_tokens": {
        "prompt": 50,
        "completion": 150
      }
    }
  }
}

Status: 402
Resolution: Top up API key balance

Cashu Token Redemption Errors

These errors are returned when a Cashu token you pay with cannot be redeemed. They apply to every endpoint that accepts a token:

  • Per-request payment via the X-Cashu header (chat completions + Responses API).
  • API key top-up via POST /v1/wallet/topup.
  • Minting an API key from a token sent in Authorization: Bearer <cashu-token>.

All three share one classifier, so the same failure yields the same HTTP status and sanitized message everywhere. All three paths also expose the same structured type and code — branch on type (or code for finer granularity) on any of them.

type Status code Retryable Meaning
token_already_spent 400 cashu_token_already_spent No The token was already redeemed.
invalid_token 400 invalid_cashu_token No The token is malformed or cannot be decoded.
mint_error 422 cashu_token_swap_fees_exceed_amount No Token value is too small to cover the mint's NUT-02 input fees.
untrusted_mint 400 cashu_untrusted_source_mint No The token was issued by a mint this node does not accept. Only the node's configured mints (PRIMARY_MINT_URL / CASHU_MINTS) are redeemable.
mint_unreachable 503 cashu_source_mint_unreachable Yes The mint that issued the token could not be reached; it cannot be redeemed at another mint.
mint_rate_limited 503 cashu_mint_rate_limited Yes The mint rate-limited the request; retry after the cooldown.
mint_timeout 503 cashu_mint_timeout Yes The mint did not respond in time; retry later.
mint_unreachable 503 cashu_mint_unreachable Yes The mint could not be reached (DNS failure, refused/reset connection). The token is fine — retry once the mint recovers.
cashu_error 400 cashu_token_redemption_failed No The token could not be redeemed for another expected reason.
cashu_error 400 cashu_token_zero_value No The token redeemed to zero (empty/dust token, or value fully consumed by fees).
token_consumed 500 cashu_token_consumed No The token was spent (melted/redeemed) but crediting it then failed. Do not retry — the token is gone; contact support to reconcile.
api_error 500 internal_error Maybe Unexpected server-side fault during redemption.

!!! important "Retry only transient mint failures" Only mint_unreachable, mint_rate_limited and mint_timeout (503) are retryable — the same token may work again later. Everything else is a permanent property of the token and must not be blindly retried. untrusted_mint is permanent: the node will never accept that mint until an operator adds it to CASHU_MINTS. Use exponential backoff for the 503 responses, and honor the mint's cooldown for mint_rate_limited. In particular, a token_consumed 500 means the mint already spent the token, so a retry would fail as token_already_spent.

Mint failures (retryable)

mint_unreachable, mint_rate_limited and mint_timeout are retryable redemption errors. For mint_rate_limited, honor the mint's cooldown before retrying.

{
  "error": {
    "type": "mint_unreachable",
    "message": "Cashu mint is unreachable",
    "code": "cashu_mint_unreachable"
  }
}

Status: 503

Resolution: The token is valid — the mint is temporarily down. Retry with backoff, or pay with a token from a different mint. If the mint returns mint_rate_limited, wait for its cooldown before retrying.

Token Already Spent

{
  "error": {
    "type": "token_already_spent",
    "message": "Cashu token already spent",
    "code": "cashu_token_already_spent"
  }
}

Status: 400

Resolution: Use a fresh, unspent token. Do not retry with the same token.

Response envelope differs by endpoint

The error object above is identical everywhere, but the surrounding envelope depends on how you paid:

  • X-Cashu header payments (chat + Responses API) return the object at the top level, alongside a request_id:

    {
      "error": { "type": "mint_unreachable", "message": "Cashu mint is unreachable", "code": "cashu_mint_unreachable" },
      "request_id": "req-abc123"
    }
    

    The original token is echoed back in the X-Cashu response header only when it is still spendable (e.g. mint_unreachable, mint_rate_limited, invalid_cashu_token, fee errors) so you can recover/retry it. It is not echoed for spent/consumed tokens (cashu_token_already_spent, cashu_token_consumed, cashu_token_zero_value, internal_error) — retrying those can never succeed.

  • Authorization: Bearer <cashu-token> (API key minting) wraps it in FastAPI's detail field:

    { "detail": { "error": { "type": "mint_unreachable", "message": "Cashu mint is unreachable", "code": "cashu_mint_unreachable" } } }
    
  • POST /v1/wallet/topup returns the same structured envelope wrapped in FastAPI's detail field, identical to the bearer path:

    { "detail": { "error": { "type": "mint_unreachable", "message": "Cashu mint is unreachable", "code": "cashu_mint_unreachable" } } }
    

Validation Errors

Invalid Parameters

{
  "error": {
    "type": "invalid_request",
    "message": "Invalid request parameters",
    "code": "validation_error",
    "details": {
      "errors": [
        {
          "field": "temperature",
          "message": "Must be between 0 and 2",
          "value": 3.5
        },
        {
          "field": "model",
          "message": "Model 'gpt-5' not found",
          "value": "gpt-5"
        }
      ]
    }
  }
}

Status: 422
Resolution: Fix parameter values

Missing Required Fields

{
  "error": {
    "type": "invalid_request",
    "message": "Missing required fields",
    "code": "missing_fields",
    "details": {
      "missing": ["model", "messages"]
    }
  }
}

Status: 400
Resolution: Include all required fields

Rate Limiting

Rate Limit Exceeded

{
  "error": {
    "type": "rate_limit_exceeded",
    "message": "Too many requests",
    "code": "rate_limit",
    "details": {
      "limit": 100,
      "window": "1 minute",
      "retry_after": 45
    }
  }
}

Status: 429
Headers:

X-RateLimit-Limit: 100
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 1640995200
Retry-After: 45

Resolution: Wait for retry_after seconds

Upstream Errors

Upstream Unavailable

A provider returned a 5xx (overloaded, bad gateway, timeout, or a provider-side outage). This node is healthy and your reservation has been reverted.

{
  "error": {
    "type": "upstream_error",
    "message": "Model is currently overloaded",
    "code": "UPSTREAM_UNAVAILABLE",
    "upstream_status": 503,
    "details": {
      "model": "gpt-4",
      "retry_after": 5
    }
  }
}

Status: 424
Header: X-Routstr-Error-Scope: upstream
Resolution: Retry after a short backoff. If the node is configured with alternative providers for the model, it already retried them before answering — try another model or provider path if the failure persists.

Upstream Timeout

{
  "error": {
    "type": "upstream_error",
    "message": "Request to upstream API timed out",
    "code": "upstream_timeout",
    "details": {
      "timeout": 30,
      "endpoint": "chat/completions"
    }
  }
}

Status: 424 (deliberately non-5xx: the timeout happened on the provider hop, not on this node)
Header: X-Routstr-Error-Scope: upstream
Resolution: Retry with shorter prompt or max_tokens

Content Policy

Content Filtered

{
  "error": {
    "type": "content_policy_violation",
    "message": "Content filtered due to policy violation",
    "code": "content_filtered",
    "details": {
      "reason": "harmful_content",
      "categories": ["violence", "hate"]
    }
  }
}

Status: 400
Resolution: Modify prompt to comply with policies

Error Handling Best Practices

Retry Logic

Implement exponential backoff with jitter:

import time
import random
from typing import Optional, Callable

def retry_with_backoff(
    func: Callable,
    max_retries: int = 3,
    base_delay: float = 1.0,
    max_delay: float = 60.0
) -> Optional[Any]:
    """Retry function with exponential backoff."""
    
    for attempt in range(max_retries):
        try:
            return func()
        except Exception as e:
            if attempt == max_retries - 1:
                raise
            
            # Check if error is retryable
            if hasattr(e, 'status_code'):
                # 424 is the upstream-attribution status: it is retryable
                # exactly like the 5xx statuses it replaced.
                if e.status_code in [424, 429, 502, 503, 504]:
                    # Calculate delay with jitter
                    delay = min(
                        base_delay * (2 ** attempt) + random.uniform(0, 1),
                        max_delay
                    )
                    
                    # Use retry_after if provided
                    if hasattr(e, 'retry_after'):
                        delay = e.retry_after
                    
                    time.sleep(delay)
                else:
                    # Non-retryable error
                    raise

Error Categories

Group errors for handling:

class ErrorHandler:
    # Errors that should be retried
    RETRYABLE_ERRORS = {
        'UPSTREAM_UNAVAILABLE',  # upstream 5xx, reported as HTTP 424
        'UPSTREAM_RATE_LIMIT',   # HTTP 429
        'UPSTREAM_TIMEOUT',      # EHBP upstream timeout, reported as HTTP 424
        'rate_limit',
        'upstream_timeout',
        'model_overloaded',
        'cashu_mint_unreachable'
    }
    
    # Errors requiring user action
    USER_ACTION_ERRORS = {
        'insufficient_balance',
        'invalid_api_key',
        'key_expired'
    }
    
    # Errors requiring code changes
    CLIENT_ERRORS = {
        'validation_error',
        'missing_fields',
        'invalid_request'
    }
    
    @classmethod
    def handle_error(cls, error_response: dict) -> None:
        error_code = error_response['error']['code']
        
        if error_code in cls.RETRYABLE_ERRORS:
            # Implement retry logic
            pass
        elif error_code in cls.USER_ACTION_ERRORS:
            # Alert user
            pass
        elif error_code in cls.CLIENT_ERRORS:
            # Log for debugging
            pass

Graceful Degradation

Handle errors without breaking application flow:

async def get_ai_response(prompt: str) -> str:
    """Get AI response with fallback handling."""
    try:
        # Try primary model
        response = await client.chat.completions.create(
            model="gpt-4",
            messages=[{"role": "user", "content": prompt}]
        )
        return response.choices[0].message.content
    
    except InsufficientBalanceError:
        # Fall back to cheaper model
        try:
            response = await client.chat.completions.create(
                model="gpt-3.5-turbo",
                messages=[{"role": "user", "content": prompt}],
                max_tokens=100  # Limit tokens
            )
            return response.choices[0].message.content
        except Exception as e:
            logger.error(f"Fallback failed: {e}")
            return "Service temporarily unavailable"
    
    except Exception as e:
        logger.error(f"Unexpected error: {e}")
        return "An error occurred processing your request"

Logging Errors

Structure error logs for debugging:

import logging
import json

def log_api_error(error_response: dict, context: dict) -> None:
    """Log API errors with context."""
    logger = logging.getLogger(__name__)
    
    error_data = {
        'timestamp': datetime.utcnow().isoformat(),
        'error': error_response['error'],
        'context': {
            'endpoint': context.get('endpoint'),
            'api_key_id': context.get('api_key_id'),
            'request_id': context.get('request_id'),
            'model': context.get('model')
        }
    }
    
    logger.error(
        "API Error",
        extra={'structured_data': json.dumps(error_data)}
    )

User-Friendly Messages

Map technical errors to user messages:

ERROR_MESSAGES = {
    'insufficient_balance': "Your account balance is too low. Please add funds to continue.",
    'invalid_api_key': "Invalid API key. Please check your configuration.",
    'rate_limit': "Too many requests. Please wait a moment and try again.",
    'model_overloaded': "The AI service is busy. Please try again in a few seconds.",
    'validation_error': "Invalid request. Please check your input and try again."
}

def get_user_message(error_code: str) -> str:
    """Get user-friendly error message."""
    return ERROR_MESSAGES.get(
        error_code,
        "An unexpected error occurred. Please try again later."
    )

Common Scenarios

Handling Balance Errors

async def make_request_with_balance_check():
    try:
        # Check balance first
        balance_info = await client.get("/v1/wallet/balance")
        
        # Estimate cost
        estimated_cost = calculate_cost(model, prompt_length)
        
        if balance_info['balance'] < estimated_cost * 1.1:  # 10% buffer
            # Proactively top up
            await top_up_balance()
        
        # Make request
        return await client.chat.completions.create(...)
        
    except InsufficientBalanceError as e:
        # Handle insufficient balance
        shortfall = e.details['shortfall']
        await top_up_balance(amount=shortfall * 2)
        # Retry request

Handling Rate Limits

from datetime import datetime, timedelta

class RateLimitTracker:
    def __init__(self):
        self.reset_times = {}
    
    def is_limited(self, endpoint: str) -> bool:
        reset_time = self.reset_times.get(endpoint)
        if reset_time and datetime.now() < reset_time:
            return True
        return False
    
    def set_limit(self, endpoint: str, reset_timestamp: int):
        self.reset_times[endpoint] = datetime.fromtimestamp(reset_timestamp)
    
    def wait_time(self, endpoint: str) -> float:
        reset_time = self.reset_times.get(endpoint)
        if reset_time:
            return max(0, (reset_time - datetime.now()).total_seconds())
        return 0

Testing Error Handling

Unit Tests

import pytest
from unittest.mock import Mock

async def test_insufficient_balance_handling():
    # Mock API client
    mock_client = Mock()
    mock_client.chat.completions.create.side_effect = InsufficientBalanceError(
        required=100,
        available=50
    )
    
    # Test error handling
    handler = ErrorHandler(mock_client)
    result = await handler.safe_request(
        model="gpt-4",
        messages=[{"role": "user", "content": "test"}]
    )
    
    # Verify fallback behavior
    assert result.fallback_used is True
    assert result.model == "gpt-3.5-turbo"

Integration Tests

async def test_real_error_scenarios():
    # Test with invalid API key
    invalid_client = OpenAI(
        api_key="sk-invalid",
        base_url=test_url
    )
    
    with pytest.raises(AuthenticationError) as exc_info:
        await invalid_client.chat.completions.create(
            model="gpt-3.5-turbo",
            messages=[{"role": "user", "content": "test"}]
        )
    
    assert exc_info.value.status_code == 401
    assert "invalid_api_key" in str(exc_info.value)

Monitoring Errors

Track error rates and patterns:

class ErrorMetrics:
    def __init__(self):
        self.error_counts = defaultdict(int)
        self.error_timestamps = defaultdict(list)
    
    def record_error(self, error_code: str):
        self.error_counts[error_code] += 1
        self.error_timestamps[error_code].append(datetime.now())
    
    def get_error_rate(self, error_code: str, window_minutes: int = 60) -> float:
        cutoff = datetime.now() - timedelta(minutes=window_minutes)
        recent_errors = [
            ts for ts in self.error_timestamps[error_code]
            if ts > cutoff
        ]
        return len(recent_errors) / window_minutes

Next Steps