13. Errors
Every error is the same envelope:
{"error": {
"code": "auth.missing_scope",
"message": "missing required scope: catalog:read",
"details": null,
"request_id": "01a06a15-d526-7310-afb6-27482d2ac2a0",
"trace_id": "d286c5a55c5b2a710a77f15a4a3b5557"}}Branch on code, never on message. Codes are a closed set; messages
improve over time.
| Code | HTTP | Meaning |
|---|---|---|
auth.unauthorized | 401 | no credential |
auth.invalid_api_key | 401 | unknown key |
auth.key_expired / auth.key_revoked | 401 | key no longer valid |
auth.missing_scope | 403 | key lacks a required scope |
auth.forbidden | 403 | authenticated but not permitted |
resource.not_found | 404 | unknown model, or not yours |
resource.conflict | 409 | already exists / concurrent change |
request.invalid | 400 | malformed request |
request.unsupported_parameter | 400 | provider cannot honour a parameter |
model.wrong_modality | 400 | the model serves another modality than this route (an image model sent to chat); details names model, kind, expected_kind, route |
request.unprocessable | 422 | well-formed but semantically wrong |
idempotency.key_required | 400 | mutating route needs the header |
idempotency.key_conflict | 409 | key reused with a different body |
rate_limit.exceeded | 429 | your rate limit |
quota.exceeded | 429 | a plan ceiling |
budget.exceeded | 429 | monthly allowance exhausted |
upstream.rate_limited | 429 | the provider throttled us |
upstream.timeout / upstream.unavailable / upstream.bad_gateway | 502/503/504 | provider failed |
server.internal | 500 | our fault |
server.not_implemented | 501 | route reserved, not built |
A 503 upstream.unavailable names its reason in details.upstream_error_code
when the platform's own account is the problem:
upstream_error_code | Meaning | What to do |
|---|---|---|
upstream.billing | the provider refused because Nozzle's account with it is out of credit or over its paid quota | fail over to another model or provider; retrying the same one will not help |
upstream.auth_failed | every platform credential for that provider is dead or quarantined | fail over |
Neither is your request's fault, and neither is a rate limit to wait out.
A buffered inference call may run for up to 900 s before the gateway answers upstream.timeout; the router's own upstream bound is 900 s as well and must never be set below the gateway's (both are NOZZLE_HTTP_REQUEST_TIMEOUT_SECS on the gateway and NOZZLE_UPSTREAM_TIMEOUT on the router, defaults pinned to the composes by test). Bound your own per-call timeouts at or below that, and use streaming for anything that legitimately runs long.
What to do with each refusal — a contract, not a hint:
| Answer | Means | Do |
|---|---|---|
429 upstream.rate_limited | the provider throttled us, or we benched it after it did | wait details.remaining_ms, then retry the same provider |
429 rate_limit.exceeded with details.remaining_ms | a window of yours is spent: the key's requests or tokens per minute, or the plan's requests per minute or per day | wait details.remaining_ms, then retry; another model on the same route draws on the same window |
429 rate_limit.exceeded without remaining_ms | a ceiling waiting does not clear: live instances, live streams, or one request larger than the key's whole tokens-per-minute window | change the request or raise the limit |
503 upstream.unavailable (any upstream_error_code, including upstream.billing and upstream.auth_failed) | the provider cannot serve right now | fail over to another provider; waiting will not help |
429 budget.exceeded / quota.exceeded | a spend or plan ceiling | raise the limit; retrying will not help |
Every 429 upstream.rate_limited, and every 429 rate_limit.exceeded that is a
window, carries details.remaining_ms (an integer, milliseconds) and a
Retry-After header that agrees with it, rounded up to whole seconds —
guaranteed, and pinned by tests on both the router and the gateway. So one rule
handles both: if remaining_ms is there, sleep exactly that long and send the
same request again.
For upstream.rate_limited the value is the provider's own Retry-After when
it sent one; otherwise how long the platform benches that provider after a 429
(5 s by default), which is exactly when the next call can reach it. Retrying
sooner reaches the same bench. When every candidate provider is benched,
details.providers lists each one and remaining_ms is the soonest to clear.
For rate_limit.exceeded it is when your window has refilled enough to admit
one more request. The windows themselves are published per model in the
limits block of GET /v1/models (§6).