Section 13

13. Errors

Every error is the same envelope:

json
{"error": {
  "code": "auth.missing_scope",
  "message": "missing required scope: catalog:read",
  "details": null,
  "request_id": "01a06a15-d526-7310-afb6-27482d2ac2a0",
  "trace_id": "d286c5a55c5b2a710a77f15a4a3b5557"}}

Branch on code, never on message. Codes are a closed set; messages improve over time.

CodeHTTPMeaning
auth.unauthorized401no credential
auth.invalid_api_key401unknown key
auth.key_expired / auth.key_revoked401key no longer valid
auth.missing_scope403key lacks a required scope
auth.forbidden403authenticated but not permitted
resource.not_found404unknown model, or not yours
resource.conflict409already exists / concurrent change
request.invalid400malformed request
request.unsupported_parameter400provider cannot honour a parameter
model.wrong_modality400the model serves another modality than this route (an image model sent to chat); details names model, kind, expected_kind, route
request.unprocessable422well-formed but semantically wrong
idempotency.key_required400mutating route needs the header
idempotency.key_conflict409key reused with a different body
rate_limit.exceeded429your rate limit
quota.exceeded429a plan ceiling
budget.exceeded429monthly allowance exhausted
upstream.rate_limited429the provider throttled us
upstream.timeout / upstream.unavailable / upstream.bad_gateway502/503/504provider failed
server.internal500our fault
server.not_implemented501route reserved, not built

A 503 upstream.unavailable names its reason in details.upstream_error_code when the platform's own account is the problem:

upstream_error_codeMeaningWhat to do
upstream.billingthe provider refused because Nozzle's account with it is out of credit or over its paid quotafail over to another model or provider; retrying the same one will not help
upstream.auth_failedevery platform credential for that provider is dead or quarantinedfail over

Neither is your request's fault, and neither is a rate limit to wait out.

A buffered inference call may run for up to 900 s before the gateway answers upstream.timeout; the router's own upstream bound is 900 s as well and must never be set below the gateway's (both are NOZZLE_HTTP_REQUEST_TIMEOUT_SECS on the gateway and NOZZLE_UPSTREAM_TIMEOUT on the router, defaults pinned to the composes by test). Bound your own per-call timeouts at or below that, and use streaming for anything that legitimately runs long.

What to do with each refusal — a contract, not a hint:

AnswerMeansDo
429 upstream.rate_limitedthe provider throttled us, or we benched it after it didwait details.remaining_ms, then retry the same provider
429 rate_limit.exceeded with details.remaining_msa window of yours is spent: the key's requests or tokens per minute, or the plan's requests per minute or per daywait details.remaining_ms, then retry; another model on the same route draws on the same window
429 rate_limit.exceeded without remaining_msa ceiling waiting does not clear: live instances, live streams, or one request larger than the key's whole tokens-per-minute windowchange the request or raise the limit
503 upstream.unavailable (any upstream_error_code, including upstream.billing and upstream.auth_failed)the provider cannot serve right nowfail over to another provider; waiting will not help
429 budget.exceeded / quota.exceededa spend or plan ceilingraise the limit; retrying will not help

Every 429 upstream.rate_limited, and every 429 rate_limit.exceeded that is a window, carries details.remaining_ms (an integer, milliseconds) and a Retry-After header that agrees with it, rounded up to whole seconds — guaranteed, and pinned by tests on both the router and the gateway. So one rule handles both: if remaining_ms is there, sleep exactly that long and send the same request again.

For upstream.rate_limited the value is the provider's own Retry-After when it sent one; otherwise how long the platform benches that provider after a 429 (5 s by default), which is exactly when the next call can reach it. Retrying sooner reaches the same bench. When every candidate provider is benched, details.providers lists each one and remaining_ms is the soonest to clear. For rate_limit.exceeded it is when your window has refilled enough to admit one more request. The windows themselves are published per model in the limits block of GET /v1/models (§6).