7. Knowing what it cost
Every non-streaming inference response — chat, responses, embeddings, rerank, classify, speech, transcription, images (generations and edits), moderation and the Anthropic drop-in — carries:
| Header | Meaning |
|---|---|
X-Nozzle-Cost-Micro-Cents | what this call cost, integer micro-cents |
X-Nozzle-Provider | which backend actually served it |
X-Nozzle-Upstream-Model | the provider's own name for the model |
X-Request-Id | correlation id, also the ledger key |
Ratelimit-Limit / -Remaining / -Reset | your current budget |
Why micro-cents and not dollars: a typical call costs a few thousand
micro-cents — well under one cent. Any float-of-dollars encoding reports that
as 0.00. Integer micro-cents is lossless and reconciles exactly against
GET /v1/billing/costs.
0 is a real answer, not a missing one: a dedicated pod prices tokens at zero
on purpose, because you are paying per GPU-hour instead.
Each ledger row from GET /v1/billing/costs carries cost_micro_cents,
request_id, the token counts, model (as ref_id), provider,
virtual_key_id and estimated, so a call's header reconciles to its row by
request_id to the unit. cost_cents on the same row is the whole-cent
rounding and reads 0 for most single calls. provider is null on rows
written before 2026-09-22.
Look up one call — the exact reconciliation for a streamed request:
curl -s "$NOZZLE/v1/billing/costs/$REQUEST_ID" -H "Authorization: Bearer $KEY"
# → [ { "request_id": "...", "cost_micro_cents": 3300, "provider": "anthropic", ... } ]An array (a call that was billed and later corrected carries more than one
row); 404 resource.not_found for an id your tenant never made.
Page through the ledger — newest first, 200 rows by default, up to 1000
with limit. The body stays a JSON array; the next page's cursor rides the
X-Nozzle-Next-Cursor response header and is absent on the last page:
rows, cursor = [], None
while True:
params = {"limit": 1000, "from": "2026-09-01T00:00:00Z"}
if cursor: params["cursor"] = cursor
r = httpx.get(f"{NOZZLE}/v1/billing/costs", params=params, headers=auth)
rows += r.json()
cursor = r.headers.get("x-nozzle-next-cursor")
if not cursor: breakFilters, all optional and combined with AND: from (inclusive) and to
(exclusive) as RFC 3339, kind, model, provider, key_id, project_id.
An unknown parameter is refused with 400, never ignored. The gateway mints
the request_id it stamps on the row — reconcile by the X-Request-Id
response header, never by an id you sent.
Chart your traffic — GET /v1/analytics/usage (scope usage:read) groups
in SQL, so a dashboard never downloads raw rows to add them up:
# spend and calls per model per day, last 30 days
curl -s "$NOZZLE/v1/analytics/usage?group_by=model&interval=day&from=2026-08-23T00:00:00Z" \
-H "Authorization: Bearer $KEY"Give group_by (model, provider, key, project, kind), interval
(hour, day), or both for one series per key. The window is from/to
(RFC 3339, default the last 7 days, at most 400); the ledger's filters apply
too (kind, model, provider, key_id, project_id). Each row is
{bucket, key, calls, error_count, billed, prompt_tokens, completion_tokens, cached_tokens, cost_micro_cents, p50_latency_ms, p95_latency_ms} and the
response adds window totals and truncated (rows cap at 5,000).
Money and tokens are summed from the ledger, over its whole history.
calls, error_count and latency come from the call log, which records
every metered call including refused ones and starts on 2026-09-22 —
an older window shows spend with zero calls. Latency is time to the response
head: the whole call when buffered, time to first byte when streamed.
The gateway stamps, the client reads, and nothing re-derives a price.
Streaming has no cost header and cannot. The cost is known only when the
terminal usage frame arrives, long after headers flush. HTTP trailers exist for
this but the OpenAI and Anthropic SDKs cannot read them, so promising cost
there would promise something most clients cannot collect. Streamed callers
reconcile through X-Request-Id.