Section 7

7. Knowing what it cost

Every non-streaming inference response — chat, responses, embeddings, rerank, classify, speech, transcription, images (generations and edits), moderation and the Anthropic drop-in — carries:

HeaderMeaning
X-Nozzle-Cost-Micro-Centswhat this call cost, integer micro-cents
X-Nozzle-Providerwhich backend actually served it
X-Nozzle-Upstream-Modelthe provider's own name for the model
X-Request-Idcorrelation id, also the ledger key
Ratelimit-Limit / -Remaining / -Resetyour current budget

Why micro-cents and not dollars: a typical call costs a few thousand micro-cents — well under one cent. Any float-of-dollars encoding reports that as 0.00. Integer micro-cents is lossless and reconciles exactly against GET /v1/billing/costs.

0 is a real answer, not a missing one: a dedicated pod prices tokens at zero on purpose, because you are paying per GPU-hour instead.

Each ledger row from GET /v1/billing/costs carries cost_micro_cents, request_id, the token counts, model (as ref_id), provider, virtual_key_id and estimated, so a call's header reconciles to its row by request_id to the unit. cost_cents on the same row is the whole-cent rounding and reads 0 for most single calls. provider is null on rows written before 2026-09-22.

Look up one call — the exact reconciliation for a streamed request:

bash
curl -s "$NOZZLE/v1/billing/costs/$REQUEST_ID" -H "Authorization: Bearer $KEY"
# → [ { "request_id": "...", "cost_micro_cents": 3300, "provider": "anthropic", ... } ]

An array (a call that was billed and later corrected carries more than one row); 404 resource.not_found for an id your tenant never made.

Page through the ledger — newest first, 200 rows by default, up to 1000 with limit. The body stays a JSON array; the next page's cursor rides the X-Nozzle-Next-Cursor response header and is absent on the last page:

python
rows, cursor = [], None
while True:
    params = {"limit": 1000, "from": "2026-09-01T00:00:00Z"}
    if cursor: params["cursor"] = cursor
    r = httpx.get(f"{NOZZLE}/v1/billing/costs", params=params, headers=auth)
    rows += r.json()
    cursor = r.headers.get("x-nozzle-next-cursor")
    if not cursor: break

Filters, all optional and combined with AND: from (inclusive) and to (exclusive) as RFC 3339, kind, model, provider, key_id, project_id. An unknown parameter is refused with 400, never ignored. The gateway mints the request_id it stamps on the row — reconcile by the X-Request-Id response header, never by an id you sent.

Chart your traffic — GET /v1/analytics/usage (scope usage:read) groups in SQL, so a dashboard never downloads raw rows to add them up:

bash
# spend and calls per model per day, last 30 days
curl -s "$NOZZLE/v1/analytics/usage?group_by=model&interval=day&from=2026-08-23T00:00:00Z" \
  -H "Authorization: Bearer $KEY"

Give group_by (model, provider, key, project, kind), interval (hour, day), or both for one series per key. The window is from/to (RFC 3339, default the last 7 days, at most 400); the ledger's filters apply too (kind, model, provider, key_id, project_id). Each row is {bucket, key, calls, error_count, billed, prompt_tokens, completion_tokens, cached_tokens, cost_micro_cents, p50_latency_ms, p95_latency_ms} and the response adds window totals and truncated (rows cap at 5,000).

Money and tokens are summed from the ledger, over its whole history. calls, error_count and latency come from the call log, which records every metered call including refused ones and starts on 2026-09-22 — an older window shows spend with zero calls. Latency is time to the response head: the whole call when buffered, time to first byte when streamed.

The gateway stamps, the client reads, and nothing re-derives a price.

Streaming has no cost header and cannot. The cost is known only when the terminal usage frame arrives, long after headers flush. HTTP trailers exist for this but the OpenAI and Anthropic SDKs cannot read them, so promising cost there would promise something most clients cannot collect. Streamed callers reconcile through X-Request-Id.