Section 12
12. Usage, billing and quotas
Every route below resolves your tenant from the credential — no ids in URLs.
| Endpoint | Answers |
|---|---|
GET /v1/usage | spend over 24h / 7d / 30d / all-time, plus per-kind breakdown |
GET /v1/billing/costs | per-request ledger rows, newest first; paged by X-Nozzle-Next-Cursor, filterable by time, kind, model, provider, key, project |
GET /v1/billing/costs/{request_id} | the ledger rows one call produced |
GET /v1/analytics/usage | calls, errors, tokens, cost and p50/p95 latency grouped by time and/or model, provider, key, project or kind |
GET /v1/billing/balance | prepaid credit balance |
GET /v1/billing/credits | credit ledger, paginated |
GET /v1/billing/entitlement | your plan and any overrides |
GET /v1/quotas | one slot per quota kind: usage and ceiling in one stated unit (see below) |
GET /v1/analytics/cost | per-day cost series |
GET /v1/analytics/instances | per-instance lifetime and spend |
GET /v1/analytics/margin | inference versus compute spend |
Each has an addressed twin at /v1/tenants/{tenant_id}/… for principals that
span tenants. They return identical data; the caller-scoped form exists because
the id could only ever have been your own.
Quotas
GET /v1/quotas always returns exactly one slot per quota kind, read through
the same code path that refuses a request, so what you read is what you are
enforced on. Within a slot, value and ceiling are in the same unit.
json
{"tenant_id":"…","plan_id":"enterprise","slots":[
{"kind":"monthly_credits","period_key":"2026-09","value":155640,"ceiling":null,
"unit":"micro_cents","measure":"spend","unlimited":true,"source":"api.cost_ledger","enforced":true},
{"kind":"max_concurrent_instances","period_key":"","value":0,"ceiling":null,
"unit":"instances","measure":"gauge","unlimited":true,"source":"compute.instances","enforced":true},
{"kind":"per_minute","period_key":"2026-09-21T14:05","value":12,"ceiling":600,
"unit":"requests","measure":"counter","unlimited":false,"source":"redis:quota_bucket","enforced":true},
{"kind":"max_model_size_b","period_key":"","value":0,"ceiling":9999,
"unit":"params_b","measure":"ceiling_only","unlimited":false,"source":"api.plans","enforced":true}
]}monthly_credits— this calendar month's spend summed from your cost ledger, against your plan allowance, both in micro-cents (1¢ = 1,000,000 µ¢).max_concurrent_instances— a live gauge of instances holding hardware.per_minute/per_day— your plan's request rate across every inference call from every key in the tenant, as rolling windows;valueis how much of the window is in use now. Refused with429 rate_limit.exceededanddetails.remaining_ms. The same ceilings appear in thelimitsblock ofGET /v1/models.max_model_size_b— bounds a launch request;valueis always0and is not a measurement (measure: "ceiling_only").ceiling: nullappears only withunlimited: true(an enterprise override).
Plans
| Plan | Monthly credits | Max concurrent instances | Requests / minute | Requests / day |
|---|---|---|---|---|
| free | 0¢ | 1 | 30 | 1,000 |
| hobby | 2,000¢ | 1 | 60 | 5,000 |
| pro | 20,000¢ | 5 | 600 | 50,000 |
| team | 100,000¢ | 20 | 6,000 | 500,000 |
| enterprise | 1,000,000¢ | 9,999 | unlimited | unlimited |
Exceeding the monthly allowance returns 429 budget.exceeded. The prepaid
balance is deliberately not enforced per call.