6. Calling models
Available surfaces
| Endpoint | Shape |
|---|---|
POST /v1/chat/completions | OpenAI chat, streaming supported |
POST /v1/responses | OpenAI Responses API — reasoning items and function tools together, streaming supported (see below) |
POST /v1/completions | OpenAI legacy text |
POST /v1/embeddings | OpenAI embeddings |
POST /v1/rerank | rerank documents against a query |
POST /v1/classify | classification |
POST /v1/moderations | content moderation |
POST /v1/judgments | typed judgments — choice / score / noul questions about a state, calibrated answers (TypeSafe's System One grammar; see below) |
POST /v1/images/generations | image generation — DALL·E or the token-billed gpt-image dialect. Bound today: gpt-image-2, gpt-image-2.5-flare, gpt-image-2.5-sunburst. size is `1024x1024 |
POST /v1/images/edits | image edits (multipart): up to 16 reference image[] parts, optional mask; gpt-image dialect only. `input_fidelity low |
POST /v1/audio/speech | text to speech |
POST /v1/audio/transcriptions | transcription (multipart) — json, text, verbose_json, srt, vtt |
POST /v1/audio/translations | the same, with the output pinned to English (whisper only) |
POST /anthropic/v1/messages | Anthropic Messages drop-in |
Discovering models
curl "$BASE/v1/models" -H "Authorization: Bearer $KEY"Returns the OpenAI-shaped list of everything you can call — platform models
plus any you registered yourself, plus any dedicated pods running in your
tenant. If a model answers /v1/chat/completions, it appears here.
Each row (and GET /v1/models/{id}) also carries limits: the ceilings the
gateway will hold this credential to when it calls that model, read through
the same code that refuses the call — size your parallelism from these instead
of discovering them as 429s. null means no ceiling of that kind.
{"id": "gpt-4.1-mini", "object": "model", "created": 0, "owned_by": "nozzle",
"limits": {"requests_per_minute": 6000, "tokens_per_minute": null,
"tenant_requests_per_minute": null, "tenant_requests_per_day": null}}| Field | Scope | Counted |
|---|---|---|
requests_per_minute | this key (a session gets the platform default) | separately per endpoint path — /v1/chat/completions and /v1/embeddings are two windows, and switching models on one path shares its window |
tokens_per_minute | this key | across every inference call, charged at an upper-bound estimate before the call and settled to actual usage after |
tenant_requests_per_minute | your plan, shared by every key in the tenant | across every inference call, rolling minute |
tenant_requests_per_day | your plan, shared by every key in the tenant | across every inference call, rolling 24 hours |
All four are refill windows, not fixed buckets that reset on the minute: a
burst up to the ceiling is admitted at once, then capacity returns at
ceiling/60 per second (ceiling/86,400 for the day). Exceeding any of them is
429 rate_limit.exceeded with details.remaining_ms (§13). The rows carry the
same numbers today; per-model ceilings would land in the same block.
What is not here, on purpose: the providers' own limits, and how the router
treats a provider after one (it benches a provider that answered 429 for a few
seconds and rotates across the platform's keys for it). Those live in the
router's configuration, which this answer cannot read, and publishing a guess
would be a number nothing enforces. A provider's limit reaches you as
429 upstream.rate_limited carrying the exact wait in details.remaining_ms.
For richer detail:
GET /v1/catalog/models— the full browsable catalog (3,500+ rows, 100 per page,?limit=up to 500,?offset=). Each row carries list pricing, the deployability tier, and what you resolve to:bound(can this key call it now),provider,upstream_model,kind(which surface serves it —openai_chat,openai_embed,openai_transcription,openai_image, …) andwire_capabilities, the same booleans the capabilities door answers. Those four arenullwhenboundis false. Filters, all AND-ed:?provider=openai,?modality=embedding,?q=qwen(substring of id or display name),?min_context=128000,?max_input_cents_per_mtok=100,?bound=true,?deployability=callable. A picker needs one call:GET /v1/catalog/models?bound=true.GET /v1/models/{id}/capabilities— the exact parameter support the router enforces for that model, answered for its binding kind: an embedding or image model does not advertise tools, vision or reasoning, and a transcription model advertises onlyaudio_in,timestampsandtranslation. The answer names thekind.GET /v1/models/{id}/health— what that model's serving lane did last:statusishealthy,degraded,deadorunknown, besidelast_success_at,last_failure_at,last_failure_codeandconsecutive_failures. Derived from every call the gateway forwards, so a lane whose provider stopped answering readsdeadwithout anyone probing it. Only the lane's own failures count (unreachable, timed out, refused our credential, 5xx); your own 400s, rate limits and budget refusals do not.
Capabilities are refused by name, not silently dropped
If you send a parameter a provider does not support, Nozzle refuses with a typed error naming the parameter — rather than stripping it and returning a plausible answer computed under different settings.
{"error": {
"code": "request.unsupported_parameter",
"message": "provider does not support these parameters",
"details": {"provider": "cerebras", "model": "gemma-4-31b",
"unsupported_parameters": ["messages[].content.image_url"]}}}This matters most for prompt caching: silently dropping a cache_control
marker costs you the entire cache discount with no way to notice.
Transcription
Send multipart: a file part plus model, and optionally language, prompt,
temperature, response_format and timestamp_granularities[].
curl -X POST "$BASE/v1/audio/transcriptions" -H "Authorization: Bearer $KEY" \
-F file=@meeting.m4a -F model=whisper-1 -F response_format=srtUploads are capped at 25 MB, refused with request.payload_too_large carrying
details.limit_bytes. Containers: flac, m4a, mp3, mp4, mpeg, mpga, ogg, wav,
webm.
Nozzle renders every output format itself from one upstream answer, so srt
and vtt are byte-identical in structure whichever backend served you — a
hosted provider, our shared pod, or a model you registered. verbose_json
relays the provider's own object so its extras survive.
Not every model can time a transcript. OpenAI's gpt-4o-transcribe family
and gpt-transcribe answer only json and text; whisper-1 also serves
verbose_json, srt, vtt, word timestamps and translations. Ask a model
before you send:
curl "$BASE/v1/models/gpt-4o-transcribe/capabilities" -H "Authorization: Bearer $KEY"
# → "capabilities": { …, "timestamps": false, "translation": false }A format a model cannot produce is refused by name with
request.unsupported_parameter and
details.unsupported_parameters: ["response_format.srt"] — never an empty
subtitle file. The same applies to timestamp_granularities and to sending a
json-only model to /v1/audio/translations.
Billing is per minute of audio, on every response format including text.
The length comes from the provider when it reports one, and otherwise from the
upload itself, which Nozzle measures. A model that reports no length and an
upload that cannot be measured is refused (request.invalid, parameter: file)
rather than served free. X-Nozzle-Cost-Micro-Cents is on the response, and it
reconciles exactly against GET /v1/billing/costs.
Streaming
Set "stream": true. Frames relay byte-for-byte from the provider, so your
SDK parses exactly what that provider produced. Add
"stream_options": {"include_usage": true} for a terminal usage frame.
Disconnecting aborts the upstream request — the model stops generating and stops costing money.
Vision, embeddings and reasoning models
- Vision is
image_urlparts (adata:URI or anhttpsURL) ongpt-4.1-mini,gemini-2.5-flashand every Claude model; through the Anthropic drop-in an image is a base64 orurlsourceblock. Cerebras models refuse it by name.GET /v1/models/{id}/capabilitiesreportsvisionper model, honouring provider overrides. Files, PDFs and audio: next section. - Embeddings are
text-embedding-3-smallby default: 1536 dimensions, 8191-token input.dimensionsshortens every vector to the width you ask for — passed to OpenAI and to self-hosted pods as-is, and to Vertex as itsoutputDimensionality. A provider that cannot shorten vectors (Cohere) refuses the call with400 request.unsupported_parameternamingdimensions; it is never dropped, because a vector of the wrong width silently corrupts every index it is written into.dimensions: 0is a 400. - Reasoning models (
cerebras/qwen-3.8-27b,cerebras/gpt-oss-120b, the gpt-5.x family, Claude withthinking) spendmax_tokenson reasoning first. Give them at least 300 or you get an empty200withfinish_reason: lengththat still bills. Omit the ceiling entirely and the model's ownmax_output_tokensfrom the catalog row is used, not an SDK's hidden 4096. - gpt-5.x with tools: set
reasoning_effortexplicitly on/v1/chat/completionsor the upstream refuses; tools and reasoning together live on/v1/responses(below).
Sending files, PDFs, images and audio
Attach a file as a content part on a user message, beside your text, in the
order you want the model to read them. Send the bytes as a base64 data: URI,
or give a public https URL and Nozzle fetches it for the models that cannot.
The inline ceiling is 32 MB per request; past it the answer is
413 request.payload_too_large.
import base64
from openai import OpenAI
client = OpenAI(base_url="https://api.opennozzle.com/v1", api_key=NOZZLE_KEY)
pdf = base64.b64encode(open("report.pdf", "rb").read()).decode()
r = client.chat.completions.create(
model="gemini-2.5-flash", # or claude-haiku-4-5, gpt-4.1-mini
messages=[{"role": "user", "content": [
{"type": "file", "file": {"filename": "report.pdf",
"file_data": f"data:application/pdf;base64,{pdf}"}},
{"type": "text", "text": "Summarize the findings."},
]}],
)| Part | Shape | Read by |
|---|---|---|
| Image | {"type":"image_url","image_url":{"url":"data:image/png;base64,…" | "https://…"}} | every vision model |
{"type":"file","file":{"filename":"a.pdf","file_data":"data:application/pdf;base64,…" | "https://…"}} | Claude, Gemini, OpenAI | |
| Text file | file part with text/plain, text/markdown, text/csv … | Claude, Gemini (Bedrock also reads Office formats) |
| Audio | {"type":"input_audio","input_audio":{"data":"<base64>","format":"wav"|"mp3"}}, or a file part with an audio/* type | Gemini; OpenAI audio models |
| Video | a file part with a video/* type (data: URI or URL) | Gemini |
A file part's type decides what it is: an audio/* or video/* file is
audio or video, whatever it is called. The type is read from the bytes when
they carry a signature, so a PNG labelled image/jpeg still goes as a PNG. A
model that cannot read what you sent refuses the request by name
(messages[].content.file, …input_audio, …video) — it is never answered as
if the attachment were not there. GET /v1/models/{id}/capabilities reports
pdf_in, audio_in and video_in. file_id references are refused until
uploads exist; send the file inline.
Through the Anthropic drop-in, a PDF is a Messages document block:
curl -s https://api.opennozzle.com/anthropic/v1/messages \
-H "x-api-key: $NOZZLE_KEY" -H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" -d '{
"model": "claude-haiku-4-5", "max_tokens": 300,
"messages": [{"role": "user", "content": [
{"type": "document", "source": {"type": "base64", "media_type": "application/pdf", "data": "'"$PDF_B64"'"}},
{"type": "text", "text": "Summarize this."}]}]}'url and text document sources work too; citations is refused by name.
Prompt caching
Anthropic cache_control markers survive only through the drop-in
(/anthropic/v1/messages): a 5.4k-token system block wrote
cache_creation_input_tokens: 5411 on the first call and read
cache_read_input_tokens: 5411 on the second, buffered and streamed,
incremental across a tool loop (measured 2026-09-10 and 2026-09-17). OpenAI and
Cerebras cache automatically with no marker and report cached_tokens. The
catalog publishes the cache rates beside input and output.
Using the litellm SDK
import litellm
litellm.drop_params = True # the drop-in denies unknown fields; pin your SDK version
BASE = "https://api.opennozzle.com"
# Claude: the /anthropic base is the only path where cache_control survives
litellm.completion(model="anthropic/claude-haiku-4-5", api_base=f"{BASE}/anthropic",
api_key="pk_live_…", messages=[...])
# everything else, OpenAI-shaped
litellm.completion(model="openai/gpt-4.1-mini", api_base=f"{BASE}/v1",
api_key="pk_live_…", messages=[...])The
anthropic/andopenai/prefixes above are litellm's, not Nozzle's. They tell the SDK which dialect to speak, and the SDK strips them before the request leaves your process. Nozzle resolves the model name verbatim, so a raw call must send the plain id:{"model": "claude-haiku-4-5"}toPOST /anthropic/v1/messages. Sendinganthropic/claude-haiku-4-5on the wire is404 resource.not_found— there is no binding by that name.
The cost headers surface as
response._hidden_params["additional_headers"]["llm_provider-x-nozzle-cost-micro-cents"];
the unprefixed name returns None.
The Responses API
POST /v1/responses serves OpenAI's Responses grammar — input items,
instructions, function tools, reasoning: {effort}, previous_response_id,
max_output_tokens — relayed verbatim to the model's upstream. It exists
because OpenAI's reasoning models (the gpt-5.x family) refuse function tools
together with reasoning_effort on /v1/chat/completions and serve both
only here, so an agent that wants tools and reasoning has no other door.
curl "$BASE/v1/responses" -H "Authorization: Bearer $KEY" \
-H "Content-Type: application/json" -d '{
"model": "gpt-5.6-luna",
"input": "What is the weather in Paris? Use the tool.",
"reasoning": {"effort": "low"},
"tools": [{"type": "function", "name": "get_weather",
"parameters": {"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"]}}]
}'What is Nozzle's and what is not:
- Which models answer. Any chat-kind model whose provider serves
/v1/responses: OpenAI proper by default, or any provider row an operator markedresponses: true. AskGET /v1/models/{id}/capabilities— theresponsesaxis is the same fact the router registers on. A model on a provider that does not serve it is refused asrequest.unsupported_parameteronmodel, naming/v1/chat/completions. - Two edits, nothing else.
modelis rewritten to the upstream's own name and the binding'stemperature/top_p/max_tokensoverrides are applied in the Responses spelling (max_tokens→max_output_tokens). Everything else in the body reaches the upstream as you wrote it, and the reply comes back as the upstream produced it. - Streaming.
"stream": truereturnstext/event-streamin the Responses event vocabulary: every event is named by itstype(response.created,response.output_text.delta,response.function_call_arguments.delta, …) and the stream ends atresponse.completed, which carriesusage. There is no[DONE]sentinel — the grammar does not define one. - No fallbacks. A Responses call is stateful across turns
(
previous_response_id, reasoning items) and bound to the upstream that minted that state, so inlinefallbacksare refused by name rather than ignored. - Billing is identical to chat:
usage.input_tokens/usage.output_tokens(output includes reasoning) at the model's chat rates, with the cached-input discount frominput_tokens_details.cached_tokens. A buffered call carriesX-Nozzle-Cost-Micro-Cents; a streamed one reconciles byX-Request-IdagainstGET /v1/billing/costs, where the row'skindisinference.responses.
Typed judgments
POST /v1/judgments answers typed questions about a state with calibrated
probabilities instead of generated text. The grammar is TypeSafe's System One,
served verbatim — a System One client switches by sending the same body
to $BASE/v1/judgments with a Nozzle key in place of its TypeSafe key. Bound today: jev-latest
(TypeSafe's alias for its current stable Jev; the reply's model names the
version that answered).
curl "$BASE/v1/judgments" -H "Authorization: Bearer $KEY" \
-H "Content-Type: application/json" -d '{
"model": "jev-latest",
"state": "Help! My payouts have been failing for 3 days.",
"questions": {
"department": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "Payments, refunds", "technical": "Bugs, outages",
"sales": null}},
"is_urgent": {"type": "noul", "instructions": "Does this convey urgency?"}
}
}'
# {"model":"jev-1.13.0",
# "answers":{"department":{"type":"choice","choice":"billing",
# "probabilities":{"billing":0.88,"technical":0.12,"sales":0.0},
# "confidence":0.81},
# "is_urgent":{"type":"noul","noul":0.95}},
# "usage":{"input_tokens":318,"output_tokens":34}}- Questions.
questionsmaps an id you choose to{type, instructions, criteria}.noulis yes/no (optionalcriteria: {"true": …, "false": …}) and answers{noul}in [0, 1].choicemaps 1–255 options to a description ornulland answers{choice, probabilities, confidence}.scoretakes an ordered array of 2–10 levels and answers{score, legend, probabilities, confidence}.state,instructionsand every description may be a string, an object or an array, and reach the model exactly as sent — key order and option order included. - Refused at the edge. An unknown field, a
typeoutsidechoice|score|noul, or criteria of the wrong shape or size is a 400request.invalidnaming the field (questions.<id>.criteria) — no upstream call, no charge. - Checked on the way back. Every answer is checked against its question
before you see it: all answered, each as the type asked, a
choiceone of the options offered, every probability in [0, 1]. An upstream that breaks that is a 503upstream.unavailablenaming the question — never an answer you would act on. - Billing. Input tokens only (the state plus every question), at the
model's input rate —
jev-latestis 4.2 ¢ per million, so a 500-token judgment costs 2,100 µ¢ (0.0021 ¢). The reply carriesX-Nozzle-Cost-Micro-Cents; the ledger row'skindisinference.judgment. - Rate limits. TypeSafe's own 429 and 529 surface as
upstream.rate_limited(withdetails.remaining_msandRetry-Afterwhen it sent one) andupstream.unavailable; wait and retry the same call.