service_tier field OpenAI uses, on every provider that offers them.
Request a tier
Addservice_tier to the request body:
chat/completions, responses, openresponses and v1/messages. On v1/messages, Anthropic’s speed: "fast" is accepted too and means priority.
auto, default, standard, standard_only and scale all mean the standard tier, so clients that send them (Codex does, for example) keep working. Any other value is a 400.
Or name the tier endpoint
Each tier is its own endpoint in the catalog, so you can also call it by id, with noservice_tier field:
<provider>[:<location>]:<tier>/<model>, for example openai:priority/gpt-5.4, gemini:flex/gemini-3.8-flash or vertexai:eu:priority/gemini-3.6-flash. Tier endpoints are never picked for a request that doesn’t ask for a tier: calling gpt-5.4-mini without service_tier always runs at the standard tier.
What happens when the tier isn’t available
If the model has no endpoint at the tier you asked for, the request runs at the standard tier and the response says
default. Asking for flex never fails just because a model lacks it.
Fallbacks
Each tier is its own model id (openai:flex/gpt-5.4-mini), so it goes into your usual fallbacks like any other model:
-
The
modelsarray (Chat Completions): flex first, the same model at standard if flex has no capacity. - A dynamic route, on every endpoint: put both endpoints in a Pool ordered Cheapest first. The flex endpoint ranks first on price and the standard one takes over when flex is out of capacity.
service_tier in a chain: with service_tier set, every entry moves to its own tier endpoint when it has one (as on OpenAI and OpenRouter), so the fallback above would run on flex too.
See which tier served the request
You are billed at the tier that actually ran. Providers sometimes serve a priority request at the standard tier when they’re out of priority capacity; Opper reads that back from the provider and bills the standard price.- Header:
X-Opper-Served-Service-Tier: flex | priority | ultrafast | default, on every response. - Body:
service_tieronchat/completions,responsesandopenresponseswhen you asked for a tier or named a tier endpoint. Chat streams carry it on every chunk;meta.routing.served.service_tiernames it too. - On a stream, the header is written before the provider confirms the tier, so treat the final chunk (or
response.completed) as the authoritative answer.
Timeouts for flex
Flex requests can wait in the provider’s queue for minutes before the first token (Google quotes 1 to 15 minutes). Opper waits up to 10 minutes for the first token of a flex request instead of its usual first-token limit, so set your client timeout to at least600 seconds. Use flex for background work: evaluations, enrichment, batch-like jobs that still need a synchronous answer.
Where tiers are available
To list the tier endpoints, look for the
service_tier field in GET /v3/models, or open the Models page in the Opper platform, where tier endpoints carry a tier mark and their own price. The public catalog on opper.ai lists standard endpoints only.
Tier endpoints follow your compliance rules like any other endpoint. If your project only allows EU processing, OpenAI and Gemini API tiers are blocked; the Vertex AI and xAI EU priority endpoints are available.