Skip to main content
Several providers sell the same model at more than one processing tier. Flex runs your request on spare capacity for about half the price, in exchange for slower and less predictable responses. Priority puts your request ahead of standard traffic for faster tokens, at a premium. Opper exposes these tiers with the same service_tier field OpenAI uses, on every provider that offers them.

Request a tier

Add service_tier to the request body:
The field works on chat/completions, responses, openresponses and v1/messages. On v1/messages, Anthropic’s speed: "fast" is accepted too and means priority. auto, default, standard, standard_only and scale all mean the standard tier, so clients that send them (Codex does, for example) keep working. Any other value is a 400.

Or name the tier endpoint

Each tier is its own endpoint in the catalog, so you can also call it by id, with no service_tier field:
Tier endpoint ids follow <provider>[:<location>]:<tier>/<model>, for example openai:priority/gpt-5.4, gemini:flex/gemini-3.8-flash or vertexai:eu:priority/gemini-3.6-flash. Tier endpoints are never picked for a request that doesn’t ask for a tier: calling gpt-5.4-mini without service_tier always runs at the standard tier.

What happens when the tier isn’t available

If the model has no endpoint at the tier you asked for, the request runs at the standard tier and the response says default. Asking for flex never fails just because a model lacks it.

Fallbacks

Each tier is its own model id (openai:flex/gpt-5.4-mini), so it goes into your usual fallbacks like any other model:
  • The models array (Chat Completions): flex first, the same model at standard if flex has no capacity.
  • A dynamic route, on every endpoint: put both endpoints in a Pool ordered Cheapest first. The flex endpoint ranks first on price and the standard one takes over when flex is out of capacity.
Name the tier endpoints rather than setting service_tier in a chain: with service_tier set, every entry moves to its own tier endpoint when it has one (as on OpenAI and OpenRouter), so the fallback above would run on flex too.

See which tier served the request

You are billed at the tier that actually ran. Providers sometimes serve a priority request at the standard tier when they’re out of priority capacity; Opper reads that back from the provider and bills the standard price.
  • Header: X-Opper-Served-Service-Tier: flex | priority | ultrafast | default, on every response.
  • Body: service_tier on chat/completions, responses and openresponses when you asked for a tier or named a tier endpoint. Chat streams carry it on every chunk; meta.routing.served.service_tier names it too.
  • On a stream, the header is written before the provider confirms the tier, so treat the final chunk (or response.completed) as the authoritative answer.

Timeouts for flex

Flex requests can wait in the provider’s queue for minutes before the first token (Google quotes 1 to 15 minutes). Opper waits up to 10 minutes for the first token of a flex request instead of its usual first-token limit, so set your client timeout to at least 600 seconds. Use flex for background work: evaluations, enrichment, batch-like jobs that still need a synchronous answer.

Where tiers are available

To list the tier endpoints, look for the service_tier field in GET /v3/models, or open the Models page in the Opper platform, where tier endpoints carry a tier mark and their own price. The public catalog on opper.ai lists standard endpoints only.
Tier endpoints follow your compliance rules like any other endpoint. If your project only allows EU processing, OpenAI and Gemini API tiers are blocked; the Vertex AI and xAI EU priority endpoints are available.
Last modified on October 1, 2026