> ## Documentation Index
> Fetch the complete documentation index at: https://docs.opper.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Service tiers

> Trade speed for price per request: flex for about half price, priority for faster tokens. One field, the same as on OpenAI.

Several providers sell the same model at more than one processing tier. **Flex** runs your request on spare capacity for about half the price, in exchange for slower and less predictable responses. **Priority** puts your request ahead of standard traffic for faster tokens, at a premium. Opper exposes these tiers with the same `service_tier` field OpenAI uses, on every provider that offers them.

| Tier | What you get | Typical price |
| - | - | - |
| `flex` | Spare capacity. Responses can queue for minutes, and requests can be refused when capacity runs out. | About 50% of standard |
| `priority` (alias `fast`) | Scheduled ahead of standard traffic: lower latency, faster output. | About 2x standard |
| `ultrafast` | OpenAI's fastest tier, on selected models. | About 6x standard |
| `default` | The standard tier. Same as leaving the field out. | Standard |

## Request a tier

Add `service_tier` to the request body:

<CodeGroup>
  ```python OpenAI SDK theme={null}
  import os
  from openai import OpenAI

  client = OpenAI(
      base_url="https://api.opper.ai/v3/compat",
      api_key=os.environ["OPPER_API_KEY"],
  )

  resp = client.chat.completions.create(
      model="gpt-5.4-mini",
      service_tier="flex",
      messages=[{"role": "user", "content": "Summarize this report."}],
      timeout=600,  # flex requests can queue, allow a long timeout
  )
  print(resp.service_tier)  # "flex"
  ```

  ```typescript OpenAI SDK theme={null}
  import OpenAI from "openai";

  const client = new OpenAI({
    baseURL: "https://api.opper.ai/v3/compat",
    apiKey: process.env.OPPER_API_KEY,
    timeout: 600_000, // flex requests can queue, allow a long timeout
  });

  const resp = await client.chat.completions.create({
    model: "gpt-5.4-mini",
    service_tier: "flex",
    messages: [{ role: "user", content: "Summarize this report." }],
  });
  console.log(resp.service_tier); // "flex"
  ```

  ```bash cURL theme={null}
  curl https://api.opper.ai/v3/compat/chat/completions \
    -H "Authorization: Bearer $OPPER_API_KEY" \
    -H "Content-Type: application/json" \
    --max-time 600 \
    -d '{
      "model": "gpt-5.4-mini",
      "service_tier": "flex",
      "messages": [{ "role": "user", "content": "Summarize this report." }]
    }'
  ```
</CodeGroup>

The field works on `chat/completions`, `responses`, `openresponses` and `v1/messages`. On `v1/messages`, Anthropic's `speed: "fast"` is accepted too and means `priority`.

`auto`, `default`, `standard`, `standard_only` and `scale` all mean the standard tier, so clients that send them (Codex does, for example) keep working. Any other value is a `400`.

### Or name the tier endpoint

Each tier is its own endpoint in the catalog, so you can also call it by id, with no `service_tier` field:

```json theme={null}
{"model": "openai:flex/gpt-5.4-mini", "messages": [...]}
```

Tier endpoint ids follow `<provider>[:<location>]:<tier>/<model>`, for example `openai:priority/gpt-5.4`, `gemini:flex/gemini-3.8-flash` or `vertexai:eu:priority/gemini-3.6-flash`. Tier endpoints are never picked for a request that doesn't ask for a tier: calling `gpt-5.4-mini` without `service_tier` always runs at the standard tier.

## What happens when the tier isn't available

| You ask for | Opper tries | If none answers |
| - | - | - |
| `flex` | the flex endpoint | the request fails (for example `429` or `503`). It never silently runs at the full price. |
| `priority` | the priority endpoint, then the standard one | standard is the fallback, billed at the standard price |
| `ultrafast` | ultrafast, then priority, then standard | standard is the fallback |

If the model has no endpoint at the tier you asked for, the request runs at the standard tier and the response says `default`. Asking for `flex` never fails just because a model lacks it.

### Fallbacks

Each tier is its own model id (`openai:flex/gpt-5.4-mini`), so it goes into your usual fallbacks like any other model:

* **The [`models` array](/capabilities/models#fallbacks-and-backup-chains)** (Chat Completions): flex first, the same model at standard if flex has no capacity.

  ```json theme={null}
  {
    "model": "openai:flex/gpt-5.4-mini",
    "models": ["openai/gpt-5.4-mini"],
    "messages": [...]
  }
  ```

* **A [dynamic route](/capabilities/routes/overview)**, on every endpoint: put both endpoints in a [Pool](/capabilities/routes/pool) ordered **Cheapest first**. The flex endpoint ranks first on price and the standard one takes over when flex is out of capacity.

Name the tier endpoints rather than setting `service_tier` in a chain: with `service_tier` set, every entry moves to its own tier endpoint when it has one (as on OpenAI and OpenRouter), so the fallback above would run on flex too.

## See which tier served the request

You are billed at the tier that actually ran. Providers sometimes serve a priority request at the standard tier when they're out of priority capacity; Opper reads that back from the provider and bills the standard price.

* Header: `X-Opper-Served-Service-Tier: flex | priority | ultrafast | default`, on every response.
* Body: `service_tier` on `chat/completions`, `responses` and `openresponses` when you asked for a tier or named a tier endpoint. Chat streams carry it on every chunk; `meta.routing.served.service_tier` names it too.
* On a stream, the header is written before the provider confirms the tier, so treat the final chunk (or `response.completed`) as the authoritative answer.

## Timeouts for flex

Flex requests can wait in the provider's queue for minutes before the first token (Google quotes 1 to 15 minutes). Opper waits up to 10 minutes for the first token of a flex request instead of its usual first-token limit, so set your client timeout to at least `600` seconds. Use flex for background work: evaluations, enrichment, batch-like jobs that still need a synchronous answer.

## Where tiers are available

| Provider | Flex | Priority | Region |
| - | - | - | - |
| OpenAI | ✓ (GPT-5 family, GPT-6, o3, o4-mini) | ✓, and `ultrafast` on GPT-6 Astra | US |
| Google Gemini API | ✓ | ✓ | US |
| Google Vertex AI | ✓ (global endpoint) | ✓ | Global, and **EU** for priority |
| xAI | | ✓ (every Grok text model) | US, and **EU** for the models served in the EU |
| Azure OpenAI | | ✓ (GPT-6 Sol, GPT-5.6 Sol and Terra) | Global |
| BytePlus | ✓ (DeepSeek V4 Flash) | | Asia-Pacific |

To list the tier endpoints, look for the `service_tier` field in [`GET /v3/models`](/v3-api-reference/models/list-models), or open the Models page in the [Opper platform](https://platform.opper.ai), where tier endpoints carry a tier mark and their own price. The public catalog on opper.ai lists standard endpoints only.

<Note>
  Tier endpoints follow your [compliance rules](/control-plane/rules/overview) like any other endpoint. If your project only allows EU processing, OpenAI and Gemini API tiers are blocked; the Vertex AI and xAI EU priority endpoints are available.
</Note>
