AI API prices compared — official vs discounted
Every price below is listed twice: what the vendor charges directly, and what the same model costs through the discounted route. Verified 2026-09-16.
| Model | Official (in / out) | Gateway (in / out) | Saving |
|---|---|---|---|
| Claude Opus 5 Anthropic |
$5 / $25 | $2.5 / $12.5 | -50% |
| Claude Sonnet 5 Anthropic |
$2 / $10 | $1 / $5 | -50% |
| Claude Fable 5 Anthropic |
$10 / $50 | $5 / $25 | -50% |
| Claude Haiku 4.5 Anthropic |
$1 / $5 | $0.5 / $2.5 | -50% |
| GPT-5.6 Sol OpenAI |
$4 / $20 | $2.5 / $15 | -27% |
| GPT-5.6 Terra OpenAI |
$2 / $12 | $1 / $6 | -50% |
| GPT-5.6 Luna OpenAI |
$0.2 / $1.2 | $0.1 / $0.6 | -50% |
| GPT-6 Astra OpenAI |
$10 / $50 | $5 / $25 | -50% |
| Gemini 3.7 Flash |
not verified | $0.375 / $1.875 | n/a |
| DeepSeek V4 Pro DeepSeek |
not verified | $0.66 / $1.98 | n/a |
| DeepSeek V4 Flash DeepSeek |
not verified | $0.22 / $0.66 | n/a |
| GLM-5.3 Zhipu |
not verified | $0.7 / $2.2 | n/a |
| MiniMax M3 MiniMax |
not verified | $0.15 / $0.6 | n/a |
| Kimi K3 Moonshot |
not verified | $1.5 / $7.5 | n/a |
| Qwen3.8 Max Alibaba |
not verified | $1 / $3 | n/a |
| Grok 4.1 xAI |
not verified | $1 / $5 | n/a |
Per 1M tokens, USD, standard tier. Synced 2026-09-16. Savings are shown only where we could verify the official list price against vendor documentation — see each model page for the source. Where a cell reads not verified we make no comparison claim.
How to read the table
- Input is what you send (prompt, context, system message).
- Output is what the model generates, and typically costs 3–6× more per token.
- Saving is the blended difference across input and output at a 1:1 ratio. Your real ratio changes it — use the calculator.
Pick a model
Claude Opus 5
Anthropic's strongest model for long-horizon agentic coding and enterprise work.
Claude Sonnet 5
The daily driver — near-flagship coding quality at a fifth of Opus's output rate.
Claude Fable 5
Frontier Claude for demanding reasoning and long agent runs.
Claude Haiku 4.5
Ultra-fast classification, extraction and tight loops.
GPT-5.6 Sol
OpenAI's heaviest 5.6 tier — deepest reasoning, widest tool use.
GPT-5.6 Terra
The middle tier most teams settle on.
GPT-5.6 Luna
Lowest-cost tier in the GPT-5.6 family — drafts, routing, high-frequency calls.
GPT-6 Astra
OpenAI's newest frontier tier.
Gemini 3.7 Flash
Google's fast multimodal model with a 1M context window.
DeepSeek V4 Pro
DeepSeek's flagship — strong reasoning at open-weight economics.
DeepSeek V4 Flash
Peak/off-peak pricing — schedule batch work off-peak to cut cost sharply.
GLM-5.3
Zhipu's current flagship.
MiniMax M3
MiniMax's latest — very low price for the capability.
Kimi K3
Moonshot's latest flagship — long-context agentic work.
Qwen3.8 Max
Alibaba's largest Qwen tier.
Grok 4.1
xAI's flagship model.