AI API prices compared — official vs discounted

Every price below is listed twice: what the vendor charges directly, and what the same model costs through the discounted route. Verified 2026-09-16.

ModelOfficial (in / out)Gateway (in / out)Saving
Claude Opus 5
Anthropic
$5 / $25 $2.5 / $12.5 -50%
Claude Sonnet 5
Anthropic
$2 / $10 $1 / $5 -50%
Claude Fable 5
Anthropic
$10 / $50 $5 / $25 -50%
Claude Haiku 4.5
Anthropic
$1 / $5 $0.5 / $2.5 -50%
GPT-5.6 Sol
OpenAI
$4 / $20 $2.5 / $15 -27%
GPT-5.6 Terra
OpenAI
$2 / $12 $1 / $6 -50%
GPT-5.6 Luna
OpenAI
$0.2 / $1.2 $0.1 / $0.6 -50%
GPT-6 Astra
OpenAI
$10 / $50 $5 / $25 -50%
Gemini 3.7 Flash
Google
not verified $0.375 / $1.875 n/a
DeepSeek V4 Pro
DeepSeek
not verified $0.66 / $1.98 n/a
DeepSeek V4 Flash
DeepSeek
not verified $0.22 / $0.66 n/a
GLM-5.3
Zhipu
not verified $0.7 / $2.2 n/a
MiniMax M3
MiniMax
not verified $0.15 / $0.6 n/a
Kimi K3
Moonshot
not verified $1.5 / $7.5 n/a
Qwen3.8 Max
Alibaba
not verified $1 / $3 n/a
Grok 4.1
xAI
not verified $1 / $5 n/a

Per 1M tokens, USD, standard tier. Synced 2026-09-16. Savings are shown only where we could verify the official list price against vendor documentation — see each model page for the source. Where a cell reads not verified we make no comparison claim.

Check live rates on the platform →Free to sign up · $1 minimum top-up · No prepayment
Read this before you compare: the discount applies to US frontier models. DeepSeek, GLM, Kimi, Qwen and MiniMax are billed at their official list price — the point of listing them is access, not saving.

How to read the table

Pick a model

Claude Opus 5

Anthropic's strongest model for long-horizon agentic coding and enterprise work.

Claude Sonnet 5

The daily driver — near-flagship coding quality at a fifth of Opus's output rate.

Claude Fable 5

Frontier Claude for demanding reasoning and long agent runs.

Claude Haiku 4.5

Ultra-fast classification, extraction and tight loops.

GPT-5.6 Sol

OpenAI's heaviest 5.6 tier — deepest reasoning, widest tool use.

GPT-5.6 Terra

The middle tier most teams settle on.

GPT-5.6 Luna

Lowest-cost tier in the GPT-5.6 family — drafts, routing, high-frequency calls.

GPT-6 Astra

OpenAI's newest frontier tier.

Gemini 3.7 Flash

Google's fast multimodal model with a 1M context window.

DeepSeek V4 Pro

DeepSeek's flagship — strong reasoning at open-weight economics.

DeepSeek V4 Flash

Peak/off-peak pricing — schedule batch work off-peak to cut cost sharply.

GLM-5.3

Zhipu's current flagship.

MiniMax M3

MiniMax's latest — very low price for the capability.

Kimi K3

Moonshot's latest flagship — long-context agentic work.

Qwen3.8 Max

Alibaba's largest Qwen tier.

Grok 4.1

xAI's flagship model.

Get access at these prices