AI API pricing by vendor
Every model a vendor ships, with the rate it actually bills at. Use this when you are choosing inside one ecosystem; use the full price index when you are comparing across them.
All vendors
Anthropic API pricing
38 billable models · USD per 1M tokens
Claude is the default choice for agentic coding, long-context document work, and anything where instruction-following quality matters more than raw token cost.
OpenAI API pricing
150 billable models · USD per 1M tokens
GPT-5.6 Luna is the volume workhorse — cheap enough for classification and routing, strong enough for most user-facing chat.
Google API pricing
80 billable models · USD per 1M tokens
Gemini wins on long-context and multimodal work. Flash tiers are among the cheapest capable models available anywhere.
DeepSeek API pricing
40 billable models · USD per 1M tokens
DeepSeek is the default choice when output volume dominates — R1-class reasoning at a fraction of US frontier pricing.
Alibaba API pricing
88 billable models · USD per 1M tokens
Qwen coder models are strong enough for agentic work and priced well below comparable US models.
xAI API pricing
32 billable models · USD per 1M tokens
Grok is competitive on real-time and search-adjacent workloads.
Zhipu API pricing
32 billable models · USD per 1M tokens
GLM-5.x is a strong general-purpose Chinese-language model with competitive agentic performance.
Moonshot API pricing
17 billable models · USD per 1M tokens
Kimi is the usual pick when you need a long context window without frontier-model pricing.
Meta API pricing
26 billable models · USD per 1M tokens
Llama 405B is the usual open-weight choice when you need on-prem-equivalent capability without a frontier price tag.
MiniMax API pricing
6 billable models · USD per 1M tokens
MiniMax is competitive for high-volume chat and character-style applications.
ByteDance API pricing
46 billable models · USD per 1M tokens
Doubao is widely used for high-volume Chinese-language workloads.
Why the vendor pages exist
A vendor's own pricing page is usually a marketing summary — it shows the flagship model and leaves out the legacy tiers still running in production, the regional variants, and the output multiples that dominate real bills. These pages list the full billable catalogue instead, because that is the only way to tell whether the tier you picked is the tier you meant to pick.
Rates are synced from the gateway's pricing endpoint and verified against vendor documentation where a stable published source exists. Where we could not verify, we say so on the row rather than inventing a comparison.