Llama API pricing: every model, every rate
Llama is open-weight, so the same model is hosted by many providers at wildly different rates. Below is what each variant bills at here.
Full rate card
26 billable models. 0 carry a rate verified against the vendor's own published pricing; the rest are gateway rates only. USD per 1M tokens, verified 2026-09-16.
| Model | Input | Output | Out/In | Official (in / out) |
|---|---|---|---|---|
Llama-3.1-405B | $3 | $6 | 2.0× | not verified |
Meta-Llama-3.1-405B-Instruct | $3 | $6 | 2.0× | not verified |
llama-3.1-405b-instruct | $3 | $6 | 2.0× | not verified |
llama-3.2-90b-vision-instruct | $3 | $9 | 3.0× | not verified |
llama-3-70b | $2 | $4 | 2.0× | not verified |
llama-3.1-70b | $2 | $4 | 2.0× | not verified |
llama-3.1-70b-instruct | $2 | $2 | 1.0× | not verified |
meta-llama/llama-3.1-70b-instruct | $2 | $2 | 1.0× | not verified |
meta-llama/llama-4-maverick | $1.25 | $5 | 4.0× | not verified |
meta-llama/llama-4-scout | $1.25 | $5 | 4.0× | not verified |
llama-2-13b | $1 | $1 | 1.0× | not verified |
llama-2-70b | $1 | $1 | 1.0× | not verified |
llama-2-7b | $1 | $1 | 1.0× | not verified |
llama-3-8b | $1 | $1 | 1.0× | not verified |
llama-3.1-8b | $1 | $1 | 1.0× | not verified |
llama-3.2-11b-vision-instruct | $1 | $1 | 1.0× | not verified |
llama-3-sonar-large-32k-chat | $0.75 | $0.75 | 1.0× | not verified |
llama-3-sonar-small-32k-chat | $0.75 | $0.75 | 1.0× | not verified |
llama-3.2-3b-instruct | $0.5 | $0.25 | 0.5× | not verified |
llama-3.3-70b-instruct | $0.36 | $0.36 | 1.0× | not verified |
llama-3.1-70b-instruct-turbo | $0.25 | $1 | 4.0× | not verified |
llama-3.2-1b-instruct | $0.25 | $0.0625 | 0.2× | not verified |
llama-3.2-90b-vision | $0.17 | $0.1938 | 1.1× | not verified |
llama-3.1-8b-instruct | $0.125 | $0.5 | 4.0× | not verified |
llama-4-maverick | $0.07 | $0.35 | 5.0× | not verified |
llama-3.3-70b | $0.05 | $0.148 | 3.0× | not verified |
Output is billed at a multiple of input on most models. The Out/In column is that multiple — it matters more than the input price once output dominates your bill.
Across 26 billable models the rate spans $0.05 (llama-3.3-70b) to $3 (Llama-3.1-405B) per million input tokens — a 60× spread. Picking the wrong tier is usually the single most expensive mistake here.
What Meta is good at
Llama 405B is the usual open-weight choice when you need on-prem-equivalent capability without a frontier price tag.
See all 827 models → or use the cost calculator with your own token split.