Llama API pricing: every model, every rate

Llama is open-weight, so the same model is hosted by many providers at wildly different rates. Below is what each variant bills at here.

Check current 0rates → Free to sign up · $1 minimum top-up · No prepayment

Full rate card

26 billable models. 0 carry a rate verified against the vendor's own published pricing; the rest are gateway rates only. USD per 1M tokens, verified 2026-09-16.

ModelInputOutputOut/InOfficial (in / out)
Llama-3.1-405B$3$62.0×not verified
Meta-Llama-3.1-405B-Instruct$3$62.0×not verified
llama-3.1-405b-instruct$3$62.0×not verified
llama-3.2-90b-vision-instruct$3$93.0×not verified
llama-3-70b$2$42.0×not verified
llama-3.1-70b$2$42.0×not verified
llama-3.1-70b-instruct$2$21.0×not verified
meta-llama/llama-3.1-70b-instruct$2$21.0×not verified
meta-llama/llama-4-maverick$1.25$54.0×not verified
meta-llama/llama-4-scout$1.25$54.0×not verified
llama-2-13b$1$11.0×not verified
llama-2-70b$1$11.0×not verified
llama-2-7b$1$11.0×not verified
llama-3-8b$1$11.0×not verified
llama-3.1-8b$1$11.0×not verified
llama-3.2-11b-vision-instruct$1$11.0×not verified
llama-3-sonar-large-32k-chat$0.75$0.751.0×not verified
llama-3-sonar-small-32k-chat$0.75$0.751.0×not verified
llama-3.2-3b-instruct$0.5$0.250.5×not verified
llama-3.3-70b-instruct$0.36$0.361.0×not verified
llama-3.1-70b-instruct-turbo$0.25$14.0×not verified
llama-3.2-1b-instruct$0.25$0.06250.2×not verified
llama-3.2-90b-vision$0.17$0.19381.1×not verified
llama-3.1-8b-instruct$0.125$0.54.0×not verified
llama-4-maverick$0.07$0.355.0×not verified
llama-3.3-70b$0.05$0.1483.0×not verified

Output is billed at a multiple of input on most models. The Out/In column is that multiple — it matters more than the input price once output dominates your bill.

Across 26 billable models the rate spans $0.05 (llama-3.3-70b) to $3 (Llama-3.1-405B) per million input tokens — a 60× spread. Picking the wrong tier is usually the single most expensive mistake here.

What Meta is good at

Llama 405B is the usual open-weight choice when you need on-prem-equivalent capability without a frontier price tag.

Before you compare: Because Llama is open-weight, rate differences between hosts are usually infrastructure margin rather than model difference. Worth shopping.

Create free account

See all 827 models → or use the cost calculator with your own token split.