Meta Llama Pricing
Meta Llama charges between $0.020 and $4.00 per 1 million output tokens, depending on the model. Llama 3.2 3B Instruct is the cheapest at $0.020/1M output and $0.020/1M input; Llama 3.1 405B (base) is the most expensive at $4.00/1M output. The largest context window is 1.0M tokens (Llama 4 Maverick). Prices are USD, current as of July 20, 2026.
Llama models are open-weight, so "price" depends entirely on who hosts them. The rates below are hosted-inference prices; self-hosting trades them for GPU-hour cost, which is usually only cheaper above sustained high utilisation.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.020/1M
Llama 3.1 8B Instruct
Cheapest output
$0.020/1M
Llama 3.2 3B Instruct
Longest context
1.0M
Llama 4 Maverick
Models priced
16
Avg $0.674/1M output
Pricing by model
All prices in USD per 1 million tokens. All 16 priced models, cheapest first.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| Llama 3.2 3B Instruct | $0.020 | $0.020 | 131K |
| Llama 3.1 8B Instruct | $0.020 | $0.030 | 131K |
| Llama 3.2 11B Vision Instruct | $0.049 | $0.049 | 131K |
| Llama 3 8B Instruct | $0.030 | $0.060 | 8K |
| Llama Guard 3 8B | $0.020 | $0.060 | 131K |
| Llama Guard 4 12B | $0.180 | $0.180 | 164K |
| Llama 3.2 1B Instruct | $0.027 | $0.200 | 60K |
| LlamaGuard 2 8B | $0.200 | $0.200 | 8K |
| Llama 4 Scout | $0.080 | $0.300 | 328K |
| Llama 3.3 70B Instruct | $0.130 | $0.380 | 131K |
| Llama 3 70B Instruct | $0.300 | $0.400 | 8K |
| Llama 3.1 70B Instruct | $0.400 | $0.400 | 131K |
| Llama 3.2 90B Vision Instruct | $0.350 | $0.400 | 33K |
| Llama 4 Maverick | $0.150 | $0.600 | 1.0M |
| Llama 3.1 405B Instruct | $3.50 | $3.50 | 131K |
| Llama 3.1 405B (base) | $4.00 | $4.00 | 33K |
Frequently asked questions
How much does Meta Llama cost per 1M tokens?
Meta Llama pricing ranges from $0.020 to $4.00 per 1 million output tokens across 16 models. Input tokens are cheaper than output on every model. Rates current as of July 20, 2026.
What is the cheapest Meta Llama model?
Llama 3.2 3B Instruct is the cheapest Meta Llama model on output tokens at $0.020 per 1M, while Llama 3.1 8B Instruct is cheapest on input at $0.020 per 1M. Which wins for you depends on your input-to-output ratio.
What is the largest Meta Llama context window?
Llama 4 Maverick has the largest context window in the Meta Llama catalog at 1.0M input tokens, with up to 16K output tokens per response.
How do I calculate my actual Meta Llama bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Head-to-head comparisons
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator