Perplexity Pricing
Perplexity charges from $0.200 per 1 million output tokens, depending on the model. Llama 3.1 8B Instruct is the cheapest at $0.200/1M output and $0.200/1M input; its current flagship Sonar Pro costs $15.00/1M output. The largest context window is 200K tokens (Sonar Pro). Prices are USD, current as of August 11, 2026.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.070/1M
Mistral 7B Instruct
Cheapest output
$0.200/1M
Llama 3.1 8B Instruct
Longest context
200K
Sonar Pro
Models priced
22
Avg $3.59/1M output
Pricing by model
All prices in USD per 1 million tokens. All 22 priced models, cheapest first.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| Llama 3.1 8B Instruct | $0.200 | $0.200 | 131K |
| Mistral 7B Instruct | $0.070 | $0.280 | 4K |
| Mixtral 8X7B Instruct | $0.070 | $0.280 | 4K |
| Pplx 7B Chat | $0.070 | $0.280 | 8K |
| Pplx 7B Online | — | $0.280 | 4K |
| Sonar Small Chat | $0.070 | $0.280 | 16K |
| Sonar Small Online | — | $0.280 | 12K |
| Llama 3.1 70B Instruct | $1.00 | $1.00 | 131K |
| Sonar | $1.00 | $1.00 | 127K |
| Codellama 34B Instruct | $0.350 | $1.40 | 16K |
| Sonar Medium Chat | $0.600 | $1.80 | 16K |
| Sonar Medium Online | — | $1.80 | 12K |
| Codellama 70B Instruct | $0.700 | $2.80 | 16K |
| Llama 2 70B Chat | $0.700 | $2.80 | 4K |
| Pplx 70B Chat | $0.700 | $2.80 | 4K |
| Pplx 70B Online | — | $2.80 | 4K |
| Sonar Reasoning | $1.00 | $5.00 | 127K |
| Perplexity Sonar Deep Research Reasoning | $2.00 | $8.00 | 128K |
| Sonar Deep Research | $2.00 | $8.00 | 128K |
| Sonar Reasoning Pro | $2.00 | $8.00 | 128K |
| Sonar Pro | $3.00 | $15.00 | 200K |
| Sonar Pro Search | $3.00 | $15.00 | 200K |
Frequently asked questions
How much does Perplexity cost per 1M tokens?
Perplexity pricing starts at $0.200 per 1 million output tokens (Llama 3.1 8B Instruct) across 22 models, with its current flagship Sonar Pro at $15.00. Older premium models in the catalog list higher. Rates current as of August 11, 2026.
What is the cheapest Perplexity model?
Llama 3.1 8B Instruct is the cheapest Perplexity model on output tokens at $0.200 per 1M, while Mistral 7B Instruct is cheapest on input at $0.070 per 1M. Which wins for you depends on your input-to-output ratio.
What is the largest Perplexity context window?
Sonar Pro has the largest context window in the Perplexity catalog at 200K input tokens, with up to 16K output tokens per response.
Which Perplexity models support reasoning or vision?
The Perplexity catalog includes 1 reasoning models. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.
How do I calculate my actual Perplexity bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator