API pricing

Perplexity Pricing

Perplexity charges between $0.200 and $15.00 per 1 million output tokens, depending on the model. Llama 3.1 8B Instruct is the cheapest at $0.200/1M output and $0.200/1M input; Sonar Pro is the most expensive at $15.00/1M output. The largest context window is 200K tokens (Sonar Pro). Prices are USD, current as of July 20, 2026.

Last updated · synced weekly from the upstream model catalog

Cheapest input

—/1M

Pplx 70B Online

Cheapest output

$0.200/1M

Llama 3.1 8B Instruct

Longest context

200K

Sonar Pro

Models priced

22

Avg $3.59/1M output

Pricing by model

All prices in USD per 1 million tokens. All 22 priced models, cheapest first.

ModelInput / 1MOutput / 1MContext
Llama 3.1 8B Instruct$0.200$0.200131K
Mistral 7B Instruct$0.070$0.2804K
Mixtral 8X7B Instruct$0.070$0.2804K
Pplx 7B Chat$0.070$0.2808K
Pplx 7B Online$0.2804K
Sonar Small Chat$0.070$0.28016K
Sonar Small Online$0.28012K
Llama 3.1 70B Instruct$1.00$1.00131K
Sonar$1.00$1.00127K
Codellama 34B Instruct$0.350$1.4016K
Sonar Medium Chat$0.600$1.8016K
Sonar Medium Online$1.8012K
Codellama 70B Instruct$0.700$2.8016K
Llama 2 70B Chat$0.700$2.804K
Pplx 70B Chat$0.700$2.804K
Pplx 70B Online$2.804K
Sonar Reasoning$1.00$5.00127K
Perplexity Sonar Deep Research
Reasoning
$2.00$8.00128K
Sonar Deep Research$2.00$8.00128K
Sonar Reasoning Pro$2.00$8.00128K
Sonar Pro$3.00$15.00200K
Sonar Pro Search$3.00$15.00200K

Frequently asked questions

How much does Perplexity cost per 1M tokens?

Perplexity pricing ranges from $0.200 to $15.00 per 1 million output tokens across 22 models. Input tokens are cheaper than output on every model. Rates current as of July 20, 2026.

What is the cheapest Perplexity model?

Llama 3.1 8B Instruct is the cheapest Perplexity model on output tokens at $0.200 per 1M, while Pplx 70B Online is cheapest on input at — per 1M. Which wins for you depends on your input-to-output ratio.

What is the largest Perplexity context window?

Sonar Pro has the largest context window in the Perplexity catalog at 200K input tokens, with up to 16K output tokens per response.

Which Perplexity models support reasoning or vision?

The Perplexity catalog includes 1 reasoning models. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.

How do I calculate my actual Perplexity bill?

Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.

Other providers

What will this actually cost you?

Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.

Open the free cost calculator