API pricing

Replicate Pricing

Replicate charges between $0.250 and $60.00 per 1 million output tokens, depending on the model. Ibm Granite Granite 3.3 8B Instruct is the cheapest at $0.250/1M output and $0.030/1M input; OpenAI o1 is the most expensive at $60.00/1M output. The largest context window is 164K tokens (DeepSeek Ai DeepSeek V3.1). Prices are USD, current as of July 20, 2026.

Last updated · synced weekly from the upstream model catalog

Cheapest input

$0.030/1M

Ibm Granite Granite 3.3 8B Instruct

Cheapest output

$0.250/1M

Ibm Granite Granite 3.3 8B Instruct

Longest context

164K

DeepSeek Ai DeepSeek V3.1

Models priced

40

Avg $6.40/1M output

Pricing by model

All prices in USD per 1 million tokens. All 40 priced models, cheapest first.

ModelInput / 1MOutput / 1MContext
Ibm Granite Granite 3.3 8B Instruct$0.030$0.2500
Meta Llama 2 7B$0.050$0.2504K
Meta Llama 2 7B Chat$0.050$0.2504K
Meta Llama 3 8B$0.050$0.2508K
Meta Llama 3 8B Instruct$0.050$0.2508K
Mistralai Mistral 7B Instruct V0.2$0.050$0.2504K
Mistralai Mistral 7B V0.1$0.050$0.2504K
GPT Oss 20B$0.090$0.3600
OpenAI GPT 4.1 Nano$0.100$0.4000
OpenAI GPT 5 Nano$0.050$0.4000
Meta Llama 2 13B$0.100$0.5004K
Meta Llama 2 13B Chat$0.100$0.5004K
OpenAI GPT 4o Mini$0.150$0.6000
OpenAI GPT Oss 120B$0.180$0.7200
Mistralai Mixtral 8X7B Instruct V0.1$0.300$1.004K
Qwen Qwen3 235B A22b Instruct 2507$0.264$1.060
DeepSeek Ai DeepSeek$1.45$1.4566K
OpenAI GPT 4.1 Mini$0.400$1.600
OpenAI GPT 5 Mini$0.250$2.000
DeepSeek Ai DeepSeek V3.1$0.672$2.02164K
Google Gemini 2.5 Flash$2.50$2.500
Meta Llama 2 70B$0.650$2.754K
Meta Llama 2 70B Chat$0.650$2.754K
Meta Llama 3 70B$0.650$2.758K
Meta Llama 3 70B Instruct$0.650$2.758K
OpenAI o4 Mini$1.00$4.000
OpenAI o1 Mini$1.10$4.400
Anthropic Claude 3.5 Haiku$1.00$5.000
Anthropic Claude 4.5 Haiku$1.00$5.000
OpenAI GPT 4.1$2.00$8.000
DeepSeek Ai DeepSeek R1$3.75$10.0066K
OpenAI GPT 4o$2.50$10.000
OpenAI GPT 5$1.25$10.000
Google Gemini 3 Pro$2.00$12.000
Anthropic Claude 3.7 Sonnet$3.00$15.000
Anthropic Claude 4 Sonnet$3.00$15.000
Anthropic Claude 4.5 Sonnet$3.00$15.000
Anthropic Claude 3.5 Sonnet$3.75$18.750
Xai Grok 4$7.20$36.000
OpenAI o1$15.00$60.000

Frequently asked questions

How much does Replicate cost per 1M tokens?

Replicate pricing ranges from $0.250 to $60.00 per 1 million output tokens across 40 models. Input tokens are cheaper than output on every model. Rates current as of July 20, 2026.

What is the cheapest Replicate model?

Ibm Granite Granite 3.3 8B Instruct is the cheapest Replicate model on both axes — $0.030 per 1M input tokens and $0.250 per 1M output. It is the right default for high-volume, low-complexity work such as classification, extraction and routing.

What is the largest Replicate context window?

DeepSeek Ai DeepSeek V3.1 has the largest context window in the Replicate catalog at 164K input tokens, with up to 164K output tokens per response.

How do I calculate my actual Replicate bill?

Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.

Other providers

What will this actually cost you?

Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.

Open the free cost calculator