Replicate Pricing
Replicate charges between $0.250 and $60.00 per 1 million output tokens, depending on the model. Ibm Granite Granite 3.3 8B Instruct is the cheapest at $0.250/1M output and $0.030/1M input; OpenAI o1 is the most expensive at $60.00/1M output. The largest context window is 164K tokens (DeepSeek Ai DeepSeek V3.1). Prices are USD, current as of July 20, 2026.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.030/1M
Ibm Granite Granite 3.3 8B Instruct
Cheapest output
$0.250/1M
Ibm Granite Granite 3.3 8B Instruct
Longest context
164K
DeepSeek Ai DeepSeek V3.1
Models priced
40
Avg $6.40/1M output
Pricing by model
All prices in USD per 1 million tokens. All 40 priced models, cheapest first.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| Ibm Granite Granite 3.3 8B Instruct | $0.030 | $0.250 | 0 |
| Meta Llama 2 7B | $0.050 | $0.250 | 4K |
| Meta Llama 2 7B Chat | $0.050 | $0.250 | 4K |
| Meta Llama 3 8B | $0.050 | $0.250 | 8K |
| Meta Llama 3 8B Instruct | $0.050 | $0.250 | 8K |
| Mistralai Mistral 7B Instruct V0.2 | $0.050 | $0.250 | 4K |
| Mistralai Mistral 7B V0.1 | $0.050 | $0.250 | 4K |
| GPT Oss 20B | $0.090 | $0.360 | 0 |
| OpenAI GPT 4.1 Nano | $0.100 | $0.400 | 0 |
| OpenAI GPT 5 Nano | $0.050 | $0.400 | 0 |
| Meta Llama 2 13B | $0.100 | $0.500 | 4K |
| Meta Llama 2 13B Chat | $0.100 | $0.500 | 4K |
| OpenAI GPT 4o Mini | $0.150 | $0.600 | 0 |
| OpenAI GPT Oss 120B | $0.180 | $0.720 | 0 |
| Mistralai Mixtral 8X7B Instruct V0.1 | $0.300 | $1.00 | 4K |
| Qwen Qwen3 235B A22b Instruct 2507 | $0.264 | $1.06 | 0 |
| DeepSeek Ai DeepSeek | $1.45 | $1.45 | 66K |
| OpenAI GPT 4.1 Mini | $0.400 | $1.60 | 0 |
| OpenAI GPT 5 Mini | $0.250 | $2.00 | 0 |
| DeepSeek Ai DeepSeek V3.1 | $0.672 | $2.02 | 164K |
| Google Gemini 2.5 Flash | $2.50 | $2.50 | 0 |
| Meta Llama 2 70B | $0.650 | $2.75 | 4K |
| Meta Llama 2 70B Chat | $0.650 | $2.75 | 4K |
| Meta Llama 3 70B | $0.650 | $2.75 | 8K |
| Meta Llama 3 70B Instruct | $0.650 | $2.75 | 8K |
| OpenAI o4 Mini | $1.00 | $4.00 | 0 |
| OpenAI o1 Mini | $1.10 | $4.40 | 0 |
| Anthropic Claude 3.5 Haiku | $1.00 | $5.00 | 0 |
| Anthropic Claude 4.5 Haiku | $1.00 | $5.00 | 0 |
| OpenAI GPT 4.1 | $2.00 | $8.00 | 0 |
| DeepSeek Ai DeepSeek R1 | $3.75 | $10.00 | 66K |
| OpenAI GPT 4o | $2.50 | $10.00 | 0 |
| OpenAI GPT 5 | $1.25 | $10.00 | 0 |
| Google Gemini 3 Pro | $2.00 | $12.00 | 0 |
| Anthropic Claude 3.7 Sonnet | $3.00 | $15.00 | 0 |
| Anthropic Claude 4 Sonnet | $3.00 | $15.00 | 0 |
| Anthropic Claude 4.5 Sonnet | $3.00 | $15.00 | 0 |
| Anthropic Claude 3.5 Sonnet | $3.75 | $18.75 | 0 |
| Xai Grok 4 | $7.20 | $36.00 | 0 |
| OpenAI o1 | $15.00 | $60.00 | 0 |
Frequently asked questions
How much does Replicate cost per 1M tokens?
Replicate pricing ranges from $0.250 to $60.00 per 1 million output tokens across 40 models. Input tokens are cheaper than output on every model. Rates current as of July 20, 2026.
What is the cheapest Replicate model?
Ibm Granite Granite 3.3 8B Instruct is the cheapest Replicate model on both axes — $0.030 per 1M input tokens and $0.250 per 1M output. It is the right default for high-volume, low-complexity work such as classification, extraction and routing.
What is the largest Replicate context window?
DeepSeek Ai DeepSeek V3.1 has the largest context window in the Replicate catalog at 164K input tokens, with up to 164K output tokens per response.
How do I calculate my actual Replicate bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator