Replicate Pricing
Replicate charges from $0.250 per 1 million output tokens, depending on the model. Ibm Granite Granite 3.3 8B Instruct is the cheapest at $0.250/1M output and $0.030/1M input; its current flagship OpenAI o1 costs $60.00/1M output. The largest context window is 164K tokens (DeepSeek Ai DeepSeek V3.1). Prices are USD, current as of August 11, 2026.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.030/1M
Ibm Granite Granite 3.3 8B Instruct
Cheapest output
$0.250/1M
Ibm Granite Granite 3.3 8B Instruct
Longest context
164K
DeepSeek Ai DeepSeek V3.1
Models priced
40
Avg $6.40/1M output
Pricing by model
All prices in USD per 1 million tokens. All 40 priced models, cheapest first.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| Ibm Granite Granite 3.3 8B Instruct | $0.030 | $0.250 | 0 |
| Meta Llama 2 7B | $0.050 | $0.250 | 4K |
| Meta Llama 2 7B Chat | $0.050 | $0.250 | 4K |
| Meta Llama 3 8B | $0.050 | $0.250 | 8K |
| Meta Llama 3 8B Instruct | $0.050 | $0.250 | 8K |
| Mistralai Mistral 7B Instruct V0.2 | $0.050 | $0.250 | 4K |
| Mistralai Mistral 7B V0.1 | $0.050 | $0.250 | 4K |
| OpenAI GPT Oss 20B | $0.090 | $0.360 | 0 |
| OpenAI GPT 4.1 Nano | $0.100 | $0.400 | 0 |
| OpenAI GPT 5 Nano | $0.050 | $0.400 | 0 |
| Meta Llama 2 13B | $0.100 | $0.500 | 4K |
| Meta Llama 2 13B Chat | $0.100 | $0.500 | 4K |
| OpenAI GPT 4o Mini | $0.150 | $0.600 | 0 |
| OpenAI GPT Oss 120B | $0.180 | $0.720 | 0 |
| Mistralai Mixtral 8X7B Instruct V0.1 | $0.300 | $1.00 | 4K |
| Qwen Qwen3 235B A22b Instruct 2507 | $0.264 | $1.06 | 0 |
| DeepSeek Ai DeepSeek | $1.45 | $1.45 | 66K |
| OpenAI GPT 4.1 Mini | $0.400 | $1.60 | 0 |
| OpenAI GPT 5 Mini | $0.250 | $2.00 | 0 |
| DeepSeek Ai DeepSeek V3.1 | $0.672 | $2.02 | 164K |
| Google Gemini 2.5 Flash | $2.50 | $2.50 | 0 |
| Meta Llama 2 70B | $0.650 | $2.75 | 4K |
| Meta Llama 2 70B Chat | $0.650 | $2.75 | 4K |
| Meta Llama 3 70B | $0.650 | $2.75 | 8K |
| Meta Llama 3 70B Instruct | $0.650 | $2.75 | 8K |
| OpenAI o4 Mini | $1.00 | $4.00 | 0 |
| OpenAI o1 Mini | $1.10 | $4.40 | 0 |
| Anthropic Claude 3.5 Haiku | $1.00 | $5.00 | 0 |
| Anthropic Claude 4.5 Haiku | $1.00 | $5.00 | 0 |
| OpenAI GPT 4.1 | $2.00 | $8.00 | 0 |
| DeepSeek Ai DeepSeek R1 | $3.75 | $10.00 | 66K |
| OpenAI GPT 4o | $2.50 | $10.00 | 0 |
| OpenAI GPT 5 | $1.25 | $10.00 | 0 |
| Google Gemini 3 Pro | $2.00 | $12.00 | 0 |
| Anthropic Claude 3.7 Sonnet | $3.00 | $15.00 | 0 |
| Anthropic Claude 4 Sonnet | $3.00 | $15.00 | 0 |
| Anthropic Claude 4.5 Sonnet | $3.00 | $15.00 | 0 |
| Anthropic Claude 3.5 Sonnet | $3.75 | $18.75 | 0 |
| Xai Grok 4 | $7.20 | $36.00 | 0 |
| OpenAI o1 | $15.00 | $60.00 | 0 |
Frequently asked questions
How much does Replicate cost per 1M tokens?
Replicate pricing starts at $0.250 per 1 million output tokens (Ibm Granite Granite 3.3 8B Instruct) across 40 models, with its current flagship OpenAI o1 at $60.00. Older premium models in the catalog list higher. Rates current as of August 11, 2026.
What is the cheapest Replicate model?
Ibm Granite Granite 3.3 8B Instruct is the cheapest Replicate model on both axes — $0.030 per 1M input tokens and $0.250 per 1M output. It is the right default for high-volume, low-complexity work such as classification, extraction and routing.
What is the largest Replicate context window?
DeepSeek Ai DeepSeek V3.1 has the largest context window in the Replicate catalog at 164K input tokens, with up to 164K output tokens per response.
How do I calculate my actual Replicate bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator