Together AI Pricing
Together AI charges from $0.100 per 1 million output tokens, depending on the model. Arize Ai Qwen 2 1.5b Instruct is the cheapest at $0.100/1M output and $0.100/1M input; its current flagship Kimi K3 costs $15.00/1M output. The largest context window is 1.0M tokens (DeepSeek Ai DeepSeek V4 Flash 0731). Prices are USD, current as of August 31, 2026.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.030/1M
LFM2-24B-A2B
Cheapest output
$0.100/1M
Arize Ai Qwen 2 1.5b Instruct
Longest context
1.0M
DeepSeek Ai DeepSeek V4 Flash 0731
Models priced
93
Avg $2.34/1M output
Pricing by model
All prices in USD per 1 million tokens. Showing 80 of 93 priced models, cheapest first.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| Arize Ai Qwen 2 1.5b Instruct | $0.100 | $0.100 | 33K |
| Together Ai Up To 4B | $0.100 | $0.100 | 0 |
| Gemma 3N E4B Instruct | $0.060 | $0.120 | 33K |
| Google Gemma 3n E4b It | $0.060 | $0.120 | 33K |
| LFM2-24B-A2B | $0.030 | $0.120 | 33K |
| Meta Llama 3 8B Instruct Lite | $0.140 | $0.140 | 8K |
| Rnj-1 Instruct Tools | $0.150 | $0.150 | 33K |
| Meta Llama Meta Llama 3.1 8B Instruct Turbo | $0.180 | $0.180 | 0 |
| GPT OSS 20B ReasoningTools | $0.050 | $0.200 | 131K |
| Meta Llama Llama Guard 4 12B | $0.200 | $0.200 | 1.0M |
| OpenAI GPT Oss 20B | $0.050 | $0.200 | 131K |
| Together Ai 4.1b 8B | $0.200 | $0.200 | 0 |
| Qwen Qwen3.5 9B | $0.170 | $0.250 | 262K |
| Qwen3.5 9B VisionReasoningTools | $0.170 | $0.250 | 262K |
| DeepSeek Ai DeepSeek V4 Flash 0731 | $0.140 | $0.280 | 1.0M |
| DeepSeek V4 Flash 0731 ReasoningToolsCache | $0.140 | $0.280 | 1.0M |
| Qwen 2.5 7B Instruct Turbo Tools | $0.300 | $0.300 | 33K |
| Together Ai 8.1b 21B | $0.300 | $0.300 | 1K |
| GLM-5.3-Flash VisionReasoningToolsCache | $0.150 | $0.500 | 1.0M |
| Zai Org Glm 5.3 Flash | $0.150 | $0.500 | 1.0M |
| Meta Llama Llama 4 Scout 17B 16e Instruct | $0.180 | $0.590 | 0 |
| GPT OSS 120B ReasoningTools | $0.150 | $0.600 | 131K |
| Mistralai Mixtral 8X7B Instruct V0.1 | $0.600 | $0.600 | 0 |
| OpenAI GPT Oss 120B | $0.150 | $0.600 | 131K |
| Qwen Qwen3 235B A22b Fp8 Tput | $0.200 | $0.600 | 40K |
| Qwen3 235B A22B Instruct 2507 FP8 Tools | $0.200 | $0.600 | 262K |
| Together Ai 21.1b 41B | $0.800 | $0.800 | 0 |
| Meta Llama Llama 4 Maverick 17B 128e Instruct Fp8 | $0.270 | $0.850 | 0 |
| Pearl AI Gemma 4 31B Instruct Reasoning | $0.280 | $0.860 | 32K |
| Pearl Ai Gemma 4 31B It | $0.280 | $0.860 | 262K |
| Meta Llama Meta Llama 3.1 70B Instruct Turbo | $0.880 | $0.880 | 0 |
| Together Ai 41.1b 80B | $0.900 | $0.900 | 0 |
| Gemma 4 31B Instruct VisionReasoningTools | $0.390 | $0.970 | 262K |
| Google Gemma 4 31B It | $0.390 | $0.970 | 262K |
| Llama 3.3 70B Tools | $1.04 | $1.04 | 131K |
| Meta Llama Llama 3.3 70B Instruct Turbo | $1.04 | $1.04 | 131K |
| Zai Org Glm 4.5 Air Fp8 | $0.200 | $1.10 | 128K |
| MiniMax-M2.5 ReasoningToolsCache | $0.300 | $1.20 | 205K |
| MiniMax-M2.7 ReasoningToolsCache | $0.300 | $1.20 | 203K |
| MiniMax-M3 VisionReasoningToolsCache | $0.300 | $1.20 | 524K |
| Minimaxai Minimax M3 | $0.300 | $1.20 | 524K |
| Qwen3 Coder Next FP8 Tools | $0.500 | $1.20 | 262K |
| Thinkingmachines Inkling Small | $0.500 | $1.20 | 524K |
| Cogito v2.1 671B Reasoning | $1.25 | $1.25 | 164K |
| DeepSeek Ai DeepSeek V3 | $1.25 | $1.25 | 66K |
| DeepSeek-V3 Tools | $1.25 | $1.25 | 131K |
| Qwen Qwen3.7 Plus | $0.320 | $1.28 | 1.0M |
| Meta Models Muse Glimmer 30B | $0.350 | $1.50 | 131K |
| Qwen Qwen3 Next 80B A3b Instruct | $0.150 | $1.50 | 262K |
| Qwen Qwen3 Next 80B A3b Thinking | $0.150 | $1.50 | 262K |
| DeepSeek Ai DeepSeek V3.1 | $0.600 | $1.70 | 128K |
| DeepSeek V3.1 ReasoningTools | $0.600 | $1.70 | 131K |
| Together Ai 81.1b 110B | $1.80 | $1.80 | 0 |
| Qwen Qwen3 Coder 480B A35b Instruct Fp8 | $2.00 | $2.00 | 256K |
| Qwen3 Coder 480B A35B Instruct Tools | $2.00 | $2.00 | 262K |
| Zai Org Glm 4.7 | $0.450 | $2.00 | 200K |
| DeepSeek Ai DeepSeek R1 0528 Tput | $0.550 | $2.19 | 128K |
| Zai Org Glm 4.6 | $0.600 | $2.20 | 200K |
| Kimi K2.5 VisionReasoningTools | $0.500 | $2.80 | 262K |
| Moonshotai Kimi K2.5 | $0.500 | $2.80 | 256K |
| Moonshotai Kimi K2 Instruct | $1.00 | $3.00 | 0 |
| Moonshotai Kimi K2 Instruct 0905 | $1.00 | $3.00 | 262K |
| Qwen Qwen3 235B A22b Thinking 2507 | $0.650 | $3.00 | 256K |
| Qwen Qwen3.6 Plus | $0.500 | $3.00 | 1.0M |
| Qwen3.6 Plus ReasoningTools | $0.500 | $3.00 | 1.0M |
| GLM-5 ReasoningTools | $1.00 | $3.20 | 203K |
| DeepSeek Ai DeepSeek V4 Pro | $1.74 | $3.48 | 512K |
| DeepSeek V4 Pro ReasoningToolsCache | $1.74 | $3.48 | 512K |
| Meta Llama Meta Llama 3.1 405B Instruct Turbo | $3.50 | $3.50 | 0 |
| Nemotron 3 Ultra 550B A55B ReasoningToolsCache | $0.600 | $3.60 | 512K |
| Nvidia Nemotron 3 Ultra 550B A55b | $0.600 | $3.60 | 512K |
| Qwen Qwen3.5 397B A17b | $0.600 | $3.60 | 262K |
| Qwen3.5 397B A17B VisionReasoningToolsCache | $0.600 | $3.60 | 262K |
| Qwen3.7 Max ToolsCache | $1.25 | $3.75 | 1.0M |
| DeepSeek Ai DeepSeek V4 Pro 0813 | $1.32 | $3.96 | 1.0M |
| DeepSeek V4 Pro 0813 ReasoningToolsCache | $1.32 | $3.96 | 1.0M |
| Kimi K2.7 Code ReasoningToolsCache | $0.950 | $4.00 | 262K |
| Moonshotai Kimi K2.7 Code | $0.950 | $4.00 | 262K |
| Inkling VisionReasoningToolsCache | $1.00 | $4.05 | 524K |
| Moonshotai Kimi K3 | $3.00 | $15.00 | 1.0M |
Frequently asked questions
How much does Together AI cost per 1M tokens?
Together AI pricing starts at $0.100 per 1 million output tokens (Arize Ai Qwen 2 1.5b Instruct) across 93 models, with its current flagship Kimi K3 at $15.00. Older premium models in the catalog list higher. Rates current as of August 31, 2026.
What is the cheapest Together AI model?
Arize Ai Qwen 2 1.5b Instruct is the cheapest Together AI model on output tokens at $0.100 per 1M, while LFM2-24B-A2B is cheapest on input at $0.030 per 1M. Which wins for you depends on your input-to-output ratio.
What is the largest Together AI context window?
DeepSeek Ai DeepSeek V4 Flash 0731 has the largest context window in the Together AI catalog at 1.0M input tokens, with up to 1.0M output tokens per response.
Does Together AI support prompt caching?
Yes. DeepSeek V4 Flash 0731 reads cached input at $0.030 per 1M tokens, against $0.140 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.
Which Together AI models support reasoning or vision?
The Together AI catalog includes 27 reasoning models, 9 vision models, 32 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.
How do I calculate my actual Together AI bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator