API pricing

Together AI Pricing

Together AI charges between $0.100 and $4.40 per 1 million output tokens, depending on the model. Together Ai Up To 4B is the cheapest at $0.100/1M output and $0.100/1M input; GLM-5.1 is the most expensive at $4.40/1M output. The largest context window is 1.0M tokens (Qwen3.6 Plus). Prices are USD, current as of July 20, 2026.

Last updated · synced weekly from the upstream model catalog

Cheapest input

$0.030/1M

LFM2-24B-A2B

Cheapest output

$0.100/1M

Together Ai Up To 4B

Longest context

1.0M

Qwen3.6 Plus

Models priced

65

Avg $1.92/1M output

Pricing by model

All prices in USD per 1 million tokens. All 65 priced models, cheapest first.

ModelInput / 1MOutput / 1MContext
Together Ai Up To 4B$0.100$0.1000
Gemma 3N E4B Instruct$0.060$0.12033K
LFM2-24B-A2B$0.030$0.12033K
Meta Llama 3 8B Instruct Lite$0.140$0.1408K
Rnj-1 Instruct
Tools
$0.150$0.15033K
Meta Llama Meta Llama 3.1 8B Instruct Turbo$0.180$0.1800
GPT OSS 20B
ReasoningTools
$0.050$0.200131K
OpenAI GPT Oss 20B$0.050$0.200128K
Together Ai 4.1b 8B$0.200$0.2000
Qwen3.5 9B
VisionReasoningTools
$0.170$0.250262K
Qwen 2.5 7B Instruct Turbo
Tools
$0.300$0.30033K
Together Ai 8.1b 21B$0.300$0.3001K
Meta Llama Llama 4 Scout 17B 16e Instruct$0.180$0.5900
GPT OSS 120B
ReasoningTools
$0.150$0.600131K
Mistralai Mixtral 8X7B Instruct V0.1$0.600$0.6000
OpenAI GPT Oss 120B$0.150$0.600131K
Qwen Qwen3 235B A22b Fp8 Tput$0.200$0.60040K
Qwen3 235B A22B Instruct 2507 FP8
Tools
$0.200$0.600262K
Together Ai 21.1b 41B$0.800$0.8000
Meta Llama Llama 4 Maverick 17B 128e Instruct Fp8$0.270$0.8500
Pearl AI Gemma 4 31B Instruct
Reasoning
$0.280$0.86032K
Meta Llama Llama 3.3 70B Instruct Turbo$0.880$0.8800
Meta Llama Meta Llama 3.1 70B Instruct Turbo$0.880$0.8800
Together Ai 41.1b 80B$0.900$0.9000
Gemma 4 31B Instruct
VisionReasoningTools
$0.390$0.970262K
Llama 3.3 70B
Tools
$1.04$1.04131K
Zai Org Glm 4.5 Air Fp8$0.200$1.10128K
MiniMax-M2.5
ReasoningToolsCache
$0.300$1.20205K
MiniMax-M2.7
ReasoningToolsCache
$0.300$1.20203K
MiniMax-M3
VisionReasoningToolsCache
$0.300$1.20524K
Qwen3 Coder Next FP8
Tools
$0.500$1.20262K
Cogito v2.1 671B
Reasoning
$1.25$1.25164K
DeepSeek Ai DeepSeek V3$1.25$1.2566K
DeepSeek-V3
Tools
$1.25$1.25131K
Qwen Qwen3 Next 80B A3b Instruct$0.150$1.50262K
Qwen Qwen3 Next 80B A3b Thinking$0.150$1.50262K
DeepSeek Ai DeepSeek V3.1$0.600$1.70128K
DeepSeek V3.1
ReasoningTools
$0.600$1.70131K
Together Ai 81.1b 110B$1.80$1.800
Qwen Qwen3 Coder 480B A35b Instruct Fp8$2.00$2.00256K
Qwen3 Coder 480B A35B Instruct
Tools
$2.00$2.00262K
Zai Org Glm 4.7$0.450$2.00200K
DeepSeek Ai DeepSeek R1 0528 Tput$0.550$2.19128K
Zai Org Glm 4.6$0.600$2.20200K
Kimi K2.5
VisionReasoningTools
$0.500$2.80262K
Moonshotai Kimi K2.5$0.500$2.80256K
Moonshotai Kimi K2 Instruct$1.00$3.000
Moonshotai Kimi K2 Instruct 0905$1.00$3.00262K
Qwen Qwen3 235B A22b Thinking 2507$0.650$3.00256K
Qwen3.6 Plus
ReasoningTools
$0.500$3.001.0M
GLM-5
ReasoningTools
$1.00$3.20203K
DeepSeek V4 Pro
ReasoningToolsCache
$1.74$3.48512K
Meta Llama Meta Llama 3.1 405B Instruct Turbo$3.50$3.500
Nemotron 3 Ultra 550B A55B
ReasoningToolsCache
$0.600$3.60512K
Qwen Qwen3.5 397B A17b$0.600$3.60262K
Qwen3.5 397B A17B
VisionReasoningTools
$0.600$3.60262K
Qwen3.7 Max
Tools
$1.25$3.751.0M
Kimi K2.7 Code
ReasoningToolsCache
$0.950$4.00262K
Inkling
VisionReasoningToolsCache
$1.00$4.05524K
GLM-5.1
ReasoningTools
$1.40$4.40203K
GLM-5.2
ReasoningToolsCache
$1.40$4.40262K
Kimi K2.6
VisionReasoningToolsCache
$1.20$4.50262K
Qwen Qwen3 235B A22b Instruct 2507 Tput$0.200$6.00262K
DeepSeek Ai DeepSeek R1$3.00$7.00128K
DeepSeek-R1
Reasoning
$3.00$7.00164K

Frequently asked questions

How much does Together AI cost per 1M tokens?

Together AI pricing ranges from $0.100 to $4.40 per 1 million output tokens across 65 models. Input tokens are cheaper than output on every model. Rates current as of July 20, 2026.

What is the cheapest Together AI model?

Together Ai Up To 4B is the cheapest Together AI model on output tokens at $0.100 per 1M, while LFM2-24B-A2B is cheapest on input at $0.030 per 1M. Which wins for you depends on your input-to-output ratio.

What is the largest Together AI context window?

Qwen3.6 Plus has the largest context window in the Together AI catalog at 1.0M input tokens, with up to 500K output tokens per response.

Does Together AI support prompt caching?

Yes. MiniMax-M2.5 reads cached input at $0.060 per 1M tokens, against $0.300 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.

Which Together AI models support reasoning or vision?

The Together AI catalog includes 22 reasoning models, 7 vision models, 27 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.

How do I calculate my actual Together AI bill?

Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.

Other providers

What will this actually cost you?

Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.

Open the free cost calculator