Novita Pricing
Novita charges from $0.020 per 1 million output tokens, depending on the model. Meta Llama Llama 3.2 1B Instruct is the cheapest at $0.020/1M output and $0.020/1M input; its current flagship Kimi K3 costs $15.00/1M output. The largest context window is 1.0M tokens (DeepSeek DeepSeek V4 Flash). Prices are USD, current as of August 31, 2026.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.020/1M
Llama 3.1 8B Instruct
Cheapest output
$0.020/1M
Meta Llama Llama 3.2 1B Instruct
Longest context
1.0M
DeepSeek DeepSeek V4 Flash
Models priced
230
Avg $1.53/1M output
Pricing by model
All prices in USD per 1 million tokens. Showing 80 of 230 priced models, cheapest first.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| Meta Llama Llama 3.2 1B Instruct | $0.020 | $0.020 | 131K |
| PaddleOCR-VL Vision | $0.020 | $0.020 | 16K |
| Paddlepaddle Paddleocr VL | $0.020 | $0.020 | 16K |
| DeepSeek DeepSeek Ocr | $0.030 | $0.030 | 8K |
| DeepSeek DeepSeek Ocr 2 | $0.030 | $0.030 | 8K |
| DeepSeek-OCR Vision | $0.030 | $0.030 | 8K |
| Qwen Qwen3 4B Fp8 | $0.030 | $0.030 | 128K |
| Qwen3 4B Reasoning | $0.030 | $0.030 | 128K |
| Llama 3 8B Instruct | $0.040 | $0.040 | 8K |
| Meta Llama Llama 3 8B Instruct | $0.040 | $0.040 | 8K |
| L3 8B Stheno V3.2 Tools | $0.050 | $0.050 | 8K |
| Llama 3.1 8B Instruct | $0.020 | $0.050 | 16K |
| Llama 3.2 3B Instruct | $0.030 | $0.050 | 33K |
| Meta Llama Llama 3.1 8B Instruct | $0.020 | $0.050 | 16K |
| Meta Llama Llama 3.2 3B Instruct | $0.030 | $0.050 | 33K |
| Sao10k L3 8B Lunaris | $0.050 | $0.050 | 8K |
| Sao10k L3 8B Stheno V3.2 | $0.050 | $0.050 | 8K |
| Baichuan Baichuan M2 32B | $0.070 | $0.070 | 131K |
| baichuan-m2-32b | $0.070 | $0.070 | 131K |
| Qwen Qwen2.5 7B Instruct | $0.070 | $0.070 | 32K |
| Qwen2.5 7B Instruct Tools | $0.070 | $0.070 | 32K |
| DeepSeek DeepSeek R1 0528 Qwen3 8B | $0.060 | $0.090 | 128K |
| DeepSeek R1 0528 Qwen3 8B Reasoning | $0.060 | $0.090 | 128K |
| Gryphe Mythomax L2 13B | $0.090 | $0.090 | 4K |
| Mythomax L2 13B | $0.090 | $0.090 | 4K |
| Gemma 3 12B Vision | $0.050 | $0.100 | 131K |
| Google Gemma 3 12B It | $0.050 | $0.100 | 131K |
| AutoGLM-Phone-9B-Multilingual Vision | $0.035 | $0.138 | 66K |
| Qwen Qwen3 8B Fp8 | $0.035 | $0.138 | 128K |
| Qwen3 8B Reasoning | $0.035 | $0.138 | 128K |
| Zai Org Autoglm Phone 9B Multilingual | $0.035 | $0.138 | 66K |
| Hermes 2 Pro Llama 3 8B | $0.140 | $0.140 | 8K |
| Nousresearch Hermes 2 Pro Llama 3 8B | $0.140 | $0.140 | 8K |
| DeepSeek DeepSeek R1 Distill Qwen 14B | $0.150 | $0.150 | 33K |
| DeepSeek R1 Distill Qwen 14B | $0.150 | $0.150 | 33K |
| OpenAI GPT Oss 20B | $0.040 | $0.150 | 131K |
| Mistral Nemo | $0.040 | $0.170 | 60K |
| Mistralai Mistral Nemo | $0.040 | $0.170 | 60K |
| Inclusionai Ling 3.0 Flash | $0.060 | $0.180 | 262K |
| Inclusionai Ling 3.0 Flash Fast | $0.060 | $0.180 | 262K |
| Gemma 3 27B Vision | $0.119 | $0.200 | 98K |
| Google Gemma 3 27B It | $0.119 | $0.200 | 98K |
| Nvidia Nemotron 3 Nano 30B A3b | $0.050 | $0.200 | 262K |
| OpenAI GPT OSS 120B VisionReasoningTools | $0.050 | $0.250 | 131K |
| Qwen Qwen3 Coder 30B A3b Instruct | $0.070 | $0.270 | 160K |
| Qwen3 Coder 30b A3B Instruct Tools | $0.070 | $0.270 | 160K |
| Baidu Ernie 4.5 21B A3b | $0.070 | $0.280 | 120K |
| Baidu Ernie 4.5 21B A3b Thinking | $0.070 | $0.280 | 131K |
| DeepSeek DeepSeek V4 Flash | $0.140 | $0.280 | 1.0M |
| DeepSeek V4 Flash ReasoningToolsCache | $0.140 | $0.280 | 1.0M |
| ERNIE 4.5 21B A3B Tools | $0.070 | $0.280 | 120K |
| ERNIE-4.5-21B-A3B-Thinking Reasoning | $0.070 | $0.280 | 131K |
| DeepSeek DeepSeek R1 Distill Qwen 32B | $0.300 | $0.300 | 64K |
| DeepSeek R1 Distill Qwen 32B | $0.300 | $0.300 | 64K |
| Ling-2.6-flash ToolsCache | $0.100 | $0.300 | 262K |
| XiaomiMiMo/MiMo-V2-Flash ReasoningToolsCache | $0.100 | $0.300 | 262K |
| Xiaomimimo Mimo V2 Flash | $0.110 | $0.330 | 262K |
| Xiaomimimo Mimo V2.5 | $0.168 | $0.336 | 1.0M |
| Baidu Ernie 4.5 VL 28B A3b Thinking | $0.390 | $0.390 | 131K |
| ERNIE-4.5-VL-28B-A3B-Thinking VisionReasoningTools | $0.390 | $0.390 | 131K |
| DeepSeek DeepSeek V3.2 | $0.269 | $0.400 | 164K |
| Deepseek V3.2 ReasoningToolsCache | $0.269 | $0.400 | 164K |
| Gemma 4 26B A4B VisionReasoningTools | $0.130 | $0.400 | 262K |
| Gemma 4 31B VisionReasoningTools | $0.140 | $0.400 | 262K |
| GLM-4.7-Flash ReasoningToolsCache | $0.070 | $0.400 | 200K |
| Google Gemma 4 26B A4b It | $0.130 | $0.400 | 262K |
| Google Gemma 4 31B It | $0.140 | $0.400 | 262K |
| Llama 3.3 70B Instruct Tools | $0.135 | $0.400 | 131K |
| Meta Llama Llama 3.3 70B Instruct | $0.135 | $0.400 | 12K |
| Qwen 2.5 72B Instruct Tools | $0.380 | $0.400 | 32K |
| Qwen Qwen 2.5 72B Instruct | $0.380 | $0.400 | 32K |
| Zai Org Glm 4.7 Flash | $0.070 | $0.400 | 200K |
| DeepSeek DeepSeek V3.2 Exp | $0.270 | $0.410 | 164K |
| Deepseek V3.2 Exp ReasoningTools | $0.270 | $0.410 | 164K |
| Qwen Qwen3 30B A3b Fp8 | $0.090 | $0.450 | 41K |
| Qwen Qwen3 32B Fp8 | $0.100 | $0.450 | 41K |
| Qwen3 30B A3B Reasoning | $0.090 | $0.450 | 41K |
| Qwen3 32B Reasoning | $0.100 | $0.450 | 41K |
| Qwen Qwen3 VL 8B Instruct | $0.080 | $0.500 | 131K |
| Moonshotai Kimi K3 | $3.00 | $15.00 | 1.0M |
Frequently asked questions
How much does Novita cost per 1M tokens?
Novita pricing starts at $0.020 per 1 million output tokens (Meta Llama Llama 3.2 1B Instruct) across 230 models, with its current flagship Kimi K3 at $15.00. Older premium models in the catalog list higher. Rates current as of August 31, 2026.
What is the cheapest Novita model?
Meta Llama Llama 3.2 1B Instruct is the cheapest Novita model on output tokens at $0.020 per 1M, while Llama 3.1 8B Instruct is cheapest on input at $0.020 per 1M. Which wins for you depends on your input-to-output ratio.
What is the largest Novita context window?
DeepSeek DeepSeek V4 Flash has the largest context window in the Novita catalog at 1.0M input tokens, with up to 393K output tokens per response.
Does Novita support prompt caching?
Yes. MiMo-V2.5-Pro reads cached input at $0.0043 per 1M tokens, against $0.522 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.
Which Novita models support reasoning or vision?
The Novita catalog includes 57 reasoning models, 33 vision models, 72 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.
How do I calculate my actual Novita bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator