Novita Pricing
Novita charges between $0.020 and $6.00 per 1 million output tokens, depending on the model. PaddleOCR-VL is the cheapest at $0.020/1M output and $0.020/1M input; MiMo-V2-Pro is the most expensive at $6.00/1M output. The largest context window is 1.0M tokens (DeepSeek V4 Flash). Prices are USD, current as of July 20, 2026.
Last updated · synced weekly from the upstream model catalog
Cheapest input
$0.020/1M
Llama 3.1 8B Instruct
Cheapest output
$0.020/1M
PaddleOCR-VL
Longest context
1.0M
DeepSeek V4 Flash
Models priced
178
Avg $1.19/1M output
Pricing by model
All prices in USD per 1 million tokens. Showing 80 of 178 priced models, cheapest first.
| Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| PaddleOCR-VL Vision | $0.020 | $0.020 | 16K |
| Paddlepaddle Paddleocr VL | $0.020 | $0.020 | 16K |
| DeepSeek DeepSeek Ocr | $0.030 | $0.030 | 8K |
| DeepSeek-OCR Vision | $0.030 | $0.030 | 8K |
| deepseek/deepseek-ocr-2 Vision | $0.030 | $0.030 | 8K |
| Qwen Qwen3 4B Fp8 | $0.030 | $0.030 | 128K |
| Qwen3 4B Reasoning | $0.030 | $0.030 | 128K |
| Llama 3 8B Instruct | $0.040 | $0.040 | 8K |
| Meta Llama Llama 3 8B Instruct | $0.040 | $0.040 | 8K |
| L3 8B Stheno V3.2 Tools | $0.050 | $0.050 | 8K |
| Llama 3.1 8B Instruct | $0.020 | $0.050 | 16K |
| Llama 3.2 3B Instruct | $0.030 | $0.050 | 33K |
| Meta Llama Llama 3.1 8B Instruct | $0.020 | $0.050 | 16K |
| Meta Llama Llama 3.2 3B Instruct | $0.030 | $0.050 | 33K |
| Sao10k L3 8B Lunaris | $0.050 | $0.050 | 8K |
| Sao10k L3 8B Stheno V3.2 | $0.050 | $0.050 | 8K |
| Baichuan Baichuan M2 32B | $0.070 | $0.070 | 131K |
| baichuan-m2-32b | $0.070 | $0.070 | 131K |
| Qwen Qwen2.5 7B Instruct | $0.070 | $0.070 | 32K |
| Qwen2.5 7B Instruct Tools | $0.070 | $0.070 | 32K |
| DeepSeek DeepSeek R1 0528 Qwen3 8B | $0.060 | $0.090 | 128K |
| DeepSeek R1 0528 Qwen3 8B Reasoning | $0.060 | $0.090 | 128K |
| Gryphe Mythomax L2 13B | $0.090 | $0.090 | 4K |
| Mythomax L2 13B | $0.090 | $0.090 | 4K |
| Gemma 3 12B Vision | $0.050 | $0.100 | 131K |
| Google Gemma 3 12B It | $0.050 | $0.100 | 131K |
| AutoGLM-Phone-9B-Multilingual Vision | $0.035 | $0.138 | 66K |
| Qwen Qwen3 8B Fp8 | $0.035 | $0.138 | 128K |
| Qwen3 8B Reasoning | $0.035 | $0.138 | 128K |
| Zai Org Autoglm Phone 9B Multilingual | $0.035 | $0.138 | 66K |
| Hermes 2 Pro Llama 3 8B | $0.140 | $0.140 | 8K |
| Nousresearch Hermes 2 Pro Llama 3 8B | $0.140 | $0.140 | 8K |
| DeepSeek DeepSeek R1 Distill Qwen 14B | $0.150 | $0.150 | 33K |
| DeepSeek R1 Distill Qwen 14B | $0.150 | $0.150 | 33K |
| OpenAI GPT Oss 20B | $0.040 | $0.150 | 131K |
| Mistral Nemo | $0.040 | $0.170 | 60K |
| Mistralai Mistral Nemo | $0.040 | $0.170 | 60K |
| Gemma 3 27B Vision | $0.119 | $0.200 | 98K |
| Google Gemma 3 27B It | $0.119 | $0.200 | 98K |
| OpenAI GPT OSS 120B VisionReasoningTools | $0.050 | $0.250 | 131K |
| Qwen Qwen3 Coder 30B A3b Instruct | $0.070 | $0.270 | 160K |
| Qwen3 Coder 30b A3B Instruct Tools | $0.070 | $0.270 | 160K |
| Baidu Ernie 4.5 21B A3b | $0.070 | $0.280 | 120K |
| Baidu Ernie 4.5 21B A3b Thinking | $0.070 | $0.280 | 131K |
| DeepSeek V4 Flash ReasoningToolsCache | $0.140 | $0.280 | 1.0M |
| ERNIE 4.5 21B A3B Tools | $0.070 | $0.280 | 120K |
| ERNIE-4.5-21B-A3B-Thinking Reasoning | $0.070 | $0.280 | 131K |
| DeepSeek DeepSeek R1 Distill Qwen 32B | $0.300 | $0.300 | 64K |
| DeepSeek R1 Distill Qwen 32B | $0.300 | $0.300 | 64K |
| Ling-2.6-flash ToolsCache | $0.100 | $0.300 | 262K |
| Xiaomimimo Mimo V2 Flash | $0.100 | $0.300 | 262K |
| Baidu Ernie 4.5 VL 28B A3b Thinking | $0.390 | $0.390 | 131K |
| ERNIE-4.5-VL-28B-A3B-Thinking VisionReasoningTools | $0.390 | $0.390 | 131K |
| DeepSeek DeepSeek V3.2 | $0.269 | $0.400 | 164K |
| Deepseek V3.2 ReasoningToolsCache | $0.269 | $0.400 | 164K |
| Gemma 4 26B A4B VisionReasoningTools | $0.130 | $0.400 | 262K |
| Gemma 4 31B VisionReasoningTools | $0.140 | $0.400 | 262K |
| GLM-4.7-Flash ReasoningToolsCache | $0.070 | $0.400 | 200K |
| Llama 3.3 70B Instruct Tools | $0.135 | $0.400 | 131K |
| Meta Llama Llama 3.3 70B Instruct | $0.135 | $0.400 | 131K |
| Qwen 2.5 72B Instruct Tools | $0.380 | $0.400 | 32K |
| Qwen Qwen 2.5 72B Instruct | $0.380 | $0.400 | 32K |
| DeepSeek DeepSeek V3.2 Exp | $0.270 | $0.410 | 164K |
| Deepseek V3.2 Exp ReasoningTools | $0.270 | $0.410 | 164K |
| Qwen Qwen3 30B A3b Fp8 | $0.090 | $0.450 | 41K |
| Qwen Qwen3 32B Fp8 | $0.100 | $0.450 | 41K |
| Qwen3 30B A3B Reasoning | $0.090 | $0.450 | 41K |
| Qwen3 32B Reasoning | $0.100 | $0.450 | 41K |
| Qwen Qwen3 VL 8B Instruct | $0.080 | $0.500 | 131K |
| Baidu Ernie 4.5 VL 28B A3b | $0.140 | $0.560 | 30K |
| ERNIE 4.5 VL 28B A3B VisionReasoningTools | $0.140 | $0.560 | 30K |
| Qwen Qwen3 235B A22b Instruct 2507 | $0.090 | $0.580 | 131K |
| Qwen3 235B A22B Instruct 2507 Tools | $0.090 | $0.580 | 131K |
| Llama 4 Scout Instruct Vision | $0.180 | $0.590 | 131K |
| Meta Llama Llama 4 Scout 17B 16e Instruct | $0.180 | $0.590 | 131K |
| Skywork R1v4 Lite | $0.200 | $0.600 | 262K |
| Microsoft Wizardlm 2 8X22B | $0.620 | $0.620 | 66K |
| Wizardlm 2 8x22B | $0.620 | $0.620 | 66K |
| Qwen Qwen3 VL 30B A3b Instruct | $0.200 | $0.700 | 131K |
| Qwen3 Max Tools | $2.11 | $8.45 | 262K |
Frequently asked questions
How much does Novita cost per 1M tokens?
Novita pricing ranges from $0.020 to $6.00 per 1 million output tokens across 178 models. Input tokens are cheaper than output on every model. Rates current as of July 20, 2026.
What is the cheapest Novita model?
PaddleOCR-VL is the cheapest Novita model on output tokens at $0.020 per 1M, while Llama 3.1 8B Instruct is cheapest on input at $0.020 per 1M. Which wins for you depends on your input-to-output ratio.
What is the largest Novita context window?
DeepSeek V4 Flash has the largest context window in the Novita catalog at 1.0M input tokens, with up to 393K output tokens per response.
Does Novita support prompt caching?
Yes. MiMo-V2.5-Pro reads cached input at $0.0043 per 1M tokens, against $0.522 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.
Which Novita models support reasoning or vision?
The Novita catalog includes 55 reasoning models, 31 vision models, 70 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.
How do I calculate my actual Novita bill?
Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.
Other providers
What will this actually cost you?
Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.
Open the free cost calculator