API pricing

Novita Pricing

Novita charges between $0.020 and $6.00 per 1 million output tokens, depending on the model. PaddleOCR-VL is the cheapest at $0.020/1M output and $0.020/1M input; MiMo-V2-Pro is the most expensive at $6.00/1M output. The largest context window is 1.0M tokens (DeepSeek V4 Flash). Prices are USD, current as of July 20, 2026.

Last updated · synced weekly from the upstream model catalog

Cheapest input

$0.020/1M

Llama 3.1 8B Instruct

Cheapest output

$0.020/1M

PaddleOCR-VL

Longest context

1.0M

DeepSeek V4 Flash

Models priced

178

Avg $1.19/1M output

Pricing by model

All prices in USD per 1 million tokens. Showing 80 of 178 priced models, cheapest first.

ModelInput / 1MOutput / 1MContext
PaddleOCR-VL
Vision
$0.020$0.02016K
Paddlepaddle Paddleocr VL$0.020$0.02016K
DeepSeek DeepSeek Ocr$0.030$0.0308K
DeepSeek-OCR
Vision
$0.030$0.0308K
deepseek/deepseek-ocr-2
Vision
$0.030$0.0308K
Qwen Qwen3 4B Fp8$0.030$0.030128K
Qwen3 4B
Reasoning
$0.030$0.030128K
Llama 3 8B Instruct$0.040$0.0408K
Meta Llama Llama 3 8B Instruct$0.040$0.0408K
L3 8B Stheno V3.2
Tools
$0.050$0.0508K
Llama 3.1 8B Instruct$0.020$0.05016K
Llama 3.2 3B Instruct$0.030$0.05033K
Meta Llama Llama 3.1 8B Instruct$0.020$0.05016K
Meta Llama Llama 3.2 3B Instruct$0.030$0.05033K
Sao10k L3 8B Lunaris$0.050$0.0508K
Sao10k L3 8B Stheno V3.2$0.050$0.0508K
Baichuan Baichuan M2 32B$0.070$0.070131K
baichuan-m2-32b$0.070$0.070131K
Qwen Qwen2.5 7B Instruct$0.070$0.07032K
Qwen2.5 7B Instruct
Tools
$0.070$0.07032K
DeepSeek DeepSeek R1 0528 Qwen3 8B$0.060$0.090128K
DeepSeek R1 0528 Qwen3 8B
Reasoning
$0.060$0.090128K
Gryphe Mythomax L2 13B$0.090$0.0904K
Mythomax L2 13B$0.090$0.0904K
Gemma 3 12B
Vision
$0.050$0.100131K
Google Gemma 3 12B It$0.050$0.100131K
AutoGLM-Phone-9B-Multilingual
Vision
$0.035$0.13866K
Qwen Qwen3 8B Fp8$0.035$0.138128K
Qwen3 8B
Reasoning
$0.035$0.138128K
Zai Org Autoglm Phone 9B Multilingual$0.035$0.13866K
Hermes 2 Pro Llama 3 8B$0.140$0.1408K
Nousresearch Hermes 2 Pro Llama 3 8B$0.140$0.1408K
DeepSeek DeepSeek R1 Distill Qwen 14B$0.150$0.15033K
DeepSeek R1 Distill Qwen 14B$0.150$0.15033K
OpenAI GPT Oss 20B$0.040$0.150131K
Mistral Nemo$0.040$0.17060K
Mistralai Mistral Nemo$0.040$0.17060K
Gemma 3 27B
Vision
$0.119$0.20098K
Google Gemma 3 27B It$0.119$0.20098K
OpenAI GPT OSS 120B
VisionReasoningTools
$0.050$0.250131K
Qwen Qwen3 Coder 30B A3b Instruct$0.070$0.270160K
Qwen3 Coder 30b A3B Instruct
Tools
$0.070$0.270160K
Baidu Ernie 4.5 21B A3b$0.070$0.280120K
Baidu Ernie 4.5 21B A3b Thinking$0.070$0.280131K
DeepSeek V4 Flash
ReasoningToolsCache
$0.140$0.2801.0M
ERNIE 4.5 21B A3B
Tools
$0.070$0.280120K
ERNIE-4.5-21B-A3B-Thinking
Reasoning
$0.070$0.280131K
DeepSeek DeepSeek R1 Distill Qwen 32B$0.300$0.30064K
DeepSeek R1 Distill Qwen 32B$0.300$0.30064K
Ling-2.6-flash
ToolsCache
$0.100$0.300262K
Xiaomimimo Mimo V2 Flash$0.100$0.300262K
Baidu Ernie 4.5 VL 28B A3b Thinking$0.390$0.390131K
ERNIE-4.5-VL-28B-A3B-Thinking
VisionReasoningTools
$0.390$0.390131K
DeepSeek DeepSeek V3.2$0.269$0.400164K
Deepseek V3.2
ReasoningToolsCache
$0.269$0.400164K
Gemma 4 26B A4B
VisionReasoningTools
$0.130$0.400262K
Gemma 4 31B
VisionReasoningTools
$0.140$0.400262K
GLM-4.7-Flash
ReasoningToolsCache
$0.070$0.400200K
Llama 3.3 70B Instruct
Tools
$0.135$0.400131K
Meta Llama Llama 3.3 70B Instruct$0.135$0.400131K
Qwen 2.5 72B Instruct
Tools
$0.380$0.40032K
Qwen Qwen 2.5 72B Instruct$0.380$0.40032K
DeepSeek DeepSeek V3.2 Exp$0.270$0.410164K
Deepseek V3.2 Exp
ReasoningTools
$0.270$0.410164K
Qwen Qwen3 30B A3b Fp8$0.090$0.45041K
Qwen Qwen3 32B Fp8$0.100$0.45041K
Qwen3 30B A3B
Reasoning
$0.090$0.45041K
Qwen3 32B
Reasoning
$0.100$0.45041K
Qwen Qwen3 VL 8B Instruct$0.080$0.500131K
Baidu Ernie 4.5 VL 28B A3b$0.140$0.56030K
ERNIE 4.5 VL 28B A3B
VisionReasoningTools
$0.140$0.56030K
Qwen Qwen3 235B A22b Instruct 2507$0.090$0.580131K
Qwen3 235B A22B Instruct 2507
Tools
$0.090$0.580131K
Llama 4 Scout Instruct
Vision
$0.180$0.590131K
Meta Llama Llama 4 Scout 17B 16e Instruct$0.180$0.590131K
Skywork R1v4 Lite$0.200$0.600262K
Microsoft Wizardlm 2 8X22B$0.620$0.62066K
Wizardlm 2 8x22B$0.620$0.62066K
Qwen Qwen3 VL 30B A3b Instruct$0.200$0.700131K
Qwen3 Max
Tools
$2.11$8.45262K

Frequently asked questions

How much does Novita cost per 1M tokens?

Novita pricing ranges from $0.020 to $6.00 per 1 million output tokens across 178 models. Input tokens are cheaper than output on every model. Rates current as of July 20, 2026.

What is the cheapest Novita model?

PaddleOCR-VL is the cheapest Novita model on output tokens at $0.020 per 1M, while Llama 3.1 8B Instruct is cheapest on input at $0.020 per 1M. Which wins for you depends on your input-to-output ratio.

What is the largest Novita context window?

DeepSeek V4 Flash has the largest context window in the Novita catalog at 1.0M input tokens, with up to 393K output tokens per response.

Does Novita support prompt caching?

Yes. MiMo-V2.5-Pro reads cached input at $0.0043 per 1M tokens, against $0.522 per 1M uncached. Caching pays off when a long system prompt or document is reused across many requests.

Which Novita models support reasoning or vision?

The Novita catalog includes 55 reasoning models, 31 vision models, 70 with tool calling. Reasoning models bill their internal thinking as output tokens, so they cost more per visible response than the headline rate suggests.

How do I calculate my actual Novita bill?

Multiply your input tokens by the input rate and your output tokens by the output rate, both per million, then add retries and any conversation history resent on each turn. Calcaas does this against your real usage pattern and shows the margin left at your price point.

Other providers

What will this actually cost you?

Per-token rates only tell you half the story. Calcaas multiplies them by your real usage — prompt size, response length, retries and conversation history — then shows the margin left at your price point.

Open the free cost calculator