Back to Home
API Pricing

DeepSeek V4 Pro API Pricing

¥9.00 input · ¥27.00 output per 1M tokens at peak — 50% off off-peak. Structured reasoning with o1-level thinking.

DeepSeek V4 Pro is the flagship of the DeepSeek line: structured reasoning, complex code generation, and long-context mastery (1M tokens) with a Thinking Mode that produces chain-of-thought reasoning comparable to o1-level models — at a price point that makes it viable for high-volume production workloads.

Estimate My CostsPrices as of September 2026

Price per 1M tokens

DeepSeek V4 Pro · CNY official list price

TierRate
InputPeak ¥9.00 · Off-peak ¥4.50
OutputPeak ¥27.00 · Off-peak ¥13.50
Cache hitPeak ¥0.30 · Off-peak ¥0.15

DeepSeek lists official prices in CNY; USD uses ≈ 6.9 exchange rate


Best for

1

o1-level thinking

Thinking Mode produces chain-of-thought reasoning comparable to OpenAI o1-class models.

2

1M token context

Long-context tasks, whole-repo coding, and deep agent workflows fit comfortably.

3

~3.5× Flash output

About 3.5× V4 Flash on output (4.5× on input) at every hour — the extra quality is for the hard 10% of requests.


Best for

Complex reasoning pipelines, whole-repo code generation, and production workloads where structured, verifiable output offsets the ~3.5× price delta over V4 Flash.


Pricing notes

Peak / off-peak billing

Peak (Beijing 09:00–12:00, 14:00–18:00) bills at full price; all other hours are 50% off. Cached input is ¥0.30 peak / ¥0.15 off-peak — 1/30 of the uncached rate.


vs the competition

V4 Pro vs the competition

Against GPT-5.6 Sol ($5/$30) and Claude Opus 5 ($5/$25), V4 Pro at ¥9.00/$27.00 peak competes on output quality for a fraction of the cost — and its ¥4.50 off-peak output undercuts every flagship on the market.


Frequently Asked Questions

What is the price of DeepSeek V4 Pro per million tokens?

DeepSeek V4 Pro bills ¥9.00 per million input tokens and ¥27.00 per million output tokens at peak (Beijing 09:00–12:00, 14:00–18:00); off-peak is 50% off at ¥4.50 / ¥13.50. Cache-hit input is ¥0.30 peak / ¥0.15 off-peak, again about 1/30 of the uncached rate.

When should I use V4 Pro instead of V4 Flash?

V4 Pro targets structured reasoning, complex code generation, and long-context tasks where output quality offsets the 3× price delta. The real question is whether Pro cuts retries, human review, or downstream errors enough to justify the gap for your workload.

Does V4 Pro support thinking mode and does it cost extra?

V4 Pro's Thinking Mode produces chain-of-thought reasoning comparable to o1-level models. Reasoning tokens are billed at the output rate, so factor them into your estimates. Routing batch jobs and cron workloads to Beijing evening/night windows cuts the bill further thanks to off-peak pricing.


Related model pricing

Compare this model against the rest of the lineup.


Estimate your exact cost

Model your own workload — input/output tokens, cache hit rate, and peak-hour share — against any model mix with the interactive pricing calculator.

Open the Pricing Calculator

Recommended Tools We ARE USING

We are using these tools ourselves for Development / Deployment. Check out for more details.