¥9.00 input · ¥27.00 output per 1M tokens at peak — 50% off off-peak. Structured reasoning with o1-level thinking.
DeepSeek V4 Pro is the flagship of the DeepSeek line: structured reasoning, complex code generation, and long-context mastery (1M tokens) with a Thinking Mode that produces chain-of-thought reasoning comparable to o1-level models — at a price point that makes it viable for high-volume production workloads.
DeepSeek V4 Pro · CNY official list price
| Tier | Rate |
|---|---|
| Input | Peak ¥9.00 · Off-peak ¥4.50 |
| Output | Peak ¥27.00 · Off-peak ¥13.50 |
| Cache hit | Peak ¥0.30 · Off-peak ¥0.15 |
DeepSeek lists official prices in CNY; USD uses ≈ 6.9 exchange rate
Thinking Mode produces chain-of-thought reasoning comparable to OpenAI o1-class models.
Long-context tasks, whole-repo coding, and deep agent workflows fit comfortably.
About 3.5× V4 Flash on output (4.5× on input) at every hour — the extra quality is for the hard 10% of requests.
Complex reasoning pipelines, whole-repo code generation, and production workloads where structured, verifiable output offsets the ~3.5× price delta over V4 Flash.
Peak (Beijing 09:00–12:00, 14:00–18:00) bills at full price; all other hours are 50% off. Cached input is ¥0.30 peak / ¥0.15 off-peak — 1/30 of the uncached rate.
Against GPT-5.6 Sol ($5/$30) and Claude Opus 5 ($5/$25), V4 Pro at ¥9.00/$27.00 peak competes on output quality for a fraction of the cost — and its ¥4.50 off-peak output undercuts every flagship on the market.
DeepSeek V4 Pro bills ¥9.00 per million input tokens and ¥27.00 per million output tokens at peak (Beijing 09:00–12:00, 14:00–18:00); off-peak is 50% off at ¥4.50 / ¥13.50. Cache-hit input is ¥0.30 peak / ¥0.15 off-peak, again about 1/30 of the uncached rate.
V4 Pro targets structured reasoning, complex code generation, and long-context tasks where output quality offsets the 3× price delta. The real question is whether Pro cuts retries, human review, or downstream errors enough to justify the gap for your workload.
V4 Pro's Thinking Mode produces chain-of-thought reasoning comparable to o1-level models. Reasoning tokens are billed at the output rate, so factor them into your estimates. Routing batch jobs and cron workloads to Beijing evening/night windows cuts the bill further thanks to off-peak pricing.
Compare this model against the rest of the lineup.
Model your own workload — input/output tokens, cache hit rate, and peak-hour share — against any model mix with the interactive pricing calculator.
Open the Pricing CalculatorWe are using these tools ourselves for Development / Deployment. Check out for more details.
You will receive $5 when subscribed, which directly offsets the 1st month's Go fee.
Could be the cheapest option for calling DeepSeek V4 Flash model in the world.
RReceive $300 to test out VVultr platform
Referred user must be active 30+ days and spend $10–$25
Deploy anything without the complexity
Connect your repo, Railway handles the rest.
Cloud servers, cloud databases, COS, CDN, SMS and other cloud products are on special offer now.
AWS high-quality alternative, we use COS as a replacement for S3
¥16 universal voucher for all platforms, valid for 180 days from the date of receipt
A comprehensive product matrix supporting the full-process implementation of AI applications.
One of the best AI terminal tools in the world
Interestingly, we can directly use various top-tier models in Warp without leaving the terminal.