¥2.00 input · ¥8.00 output per 1M tokens at peak — 50% off off-peak. The most-used coding model on Earth.
DeepSeek V4 Flash has quietly become the most-used coding model on Earth — on OpenCode Go it holds a 69% token share. It combines a 1M-token context window, a 96% real-world cache-hit rate, and a reasoning score that rivals flagship models, at prices that make it the default workhorse for high-volume AI workloads.
DeepSeek V4 Flash · CNY official list price
| Tier | Rate |
|---|---|
| Input | Peak ¥2.00 · Off-peak ¥1.00 |
| Output | Peak ¥8.00 · Off-peak ¥4.00 |
| Cache hit | Peak ¥0.04 · Off-peak ¥0.02 |
DeepSeek lists official prices in CNY; USD uses ≈ 6.9 exchange rate
Real-world cache-hit rates on OpenCode Go sit at 96%, so most input rides the cached lane at ¥0.04/1M peak.
A 1M-token context window lets whole codebases and long conversations fit in a single request.
Normalized reasoning score of 100/100 — it reasons like a flagship model at a fraction of the price.
High-volume coding agents, batch summarization, RAG pipelines, and any repetitive workflow where cost per token matters more than absolute ceiling. Still ~4× cheaper than V4 Pro on input (and ~3.4× on output) at every hour, it's the best everyday agent model on the market.
Peak hours are Beijing time 09:00–12:00 and 14:00–18:00 and bill at full price; everything else is 50% off. Cached input is billed at 1/50 of the uncached rate (¥0.04 vs ¥2.00 at peak).
At off-peak rates, V4 Flash output costs ¥4.00 ($0.58) per million tokens — lower than GPT-5.6 Luna's $1.20 and a fraction of V4 Pro's ¥8.00 peak. It wins on cached input economics and context length against both OpenAI and Anthropic.
DeepSeek V4 Flash bills ¥2.00 per million input tokens and ¥8.00 per million output tokens at peak hours (Beijing 09:00–12:00 and 14:00–18:00). Off-peak hours are 50% off: ¥1.00 input / ¥4.00 output. Cached input is charged at ¥0.04 peak / ¥0.02 off-peak — about 1/50 of the uncached rate.
Yes. With a 1M-token context window, a 96% real-world cache-hit rate, and reasoning scores that rival flagship models, V4 Flash has become the most-used coding model on Earth (69% of OpenCode Go's token share). For high-volume routine work it sets the floor for cost planning.
DeepSeek splits the day in half: peak hours are Beijing time 09:00–12:00 and 14:00–18:00 and bill at full price; everything else is billed at 50% off. A nightly batch job generating 100M output tokens costs ¥800 at peak but only ¥400 after 8 PM — same tokens, zero code changes.
Compare this model against the rest of the lineup.
Model your own workload — input/output tokens, cache hit rate, and peak-hour share — against any model mix with the interactive pricing calculator.
Open the Pricing CalculatorWe are using these tools ourselves for Development / Deployment. Check out for more details.
You will receive $5 when subscribed, which directly offsets the 1st month's Go fee.
Could be the cheapest option for calling DeepSeek V4 Flash model in the world.
RReceive $300 to test out VVultr platform
Referred user must be active 30+ days and spend $10–$25
Deploy anything without the complexity
Connect your repo, Railway handles the rest.
Cloud servers, cloud databases, COS, CDN, SMS and other cloud products are on special offer now.
AWS high-quality alternative, we use COS as a replacement for S3
¥16 universal voucher for all platforms, valid for 180 days from the date of receipt
A comprehensive product matrix supporting the full-process implementation of AI applications.
One of the best AI terminal tools in the world
Interestingly, we can directly use various top-tier models in Warp without leaving the terminal.