Back to Home
API Pricing

DeepSeek V4 Flash API Pricing

¥2.00 input · ¥8.00 output per 1M tokens at peak — 50% off off-peak. The most-used coding model on Earth.

DeepSeek V4 Flash has quietly become the most-used coding model on Earth — on OpenCode Go it holds a 69% token share. It combines a 1M-token context window, a 96% real-world cache-hit rate, and a reasoning score that rivals flagship models, at prices that make it the default workhorse for high-volume AI workloads.

Estimate My CostsPrices as of September 2026

Price per 1M tokens

DeepSeek V4 Flash · CNY official list price

TierRate
InputPeak ¥2.00 · Off-peak ¥1.00
OutputPeak ¥8.00 · Off-peak ¥4.00
Cache hitPeak ¥0.04 · Off-peak ¥0.02

DeepSeek lists official prices in CNY; USD uses ≈ 6.9 exchange rate


Best for

1

96% cache hit rate

Real-world cache-hit rates on OpenCode Go sit at 96%, so most input rides the cached lane at ¥0.04/1M peak.

2

1M token context

A 1M-token context window lets whole codebases and long conversations fit in a single request.

3

100/100 reasoning

Normalized reasoning score of 100/100 — it reasons like a flagship model at a fraction of the price.


Best for

High-volume coding agents, batch summarization, RAG pipelines, and any repetitive workflow where cost per token matters more than absolute ceiling. Still ~4× cheaper than V4 Pro on input (and ~3.4× on output) at every hour, it's the best everyday agent model on the market.


Pricing notes

Peak / off-peak billing

Peak hours are Beijing time 09:00–12:00 and 14:00–18:00 and bill at full price; everything else is 50% off. Cached input is billed at 1/50 of the uncached rate (¥0.04 vs ¥2.00 at peak).


vs the competition

V4 Flash vs the competition

At off-peak rates, V4 Flash output costs ¥4.00 ($0.58) per million tokens — lower than GPT-5.6 Luna's $1.20 and a fraction of V4 Pro's ¥8.00 peak. It wins on cached input economics and context length against both OpenAI and Anthropic.


Frequently Asked Questions

What is the price of DeepSeek V4 Flash per million tokens?

DeepSeek V4 Flash bills ¥2.00 per million input tokens and ¥8.00 per million output tokens at peak hours (Beijing 09:00–12:00 and 14:00–18:00). Off-peak hours are 50% off: ¥1.00 input / ¥4.00 output. Cached input is charged at ¥0.04 peak / ¥0.02 off-peak — about 1/50 of the uncached rate.

Is DeepSeek V4 Flash actually cheap enough for production workloads?

Yes. With a 1M-token context window, a 96% real-world cache-hit rate, and reasoning scores that rival flagship models, V4 Flash has become the most-used coding model on Earth (69% of OpenCode Go's token share). For high-volume routine work it sets the floor for cost planning.

How do DeepSeek peak and off-peak prices work for V4 Flash?

DeepSeek splits the day in half: peak hours are Beijing time 09:00–12:00 and 14:00–18:00 and bill at full price; everything else is billed at 50% off. A nightly batch job generating 100M output tokens costs ¥800 at peak but only ¥400 after 8 PM — same tokens, zero code changes.


Related model pricing

Compare this model against the rest of the lineup.


Estimate your exact cost

Model your own workload — input/output tokens, cache hit rate, and peak-hour share — against any model mix with the interactive pricing calculator.

Open the Pricing Calculator

Recommended Tools We ARE USING

We are using these tools ourselves for Development / Deployment. Check out for more details.