DeepSeek's Pricing History: From V3 to the Sep 2026 V4 Flash Cut
The September 10 Reprice
On September 10, 2026 at 12:00 Beijing time, DeepSeek cut V4 Flash prices again. Off-peak now bills ¥1.00 per million input tokens (cache miss), ¥4.00 per million output tokens, and ¥0.02 per million cached input tokens; peak is exactly double that (¥2.00 / ¥8.00 / ¥0.04). It is the second pricing change in under a month — and the latest chapter in a pricing story that has swung wildly since DeepSeek first opened its API.
That story is worth understanding, because the numbers that are “cheap” today have been cut, raised, and restructured repeatedly — and each change reshapes which model is the value champion. Here is the full arc, from the 2024 V2 era to today's Flash/Pro line.
The Price History, At a Glance
A condensed timeline of DeepSeek's headline pricing changes. Unless noted, prices are per 1M tokens; DeepSeek's own models are in CNY with a peak/off-peak pair since August 2026.
2024: Context Caching Arrives
In August 2024, DeepSeek launched disk-based context caching, default-on for everyone: $0.014/M cache-hit input versus $0.14/M cache-miss input. The mechanism — caching the computation for prompt prefixes and reusing it on matching requests — is the same idea that still powers today's cache-hit pricing, and it established caching as a first-class cost lever from the very beginning.
That December, V3 set the 2025 baseline: $0.27/M input (miss), $0.07/M (hit), $1.10/M output from February 8, 2025.
2025: The R1 Era and Time-of-Day Discounts
January 2025 brought R1, a dedicated reasoning model billed separately from the general model, at $0.55/M input and $2.19/M output. In February 2025, DeepSeek introduced its first verifiable off-peak discount: R1 cost up to 75% less and V3 50% less during off-peak windows (16:30–00:30 UTC).
That off-peak discount ended on September 5, 2025, when V3.1 unified the general and reasoning models into one model with two modes. On September 29, 2025, the experimental V3.2-Exp used a new sparse-attention architecture to cut prices 50%+ immediately — cache-hit input fell to $0.028/M, cache-miss to $0.28/M, and output to $0.42/M.
2026: V4, Flash/Pro Layering, and Peak/Off-Peak
In April 2026, DeepSeek previewed V4 and replaced the old chat/reasoner split with a Flash/Pro capability tier (both 1M context, both thinking and non-thinking modes). A 75% launch discount on Pro ran through early May, and cache-hit pricing across the API was cut to a tenth.
On August 16, 2026 at 16:00 UTC, DeepSeek introduced formal peak/off-peak pricing: off-peak is exactly half of peak, with peak windows at 09:00–12:00 and 14:00–18:00 Beijing time (Mon–Fri). This was a raise relative to the spring's flat pricing — Reuters reported increases of 50% to over 1,100% depending on the model and token type — even though off-peak remained cheaper than peak.
Now, September 10, 2026 reverses direction for Flash: the new off-peak input (¥1.00), output (¥4.00), and cached input (¥0.02) are each lower than the August off-peak rates, and the cache-hit discount widened from 1/30 to 1/50. Pro's pricing is unchanged.
What Drives the Price: Three Levers
- Caching — DeepSeek has priced cached input at a fraction of uncached input since 2024. Today that gap is ~50× on Flash. It rewards stable prompt prefixes and heavily reused context.
- Flash/Pro + thinking effort — the old “two models” became one family with a capability tier and a thinking-effort dial. Flash is the throughput/value layer; Pro is the ceiling for the hard requests.
- Time-of-day pricing — DeepSeek has repeatedly used peak/off-peak windows (2025 discount, then the 2026 peak/off-peak system) to move deferrable load off the busy windows. Off-peak has always been the cheap lane.
The net signal across two years: DeepSeek is managing per-unit economics with cache pricing, demand shape with time-of-day pricing, and product fit with Flash/Pro layering — rather than one flat number.
How to Read Today's Prices
For V4 Flash today: off-peak ¥1.00 input / ¥4.00 output / ¥0.02 cached input per 1M tokens; peak is double. Compare that to V4 Pro (¥4.50 / ¥13.50 off-peak, ¥9.00 / ¥27.00 peak) and the July-30-cut GPT-5.6 Luna ($0.20 / $1.20). Flash's output is now below Luna's at every hour of the day, and its cached input economics are the cheapest of any model on the market.
The practical playbook has not changed: keep a stable prompt prefix to ride the cache lane, schedule batch and cron work off-peak, and default to Flash non-thinking for high-volume work — upgrading to Pro or higher thinking effort only where output quality pays for itself. See the current DeepSeek V4 Flash pricing for the live numbers.
The Bottom Line
DeepSeek's pricing has moved from a flat V2/V3 model, through R1's separate reasoning SKU and 2025's off-peak discounts, to today's Flash/Pro peak-off-peak system. The September 2026 Flash cut is the latest swing of that pendulum — and a reminder that any single price is a snapshot, not a promise. The strategy that survives these changes is the one that treats caching, scheduling, and model routing as design inputs rather than fixed assumptions.
| Event | Input (miss) | Output | Cache hit | Notes |
|---|---|---|---|---|
| Aug 2024 · Disk caching | $0.14 | — | $0.014 | Context caching introduced |
| Feb 2025 · V3 | $0.27 | $1.10 | $0.07 | V3 baseline |
| Jan 2025 · R1 | $0.55 | $2.19 | — | Separate reasoning model |
| Sep 2025 · V3.2-Exp | $0.28 | $0.42 | $0.028 | 50%+ cut via sparse attention |
| Aug 2026 · V4 Flash off-peak | ¥1.50 | ¥4.50 | ¥0.05 | Peak/off-peak launched |
| Sep 2026 · V4 Flash off-peak | ¥1.00 | ¥4.00 | ¥0.02 | Sep 10 repricing (peak = 2×) |
Recommended Tools We ARE USING
We are using these tools ourselves for Development / Deployment. Check out for more details.
Opencode Go
You will receive $5 when subscribed, which directly offsets the 1st month's Go fee.
Could be the cheapest option for calling DeepSeek V4 Flash model in the world.
Vultr
RReceive $300 to test out VVultr platform
Referred user must be active 30+ days and spend $10–$25
Railway
Deploy anything without the complexity
Connect your repo, Railway handles the rest.
腾讯云 Tencent Cloud
Cloud servers, cloud databases, COS, CDN, SMS and other cloud products are on special offer now.
AWS high-quality alternative, we use COS as a replacement for S3
硅基流动 SiliconFlow
¥16 universal voucher for all platforms, valid for 180 days from the date of receipt
A comprehensive product matrix supporting the full-process implementation of AI applications.
Warp
One of the best AI terminal tools in the world
Interestingly, we can directly use various top-tier models in Warp without leaving the terminal.