Blog

DeepSeek's Pricing History: From V3 to the Sep 2026 V4 Flash Cut

September 9, 2026Gavin ChenMindRose Team

The September 10 Reprice

On September 10, 2026 at 12:00 Beijing time, DeepSeek cut V4 Flash prices again. Off-peak now bills ¥1.00 per million input tokens (cache miss), ¥4.00 per million output tokens, and ¥0.02 per million cached input tokens; peak is exactly double that (¥2.00 / ¥8.00 / ¥0.04). It is the second pricing change in under a month — and the latest chapter in a pricing story that has swung wildly since DeepSeek first opened its API.

That story is worth understanding, because the numbers that are “cheap” today have been cut, raised, and restructured repeatedly — and each change reshapes which model is the value champion. Here is the full arc, from the 2024 V2 era to today's Flash/Pro line.

The Price History, At a Glance

A condensed timeline of DeepSeek's headline pricing changes. Unless noted, prices are per 1M tokens; DeepSeek's own models are in CNY with a peak/off-peak pair since August 2026.

2024: Context Caching Arrives

In August 2024, DeepSeek launched disk-based context caching, default-on for everyone: $0.014/M cache-hit input versus $0.14/M cache-miss input. The mechanism — caching the computation for prompt prefixes and reusing it on matching requests — is the same idea that still powers today's cache-hit pricing, and it established caching as a first-class cost lever from the very beginning.

That December, V3 set the 2025 baseline: $0.27/M input (miss), $0.07/M (hit), $1.10/M output from February 8, 2025.

2025: The R1 Era and Time-of-Day Discounts

January 2025 brought R1, a dedicated reasoning model billed separately from the general model, at $0.55/M input and $2.19/M output. In February 2025, DeepSeek introduced its first verifiable off-peak discount: R1 cost up to 75% less and V3 50% less during off-peak windows (16:30–00:30 UTC).

That off-peak discount ended on September 5, 2025, when V3.1 unified the general and reasoning models into one model with two modes. On September 29, 2025, the experimental V3.2-Exp used a new sparse-attention architecture to cut prices 50%+ immediately — cache-hit input fell to $0.028/M, cache-miss to $0.28/M, and output to $0.42/M.

2026: V4, Flash/Pro Layering, and Peak/Off-Peak

In April 2026, DeepSeek previewed V4 and replaced the old chat/reasoner split with a Flash/Pro capability tier (both 1M context, both thinking and non-thinking modes). A 75% launch discount on Pro ran through early May, and cache-hit pricing across the API was cut to a tenth.

On August 16, 2026 at 16:00 UTC, DeepSeek introduced formal peak/off-peak pricing: off-peak is exactly half of peak, with peak windows at 09:00–12:00 and 14:00–18:00 Beijing time (Mon–Fri). This was a raise relative to the spring's flat pricing — Reuters reported increases of 50% to over 1,100% depending on the model and token type — even though off-peak remained cheaper than peak.

Now, September 10, 2026 reverses direction for Flash: the new off-peak input (¥1.00), output (¥4.00), and cached input (¥0.02) are each lower than the August off-peak rates, and the cache-hit discount widened from 1/30 to 1/50. Pro's pricing is unchanged.

What Drives the Price: Three Levers

  • Caching — DeepSeek has priced cached input at a fraction of uncached input since 2024. Today that gap is ~50× on Flash. It rewards stable prompt prefixes and heavily reused context.
  • Flash/Pro + thinking effort — the old “two models” became one family with a capability tier and a thinking-effort dial. Flash is the throughput/value layer; Pro is the ceiling for the hard requests.
  • Time-of-day pricing — DeepSeek has repeatedly used peak/off-peak windows (2025 discount, then the 2026 peak/off-peak system) to move deferrable load off the busy windows. Off-peak has always been the cheap lane.

The net signal across two years: DeepSeek is managing per-unit economics with cache pricing, demand shape with time-of-day pricing, and product fit with Flash/Pro layering — rather than one flat number.

How to Read Today's Prices

For V4 Flash today: off-peak ¥1.00 input / ¥4.00 output / ¥0.02 cached input per 1M tokens; peak is double. Compare that to V4 Pro (¥4.50 / ¥13.50 off-peak, ¥9.00 / ¥27.00 peak) and the July-30-cut GPT-5.6 Luna ($0.20 / $1.20). Flash's output is now below Luna's at every hour of the day, and its cached input economics are the cheapest of any model on the market.

The practical playbook has not changed: keep a stable prompt prefix to ride the cache lane, schedule batch and cron work off-peak, and default to Flash non-thinking for high-volume work — upgrading to Pro or higher thinking effort only where output quality pays for itself. See the current DeepSeek V4 Flash pricing for the live numbers.

The Bottom Line

DeepSeek's pricing has moved from a flat V2/V3 model, through R1's separate reasoning SKU and 2025's off-peak discounts, to today's Flash/Pro peak-off-peak system. The September 2026 Flash cut is the latest swing of that pendulum — and a reminder that any single price is a snapshot, not a promise. The strategy that survives these changes is the one that treats caching, scheduling, and model routing as design inputs rather than fixed assumptions.

EventInput (miss)OutputCache hitNotes
Aug 2024 · Disk caching$0.14$0.014Context caching introduced
Feb 2025 · V3$0.27$1.10$0.07V3 baseline
Jan 2025 · R1$0.55$2.19Separate reasoning model
Sep 2025 · V3.2-Exp$0.28$0.42$0.02850%+ cut via sparse attention
Aug 2026 · V4 Flash off-peak¥1.50¥4.50¥0.05Peak/off-peak launched
Sep 2026 · V4 Flash off-peak¥1.00¥4.00¥0.02Sep 10 repricing (peak = 2×)

Recommended Tools We ARE USING

We are using these tools ourselves for Development / Deployment. Check out for more details.