Back to Home
Free Tool

Estimate Your DeepSeek API Costs Instantly

Free DeepSeek API Pricing Calculator — Slide & Estimate Peak vs. Off-Peak Costs

Compare DeepSeek V4 Flash, V4 Pro, and competitor model pricing. Adjust input/output token counts, cache hit rate, and peak-hour share to see real cost projections.


How to Estimate Monthly Usage

1M50000M
100K5000M
%
0%100%

Higher hit rate = lower cost. DeepSeek bills cached input at 1/50 of the uncached input price (peak and off-peak alike).

%
0%100%

Peak hours are Beijing time 09:00–12:00 and 14:00–18:00; all other times are off-peak and billed at 50% of the peak price. Nightly batch jobs mostly run off-peak, interactive traffic leans peak.

DeepSeek V4 Flash¥14.11
DeepSeek V4 Pro¥57.54
GPT-5.6 Sol¥427.80
Flash ×4.4Pro ×1.1
GPT-5.6 Terra¥171.12
Flash ×1.8Pro ×0.4
GPT-5.6 Luna¥17.11
Flash ×0.2Pro ×0.0
Claude Opus 5¥393.30
Flash ×4.0Pro ×1.0
Claude Sonnet 5¥157.32
Flash ×1.6Pro ×0.4
Claude Haiku 4.5¥78.66
Flash ×0.8Pro ×0.2

How to Estimate Monthly Usage

If you do not already have billing exports, start from product traffic: requests per day, average prompt size, average response length, and expected cache reuse.

1

1. Estimate request volume

Count how many requests your app will send in a normal day, then multiply by 30. Separate peak campaigns from baseline traffic.

2

2. Estimate prompt and response size

Use a few representative prompts to approximate average input and output tokens. For rough planning, consistency matters more than absolute precision.

3

3. Add a cache assumption

If your system prompt and context are reused heavily, model a higher hit rate. If every request is unique, start closer to 0% and treat savings as upside.

4

4. Estimate your peak-hour share

Requests during Beijing peak windows (09:00–12:00, 14:00–18:00) are billed at full price; everything else is half price. Nightly batch jobs are mostly off-peak, interactive tools lean peak.


How the Billing Model Works

Uncached input tokens

These are charged at the model's standard input rate. Large prompts dominate cost when the cache hit rate is low.

Cached input tokens

When the prefix matches a previous request, DeepSeek charges the reduced cache-hit price instead of the full input rate.

Output tokens

Generated tokens are billed separately and can become the main cost driver in agentic or reasoning-heavy workflows.

Peak vs. off-peak rates

DeepSeek charges full price during peak hours (Beijing time 09:00–12:00 and 14:00–18:00) and 50% off at all other times. If your traffic is mostly batch jobs or night cron runs, off-peak pricing cuts cost significantly.


Based on published API pricing as of September 2026 (USD / 1M tokens, standard tier). GPT-5.6 Luna was cut 80% and Terra 20% on July 30, 2026; OpenAI and Anthropic both offer a 50% Batch API discount for offline workloads. DeepSeek V4 Flash was repriced on September 10, 2026 (peak ¥2.00 input / ¥8.00 output / ¥0.04 cached, off-peak halved). DeepSeek prices are per million tokens in CNY (peak / off-peak). Actual costs vary with usage patterns and caching efficiency.

ModelInput / 1M tokensOutput / 1M tokensCache Hit / 1M tokensNotes
DeepSeek V4 FlashPeak ¥2.00Off-peak ¥1.00Peak ¥8.00Off-peak ¥4.00Peak ¥0.04Off-peak ¥0.02Peak: Beijing 09–12 & 14–18; off-peak 50% off
DeepSeek V4 ProPeak ¥9.00Off-peak ¥4.50Peak ¥27.00Off-peak ¥13.50Peak ¥0.30Off-peak ¥0.15Peak: Beijing 09–12 & 14–18; off-peak 50% off
GPT-5.6 Sol¥34.50¥207.00¥3.45
GPT-5.6 Terra¥13.80¥82.80¥1.38
GPT-5.6 Luna¥1.38¥8.28¥0.14
Claude Opus 5¥34.50¥172.50¥3.45
Claude Sonnet 5¥13.80¥69.00¥1.38
Claude Haiku 4.5¥6.90¥34.50¥0.69

How to Read the Result

Use Flash for high-volume routine work

If the workload is repetitive and latency-sensitive, DeepSeek V4 Flash usually sets the floor for cost planning.

Use Pro when quality offsets the delta

The real question is not whether Pro is more expensive, but whether it cuts retries, human review, or downstream errors enough to justify the gap.

Compare competitors with your actual mix

A model that looks expensive on list price can still make sense for narrow tasks. Re-run the calculator with your own token mix before locking the stack.


Full pricing pages per model

Each model has a dedicated pricing page with peak/off-peak rates, cache-hit pricing, best-use guidance, and FAQs.


Recommended Tools We ARE USING

We are using these tools ourselves for Development / Deployment. Check out for more details.