Free DeepSeek API Pricing Calculator — Slide & Estimate Peak vs. Off-Peak Costs
Compare DeepSeek V4 Flash, V4 Pro, and competitor model pricing. Adjust input/output token counts, cache hit rate, and peak-hour share to see real cost projections.
Higher hit rate = lower cost. DeepSeek bills cached input at 1/50 of the uncached input price (peak and off-peak alike).
Peak hours are Beijing time 09:00–12:00 and 14:00–18:00; all other times are off-peak and billed at 50% of the peak price. Nightly batch jobs mostly run off-peak, interactive traffic leans peak.
If you do not already have billing exports, start from product traffic: requests per day, average prompt size, average response length, and expected cache reuse.
Count how many requests your app will send in a normal day, then multiply by 30. Separate peak campaigns from baseline traffic.
Use a few representative prompts to approximate average input and output tokens. For rough planning, consistency matters more than absolute precision.
If your system prompt and context are reused heavily, model a higher hit rate. If every request is unique, start closer to 0% and treat savings as upside.
Requests during Beijing peak windows (09:00–12:00, 14:00–18:00) are billed at full price; everything else is half price. Nightly batch jobs are mostly off-peak, interactive tools lean peak.
These are charged at the model's standard input rate. Large prompts dominate cost when the cache hit rate is low.
When the prefix matches a previous request, DeepSeek charges the reduced cache-hit price instead of the full input rate.
Generated tokens are billed separately and can become the main cost driver in agentic or reasoning-heavy workflows.
DeepSeek charges full price during peak hours (Beijing time 09:00–12:00 and 14:00–18:00) and 50% off at all other times. If your traffic is mostly batch jobs or night cron runs, off-peak pricing cuts cost significantly.
Based on published API pricing as of September 2026 (USD / 1M tokens, standard tier). GPT-5.6 Luna was cut 80% and Terra 20% on July 30, 2026; OpenAI and Anthropic both offer a 50% Batch API discount for offline workloads. DeepSeek V4 Flash was repriced on September 10, 2026 (peak ¥2.00 input / ¥8.00 output / ¥0.04 cached, off-peak halved). DeepSeek prices are per million tokens in CNY (peak / off-peak). Actual costs vary with usage patterns and caching efficiency.
| Model | Input / 1M tokens | Output / 1M tokens | Cache Hit / 1M tokens | Notes |
|---|---|---|---|---|
| DeepSeek V4 Flash | Peak ¥2.00Off-peak ¥1.00 | Peak ¥8.00Off-peak ¥4.00 | Peak ¥0.04Off-peak ¥0.02 | Peak: Beijing 09–12 & 14–18; off-peak 50% off |
| DeepSeek V4 Pro | Peak ¥9.00Off-peak ¥4.50 | Peak ¥27.00Off-peak ¥13.50 | Peak ¥0.30Off-peak ¥0.15 | Peak: Beijing 09–12 & 14–18; off-peak 50% off |
| GPT-5.6 Sol | ¥34.50 | ¥207.00 | ¥3.45 | — |
| GPT-5.6 Terra | ¥13.80 | ¥82.80 | ¥1.38 | — |
| GPT-5.6 Luna | ¥1.38 | ¥8.28 | ¥0.14 | — |
| Claude Opus 5 | ¥34.50 | ¥172.50 | ¥3.45 | — |
| Claude Sonnet 5 | ¥13.80 | ¥69.00 | ¥1.38 | — |
| Claude Haiku 4.5 | ¥6.90 | ¥34.50 | ¥0.69 | — |
If the workload is repetitive and latency-sensitive, DeepSeek V4 Flash usually sets the floor for cost planning.
The real question is not whether Pro is more expensive, but whether it cuts retries, human review, or downstream errors enough to justify the gap.
A model that looks expensive on list price can still make sense for narrow tasks. Re-run the calculator with your own token mix before locking the stack.
Each model has a dedicated pricing page with peak/off-peak rates, cache-hit pricing, best-use guidance, and FAQs.
We are using these tools ourselves for Development / Deployment. Check out for more details.
You will receive $5 when subscribed, which directly offsets the 1st month's Go fee.
Could be the cheapest option for calling DeepSeek V4 Flash model in the world.
RReceive $300 to test out VVultr platform
Referred user must be active 30+ days and spend $10–$25
Deploy anything without the complexity
Connect your repo, Railway handles the rest.
Cloud servers, cloud databases, COS, CDN, SMS and other cloud products are on special offer now.
AWS high-quality alternative, we use COS as a replacement for S3
¥16 universal voucher for all platforms, valid for 180 days from the date of receipt
A comprehensive product matrix supporting the full-process implementation of AI applications.
One of the best AI terminal tools in the world
Interestingly, we can directly use various top-tier models in Warp without leaving the terminal.