$0.20 input · $1.20 output per 1M tokens — 'good enough' at a fraction of the price.
GPT-5.6 Luna charges $0.20 input and $1.20 output per million tokens — one-eleventh of Sol's output price — yet OpenAI's own benchmarks show it nearly matching GPT-5.5's peak performance at less than half the cost, and outperforming Claude Opus 4.8 on coding. The July 30 price cut (80% off) turned it from 'cheap' into 'almost free'.
GPT-5.6 Luna · USD official list price
| Tier | Rate |
|---|---|
| Input | $0.20 |
| Output | $1.20 |
| Cache hit | $0.02 |
DeepSeek lists official prices in CNY; USD uses ≈ 6.9 exchange rate
Ultra-cheap input and output per million tokens.
The July 30, 2026 cut turned Luna into the budget champion.
Cached input is effectively free at two cents per million tokens.
High-volume traffic with mostly uncached input: default interactive workhorses, bulk summarization, and budget-sensitive startups. Combined with DeepSeek V4 Flash off-peak, it forms the embarrassingly cheap 2026 stack.
Published API pricing as of September 2026, standard tier, USD per million tokens. Also qualifies for the 50% Batch API discount — Luna at $0.60 output per million tokens is nearly free for offline jobs.
Luna's $1.20 output now loses to DeepSeek V4 Flash at every hour — Flash output is ¥8.00 peak ($1.16) and ¥4.00 off-peak ($0.58), both below $1.20, and its cached input economics are far cheaper. Luna only wins if you must stay entirely on OpenAI infrastructure.
GPT-5.6 Luna bills $0.20 per million input tokens and $1.20 per million output tokens, with cached input at just $0.02 per million. The July 30, 2026 price cut (80% off) turned it from 'cheap' into 'almost free'.
OpenAI's own benchmarks show Luna nearly matching GPT-5.5's peak performance at less than half the cost, and outperforming Claude Opus 4.8 on coding. It's the definition of 'good enough, at a fraction of the price'.
Luna's edge assumes mostly uncached input — its $0.20 input rate already counts as cheap. If your workload is output-heavy agent loops, a model with cheaper effective output per token may win; measure your real mix before locking anything in.
Compare this model against the rest of the lineup.
Model your own workload — input/output tokens, cache hit rate, and peak-hour share — against any model mix with the interactive pricing calculator.
Open the Pricing CalculatorWe are using these tools ourselves for Development / Deployment. Check out for more details.
You will receive $5 when subscribed, which directly offsets the 1st month's Go fee.
Could be the cheapest option for calling DeepSeek V4 Flash model in the world.
RReceive $300 to test out VVultr platform
Referred user must be active 30+ days and spend $10–$25
Deploy anything without the complexity
Connect your repo, Railway handles the rest.
Cloud servers, cloud databases, COS, CDN, SMS and other cloud products are on special offer now.
AWS high-quality alternative, we use COS as a replacement for S3
¥16 universal voucher for all platforms, valid for 180 days from the date of receipt
A comprehensive product matrix supporting the full-process implementation of AI applications.
One of the best AI terminal tools in the world
Interestingly, we can directly use various top-tier models in Warp without leaving the terminal.