The Tech Docket Daily signal on tech & AI — what the world is searching, and why
AI Models

GPT-5.6 Luna Price Cut: OpenAI Slashes API Rates by 80%

OpenAI cut GPT-5.6 Luna API prices 80% to $0.20 per million input tokens and Terra 20% on July 30. New rates vs Gemini, Claude, DeepSeek — and the India cost math.
Falling price tags beside stacks of glowing tokens on a dark pricing dashboard, symbolising cheaper AI model rates.

OpenAI cut its GPT-5.6 API prices on July 30, 2026 — just three weeks after the models launched, per its official announcement. GPT-5.6 Luna, the high-volume tier, dropped 80%, from $1.00 to $0.20 per million input tokens and from $6.00 to $1.20 per million output tokens. GPT-5.6 Terra fell 20%, to $2.00 and $12.00. The flagship, GPT-5.6 Sol, keeps its price. The cut applies to the API, ChatGPT Work and Codex; consumer subscription prices are unchanged. Here is what changed, how the new rates stack up against Google, Anthropic and DeepSeek, and what the cheaper tokens mean for developers in India.

What changed on July 30

The new rate card comes straight from OpenAI’s announcement post, with the original July 9 launch prices on OpenAI’s GPT-5.6 introduction page for comparison. All figures are US dollars per million tokens.

Model Old input New input Old output New output Change
GPT-5.6 Luna $1.00 $0.20 $6.00 $1.20 −80%
GPT-5.6 Terra $2.50 $2.00 $15.00 $12.00 −20%
GPT-5.6 Sol $5.00 $5.00 $30.00 $30.00 unchanged

Sol did get one change: a new Fast mode in the API that runs up to 2.5 times faster at double the standard price, replacing the old Priority Processing option, per InfoWorld’s report on the announcement. All three models keep the family’s roughly 1-million-token context window, 128,000-token maximum output and February 2026 knowledge cutoff, as listed in OpenAI’s model documentation.

The scope matters as much as the numbers. The lower unit costs flow through to the API, ChatGPT Work and Codex, and because Luna and Terra now consume less quota per task, ChatGPT Work and Codex users effectively get more usage from the same subscription, 9to5Mac explains. Sticker prices for Free, Go, Plus and Pro plans are untouched. One note for anyone checking OpenAI’s own site: as of August 1 the Terra developer-docs page still showed the old rate — the announcement post carries the authoritative number.

Why OpenAI is cutting, in its own words and everyone else’s

OpenAI frames the cut purely as passed-through efficiency: “Better routing keeps hardware productive, optimized production software generates tokens more efficiently, and smarter context management helps agents avoid repeating completed work,” the announcement says. The company puts numbers on it — an end-to-end serving-cost reduction of about 20% and a token-generation efficiency gain of more than 15% — as reported by Axios. No competitor is named anywhere in the official post.

The press supplies the competitive read: cutting prices three weeks after launch is unusually fast, Axios notes. A Yahoo Finance analysis argues the pattern — slash the mid and budget tiers, leave the flagship untouched — is a defence of OpenAI’s volume business where cheap rivals bite hardest, while flagship margins are preserved; it also relays a CNBC finding that Chinese models had captured 46% of US enterprise token usage on the OpenRouter marketplace, per that analysis. The pressure has a history: DeepSeek’s R1 launched in January 2025 at roughly 90–95% below comparable OpenAI and Anthropic rates, Silicon Canals reported at the time, Google answered within weeks with a budget Gemini 2.0 Flash priced below DeepSeek, per Cybernews, and DeepSeek priced its V4 beta about 97% below OpenAI’s then-current GPT-5.5 this spring, the South China Morning Post reported. We traced how those Chinese models translated cheap tokens into real usage share in our look at the Kimi and Qwen surge.

The buyer side is pushing in the same direction. Forbes ties the cut to enterprise AI bills spiralling past budgets — Uber reportedly exhausted its entire 2026 AI budget in the first quarter — and cites a Harness survey of about 700 engineering and finance leaders in which 29% of organisations attribute more than a quarter of total cloud spend to AI, in its report on the cut. Zoom out and the direction of travel is stark: frontier token prices have fallen from around $60 per million in 2021 to a few cents at the budget end, per EpochAI data cited by Forbes — the economics we unpacked in our explainer on what running AI actually costs.

Early developer verdicts are enthusiastic — with the caveat that the most-quoted surfaced alongside OpenAI’s announcement push. Replit’s Michele Catasta called Luna “the closest we’ve come to intelligence too cheap to meter,” and Dust co-founder Stanislas Polu said Luna runs about 40% faster and 40% cheaper on agentic tasks, both quoted by TechTimes. Analyst Pareekh Jain’s take is the sober version: “Lower prices make it easier to move pilots into production, expand AI across more employees and business processes,” he told InfoWorld.

How Luna and Terra now compare with rivals

The comparison below uses each vendor’s own current pricing page — Google’s Gemini API rates, Anthropic’s Claude rates and DeepSeek’s rate card — against OpenAI’s new prices. List prices, US dollars per million tokens, as of August 1, 2026.

Provider Model Input $/1M Output $/1M
OpenAI GPT-5.6 Luna $0.20 $1.20
OpenAI GPT-5.6 Terra $2.00 $12.00
Google Gemini 3.5 Flash-Lite $0.30 $2.50
Google Gemini 3.6 Flash $1.50 $7.50
Anthropic Claude Haiku 4.5 $1.00 $5.00
Anthropic Claude Sonnet 5 (intro rate) $2.00 $10.00
DeepSeek V4-Flash $0.14 $0.28
DeepSeek V4-Pro $0.435 $0.87

Three things stand out. First, Luna’s input price now undercuts Google’s budget tier — $0.20 against Flash-Lite’s $0.30 — a reversal of the usual order in which Google owns the cheap end; where Gemini’s speed tier fits in Google’s line-up is something we covered when 3.6 Flash launched. Second, DeepSeek still holds the floor: cheaper than Luna on output by a wide margin ($0.87 versus $1.20 on V4-Pro, with V4-Flash cheaper on every line), plus cache-hit input rates below a cent, per DeepSeek’s pricing page. Third, the mid-tier is converging: Terra now sits almost exactly on Anthropic’s introductory Sonnet 5 rate of $2.00/$10.00, which lasts only through August 31, per its pricing documentation. OpenAI, for its part, claims Luna “outperforms Opus 4.8 at roughly one-quarter the cost” — a vendor benchmark claim worth independent testing, made on its launch page.

One caveat: list prices are not effective prices — caching discounts, batch tiers and per-workload token efficiency move real bills, so the cheapest list price is not automatically the cheapest system for a given job.

The India angle

For Indian developers billed in dollars, the rupee arithmetic is the story. At about ₹95.4 to the dollar, the prevailing rate in late July 2026, a million input tokens on Luna falls from roughly ₹95 to about ₹19, and a million output tokens from about ₹572 to ₹114. Terra drops to roughly ₹191 in and ₹1,145 out. For a startup pushing billions of tokens a month through classification or support-bot workloads, that is the difference between a lakh-scale and a crore-scale annual API bill.

The cut also fits a pattern: India is where OpenAI has been most willing to compete on price. ChatGPT Go launched in India first at ₹399 a month with UPI support before expanding to 170 more countries, OpenAI’s own post records — a plan TechCrunch called OpenAI’s cheapest anywhere at launch. The commercial logic is scale: Sam Altman put India at 100 million weekly ChatGPT users this February — OpenAI’s second-largest market — as The Hans India reported, and it has opened a Delhi office and named former Uber India head Prabhjeet Singh its first India managing director, per TechCrunch. Not every rival localises this way — an industry comparison found Anthropic’s India prices are straight currency conversions without UPI support, while OpenAI and Cursor built India-specific plans, The New Stack noted.

The counter-current is sovereignty. The day after OpenAI’s cut, IBM and Sarvam announced a partnership to put Sarvam’s India-built models and voice stack on IBM’s governance layer for government and regulated-sector deployments, per IBM’s announcement, and Sarvam has said it plans a trillion-parameter domestically trained model while pricing below global rivals, Digit reports. Cheaper OpenAI tokens raise the bar for that pitch on cost alone — which is precisely why the sovereign case is shifting to control, data residency and Indian-language depth rather than price. Both trends can be true at once: Indian startups get cheaper frontier tokens, and the state builds its own stack anyway.

What to watch

Whether the flagship joins next. Sol kept its $5.00/$30.00 rate while gaining a paid Fast mode, per OpenAI’s announcement — if the efficiency claims keep compounding, a Sol cut is the natural sequel; if not, the two-speed pricing shows where OpenAI thinks its margin lives.

Anthropic’s September 1 decision. Sonnet 5’s introductory $2.00/$10.00 rate is scheduled to become $3.00/$15.00 from September, per Anthropic’s pricing page. Raising prices into a market OpenAI just cut would signal confidence; extending the intro rate would be an answer to Luna.

DeepSeek’s response. Every round of this price war has drawn a counter-cut within weeks, as the spring V4 pricing showed.

Whether cheaper tokens reach consumers. Nothing in this round changes Free, Go, Plus or Pro pricing, 9to5Mac confirms — but unit costs falling 80% at the volume tier is exactly the kind of change that eventually shows up in quotas, free-tier limits or another India-first pricing experiment. Follow our AI models coverage as the next moves land.

Frequently asked questions

Does the GPT-5.6 price cut change ChatGPT Plus or ChatGPT Go prices?

No. The July 30 cut applies to API usage, ChatGPT Work and Codex — the unit cost of tokens, not subscription stickers. Free, Go, Plus and Pro prices are unchanged. The indirect benefit is that Luna and Terra now consume less of a ChatGPT Work quota per task, so the same subscription stretches further.

GPT-5.6 Sol vs Terra vs Luna — which one should you use?

OpenAI positions Sol as the flagship for the hardest reasoning work, Terra as the balanced everyday model, and Luna as the fast, cheap tier built for high-volume and agentic workloads. After the cut, the practical default for most automated pipelines is Luna, stepping up to Terra or Sol only where quality measurably improves results.

Is GPT-5.6 Luna now cheaper than Gemini Flash and DeepSeek?

On input pricing Luna now undercuts Google's budget Gemini Flash-Lite tier — a first for OpenAI at the cheap end. DeepSeek remains cheaper overall, with lower output and cached-input rates across its V4 models, so on list price OpenAI has closed most of the gap rather than won outright.

Why did OpenAI cut GPT-5.6 prices just three weeks after launch?

OpenAI's official reason is efficiency: better routing, optimised inference software and smarter context management cut its serving costs. Press coverage adds the competitive read — cheap Chinese open-weight models and cost-conscious enterprise buyers were eroding OpenAI's mid-tier volume, which is exactly where this cut lands.

Sources & further reading

  1. Advancing the price-performance frontier with GPT-5.6 — OpenAI announcement (primary source)
  2. Introducing GPT-5.6 — OpenAI launch post (original pricing) (primary source)
  3. GPT-5.6 Sol — OpenAI API model documentation (primary source)
  4. Gemini API pricing — Google AI for Developers (primary source)
  5. Claude API pricing — Anthropic documentation (primary source)
  6. DeepSeek API pricing — DeepSeek documentation (primary source)
  7. IBM and Sarvam collaborate to advance AI sovereignty — IBM India newsroom (primary source)
  8. Introducing ChatGPT Go — OpenAI (primary source)
  9. OpenAI drops GPT-5.6 Luna and Terra API prices by up to 80% — InfoWorld
  10. OpenAI cuts GPT-5.6 Luna API prices by 80%, Terra 20% — eWeek
  11. OpenAI discounts GPT-5.6 Luna and Terra — Axios
  12. OpenAI makes two GPT-5.6 models cheaper, expanding usage in ChatGPT — 9to5Mac
  13. OpenAI just cut GPT-5.6 Luna's price by 80 percent — Yahoo Finance
  14. OpenAI cuts GPT-5.6 pricing up to 80% as AI costs come under scrutiny — Forbes
  15. As token costs plunge, enterprise AI providers face a new margin squeeze — Forbes
  16. OpenAI cuts Luna 80% after Sol rewrote its own inference stack — TechTimes
  17. India emerges as OpenAI's second-largest market with 100 million weekly ChatGPT users — The Hans India
  18. OpenAI launches a sub-$5 ChatGPT plan in India — TechCrunch
  19. OpenAI poaches Uber India chief to lead its biggest market outside the US — TechCrunch
  20. OpenAI makes GPT-5.6 dramatically cheaper, putting pressure on rivals — Digit
  21. How Cursor, Anthropic and OpenAI price for India — The New Stack
  22. USD to INR exchange rate — X-Rates
  23. China's DeepSeek prices new V4 AI model 97% below OpenAI's GPT-5.5 — South China Morning Post
  24. Google releases Gemini with lower price tag than DeepSeek — Cybernews
How this article was made: topic selected from same-day search-trend and community-momentum data across India and the US; researched, drafted and fact-checked with AI assistance under the site's automated quality gates (source citations, originality, no-clickbait and accuracy checks), on the editorial standards set by Saurab Jain. Details in our editorial policy. Spotted an error? Email a correction.

More of today, in 60 seconds: Today's Docket →