DeepSeek API Price Hike: New Rates, Peak Hours, Real Math
DeepSeek’s API gets more expensive today. At 16:00 UTC on August 16 — 9:30 pm IST — the company replaces its flat rate card for the V4 models with peak and off-peak billing, and every line moves up: increases run from about 50% on the cheapest off-peak rows to just over 1,100% on cached input at peak hours. The company that spent two years forcing AI prices down is the one raising them — and even its new discounted off-peak rates sit above the prices it charged yesterday.
What changes at 16:00 UTC
DeepSeek warned on August 6 that prices would rise “by a relatively large margin,” without saying by how much. The numbers arrived a week later, and they apply to both current models — V4-Flash, the high-volume workhorse that absorbed the old deepseek-chat and deepseek-reasoner endpoints in July, and V4-Pro, the heavier tier that hit general availability on August 13, the same week the new rate card landed.
Here is the full old-versus-new table, in US dollars per million tokens, from DeepSeek’s official pricing docs as reported by Quartz, TechNode and Bloomberg; the two cache-hit rows carry TechApple’s breakdown, the one outlet we found publishing them, and are consistent with the official rule that off-peak is half of peak:
| Rate (per 1M tokens) | Old (flat) | New off-peak | New peak | Peak vs old |
|---|---|---|---|---|
| V4-Flash input, cache hit | $0.0028 | $0.007 | $0.014 | 5x |
| V4-Flash input, cache miss | $0.14 | $0.22 | $0.44 | ~3.1x |
| V4-Flash output | $0.28 | $0.66 | $1.32 | ~4.7x |
| V4-Pro input, cache hit | $0.003625 | $0.022 | $0.044 | ~12x |
| V4-Pro input, cache miss | $0.435 | $0.66 | $1.32 | ~3x |
| V4-Pro output | $0.87 | $1.98 | $3.96 | ~4.6x |
The “up to 1,100%” in the headlines is that V4-Pro cache-hit line: $0.003625 rising to $0.044 at peak is a twelvefold jump, which matches the 50%-to-1,100% range Reuters reported. The floor of that range is real too: the gentlest change on the card, V4-Pro’s off-peak cache-miss input, still rises about 52% from $0.435 to $0.66. There is no row on which a DeepSeek customer pays the same tomorrow as today.
Peak hours: when the meter runs hot
The new peak windows are 01:00–04:00 and 06:00–10:00 UTC — 6:30–9:30 am and 11:30 am–3:30 pm in India, 9 am–noon and 2–6 pm in Beijing. Everything outside those seven hours bills at off-peak, which the company’s announcement post defines plainly: “Off-peak rates are 50% lower than peak, enabling more flexible workload scheduling.” The windows track the Chinese working day, which tells you whose congestion DeepSeek is managing.
The company’s stated rationale, as reported by Quartz, is “to allocate resources more reasonably” — pricing as a queue-management tool, nudging batch jobs into the quiet hours. This is not DeepSeek’s first experiment with the idea: it had trialed a peak-hour surcharge on V4-Pro earlier this summer, per the South China Morning Post. What changes today is that the split becomes the default for the whole V4 lineup, on a much higher base.
For developers who can shift work, the damage is smaller than the headline. Sanchit Vir Gogia, chief analyst at Greyhound Research, told InfoWorld that “the schedule’s own clock and cache hand most of it back to any buyer paying attention” — though at face-value peak rates, he noted, DeepSeek’s price advantage “does disappear, and in places inverts.”
Why the cheapest AI just got pricier
DeepSeek’s own explanation never mentions capacity, demand or GPUs. The demand evidence is public anyway. V4-Flash topped OpenRouter’s weekly usage ranking in early August, processing 7.22 trillion tokens through that one router alone, per TechNode — and InfoWorld notes it became the fastest-growing model ever measured on Ollama for local deployment interest. The SCMP’s framing is that the hike lands “amid surge in demand for low-cost AI models”; InfoWorld’s headline calls it demand straining capacity. Selling intelligence below everyone else’s price turned out to work — well enough that the price had to go up.
Bloomberg adds one more frame: the increase comes ahead of a potential public listing, and even after it, DeepSeek’s rates “remain well below” the major Western flagships. Both halves of that sentence matter. This is a company converting an underpriced loss-leader into something that looks like a business — without, on most rows, surrendering the discount that made it famous.
What it means
The remarkable part is the direction of travel. In under three weeks, OpenAI cut GPT-5.6 Luna by 80%, we covered the cut’s math and then DeepSeek’s counterpunch — a model one benchmark point behind Luna at a fraction of its price. Google answered on August 13 with Gemini 3.7 Flash at half price until December 31, and Anthropic made Claude Sonnet 5’s discounted rate permanent on August 10. Every Western lab is cutting. The company that started the race to the bottom is the only one climbing back up.
That produces genuine inversions, visible in the post-hike list prices. All figures are each vendor’s own published rates, US dollars per million tokens:
| Model | Input | Output | Source |
|---|---|---|---|
| DeepSeek V4-Flash (off-peak) | $0.22 | $0.66 | DeepSeek docs |
| DeepSeek V4-Flash (peak) | $0.44 | $1.32 | DeepSeek docs |
| GPT-5.6 Luna | $0.20 | $1.20 | OpenAI |
| DeepSeek V4-Pro (off-peak) | $0.66 | $1.98 | DeepSeek docs |
| DeepSeek V4-Pro (peak) | $1.32 | $3.96 | DeepSeek docs |
| Gemini 3.7 Flash (intro, to Dec 31) | $0.75 | $3.75 | |
| Claude Haiku 4.5 | $1.00 | $5.00 | Anthropic |
| Claude Sonnet 5 | $2.00 | $10.00 | Anthropic |
| Qwen3.8-Max | $2.00 | $6.00 | reported |
Read the Luna row against the two V4-Flash rows. During DeepSeek’s peak hours, OpenAI’s budget tier is now cheaper than DeepSeek on both input and output — a sentence that would have been absurd in July. Off-peak, V4-Flash keeps a roughly 45% output advantage over Luna while conceding a hair on input. At the Pro tier, Google’s introductory Gemini 3.7 Flash rate undercuts V4-Pro’s peak price on both sides of the ledger, though DeepSeek wins both back off-peak. “DeepSeek is cheapest” has stopped being a fact and become a scheduling question. It is the same lesson our AI Models coverage keeps returning to: in this market, price is a feature that ships weekly, in both directions — and speed tiers like OpenAI’s Ultrafast mode are becoming the next axis after price.
What developers can do
The obvious lever is the clock. The two peak windows cover seven hours; the other seventeen bill at half rate. Batch summarization, evals, embedding refreshes and overnight agent runs belong outside 6:30–9:30 am and 11:30 am–3:30 pm IST from today. The second lever is cache discipline: even after the hike, a V4-Flash cache-hit costs $0.007 per million tokens off-peak — repeated system prompts and shared context remain nearly free next to cache-miss rates.
The third lever exists because the V4 models are open-weight under an MIT license: DeepSeek’s own API is not the only place to run them. Third-party hosts price V4-Pro at flat rates with no peak windows — DeepInfra, Fireworks and Novita list $1.74 input and $3.48 output per million tokens. Against DeepSeek’s new peak rate, that is more expensive input but cheaper output, and it is immune to the windows entirely. The open weights that DeepSeek used as a marketing weapon now cap its own pricing power — developers on Hacker News spent the week doing exactly this arithmetic in public.
What to watch
Three things over the next month. Whether DeepSeek keeps its OpenRouter crown once the new rates bite — the weekly rankings will show any migration in near real time. Whether the Western cuts hold: Google’s Gemini 3.7 Flash discount is scheduled to end December 31, which puts a second price rise on the calendar already. And whether peak/off-peak spreads: if the economics of serving open-weight models at scale forces time-of-day pricing on other providers, DeepSeek will have started a second industry norm — this time in the opposite direction from the first.
Frequently asked questions
When does the DeepSeek price increase take effect?
At 16:00 UTC on August 16, 2026 — 9:30 pm in India, and midnight August 17 in Beijing, which is why some coverage dates it a day later. Requests billed after that moment use the new peak or off-peak rate for whichever window they land in.
Why is DeepSeek raising its API prices?
The company's own explanation is about allocating resources more sensibly: it wants developers to shift flexible workloads into off-peak hours, which cost half as much as peak hours. Press coverage adds the context the company left out — usage of the V4 models has surged to the top of third-party rankings, and serving that demand cheaply at all hours had become the constraint.
Is DeepSeek still cheaper than GPT-5.6, Gemini or Claude?
Mostly, but no longer always. Off-peak, DeepSeek's V4 models still undercut every Western rival on output. During its defined peak hours, though, OpenAI's budget GPT-5.6 tier is now cheaper than V4-Flash on both input and output, and Google's introductory Gemini 3.7 Flash rate beats V4-Pro on both. The blanket assumption that DeepSeek is the cheapest option is over; it now depends on when you run.
Can developers avoid the higher rates?
Partly. Scheduling batch and background work outside the two peak windows halves the new rates, and DeepSeek's prompt caching still discounts repeated context heavily. Because the V4 models are open-weight, third-party hosts also serve them at flat rates with no peak windows, which can beat DeepSeek's own peak pricing for some workloads.
Sources & further reading
- API pricing update — @deepseek_ai official announcement post (via xcancel mirror) (primary source)
- Models & Pricing — DeepSeek API Docs (official rate card) (primary source)
- Pricing details (USD) — DeepSeek API Docs (official) (primary source)
- DeepSeek Plans 'Significant' Price Increase for Its AI Services — Bloomberg (Aug 6)
- DeepSeek Increases Prices for AI Services by Multiple Times — Bloomberg (Aug 13)
- DeepSeek raises API pricing for its V4 models — Reuters via Investing.com
- DeepSeek is raising API prices by up to 1,100% — Quartz
- DeepSeek to introduce peak and off-peak pricing for its API — TechNode
- DeepSeek API price increase confirmed: peak and off-peak billing from August 16 — TechApple
- DeepSeek signals 'significant' price hike amid surge in demand for low-cost AI models — SCMP
- After triggering price war, DeepSeek reverses course with surcharge on peak-hour API use — SCMP
- DeepSeek Launches V4-Pro and Raises API Prices by as Much as 1,100% — Caixin Global
- DeepSeek raises some V4 prices by more than 10x as AI demand strains capacity — InfoWorld
- DeepSeek V4 Flash tops OpenRouter weekly ranking with 7.22 trillion tokens — TechNode
- Advancing the price-performance frontier with GPT-5.6 — OpenAI (official) (primary source)
- Introducing Gemini 3.7 Flash — Google (official) (primary source)
- Claude API pricing — Anthropic docs (official) (primary source)
- Sonnet 5 pricing is now permanent — @claudeai official post (primary source)
- DeepSeek V4 Pro pricing guide: providers cost analysis — DeepInfra (host's own pricing)
- DeepSeek up to 1000% price hike is live — Hacker News discussion
More of today, in 60 seconds: Today's Docket →