DeepSeek V4 Flash 0731: Big Benchmark Jump, Same $0.28 Price
DeepSeek released V4 Flash 0731 on July 31, 2026 — the official version of its budget model, retrained rather than redesigned — and held API prices at $0.14 per million input tokens and $0.28 per million output, per its Hugging Face model card and official pricing page. Independent testing by Artificial Analysis scores it 50, a 10-point jump that lands it one point behind OpenAI’s GPT-5.6 Luna — days after OpenAI cut Luna’s prices by 80%. Here is what actually changed, which numbers are independently verified, and what the cheapest near-frontier model yet means for developers, rivals and India.
What DeepSeek shipped on July 31
V4 Flash 0731 is not a new model so much as a much better-trained one. The Hugging Face model card describes it as “the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities” — agentic meaning multi-step work like operating a terminal, editing repositories and calling tools, not just answering questions. Caixin Global reports the architecture is unchanged from the April preview, with the gains coming from additional post-training — the phase where a model is refined on curated tasks after its initial bulk training.
The underlying design carries over from April’s launch: a mixture-of-experts model with 284 billion total parameters, of which only 13 billion activate per token, per DeepSeek’s launch documentation — a structure that keeps serving costs low because most of the network stays idle on any given word. The official pricing page lists a 1-million-token context window, up to 384K output tokens, and a new three-level reasoning_effort control (low, high, max) that trades thinking time for quality.
Two details matter for anyone planning to use it. First, the weights are on Hugging Face under the MIT licence — one of the most permissive open-source licences, allowing free commercial use and modification — with vLLM and SGLang deployment recipes for self-hosting; developer Simon Willison, who tested it at launch, puts the download at roughly 167 GB. Second, the upgrade applied to the API only at launch: DeepSeek said its consumer chat app and website stay on the older models for now, with the flagship V4 Pro’s official release still to come.
Verified benchmarks vs company claims
The number to trust first is the independent one. Artificial Analysis measured V4 Flash 0731 at 50 on its Intelligence Index — up 10 points from the April preview’s 40, six ahead of DeepSeek’s own bigger V4 Pro (44), one behind GPT-5.6 Luna (51), level with Gemini 3.6 Flash (50), and seven behind fellow Chinese lab Moonshot’s Kimi K3 (57), the open-weights leader we profiled in our look at China’s surging AI models. The same analysis places it among the top three open-weights models overall, and its Terminal-Bench 2.1 result rose 17 points to 79% in Artificial Analysis’s own harness. Accuracy on factual recall remains the weak spot: the firm’s Omniscience score, which penalises confident wrong answers, improved from -23 to -16 — better, still negative.
DeepSeek’s own claims go further. The model card’s benchmark table reports Terminal Bench 2.1 at 82.7 (from the preview’s 61.8) and DeepSWE, a hard software-engineering agent test, at 54.4 (from 7.3) — within a few points of Anthropic’s premium Claude Opus 4.8 on both. Treat those with care: Tech Times notes DeepSeek’s benchmark harness is unreleased, so the self-reported table cannot be independently replicated, and earlier V4-Pro self-reported scores had outrun independent measurements. The honest read: the independent 10-point jump is real and large; the “nearly Opus-grade” framing rests on company-run tests.
Here is how the field looks on independent scores and list prices:
| Model | Intelligence Index (AA) | Input $/1M | Output $/1M | Open weights |
|---|---|---|---|---|
| Kimi K3 (Moonshot) | 57 | $3.00 | $15.00 | Yes |
| GPT-5.6 Luna (OpenAI) | 51 | $0.20 | $1.20 | No |
| DeepSeek V4 Flash 0731 | 50 | $0.14 | $0.28 | Yes (MIT) |
| Gemini 3.6 Flash (Google) | 50 | $1.50 | $7.50 | No |
| DeepSeek V4 Pro | 44 | $0.435 | $0.87 | Yes (MIT) |
Scores from Artificial Analysis (Kimi K3 and pricing per its model pages); DeepSeek prices from its official page; Gemini 3.6 Flash pricing from Google’s announcement; Luna’s post-cut rate from Tech Times. Speed is Luna’s clear win — 215 tokens per second against roughly 116 for V4 Flash, per Artificial Analysis’s model pages — and V4 Flash is text-only, with no image input.
Where $0.28 lands in the price war
The timing was the message. OpenAI cut GPT-5.6 Luna by 80% on July 30 — we covered the cut and its India math — and DeepSeek’s answer arrived within a day: a model one point behind Luna at 70% of its input price and less than a quarter of its output price, with a cache-hit input rate of $0.0028 per million — about a 98% discount when repeated prompt content is served from cache.
This is the fourth escalation in a cycle DeepSeek started in April, when it priced the V4 preview about 97% below OpenAI’s then-flagship on cached input, per the South China Morning Post. In May it made a 75% V4 Pro cut permanent, per InfoWorld; by June, SCMP reported Xiaomi slashing its MiMo model prices by up to 99% in response. Axios now frames the whole market as a “race to zero”, noting Anthropic as the clearest holdout on premium pricing — its Claude Opus 4.8 costs about $25 per million output tokens in Axios’s comparison, roughly 90 times V4 Flash’s rate.
Demand context explains DeepSeek’s confidence. Caixin reports the V4 Flash preview was the most-used model on routing platform OpenRouter for seven straight weeks earlier this summer, and that DeepSeek has raised roughly $7.4 billion at a 350-billion-yuan valuation with Tencent and NetEase among backers — funding a plan to double headcount and push deeper into agents. Developer sentiment matches the usage data: the release led Hacker News, where one user reported a month of production usage — 3,467 API requests and 323 million tokens — costing $4.55.
The India angle
DeepSeek bills in dollars with no India-specific pricing, so the rupee math is simple: at the roughly ₹95.4-per-dollar rate on July 31, V4 Flash costs about ₹13.4 per million input tokens and ₹26.7 per million output — cache hits drop input to around ₹0.27. For a small dev shop, a hundred million tokens of monthly output is under ₹2,700 — the kind of budget that made India the largest source of DeepSeek app downloads worldwide in early 2025 at 15.6%, per Business Standard.
The trust questions have not gone away. India’s Finance Ministry told employees in February 2025 to keep ChatGPT and DeepSeek off office devices on data-security grounds, and a public-interest petition alleging DeepSeek transfers Indian users’ data to Chinese servers is before the Delhi High Court, where MeitY said in late 2025 it is drafting AI governance guidelines under the DPDP Act and weighing a licensing framework for foreign AI platforms. India’s data-protection law adds context we’ve mapped before in our DPDP Act explainer: the rules were notified in November 2025, but full enforcement with penalties is not expected until 2027, per India Briefing.
The open weights are the escape hatch. Because V4 Flash is MIT-licensed, Indian firms can run it on Indian servers with no data leaving the country — the route Ola’s Krutrim took in February 2025 when it hosted DeepSeek R1 domestically at an introductory ₹1 per million tokens, per OfficeChai. No Indian cloud has announced V4 Flash hosting yet. The strategic discomfort is sharper: a June Bernstein analysis via Business Today argued India still lacks a globally competitive foundation model — Chinese models processed 21.37 trillion tokens in a single June week against 5.76 trillion for leading US models, while India’s best-funded contender, Sarvam AI, closed a $234-million funding tranche in June, per Medianama. Cheap, capable, self-hostable Chinese models are simultaneously Indian developers’ best bargain and Indian policy’s hardest problem.
What to watch
Three markers over the next month. First, V4 Pro: DeepSeek says the official release of its flagship is coming soon, and if the Flash retrain is a preview of the method, the top of the comparison table could move again. Second, verification: whether independent harnesses reproduce anything close to the self-reported agent scores — the gap between the independent and self-reported Terminal-Bench scores is small today, but DeepSeek’s history of optimistic self-reporting makes every unreplicated claim worth holding loosely. Third, the responses: OpenAI just showed it will cut prices within weeks when pressed, Google’s Gemini 3.6 Flash now looks expensive for a tied score, and Anthropic’s premium position gets harder to hold as the floor falls. We track all of it daily in our AI models coverage.
For developers, the practical takeaway is simpler: near-Luna capability at a quarter of the output price, with a legitimate self-hosting option, moves the default choice for high-volume, text-only workloads — agents, batch processing, code assistance — toward the cheapest tier yet. The caveats are real (slower output, no image input, weak factual-recall scores, an unresolved Indian court case), but the price-performance frontier moved on July 31, and it moved down.
Frequently asked questions
Is DeepSeek V4 Flash free to use?
The weights are free: DeepSeek published V4 Flash 0731 on Hugging Face under the MIT licence, so anyone can download and self-host it. The hosted API is paid and billed per million tokens, at rates far below Western rivals — the exact figures are in the article's comparison table. DeepSeek's own chat app and website were not switched to the 0731 build at launch, per the company.
Is DeepSeek V4 Flash better than GPT-5.6 Luna?
On Artificial Analysis's independent Intelligence Index they are nearly tied — Luna scores 51, V4 Flash 0731 scores 50. Luna generates faster and sits inside OpenAI's ecosystem; V4 Flash is roughly four times cheaper on output tokens, handles a 1-million-token context, and its open weights let you run it on your own hardware.
How much does the DeepSeek API cost in India?
DeepSeek bills in US dollars and has no India-specific pricing. Converted at end-July exchange rates, a million output tokens comes to well under thirty rupees, and cached repeat input falls to under one rupee per million tokens. The article's India section has the full rupee-denominated math and the trust caveats that go with it.
Can I run DeepSeek V4 Flash on my own computer?
Technically yes, practically only with serious hardware. The MIT-licensed weights are a roughly 167 GB download, and DeepSeek publishes vLLM and SGLang deployment recipes. That puts self-hosting in reach of workstations and small clusters with large GPU memory, not typical laptops — although quantised builds and cloud GPUs narrow the gap.
Sources & further reading
- DeepSeek-V4-Flash-0731 model card — Hugging Face (official) (primary source)
- DeepSeek API pricing — official documentation (primary source)
- DeepSeek-V4 launch announcement (April 2026) — official API docs (primary source)
- DeepSeek V4 Flash 0731 scores 50 on the Intelligence Index — Artificial Analysis (primary source)
- Gemini 3.6 Flash announcement and pricing — Google (primary source)
- Kimi K3 model page — Artificial Analysis
- DeepSeek releases official V4-Flash model — Caixin Global
- DeepSeek retrained V4-Flash beats its flagship Pro on nine agent benchmarks — Tech Times
- DeepSeek's new bargain model accelerates AI's race to zero — Axios
- OpenAI cuts Luna 80% after Sol rewrote its own inference stack — Tech Times
- DeepSeek V4 forces rivals to slash prices, rattling China's cloud providers — South China Morning Post
- DeepSeek's steep V4-Pro price cut escalates AI pricing war — InfoWorld
- DeepSeek-V4-Flash-0731 — Simon Willison
- DeepSeek-V4-Flash Update — Hacker News discussion
- DeepSeek tops India app charts amid security concerns — Business Standard
- Ola Krutrim makes DeepSeek available at ₹1 per million tokens — OfficeChai
- Finance Ministry asks employees to avoid ChatGPT, DeepSeek on office devices — Onmanorama
- DeepSeek scrutiny deepens as court seeks Indian government's response — MLex
- India DPDP compliance timeline and enforcement — India Briefing
- Government may take minority stake in Sarvam under IndiaAI Mission — Medianama
- China's AI price war is exposing India's biggest weakness — Business Today
- USD/INR exchange rate — Trading Economics
More of today, in 60 seconds: Today's Docket →