GPT-5.6 Sol Ultrafast: OpenAI's 14x Speed Mode on Cerebras
OpenAI has switched on Ultrafast mode, a new service tier that serves GPT-5.6 Sol — its flagship model — at up to 750 output tokens per second, which the company says is as much as 14 times faster than its Standard processing. The speed comes from running the unchanged model on Cerebras’ dinner-plate-sized wafer chips instead of GPUs. It launched on August 13 in the OpenAI API as a limited preview: waitlist only, no published price, and a central claim — same intelligence, just faster — that no independent tester has yet measured.
What OpenAI and Cerebras announced
OpenAI’s announcement, titled “Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed”, and Cerebras’ companion post describe the same product from two sides: GPT-5.6 Sol, the top model in the family OpenAI shipped in July, served from Cerebras wafer-scale hardware at up to 750 output tokens per second.
Access is deliberately narrow for now. The preview is open to a select group of customers, with a waitlist that asks businesses to describe their workload, latency needs and expected usage — Decrypt bluntly called it “invite-only”. OpenAI says early testers span coding, commerce, financial research and customer support, and that access will expand “as capacity grows.” Sachin Katti, who runs compute strategy at OpenAI, framed the preview as an experiment: the company wants to “learn where that speed creates meaningful value” before deciding how to expand the service. Cerebras CEO Andrew Feldman was less restrained: “GPT-5.6 Sol on Ultrafast is proof that speed and intelligence are no longer mutually exclusive.”
Here is where the new tier sits in the GPT-5.6 Sol line-up:
| GPT-5.6 Sol tier | Claimed speed | Price (per 1M tokens, in/out) |
|---|---|---|
| Standard | ~53 tokens/sec baseline | $5 / $30 |
| Fast | up to ~2.5x Standard | ~2x the Standard rate |
| Ultrafast (new) | up to 750 tokens/sec, “up to 14x” | not yet published |
The Standard price is listed at $5 and $30 per million tokens; the Fast tier’s speed and premium come from third-party pricing guides, and the Ultrafast row is OpenAI’s and Cerebras’ own claim. The missing price is the detail to watch — nothing about this launch says the speed will come cheap.
The benchmark claims, and what they actually measure
Two headline numbers travel with this launch, and they are not the same kind of number. The 14x figure is raw generation speed: 750 tokens a second against a Standard baseline of roughly 53. The 5.6x figure is different — on GDPval, OpenAI’s benchmark of economically valuable knowledge work built from 1,320 real deliverables across 44 occupations, the companies say Ultrafast produced “a 5.6x end-to-end speedup with no loss in quality” on tasks such as legal briefs, financial models and engineering reports. End-to-end time includes reading documents and calling tools, not just writing output, so real workflows gain less than the raw multiplier suggests. Treat the two figures as answers to different questions.
The showier claims deserve more caution. Cerebras says the accelerated model worked through Humanity’s Last Exam — a 2,500-question test spanning graduate-level chemistry, economics and literature — in just over 11 hours, against more than three days for a rival frontier model, and pitches Ultrafast as several times faster than Anthropic’s comparable tiers. Those comparisons come from the vendors’ own materials, with no published methodology, and Anthropic has said nothing. Independent speed-measurement outfits such as Artificial Analysis — whose numbers have run well below vendor claims before — have not yet published Ultrafast data. Until they do, “same intelligence, 14x faster” is a promise, not a result.
How a dinner-plate chip gets to 750 tokens a second
The physics of the speedup is the most interesting part of the story, and it is not new physics. A GPU keeps a model’s weights in memory chips that sit next to the processor, and generating text means re-reading those weights for every token. That round trip — compute to memory and back, billions of times — is the bottleneck that caps how fast even the biggest GPU clusters can generate.
Cerebras builds around the bottleneck by refusing to cut the wafer. Its Wafer-Scale Engine is a single chip about 57 times the area of a flagship GPU die, carrying 4 trillion transistors, 900,000 cores and 44GB of memory on the silicon itself. The weights live on the chip, next to the compute, so the round trip largely disappears. The company has been setting speed records with this design for two years: 969 tokens a second on Llama 3.1 405B in late 2024, then 3,000 tokens a second on OpenAI’s open-weight gpt-oss-120B last year.
One correction to a framing you will see elsewhere: this is not the first OpenAI model on non-Nvidia silicon. That happened in February, when GPT-5.3-Codex-Spark launched on Cerebras chips at over 1,000 tokens a second — but Spark was a slimmed-down coding model built for speed. What changed on August 13 is the class of model getting this speed: Sol is OpenAI’s full flagship, the model the company points at its hardest reasoning work, not a distilled variant.
What it means
For OpenAI, this is a January contract becoming a product. OpenAI signed a multi-year deal reported at over $10 billion for up to 750 megawatts of Cerebras compute through 2028, with an option on more through 2030 — part of a deliberate spread across AMD, Broadcom and its own custom silicon to loosen dependence on Nvidia. Anthropic made a similar move toward AMD earlier this month. Ultrafast is the first consumer-visible answer to what all that non-Nvidia capacity is actually for. It also lands in a month when OpenAI has been tuning the whole GPT-5.6 range — an 80% price cut on its smallest model, Luna, and a tightly gated security-focused sibling — which reads as a model family being segmented by speed, price and permission, not just capability.
For Nvidia, it is a race it has already joined. In December, Nvidia paid about $20 billion to license the technology of Groq, the other famously fast inference chipmaker, and hire its founders. At its March developer conference it unveiled the Groq 3 LPU, an on-chip-memory inference part targeting 1,500 tokens a second, due to ship this quarter. Jensen Huang’s own framing at that event — that response speed is what “determines whether an AI product makes money” — is effectively the thesis of the OpenAI-Cerebras launch, stated by the incumbent it pressures. Speed is becoming a product line of its own, and the reason is agents: when software consumes a model’s output and acts on it in a loop, with no human reading along, throughput stops being a comfort feature and becomes the price of doing more steps per hour.
The constraint is capacity, not demand. Wafer-scale chips are made in far smaller volumes than GPUs, which is why the preview is small and TechCrunch notes wider rollout will not be a matter of flipping a switch. Cerebras — which went public on Nasdaq in May at an implied value near $56 billion — now has to build wafers as fast as OpenAI can sell speed.
The India angle
Nothing in the announcement restricts Ultrafast by country — access is gated by the waitlist, not geography, though no source we found confirms Indian developers are in the first cohort. The context that matters is that India is OpenAI’s second-largest market, with 100 million weekly ChatGPT users, and home to exactly the latency-obsessed workloads Ultrafast is aimed at: voice agents. Indian voice-AI firms such as Bolna already advertise sub-half-second response times on OpenAI models for tens of thousands of daily calls — for a phone conversation in Hindi or English, every saved millisecond is product quality.
The same day Ultrafast launched, India’s serving capacity took its own step up. Larsen & Toubro won an order it classes at ₹10,000-15,000 crore — up to about $1.57 billion — to build what it calls India’s largest single-cluster AI facility: 10,000 Nvidia B300 GPUs at its Chennai campus for the American inference cloud Together AI, on a site designed for 250 megawatts in its first phase. It joins a crowded buildout: Yotta’s 20,736 Blackwell Ultra GPUs targeted to go live this month, the IndiaAI Mission’s 38,000-plus subsidised GPUs at roughly ₹65 an hour for startups, and OpenAI’s own 100-megawatt Stargate India commitment with Tata, scaling toward a gigawatt. The pattern worth naming: the ultra-fast tier itself runs on wafers in American data centers, while India’s boom — not without local friction — is overwhelmingly GPU-based. If wafer-speed inference becomes a competitive edge, it is one more layer of the stack India currently imports.
What to watch
Four things will tell us whether Ultrafast is a milestone or a demo. First, the price — none has been published, and speed premiums tend to be steep. Second, independent numbers: Artificial Analysis and its peers measuring whether 750 tokens a second survives real context lengths and load, and whether quality really is unchanged. Third, Nvidia’s Groq 3 LPU shipping this quarter, which would give every other lab a fast-inference answer without a Cerebras contract. Fourth, whether the tier spreads — to Terra and Luna, to Codex, and to the enterprise agents OpenAI is openly courting. We track each of these threads daily in our AI models coverage.
The honest summary: the hardware story is real and two years deep, the flagship-model milestone is genuine, and every number attached to it is still the vendors’ homework. Speed has become the frontier labs’ favourite new axis of competition — which means the next few months of independent measurement matter more than launch day did.
Frequently asked questions
What is GPT-5.6 Sol Ultrafast?
It is a new serving tier in the OpenAI API, not a new model. The same GPT-5.6 Sol flagship is run on Cerebras wafer-scale chips instead of GPUs, which OpenAI says pushes output speed to as much as 750 tokens a second — up to 14 times its Standard tier.
How fast is Ultrafast in practice?
Two different numbers matter. The 14x figure describes raw token generation. On full working tasks — where a model also reads documents and calls tools — OpenAI and Cerebras claim a 5.6x end-to-end speedup on the GDPval benchmark. Both are vendor figures; independent testers have not yet published measurements.
How much does Ultrafast mode cost?
OpenAI has not published a price yet. The Standard tier lists at five dollars per million input tokens and thirty per million output tokens, and the existing Fast tier is listed by third-party pricing guides at roughly double the Standard rate — so a speed premium for Ultrafast would fit the established pattern.
Can developers in India use Ultrafast?
There is no stated country restriction — access is gated by a waitlist for a small preview group, and OpenAI says it will expand as capacity grows. OpenAI's API is otherwise fully available in India, which the company calls its second-largest market, with local data residency offered since 2025.
Sources & further reading
- Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed — OpenAI (official) (primary source)
- Accelerating GPT-5.6 Sol Ultrafast with OpenAI — Cerebras (official blog) (primary source)
- Cerebras Powers Ultrafast Mode for OpenAI's GPT-5.6 Sol — GlobeNewswire (official release) (primary source)
- OpenAI introduces 'Ultrafast,' a new mode that makes GPT-5.6 Sol work at 14x the speed — TechCrunch
- OpenAI previews 'Ultrafast' GPT-5.6 Sol running up to 14 times faster — 9to5Mac
- Gemini 3.7 Flash Is Out, But GPT-5.6 Sol Ultrafast Is Invite-Only — Decrypt
- Cerebras scores OpenAI deal worth over $10 billion — CNBC
- OpenAI Debuts First Model Using Chips From Nvidia Rival Cerebras — Bloomberg
- OpenAI launches GPT-5.3-Codex-Spark on Cerebras chips — Tom's Hardware
- Cerebras launches OpenAI's gpt-oss-120B at 3,000 tokens/sec — Cerebras (official blog)
- gpt-oss-120B providers, independent speed measurements — Artificial Analysis
- Cerebras stock debut on Nasdaq — CNBC
- Nvidia buying AI chip startup Groq's technology for about $20 billion — CNBC
- Nvidia announces Groq 3 LPU AI inference chip — DataCenterDynamics
- GDPval — OpenAI benchmark paper (arXiv:2510.04374) (primary source)
- India has 100M weekly active ChatGPT users, Sam Altman says — TechCrunch
- OpenAI taps Tata for 100MW AI data center capacity in India, eyes 1GW — TechCrunch
- L&T to set up Nvidia B300 AI Factory for Together AI — Business Standard
- India's Larsen & Toubro secures order worth up to $1.57 billion — Reuters via Yahoo Finance
- India to add 20,000 GPUs beyond existing 38,000 — PIB (official) (primary source)
- Yotta to deploy 20,736 Nvidia Blackwell Ultra GPUs — Yotta (official release) (primary source)
- GPT-5.6 Sol pricing and specs — OpenRouter listing
More of today, in 60 seconds: Today's Docket →