OpenAI’s GPT 6.1 Sol Ultrafast Is Not Running On Cerebras But On NVIDIA GPUs, Says Semi Analysis

OpenAI’s new Ultrafast tier for GPT-6.1 Sol isn’t being served by the chipmaker most closely associated with OpenAI’s push for faster inference. That’s according to Semi Analysis, which said on X that GPT-6.1 Sol Ultrafast is “NOT running on Cerebras” and is instead running at a low batch size on NVIDIA GPUs. The firm also posed two pointed questions to its followers: what does this say about Cerebras, and will Cerebras serve the model in the future?

NVIDIA’s official AI account appeared to enjoy the moment, posting nothing but a pair of eyes emoji in response. The post had crossed 100,000 views within hours.

What OpenAI announced

OpenAI unveiled GPT-6.1 Sol at DevDay on September 29, pitching it as a model that nearly matches its flagship GPT-6 Astra on agentic coding, computer use and professional work at roughly a fifth of the price. Alongside it, the company said an Ultrafast version was coming within days, promising up to 8x faster token generation in Codex.

Ultrafast is sold as a premium lane. OpenAI says it delivers up to 6x faster output in the API at 6x the standard price, with speeds of roughly 300 tokens per second. Notably, those speeds are well below the 750 tokens per second OpenAI touted for GPT-5.6 Sol on Cerebras in August.

Ultrafast is already live for GPT-6 Astra on the $500 Pro tier and Enterprise plans. Ultrafast isn’t cheap — it costs $60/million input and $300/million output tokens.

openai ultrafast pricing

Why “low batch size” matters

Batch size refers to how many user requests a GPU processes together. Large batches keep the hardware busy and make each token cheap, but every individual user waits longer. Running at a low batch size does the opposite: each user gets faster output, but the GPUs are used less efficiently and each token costs more to produce.

That trade-off lines up with how Ultrafast is priced. A 6x premium is what you might expect if OpenAI is effectively reserving a larger slice of expensive GPU capacity for each request. It is an inference from the pricing and Semi Analysis’s claim, however, not something OpenAI has confirmed. OpenAI has not said publicly which hardware serves the 6.1 Sol Ultrafast tier, and Cerebras has not commented on the claim at the time of writing.

OpenAI and Cerebras: a deep relationship

The claim lands on one of the more closely watched partnerships in AI. OpenAI and Cerebras have been intertwined for years: Greg Brockman and Sam Altman both held personal stakes in Cerebras when OpenAI explored acquiring the company in 2017, a detail that surfaced during the Musk trial.

The relationship turned commercial in a big way in January 2026, when OpenAI signed a multi-year deal for 750 megawatts of Cerebras compute through 2028, initially reported as worth more than $10 billion. Later reporting put the figure above $20 billion, alongside roughly $1 billion from OpenAI toward Cerebras data centers and warrants that could give OpenAI a meaningful equity stake.

The first product came in February with GPT-5.3-Codex-Spark, a smaller coding model running at over 1,000 tokens per second on Cerebras’s wafer-scale chips. At the time, OpenAI described it as the first step toward putting larger frontier models on the hardware as capacity ramped up, while stressing that GPUs remain the backbone of its training and most cost-effective inference.

Cerebras then went public on Nasdaq in May, pricing at $185 a share and raising about $5.5 billion. In August, it got the endorsement investors had been waiting for when OpenAI previewed Ultrafast for GPT-5.6 Sol, running the full flagship at up to 750 tokens per second on Cerebras chips. Shares jumped on the news, with analysts reading it as a sign the partnership was progressing.

What this says about Cerebras

It would be premature to read the Semi Analysis claim as a verdict on Cerebras. Its chips keep model weights in on-chip memory, about 44 GB per wafer, which is the source of its speed advantage. Very large models have to be spread across multiple wafers, which adds complexity. Whether GPT-6.1 Sol and Astra-class models are simply harder to serve that way, or whether capacity and timing are the constraint, is something only OpenAI and Cerebras can answer.

There are also simpler explanations. Cerebras capacity is still ramping, the GPT-5.6 Sol Ultrafast tier remains in limited preview, and OpenAI may be using GPUs to get a new model to market quickly while Cerebras deployments catch up. OpenAI has been explicit that it runs a mixed fleet across NVIDIA, AMD, Broadcom and Cerebras silicon.

For investors, the concern is concentration. A large share of Cerebras’s inference backlog is tied to a single customer, so any sign that OpenAI’s newest models are launching elsewhere invites questions. Whether Cerebras ends up serving GPT-6.1 Sol Ultrafast later remains an open question, and one neither company has addressed.

For NVIDIA, meanwhile, the eyes emoji says plenty. Even in the category where Cerebras has built its reputation, a GPU-based setup appears to be good enough for OpenAI’s newest speed tier.

Posted in AI