Executive Summary
Cerebras and Gimlet Labs plan to deploy up to 100 megawatts of Cerebras powered inference capacity through Gimlet Cloud, with the first site due online later this year and a target of 3,000 tokens per second at production scale. The interesting part is not the number. It is that Gimlet is building a multisilicon cloud, pairing wafer-scale engines with GPUs inside a single request rather than picking one architecture.
Prefill and decode stress different parts of an accelerator, so Gimlet splits them and routes each phase to the silicon that suits it. The company reports 3 to 10 times better interactivity at a given throughput efficiency target across frontier workloads, which is a vendor figure and not independently measured. The bet is that the orchestration layer, not the chip, becomes the product. If it holds, single architecture inference clouds are leaving performance on the table.
Gimlet Cloud is the first inference platform to sell wafer-scale silicon and GPUs as one system. The two companies put 100 megawatts of Cerebras capacity behind a disaggregated inference cloud, and they expect the first Cerebras powered datacenter to come online later this year.
The reason to care is latency compounding. A task that takes ten minutes at 100 tokens per second can finish in about twenty seconds at 3,000, because an agent makes many sequential calls. Faster tokens do not just feel better. They change how much work a single agent can complete.
Prefill and decode want different hardware
An inference request runs in two phases, and they are not alike. Prefill ingests the prompt, which is compute heavy and batch friendly. Decode emits tokens one at a time, which is bound by memory bandwidth and latency. Running both on the same silicon means optimising for an average of two different problems.
Gimlet takes the disaggregation further than most. Its software splits prefill from decode, splits attention from the feed forward block, and applies speculative decoding, then maps each piece to the hardware that fits. A Cerebras wafer-scale engine brings large on-wafer SRAM and high memory bandwidth, which suits the latency sensitive decode stage. GPUs bring aggregate throughput, which suits the batch heavy work.

Developers see none of this. They call one inference API and the orchestration layer decides where each stage runs. That abstraction is the actual product, and it is the architecture the company is selling.
100 megawatts is the size of the bet
Capacity is the part that makes this more than a software claim. The two companies are putting up to 100 megawatts of Cerebras powered inference capacity behind the platform, which is a real datacenter commitment rather than a benchmark result. They say joint customer work has been running since last year, with an integrated setup already serving inference traffic in private deployments.
For buyers, this maps onto a familiar question about inference economics. Homogeneous fleets force a trade between interactivity and throughput per kilowatt, and most operators pick one and live with the other. If disaggregation genuinely widens that frontier, the gain shows up as lower cost per token at the latency a product needs.
Treat the performance claims accordingly. The 3,000 tokens per second figure and the 3 to 10 times interactivity range are company reported and vary with workload, model, and configuration. Nobody outside the two firms has reproduced them yet.
CS-4 reaches customers through Gimlet in 2027
The partnership also makes Gimlet a launch partner for Cerebras CS-4. Gimlet Cloud customers are expected to get direct access to the newer system in 2027, which puts a date on the roadmap rather than leaving it at a press release. Cerebras will supply the systems over one to two years, and Gimlet will operate and maintain them.
That division of labour is the part worth watching. The vendor supplies silicon, the inference cloud supplies orchestration, operations, and the developer contract. It is the same split we saw in the managed inference deals from earlier this year, where a model lab hands operations to a specialist and keeps the API in front of the customer.
The pattern underneath all of it is that inference capacity is becoming a portfolio problem. Power and siting gate the build, as our AI factory power analysis argues, and the serving stack decides whether that capacity earns its keep. Three questions for anyone running inference at scale. Which stages of your workload are latency bound and which are throughput bound? Are you paying for one architecture to do both jobs? And what would you save if the boring half of your traffic ran on cheaper silicon?
Get the next one before it is old news
Independent analysis of cloud-native infrastructure, Kubernetes and data center economics. No vendor spin.
