Source: Unite.AI
Cerebras Systems said on October 1, 2026 that it increased inference throughput by 5x in early results using a technique called disaggregation, with the same number of Cerebras systems and no loss in token generation speeds. The disclosure came in Disaggregated Inference From the Ground Up, a company blog post by Isaac Tai and Zhenwei Gao that opens a planned series on the subject.
The post frames the series for readers who have heard the term disaggregation, or the claim that prefill is compute-bound and decode is memory-bound, and wondered what either actually means. It builds the explanation from the ground up, beginning with how accelerators balance arithmetic against data movement.
Prefill and Decode Place Different Demands on Hardware
The post defines arithmetic intensity as the number of floating-point operations divided by the number of bytes transferred between memory and an accelerator’s compute units. In one of its examples, adding two matrices performs 1 FLOP for every 6 bytes moved, an arithmetic intensity of 0.167 FLOP per byte, and that ratio stays constant as the matrices grow. Matrix multiplication behaves differently: each output value is built from an entire row of one input and an entire column of the other, so loaded values contribute to more outputs and arithmetic intensity grows with input size.
Inference, the post explains, is a chain of such matrix multiplications between a model’s fixed weights and its input tokens, and it runs in two phases with different intensity profiles. During prefill, the entire prompt is processed in parallel as a large matrix operation, and real-world prompts can contain thousands or even hundreds of thousands of tokens. During decode, tokens are generated one at a time, and the model relies on the KV cache, which stores keys and values computed for earlier tokens so they are reused rather than recomputed.
Both phases must still move all of the model’s weights, potentially hundreds of gigabytes or terabytes, to the compute units for every token generated. The post shows memory movement staying nearly constant while arithmetic intensity falls after prefill, and it cites that gap as the reason adding raw compute capacity does not necessarily make tokens arrive faster during decode.
Disaggregation Splits Inference Into Separate Pools
In production, an inference server typically handles many requests concurrently, and when prefill and decode run on the same hardware, compute-intensive prefill can stall active decode requests. Schedulers then have to choose among getting new requests to their first token quickly, keeping active responses streaming smoothly, and maximizing total throughput. Batching lets concurrent requests share the work of reading model weights, but larger batches can make each decode step take longer, so total throughput can rise while each user receives tokens more slowly.
The post describes disaggregation as a systems design pattern that runs the two phases in separate hardware pools. Once the stages are separated, operators can allocate hardware, set batching policies, and prioritize latency or throughput for each stage independently: a system with strict time-to-first-token targets can reserve more capacity for prefill, while one built around smooth streaming can give decode a larger or more tightly scheduled pool. The pools can also be scaled individually.
Separation introduces a new requirement. After prefill builds the KV cache, that request-specific state must be transferred to the decode pool, where it is loaded into memory before generation can continue, while the model weights are already loaded in both pools. The post notes that the handoff adds network and coordination overhead, that either pool can sit idle if capacities do not match demand, and that the added latency depends on whether the cache moves across colocated machines or across regions. It argues disaggregation is most compelling at scale, where gains from independently sizing and scheduling the pools can outweigh the transfer and operating costs, and that it changes the control interface of the serving system rather than only smoothing streaming.
Heterogeneous Hardware, Early Results, and Partnerships
Cerebras said it is leading the development of heterogeneous disaggregation, combining multiple types of chips in one inference system and assigning different hardware to the segments that are memory-bound or compute-bound. The post contrasts the company’s wafer-scale design, which distributes SRAM alongside compute across the entire wafer, with GPUs, which stage model data from high-bandwidth memory through smaller on-chip memories and caches.
A published-peak memory-bandwidth chart in the post, dated September 10, 2026, lists Cerebras WSE-3 on-chip SRAM at 21,000 TB/s per wafer, alongside an unnamed on-chip SRAM accelerator at 150 TB/s and HBM4 GPUs at 23.3 and 22 TB/s. The chart cautions that SRAM figures sum local memory bandwidth across a processor while HBM figures measure traffic from off-chip memory, so the figures describe different memory tiers rather than measured token speeds.
The post also reproduces Artificial Analysis data from September 10, 2026 for GPT-oss-120B running high reasoning with 10,000 input tokens. It lists Cerebras at 1,669 output tokens per second, SambaNova at 708, Groq at 475, Microsoft Azure at 319, Nebius at 294, and Baseten at 293.
Cerebras said that in a traditional aggregated system, increasing capacity meant deploying more hardware, and that by leveraging partner accelerators to handle prompt processing it increased capacity by 5x in early tests with the same WSE footprint. The company said it has announced partnerships with multiple hardware partners to bring more ultrafast tokens to market, and a diagram in the post shows AWS Trainium and AMD Helios Instinct GPU systems among the prefill hardware options feeding a Cerebras decode pool.
The post identifies agentic applications as a compelling fit for heterogeneous disaggregation, since they often involve long, multi-turn workflows in which context grows across model calls and delays at each step compound. Cerebras said the next installments will cover the hardware and software stacks involved and the economic trade-offs of deploying disaggregated inference at scale.
