Cerebras has introduced the CS-4, a new rack-scale AI accelerator that takes direct aim at the GPU bottleneck in AI inference. According to the company, the CS-4 is designed for frontier AI models and delivers up to 30x faster inference than production GPU systems.
At the core of the system are three WSE-3 Turbo wafers. Cerebras states that each wafer delivers twice the speed of its previous generation, shifting the inference baseline to produce up to 10x more throughput per watt than the CS-3. The raw output is significant: on models exceeding 10 trillion parameters, the CS-4 can push out more than 1,000 tokens per second.
Hardware Built for Hyperscale
The CS-4 relies on a new modular architecture that physically separates the stable power and cooling layer from the compute units. Cerebras calls the compute unit a "Wafer-Scale Backpack." It is a self-contained 3D package that folds together the wafer, power conversion, direct liquid cooling, and high-speed I/O.
This backpack redesign uses 50% fewer components than previous setups. By allowing data centers to install and qualify the power and cooling rack before the compute arrives, Cerebras claims deployment times are reduced from days to hours.
Much of the CS-4's performance jump comes from rethinking power delivery and data transfer. In the new system, power delivery is positioned just 0.5 millimeters away from the processor. This is roughly 100 times closer than the 50-millimeter distance typical of conventional GPU boards. Shrinking this gap nearly eliminates board-level power loss, pushing twice as much power to the WSE-3T to enable higher operating frequencies.
Additionally, a programmable I/O subsystem enables wafer-to-wafer interconnect latency as low as two microseconds. Cerebras highlights this low latency as a key factor in maintaining interactive speeds on massive models.
What This Means for SMB Operations
Small and mid-sized businesses will likely never purchase a Cerebras CS-4 rack. But the hardware running inside hyperscale data centers directly dictates the cost, speed, and capability of the AI tools businesses use to run their operations.
For companies using AI to automate administration and operational workflows, the CS-4’s 30x leap in inference speed resolves a critical bottleneck. Current AI automations—like parsing a complex supplier contract, extracting key clauses, cross-referencing an inventory database, and drafting an email response—often rely on agentic workflows. These workflows require the AI model to execute multiple hidden reasoning steps before delivering a final output.
When data centers run on slower hardware, token generation limits the speed of multi-step workflows. A complex operational task can take 30 to 60 seconds to execute. That latency makes real-time, customer-facing automations impossible and bogs down internal data processing at scale.
If data centers can generate tokens up to 30 times faster using hardware like the CS-4, those multi-step reasoning processes execute in a fraction of a second. Fast inference turns sluggish, batch-processed workflows into instant, interactive operations. Instead of waiting for an AI agent to process a customer request, the system can complete the logic and trigger the appropriate software action instantly.
Furthermore, the CS-4’s claim of greater throughput per watt drives down the compute cost per token. For a small business running hundreds of automated tasks a day—from categorizing helpdesk tickets and routing leads to updating CRM records—cheaper inference means deploying more highly capable models to handle mundane tasks without inflating the software budget.
First shipments of the CS-4 begin this quarter. As this type of wafer-scale hardware comes online at the hyperscale level, expect the underlying API costs for frontier models to drop—and the speed of business automation to rise.