
Cerebras Systems
AI Hardware · Semiconductors · Inference Chips
Cerebras unveils CS-4, claiming 30x faster inference than GPUs
August 19, 2026
The system packs power delivery within half a millimeter of the chip, a redesign built for gigawatt-scale AI data centers.
- Cerebras unveiled its CS-4 rack-scale AI system at its Supernova event in San Francisco, built around three new WSE-3 Turbo wafer-scale chips and a new Nexus rack architecture, claiming up to 30x faster inference than GPU-based systems.
- Cerebras (Nasdaq: CBRS) builds wafer-scale accelerators that fit an entire chip on a single silicon wafer rather than small dies, positioning itself as a speed-focused inference alternative to Nvidia and AMD GPUs.
- The WSE-3 Turbo doubles memory bandwidth to 43.2 petabytes/sec and I/O to 2.4 Tbps versus its predecessor, achieved on the same TSMC 5nm silicon by pushing clock speeds roughly twice as high.
- A modular 'wafer-scale backpack' design folds power conversion, liquid cooling and I/O into a single compact assembly with 50% fewer components, cutting deployment time from days to hours.
- By cutting wafer-to-wafer interconnect latency to 2 microseconds, CS-4 can generate more than 1,000 tokens per second on models exceeding 10 trillion parameters.
- The launch follows Cerebras reporting Q2 core revenue up 103% year-over-year on $180.1 million in sales and a $6.9 million adjusted loss, as it targets 600 megawatts of deployed compute by the end of 2027.
- CS-4's power-delivery and modular packaging redesign targets the real economics bottleneck in AI infrastructure, throughput per watt and speed of deployment, as hyperscalers plan compute at the gigawatt scale.