Same wafer, twice the clock, three per rack
Cerebras doubles CS-3 throughput without a new chip — and takes rack power from 23kW to about 130kW to do it.
2 minDePIN Compute & Bandwidth
The CS-4 uses the same 5nm WSE-3 wafer as the CS-3. What changed is everything around it: power delivery and cooling were rebuilt to roughly double clock speeds, off-wafer I/O went from 1.2Tb/s to 2.4Tb/s, and a rack now holds three wafers instead of two.
The claimed result is about 4,000 tokens per second per user on frontier models against 2,000 for the CS-3, with 43 PB/s of on-chip memory bandwidth and up to 30× the interactivity of GPUs. SRAM per wafer is unchanged at 44GB. Network latency drops from 5 to 3 microseconds, and direct wafer-to-wafer links reach 2.
Read the power line twice
A CS-3 wafer draws about 23kW. A CS-4 rack is quoted at 125 to 135kW. Even allowing for three wafers instead of one, that is a large step, and it lands on the same datacentre constraints — cooling capacity and grid connection — that decide whether a site can host the machine at all.
The analysis argues total cost of ownership stays similar to the CS-3 despite the doubled performance. That claim rests on power being available at a price, which is the assumption most likely to fail first.
Neither pricing nor general availability is stated. For anyone comparing rented inference capacity, the number that will actually show up in a quote is the one still missing.
Retold from SemiAnalysis. This is a summary in our own words; follow the link for the original reporting.