Cerebras CS-4 Three Wafers Double Power and IO With Margin Proof Pending [2026]
Monday, September 21, 2026By Indie Kings | September 21, 2026
Updated September 21, 2026: Cerebras CS-4 packs three WSE-3 Turbo wafers into one Nexus rack with doubled power and I/O per wafer, per Cerebras and Futurum analysis of the August 18 Supernova 2026 event. First systems ship this quarter, but speed claims await independent benchmarks and margin questions stay open.
Image: Cerebras system art. Credit: Cerebras.
What CS-4 is
Per Cerebras investor release via Futurum on August 20, CS-4 is the fourth wafer-scale system and the first with three wafers in one rack, built on WSE-3 Turbo processors. Each wafer holds 900,000 AI cores with 44 GB on-wafer SRAM and 43.2 PB per second of memory bandwidth, per the same report.
Per Cerebras Hot Chips 2026 blog, the Nexus rack platform houses each wafer in a rear compute backpack with modular power, cooling and I/O, with a path to double token speed yearly. System totals per Futurum are 750 PFLOPs of compute with 7.2 Tb per second system I/O and 2-microsecond wafer-to-wafer latency.
Per Futurum, first shipments begin this quarter with CEO Andrew Feldman quote that speed is productivity. Company is NASDAQ CBRS, per the same report.
| CS-4 fact per Cerebras via Futurum | Figure |
|---|---|
| Wafers per rack | 3x WSE-3 Turbo |
| Cores per wafer | 900,000 AI cores |
| On-wafer SRAM per wafer | 44 GB, same as 2024 silicon |
| Memory bandwidth per wafer | 43.2 PB per second |
| System I/O | 7.2 Tb per second |
| Rack power | 120 to 140 kW |
Power does the work of a new chip
Per Cerebras Hot Chips blog, AC and DC converters sit about 0.5 millimeters from the wafer, 100 times closer than GPU norms near 50 millimeters, with no board in the final path. That feeds nearly twice the power at almost the same voltage, per the company.
Per Futurum analyst take, the WSE-3 Turbo is a tuned version of 2024 silicon with identical SRAM, so gains come from the rack: doubled power per wafer, direct liquid cooling in a rear Wafer-Scale Backpack, plus networking overhaul. Backpack claims are 50 percent fewer components with 60 percent more automated manufacturing and hours-scale deployment, per Cerebras via Futurum.
Per Futurum, 120 to 140 kW per rack lands near half the 240 to 250 kW expected of AMD Helios and NVIDIA Vera Rubin class racks. Treat rival figures as per-current-disclosures estimates, not measured head-to-head.
- Feed: Up to 30 air-cooled power modules per backpack, 277V AC to 54.5V DC, per Cerebras.
- Cool: Per-backpack water conditioning with quick-disconnect and leak safe-state, per Cerebras.
- Risk: Double wattage into unswappable silicon plus 10x ramp across three new contract makers, per Futurum bear case.
I/O doubling and the clock read
Per Futurum, the new Wafer I/O Module doubles off-wafer bandwidth to 2.4 Tb per second per wafer over standard RoCE v2 Ethernet, with Direct Wafer Links joining wafers at 2 microseconds without a switch. Arista Etherlink switches build the between-rack fabric shown at Supernova, per the same report.
Per Irrational Analysis investor recap of Supernova, each wafer doubled I/O with a single I/O card on top, which the author reads as a clock-speed doubling with zero redesign plus FPGA-side latency gains. That clock read is author inference, not a Cerebras statement, and a possible IO ASIC swap is flagged by the author as unconfirmed.
Per Cerebras, on-wafer fabric hits 53.5 PB per second aggregate against 260 TB per second NVLink in a 72-GPU Rubin rack. Vendor comparison with selected configs, pending independent runs.
| Claim | Grade at draft time |
|---|---|
| 2x faster plus 10x token capacity vs CS-3 | Vendor claim, configs undisclosed |
| Up to 30x vs GPU systems | Vendor claim vs selected configs |
| 4,400 tok per second per user on GPT-OSS-120B | Artificial Analysis plus internal, batch and context unstated per Futurum |
| 1,000 plus tok per second on 10T plus param models | Vendor target claim |
| I/O gain from clock doubling | Investor inference, unconfirmed |
Yield and margin unknowns
Per Irrational Analysis, catastrophic yield is solved with 100 percent of wafers functional, but parametric yield at clock and power is unclear, modeled by the author at 20 percent total compound. The author shared an editable model and invites corrections, which is admirable process but still one investor is spreadsheet, not fab data.
Per Futurum financial read, Q2 core revenue hit $209.9 million up 103 percent yearly with cloud up 281 percent to $126 million while hardware halved to $54.1 million. Backlog sits at $25.4 billion in remaining obligations with a 10x manufacturing ramp and 600 MW of capacity targeted by end of 2027, per the same report.
Whether CS-4 converts backlog to revenue shows in Q3 hardware prints plus this-quarter shipment proof, per Futurum watch list. Roadmaps beyond sit as targets: CS-5 in 2027 up to 10,000 tokens per second per user on open models, CS-6 with 3D-stacked DRAM, per Cerebras. Forward-looking statements carry the company is own SEC-risk framing.
- Watch: This-quarter shipments plus independent trillion-param benches with disclosed configs.
- Watch: Q3 hardware revenue rebound in November earnings.
- Watch: Production disagg clusters with AMD Helios or AWS Trainium prefill outside vendor labs.
FAQ
What launched at Supernova 2026?
Per Cerebras via Futurum, CS-4 on August 18 with three WSE-3 Turbo wafers in a Nexus rack, shipping this quarter.
Is the wafer new?
Per Futurum, WSE-3 Turbo is tuned 2024 silicon with same 44 GB SRAM. Gains come from power, cooling and networking.
What did the investor recap add?
Per Irrational Analysis, a clock-doubling read on I/O gains plus solved catastrophic yield against unknown parametric yield. Inference, not company guidance.
Are speed claims proven?
No. Per Futurum, configs behind headline numbers are undisclosed. Await independent benches.
What is disaggregated inference here?
Per Cerebras and Futurum, Cerebras decode paired with AMD Helios or AWS Trainium prefill over Ethernet. Combined speed claims are pre-production.
Is this financial advice?
No. Hardware analysis only. Revenue and backlog figures are per Futurum reads of company prints.
Bottom Line
CS-4 is a rack doing a chip is job: same silicon, double power and I/O, triple wafers. Believe the packaging, verify the speed claims, and watch shipments plus Q3 hardware revenue before calling the backlog converted.
Related: Intel confirms Arc future with Xe3P | Intel Arc Battlemage failure or future from 2025