Cerebras CS-4 Three Wafers Double Power and IO With Margin Proof Pending [2026]

By Indie Kings | September 21, 2026

Updated September 21, 2026: Cerebras CS-4 packs three WSE-3 Turbo wafers into one Nexus rack with doubled power and I/O per wafer, per Cerebras and Futurum analysis of the August 18 Supernova 2026 event. First systems ship this quarter, but speed claims await independent benchmarks and margin questions stay open.

Cerebras AI system promotional art

Image: Cerebras system art. Credit: Cerebras.

What CS-4 is

Per Cerebras investor release via Futurum on August 20, CS-4 is the fourth wafer-scale system and the first with three wafers in one rack, built on WSE-3 Turbo processors. Each wafer holds 900,000 AI cores with 44 GB on-wafer SRAM and 43.2 PB per second of memory bandwidth, per the same report.

Per Cerebras Hot Chips 2026 blog, the Nexus rack platform houses each wafer in a rear compute backpack with modular power, cooling and I/O, with a path to double token speed yearly. System totals per Futurum are 750 PFLOPs of compute with 7.2 Tb per second system I/O and 2-microsecond wafer-to-wafer latency.

Per Futurum, first shipments begin this quarter with CEO Andrew Feldman quote that speed is productivity. Company is NASDAQ CBRS, per the same report.

CS-4 fact per Cerebras via FuturumFigure
Wafers per rack3x WSE-3 Turbo
Cores per wafer900,000 AI cores
On-wafer SRAM per wafer44 GB, same as 2024 silicon
Memory bandwidth per wafer43.2 PB per second
System I/O7.2 Tb per second
Rack power120 to 140 kW

Power does the work of a new chip

Per Cerebras Hot Chips blog, AC and DC converters sit about 0.5 millimeters from the wafer, 100 times closer than GPU norms near 50 millimeters, with no board in the final path. That feeds nearly twice the power at almost the same voltage, per the company.

Per Futurum analyst take, the WSE-3 Turbo is a tuned version of 2024 silicon with identical SRAM, so gains come from the rack: doubled power per wafer, direct liquid cooling in a rear Wafer-Scale Backpack, plus networking overhaul. Backpack claims are 50 percent fewer components with 60 percent more automated manufacturing and hours-scale deployment, per Cerebras via Futurum.

Per Futurum, 120 to 140 kW per rack lands near half the 240 to 250 kW expected of AMD Helios and NVIDIA Vera Rubin class racks. Treat rival figures as per-current-disclosures estimates, not measured head-to-head.

  • Feed: Up to 30 air-cooled power modules per backpack, 277V AC to 54.5V DC, per Cerebras.
  • Cool: Per-backpack water conditioning with quick-disconnect and leak safe-state, per Cerebras.
  • Risk: Double wattage into unswappable silicon plus 10x ramp across three new contract makers, per Futurum bear case.

I/O doubling and the clock read

Per Futurum, the new Wafer I/O Module doubles off-wafer bandwidth to 2.4 Tb per second per wafer over standard RoCE v2 Ethernet, with Direct Wafer Links joining wafers at 2 microseconds without a switch. Arista Etherlink switches build the between-rack fabric shown at Supernova, per the same report.

Per Irrational Analysis investor recap of Supernova, each wafer doubled I/O with a single I/O card on top, which the author reads as a clock-speed doubling with zero redesign plus FPGA-side latency gains. That clock read is author inference, not a Cerebras statement, and a possible IO ASIC swap is flagged by the author as unconfirmed.

Per Cerebras, on-wafer fabric hits 53.5 PB per second aggregate against 260 TB per second NVLink in a 72-GPU Rubin rack. Vendor comparison with selected configs, pending independent runs.

ClaimGrade at draft time
2x faster plus 10x token capacity vs CS-3Vendor claim, configs undisclosed
Up to 30x vs GPU systemsVendor claim vs selected configs
4,400 tok per second per user on GPT-OSS-120BArtificial Analysis plus internal, batch and context unstated per Futurum
1,000 plus tok per second on 10T plus param modelsVendor target claim
I/O gain from clock doublingInvestor inference, unconfirmed

Yield and margin unknowns

Per Irrational Analysis, catastrophic yield is solved with 100 percent of wafers functional, but parametric yield at clock and power is unclear, modeled by the author at 20 percent total compound. The author shared an editable model and invites corrections, which is admirable process but still one investor is spreadsheet, not fab data.

Per Futurum financial read, Q2 core revenue hit $209.9 million up 103 percent yearly with cloud up 281 percent to $126 million while hardware halved to $54.1 million. Backlog sits at $25.4 billion in remaining obligations with a 10x manufacturing ramp and 600 MW of capacity targeted by end of 2027, per the same report.

Whether CS-4 converts backlog to revenue shows in Q3 hardware prints plus this-quarter shipment proof, per Futurum watch list. Roadmaps beyond sit as targets: CS-5 in 2027 up to 10,000 tokens per second per user on open models, CS-6 with 3D-stacked DRAM, per Cerebras. Forward-looking statements carry the company is own SEC-risk framing.

  • Watch: This-quarter shipments plus independent trillion-param benches with disclosed configs.
  • Watch: Q3 hardware revenue rebound in November earnings.
  • Watch: Production disagg clusters with AMD Helios or AWS Trainium prefill outside vendor labs.

FAQ

What launched at Supernova 2026?
Per Cerebras via Futurum, CS-4 on August 18 with three WSE-3 Turbo wafers in a Nexus rack, shipping this quarter.

Is the wafer new?
Per Futurum, WSE-3 Turbo is tuned 2024 silicon with same 44 GB SRAM. Gains come from power, cooling and networking.

What did the investor recap add?
Per Irrational Analysis, a clock-doubling read on I/O gains plus solved catastrophic yield against unknown parametric yield. Inference, not company guidance.

Are speed claims proven?
No. Per Futurum, configs behind headline numbers are undisclosed. Await independent benches.

What is disaggregated inference here?
Per Cerebras and Futurum, Cerebras decode paired with AMD Helios or AWS Trainium prefill over Ethernet. Combined speed claims are pre-production.

Is this financial advice?
No. Hardware analysis only. Revenue and backlog figures are per Futurum reads of company prints.

Bottom Line

CS-4 is a rack doing a chip is job: same silicon, double power and I/O, triple wafers. Believe the packaging, verify the speed claims, and watch shipments plus Q3 hardware revenue before calling the backlog converted.

Related: Intel confirms Arc future with Xe3P | Intel Arc Battlemage failure or future from 2025

Share