Cerebras Systems: The Ten-Year Rise of Wafer-Scale Chips
The Startup Story: From "Impossible" to Reality
Cerebras Systems was born in 2015 from a bold vision: build a wafer-scale chip dozens of times larger than traditional GPUs to break through AI computing bottlenecks. Wafer-scale integration (WSI) had never been successfully commercialized in 75 years, and many dismissed the idea as fantasy. Co-founder and CEO Andrew Feldman insisted the company's mission was to "do what others had tried and failed," aiming to write its own chapter in computing history.
After four years of secretive development, Cerebras unveiled the Wafer Scale Engine 1 (WSE-1) in 2019 — 1.2 trillion transistors across 46,225 mm², the world's largest chip at the time. Iterations followed:
- WSE-2 (2021): 7nm process, 2.6 trillion transistors
- WSE-3 (2024): 5nm process, 4 trillion transistors, 125 PFLOPS of AI compute
- Hardware and cloud: Cerebras sells CS-series systems directly to data centers and research institutions, while Cerebras Cloud offers on-demand or dedicated capacity. Integration with AWS and Azure extends its reach. In 2025, Cerebras signed a supercomputing deal with OpenAI potentially worth $30 billion, providing 750 MW of AI inference capacity with an option for an additional 1.25 GW.
- Product line and services: The flagship CS-3 delivers 125 PFLOPS peak; multiple systems can be clustered into AI supercomputers. End-to-end services include data preparation, model architecture design, training management, and inference optimization. Cerebras reports its top ten customers increase spending by 80% on average within 12 months.
- Financials and valuation: 2025 revenue reached $510 million, up 76% year-over-year, with $87.9 million net profit — a rare profitability milestone for a chip startup. February 2026's Series H valued the company at $23 billion; the IPO aims to raise a further $200 million for capacity and data center expansion.
- Ecosystem lock-in: NVIDIA's CUDA ecosystem has strong developer inertia; Cerebras mitigates this with PyTorch support.
- Capacity and cost: WSE-3 currently can only be produced on TSMC's 5nm process; improving yields and costs is critical.
- Competition: AMD and Intel are accelerating AI accelerator efforts, while Groq and SambaNova pursue alternative architectures.
The journey was not smooth: a 2024 IPO filing was withdrawn due to a CFIUS review tied to investment from the UAE's G42 Group. After adjustments, Cerebras refiled its S-1 in April 2026 to list on Nasdaq. Customers now include the Mayo Clinic and AstraZeneca.
Technical Barriers: Wafer-Scale Breakthroughs
The core innovation is treating an entire wafer as a single chip. The WSE-3 is 58 times the die area of NVIDIA's flagship B200, with 900,000 AI-optimized cores, 44GB of on-chip SRAM, and 21 PB/s memory bandwidth — 2,625 times the B200's packaged memory bandwidth.
1. Ultra-high memory bandwidth and capacity: AI workloads are often bottlenecked by the "memory wall." By integrating storage and compute on the same wafer, the WSE-3 avoids frequent external memory reads. Cerebras claims its CS-3 system achieves 15x the inference speed of traditional GPUs.
2. Fault tolerance and yield breakthroughs: Cerebras invented redundant die-level interconnect and fault-tolerant architecture: the wafer is divided into many small dies linked by high-speed interconnects. If a die is defective, traffic simply routes around it — the wafer does not need to be discarded. This means Cerebras does not require 100% yields, making wafer-scale production viable for the first time.
A supporting software stack (CSoft compiler, inference services, cluster manager) lets developers map models onto the WSE without writing CUDA or managing distributed clusters, though PyTorch compatibility lowers migration barriers.
Business Logic: Hardware Plus Services
| Metric | 2024 (estimated) | 2025 | |---|---|---| | Revenue | ~$289M | $510M | | Net profit | Loss | $87.9M |
Reshaping the AI Chip Landscape: Challenges and Outlook
Cerebras has demonstrated speed advantages in inference: over 2x faster than NVIDIA Blackwell GPUs on Llama 4 tasks and reportedly 5x faster on GPT-OSS-class models, which helped it win fast-inference business from NVIDIA.
Key challenges remain: