Advertisement


Home Server Accelerators Cerebras Intros Faster WSE-3 Turbo Processor and First Rack-Scale CS-4 System

Cerebras Intros Faster WSE-3 Turbo Processor and First Rack-Scale CS-4 System

0
Cerebras Cs 4 Rack Hero Image
Cerebras Cs 4 Rack Hero Image

Back in 2019, Cerebras Systems introduced the first version of their rather unique AI accelerator, the Wafer Scale Engine. True to its name, the wafer-scale engine was a wafer-sized processor that promised to deliver significantly better performance by building a much bigger chip. Rather than slicing the wafer into individual chips, the wafer-scale engine design used almost the entire wafer as a single accelerator, effectively allowing the company to make the largest possible processor from a 300mm wafer.

Over time, Cerebras has produced three iterations of the wafer-scale engine, each boosting processor performance and attracting more developer interest and sales momentum. The most recent version of the wafer-scale engine, the WSE-3, was launched in 2024 and offers 125 PFLOPS of sparse FP16 performance within a single processor.

Now in 2026, the company is launching the next iteration of their hardware stack. At the heart of it is an updated version of the WSE-3, dubbed the WSE-3 Turbo, which promises to double the original WSE-3’s performance. Cerebras is not stopping there: for this generation, they are also launching a wholly new CS-4 rack system and networking architecture to house the WSE-3 Turbo, enabling multiple WSEs to work together within a single rack. If you want to get a sense of what the new CS-4 is replacing, here is the video we did with the CS-2’s “engine block” in 2022:

With the AI market continuing to boom, Cerebras is looking to not only continue to position their WSE products as a more capable rival to GPUs. But the company is also looking to lay the groundwork for the future, both in terms of the modern paradigm of scale-up systems, and where WSE products may fit into larger, disaggregated data centers. In short, Cerebras is having its rack-scale architecture moment, looking to go toe-to-toe with recent developments from both NVIDIA and AMD.

WSE-3 Turbo: Doubling Performance By Doubling Clocks

We will start things off with the WSE-3 Turbo, Cerebras’s new wafer-scale engine processor.

Cerebras Wse 3 Turbo Wafers
Cerebras WSE-3 Turbo Wafers

The heart of Cerebras’s next-generation rack-scale efforts, the WSE-3 Turbo is, unusually enough, the simpler of today’s product announcements. The WSE-3 Turbo is not an all-new WSE design. Rather, it is a faster version of the original WSE-3 launched a couple of years ago. On paper, it is designed to deliver twice the performance of the original WSE-3, and in practice, it comes very close to a WSE-3 running twice as fast. Traditionally, Cerebras WSE generations follow TSMC process generations, but not in this case.

Cerebras Wafer Scale Engine Generations
WSE-3 Turbo WSE-3 WSE-2
AI Cores 900,000 900,000 850,000
FLOPS (Sparse FP16) 250 PFLOPS 125 PFLOPS 75 PFLOPS
SRAM 44GB 44GB 40GB
Memory Bandwidth 43.2PB/sec 21PB/sec 20PB/sec
Fabric Bandwidth 53.5PB/sec 26.8PB/sec 27.5PB/sec
Network Bandwidth 300GB/sec 150GB/sec ?
Transistor Count 4 Trillion 4 Trillion 2.6 Trillion
Process Node TSMC 5nm TSMC 5nm TSMC 7nm
Power Consumption ? ~27kW ~23kW

By the numbers, we are once again looking at a wafer-sized processor comprising 900,000 of Cerebras’s AI cores, which are paired with 44GB of on-processor SRAM. Like the original WSE-3, the Turbo version comprises 4 trillion transistors altogether and is fabbed on TSMC’s 5nm process.

What is new this time around, then, is that Cerebras has seemingly cranked up the clock speeds across every aspect of the WSE-3. Not only is the rated compute throughput twice as high as before, at 250 PFLOPS for highly sparse FP16 data, but the on-wafer SRAM memory bandwidth has also doubled to 42.2PB/second. So, for that matter, the mesh fabric bandwidth and the off-processor network bandwidth are now 53.5PB/second and 300GB/second, respectively. With virtually every critical element of the WSE-3 running at twice the speed, Cerebras has a very straightforward path to delivering twice the performance with the WSE-3 Turbo.

Performance aside, the company has not yet disclosed the power consumption of the faster WSE. Doubling the clock speeds on a WSE-3 would significantly increase the WSE’s power consumption by pushing it farther up the voltage/frequency curve, and potentially outside its sweet spot. At present, Cerebras says that the CS-4 rack the WSE-3 Turbo was designed to go into “enables the delivery of twice as much power to the WSE-3 Turbo,” which implies that the Turbo WSE has doubled its power consumption as well as its performance, bringing it to around 54kW. If those figures are accurate, then that in and of itself would be an impressive development on Cerebras’s part, as it means they have tamed their voltage/frequency curve enough to double clockspeeds without encountering a superlinear increase in power consumption.

Cerebras CS-4 Rack-scale System: Scaling Up With 3 WSEs In a Rack

The second hardware announcement from Cerebras is the CS-4 rack-scale compute system, which is the company’s first true rack-scale system. Succeeding the CS-3 systems that housed the WSE-3 processors, the CS-4 is a radical rearchitecture of all the hardware that surrounds a single wafer-scale engine, with Cerebras laying the groundwork for scaling well beyond a single wafer-scale engine.

To provide some base context here, the CS-3 was a single WSE system, and essentially the smallest unit of hardware for the WSE ecosystem. Each CS-3 was a 16U liquid-cooled system designed to be mounted in a server rack, and while Cerebras offered a full rack configuration for the CS-3 (the aptly named “CS-3 Rack”), this was little more than two independent CS-3 systems in a single rack. More to the point, while CS-3 Racks could be deployed in larger clusters, the systems within did not function in anything approaching the current paradigm of a scale-up system.

Cerebras Cs 3 Rack
Cerebras CS-3 Rack – Two Single-WSE Systems In a Rack

To that end, for their CS-4 system, Cerebras has redesigned their entire server and rack architecture to enable better performance and scalability. The company needed a server architecture that could not only accommodate the faster (and more power-hungry) WSE-3 Turbo, but also enable better networking and scaling between nodes. The CS-4, in turn, is designed to meet all of those needs and more.

Cerebras System Generations
CS-4 CS-3
Wafers 3x WSE-3 Turbo 1x WSE-3
FLOPS (Sparse FP16) 750 PFLOPS 125 PFLOPS
SRAM 132GB 44GB
Memory Bandwidth 129.6PB/sec 21.6PB/sec
Fabric Bandwidth 160.5PB/sec 26.7PB/sec
I/O Bandwidth 900GB/sec 150GB/sec
I/O Latency 2µs 5µs

A full rack-sized system at the heart of the CS-4 now consists of three WSE-3 Turbo processors. Compared to a single WSE-3 (as Cerebras uses for their math), this is three times as many WSE processors as before. Even if we are more generous and compare it to the larger CS-3 Rack, this is still a 50% increase in WSEs per rack. Coupled with the WSE-3 Turbo’s doubled performance, a CS-4 rack can deliver six times the performance of a CS-3 system, or around 3x the performance of a CS-3 rack.

Cerebras Cs 4 Rack Stage Presentation
Cerebras CS-4 Rack Stage Presentation

But the bigger story in terms of hardware is everything else that goes around those WSEs. To house a larger number of hotter WSEs in a single rack, Cerebras designed a new rack-scale architecture that could not only handle the WSE-3 Turbo but also be modular enough to accommodate future WSEs and other technologies. This resulted in what Cerebras calls its Nexus platform.

Cerebras Nexus Rack Scale Server Architecture
Cerebras Nexus Rack Scale Server Architecture

The Nexus platform is designed to be highly modular and is broadly split into placing power supplies, fans, and other supporting hardware at the front of the rack, while the WSEs go at the rear. Nexus is designed to be a multi-generation design, allowing Cerebras (and its customers) to reuse Nexus racks in the future by swapping out WSEs and other equipment as needed, whether to repair systems or to swap parts for newer equipment. Cerebras is placing particular emphasis on how generic the design is, which is meant to give them the flexibility to incorporate newer technologies in the future as needed.

As for the WSSes in a Nexus rack, each is housed in what Cerebras calls a “backpack,” a self-contained (and vertically standing) WSE system designed to slot into the Nexus rack. The CS-4 “backpack” is the new “engine block” found in CS-1 to CS-3 (video earlier in this article).

Cerebras Wse Backpack
Cerebras WSE Backpack

The backpack is responsible not only for connecting the WSE to the rack’s power and liquid-cooling loops, but also for connecting it to the networking and other I/O hardware running throughout the rack.

Cerebras Backpack Exploded View
Cerebras Backpack Exploded View

Speaking of networking, this is another area where the CS-4 system is taking a big step up from the earlier CS-3 system. From a raw bandwidth standpoint, the WSE-3 Turbo has also doubled its external fabric/networking bandwidth, from 150GB/second to 300GB/second, which allows for both greater bandwidth and lower latencies between wafers. Cerebras is also using the Nexus platform introduction to introduce more (and more complex) networking hardware overall.

Cerebras Wafer Io Module
Cerebras Wafer I/O Module

Paired with each WSE in a backpack are wafer I/O modules, which house all the networking hardware. By using modular I/O, Cerebras can partially decouple the networking hardware from the WSEs, allowing them to update and upgrade the networking hardware independently of the WSEs. As envisioned, the goal of this design is to enable Cerebras to add support for newer networking technologies down the line, and, in the process, underscore just how modular the Nexus platform really is.

Cerebras Cs 4 Networking Topologies
Cerebras CS-4 Networking Topologies

For their initial CS-4 rack, Cerebras will use RDMA over Converged Ethernet v2 (RoCEv2) for scale-up/scale-out networking. However, with the degree of flexibility they are aiming for, it is not too hard to imagine them using something like UALink or Ultra Ethernet in future racks, as that hardware becomes available.

Meanwhile, Cerebras is offering a rather unusual secondary/optional networking topology for their CS-4 racks. In this second mode, rather than relying on Ethernet switches to establish a hub-and-spoke topology, the company is connecting the backpacks in a chain topology. According to Cerebras, this is intended to minimize latency between each wafer, bringing it down to 2 microseconds in total. This also simplifies the design of the Nexus racks somewhat and better ensures their modularity, as it means there are no major networking components within a rack other than those found in a backpack. What is left, then, is a series of direct connections between backpacks that foregoes the need for switches, or for that matter, a carefully coordinated collection of cables to attach backpacks to switches.

Cerebras Nexus Future Roadmap
Cerebras Nexus Future Roadmap

Looking towards the future, the Nexus platform will serve as the basis for the next few generations of Cerebras rack-scale systems. Along with the CS-4, the company has already committed to using it for the CS-5 and CS-6 systems that will come out later this decade.

Next, let us get to a disaggregated future for Cerebras.

LEAVE A REPLY

Please enter your comment!
Please enter your name here

This site uses Akismet to reduce spam. Learn how your comment data is processed.