AMD Helios Compute Tray: Venice + MI455X
Digging down a layer, here is a more detailed look at the compute trays for Helios. The sizable double-wide trays weigh 170 pounds, and some 265 pounds of force is needed to insert them. If you saw our How Liquid Cooling is Prototyped and Tested in the CoolIT Liquid Lab Tour, holding an HPE Cray Shasta liquid-cooled node, this is heavier. That Shasta node we showed you years ago, in many ways, feels like a spiritual ancestor to the modern Helios compute nodes.

The single 96-core EPYC “Venice” CPU and four MI455X make up the core compute hardware of the aforementioned compute tray. AMD has been developing four-way and eight-way GPU systems for a couple of generations now, so there are few surprises here. The biggest change here is that all of the GPUs, or, as AMD technically classifies them, enhanced accelerator modules, are attached to each other via what is physically the Ultra Accelerator Link over Ethernet (UALoE), with AMD’s Infinity Fabric protocol running on top of it. Altogether, there are 36 UALoE links available from each GPU, each offering 400Gbps of bandwidth (or more specifically, two 200Gbps teamed links).

Meanwhile, the EPYC CPU has its own classic Infinity Fabric link to each GPU to provide full memory coherency across all the chips. As for its own memory, AMD is running a fully populated configuration here, with 16 RDIMM slots, each filled with 64GB DDR5-12800 ECC MRDIMMs for a total of 1TB of DDR5 memory supplying 1.6TB/second of memory bandwidth per tray. A notable design difference between the AMD and NVIDIA designs is that NVIDIA is using twice as many Vera CPUs to handle four GPUs in its Vera Rubin compute trays.

A good deal of those UALoE links, in turn, are allocated to the scale-up network that connects the GPUs. The rest of the scale-out network, which goes to the AMD Vulcano 800 AI NICs in the system. Each GPU can have up to three such NICs attached to it, though at 600Gb/sec per GPU in scale-out bandwidth, this still pales in comparison to the 3.6Tb/sec of scale-up network bandwidth.

We were told that customers can elect to use either three NICs per GPU or two NICs per GPU, which means four or six NICs per side. While AMD is focused on showing Vulcano, there are designs with OCP NIC 3.0 slots instead, which might be what you want if you are a customer looking for Broadcom NICs.

Each compute tray is also equipped with one of AMD’s Salina 400 DPUs for the front-end network.

Finally, while not on AMD’s official diagram, each compute tray also features five E1.S SSD slots for local storage. Next, let us get into the networking tray.


