Advertisement


Home Server Accelerators AMD Instinct MI455X Deep Dive: CDNA 5 Marks The Next Era of...

AMD Instinct MI455X Deep Dive: CDNA 5 Marks The Next Era of Instinct

0

The Backbone for Scaling: Much More Interconnect Bandwidth with UALink

Despite AMD’s major architectural overhaul of the core compute architecture for MI455X/CDNA 5, there is a very good argument to be had that it is not even the most significant upgrade to AMD’s overall GPU architecture. Rather, that goes to the off-chip I/O capabilities of the chip, which have been massively increased.

Owing in large part to AMD’s need to support larger scale-up domains, AMD has almost completely tossed all of the old I/O capabilities of the MI3xx generations in favor of a new I/O subsystem that can meet AMD’s bandwidth needs while supporting the latest in I/O and networking technologies. And all of this functionality has been placed in its own dedicated silicon: the new IOD.

With regard to features, the MI455X brings support for a bevy of new I/O technologies. The big news here is that AMD has added support for Ultra Accelerator Link, which is the open industry standard for high-bandwidth, cache-coherent connections between accelerators for scale-up networking. In a nutshell, it is the rest of the major industry players’ answer to NVIDIA’s NVLink.

AMD AAI 2026 Helios Architecture Scale Up Topology
AMD AAI 2026 Helios Architecture Scale Up Topology

UAL support, in turn, forms the backbone of MI455X’s connectivity in two ways. First, UAL over Ethernet is the basis of all of the scale-up networking in the system, with AMD’s IODs able to provide complete UALoE connectivity on their own. Meanwhile, UAL is used to provide connectivity to the Pensando NICs used for scale-out networking. In principle, the use of UAL could also allow vendors to eventually mix and match accelerators, and while no one is there at this time, AMD is doing the next best thing by using them to connect to NICs in Helios.

CDNA5 Architecture MI455X Link Bandwidth
CDNA5 Architecture MI455X Link Bandwidth

Alongside UAL, the new IODs also bring support for both a new generation of AMD’s xGMI link, as well as PCIe. On the former, xGMI, which was previously used for GPU-to-GPU connectivity in the MI3xx generations, is now being used for CPU-to-GPU connectivity instead. This marks the first time that AMD has supported connecting a discrete Instinct accelerator to a discrete CPU in this fashion (and the first time a CPU has supported the reciprocal connection), with the xGMI link effectively functioning as a dedicated PCIe Gen6 connection with cache coherency. For all other needs, the IODs can support plain PCIe Gen6 as well, an upgrade over the PCIe Gen5 support in MI350.

By the numbers, all of these links amount to a massive amount of bandwidth. For the UALoE links, AMD has the chip organized as 36 links, each of which is implemented as two 200GbE lanes. This makes for a total bandwidth of 400Gbps per UALoE link, or an aggregate chip bandwidth of 3.6TB/second (bi-directional). This is around 3.4x the amount of GPU-to-GPU bandwidth as MI350, and spread out over a far larger number of links to support a larger networking mesh.

While UAL and PCIe connectivity is a bit more interesting, as there are multiple configurations available. If an MI455X is configured to use an all-UALink configuration, as is the case with Helios, then an IOD can be configured with three x8 UALinks, each operating at 128Gbps/lane (twice the rate of PCIe Gen6). This gives MI455X 128GB/second of bandwidth per UALink, or an aggregate UALink bandwidth of 384GB/second. Otherwise, if the IOD is configured for PCIe Gen6, then it functions as a pair of x16 links (each lane running at 64Gbps) for an aggregate bandwidth of 256GB/second.

Either mode represents a significant improvement over the amount of scale-out bandwidth available for the Instinct series. Whereas MI350 had just one 400Gbps NIC, MI455X can drive three 800Gbps NICs, for six times the scale-out bandwidth.

Finally, the xGMI link, which is an x16 link running at a data rate of 64Gbps, offers a further 128GB/second of bandwidth between the GPU and the host CPU.

CDNA5 Architecture UALoE Scale Up Networking
CDNA5 Architecture UALoE Scale Up Networking

Taken altogether, this means that each MI455X chip offers just over 2.3TB/second of bandwidth going in each direction. Or in AMD’s preferred bidirectional figures, about 4.6TB/second of bandwidth.

To be sure, this is still a drop in the bucket compared to the amount of HBM4 bandwidth at hand, never mind the internal Infinity Fabric itself. But it is multiple TB/second more bandwidth than was available on MI355X. It is not too big of a stretch to say that one of the major innovations of MI455X is that AMD has been able to stuff several dozen network controllers into each chip.

LEAVE A REPLY

Please enter your comment!
Please enter your name here

This site uses Akismet to reduce spam. Learn how your comment data is processed.