Advertisement


Home Server Server CPUs AMD Takes the Lid off of Next-Gen EPYC 9006 Venice As Zen...

AMD Takes the Lid off of Next-Gen EPYC 9006 Venice As Zen 6 Comes to Servers

1
AMD EPYC 9006 SP7 Hero Image
AMD EPYC 9006 SP7 Hero Image

AMD has been on a run of enterprise product announcements this summer. Building off of the company’s momentum over the past couple of years in not only the AI space but the traditional server space, AMD has been preparing a whole slate of new products for the fall.

So far, we have looked at their announcements for the Instinct MI450 series of accelerators. As well as the Helios rackscale system that those accelerators will go in. But today it is time to talk about the most fundamental component of AMD’s enterprise hardware stack: CPUs.

EPYC has been at the heart of AMD’s resurgence in the server and data center market over the last decade. After losing all momentum in the market, AMD has slowly but surely turned things around, continuing to pick up an ever-larger share of the server CPU market. With AMD’s server revenue now handily eclipsing its consumer revenue, the EPYC CPU is a central part of AMD’s product offerings.

With the launch of the new EPYC 9006 processors quickly approaching, AMD has a lot to say about the upcoming chips. While stopping short of a deep dive into the Zen 6 architecture itself, AMD has laid out the product stack for the next generation of EPYC CPUs in great detail through its Advancing AI 2026 event and a recent whitepaper. With four major chips, AMD is planning its biggest EPYC lineup yet.

AMD AAI EPYC Overview Introducing Venice
AMD AAI EPYC Overview: Introducing Venice

With the EPYC 9006 family, AMD is giving the entire family of chips a top-to-bottom overhaul. Along with the latest and greatest Zen 6 CPU architecture at the core of the CPUs, AMD is bringing to market a new I/O die with vastly faster memory support, new sockets for everything, the return of 3D stacked L3 cache, and later in 2027, even an AI-focused chip with LPDDR support, a first for an AMD EPYC processor. It is a significant technical upgrade that, by the end, will dwarf the scale of AMD’s previous releases.

These upgrades and improvements are intended to help AMD maintain momentum in the server market and meet the needs of the burgeoning AI market. With a particular emphasis on AMD’s first-generation Helios rackscale system, AMD is looking to give customers the hardware and features they need to deploy EPYC CPUs more widely than ever before, from classic servers right up to the latest and greatest in AI systems.

There is a lot to cover, so let us dive into the EPYC 9006 hardware.

EPYC 9006 Platform: Zen 6 For Servers

Starting with AMD’s Zen 6 silicon plans, the company will once again produce two types of CCD chiplets for its server processors. As with the Zen 5-based Turin family, this will include a high-performance Zen 6 chiplet and a high-density Zen 6c chiplet. At this point, AMD has not gone into depth on the Zen 6 architecture, so we do not have significant details on whether and how Zen 6 and Zen 6c differ, but for now we are assuming that Zen 6c follows the same archetype of being a denser core configuration that in turn will not reach the same clock speeds as the regular Zen 6 cores.

AMD AAI EPYC Overview Generational Improvements
AMD AAI EPYC Overview Generational Improvements

AMD is producing both CCD variants on the same process node, which is TSMC 2nm. Unlike the Turin generation, the two CCDs share the same process node, which should give us a more apples-to-apples comparison between Zen 6 and Zen 6c in the long run.

One interesting aspect of the Zen 6 generation is that there will be a much wider gap between Zen 6 and Zen 6c in core counts. Thanks to the density improvements of TSMC 2nm, AMD is going to be able to load up dense Venice chips with up to 256 CPU cores, which is 64 more (or 33% more) than 192-core Turin Dense chips. At the other end of the spectrum, the highest-core-count Zen 6 chip on the disclosures thus far will feature 96 CPU cores, down from 128 cores in the Turin generation. Consequently, dense Venice chip configurations can now have up 2.66x as many CPU cores as the high-performance configurations, a significantly larger advantage than the 1.5x core count advantage that dense Turin chips brought.

AMD AAI Venice SP7 Delidded
AMD AAI Venice SP7 Delidded

Meanwhile, as AMD has already published SKU lists for the near-term EPYC 9006 family, we have a very good idea of where clock speeds will land. The fastest Venice Zen 6 chips will top out at 5.0GHz (the same as Turin), while the dense Zen 6c chips will top out at 4.1GHz, which is some 400MHz faster than where Turin Dense chips topped out. As a result, not only is the core-count gap growing between regular and dense chips, but the clock-speed gap is shrinking at the same time.

Speaking of core counts, at this point AMD is not directly disclosing finer details such as the number of CPU cores in each CCD. But the company has been readily showing off an assembled Venice Zen 6 chip, which sports 8 CCDs and 2 IODs. From that we can reasonably infer that Zen 6 CCDs (for Venice, at any rate) will feature 12 CPU cores per CCD, which is up from 8 cores per CCD in the Zen 5 generation.

Beyond packing in more cores and higher IPCs, AMD is also leveraging transistor gains from TSMC’s 2nm process by adding more L3 cache to some CPUs. With the Zen 6 architecture family, every CCD is paired with 4MB of L3 cache per core, and this applies to both Zen 6 and Zen 6c chiplets. This means high-performance EPYC 9006 processors will come with up to 384MB of L3 cache, while dense chips can ship with up to 1024 MB. Compared to Zen 5, this doubles the L3 cache per core in AMD’s dense chiplets while maintaining the status quo on regular chiplets.

New I/O Die Brings MRDIMMs, PCIe Gen6, and More

CPU cores aside, the other major improvements for the EPYC 9006 generation come courtesy of AMD’s new IOD. AMD has invested heavily in memory support for the Zen 6 generation, and that all comes to roost on Venice chips, which will support up to 16 memory channels, up from 12 on Turin. Equally significant, Venice will support much higher memory frequencies, with DDR5 RDIMMs up to 8000 MT/sec, while MRDIMM support lets Venice push memory bus frequencies as high as 12,800 MT/sec with speedy (and spendy) multi-ranked DIMMs. This significantly boosts the platform’s total memory bandwidth, allowing Venice to offer up to 1TB/second with DDR5 RDIMMs and 1.6TB/second with MRDIMMs.

Samsung Multiplexer Rank DIMM Gen2 Close Up
Samsung Multiplexer Rank DIMM Gen2 Close Up

The new IODs also bring support for the latest generations of PCIe and CXL technologies. Specifically, Venice supports both PCIe 6 and its related Compute Express Link offshoot, CXL 3.1, doubling I/O bandwidth over the Turin generation. In a 1P configuration, a single Venice chip can support up to 128 lanes of PCIe/CXL.

Meanwhile, in a factoid largely underplayed by AMD, Venice’s IOD also brings with it a major improvement to the chip’s xGMI interconnect. Better known as AMD’s Infinity Fabric-based high-speed chip-to-chip interconnect, AMD has used various iterations of xGMI over the years to provide a cache-coherent link between multiple GPU configurations or multiple CPU configurations.

With Venice (and going hand-in-hand with MI455X), AMD’s CPUs and GPUs can now speak the same flavor of xGMI, thereby allowing AMD’s CPUs to share in the cache-coherent memory domain. Prior to this, AMD could only link up its CPUs and GPUs via PCIe, which had similar bandwidth as xGMI for a given generation (e.g., 32 Gbps per lane for EPYC 9005/MI355X) but without that cache coherency. The net result is that Venice can now couple much more tightly with AMD’s GPUs than any previous EPYC generation. This has been something of a long-term goal for AMD, and the addition of cache coherency should pay dividends for both system performance, including letting the GPUs have a coherent window into the CPU’s massive memory pool, as well as simplifying data movement and synchronization for programmers.

AMD AAI Venice SP7 Rear
AMD AAI Venice SP7 Rear – A Lot of Pads for a lot of I/O and Memory Bandwidth

Features aside, the new IODs’ hardware design is also particularly notable here, with an emphasis on the plural. For high-end EPYC 9006 chips, AMD is now using two IODs on the chip rather than a single IOD. At the moment, AMD is not calling much attention to this design choice (this being another aspect we expect the eventual architectural deep-dive to better cover), so it is unclear what the full ramifications of this split are for chip performance and how NUMA domains will function. Meanwhile, it gives AMD a potential route to offer cut-down EPYC chips with fewer of its existing IODs, rather than needing to produce multiple distinct versions of the IOD.

Somewhat surprisingly, AMD has also kept things fairly conservative on the silicon side with the Venice IOD. The chiplet is produced on TSMC’s 6nm process, the same process node as Turin’s IOD, meaning that the improvements to AMD’s IOD in this generation are all on the design side of matters, and AMD is not getting a performance or density boost from a newer process node. One thing it gets, and an important one in today’s environment, is the ability to produce part of the chip that does not benefit as much from a newer process node on an older process node. TSMC largely controls wafer allocations and thus who can make chips in a supply-constrained environment. Using a more mature process node for the IODs and a newer process node for the CCDs likely helps AMD produce more CPUs.

All of this extra CPU performance, memory bandwidth, and I/O will come at a power cost, however. 192- and 256-core Venice chips will top out at a TDP of 600W, a full 100W higher than Turin’s highest TDP. Things will be a bit more restrained for the high-frequency Zen 6 chips, but those will still draw as much as 500W, which is higher on a watts-per-core basis than Turin chips. Venice of all flavors should still deliver a significant uplift in performance-per-watt compared to Turin, but it is clear server processor TDPs will continue their slow creep higher in the EPYC 9006 generation.

1 COMMENT

  1. “AMD will be adding AI compute extensions, which they are calling ACE. Presumably, these will be something akin to Intel’s AMX or Arm’s SME”

    I’m surprised at this statement. We know exactly what ACE is since its whitepaper is half a year old, and compiler enablement is well underway for both GCC and LLVM. It could be a nice article to explain it here though.

LEAVE A REPLY

Please enter your comment!
Please enter your name here

This site uses Akismet to reduce spam. Learn how your comment data is processed.