We now have the first benchmarks for the NVIDIA Vera CPUs, as the company released its architectural paper with them. Ryan covered the release in Diving Deeper on NVIDIA’s Vera CPU: New Architectural Details and SPEC CPU 2026 Benchmarks. From the figures released so far, NVIDIA Vera appears to be an interesting Arm-compatible server CPU with its custom Olympus cores. It took a few minutes after NVIDIA published, and then I started fielding questions like how NVIDIA’s multiples of the AMD EPYC speed. Like virtually every company that produces products, NVIDIA crafted favorable comparisons. I wanted to provide a more normalized framework for the Vera to Turin comparison so folks who have not been doing this for a decade and a half or more can have a sense of where we would expect Vera and Turin to fall on a more normalized basis.
A few quick disclosures that we work with both AMD and NVIDIA. Also, we published a more extensive version of the below on our paid Substack earlier, aimed at a different audience.
I still wanted to provide a framework to our STH readers since we have attempted to provide a more balanced view over the years.
Normalizing the NVIDIA Vera Whitepaper Numbers to AMD EPYC Turin
Since NVIDIA dropped its Vera whitepaper just before AMD’s Advancing AI 2026, where we expect to learn more about AMD’s next-generation “Venice” CPU. Companies typically do this because the valid comparison is versus the outgoing generation rather than the concurrent generation. Since we are still under embargo for AMD’s Venice, I wanted to provide a framework to help our readers create a useful comparison framework for NVIDIA’s Vera results. The key points I want to cover are:
- Socket and per-core memory bandwidth gains
- Core-to-Core Latency, Chiplet Construction, and Glue
- Important Cores and SMT Threads
- Then we will have a big discussion on normalizing the SPEC CPU2026 performance, and a framework for evaluating AMD EPYC Turin versus NVIDIA Vera through a fairer lens.
These are all framing techniques we have seen in the industry before, so this may not be ground-breaking for many STH readers. At the same time, we want to help folks get more context about what was presented, why, and how to get to an alternate comparison point.
Socket memory bandwidth and per-core memory bandwidth
Memory bandwidth is mostly simple math. It is the number of channels per socket multiplied by the speed of those channels, and faster memory arrives over time. That means a large portion of any bandwidth comparison between CPUs from different years is really a comparison of memory vintages. Twelve channels of DDR5-6400 on the 2024-era EPYC 9755 deliver roughly 0.6TB/s of peak bandwidth. Vera’s LPDDR5X-9600 is up to 1.2TB/s. AMD has already said Venice is 1.6TB/s of memory bandwidth, and we have already shown 16-channel Venice systems. Today, the comparison point is still Turin. Here is NVIDIA’s comparison chart from its Vera whitepaper:

The whitepaper’s memory bandwidth per core figure of 12.7GB/s versus 3.1GB/s seems shocking and significantly more than the raw bandwidth figures. One reason is the number of cores in the EPYC 9755. It divides newer, faster memory by 88 cores and older, slower memory by 128 cores. NVIDIA is using a 45% larger denominator on AMD’s side than on its own Vera CPU. A 96-core EPYC 9655 would have been the more relevant denominator today simply because it is closer to NVIDIA’s 88 cores. Changing the denominator would change the ratio, but the 2026 Vera has a newer memory subsystem than the 2024 Turin, giving it a bigger numerator. That is to be expected, and AMD has shown this type of generational transition before.

NVIDIA’s chart may look shocking, but AMD has shown something similar when increasing the number of memory channels and new memory speeds that yield significant gains. Here is the AMD slide on the EPYC 7003 “Milan” (8ch DDR4-3200) to EPYC 9004 Genoa (12ch DDR5-4800). Whenever there is a big channel and memory speed bump, often accompanying a PCIe generation increment, the memory bandwidth numerator also jumps, causing the large jump on these charts. Once you change the 128-core denominator to 88 or 96, the chart looks a lot more like a Milan-to-Genoa route.
I think AMD would concede that Vera has more memory bandwidth and per-core bandwidth than Turin. It would also say this is an area that its 2026 Venice part will see a big jump, like Milan to Genoa, and that will dramatically change the comparison chart.
Core-to-Core Latency, Chiplet Construction, and Glue
Part of the whitepaper felt a bit like when I was sitting front row for a previous glued-together comment many years ago. The whitepaper’s core-to-core latency heat map shows AMD mostly red outside the CCX and CCD structures, with Vera mostly green. Directionally, that is accurate. Chiplets add hops, and hops add latency. The question is whether the red cells matter for the workload Vera targets.

Generally, agentic workloads run in VMs, containers, or sandboxes that use 1-4 cores each. Those fit entirely within a single EPYC CCX so long as the operating system schedules them properly, which mainstream schedulers and container runtimes do as a matter of course. A workload that stays inside a CCX lives in the green cells of AMD’s own chart, and a single core sandbox is barely exposed to core-to-core latency at all. It is true that if a scheduler fragments a 4-core workload by placing threads across the CPU, a large, monolithic die is a safer topology. That is a legitimate Vera advantage in poorly tuned environments.
The conveniently incomplete part of the story sits on the other side. By picking a 128-core comparison point, NVIDIA set up a situation where a like-for-like 128-core test on Vera would need cores 89 through 128 to come from the second socket. Those 40 cores are not just across a chiplet boundary on a package. They are across socket-to-socket links, and we would expect much higher latency and jitter on those links simply because the signals need to travel farther. The monolithic latency story holds through 88 cores and inverts beyond them. Going a step further, beyond 176 cores, the hop is not just within a node. Cores 177 to 256 are now a PCIe hop to a NIC, a NIC hop to a switch, a switch hop to a NIC, then a PCIe hop to a CPU away.
At some point, if you think single-thread dominates, then core-to-core matters less. If you think core-to-core matters more (beyond 8 cores), you start wanting an architecture with more cores. On the subject of cores, we need to talk about cores and SMT threads.
Cores and SMT threads
When a vendor publishes single-thread or single-core results for an SMT-capable CPU, the two terms are used intentionally, and they do not mean the same thing.
- A single-threaded result on an EPYC part is one hardware thread of two
- A single-core result is the full core, with both threads contributing to what is measured
On older CPUs, the difference was modest. On newer parts with better SMT implementations, it is significant. In our testing, loading the second thread adds roughly 12% on a Skylake-era Xeon (with default side channel mitigations active), 23-25% on Sapphire Rapids, 29-32% on Xeon 6, and 32-33% on Zen 5, with Zen 5c a few points lower because it runs into L3 cache misses more often. This is a preview of something we will discuss later this week in what we hope will be our next video. Using a single core with SMT often gives 30% better performance than the core alone.

So if one core is 1.0, one core with its SMT thread is 1.3, and one thread is 0.65. In the past, we have seen non-SMT vendors use the “per thread” metric because of this effect. As we move into a world of more diverse CPU offerings, this is something to always keep in mind.
The big lever, however, is really how comparisons are crafted. Here, we wanted to provide a more detailed breakdown of NVIDIA’s comparison points and offer guidance on how to frame the Vera to Turin comparison more neutrally.



