Advertisement


Home Server Server CPUs Normalizing NVIDIA Vera Benchmarks to AMD EPYC Turin A Framework

Normalizing NVIDIA Vera Benchmarks to AMD EPYC Turin A Framework

0

Normalizing Turin Performance Based on NVIDIA’s Data

If we were picking AMD parts to compare against an 88-core Vera, the EPYC 9755 that NVIDIA used would not be on the shortlist. The AMD EPYC 9575F is AMD’s 64-core frequency-optimized part, and it is the closest analog to how Vera is tuned. AMD sells the F series specifically in the performance-per-core segment, and those parts run notably higher clock speeds under load. The AMD EPYC 9655 is a 96-core part, which is a far closer match to Vera’s 88 cores than a 128-core flagship. Either of those would have been a highly defensible comparison point. NVIDIA instead selected the one part of the stack that is optimized for neither frequency nor a Vera-class core count.

This matters because performance per core is inversely correlated with core count. As vendors add cores, they lower frequencies to stay inside a socket power budget, so per-core figures fall even as socket throughput rises. Comparing loaded clocks, the EPYC 9755 runs roughly 10-28% below the EPYC 9655 or 9575F, depending on the workload. When the headline metric is performance per core under full load, that clock difference flows almost directly into the result. Perhaps another way to look at it is that an EPYC 9755 is running only 88 cores and using roughly 69% of its capacity. Running on fewer cores means more power is available to allocate to the active 88 cores, allowing them to run at higher clock speeds.

SPEC CPU2026 Configurations NVIDIA Vera Whitepaper
SPEC CPU2026 Configurations NVIDIA Vera Whitepaper

My sense of how NVIDIA came to this comparison point, instead of using an EPYC 9575F, which is AMD’s current-generation, higher-core-count frequency optimized part, or the 96-core EPYC 9655, is that it was able to show similar total socket SPECrate2026_int_base scores at a socket level using its testing methodology instead of using the official AMD results. Typically, Arm CPU vendors utilize gcc as a baseline because gcc is enormously popular in the real world and because that is where Arm and Arm CPU vendors have put their optimization efforts into. This is similar to the Cavium ThunderX (1) days in 2016 and is true a decade later.

The challenge with this methodology is that it reports lower results for Intel Xeon and AMD EPYC CPUs that have other optimization paths. Granted, Intel has an entire series of CPU2017 results with footnotes due to the aggressiveness of its compiler optimizations in previous versions. Debating gcc versus other compilers is an almost religious debate in the industry. From a logical standpoint, using a least common denominator means another test metric is standardized in the comparison. The other camp would argue that the benchmark submissions are highly optimized by all vendors and then reviewed in their submission process for official results. Personally, I have said for years that I think both methodologies have merit. Since SPEC CPU benchmarks are widely used for RFP criteria, the official results are probably the most impactful. Said another way, the official figures are the real scores, while other scores are estimates that may be the result of a valid methodology.

Arm vendors generally suggest that today, agentic workloads are likely to compile code using gcc. I think that statement is true. We use a lot of gcc. On the other hand, if there are real performance gains, a good agentic AI workflow in the future would use LLVM/ clang if there were benefits. Something the industry will have to wrestle with is whether we expect AI agents to consider the performance impacts of compilers in the future. We are not there yet, but it is worth considering, since AI agents are proving useful for performance optimization.

SPEC CPU2026 Per Socket With Official Intel And AMD Results And NVIDIA Vera Whitepaper Figures
SPEC CPU2026 Per Socket With Official Intel And AMD Results And NVIDIA Vera Whitepaper Figures

The impact of eschewing the official submitted and reviewed benchmark figures for the AMD EPYC 9755 and using an alternate methodology to estimate a different result is that NVIDIA’s figure for the EPYC is about 16% off the official result. Just a quick note: the CPU2026 data shows that 2P results tend to be better than twice the 1P results across many configurations. We will see if that holds as the benchmark gets more results. Also, the Intel Xeon Clearwater Forest does well in throughput, but it is a different class of per-core performance. One reason we are not hearing more about the Xeon 6990E+ is likely that the delayed 2026 line is not beating the 2024 AMD EPYC 9965 here.

If you then look at NVIDIA’s results as its highest thus far, which, if submitted and approved, could be their number for Vera, against the official EPYC results, some portions where it is showing 20% per core gains, the world looks a bit different.

SPEC CPU2026 Per Core With Official Intel And AMD Results And NVIDIA Vera Whitepaper Figures
SPEC CPU2026 Per Core With Official Intel And AMD Results And NVIDIA Vera Whitepaper Figures

According to official SPEC CPU2026 data, the EPYC 9575F lands at roughly 4.8-5.0 points per core, compared with roughly 5.3 points per core in NVIDIA’s Vera estimate. That is still a Vera win, but it is a 6-10% class win against the parts AMD currently sells, optimized for per-core rather than per-socket performance, not the 1.5x and larger figures the whitepaper charts suggest.

Reframing the results with the official EPYC numbers, where we have additional data, we can then look at a more detailed comparison.

SPEC CPU2026 Subtests 2P Chart NVIDIA Vera Whitepaper And AMD EPYC Official
SPEC CPU2026 Subtests 2P Chart NVIDIA Vera Whitepaper And AMD EPYC Official

Here are the 2P results details, swapping in official results instead of NVIDIA’s estimates at a system level, broken down by subtests.

SPEC CPU2026 Subtests 2P Table NVIDIA Vera Whitepaper And AMD EPYC Official
SPEC CPU2026 Subtests 2P Table NVIDIA Vera Whitepaper And AMD EPYC Official

At a system level, if we use official EPYC results rather than NVIDIA’s best estimate of its own performance, the 2024-era 128-core processor is notching many wins, despite an older memory subsystem.

The 96-core AMD EPYC 9655 does not have an official 2P result, but does have an official 1P result. If we wanted to add that in, here is what the table would look like per socket.

SPEC CPU2026 Subtests Per Socket Table NVIDIA Vera Whitepaper And AMD EPYC Official
SPEC CPU2026 Subtests Per Socket Table NVIDIA Vera Whitepaper And AMD EPYC Official

If we add that 1P result, normalize everything to per-socket, then down to per-core, we get a really neat, detailed view of what NVIDIA presented versus a portfolio of AMD EPYC Turin. We are removing the high-core-count Turin 9965 and Clearwater Forest, since they are highly tuned for socket throughput rather than per-core performance.

SPEC CPU2026 Subtests Per Core Chart NVIDIA Vera Whitepaper And AMD EPYC Official
SPEC CPU2026 Subtests Per Core Chart NVIDIA Vera Whitepaper And AMD EPYC Official

Here is an interesting one: AMD’s 2024 frequency-optimized Turin can notch wins against the 2026 NVIDIA Vera in some limited cases. Note that we are using the core-for-AMD equals two-threads methodology, since that is how the base numbers are run.

SPEC CPU2026 Subtests Per Core Table NVIDIA Vera Whitepaper And AMD EPYC Official
SPEC CPU2026 Subtests Per Core Table NVIDIA Vera Whitepaper And AMD EPYC Official

Vera still leads in most tests, though not all, and the important context here is that it is a 2026 CPU with a 2026 memory subsystem being compared to a 2024 CPU with a 2024 memory subsystem. Again, the more relevant comparisons will be against Venice, the Arm AGI CPU, and Diamond Rapids when it launches.

Final Words

I have been through every major CPU launch since 2010 and seen how different companies market chips. NVIDIA followed a common framework for how an Arm-based CPU has a generational advantage over a current-generation part, especially on the memory side. At 88 cores of NVIDIA Vera to 88-96 cores of AMD EPYC Turin, NVIDIA shows why it is a next-generation part. At the same time, CPUs in the data center are used for many different tasks. AMD has not just the 88-96-core class of chips, but also much larger 192-core parts like the EPYC 9965, which is designed for a different workload-optimization point.

Hopefully, this helps you move from marketing comparisons to a more balanced evaluation framework as we get into the next generation of PCIe Gen6 CPUs with dramatically more memory bandwidth. I will also just add, if you get excited about comparisons, stay tuned over the next few days, months, and quarters.

LEAVE A REPLY

Please enter your comment!
Please enter your name here

This site uses Akismet to reduce spam. Learn how your comment data is processed.