Advertisement


Home AI Diving Deeper on NVIDIA’s Vera CPU: New Architectural Details and SPEC CPU...

Diving Deeper on NVIDIA’s Vera CPU: New Architectural Details and SPEC CPU 2026 Benchmarks

5

Vera Performance: Favoring Single-Threaded Performance Over Total Throughput

All of that architecture chatter aside, what does this mean for the performance of NVIDIA’s new CPU and its competitive positioning? As we noted at the start of the article, NVIDIA has released a selection of SPEC CPU 2026 benchmarks, so we can finally begin to answer that question with some industry-standard data.

With that said, before we dive into the performance figures it is important to note that like other vendor CPU data reveals, we are looking at cherry-picked data. Everyone wants to put their best foot forward, and NVIDIA is no different in that regard. This shows for both their selection of benchmarks – NVIDIA has slides for single-threaded performance, but not multi-threaded throughput – as well as NVIDIA’s point of competitive comparison: AMD’s EPYC 9755 “Turin.”

As Patrick was very quick to point out after seeing NVIDIA’s data, The EPYC 9755 is not a high single-threaded performance part. In the AMD processor hierarchy, those are their F-suffix parts. Instead, the 9755 is AMD’s highest core count Turin part (not to be confused with Turin Dense) that is designed for balanced performance, offering 128 Zen 5 CPU cores at a modest boost clockspeed of 4.1GHz to keep the complete chip within a manageable 500W TDP. From a silicon standpoint it is AMD’s best (or at least, most expensive) Turin offering, but it is not a great representation of what the architecture can do when not held back on clockspeeds, where EPYC F parts can clock as high as 5.0GHz.  Nor is it the best example of AMD’s overall multi-threaded throughput, however, since that is where the 192 core Turin Dense parts shine.

NVIDIA Vera CPU Die Uncapped
NVIDIA Vera CPU Die Uncapped

The end result is that NVIDIA’s benchmarks and associated competitive comparison offer us some valuable insight into how Vera performs, but it does not offer a great comparison to rival hardware. It is still apples-to-apples, but perhaps comparing red apples to green apples. In any case, it also gives us an idea of who NVIDIA expects to be their chief rival in the CPU space for the Vera generation, with AMD’s EPYC earning that distinction.

For today’s release, NVIDIA has published a full set of SPEC CPU 2026 integer performance figures in their Vera whitepaper, as well as a limited set of slides. These performance figures are officially classified as estimates as they come from a dual socket pre-production system.

NVIDIA Vera SPEC CPU 2026 Integer Rate Performance (1T)
NVIDIA Vera SPEC CPU 2026 Integer Rate Performance (1T)

Keeping in mind the relatively low clockspeed of the EPYC 9755 processor here, it is still a strong showing for Vera. Even accounting for clockspeed differences (e.g. an EPYC F chip running at 5.0GHz, about 20% higher), these figures put Vera in the lead in every single SPEC CPU 2026 integer benchmark for single-threaded workloads. NVIDIA said they designed the Olympus CPU core to be an IPC monster, and while we do not have the complete picture quite yet (NVIDIA has not disclosed clockspeeds), the results are certainly pointing in that direction with these single-threaded results.

There are a couple of other graphs from NVIDIA’s whitepaper that also show a strong single-threaded architectural efficiency for Olympus. Taking a subset of the SPEC CPU 2026 benchmarks (where NVIDIA has the largest lead over AMD), the company is showcasing anywhere between a 1.9x and 2.4x increase in instruction fetch operations per cycle.

NVIDIA Vera Instruciton Fetch Ops Per Cycle
NVIDIA Vera Instruciton Fetch Ops Per Cycle

And on the backend of things, NVIDIA is claiming anywhere between a 2.3x and 4.3x increase in backend operations per cycle on those same tests.

NVIDIA Vera Backend Ops Per Cycle
NVIDIA Vera Backend Ops Per Cycle

Finally, I thought this set of benchmarks for graph traversal core scaling with Google’s PageRank algorithm was quite interesting given NVIDIA’s focus on optimizing memory prefetching for graph-like data structures.

NVIDIA Vera Graph Traversal Core Scaling
NVIDIA Vera Graph Traversal Core Scaling

That benchmark set has Vera reaching up to 29.3x its baseline (single core) performance with 32 CPU cores, whereas the AMD EPYC comparison system sputters out at a bit over a 10x increase.

What is not said – and where the rails start getting wobbly for NVIDIA – is total chip throughput rather than single-core performance. NVIDIA’s slide deck does not call attention to the issue, and even their whitepaper primarily focuses on single-threaded or modestly-threaded workloads. It is only at the very end of the whitepaper, when NVIDIA is outlining their SPEC CPU 2026 integer results, that they talk about total chip throughput. And the results are not as rosy.

NVIDIA Vera SPEC CPU 2026 Integer Rate Performance (nT)
NVIDIA Vera SPEC CPU 2026 Integer Rate Performance (nT)

As per NVIDIA’s numbers, the dual-socket Vera system achieves a SPECrate2026_int_base score of 925, versus 898 for the AMD EPYC 9755 system. To be sure, this does put NVIDIA in the lead in their carefully constructed competitive comparison, but not by much: Vera is only beating the Turin chip by 3% here.

And if we go spelunking into the official SPEC CPU 2026 integer rate results to get a better idea of how Vera compares to other published systems, we find multiple dual-socket systems that are ahead of Vera, including EPYC 9755 systems hitting over 1000 points using AMD’s own compiler. Meanwhile there are EPYC 99×5 (Turin Dense) systems scoring over 1200 points in the same scenario.

NVIDIA has also not published any dedicated floating point benchmark data for Vera at this time. So we do not have a good baseline to see how it performs there, even in vector-friendly workloads.

All of which is to say that, based on these initial numbers from NVIDIA, it is fair to say that NVIDIA’s early characterization of Vera is correct: this is a CPU and architecture that is optimized for high CPU core performance. But the converse of that is that the focus on high IPCs and lower core counts means that this is not a chip that is optimized for high overall throughput. NVIDIA built a chip that is designed to go toe-to-toe with rival P-core chips – and even then, many of them feature a larger number of CPU cores.

5 COMMENTS

  1. It looks to me like someone set the colour scale on the core to core latency table so the entire table was green. Even so, the chip seems well designed for tightly coupled parallel workloads. Since AI is different than the cloud-style micro-service throughput targeted by existing processors, I think it makes sense for Nvidia to fabricate their own.

  2. The first diagram shows only 16 lanes of PCIe 6. Surely that’s not right. How could a server CPU in 2026 ship with fewer lanes than a desktop CPU, even if they are a newer generation?

  3. It should be noted that if Vera had been compared to the 9575F for the single-core comparisons, which you suggested would have been more appropriate, then Vera would have held a sizable advantage in the estimated SPECrate2026_int_base score. More importantly, Nvidia doesn’t seem to intend this CPU to be a general purpose CPU to go head-to-head with established processors in most data center workloads. They are targeting it for AI servers, and in particular agentic AI servers. Given how much Vera is ahead in various single core performance metrics, and considering the normal generation-on-generation IPC uplift, and that other CPU manufacturers are unlikely to have engineered a special single core monster for this upcoming generation, it seems very likely that Vera will keep an advantage for these single core performance workloads even against the coming generation of server chips. The whole reason the future CPU server TAM has recently been revised sharply upward is because of the agentic AI workload. Nvidia’s Vera CPU will live and die by how it performs against its coming-generation CPU competition in agentic AI workloads.

  4. The microarchitecture block diagram contains multiple errors (LLM-esque). If it comes directly from nvidia (judging by the color scheme), then one has to question the accuracy of other information provided…

  5. NVIDIA’s latest disclosure unfortunately does not go into any further detail on the branch predictor, but it does give us our first look at the instruction fetch unit it feeds, as well as the path into the 10-wide decoder. In short, Olympus’s instruction fetch unit can feed as many as 16 instructions – 64 bits each – into the decode queue. The queue can hold 48 instructions altogether, and can spit out up to 10 fused instructions to be consumed by the actual decoder.

    I thought ARM ISA was 4-bytes (32-bits instruction, fixed length).
    add x14,x15,x16 (4-bytes)
    mul w3,w2,w4 (4-bytes)

    Is that a custom ARM ISA that Nvidia is using?

LEAVE A REPLY

Please enter your comment!
Please enter your name here

This site uses Akismet to reduce spam. Learn how your comment data is processed.