Gigabyte W775-V10-L01 NVIDIA Grace Performance
The NVIDIA Grace CPU has been out for some time, but this is a 300W version where we are seeing clocks in excess of 3.3GHz. We wanted to see how the 72-core Arm Neoverse-V2-based CPU fares, so let us start with the core-to-core latency.

The Grace CPU is a monolithic die, so we get a fairly predictable pattern. Perhaps most useful there is that the core-to-core latency is within a fairly narrow window of around 42ns. Other designs can have triple-digit ns variances between core latencies. We also ran lmbench on the chip.

Again, this is what we are accustomed to seeing for the platform.
Gigabyte W775 AgentSTH V7 Performance
NVIDIA’s flagship agentic AI CPU is soon to be the NVIDIA Vera, but for NVIDIA Grace we still get very good performance. Let us start with some of the per-core performance of workstation CPUs.

One of the major challenges with this comparison is that the Grace portion of the Gigabyte W775 is running at a 300W power limit. That is somewhere in the range of the Intel Xeon 658X (250W, 300W turbo) at 24 cores, and is a bit less than a 64 core/ 128 thread AMD Ryzen Threadripper 9980X. We wanted to see what the performance per core would be of a 72-core design against some workstation chips. These are not the perfect comparison points, but they are at least useful in this generation. SMT helps on a number of tasks (but not all), so the x86 processors do well when we look at single-core performance.
That is only part of the story. Meta Muse is allegedly running on 2 vCPUs, which is roughly the above single x86 core with a two-thread view. But where the Neoverse line does well is scaling. If you are buying a GB300 AI Station, you should be seeing maximum performance. So you may use four cores, or eight cores for an agent VM.

72 cores is a lot, and more than the majority of agentic AI workloads will see. Usually, larger AI shops tell us they are targeting 2-4 cores for AI agents, so the Grace CPU’s architecture actually scales very well. I wanted to swap out the workstation processors for more of lower-power server processors to get the per-core power a bit closer when we look at scaling. Here is what that typical 2-4 core range looks like, but expanded to 1-8 cores and how the scaling happens even in those common bands.

What you can see here is that the NVIDIA Grace Neoverse V2 cores are scaling much better than the Intel Xeon PCIe Gen5 generation bookends with Intel Xeon 6 and 4th Gen Intel Xeon Scalable. AMD is a bit closer on scaling in this range. One could argue that the AMD EPYC 8653P with 84 cores in a 225W TDP is disadvantaged here, but we wanted to get a similar core count in a similar power band. Of the 147 configurations we have tested at this point, this is the closest that we have.
Gigabyte W775 Estimated SPEC CPU2026 -O3 Performance
Given that these are designs that have been on the market for some time, you can look up many of the standard benchmarks which are highly optimized. Vendors are very good at optimizing compiler flags and systems for the official SPEC CPU runs. At the same time, we wanted to offer a different view, getting to a lower optimization point using an -O3 flag instead. Our goal is really to create something that is different from the official results, removing many of the common optimizations. Starting with the SPEC CPU 2026 integer rate base estimated result, here is what we saw:

A quick look here, and you will see why we are using the AMD EPYC 8635P here despite its lower TDP. That AMD EPYC scores about 2.58 per physical core, while the NVIDIA Grace is about 2.71. Swapping to the floating-point side, here is what we saw:

NVIDIA does well here. AMD will probably look at this and think that they will do better with compiler optimizations, higher TDP, and perhaps a different SKU. All are fair points.
We kept this comparison a bit tighter just because, realistically, I think that you do not buy a GB300 for just the CPU, but you may buy a Xeon or EPYC for just the CPU. Instead, what you really are looking for is the performance of the CPU in agentic AI workloads, since that is the purpose of the Gigabyte W775.
On that note, let us get to the GPU.


There’s so much more in here than in the early reviews that made it sound like it’s a normal workstation. I watched and read other GB30 reviews, and I didn’t know about the 229Gbps limit. It’s a prime example of why STH is the best at this today, now that Anand is done.