ASRock Rack 4U16X-GNR2 Performance
Starting with some basic CPU performance, here is a look at a single socket Intel Xeon 6767P core-to-core latency:

Since this is a dual-socket system, we also have a dual-socket view of the dual 64-core system.

Here is the lmbench view of the memory latency on the CPU side.

With the basic latency checks out of the way, it is time to get to our agentic CPU benchmark, AgentSTH.
AgentSTH V7 Performance
Since single-core performance is dominating headlines these days, I wanted to take a quick look at the Intel Xeon 6767P. We are using a baseline of the Ampere AmpereOne A192-32X.

You can see the difference between larger P-cores and lower-power cores fairly clearly in this comparison, which is a good reason to use it. Also, we used the ASRock Rack AMPONED8-2T BCM Motherboard for the platform.

Taking a quick look at the multi-agent splits on this platform, we can see the impact of running this on a dual-socket server. The increase in performance from one to two instances is splitting the agents onto physical sockets, and the 4x split further helps since there is Hyper-Threading.

On a single socket, here is a quick look at the scaling relative to single-core performance on the Intel Xeon 6767P in the server.

Overall, the Intel Xeon 6767P is a popular option, as 64-core CPUs are common in this class of servers, as they increase throughput. For agentic workloads, they scale fairly well, although most organizations will run agent sandboxes on another server, rather than on the GPU head node.
Given when we were running this, we were still using Kimi K2.5. We might be using GLM 5.2 today, or a newer Kimi model. Still, with the massive NVIDIA Blackwell Ultra HBM3e memory pools, we were able to run Kimi K2.5 on only four GPUs. Better said, we could have vLLM running on two sets of B300 GPUs in the same system.

This clearly shows why having faster memory is so important. As an important piece of context here, the 2x4x GPU numbers were running 2 instances of K2.5, so each instance was getting single-user concurrency. Realistically, you can either double-up instances like we did, or run K2.5 and then run other models alongside it.

As a quick aside here, SGLang is faster at serving Kimi K2.5, but we are using vLLM here because that is what we have used in the past.
Next, let us discuss the power consumption.
ASRock Rack 4U16X-GNR2 Power Consumption
In terms of power, this server has ten 3kW 80Plus Titanium power supplies so that it can run in 5+5 redundancy mode.

Overall power consumption is significant compared to traditional servers as it was using over 2kW at idle. We saw a power closer to 9kW on a workload that, in a recent air-cooled HGX 8-GPU server, we would have expected over 10kW. That makes sense since instead of fans, the heat is being exchanged outside of the chassis.

Next, let us take a special look at the ZutaCore version of this server.


