Gigabyte W775-V10-L01 NVIDIA B300 Performance
Now, it is time for the big show, the NVIDIA Blackwell Ultra or B300.

We generated a ton of data across different models. Originally, we planned to run a few big models, but that quickly led us down the rabbit hole of trying different draft models. We use Qwen3.6-35B-A3B a lot to simply monitor tasks in the studio, so we wanted to see how that performed, and it was shocking.
NVIDIA GB300 on Smaller (but useful) Models
Given we have lots of HBM, we looked at interactivity quite a bit. The reason for this was simple: what if you were running a business with a bunch of monitoring tasks for different sensors, machines, and so forth? Those are tasks that can be time-consuming. Many users may have needs for this type of task, and they are great for agentic workflows.

We charted different input/output lengths, along with different concurrency rates. As you would expect, you get more performance per user at lower concurrency. At the same time, this is a big machine, and we wanted to see what would happen if you used this only for running a swarm of monitoring agents all the time.

You probably saw that data point (and the re-test) on this chart, and it is impressive. We got well over 26,000 tokens-per-second on a short 24-30 input length, 1024 output length, Qwen3.6-35B-A3B NVFP4 flash runs at 512 user concurrency. Even a larger MoE model, Nemotron3-Super-A12B NVFP4/FP8 at c=512, was around 7000 tokens-per-second, and gpt-oss-120b was ripping in terms of performance at higher concurrences. For context, the 26,000T/s result was about 2.25B tokens/day of throughput, which is wild.

If you are a business that has a bunch of people monitoring machines or sensors, watching for blips and actioning issues, or if you just have many agents where you need heartbeats, this is a crazy-fast solution.
NVIDIA GB300 Qwen3.8-27B Performance
Moving to a bigger and newer model, Qwen3.8-27B is a fairly large upgrade from the Qwen3.6-27B model, and so we wanted to see how that would perform, even at lower concurrency.

This is somewhat of a fun, but also important bit about a system like this. We had the option of running in a number of different numeric formats, draft models, and so forth. You can see even at c=4, the throughput can double depending on the setup, and there are plenty of new techniques coming out all the time.
NVIDIA GB300 Nemotron3-Super-120B-A12B Performance
NVIDIA has been pushing hard into the open model space, and its Nemotron3 models are useful and open. But also, the 120B A12B model is quite different from the Qwen models we looked at previously.

Here, I think the interesting part is that a 120B MoE model is very useful. It is not frontier intelligence, but Nemotron3 Super 120B is a notable upgrade over GPT-OSS-120B, which itself set off a lot of very useful workflows. Maybe the key thing that we learned, aside from being able to generate 86.5M tokens/day even at very reasonable concurrency levels, is that the GB300 continues to scale as it is used more. This is very powerful. If you invest in a machine like this, it likely means you believe in agentic AI. Assuming that is the case, if you can offload something with a model today, then chances are a quarter or a few quarters from now a similar-sized model will be better, so you will have new agentic workflows. The Gigabyte W775 has the capacity to add more concurrency, supporting more agents in the future, while having more memory than the smaller 128GB-class unified memory systems. After testing these systems and having the 8xGB10 cluster, I have a firm belief that over time, as models get better, you offload more tasks to machines like the Gigabyte W775.
NVIDIA GB300 GLM-5.3 Flash 321B Performance
The next step was getting much bigger. GLM-5.3 is a 321B parameter MoE model with something like 18B active parameters. When it arrived on the scene, it was being compared to GPT-5.6 Terra and Opus 4.8, so this is a different intelligence class than many of the smaller models.

Looking at throughput versus concurrent users, we can see that even at lower concurrency, we could run the model at decent speeds and increase cumulative performance while still maintaining reasonable TTFTs.

While it is not as quick to generate first tokens as many of the smaller models, it held up fairly well.

As we are writing this article, many of the providers on OpenRouter are delivering GLM-5.3 Flash at 20-40T/s, so often this system is outperforming the cloud providers.
NVIDIA GB300 Deepseek V4 Flash 0731 Performance
Deepseek V4 Flash 0731 is one that we have used quite a bit at STH, as the 284B A13B model is text-only but better at many coding tasks than GLM-5.3 Flash.

Overall, performance was slightly better, and latency was often lower than GLM-5.3 Flash.

There are benefits to both models, but this one can also support a decent number of users on the Gigabyte system.

We are in the world of agentic AI, though, so it was worth letting the systems go beyond the charts and into a real-world coding exercise.
Gigabyte W775-V10-L01 NVIDIA Blackwell Ultra Real-World Coding
In the video, we also showed the output from using GLM-5.3 and Deepseek V4 in an agentic coding exercise. The prompt was to build a kart racing game that could run in a browser without using any existing artwork.

Here, Deepseek was about 20% faster (~100T/s versus ~120T/s) than the GLM-5.3 model. Still, GLM-5.3 ended up being over twice as fast for the main coding tasks, taking roughly 2.1 hours versus around 4.7 hours for the Deepseek V4 model. The post-playtest fix round was virtually identical, but that was an interesting result. Being able to run GLM-5.3 locally meant we could use a model that was slower in tokens/second but much faster in time to completion. Also, and this is subjective since both passed acceptance criteria, we thought the GLM-5.3 game was more fun.
At this point, you may be wondering about networking, so let us get to that next.


There’s so much more in here than in the early reviews that made it sound like it’s a normal workstation. I watched and read other GB30 reviews, and I didn’t know about the 229Gbps limit. It’s a prime example of why STH is the best at this today, now that Anand is done.