F5 BNK Performance Impact
If you want to see the actual lab we are using, we toured the racks during the lab visit. We had Supermicro servers with eight H100 GPUs, each serving a Qwen3-32B model in FP8. The BNK control and EPP path ran on the DPU, while the comparison used an Envoy AI Gateway (now Agent Router) running on the host. A few servers also handled cluster storage. Testing used P90 latency over 60-minute runs, with the NVIDIA AI Perf Tool simulating varying GPU workloads to show how load affects the two paths.

We were left with lots of logs. Instead of showing folks those raw logs, and since we are talking about AI management, a bit of AI magic later and those logs turned into a visual display showing the four traffic patterns: No shared prefix, Multi-turn chat, Mixed traffic, and Heavy reuse. Then the concurrency was set to 150 with 10K input tokens per request, 150 with 20K, and 200 x 20K. The last one really stresses the KV storage because you are running at higher concurrency and more input tokens per request. You may be wondering why we are not doing this at c=1, but remember, this is a H100 cluster, and BIG-IP Next for Kubernetes is really designed for clusters instead of single users doing local AI.

Starting with the base case where we are running at relatively low c=150 and 10K input tokens, you can see that the F5 solution and the Envoy AI Gateway are fairly close.

Here is a look at some of the key metrics over the course of the run.

Now let us get to the hard case, which is the other extreme. We are increasing the concurrency by 33% (200 versus 150) and doubling the input tokens. That puts a lot more pressure on the infrastructure to place workloads on the right GPUs. In the base case above, we used only 46% of the cluster’s KV capacity, but now we are at 1.24x, which means we are oversubscribed.

This is where we get the massive 3.24x gain number because F5 is doing a much better job balancing this load.

Here is a look at the runs again across some key metrics:

As you may have noticed, this exercise generated a lot of data running four different workloads across four different concurrency and input sizes. Here are plots using those four scenarios across the three test loads and then in terms of requests completed. You will notice that the lower KV pressure on the right side is generally closer between the two gateways, while the heavier concurrency tends to favor F5 more.

In terms of raw tokens per second, here is a similar view of that:

In terms of mean TTFT, here is what that look like:

Of course, folks want to see the p99 TTFT as well, so here is what that looks like, and again, lower is better:

I think the key takeaway is that, especially at lower loads, the solutions can be relatively close together. Once the load ramps up, the F5 solution performs really well. The 3.24x number is probably closer to the extreme in terms of benefit we saw with this setup, but even if you only get 1.25x the performance from the same GPUs, that is like a buy four GPUs and get one free proposition. Actually, it is better than that because it is not just the cost of the GPU. It is also the all-in cost of running that GPU, including the servers, networking, power, and so forth. At a smaller scale, that may not be exciting, but as infrastructure scales, and also lead times for power and GPU systems increase, using the same infrastructure and getting more work done with it makes a lot of sense.
Final Words
Although we usually focus on hardware, I thought this was a fun opportunity to check out a networking company’s lab as part of our current lab tour series. Orchestration and operationalizing the AI servers we review is also something we get asked about a lot, so I wanted to learn a bit more. It also shows just how much optimization can help utilization. That will be a major topic in the coming years, given that AI is a trillion-plus-dollar endeavor for humanity. Optimizing the build-out, and what is already installed, is something that simply must happen.

Putting F5 BIG-IP for Kubernetes on BlueField-3 DPUs keeps load balancing and L4-to-L7 security off the CPU and frees up GPU cycles an inference cluster needs. To me, seeing a rack with F5 BIG-IP appliances and one running on the DPUs was neat. On our AI data center tours, we often see firewalls, VPNs, and other network appliances in dedicated connectivity racks. They have been a necessary part of shared infrastructure. Now, that runs on the DPU that is already installed in many AI servers. The idea of offloading this security and control plane to the DPUs and freeing up CPU resources is really neat, and is one of the reasons I wanted to do this piece. It is also a cool example of why you want DPUs in servers serving the double need now that server CPUs are harder to come by.
At the same time, I know folks are going to want to know about 20 different models, what about different GPUs, larger clusters, and all of the features we did not really get into in this piece. Also, I have no idea what the pricing is. Still, we wanted to at least show that something like this exists because many do not know that there is an entire industry working on these types of solutions.


