Advertisement


Home AI Touring the F5 BIG-IP Next for Kubernetes Lab to Make AI Clusters...

Touring the F5 BIG-IP Next for Kubernetes Lab to Make AI Clusters More Efficient

0

F5 BNK Performance Impact

If you want to see the actual lab we are using, we toured the racks during the lab visit. We had Supermicro servers with eight H100 GPUs, each serving a Qwen3-32B model in FP8. The BNK control and EPP path ran on the DPU, while the comparison used an Envoy AI Gateway (now Agent Router) running on the host. A few servers also handled cluster storage. Testing used P90 latency over 60-minute runs, with the NVIDIA AI Perf Tool simulating varying GPU workloads to show how load affects the two paths.

AI Lab Demo - F5 BIG-IP on NVIDIA DPU - Copy Slide 8: F5 BNK lab setup and test methodology
F5 BNK lab setup and test methodology

We were left with lots of logs. Instead of showing folks those raw logs, and since we are talking about AI management, a bit of AI magic later and those logs turned into a visual display showing the four traffic patterns: No shared prefix, Multi-turn chat, Mixed traffic, and Heavy reuse. Then the concurrency was set to 150 with 10K input tokens per request, 150 with 20K, and 200 x 20K. The last one really stresses the KV storage because you are running at higher concurrency and more input tokens per request. You may be wondering why we are not doing this at c=1, but remember, this is a H100 cluster, and BIG-IP Next for Kubernetes is really designed for clusters instead of single users doing local AI.

F5 BIG-IP Next for Kubernetes versus Envoy AI Gateway - Qwen3-32B FP8 - No Shared Prefix - 150 x 10K Requests - Setup
F5 BIG-IP Next for Kubernetes versus Envoy AI Gateway – Qwen3-32B FP8 – No Shared Prefix – 150 x 10K Requests – Setup

Starting with the base case where we are running at relatively low c=150 and 10K input tokens, you can see that the F5 solution and the Envoy AI Gateway are fairly close.

F5 BIG-IP Next for Kubernetes versus Envoy AI Gateway - Qwen3-32B FP8 - No Shared Prefix - 150 x 10K Requests - Headline Performance
F5 BIG-IP Next for Kubernetes versus Envoy AI Gateway – Qwen3-32B FP8 – No Shared Prefix – 150 x 10K Requests – Headline Performance

Here is a look at some of the key metrics over the course of the run.

F5 BIG-IP Next for Kubernetes versus Envoy AI Gateway - Qwen3-32B FP8 - No Shared Prefix - 150 x 10K Requests - Charts
F5 BIG-IP Next for Kubernetes versus Envoy AI Gateway – Qwen3-32B FP8 – No Shared Prefix – 150 x 10K Requests – Charts

Now let us get to the hard case, which is the other extreme. We are increasing the concurrency by 33% (200 versus 150) and doubling the input tokens. That puts a lot more pressure on the infrastructure to place workloads on the right GPUs. In the base case above, we used only 46% of the cluster’s KV capacity, but now we are at 1.24x, which means we are oversubscribed.

F5 BIG-IP Next for Kubernetes versus Envoy AI Gateway - Qwen3-32B FP8 - No Shared Prefix - 200 x 20K Requests - Setup
F5 BIG-IP Next for Kubernetes versus Envoy AI Gateway – Qwen3-32B FP8 – No Shared Prefix – 200 x 20K Requests – Setup

This is where we get the massive 3.24x gain number because F5 is doing a much better job balancing this load.

F5 BIG-IP Next for Kubernetes versus Envoy AI Gateway - Qwen3-32B FP8 - No Shared Prefix - 200 x 20K Requests - Headline Performance
F5 BIG-IP Next for Kubernetes versus Envoy AI Gateway – Qwen3-32B FP8 – No Shared Prefix – 200 x 20K Requests – Headline Performance

Here is a look at the runs again across some key metrics:

F5 BIG-IP Next for Kubernetes versus Envoy AI Gateway - Qwen3-32B FP8 - No Shared Prefix - 200 x 20K Requests - Charts
F5 BIG-IP Next for Kubernetes versus Envoy AI Gateway – Qwen3-32B FP8 – No Shared Prefix – 200 x 20K Requests – Charts

As you may have noticed, this exercise generated a lot of data running four different workloads across four different concurrency and input sizes. Here are plots using those four scenarios across the three test loads and then in terms of requests completed. You will notice that the lower KV pressure on the right side is generally closer between the two gateways, while the heavier concurrency tends to favor F5 more.

F5 BIG-IP Next for Kubernetes versus Envoy AI Gateway - Qwen3-32B FP8 - Requests Completed
F5 BIG-IP Next for Kubernetes versus Envoy AI Gateway – Qwen3-32B FP8 – Requests Completed

In terms of raw tokens per second, here is a similar view of that:

F5 BIG-IP Next for Kubernetes versus Envoy AI Gateway - Qwen3-32B FP8 - Output Tokens S
F5 BIG-IP Next for Kubernetes versus Envoy AI Gateway – Qwen3-32B FP8 – Output Tokens S

In terms of mean TTFT, here is what that look like:

F5 BIG-IP Next for Kubernetes versus Envoy AI Gateway - Qwen3-32B FP8 - mean TTFT
F5 BIG-IP Next for Kubernetes versus Envoy AI Gateway – Qwen3-32B FP8 – mean TTFT

Of course, folks want to see the p99 TTFT as well, so here is what that looks like, and again, lower is better:

F5 BIG-IP Next for Kubernetes versus Envoy AI Gateway - Qwen3-32B FP8 - p99 TTFT
F5 BIG-IP Next for Kubernetes versus Envoy AI Gateway – Qwen3-32B FP8 – p99 TTFT

I think the key takeaway is that, especially at lower loads, the solutions can be relatively close together. Once the load ramps up, the F5 solution performs really well. The 3.24x number is probably closer to the extreme in terms of benefit we saw with this setup, but even if you only get 1.25x the performance from the same GPUs, that is like a buy four GPUs and get one free proposition. Actually, it is better than that because it is not just the cost of the GPU. It is also the all-in cost of running that GPU, including the servers, networking, power, and so forth. At a smaller scale, that may not be exciting, but as infrastructure scales, and also lead times for power and GPU systems increase, using the same infrastructure and getting more work done with it makes a lot of sense.

Final Words

Although we usually focus on hardware, I thought this was a fun opportunity to check out a networking company’s lab as part of our current lab tour series. Orchestration and operationalizing the AI servers we review is also something we get asked about a lot, so I wanted to learn a bit more. It also shows just how much optimization can help utilization. That will be a major topic in the coming years, given that AI is a trillion-plus-dollar endeavor for humanity. Optimizing the build-out, and what is already installed, is something that simply must happen.

F5 Lab 3
F5 Lab 3

Putting F5 BIG-IP for Kubernetes on BlueField-3 DPUs keeps load balancing and L4-to-L7 security off the CPU and frees up GPU cycles an inference cluster needs. To me, seeing a rack with F5 BIG-IP appliances and one running on the DPUs was neat. On our AI data center tours, we often see firewalls, VPNs, and other network appliances in dedicated connectivity racks. They have been a necessary part of shared infrastructure. Now, that runs on the DPU that is already installed in many AI servers. The idea of offloading this security and control plane to the DPUs and freeing up CPU resources is really neat, and is one of the reasons I wanted to do this piece. It is also a cool example of why you want DPUs in servers serving the double need now that server CPUs are harder to come by.

At the same time, I know folks are going to want to know about 20 different models, what about different GPUs, larger clusters, and all of the features we did not really get into in this piece. Also, I have no idea what the pricing is. Still, we wanted to at least show that something like this exists because many do not know that there is an entire industry working on these types of solutions.

LEAVE A REPLY

Please enter your comment!
Please enter your name here

This site uses Akismet to reduce spam. Learn how your comment data is processed.