F5 BIG-IP Next for Kubernetes (BNK): The AI Service Plane
With that context, BNK sits in a layer between clients and GPU infrastructure that F5 calls the AI service plane. The diagram shows that plane bundling intelligent AI load balancing with the Endpoint Picker, LLM routing and orchestration, token governance and metering, zero-trust security, and multi-tenant isolation. It deploys cloud-native on Kubernetes over a host CPU or directly on BlueField-3 DPUs, and F5 says no model changes are required for the workloads it fronts.

In the video, we had a fun story about how F5 started as a gaming company, and its load balancing and security offerings came from being better at that part of multiplayer games than making games themselves. BNK has a load balancer (which is fun to think about how that came about) with Endpoint Picker (EPP) that balances across inference endpoints using live GPU metrics instead of static round-robin, with routing F5 describes as prefix-aware, KV-cache-aware, and load-aware. Telemetry from NVIDIA GPU infrastructure and inferencing tooling feeds EPP, so prompts land on GPUs that are ready to serve them, driving the utilization and throughput you will see in the performance section.

BIG-IP for Kubernetes intercepts prompts, sends them to NVIDIA NIM for classification, then routes the prompt to the LLM best suited to the request, sized from a full LLM down to a smaller SLM depending on the policy result.

Token governance uses the same programmable data plane to track, rate-limit, and hard-cap token usage, which has become a major topic in enterprises lately. Here, the solution counts tokens for a given user, routes a request to a more expensive LLM while applying a policy against the token limit, then steers new requests to a less expensive model once that limit is reached. We did not get to test this just due to time, but it is a feature that I know many readers are trying to figure out these days.

Ingress protection starts with an edge firewall and DDoS mitigation, then authorization and sensitive-data-leakage prevention through programmable policies, followed by egress controls that establish trusted origins to external services and observability for audits and traceability. This makes a lot of sense for folks setting up agentic workflows with MCP, which is everywhere. Even video editing is getting impacted by MCP as more of our tools add support. BNK’s coverage targets L4 to L7 traffic to and from MCP servers, which need dedicated security now that agentic AI is driving requests between hosts.

F5 is an NVIDIA Cloud Partner for the NVIDIA Common Networking Reference Architecture, so it is integrating directly into NVIDIA’s frameworks rather than sitting outside of them and trying to build everything from scratch. BIG-IP Next for Kubernetes embeds networking, security, and AI-aware control directly into the NVIDIA reference fabric rather than sitting as a separate hop.

The BNK integration also builds multi-tenant EVPN VXLAN overlays over BGP. That removes the need for manually configured VLANs, uses unnumbered BGP links with IPv6 link-local addresses for automatic peer discovery, and relies on loopback addresses for stable node identification. Supporting EVPN route types 2, 3, and 5 lets compute nodes, including DPUs running BNK, participate as full overlay citizens with route exchange.

BIG-IP has been re-engineered for Arm processors and runs on the NVIDIA BlueField-3 DPU. Arm processor compatibility matters because NVIDIA also offers Grace and Vera as host CPUs, and many hyperscalers use Arm for internal workloads.

DOCA acceleration handles L4 flows and security at full wire speed, so application delivery and security services run toward the AI cluster nodes instead of competing with inference work on the host.

BNK 2.4 is scheduled for the third quarter of CY2026 and will offer on-host, on-host accelerated, DPU-trusted, DPU zero-trust, and mixed deployment models. Today, you can run BNK on the host, but it then uses valuable host CPU cycles. Utilizing DPU cores is interesting these days because in 2026-2027 we expect a server CPU shortage, so offloading workloads from those cores is a way to use CPUs more effectively. If you are a hyperscaler, a neocloud provider, or an enterprise, this offload means you can use those server CPU cores and memory bandwidth for something else.

With that, let us fire up the lab and see the performance impact.


