F5's BIG-IP Next for Kubernetes (BNK) sits between clients and GPU infrastructure to handle LLM routing, token governance, and L4-L7 security. In a sponsored lab tour, ServeTheHome saw it running on NVIDIA BlueField-3 DPUs rather than host CPUs, which the company says avoids consuming valuable host cores. The aim is to improve GPU efficiency, getting more performance from the same power and GPU footprint.
The lab underscored the complexity of AI cluster networking. Beyond GPU-to-GPU traffic, there are networks for storage, user access, control planes, and even DPU management interfaces, often running at 400Gbps or 800Gbps per port. F5's position is that application delivery and security need to be managed across all of that, and offloading those tasks to DPUs keeps them from interfering with GPU work.
ServeTheHome toured with F5's sponsorship, and the article notes that normally you cannot walk into the lab and film. The test setup used NVIDIA H100 GPUs, Supermicro servers, and BlueField-3 DPUs, with the top 2U servers carrying eight DPUs each. STH also brought Keysight load generators capable of over 1.6Tbps to simulate many users.