Training the largest AI models has hit a wall that adding GPUs cannot solve: power. According to a sponsored interview published by The Next Platform, infrastructure teams are increasingly forced to spread a single training workload across multiple datacenters, which turns the network from a simple transport layer into part of the compute system itself.

Cisco's answer, described by SVP Rakesh Chopra, is a 'Scale-Across' approach. It calls for routing and silicon designed for long-distance fiber links, with buffering and high-speed coherent optics to handle synchronous GPU traffic. The goal is to make geographically separate facilities behave like one deterministic machine, avoiding stalls that ripple across the whole job.

Inside the cluster, Cisco's Silicon One and Intelligent Collective Networking target synchronized bursts of GPU traffic to reduce bottlenecks and improve job completion time. The interview also stresses power efficiency, noting a trade-off between network energy use and power available to GPUs, plus hardware-accelerated MACsec/IPsec for security and programmability for future AI workloads.