Amazon Web Services has introduced SageMaker HyperPod Inference Gateway, a managed add-on for Amazon EKS that brings GPU-aware routing to large language model inference. The announcement, made on September 18, 2026, positions the gateway as a single add-on that runs directly on existing HyperPod infrastructure, avoiding the need for additional cluster management.

The gateway is designed to route inference requests based on GPU availability and suitability, which can help improve utilization and reduce idle capacity in distributed inference environments. Because it is Kubernetes-native, it integrates with standard EKS workflows and operational tooling.

AWS frames the offering as a way to extend HyperPod's training-focused capabilities into production inference. The main significance is operational: teams already running HyperPod can add intelligent request routing without rearchitecting their infrastructure, potentially lowering costs and latency for LLM serving.