AWS turned GPU-aware inference routing into a single managed add-on for Amazon EKS, and the specification underneath it is open.
- Amazon SageMaker HyperPod Inference Gateway installs as one EKS add-on. No sidecars, no service mesh, no application changes.
- An Endpoint Picker scores live pods on KV cache utilization, queue depth, LoRA adapter residency, prefix cache hit rate, and running requests.
- It is built on the Kubernetes Gateway API Inference Extension, and the sample configuration names llm-d as the scheduler.
- AWS measured first-token latency at P99 down 97 percent on mixed GPU generations and 98 percent under bursty traffic, against a round-robin baseline.
- On a uniform fleet under steady traffic the gateway matched round-robin, which AWS states plainly.
Intelligent routing pays where the fleet is uneven. Tier 2, a Global Inference Router for cross-cluster failover and cost-aware traffic shaping, is not out yet.
Get the next one before it is old news
Independent analysis of cloud-native infrastructure, Kubernetes and data centre economics. No vendor spin.
