AWS released the Kubernetes addon HyperPod Inference Gateway for GPU-aware routing of LLM inference
AWS introduced Amazon SageMaker HyperPod Inference Gateway, a Kubernetes addon for GPU-aware routing of LLM requests. According to the company, it reduces time to first token by up to 82 %, in one example from 4.4 s to 800 ms, without requiring changes to applications.
AWS released Amazon SageMaker HyperPod Inference Gateway, a Kubernetes-native addon for EKS intended to address inefficient distribution of requests across GPUs when running large language models. According to the company, standard Kubernetes load balancers (round-robin, least-connections) cannot see inside GPU processes — they do not know which pods have a full KV cache, which are currently generating a long context, or which already have the required LoRA adapter in memory — so requests pile up behind busy pods while other capacity remains unused.
The gateway instead uses actual metrics from Prometheus and a weighted scoring algorithm to route each request to the most suitable pod. The architecture has two stages: the first stage is installed on an individual cluster as a managed addon (L7 proxy, model identification from the request body, scoring layer), while the second stage adds coordination across multiple clusters and regions, including cross-cluster failover and global rate limiting. According to the company, installation requires no sidecars, service mesh, or changes to application or client code — the interface remains OpenAI-compatible. The gateway can also route requests for fine-tuned LoRA adapters to pods that already have them loaded in GPU memory, and direct requests with a shared prompt prefix to pods with that prefix cached.
According to AWS, tests on four models ranging from 8 to 235 billion parameters, deployed on p5.48xlarge (H100) and g5 (A10G) instances, showed a reduction in time to first token of up to 82 %, in one reported example from 4.4 seconds to less than 800 milliseconds. According to the company, the improvement is most pronounced with uneven GPU fleets, bursts of traffic, and workloads with shared context, such as multi-turn conversations.
Details on pricing, addon availability outside AWS, and the complete table of behavior during outages can be found in the source article.
Why it matters
Teams running LLMs on GPU fleets in AWS can, according to the company, reduce wasted capacity and overprovisioning because the gateway detects GPU utilization, KV cache occupancy, and loaded LoRA adapters in real time and routes requests accordingly. Deployment as a single addon without modifying application code lowers the barrier to deployment for teams already running inference on SageMaker HyperPod or EKS.
Two audiences, two different impacts
What this means
For individuals
Engineers deploying models on Kubernetes/EKS can deploy the gateway as a single addon without modifying application code or clients, because the interface remains OpenAI-compatible.
For a business
Companies running LLMs on GPU fleets in AWS can, according to AWS, reduce GPU overprovisioning and cut costs because the gateway distributes the load more evenly than default Kubernetes load balancers.
DevelopmentCheck the original
Event sources
clearly official source · 1 publisher, 0 independent. We count feeds from the same owner only once.