Skip to content
worth noting Tools and apps

AWS released the Kubernetes addon HyperPod Inference Gateway for GPU-aware routing of LLM inference

clearly official source

AWS introduced Amazon SageMaker HyperPod Inference Gateway, a Kubernetes addon for GPU-aware routing of LLM requests. According to the company, it reduces time to first token by up to 82 %, in one example from 4.4 s to 800 ms, without requiring changes to applications.

AWS released Amazon SageMaker HyperPod Inference Gateway, a Kubernetes-native addon for EKS intended to address inefficient distribution of requests across GPUs when running large language models. According to the company, standard Kubernetes load balancers (round-robin, least-connections) cannot see inside GPU processes — they do not know which pods have a full KV cache, which are currently generating a long context, or which already have the required LoRA adapter in memory — so requests pile up behind busy pods while other capacity remains unused.

The gateway instead uses actual metrics from Prometheus and a weighted scoring algorithm to route each request to the most suitable pod. The architecture has two stages: the first stage is installed on an individual cluster as a managed addon (L7 proxy, model identification from the request body, scoring layer), while the second stage adds coordination across multiple clusters and regions, including cross-cluster failover and global rate limiting. According to the company, installation requires no sidecars, service mesh, or changes to application or client code — the interface remains OpenAI-compatible. The gateway can also route requests for fine-tuned LoRA adapters to pods that already have them loaded in GPU memory, and direct requests with a shared prompt prefix to pods with that prefix cached.

According to AWS, tests on four models ranging from 8 to 235 billion parameters, deployed on p5.48xlarge (H100) and g5 (A10G) instances, showed a reduction in time to first token of up to 82 %, in one reported example from 4.4 seconds to less than 800 milliseconds. According to the company, the improvement is most pronounced with uneven GPU fleets, bursts of traffic, and workloads with shared context, such as multi-turn conversations.

Details on pricing, addon availability outside AWS, and the complete table of behavior during outages can be found in the source article.

What changed

Why it matters

Teams running LLMs on GPU fleets in AWS can, according to the company, reduce wasted capacity and overprovisioning because the gateway detects GPU utilization, KV cache occupancy, and loaded LoRA adapters in real time and routes requests accordingly. Deployment as a single addon without modifying application code lowers the barrier to deployment for teams already running inference on SageMaker HyperPod or EKS.

Two audiences, two different impacts

What this means

01

For individuals

Engineers deploying models on Kubernetes/EKS can deploy the gateway as a single addon without modifying application code or clients, because the interface remains OpenAI-compatible.

What to do If you work with LLM deployments on SageMaker HyperPod or EKS, review the documentation for the Amazon SageMaker HyperPod Inference Gateway addon and consider installing it for existing model servers.
More practical updates →
02

For a business

Companies running LLMs on GPU fleets in AWS can, according to AWS, reduce GPU overprovisioning and cut costs because the gateway distributes the load more evenly than default Kubernetes load balancers.

Development
What to decide If your company runs LLM inference on AWS SageMaker HyperPod/EKS, consider testing Inference Gateway as a way to reduce GPU costs and latency without modifying applications.
More business impacts →
Amazon SageMaker AWS GPU inference Kubernetes LLM

Check the original

Event sources

clearly official source · 1 publisher, 0 independent. We count feeds from the same owner only once.

1
AWS Machine Learning Blog primary source · first detected Introducing Amazon SageMaker HyperPod Inference Gateway