Skip to content
context Tools and apps

AWS summarized 13 new features for Amazon SageMaker Inference, with automated deployment benchmarking as the key addition

clearly official source

AWS summarized 13 new features for Amazon SageMaker Inference released in 2026. The main addition is Inference recommendations – automated selection of an instance and model deployment configuration, which, according to AWS, replaces 2–3 weeks of manual benchmarking.

Amazon (AWS) summarized 13 new features it introduced in 2026 to Amazon SageMaker Inference, a service for deploying and running AI models. The new features cover two deployment approaches: fully managed SageMaker endpoints and Amazon SageMaker HyperPod Inference for teams that need control over Kubernetes clusters with dedicated GPUs.

The key new feature is Inference recommendations. According to AWS, manually selecting an instance type, serving container and optimization settings for a generative model typically requires two to three weeks of testing more than 1000 combinations and expertise that few teams have in-house. The customer specifies a model and an objective (cost, latency or throughput), and SageMaker automates the process; the output is a Model Package with a recommended configuration and measured metrics – time to first token (TTFT), inter-token latency (ITL), P50/P90/P99 latency percentiles, throughput and estimated cost. According to AWS, optimizing for throughput with the GPT-OSS-20B model in one cited example doubled the number of tokens per second at the same request latency. Generating recommendations is free; customers with reserved capacity (ML Reservations) can run the benchmark on that capacity at no additional cost.

Among the six other new features for managed endpoints is Capacity Aware Inference – the ability to define a prioritized list of up to five instance types that SageMaker automatically switches between during endpoint creation and when scaling up or down, so that a capacity shortage for one instance type does not lead to downtime. Other additions include support for an OpenAI-compatible API (the /openai/v1 path with Chat Completions and streaming), Container Caching for faster scaling of large containers, Observability, Async Inference Inline Payloads and Prefix-Aware Routing. For HyperPod Inference, AWS lists Simplified Operator, Tiered KV Cache, Data Capture, Performance Features, Disaggregated Prefill and Decode and Model Caching.

The source article text is only partially available and describes some of the listed features only briefly. Details can be found in the source article.

What changed

Why it matters

Automated benchmarking (Inference recommendations) saves teams deploying generative AI models weeks of manual work and the need for specialized expertise that smaller companies often lack. Instance pools (Capacity Aware Inference) also reduce the risk of production deployment downtime when GPU capacity is insufficient for a particular instance type, a practical risk when running large models at scale.

Two audiences, two different impacts

What this means

01

For individuals

ML engineers who have been manually benchmarking combinations of instances and serving containers (typically 2–3 weeks of work, according to AWS) can have this process automated and receive a configuration complete with latency, throughput and cost metrics.

What to do Try Inference recommendations to estimate the optimal deployment configuration instead of manual benchmarking.
More practical updates →
02

For a business

According to AWS, companies running generative AI models on Amazon SageMaker can reduce the time needed to select an instance and deployment configuration, and lower the risk of endpoint downtime due to insufficient capacity thanks to support for up to five fallback instance types.

Development
What to decide For teams deploying models on SageMaker, consider trying Inference recommendations instead of manually benchmarking a new deployment configuration.
More business impacts →
Amazon SageMaker Inference Container Caching HyperPod Inference Inference recommendations model deployment

Check the original

Event sources

clearly official source · 1 publisher, 0 independent. We count feeds from the same owner only once.

1
AWS Machine Learning Blog primary source · first detected Amazon SageMaker Inference: 2026 year-to-date launches in review An overview of multiple AI topics; AI Radar covers only this event.