Skip to content
worth noting New models

AWS SageMaker AI introduces instance preference lists for training and processing jobs

clearly official source

Amazon SageMaker AI now allows users to specify an ordered list of up to 5 GPU instance types for Training and Processing Jobs. The system automatically starts the job on the first available type, eliminating the need for manual retry scripts.

Amazon Web Services announced the instance preference lists feature for Amazon SageMaker AI Training Jobs and Amazon SageMaker Processing Jobs. When creating a job, users can specify an ordered list of up to five acceptable instance types; according to the company, SageMaker AI evaluates the list in priority order and starts the job on the first type with available capacity.

According to the company, the feature addresses situations where a job is tied to a single instance type that is not immediately available during peak demand — teams have previously relied on custom retry scripts that repeatedly check job status, cancel stalled requests, and resubmit them with a different instance type. According to the article, these workarounds are unreliable and are not compatible with reserved capacity through Flexible Training Plans (FTP).

The feature integrates with Flexible Training Plans: reserved capacity from FTP can be assigned to specific preferences in the list, while the others remain on on-demand capacity. The system evaluates the reservation first and moves to the next type in the list if that reservation is exhausted. If none of the listed types has available capacity at the time of evaluation, the description states that the job enters an event-driven queue and automatically attempts to start again as soon as capacity becomes available; the maximum waiting time in the queue is controlled by the MaxPendingTimeInSeconds parameter, which applies only to jobs requesting accelerated instances (the ml.p, ml.g, ml.trn families), not to CPU-only instances.

The source text does not specify a general availability date or pricing terms for the feature itself. Details can be found in the source article.

What changed

Why it matters

Teams running training or batch processing on SageMaker AI no longer need to write and maintain custom retry and polling logic to work around GPU capacity shortages. According to the company, this shortens the wait for jobs to start and reduces the risk of critical tasks, such as nightly model retraining, failing due to an InsufficientCapacityError.

Two audiences, two different impacts

What this means

01

For individuals

ML engineers and data scientists working with SageMaker Training/Processing Jobs can replace custom retry scripts by specifying an ordered list of up to 5 instance types in a single API call.

What to do When creating a SageMaker Training/Processing Job, consider specifying a list of preferred instances instead of using custom retry logic.
More practical updates →
02

For a business

Companies training models on AWS can reduce the operational overhead associated with custom systems for working around GPU capacity shortages and reduce the risk of disruptions to time-critical pipelines caused by the unavailability of a specific instance type.

Development
What to decide Check whether existing internal retry/monitoring tools for SageMaker jobs can be replaced with the native instance preference lists feature and potentially integrated with Flexible Training Plans.
More business impacts →
Amazon SageMaker AI AWS capacity management GPU scheduling instance preferences machine learning training

Check the original

Event sources

clearly official source · 1 publisher, 0 independent. We count feeds from the same owner only once.

1
AWS Machine Learning Blog primary source · first detected Announcing instance preference lists for Amazon SageMaker AI training jobs