AWS SageMaker AI introduces instance preference lists for training and processing jobs
Amazon SageMaker AI now allows users to specify an ordered list of up to 5 GPU instance types for Training and Processing Jobs. The system automatically starts the job on the first available type, eliminating the need for manual retry scripts.
Amazon Web Services announced the instance preference lists feature for Amazon SageMaker AI Training Jobs and Amazon SageMaker Processing Jobs. When creating a job, users can specify an ordered list of up to five acceptable instance types; according to the company, SageMaker AI evaluates the list in priority order and starts the job on the first type with available capacity.
According to the company, the feature addresses situations where a job is tied to a single instance type that is not immediately available during peak demand — teams have previously relied on custom retry scripts that repeatedly check job status, cancel stalled requests, and resubmit them with a different instance type. According to the article, these workarounds are unreliable and are not compatible with reserved capacity through Flexible Training Plans (FTP).
The feature integrates with Flexible Training Plans: reserved capacity from FTP can be assigned to specific preferences in the list, while the others remain on on-demand capacity. The system evaluates the reservation first and moves to the next type in the list if that reservation is exhausted. If none of the listed types has available capacity at the time of evaluation, the description states that the job enters an event-driven queue and automatically attempts to start again as soon as capacity becomes available; the maximum waiting time in the queue is controlled by the MaxPendingTimeInSeconds parameter, which applies only to jobs requesting accelerated instances (the ml.p, ml.g, ml.trn families), not to CPU-only instances.
The source text does not specify a general availability date or pricing terms for the feature itself. Details can be found in the source article.
Why it matters
Teams running training or batch processing on SageMaker AI no longer need to write and maintain custom retry and polling logic to work around GPU capacity shortages. According to the company, this shortens the wait for jobs to start and reduces the risk of critical tasks, such as nightly model retraining, failing due to an InsufficientCapacityError.
Two audiences, two different impacts
What this means
For individuals
ML engineers and data scientists working with SageMaker Training/Processing Jobs can replace custom retry scripts by specifying an ordered list of up to 5 instance types in a single API call.
For a business
Companies training models on AWS can reduce the operational overhead associated with custom systems for working around GPU capacity shortages and reduce the risk of disruptions to time-critical pipelines caused by the unavailability of a specific instance type.
DevelopmentCheck the original
Event sources
clearly official source · 1 publisher, 0 independent. We count feeds from the same owner only once.