Skip to content
context Tools and apps

AWS integrated the Ray framework into Amazon SageMaker HyperPod to simplify distributed training and serving

clearly official source

AWS integrated the Ray framework into Amazon SageMaker HyperPod. Data scientists can create and manage Ray clusters directly from SageMaker Studio without Kubernetes manifests, with automatic recovery from node failures and faster training recovery through checkpointing.

AWS announced new Ray capabilities for Amazon SageMaker HyperPod that connect the open-source framework Ray (used for distributed training through Ray Train and model serving through Ray Serve) with HyperPod infrastructure designed for training and serving foundation models on Amazon EKS. According to the company, running Ray on Kubernetes previously required writing YAML manifests, manually rebuilding Docker images whenever dependencies changed, setting up kubectl port-forward to access Ray Dashboard, and manually configuring Prometheus and Grafana for observability.

The new integration allows users to create Ray clusters, open Ray Dashboard and Amazon Managed Grafana dashboards, connect a JupyterLab or Code Editor workspace to a cluster, submit distributed jobs, and configure detection of stuck jobs directly from SageMaker Studio. A cluster can be created using a form (name, instance types for head and worker nodes, number of workers, container image), with the AWS-managed SageMaker Distribution image with Ray preinstalled used by default. For advanced use cases, an inline YAML editor and integration with HyperPod task governance remain available for setting compute quotas and scheduling priorities.

At the application level, according to the company, Ray training jobs gain automatic fault tolerance through node health monitoring and recovery within HyperPod, along with tiered checkpointing for faster training recovery through distributed HyperPod tiered storage. Integration with SageMaker JumpStart loads model weights directly into Ray Serve endpoints and offers KV cache offloading to tiered storage to serve requests with long contexts. According to the company, these capabilities are compatible with standard Ray APIs and the open-source KubeRay operator, so existing scripts and workflows run unchanged.

You can find details in the source article.

What changed

Why it matters

For teams that previously managed Ray on Kubernetes manually (writing YAML manifests, rebuilding Docker images, kubectl port-forward, configuring Prometheus/Grafana), the integration, according to AWS, removes much of this operational burden and moves cluster management into the SageMaker Studio console. This may shorten the time needed to get distributed training up and running and reduce the risk of losing work when a node fails through automatic recovery and tiered checkpointing, without having to change existing Ray scripts or workflows based on KubeRay.

Two audiences, two different impacts

What this means

01

For individuals

Data scientists and ML engineers using AWS can now create, manage, and scale Ray clusters for distributed training and serving directly from the SageMaker Studio console, without writing Kubernetes manifests, manually setting up kubectl port-forward, or configuring monitoring.

What to do Data scientists and ML engineers working with AWS can try creating a Ray cluster directly in SageMaker Studio and connecting JupyterLab or Code Editor to it without having to write Kubernetes manifests.
More practical updates →
02

For a business

For companies running distributed training and serving ML models on AWS, the integration reduces the need for specialized Kubernetes expertise and manual management of YAML manifests, Docker images, and monitoring (Prometheus, Grafana), which, according to AWS, simplifies operations and reduces the risk of outages through automatic recovery from node failures.

Development
What to decide Consider evaluating SageMaker HyperPod with Ray integration for ML/infrastructure teams using AWS as an alternative to manually managing Kubernetes clusters for distributed training and model serving.
More business impacts →
AWS distributed training KubeRay Kubernetes Ray SageMaker HyperPod

Check the original

Event sources

clearly official source · 1 publisher, 0 independent. We count feeds from the same owner only once.

1
AWS Machine Learning Blog primary source · first detected Introducing new Ray capabilities on SageMaker HyperPod