Skip to content
context Tools and apps

AWS guide describes how to manage access and shared capacity in Amazon SageMaker HyperPod

only one source so far

AWS has published a guide on making shared capacity in Amazon SageMaker HyperPod available to teams through projects in Amazon SageMaker Unified Studio. It distinguishes four layers of control and leaves infrastructure management to the infrastructure team.

AWS describes how to manage access and shared compute capacity for Amazon SageMaker HyperPod through Amazon SageMaker Unified Studio. A project can be linked to an existing cluster so that its members can run machine learning jobs, view information about the cluster and jobs, and open JupyterLab. Clusters continue to be managed through the interface and API for Amazon SageMaker AI.

The guide distinguishes four layers of control: organization, project, cluster and job. Organization policies determine the available accounts, regions and tools; project membership and roles define team access. At the cluster level, IAM permissions, role-based access control (RBAC) in Amazon EKS or Slurm rules remain in place. Priorities and rules for borrowing unused capacity govern capacity allocation to individual jobs. According to AWS, access to data, encryption keys, the network and job information must also be aligned before the cluster is made available.

AWS recommends centralizing cluster ownership and capacity management in one designated account. Projects in Amazon SageMaker Unified Studio support collaboration, but do not themselves provide a strong security boundary when jobs are running. For greater isolation, the guide recommends dedicated nodes; if legal, regulatory or security requirements call for separate infrastructure, it recommends a separate cluster or account.

What changed

Why it matters

The guide provides a concrete division of responsibilities for multiple teams using the same compute capacity. It also highlights a key limitation: linking a cluster to a project does not replace its security rules. This matters when deciding whether separate projects are sufficient for teams or whether they need separate infrastructure.

Two audiences, two different impacts

What this means

01

For individuals

A machine learning team member can run jobs on an approved linked cluster and monitor their status directly from the project environment without taking over infrastructure management.

What to do Check that an approved linked cluster is available in the project, and check its status before launching a job.
More practical updates →
02

For a business

A company can leave cluster operations to the infrastructure team and delegate project access to individual teams. In doing so, it must assess capacity allocation, permissions and the required isolation together, because separating projects alone does not ensure the security of running jobs.

Processes
What to decide Before making the shared cluster available to teams, jointly review project roles, cluster permissions, data access and capacity allocation rules.
More business impacts →
Amazon SageMaker HyperPod Amazon SageMaker Unified Studio AWS

Check the original

Event sources

only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.

1
AWS Machine Learning Blog primary source · first detected Best practices for Amazon SageMaker HyperPod administration and governance