Skip to content
worth noting Hardware

AWS outlined an architecture for sharing a GPU cluster across AI teams

only one source so far

The reference architecture from AWS connects Amazon SageMaker HyperPod and Amazon EKS. Teams share a GPU cluster but have separate work environments and permissions. The design includes compute quotas and cost allocation to individual teams.

AWS introduced a reference architecture in which multiple AI teams share a single Amazon SageMaker HyperPod cluster with Amazon EKS. The design combines centralized sign-in, team work environments, workload isolation, compute resource management and cost allocation. According to AWS, it is intended to enable efficient sharing of GPU infrastructure while preserving the operational independence of teams.

AWS IAM Identity Center handles authentication. Each team has its own SageMaker AI domain and a separate space for workloads, known as a Kubernetes namespace. Users can submit workloads through kubectl or the Amazon SageMaker Studio interface. Both routes require Amazon EKS access entries to be configured to link the relevant IAM roles to permissions restricted to the namespace of the team concerned.

HyperPod Task Governance manages compute quotas and workload scheduling priorities. Cost allocation by team namespace makes it possible to track usage and apportion shared expenses. The design also includes team and user directories in shared file storage, as well as Amazon S3 object storage, where access is controlled by the team's IAM role.

What changed

Why it matters

The architecture gives administrators a concrete design for running training, interactive development and model serving for multiple teams on shared GPU infrastructure. It connects access permissions with capacity allocation and expense tracking, so access to the cluster itself is only one of the areas it addresses.

Two audiences, two different impacts

What this means

01

For individuals

In the environment described, an AI team member can submit workloads through a graphical interface or the command line. Both routes restrict access to that member's team namespace.

What to do Verify that your permissions through both kubectl and Amazon SageMaker Studio apply only to your team's namespace.
More practical updates →
02

For a business

A company can manage the allocation of shared GPU capacity using team quotas and scheduling priorities while also allocating costs to teams based on their namespaces for workloads.

Processes
What to decide When designing a shared cluster, assess team quotas in HyperPod Task Governance together with cost allocation by team namespace.
More business impacts →
Amazon EKS Amazon SageMaker HyperPod AWS IAM Identity Center HyperPod Task Governance Kubernetes

Check the original

Event sources

only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.

1
AWS Machine Learning Blog primary source · first detected Share GPU clusters across teams with isolation and fairness using Amazon SageMaker HyperPod