Skip to content
context Tools and apps

Amazon SageMaker HyperPod and Qumulo enable AI model training across AWS regions without data replication

only one source so far

Amazon Web Services and Qumulo have described a solution that, using Amazon SageMaker HyperPod and Cloud Data Fabric, enables training AI models with data located in a different AWS region without copying it. A test on the LLaMA v3 model at 60 ms latency showed performance comparable to that with local data.

Amazon Web Services and Qumulo, according to their own description, have introduced an architecture that combines Amazon SageMaker HyperPod with the Cloud Native Qumulo layer and its Cloud Data Fabric technology. The solution is meant to allow compute capacity for AI model training to run in one AWS region while the training data remains in another region or on-premises, without needing to copy it or change the training job's code.

They verified the functionality on a LLaMA v3 model with 1.02 billion parameters, run on two ml.p5.48xlarge instances per cluster (16 H100 GPUs in total). The hub cluster ran in the us-east-2 region along with the data, while the spoke cluster in the us-west-2 region read data remotely via Cloud Data Fabric at a network latency of 60 ms. After a brief warm-up, the spoke cluster achieved speeds comparable to the hub (115–116 versus 116–117 samples per second); according to the companies, on a cold start operation was approximately 15–20% slower for the first roughly 150 batches, before performance converged to the hub cluster's level within the first epoch.

A key component is Qumulo NeuralCache, a layer using a predictive model that learns the pattern of sequential reads in 4 KB blocks and prefetches this data from the hub cluster into local NVMe storage on the spoke cluster. After warm-up, according to the companies, 94–96% of reads are served from the local cache with latency under 5 ms, and Cloud Data Fabric is meant to maintain near-full throughput even on links with round-trip delay of up to 900 ms.

The description of the architecture in the available part of the article ends in the middle of describing a configuration where the SageMaker HyperPod cluster runs in a different region than the source data. See the source article for details.

What changed

Why it matters

For companies running training of large AI models, the solution, according to the vendors, removes the need to choose between costly replication of petabytes of data across regions and accepting slower training due to network latency. This makes it possible to use available GPU capacity wherever it is accessible, even while data stays at its source.

Relevant practical impact

What this means

01

For a business

According to the described solution, companies running distributed AI model training can reduce the cost of replicating data across AWS regions while maintaining training performance comparable to local data placement, even at 60 ms latency between regions.

Development
What to decide Check with the infrastructure team whether combining Amazon SageMaker HyperPod and Qumulo Cloud Data Fabric could reduce costs or latency for our own cross-region AI model training.
More business impacts →
Amazon SageMaker HyperPod AWS Cloud Native Qumulo Qumulo trénink modelů

Check the original

Event sources

only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.

1
AWS Machine Learning Blog primary source · first detected Multi-Region training with Amazon SageMaker HyperPod and Qumulo