Amazon SageMaker HyperPod and Qumulo enable AI model training across AWS regions without data replication
Amazon Web Services and Qumulo have described a solution that, using Amazon SageMaker HyperPod and Cloud Data Fabric, enables training AI models with data located in a different AWS region without copying it. A test on the LLaMA v3 model at 60 ms latency showed performance comparable to that with local data.
Amazon Web Services and Qumulo, according to their own description, have introduced an architecture that combines Amazon SageMaker HyperPod with the Cloud Native Qumulo layer and its Cloud Data Fabric technology. The solution is meant to allow compute capacity for AI model training to run in one AWS region while the training data remains in another region or on-premises, without needing to copy it or change the training job's code.
They verified the functionality on a LLaMA v3 model with 1.02 billion parameters, run on two ml.p5.48xlarge instances per cluster (16 H100 GPUs in total). The hub cluster ran in the us-east-2 region along with the data, while the spoke cluster in the us-west-2 region read data remotely via Cloud Data Fabric at a network latency of 60 ms. After a brief warm-up, the spoke cluster achieved speeds comparable to the hub (115–116 versus 116–117 samples per second); according to the companies, on a cold start operation was approximately 15–20% slower for the first roughly 150 batches, before performance converged to the hub cluster's level within the first epoch.
A key component is Qumulo NeuralCache, a layer using a predictive model that learns the pattern of sequential reads in 4 KB blocks and prefetches this data from the hub cluster into local NVMe storage on the spoke cluster. After warm-up, according to the companies, 94–96% of reads are served from the local cache with latency under 5 ms, and Cloud Data Fabric is meant to maintain near-full throughput even on links with round-trip delay of up to 900 ms.
The description of the architecture in the available part of the article ends in the middle of describing a configuration where the SageMaker HyperPod cluster runs in a different region than the source data. See the source article for details.
Why it matters
For companies running training of large AI models, the solution, according to the vendors, removes the need to choose between costly replication of petabytes of data across regions and accepting slower training due to network latency. This makes it possible to use available GPU capacity wherever it is accessible, even while data stays at its source.
Relevant practical impact
What this means
For a business
According to the described solution, companies running distributed AI model training can reduce the cost of replicating data across AWS regions while maintaining training performance comparable to local data placement, even at 60 ms latency between regions.
DevelopmentCheck the original
Event sources
only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.