NVIDIA released the open Cosmos 3 model for physical AI, AWS described deployment on SageMaker HyperPod
NVIDIA released Cosmos 3, an open omnimodal world foundation model for robotics and autonomous vehicles, under the OpenMDW-1.1 license. AWS published a guide to building a pipeline on Amazon SageMaker HyperPod in which one model replaces three separate GPU queues.
AWS published a guide to building a so-called Physical AI model factory — a continuous pipeline for training robots and autonomous vehicles — built on the NVIDIA Cosmos 3 model and running on Amazon SageMaker HyperPod. According to NVIDIA, Cosmos 3 is an open omnimodal world foundation model that processes video, images, actions and audio as a single stream of tokens. The same transformer trunk runs in three modes: as a forward-dynamics world model for generating synthetic video, as an inverse-dynamics action labeler and as a deployable policy for controlling a robot.
Unlike the conventional approach, which combines a diffusion-transformer video generator with a separate vision-language model for text conditioning, Cosmos 3 integrates both functions into a single trunk at every layer. According to AWS, this allows three previously separate GPU queues (for data generation, post-training and evaluation) to be consolidated into one persistent, shared GPU pool under one control plane, instead of creating and tearing down a separate pool for each phase. The model was released under the Linux Foundation OpenMDW-1.1 license, and the architecture is described in the Cosmos 3 technical report.
The model comes in several sizes: Cosmos3-Nano (16 billion parameters on a dense 8B Qwen3-VL backbone), Cosmos3-Super (64 billion parameters on a dense 32B Qwen3-VL backbone) and derived task-specific variants such as Cosmos3-Nano-Policy-DROID. NVIDIA also separately released Cosmos3-Edge, a compact 4-billion-parameter variant for deployment directly on devices (tested on Jetson Thor and Orin), built on its own ~2B backbone trained from scratch, rather than derived from Qwen3-VL.
The described pipeline operates as a loop: real-world data from DROID, BridgeData2 and autonomous vehicle sensors is stored in Amazon S3 and FSx for Lustre, the Cosmos3-Super model acts as a teacher to generate synthetic data to supplement it, the combined corpus is used for post-training a deployable Cosmos3-Nano policy, and the result is tested in a closed simulation, with failures fed back into the corpus for the next round. Details on the individual phases and infrastructure configuration can be found in the source article.
Why it matters
For teams developing robots or autonomous vehicles, consolidating three GPU pools into one shared pool under one control plane offers an opportunity to reduce the cost of reserved GPU capacity — AWS describes this using the GPU goodput metric, meaning useful pipeline progress per reserved GPU-hour. The open OpenMDW-1.1 license and the available model family (Nano, Super, Edge) also allow teams to build on a ready-made architecture instead of developing their own world model from scratch.
Two audiences, two different impacts
What this means
For individuals
Engineers working on robotics or autonomous vehicles can use the open Cosmos 3 model and the published reference architecture instead of building their own pipeline for data generation, training and evaluation from scratch.
For a business
According to AWS, companies building physical AI can consolidate three previously separate GPU queues (data generation, post-training, evaluation) into a single shared infrastructure under one control plane, reducing the amount of reserved GPU capacity needed to run the entire pipeline.
DevelopmentCheck the original
Event sources
only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.