
AWS shows how to build a continuous Physical AI model factory. It uses NVIDIA Cosmos 3 on SageMaker HyperPod.
One model family handles generation, training, and evaluation.
This reduces the need for separate GPU pools per stage.
What happened
AWS published a guide to building a Physical AI model factory using NVIDIA Cosmos 3 on Amazon SageMaker HyperPod, with code in the awsome-distributed-ai GitHub repository.
Why it matters
Cosmos 3 unifies generation, post-training, and evaluation into one model running in different modes, letting a robot or AV team use a single persistent GPU pool instead of separate capacity per stage. This raises GPU goodput—the useful pipeline progress per reserved GPU-hour—which the post identifies as the real cost driver.
What to watch
The model family includes Cosmos3-Nano (16B parameters), Cosmos3-Super (64B parameters), and Cosmos3-Edge (4B tier for on-device deployment). AWS walks through the robot-policy stage on the public DROID dataset.
Ask the AI about this article →
The blog post addresses a practical bottleneck in robotics and autonomous driving development: the need to run a perpetual improvement loop, not a single training job. Traditionally, teams provision separate GPU clusters for each stage—data generation, post-training, and evaluation—each with its own setup and teardown lifecycle. AWS argues this fragments capacity and lowers GPU goodput, the measure of useful pipeline progress per reserved GPU-hour.
Cosmos 3's design is central to AWS's proposal. Because one model can operate as a forward-dynamics world model, an inverse-dynamics action labeler, and a deployable policy, all three stages can run as workloads on one persistent GPU pool. AWS states this eliminates re-provisioning steps and terabyte-scale data migrations between stages. The cluster runs on Amazon SageMaker HyperPod with EKS, with all stages sharing a single Amazon FSx for Lustre file system backed by S3.
The post walks through a complete end-to-end robot-policy stage on the public DROID dataset, with runnable code in the awsome-distributed-ai GitHub repository. AWS emphasizes that capacity should be committed to the whole loop, whether through a flexible training plan for a bounded campaign or a capacity reservation for an open-ended one, to avoid variability in GPU availability and lead times.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
BEXCO general manager Tom Choi argues that South Korea's safety-tech sector, filled with AI cameras, robots, d…

The U.S. Army has selected Palantir to build eight AI-powered TITAN trucks, as reported by Axios

Midea presented SMART MASTER, an AI-powered home ecosystem, at IFA 2026 in Berlin, built around an AI Agent th…

Eric Sivertson, vice president of the security business at Lattice Semiconductor, appeared on The Robot Report…

Nvidia confirmed the $12.9 billion acquisition of AI hosting platform Hugging Face
Lyte AI Inc. raised $165 million in Series C funding, valuing the physical AI company at $1.6 billion post-mon…
