
NVIDIA released Cosmos 3 on Hugging Face with two model sizes: Cosmos 3 Nano (8B parameter model) optimized for efficient inference on workstation-grade compute like the RTX PRO 6000 GPU, and Cosmos 3 Super (32B parameter model) designed for large-scale synthetic data generation and research on NVIDIA Hopper and Blackwell GPUs.
Cosmos 3 is built on a Mixture-of-Transformers (MoT) architecture that processes text, image, video, audio, and action in a single unified model. It replaces the previous approach where developers had to work with separate models for world generation, controlled generation, scene understanding, and policy generation.
The model supports multiple input-output combinations: text/image/video-to-video generation, text/video-to-text output (for vision language tasks), action/image/text-to-video (forward dynamics), text/video-to-action (inverse dynamics), and image/text-to-video-and-action (policy model). The release includes Diffusers integration, post-training scripts on GitHub, and open synthetic data generation datasets for physical AI.
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
CrowdStrike Holdings Inc
ASE Technology CEO and SEMI Global Board Executive Committee Chair Tien Wu said AI's early-stage growth is cre…

Fervo Energy signed a deal to supply 396 MW of electricity from its Cape Station project in Utah to a Google d…

Nvidia co-announced with Vertiv in August that it is making progress on 800-volt power infrastructure for its…

John Ternus takes over as CEO of Apple today, stepping into the role as the company confronts the AI era

Top AI-native open source projects like Flue and tldraw are refusing external pull requests, often because the…
