
Nvidia, known primarily for hardware, has become the world's largest publisher of open AI models, releasing systems across reasoning, robotics, autonomous vehicles, and scientific computing—all available for free. The company deliberately designs its models and GPUs together for efficiency, allowing it to train in low-precision 4-bit format and scale reinforcement learning cheaply, which not only proves its hardware's value but has influenced rival labs to adopt the same architectural patterns. By open-sourcing these models, Nvidia creates a developer ecosystem that relies on its chips while advancing its own research and shifting focus from data abundance to the diversity of environments models can learn from.
Summaries like this, in your inbox every morning.
Sign up free →What happened
Nvidia has become the world's largest publisher of open AI models, with its lineup spanning reasoning models (Nemotron), world models for physical AI (Cosmos), robot foundation models (Isaac GR00T), self-driving AI (Alpamayo), and specialized models for biology, quantum computing, weather, and climate. The company released Nemotron 3 across late 2025 and 2026 in three sizes (Nano, Super, Ultra), and unified its Cosmos world model line into Cosmos 3 in 2026, which includes a 4-billion-parameter Edge version for real-time robot control.
Why it matters
Nvidia's strategy of releasing powerful open models for free creates a paradox that actually benefits its core GPU business: faster, open models train better on its hardware and prove its chips' value to developers and enterprises. By designing models and hardware in tandem—such as training larger models in 4-bit NVFP4 format built for its Blackwell GPU generation—Nvidia demonstrates that progress now requires co-designing silicon and software rather than relying on Moore's Law alone. This approach has already influenced competitors; other labs including Qwen and Kimi have adopted Nvidia's hybrid Mamba-Attention architecture after it was published in 2024.
What to watch
Nvidia's focus on scaling reinforcement learning (RL) training—enabling the model to learn from over a million rollouts across parallel environments at low cost—rather than scaling data or raw compute alone, signals where the company sees the next capability breakthrough. The production-friendly commercial license added to GR00T 1.7 indicates Nvidia expects real-world robotics deployment to accelerate.
Nvidia, traditionally known for manufacturing GPUs, has quietly become the world's largest publisher of open AI models, according to Bryan Catanzaro, VP of Applied Deep Learning Research at Nvidia. The company's models rank among the most downloaded on Hugging Face and span an unusually wide range of applications: reasoning systems that think through math and coding, world models that simulate and predict physical environments, robot foundation models (called vision-language-action models), autonomous driving systems, and specialized tools for drug discovery, quantum computing, and global weather forecasting.
The reasoning model family, called Nemotron, illustrates Nvidia's iterative approach. The first Nemotron model debuted in 2023. In 2024, Nvidia released a next-generation version with 340 billion parameters, followed by a faster hybrid design in 2025. The current generation, Nemotron 3, arrived across late 2025 and 2026 in three sizes: a Nano model for quick tasks, a Super model for planning and reasoning, and an Ultra model for complex problems requiring heavy reasoning. Separately, Nvidia's CEO Jensen Huang has positioned world models—systems that predict the next state of a physical environment given current conditions and an action—as "the ChatGPT moment for physical AI." Nvidia's Cosmos world model family launched at CES in January 2025 and unified in 2026 into Cosmos 3, a single omnidirectional model that can generate scenes, reason about them, and predict outcomes. Cosmos 3 also includes a 4-billion-parameter Edge version small enough to run directly on a robot for real-time control.
Nvidia's robotics foundation model, Isaac GR00T, debuted as a project in 2024, with the first open humanoid foundation model, GR00T N1, released in early 2025 and iterated to GR00T 1.7. GR00T uses Cosmos as its reasoning backbone, enabling robots to share a foundation for perception, reasoning, and locomanipulation—the ability to move while manipulating objects. This shared backbone allows robots to break complex actions into structured plans that developers can adapt for specific tasks. The 1.7 release added a production-friendly commercial license to ease real-world deployment. A similar pattern applies to autonomous driving: Alpamayo, Nvidia's self-driving model family, also uses Cosmos as its reasoning backbone, allowing vehicles to reason through unfamiliar road situations and expose the chain of cause and effect behind decisions for developer inspection. Beyond these domains, Nvidia offers BioNeMo for biology and drug discovery (with an Agent Toolkit for autonomous research agents), Ising for quantum computing, and Earth-2 for weather and climate modeling.
The company's technical approach reveals why it can produce strong models across so many domains at speed. Nvidia pursues what Catanzaro calls the principle that "the fastest model is the smartest model." Faster models train on more data in the same timeframe, can be post-trained across more environments, and can reason longer on hard problems at the same inference cost. To achieve both speed and capability, Nvidia makes three key design choices. First, it uses a hybrid architecture combining Mamba layers—which process sequences in linear time with fixed memory—with selective Transformer attention layers that restore precise fact recall. This hybrid design, published by Nvidia in 2024, makes a million-token context window practical and has since been adopted by other labs including Qwen and Kimi. Second, Nvidia trains its larger models in 4-bit NVFP4 format from the initial pretraining step, not afterward, using less memory and less data movement to speed up training. This was only possible because Nvidia designed its Blackwell GPU generation with fast 4-bit hardware, exemplifying the company's philosophy of co-designing model and chip. Catanzaro explained the mindset: "Moore's law is no longer giving easy gains every generation, so progress has to come from designing the model and the hardware together."
Third, Nvidia uses mixture-of-experts (MoE) layers that route each token to only a subset of expert layers, activating fewer parameters per token while maintaining large total capacity—keeping the model fast without sacrificing scale. Post-training, the models follow a standard recipe: supervised fine-tuning on curated examples, followed by reinforcement learning where the model learns from its own attempts on real tasks. Nvidia's efficiency advantage allows it to scale RL at lower cost, enabling the model to practice across many parallel environments and learn from over a million rollouts. Catanzaro placed the real bottleneck for better models not on data or raw compute, but on diversity—the variety of environments a model encounters during training. By releasing open models widely, Nvidia effectively crowdsources this diversity problem across thousands of developers and deployment scenarios.
Nvidia's strategy of becoming the world's largest open AI model publisher reflects a fundamental shift in how the company sees competitive advantage in the AI era. Rather than hoarding models to sell software licenses, Nvidia uses open-source releases to demonstrate and lock in the value of its hardware. The article reveals that this is not accidental: the company designs models and chips in tandem. Its decision to train Nemotron in 4-bit NVFP4 format was possible only because Nvidia knew its next-generation Blackwell GPU would include fast 4-bit hardware, meaning the model and silicon were engineered for each other from the start. This co-design philosophy stems from the observation that Moore's Law no longer delivers easy gains, so progress requires thinking beyond chip speeds alone.
The breadth of Nvidia's model portfolio—from reasoning to robotics to climate forecasting—serves a dual purpose. It seeds a developer ecosystem that will naturally build on Nvidia infrastructure, while also positioning the company as a foundational platform for physical AI, a space Jensen Huang has called "the ChatGPT moment for physical AI." The Isaac GR00T and Cosmos lines, which share the same perception and reasoning backbones, exemplify this platform strategy: a single foundation enables developers to build both robot and autonomous vehicle applications on the same base, reducing friction and deepening lock-in to Nvidia's stack.
The article also hints at where Nvidia sees the next research frontier. Rather than framing the bottleneck as data or compute, Bryan Catanzaro places it on diversity—the variety of environments a model learns from during reinforcement learning. By releasing open models, Nvidia essentially crowdsources this diversity problem: thousands of developers will fine-tune and deploy these models across new domains, generating feedback and insights that flow back into Nvidia's research. This makes open-sourcing not a cost but an investment in a richer training signal.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion





Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime