AIToday
Large Language ModelsAI Business & IndustryTHE DECODERPublished: Apr 29, 2026, 22:01 JST1 min read

Nvidia releases Nemotron 3 Nano Omni, a 30-billion-parameter open multimodal model trained on data from competing AI labs including Qwen, OpenAI, and DeepSeek.

3 Key Points

  1. Nemotron 3 Nano Omni is an open-source multimodal model that processes text, images, video, and audio. It uses a Mamba-Transformer hybrid with Mixture-of-Experts, activating about three billion parameters per query, and supports a context window of up to 256,000 tokens.

  2. On the OSWorld benchmark for GUI agents (a type of AI that performs computer tasks autonomously), accuracy jumps from 11.1 to 47.4 points compared to the previous version. Nvidia says throughput at the same interactivity level is up to nine times higher than Qwen3-Omni.

  3. Synthetic training data comes from competing models: Qwen, OpenAI's gpt-oss-120b, Kimi-K2.5, and DeepSeek-OCR generated captions and reasoning traces. Nvidia processed roughly 717 billion tokens across seven training stages. The model ships under the NVIDIA Open Model Agreement, which allows commercial use, and Nvidia is releasing training data and training pipelines alongside the model weights.

Ask the AI about this article →

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Anthropic releases Claude Fable 5.1 and Mythos 5.1ITmedia AI+ · 42m ago
  • LLM serving: why continuous batching winsDaily Dose of Data Science · 42m ago
  • Anthropic's Claude Fable 5.1 Now on Snowflake Cortex AISnowflake AI Blog · 42m ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleAlphabet plans up to $40 billion investment in Anthropic to deepen AI collaboration