Nemotron 3 Nano Omni is an open-source multimodal model that processes text, images, video, and audio. It uses a Mamba-Transformer hybrid with Mixture-of-Experts, activating about three billion parameters per query, and supports a context window of up to 256,000 tokens.
On the OSWorld benchmark for GUI agents (a type of AI that performs computer tasks autonomously), accuracy jumps from 11.1 to 47.4 points compared to the previous version. Nvidia says throughput at the same interactivity level is up to nine times higher than Qwen3-Omni.
Synthetic training data comes from competing models: Qwen, OpenAI's gpt-oss-120b, Kimi-K2.5, and DeepSeek-OCR generated captions and reasoning traces. Nvidia processed roughly 717 billion tokens across seven training stages. The model ships under the NVIDIA Open Model Agreement, which allows commercial use, and Nvidia is releasing training data and training pipelines alongside the model weights.
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
CrowdStrike extends its Falcon platform to police AI agents at the endpoint, treating each agent as an asset w…
PlayNitride Inc., a Micro LED maker, expects its technology to enter commercial optical communications applica…

McKinsey's 2025 survey found that while 65% of companies continuously use generative AI, fewer than 5% have ac…

Anthropic announced Enterprise Frontier Safeguards (EFS) on September 1, offering enterprise customers privacy…

Anthropic announced Claude Fable 5.1 and Claude Mythos 5.1 on September 1

New Goldman Sachs analysis finds that currencies of South Korea, Taiwan, and Malaysia are outperforming those…
