AIToday
Large Language ModelsOpen-Source AIAI Business & IndustryTHE DECODERPublished: Aug 12, 2026, 01:01 JST

Nvidia's Nemotron 3.5 Lightning matches larger models at quarter the size

Nvidia's Nemotron 3.5 Lightning matches larger models at quarter the size

3 Key Points

  1. What happened

    Nvidia released Nemotron 3.5 Lightning, a compact open-weights model with 31.6 billion total parameters (only 3.6 billion active per step) that scores 24 on the Artificial Analysis Intelligence Index—matching OpenAI's gpt-oss-120b—while delivering 669 tokens per second, the fastest inference speed in its class.

  2. Why it matters

    The model demonstrates that smaller, efficient architectures can match larger competitors on reasoning tasks. For developers building agent-based systems, Lightning offers a permissively-licensed alternative (OpenMDW-1.1) that runs nearly twice as fast as Google's Gemini 3.5 Flash-Lite (386 tokens/s) while achieving parity on intelligence benchmarks, reducing computational cost and latency.

  3. What to watch

    Lightning is available now in both BF16 and NVFP4 weights, with serverless inference offered by DeepInfra, Fireworks, FriendliAI, CoreWeave, GMI Cloud, Nebius, and Crusoe. The model shows the biggest gains on agentic benchmarks, scoring an Elo rating of 824 on GDPval-AA v2 and 24.3% on Terminal-Bench v2.1.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

FAQ
How fast is Nemotron 3.5 Lightning compared to other models?
Lightning achieves nearly 670 tokens per second, almost twice as fast as Google's Gemini 3.5 Flash-Lite (386 tokens/s). A task from the Intelligence Index takes about 0.5 minutes to complete on Lightning, while Qwen3.6 35B A3B needs around 3.5 minutes and Gemma 4 31B takes roughly 5.8 minutes.
Where can I access Nemotron 3.5 Lightning?
Weights are available now in both BF16 and NVFP4 formats under the permissive OpenMDW-1.1 license. Serverless inference is offered by DeepInfra, Fireworks, FriendliAI, CoreWeave, GMI Cloud, Nebius, and Crusoe, among others. The model supports a context window of one million tokens and handles text only.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleBallmers split $8B philanthropy into three regional nonprofits