AIToday
Large Language ModelsOpen-Source AIAI Business & IndustryTHE DECODERPublished: Aug 12, 2026, 01:01 JST3 min read

Nvidia's Nemotron 3.5 Lightning matches larger models at quarter the size

Nvidia's Nemotron 3.5 Lightning matches larger models at quarter the size

Key takeaway

  • Nvidia released Nemotron 3.5 Lightning, a compact open-weights model that matches OpenAI's gpt-oss-120b on intelligence benchmarks while using a quarter of the parameters and delivering the fastest inference speed in its class at 669 tokens per second.

  • The model demonstrates that smaller, efficiency-focused architectures can achieve competitive reasoning performance, particularly on agentic tasks, making it a practical choice for developers building agent pipelines who prioritize speed and cost.

3 Key Points

  1. What happened

    Nvidia released Nemotron 3.5 Lightning, a compact open-weights model with 31.6 billion total parameters (only 3.6 billion active per step) that scores 24 on the Artificial Analysis Intelligence Index—matching OpenAI's gpt-oss-120b—while delivering 669 tokens per second, the fastest inference speed in its class.

  2. Why it matters

    The model demonstrates that smaller, efficient architectures can match larger competitors on reasoning tasks. For developers building agent-based systems, Lightning offers a permissively-licensed alternative (OpenMDW-1.1) that runs nearly twice as fast as Google's Gemini 3.5 Flash-Lite (386 tokens/s) while achieving parity on intelligence benchmarks, reducing computational cost and latency.

  3. What to watch

    Lightning is available now in both BF16 and NVFP4 weights, with serverless inference offered by DeepInfra, Fireworks, FriendliAI, CoreWeave, GMI Cloud, Nebius, and Crusoe. The model shows the biggest gains on agentic benchmarks, scoring an Elo rating of 824 on GDPval-AA v2 and 24.3% on Terminal-Bench v2.1.

In Depth

Read the full story

Nvidia announced Nemotron 3.5 Lightning, the first model in its new Nemotron 3.5 lineup, positioning it as a fast, efficient alternative to larger open-weight and proprietary models. The model retains the hybrid Mamba-Transformer architecture of its predecessor, Nemotron 3 Nano 30B A3B, but improves its performance substantially. With 31.6 billion total parameters, it activates only 3.6 billion at any given time—a sparse activation pattern that enables both speed and efficiency.

On the Artificial Analysis Intelligence Index, Lightning scores 24, a nine-point jump from Nemotron 3 Nano's 15. This score matches OpenAI's gpt-oss-120b (also 24) and sits just behind Nvidia's larger Nemotron 3 Super (26), which has approximately four times more parameters. However, the smartest models in the same compact size class—Qwen3.6 35B A3B (32 points) and Meta's Muse Glimmer (35)—still maintain a clear lead on raw intelligence.

Where Lightning shines is speed and efficiency. In pre-release tests using final NVFP4 weights, it achieves nearly 670 tokens per second, the highest measured throughput among all compared models and almost twice the speed of Google's Gemini 3.5 Flash-Lite (386 tokens/s). A task from the Intelligence Index takes approximately 0.5 minutes on Lightning, compared to around 3.5 minutes on Qwen3.6 35B A3B and roughly 5.8 minutes on Gemma 4 31B. Proprietary models maintain an overall efficiency advantage—Gemini 3.5 Flash-Lite scores 37 on the Intelligence Index with similar per-task latency, and GPT-5.6 Luna reaches 52 points in under two minutes—but Lightning offers the fastest open-weight alternative.

The most significant performance gains appear on agentic benchmarks, which measure a model's ability to reason through multi-step tasks. On GDPval-AA v2, Lightning achieves an Elo rating of 824, a 334-point improvement over Nemotron 3 Nano and exceeding both gpt-oss-120b (800) and the larger Nemotron 3 Super (698). On Terminal-Bench v2.1, its score jumps from 7 to 24.3 percent, nearly matching gpt-oss-120b at 26.2 percent. Nvidia notes that it worked with partners including CodeRabbit and Harvey on post-training to boost domain-specific performance.

Nvidia ships Lightning under the permissive OpenMDW-1.1 license, positioning it as a high-throughput model for agent-based pipelines. The model is available in both BF16 and NVFP4 weights; the NVFP4 variant maintains the 24-point Intelligence Index score with minimal quality loss, according to Artificial Analysis. It handles text only and supports a context window of one million tokens. Weights are available now, and serverless inference is offered by DeepInfra, Fireworks, FriendliAI, CoreWeave, GMI Cloud, Nebius, and Crusoe, among others. This release builds on Nvidia's earlier research argument that models with only a few billion active parameters can match much larger systems on agent workloads while reducing computational cost tenfold or more.

FAQ

How fast is Nemotron 3.5 Lightning compared to other models?
Lightning achieves nearly 670 tokens per second, almost twice as fast as Google's Gemini 3.5 Flash-Lite (386 tokens/s). A task from the Intelligence Index takes about 0.5 minutes to complete on Lightning, while Qwen3.6 35B A3B needs around 3.5 minutes and Gemma 4 31B takes roughly 5.8 minutes.
Where can I access Nemotron 3.5 Lightning?
Weights are available now in both BF16 and NVFP4 formats under the permissive OpenMDW-1.1 license. Serverless inference is offered by DeepInfra, Fireworks, FriendliAI, CoreWeave, GMI Cloud, Nebius, and Crusoe, among others. The model supports a context window of one million tokens and handles text only.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleBallmers split $8B philanthropy into three regional nonprofits

The AI news that matters, in one minute each morning.

Sign up free