
Nvidia released Nemotron 3.5 Lightning, a compact open-weights model that matches OpenAI's gpt-oss-120b on intelligence benchmarks while using a quarter of the parameters and delivering the fastest inference speed in its class at 669 tokens per second.
The model demonstrates that smaller, efficiency-focused architectures can achieve competitive reasoning performance, particularly on agentic tasks, making it a practical choice for developers building agent pipelines who prioritize speed and cost.
What happened
Nvidia released Nemotron 3.5 Lightning, a compact open-weights model with 31.6 billion total parameters (only 3.6 billion active per step) that scores 24 on the Artificial Analysis Intelligence Index—matching OpenAI's gpt-oss-120b—while delivering 669 tokens per second, the fastest inference speed in its class.
Why it matters
The model demonstrates that smaller, efficient architectures can match larger competitors on reasoning tasks. For developers building agent-based systems, Lightning offers a permissively-licensed alternative (OpenMDW-1.1) that runs nearly twice as fast as Google's Gemini 3.5 Flash-Lite (386 tokens/s) while achieving parity on intelligence benchmarks, reducing computational cost and latency.
What to watch
Lightning is available now in both BF16 and NVFP4 weights, with serverless inference offered by DeepInfra, Fireworks, FriendliAI, CoreWeave, GMI Cloud, Nebius, and Crusoe. The model shows the biggest gains on agentic benchmarks, scoring an Elo rating of 824 on GDPval-AA v2 and 24.3% on Terminal-Bench v2.1.
Nvidia announced Nemotron 3.5 Lightning, the first model in its new Nemotron 3.5 lineup, positioning it as a fast, efficient alternative to larger open-weight and proprietary models. The model retains the hybrid Mamba-Transformer architecture of its predecessor, Nemotron 3 Nano 30B A3B, but improves its performance substantially. With 31.6 billion total parameters, it activates only 3.6 billion at any given time—a sparse activation pattern that enables both speed and efficiency.
On the Artificial Analysis Intelligence Index, Lightning scores 24, a nine-point jump from Nemotron 3 Nano's 15. This score matches OpenAI's gpt-oss-120b (also 24) and sits just behind Nvidia's larger Nemotron 3 Super (26), which has approximately four times more parameters. However, the smartest models in the same compact size class—Qwen3.6 35B A3B (32 points) and Meta's Muse Glimmer (35)—still maintain a clear lead on raw intelligence.
Where Lightning shines is speed and efficiency. In pre-release tests using final NVFP4 weights, it achieves nearly 670 tokens per second, the highest measured throughput among all compared models and almost twice the speed of Google's Gemini 3.5 Flash-Lite (386 tokens/s). A task from the Intelligence Index takes approximately 0.5 minutes on Lightning, compared to around 3.5 minutes on Qwen3.6 35B A3B and roughly 5.8 minutes on Gemma 4 31B. Proprietary models maintain an overall efficiency advantage—Gemini 3.5 Flash-Lite scores 37 on the Intelligence Index with similar per-task latency, and GPT-5.6 Luna reaches 52 points in under two minutes—but Lightning offers the fastest open-weight alternative.
The most significant performance gains appear on agentic benchmarks, which measure a model's ability to reason through multi-step tasks. On GDPval-AA v2, Lightning achieves an Elo rating of 824, a 334-point improvement over Nemotron 3 Nano and exceeding both gpt-oss-120b (800) and the larger Nemotron 3 Super (698). On Terminal-Bench v2.1, its score jumps from 7 to 24.3 percent, nearly matching gpt-oss-120b at 26.2 percent. Nvidia notes that it worked with partners including CodeRabbit and Harvey on post-training to boost domain-specific performance.
Nvidia ships Lightning under the permissive OpenMDW-1.1 license, positioning it as a high-throughput model for agent-based pipelines. The model is available in both BF16 and NVFP4 weights; the NVFP4 variant maintains the 24-point Intelligence Index score with minimal quality loss, according to Artificial Analysis. It handles text only and supports a context window of one million tokens. Weights are available now, and serverless inference is offered by DeepInfra, Fireworks, FriendliAI, CoreWeave, GMI Cloud, Nebius, and Crusoe, among others. This release builds on Nvidia's earlier research argument that models with only a few billion active parameters can match much larger systems on agent workloads while reducing computational cost tenfold or more.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
IBM and Together AI have signed a $240 million agreement to build and operate an Nvidia-powered AI inference c…

Running a 122-billion-parameter model on three RTX 3090 GPUs with a 256K-token context, the author's AI agent…

NVIDIA and partners released multiple open-source AI models optimized for local execution throughout August, i…

Warren Buffett's Berkshire Hathaway holds few pure AI stocks, but its portfolio of insurance and banking busin…

Bloom Energy reported Q2 2026 record revenue of $1.1 billion (up 166% year-over-year) with gross margin expand…

Major technology companies are advocating for a new standardized framework to report incidents involving AI ag…

The AI news that matters, in one minute each morning.
Sign up free