
Nvidia published research showing that the software wrapper around an AI model—called a harness—matters far more than the model itself for long-horizon tasks.
By adding a supervising agent component to the harness, Claude Opus 5 achieved 100% on a reasoning benchmark where it scored only 30% without it.
OpenAI's own recent harness tweaks tripled scores but fell short of that result.
What happened
Nvidia researchers used a custom harness with a "supervisor" component to get Claude Opus 5 to score 100% on the ARC-AGI-3 interactive reasoning benchmark, where the same model scored only 30% without the harness. A harness is the software wrapper around an AI model—tools, memory management, and rules that let it act independently.
Why it matters
The finding challenges the assumption that an AI model's underlying intelligence is what matters most for long-horizon tasks (those requiring many chained decisions over time). Databricks research showed in July that harness choice can 2× costs even when using the same model, meaning businesses may be optimizing the wrong layer. Nvidia's result also directly addresses a weakness OpenAI struggled with; OpenAI's own harness tweaks in recent research only tripled scores but did not approach 100%.
What to watch
Most agent users today rely on single-layer harnesses like Claude Code or Codex. Nvidia produces open harness components under the Nemo brand (some commercial, some freely available), giving users more control than closed systems. Nvidia emphasizes that open harnesses let teams adjust accuracy by controlling the harness, infrastructure, and runtime.
Ask the AI about this article →
Nvidia's research arrives at a moment when the field is shifting focus from raw model capability to the systems around models. The company's 100% score on ARC-AGI-3 is particularly significant because it directly challenges OpenAI's earlier struggles on the same benchmark—OpenAI scored less than 10% and was forced to conduct its own research in response. OpenAI's subsequent harness tweaks (adjusting two settings) did improve performance substantially by tripling scores, but still fell short of Nvidia's result, suggesting that the architecture of the harness itself, not just tuning, matters. Databricks' July research showing that harness choice can 2× costs even for the same model reinforces the finding that businesses optimizing purely for model selection are missing a major lever. Nvidia's framing of the harness as the critical piece also serves the company's commercial strategy: it positions open harness components (which Nvidia offers under the Nemo brand) as strategic assets, allowing users to avoid lock-in to proprietary systems like OpenAI's Claude Code or Anthropic's offerings.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Thomson Reuters Corp. today launched Thomson, its first proprietary large language model, combining its legal…
Xiaomi is expanding its in-house semiconductor push from smartphones into AI acceleration and autonomous drivi…

Amazon told investors it now expects to spend $220 billion in 2026, which is $20 billion more than its prior c…

Thomson Reuters launched its first in-house language model, built on Alibaba's Qwen, after spending about $40…

Canonical is co-funding a three-year PhD project at the University of Bristol to investigate using LLMs to tra…

In 9 days from Aug 10, Meta (Muse Glimmer), NVIDIA (Nemotron 3.5 Lightning), and Alibaba Cloud (Qwen3.8-27B) r…
