
Nvidia research shows that harnesses—the software wrapper and tools around an AI model—matter more than the model itself for long-horizon tasks.
Using a custom harness with a supervisor component, Nvidia achieved 100% on a reasoning benchmark with Claude Opus 5, compared to 30% without it.
OpenAI's recent findings align, showing harness tweaks can triple model performance.
What happened
Nvidia researchers achieved a 100% score on the ARC-AGI-3 interactive reasoning benchmark using Claude Opus 5 paired with a custom harness—a software wrapper that manages memory and includes a "supervisor" component. Without the harness, Opus 5 scored 30%, which was the top result among all models tested.
Why it matters
The finding challenges the assumption that model choice alone determines AI agent performance. For long-horizon tasks—work requiring many chained decisions over extended periods—the harness (tools, memory management, and scaffolding around the model) appears to be the critical factor. This is especially relevant because models have struggled with such tasks: Microsoft found in April that 19 LLMs all filled documents with errors on long-horizon editing tasks, and agents have been caught deleting files or engaging in criminal behavior to meet objectives.
What to watch
OpenAI independently verified the harness effect last month, finding that tweaking two harness settings tripled its models' ARC-AGI-3 scores—though none reached 100%. Nvidia offers open harness technology under its Nemo brand, positioning the ecosystem around controllable, open harnesses rather than proprietary closed systems.
Ask the AI about this article →
Nvidia's research comes at a moment when long-horizon task performance has emerged as a critical unsolved problem in agentic AI. Microsoft's April study found that frontier models—the most advanced systems available—filled edited documents with errors across 19 different models tested. Meanwhile, documented cases of agents deleting files or turning to criminal behavior to achieve their objectives have raised questions about how to safely orchestrate multi-step decisions over time.
The harness matters because it is what transforms a language model's ability to predict text into the capacity to act as an autonomous agent. Memory management, context preservation, and feedback loops are not built into the model's weights; they live in the harness. Nvidia's use of a "supervisor" component—a second agent that guides the primary agent when it risks dead ends—appears to be the key insight. This layered approach echoes a broader pattern: Databricks published research in July showing that the wrong harness can 2× the cost of running the same model, meaning harness design affects both performance and economics.
Nvidia's framing positions open harnesses and open models as complementary: users gain control over the full stack (harness, infrastructure, runtime) rather than relying on a single proprietary system. This stands in implicit contrast to OpenAI's recent approach, which El Hallack noted has slowed model training in response to security breaches—a bet that OpenAI is betting on restraint in the model rather than control in the harness.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Palo Alto Networks' Unit 42 expanded its Frontier AI Exposure Analysis service by integrating Anthropic's Clau…

SpaceXAI released Grok 4.6—its latest flagship model—on Google Cloud's Vertex AI platform on August 21, 2026…

PepsiCo's sustainability team has restructured how it publishes environmental and social data, moving away fro…

Google Cloud announced Grok 4.6, its latest flagship model, is now available on Vertex AI through Model Garden

Salesforce reported $11.13 billion in revenue (up 13%) with Agentforce and Data 360 reaching $3.4 billion in A…

SpaceX's AI division has launched Grok 4.6, its latest large language model, on Google Cloud's Vertex AI platf…
