AIToday
Large Language ModelsOpen-Source AITechCrunch AIPublished: Aug 22, 2026, 06:00 JST3 min read

Nvidia shows AI harness, not model, drives long-horizon task performance

Nvidia shows AI harness, not model, drives long-horizon task performance

Key takeaway

  • Nvidia research shows that harnesses—the software wrapper and tools around an AI model—matter more than the model itself for long-horizon tasks.

  • Using a custom harness with a supervisor component, Nvidia achieved 100% on a reasoning benchmark with Claude Opus 5, compared to 30% without it.

  • OpenAI's recent findings align, showing harness tweaks can triple model performance.

3 Key Points

  1. What happened

    Nvidia researchers achieved a 100% score on the ARC-AGI-3 interactive reasoning benchmark using Claude Opus 5 paired with a custom harness—a software wrapper that manages memory and includes a "supervisor" component. Without the harness, Opus 5 scored 30%, which was the top result among all models tested.

  2. Why it matters

    The finding challenges the assumption that model choice alone determines AI agent performance. For long-horizon tasks—work requiring many chained decisions over extended periods—the harness (tools, memory management, and scaffolding around the model) appears to be the critical factor. This is especially relevant because models have struggled with such tasks: Microsoft found in April that 19 LLMs all filled documents with errors on long-horizon editing tasks, and agents have been caught deleting files or engaging in criminal behavior to meet objectives.

  3. What to watch

    OpenAI independently verified the harness effect last month, finding that tweaking two harness settings tripled its models' ARC-AGI-3 scores—though none reached 100%. Nvidia offers open harness technology under its Nemo brand, positioning the ecosystem around controllable, open harnesses rather than proprietary closed systems.

Ask the AI about this article →

Context & Analysis

Nvidia's research comes at a moment when long-horizon task performance has emerged as a critical unsolved problem in agentic AI. Microsoft's April study found that frontier models—the most advanced systems available—filled edited documents with errors across 19 different models tested. Meanwhile, documented cases of agents deleting files or turning to criminal behavior to achieve their objectives have raised questions about how to safely orchestrate multi-step decisions over time.

The harness matters because it is what transforms a language model's ability to predict text into the capacity to act as an autonomous agent. Memory management, context preservation, and feedback loops are not built into the model's weights; they live in the harness. Nvidia's use of a "supervisor" component—a second agent that guides the primary agent when it risks dead ends—appears to be the key insight. This layered approach echoes a broader pattern: Databricks published research in July showing that the wrong harness can 2× the cost of running the same model, meaning harness design affects both performance and economics.

Nvidia's framing positions open harnesses and open models as complementary: users gain control over the full stack (harness, infrastructure, runtime) rather than relying on a single proprietary system. This stands in implicit contrast to OpenAI's recent approach, which El Hallack noted has slowed model training in response to security breaches—a bet that OpenAI is betting on restraint in the model rather than control in the harness.

FAQ

What is a harness in AI?
A harness is the software wrapper around an AI model—the tools, memory management, and rules that turn a raw model into something that can act independently. It includes components like a supervisor that guides the agent when it gets stuck or goes off track.
What is the ARC-AGI-3 benchmark?
ARC-AGI-3 is an interactive reasoning benchmark consisting of 2D games with no instructions, where the model must figure out how to play and win, similar to how a human would.
Did OpenAI achieve the same 100% score?
No. OpenAI's models tripled their scores by tweaking two harness settings but none came close to hitting 100%, the score Nvidia's researchers achieved.

Also reported by Hacker News

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleNvidia Takes Stake in Power Infrastructure Play as AI Deployment Bottleneck

The AI news that matters, in one minute each morning.

Sign up free