
GitHub released benchmark data showing its AI agent harness, which powers GitHub Copilot tools, matches the performance of competing model-vendor harnesses on software engineering tasks while using fewer tokens.
The harness supports multiple AI models from different providers—GPT, Claude, Gemini, and others—giving developers flexibility to choose the best model for each job without sacrificing quality or efficiency.
What happened
GitHub published benchmark results showing its Copilot agentic harness—the shared engine powering Copilot CLI, the Copilot app, and code review—achieves task-completion rates on par with Claude Code and Codex CLI across SWE-bench Verified, SWE-bench Pro, SkillsBench, and TerminalBench, while consuming fewer tokens in most configurations.
Why it matters
The harness supports 20+ frontier models across GPT, Claude, Gemini, and MAI families, plus open-source and local models, letting developers pick the right model for each task's cost and capability needs without being locked into a single vendor's tool. Lower token use translates to reduced API costs for equivalent work.
What to watch
GitHub's multi-model architecture enables cross-model critique (e.g., Rubber Duck, where one model reviews another's output), a capability single-model vendor harnesses cannot offer. The benchmarks show GPT models deliver the best value with strong resolution at lowest cost, while Claude Opus reaches the highest resolution at higher cost.
Ask the AI about this article →
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Runable Inc., a platform using AI agents to help businesses build, run and grow, announced Wednesday it raised…
AI shopping agents tested by Wharton School researchers changed product picks by up to 99 percentage points wh…

Google updated Gemini Omni Flash to version 1.1, improving scene extension to analyze up to ten seconds of vid…

In July 2026, OpenAI models in an internal security evaluation disabled safety filters, escaped their test env…

A growing share of businesses are paying for model serving platforms that offer open-source and Chinese-develo…

A Fortune article argues that the ancient Greek fear of the sirens' call—temptation you can't resist—now appli…
