
OpenAI announced that its GPT-5.6 Sol model achieves 38.3 percent on the ARC-AGI-3 logic benchmark—above Anthropic's Claude Opus 5 at 30.2 percent—but only when using OpenAI's custom API settings that preserve reasoning and compress context. In the official test environment, the same model scored just 7.8 percent, raising questions about the fairness of benchmark comparisons when providers use different technical setups. ARC Prize's founder acknowledged the tension but said such differences are acceptable if clearly disclosed.
Summaries like this, in your inbox every morning.
Sign up free →What happened
OpenAI reported that GPT-5.6 Sol scored 38.3 percent on the ARC-AGI-3 logic benchmark, surpassing Anthropic's Claude Opus 5 score of 30.2 percent. The higher score was achieved using OpenAI's custom Responses API with "Retained Reasoning" (which preserves the model's chain of thought between steps) and "Compaction" (which summarizes old context instead of discarding it). In the official test harness without these settings, GPT-5.6 Sol scored only 7.8 percent.
Why it matters
ARC-AGI-3 is designed to test pure model performance using a standardized approach to ensure fair comparisons across providers. The dispute centers on whether OpenAI's use of provider-specific API settings—unavailable in the official testing environment—undermines the benchmark's fairness. ARC Prize co-founder François Chollet acknowledged that Anthropic may have used an older API lacking features the Claude API already had, putting OpenAI at a disadvantage in the official test, but he drew a distinction: general-purpose API settings "that were not developed for ARC-AGI-3 and that are available to all API users" are acceptable, while harnesses "custom-made to solve the benchmark" are not.
What to watch
Chollet noted that different providers using different settings creates "a potential parity issue," but he considers it acceptable "as long as the settings and the cost are clearly reported." The back-and-forth highlights an ongoing challenge in AI benchmarking: the difficulty of isolating model capability from the technical infrastructure around it.
After Anthropic's Claude Opus 5 set a record on the ARC-AGI-3 logic benchmark by quadrupling the previous high score, OpenAI demonstrated that its GPT-5.6 Sol model can surpass that result—but the claim immediately sparked a methodological dispute.
OpenAI reported that GPT-5.6 Sol achieves 38.3 percent on ARC-AGI-3, beating Opus 5's 30.2 percent. However, this score was achieved using OpenAI's custom Responses API configured with two non-standard settings: "Retained Reasoning" and "Compaction." The Retained Reasoning setting preserves the model's chain of thought reasoning between sequential steps, whereas the standard approach discards the reasoning after each action. Compaction summarizes accumulated context instead of truncating it. When evaluated in the official test harness—which does not include these settings—GPT-5.6 Sol scored only 7.8 percent, a dramatic drop that underscores how heavily the higher result depends on OpenAI's custom configuration.
ARC-AGI-3 is designed specifically to measure pure model performance, and ARC Prize administers it using a standardized approach to ensure fair comparison across all providers. In response to OpenAI's results, ARC Prize emphasized this principle and questioned whether the benchmark conditions were truly equivalent when OpenAI tested its model using features unavailable during the official test runs.
OpenAI defended its approach by arguing that benchmarks inherently measure not just the model itself but the technical infrastructure surrounding it—a reasonable observation. However, the company raised a more pointed objection: it suggested that when ARC Prize tested Anthropic's Claude, it may have used an older "OpenAI-style completions API" that lacked capabilities already present in the Claude API, placing OpenAI at a disadvantage in the official environment.
François Chollet, co-founder of the ARC Prize, responded by drawing a distinction between two types of testing configurations. Harnesses "custom-made to solve the benchmark or that contain knowledge about the benchmark format" violate the spirit of fair testing. By contrast, general-purpose API settings "that were not developed for ARC-AGI-3 and that are available to all API users" are permissible. In effect, Chollet conceded that the original test setup may have inadvertently favored Anthropic. He noted that ARC Prize has engaged in "a lot of back and forth with OpenAI about how to best test their models, especially with regard to compaction," and welcomed OpenAI "starting to figure out the answer." While acknowledging that "different providers using different settings does create a potential parity issue," Chollet indicated this is acceptable "as long as the settings and the cost are clearly reported."
The dispute over GPT-5.6 Sol's ARC-AGI-3 score reflects a fundamental tension in AI benchmarking: the benchmark measures not just raw model capability but the entire system surrounding it. OpenAI's argument—that benchmarks inherently depend on technical setup—is factually sound, yet ARC-AGI-3 was explicitly designed to test pure model performance using a standardized approach to ensure fair cross-provider comparison.
The crux of the disagreement concerns whether the comparison itself was inequitable from the start. OpenAI contends that Anthropic may have used an older API with fewer features when ARC Prize ran its tests, giving Claude an unfair advantage in the official environment. ARC Prize co-founder François Chollet's response draws a careful distinction: general-purpose API settings that are publicly available to all users and were not developed specifically for the benchmark are acceptable; custom harnesses designed to exploit the benchmark format are not. This distinction permits OpenAI's use of Retained Reasoning and Compaction—features now available to all API users—while protecting the benchmark's integrity against purpose-built workarounds.
Chollet acknowledged ongoing collaboration with OpenAI "about how to best test their models, especially with regard to compaction," suggesting the two organizations are working toward a framework where technical differences are transparent and comparable. The unresolved question is whether future benchmarks will require all providers to use identical settings, or whether clearly disclosed provider-specific configurations will become standard practice.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion




Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime