AIToday
Large Language ModelsAI Safety & AlignmentLessWrong AIPublished: Aug 18, 2026, 16:01 JST2 min read

Claude models show near-perfect link between raw capability and decision-theory preference

Claude models show near-perfect link between raw capability and decision-theory preference

Key takeaway

  • Anthropic's Claude models display a near-perfect correlation between raw capability and a preference against CDT (Causal Decision Theory) answers, with r=0.97 when measured by DTBench and r=0.95 by TextArena.

  • This tight link does not appear in OpenAI models, where the same correlation is only r=0.55 or r=0.44, suggesting Claude's design or training creates a unified scaling law between capability and decision-theoretic reasoning preference that is absent in GPT.

3 Key Points

  1. What happened

    Researchers found that for Anthropic's Claude models, capability (measured by DTBench or TextArena benchmarks) correlates almost perfectly with a preference against CDT (Causal Decision Theory) answers — with correlation coefficients of r=0.97 and r=0.95 respectively. For Claude flagship models, this correlation is essentially identical to the correlation with release date (r=0.97).

  2. Why it matters

    The finding suggests that for Claude, higher capability and philosophical preference for EDT/FDT/UDT over CDT are not separate properties but emerge together as the model scales. This differs sharply from OpenAI models, where the same correlations are much weaker (r=0.55 for capability vs. CDT preference, r=0.44 for TextArena, r=0.45 for release date), indicating Claude's training may systematize decision-theoretic reasoning differently than GPT.

  3. What to watch

    Anthropic has already replicated this pattern in their Opus 4.7 and Fable 5 model cards. The correlation holds across different capability measures (DTBench and TextArena), though TextArena and DTBench themselves show a higher correlation within Anthropic models than when all models are pooled.

Ask the AI about this article →

Context & Analysis

The research documents a striking divergence between how Anthropic and OpenAI models scale decision-theoretic reasoning. For Claude, the near-perfect correlation (r=0.97) between capability and EDT/FDT/UDT preference suggests that as the model becomes more capable, it naturally gravitates toward these decision theories. This is reinforced by the fact that for flagship models, this correlation mirrors the correlation with release date (r=0.97), implying that capability and decision-theoretic preference advance in lockstep as Anthropic releases newer versions.

In contrast, OpenAI models show a weak correlation (r=0.55 or lower) between capability and CDT preference, indicating that raw capability and decision-theoretic leaning are largely independent properties in GPT. This gap points to a possible difference in training objective, alignment approach, or architectural bias. Anthropic's results (including those on Opus 4.7 and Fable 5) confirm the pattern is reproducible within their model family. The finding is further grounded by the observation that TextArena and DTBench themselves correlate more strongly within Anthropic models than across the broader model population, suggesting Anthropic's models behave more consistently along both capability and decision-theory axes.

FAQ

What exactly is being measured in this correlation?
Researchers measured capability using two benchmarks (DTBench and TextArena) and correlated those scores with whether models favor EDT/FDT/UDT over CDT answers, as measured by DTBench. For Anthropic flagships, they also checked correlation with release date.
Why is Claude's pattern so different from GPT's?
For Anthropic models, the correlation between capability and CDT preference is r=0.97–0.95; for OpenAI models, it is only r=0.55–0.44. The body does not explain why, but the gap suggests Anthropic and OpenAI may train or architect their models differently with respect to decision-theoretic reasoning.
Does this pattern hold across all Claude models?
The finding holds for Anthropic's flagship models and is replicated in Opus 4.7 and Fable 5 model cards. When all Anthropic models are included, the overall correlation between capability and CDT preference remains around r=0.8.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleCursor launches Origin to host code, competes with GitHub

The AI news that matters, in one minute each morning.

Sign up free