AIToday
Large Language ModelsImport AIPublished: Aug 18, 2026, 01:00 JST3 min read

AI shows early scientific taste; Fable beats peers on discovery games

AI shows early scientific taste; Fable beats peers on discovery games

Key takeaway

  • Researchers have released two new tools showing AI systems making strides in scientific reasoning and autonomous discovery. DiG-bench, a benchmark of 70 games with hidden rules, shows that Fable 5 and Opus 5 outperform other models but still lag significantly behind humans—none solved more than a handful of the hardest tier.

  • Inherent's Faraday model, a 27B supervisor trained on replicated research papers, beat larger frontier models on majority of replication tasks, suggesting AI systems are developing "taste" for research design.

  • Both advances point toward AI systems that can autonomously discover and reason about novel problems, a capability experts see as a prerequisite for systems capable of inventing solutions without human instruction.

3 Key Points

  1. What happened

    Researchers released DiG-bench, a benchmark of 70 games designed to test whether AI systems can discover hidden rules through exploration. Fable 5 and Opus 5 (with Claude Code) performed best overall, but only Opus 5 and Fable 5 beat any tasks in Tier 7 (the hardest level), achieving a 0.2 success rate. Meanwhile, Inherent published results for Faraday, a 27B supervisory model that beat standard Opus 4.8 and GPT-5.5 on 73% of in-distribution ML replication tasks and 60% of held-out AI-for-science tasks.

  2. Why it matters

    These benchmarks isolate prerequisites for autonomous creativity—the ability to discover useful undocumented rules in novel situations. The results suggest frontier models are already capable of some impressive feats of discovery but still lag far behind humans; for instance, a 20% success rate on Tier 7 compares poorly to individual humans achieving 100% on the tests. As AI systems improve at this kind of intuitive task discovery, they may become capable of recursive self-improvement—designing and advancing their own research.

  3. What to watch

    DiG-bench's 21 public games are playable at digbench.ai with a leaderboard showing model rankings. The author estimates human parity on DiG-bench could arrive by middle of 2027, at which point recursive self-improvement is expected to "seriously kick off." Faraday's approach—using a small post-trained model to supervise frontier models—illustrates how scaling these capabilities may track advances in frontier coding models over time.

Ask the AI about this article →

Context & Analysis

The two research projects reflect a shared focus on isolating and measuring prerequisites for autonomous discovery and invention in AI systems. DiG-bench frames this as the ability to infer hidden environmental rules through exploration and curiosity—a foundational skill that humans excel at but frontier models still struggle with significantly. The benchmark's seven-tier structure reveals a steep performance cliff: models that handle Tier 4 stumble dramatically on Tiers 6 and 7, with only the highest-capability models (Opus 5 and Fable 5) achieving even minimal success at the hardest level. Inherent's Faraday takes a different but complementary approach, treating scientific research as a domain where AI systems must develop intuition about what experiments to run, how to scope them to available resources, and how to judge results. Both efforts implicitly measure what researchers call "taste"—the capacity to make nuanced judgments about which problems are worth solving and which approaches are promising—in controlled, evaluable settings.

The significance both teams attach to these tasks stems from their connection to recursive self-improvement. If AI systems learn to discover hidden rules, design novel experiments, and make taste-based judgments about research direction, they may eventually extend those capabilities to improving themselves rather than merely serving human-specified goals. The Import AI newsletter author predicts human-level performance on DiG-bench by mid-2027, framing that milestone as a point at which recursive self-improvement could "seriously kick off." Inherent's team makes a similar argument: the skills Faraday acquires—deciding what to investigate, scoping experiments, judging outcomes—may be the same skills that allow a system to advance the state of the art autonomously. Both teams present these developments as capabilities-centric contributions to understanding how far AI systems have come; neither claims the systems are at human level, but both suggest the gap is narrowing in directions that matter.

FAQ

What are the DiG-bench games and how do you play them?
DiG-bench consists of 70 text-based games, each a self-contained miniature world with hidden rules and objectives that players must uncover through interaction. All games fit within the context window of current frontier models, 21 are released publicly, and the rest are kept private to prevent AI training on them. You can play some games online at digbench.ai.
How does Faraday work and what makes it different?
Faraday is a 27B model post-trained on Qwen-3.6-27B that acts as a supervisory layer atop frontier models like Claude Opus. It was trained via GRPO on a dataset called Replica—100 ML and AI-for-science papers with key results removed—to learn how to autonomously design experiments that fill in the blanks. The model beat standard Opus 4.8 and GPT-5.5 on 73% of in-distribution ML tasks and 60% of held-out AI-for-science tasks.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleHeidi Scribe automates clinical admin work across 190 countries, 2.7M patient interactions weekly

The AI news that matters, in one minute each morning.

Sign up free