
Researchers have released two new tools showing AI systems making strides in scientific reasoning and autonomous discovery. DiG-bench, a benchmark of 70 games with hidden rules, shows that Fable 5 and Opus 5 outperform other models but still lag significantly behind humans—none solved more than a handful of the hardest tier.
Inherent's Faraday model, a 27B supervisor trained on replicated research papers, beat larger frontier models on majority of replication tasks, suggesting AI systems are developing "taste" for research design.
Both advances point toward AI systems that can autonomously discover and reason about novel problems, a capability experts see as a prerequisite for systems capable of inventing solutions without human instruction.
What happened
Researchers released DiG-bench, a benchmark of 70 games designed to test whether AI systems can discover hidden rules through exploration. Fable 5 and Opus 5 (with Claude Code) performed best overall, but only Opus 5 and Fable 5 beat any tasks in Tier 7 (the hardest level), achieving a 0.2 success rate. Meanwhile, Inherent published results for Faraday, a 27B supervisory model that beat standard Opus 4.8 and GPT-5.5 on 73% of in-distribution ML replication tasks and 60% of held-out AI-for-science tasks.
Why it matters
These benchmarks isolate prerequisites for autonomous creativity—the ability to discover useful undocumented rules in novel situations. The results suggest frontier models are already capable of some impressive feats of discovery but still lag far behind humans; for instance, a 20% success rate on Tier 7 compares poorly to individual humans achieving 100% on the tests. As AI systems improve at this kind of intuitive task discovery, they may become capable of recursive self-improvement—designing and advancing their own research.
What to watch
DiG-bench's 21 public games are playable at digbench.ai with a leaderboard showing model rankings. The author estimates human parity on DiG-bench could arrive by middle of 2027, at which point recursive self-improvement is expected to "seriously kick off." Faraday's approach—using a small post-trained model to supervise frontier models—illustrates how scaling these capabilities may track advances in frontier coding models over time.
Ask the AI about this article →
The two research projects reflect a shared focus on isolating and measuring prerequisites for autonomous discovery and invention in AI systems. DiG-bench frames this as the ability to infer hidden environmental rules through exploration and curiosity—a foundational skill that humans excel at but frontier models still struggle with significantly. The benchmark's seven-tier structure reveals a steep performance cliff: models that handle Tier 4 stumble dramatically on Tiers 6 and 7, with only the highest-capability models (Opus 5 and Fable 5) achieving even minimal success at the hardest level. Inherent's Faraday takes a different but complementary approach, treating scientific research as a domain where AI systems must develop intuition about what experiments to run, how to scope them to available resources, and how to judge results. Both efforts implicitly measure what researchers call "taste"—the capacity to make nuanced judgments about which problems are worth solving and which approaches are promising—in controlled, evaluable settings.
The significance both teams attach to these tasks stems from their connection to recursive self-improvement. If AI systems learn to discover hidden rules, design novel experiments, and make taste-based judgments about research direction, they may eventually extend those capabilities to improving themselves rather than merely serving human-specified goals. The Import AI newsletter author predicts human-level performance on DiG-bench by mid-2027, framing that milestone as a point at which recursive self-improvement could "seriously kick off." Inherent's team makes a similar argument: the skills Faraday acquires—deciding what to investigate, scoping experiments, judging outcomes—may be the same skills that allow a system to advance the state of the art autonomously. Both teams present these developments as capabilities-centric contributions to understanding how far AI systems have come; neither claims the systems are at human level, but both suggest the gap is narrowing in directions that matter.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
OpenAI's GPT-5.6 Sol, launched July 9, drove a 35 percent revenue increase this quarter, with enterprise reven…

OpenAI is previewing transparent background support for GPT-Image-2 through its API, allowing users to generat…

HP Korea has formed a partnership with Upstage, a large language model (LLM) startup, to advance its localized…

At the "AI on Chips: Semiconductor Industry Trends Forum" hosted by DIGITIMES, industry experts highlighted th…

Stripe is in acquisition talks to buy AI startup OpenRouter for more than US$7 billion, according to Bloomberg…

Anthropic launched Claude Academy on August 20, a free learning site that explains AI fundamentals and how to…
