A new benchmark called ActiveVision has exposed a significant weakness in frontier AI models' ability to handle visual perception tasks that require repeated observation rather than a single static analysis. GPT-5.5, at its highest reasoning tier, solved only 10.6% of the benchmark's 17 tasks, while Claude Fable 5 achieved 3.5%—a stark contrast to the 96.1% average achieved by three human participants. The models cannot overcome this limitation even when allowed to write their own code, pointing to a fundamental gap in how current AI approaches visual reasoning.
Summaries like this, in your inbox every morning.
Sign up free →What happened
A new benchmark called ActiveVision, containing 17 tasks across 3 categories designed to test repeated visual perception, reveals a significant gap between frontier AI models and human performance. GPT-5.5 at its highest reasoning-effort tier solves 10.6% of items and scores zero on 11 of the 17 tasks, while Claude Fable 5 manages 3.5%. Three human participants averaged 96.1%.
Why it matters
The finding highlights a specific failure mode that current reasoning models cannot overcome—even when given the ability to write their own code to solve problems. This suggests that visual reasoning tasks requiring sustained, repeated perception remain a weak point for frontier AI systems, not merely a temporary limitation that additional compute or code generation can patch.
What to watch
The ActiveVision benchmark itself may become a reference point for evaluating whether future models can close the gap on visual perception tasks. The paper is available on arXiv (2607.16165), and the detailed design of the benchmark's 17 tasks across 3 categories could inform how AI developers approach this class of problems.
A new arXiv paper (2607.16165) introduces ActiveVision, a benchmark designed to test a specific limitation in frontier vision models: their ability to handle tasks that require repeated visual perception rather than a single static description. The benchmark consists of 17 tasks distributed across 3 categories, each structured to force the model to integrate information across multiple observations rather than rely on a single-pass analysis.
When evaluated against this benchmark, GPT-5.5—tested at the highest exposed reasoning-effort tier—achieves a 10.6% solve rate and scores zero on 11 of the 17 tasks. Claude Fable 5, which the paper notes leads most reasoning and coding leaderboards, performs even worse at 3.5%. By contrast, three human participants averaged 96.1%, indicating that the tasks are well within human cognitive capability but appear to expose a fundamental mismatch between how humans and current AI models approach visual reasoning.
The paper's central observation is not simply that the models fail, but that they cannot repair this failure through code generation. Even when given the ability to write their own code to solve the problem, both models remain stuck at their low baseline performance. This suggests the limitation is not due to insufficient reasoning steps or compute, but rather stems from something more fundamental in how these models process and integrate visual information over repeated observations.
The ActiveVision benchmark reveals a specific architectural weakness in frontier vision models—not merely a gap in raw performance, but a failure mode that correlates with the task's demand for repeated visual perception rather than single-pass analysis. The paper notes that this gap persists even when models are given the ability to write code to solve the problem, suggesting that the limitation is not a shortage of computational resources or reasoning steps, but something deeper in how these models process and integrate visual information over time.
The disparity between the 10.6% and 3.5% scores achieved by GPT-5.5 and Claude Fable 5 respectively, against the 96.1% human average, illustrates a class of reasoning task where the gap has not narrowed as frontier models have scaled. This pattern—where models that lead on most reasoning and coding leaderboards still fail systematically on a specific visual-reasoning class—suggests that benchmarks designed to expose such failure modes may be necessary to identify and address remaining gaps in AI reasoning.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion



Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime