
Moonshot AI's new PerceptionBench benchmark tests how well AI models can actually "see" by isolating visual perception from reasoning across ten sub-skills.
No frontier model reached 60 percent accuracy—GPT-5.6 Sol led at 59.7 percent—and the results reveal that many failures thought to be reasoning errors actually occur when models misread images.
The benchmark exposes a consistent weakness: models diverge sharply on individual visual skills despite similar overall scores, signaling that visual perception remains a significant limitation across the industry's leading systems.
What happened
Moonshot AI released PerceptionBench, a new benchmark that isolates visual perception from reasoning by testing ten atomic sub-skills (Visual Relation, Counting, Attributes, Depth & 3D, Localization, Comparison, Fine-grained Recognition, Context Integration, OCR, and Hallucination). The 3,000-task benchmark tested 16 frontier models; the highest score was 59.7 percent by GPT-5.6 Sol, followed by Kimi K3 at 58.5 percent, Claude Fable 5 at 57.2 percent, and Gemini 3.1 Pro at 56.2 percent.
Why it matters
The benchmark reveals that many failures attributed to reasoning errors actually originate in the perception stage—when models misread images. By breaking multi-step visual tasks into perception-only sub-questions, PerceptionBench pinpoints exactly which visual skills are failing across models. This has practical implications for businesses and developers relying on multimodal AI: visual understanding remains a significant weakness, with no model yet achieving even 60 percent accuracy on basic visual tasks that the body indicates may seem trivial on the surface.
What to watch
The sharp divergence in category-level performance is striking—models with nearly identical overall scores vary dramatically on individual skills. For example, GPT-5.6 Sol scores only 26.9 percent on hallucination (detecting objects that don't exist), while Gemini 3.5 Flash, a weaker overall performer, ranks among the best at 50.6 percent on that same task. The dataset and evaluation code are available on GitHub at MoonshotAI/PerceptionBench.
Moonshot AI, the team behind the Chinese AI assistant Kimi, has introduced PerceptionBench, a new benchmark designed to isolate and measure the visual perception capabilities of multimodal language models independently from reasoning and knowledge. The benchmark consists of 3,000 publicly released tasks drawn from an internal pool of over 17,000 verified questions; 60 percent were derived from actual attributed model errors, while 40 percent were reformulated using augmented images. Unlike standard evaluations that bundle perception, knowledge, and reasoning into a single assessment, every question in PerceptionBench can be answered by looking at the image alone, with no reasoning or outside knowledge required.
The authors built the benchmark by analyzing 42 existing open-source perception benchmarks and found they contain little overlap in their error profiles—each covers a different subset of visual weaknesses. Rather than imposing categories from theory, the team extracted them from real model errors, tracing each failure back to the earliest failed step in existing benchmarks. This produced ten "skill domains": Visual Relation, Counting, Attributes, Depth & 3D, Localization, Comparison, Fine-grained Recognition, Context Integration, OCR, and Hallucination. The tasks themselves appear deceptively simple on the surface: figuring out where a symbol sits on a clock face, counting flowers inside a red box, or distinguishing which of two pencil cups shows a gray-pink combo versus a solid pink with cartoon design.
When PerceptionBench was tested across 16 frontier models, no model reached 60 percent accuracy. GPT-5.6 Sol led with 59.7 percent, followed by Kimi K3 at 58.5 percent, Claude Fable 5 at 57.2 percent, Gemini 3.1 Pro at 56.2 percent, and GPT-5.5 at 55.8 percent. Open-source models trailed significantly: Qwen3.5-397B-A17B scored 47.5 percent, and GLM-4.6V scored 32.5 percent. The gap between the top five systems spans just under four percentage points. However, the category-level results reveal a more complex picture. Models with nearly identical overall scores diverge sharply on individual skills. Across all models tested, "hallucination" is the weakest domain—the ability to recognize when an object does not exist and answer "zero." GPT-5.6 Sol, despite leading overall, scored only 26.9 percent on hallucination, while Gemini 3.5 Flash, a weaker performer overall, ranked among the best at 50.6 percent.
The authors argue that many failures typically attributed to "reasoning errors" in multimodal models actually originate at the perception stage. When a model fails a multi-step task, the first step—correctly reading the image—has often already failed. PerceptionBench makes it possible to isolate and pinpoint which specific visual ability is failing. The dataset and evaluation code are available on GitHub. This work is part of a broader research effort: the same team previously released WorldVQA, a benchmark that separates object recognition from reasoning, where Gemini 3 Pro achieved 47.4 percent—below 50 percent—and all models systematically overestimated their own confidence. Another parallel study involving Chinese institutions and Moonshot AI used the BabyVision benchmark to show that frontier models fail at basic visual tasks tied to early childhood development, such as tracing lines or counting hidden blocks; Gemini 3 Pro scored 49.7 percent while humans achieved 94.1 percent. Researchers attribute this gap to a verbalization bottleneck where visual information is translated into language and loses fidelity. Meanwhile, Kimi K3, Moonshot AI's open-source offering, has closed the gap to Claude Fable 5 and GPT-5.6 Sol to within a few points on general benchmarks, though it still lags in specialized areas like offensive cybersecurity and complex math. On visual perception as measured by PerceptionBench, K3 now performs on par with its Western rivals.
PerceptionBench addresses a fundamental gap in how the AI industry has been evaluating multimodal models. The existing landscape includes 42 open-source perception benchmarks, yet the authors found they show little overlap in their error profiles—each covers a different subset of visual weaknesses. This fragmentation meant no single test or small group could capture visual perception as a whole. By building a taxonomy from actual model errors rather than theory, the benchmark reveals what previous evaluations have obscured: the specific visual skills where frontier models systematically fail. The finding that many so-called reasoning errors actually originate in the perception stage has immediate implications for how AI teams should debug and improve their systems. If a model fails a multi-step visual task, the error may not lie in the logical reasoning that follows, but in the foundational step of accurately reading the image itself.
The category-level divergence is particularly telling. Models like GPT-5.6 Sol and Gemini 3.5 Flash have vastly different overall rankings, yet their hallucination performance differs by nearly 24 percentage points in opposite directions—Sol weak, Flash comparatively strong. This pattern suggests that visual perception is not a monolithic capability but a collection of distinct sub-skills, and current frontier models have uneven strength across them. The body notes this reflects a broader pattern: a separate study using the BabyVision benchmark showed that frontier models fail at basic visual tasks tied to early childhood development (line tracing, counting hidden blocks), with Gemini 3 Pro at 49.7 percent versus a human baseline of 94.1 percent. The authors attribute this gap to a verbalization bottleneck where visual information is translated into language and loses fidelity—a constraint that may explain why even the most capable models cannot break 60 percent on PerceptionBench.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Mastercard CEO Michael Miebach told The Motley Fool that cybersecurity damage will reach $15.6 trillion by 203…

Meta CEO Mark Zuckerberg published a 6,500-word manifesto titled "The Future is for Everyone" outlining a visi…

Snowflake announced general availability of a redesigned Observe MCP server and a new Observe CLI with full fe…

Twitch has added a new account setting that lets streamers opt out of having their content used to train Amazo…

Snowflake announced general availability of a redesigned Observe MCP server and new Observe CLI with full pari…

Snowflake announced general availability of a redesigned Observe MCP (model context protocol) server and new O…

The AI news that matters, in one minute each morning.
Sign up free