AIToday

DeepMind Kaggle Prize Criticized for Rewarding 'Nonsensical' Benchmark Work

r/MachineLearning22h ago

Key takeaway

A Google DeepMind-sponsored Kaggle competition on cognitive AI benchmarks awarded its $25K grand prize this week to a submission that a researcher claims lacks rigorous methodology and sound reasoning. The winning entry was meant to measure whether language models change their views when exposed to alternative perspectives, but critics say it became an oversized, incoherent work that neither authors nor judges appear to have properly reviewed.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    A Google DeepMind-sponsored Kaggle competition titled 'Measuring Progress Toward AGI - Cognitive Abilities' announced results this week, awarding a $25K grand prize to a submission that a critic claims is poorly reasoned and unfounded.

  • Why it matters

    The winning entry purported to test whether large language models change their assessment when presented with alternative viewpoints on claims, but the critic argues the actual work consists of unclear methodology, unfounded claims, and exceeded submission format constraints — raising questions about the quality of benchmarks being used to measure AI progress toward artificial general intelligence.

  • What to watch

    The criticism centers on whether DeepMind and Kaggle's review process adequately vetted a submission the critic describes as 'a vibed pile of spaghetti 10 times the size of the requested submission format,' suggesting potential gaps in how competitive AI benchmarking is evaluated.

In Depth

A Google DeepMind-sponsored Kaggle competition inviting novel cognitive-science-based benchmarks for measuring progress toward artificial general intelligence (AGI) has just announced its results. The grand prize winner received $25K. However, a critic with expertise in machine learning has publicly questioned the award, alleging that the winning submission is fundamentally unsound. According to the critic's analysis, the original intent of the winning work was legitimate: to test whether a large language model would modify its own assessment of five claims in a complex scenario if first presented with alternative viewpoints on those same claims from other LLMs. This is described as an 'interesting question' from a cognitive-science perspective. In practice, however, the critic argues the submission deteriorated into what they describe as 'a vibed pile of spaghetti 10 times the size of the requested submission format.' The work reportedly contains unfounded claims and unclear methodology throughout. The critic further contends that neither the authors nor the DeepMind and Kaggle judges appear to have conducted a careful review of the material, or else did not see fit to reject it despite its apparent deficiencies. The critic has documented their concerns across two posts in the competition forum, offering what they characterize as 'AI research slop detective work' in support of their argument that a submission of questionable rigor was nonetheless awarded the competition's top prize.

Context & Analysis

The Kaggle competition 'Measuring Progress Toward AGI - Cognitive Abilities' was designed to crowdsource new benchmarks for assessing how far AI systems have progressed toward artificial general intelligence. A submitted entry won the $25K grand prize, but a researcher has now publicly challenged the decision, claiming the winning work is fundamentally flawed. The intended research question — whether language models shift their conclusions when exposed to alternative LLM perspectives on the same claims — is described as 'interesting,' but the critic argues execution fell far short. The submission apparently ballooned to roughly ten times the requested size and consists largely of unclear reasoning and unsupported assertions, yet appears to have passed DeepMind and Kaggle's review process without substantial scrutiny. This incident may reflect a broader tension in automated AI benchmark design: the challenge of ensuring that novel benchmarks themselves meet rigorous standards, particularly when competition volume or submission complexity makes thorough review difficult.

FAQ

What was the competition about?
The Google DeepMind-sponsored Kaggle challenge 'Measuring Progress Toward AGI - Cognitive Abilities' asked participants to design new cognitive-science-based AI benchmarks.
What was the prize amount?
The grand prize was $25K.
What did the winning submission attempt to test?
The work aimed to present a large language model with alternative viewpoints of other LLMs on five claims regarding a tricky situation and see whether the model changed its own assessment.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime

1 minute a day. The AI essentials.

200+ sources · Email / LINE / Slack

Get it free →