AIToday

AI labs show no sign of optimizing for pelican-bicycle prompts

Simon Willison's Weblog9h ago
AI labs show no sign of optimizing for pelican-bicycle prompts

Key takeaway

A researcher tested seven major AI models systematically to determine whether AI labs have been deliberately optimizing their models to excel at drawing pelicans riding bicycles—a quirky benchmark that has circulated among AI enthusiasts. Testing 48 different animal-vehicle combinations across the models, the analysis found no meaningful evidence of such optimization, with pelicans and bicycles drawing no better in combination than their individual performance would suggest.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    A researcher named Dylan Castillo conducted a systematic test across 7 models (GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen3.7-Max, GLM-5.2, and DeepSeek V4 Pro) using 48 prompts combining 8 animals and 6 vehicles, each run three times, to investigate whether AI labs have deliberately optimized models to draw pelicans riding bicycles. The evaluation found no meaningful evidence of such optimization.

  • Why it matters

    The analysis settles a long-standing informal question about whether AI labs might have tuned their models to perform better on this specific, whimsical prompt combination. The systematic methodology—using 144 total test runs across seven models and evaluating results with additional models—provides far more rigor than previous spot-checks, allowing a more confident conclusion that no special optimization has occurred.

  • What to watch

    Readers can explore the detailed results through a filter view that breaks down performance by animal, vehicle, and the combination thereof, showing that pelican-bicycle scenes are not drawn any better than general pelican or bicycle performance would predict.

In Depth

Simon Willison's blog post discusses a detailed investigation into a playful but persistent question in the AI community: whether major AI labs have been deliberately training their models to excel at generating images of pelicans riding bicycles. The original benchmark, which Willison himself posed, is admittedly deeply unscientific, but it has become a cultural touchstone for informal AI testing.

To settle this question with real rigor, researcher Dylan Castillo designed and executed a comprehensive experiment. He created 48 test prompts by combining 8 different animals with 6 different types of vehicles (8 animals × 6 vehicles = 48 combinations). He then ran each prompt three times through seven different AI models: GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen3.7-Max, GLM-5.2, and DeepSeek V4 Pro. To evaluate the results fairly, he used two additional models—GPT-5.6 Luna and Gemini 3.1 Flash-Lite—as evaluators. This systematic approach generated 144 total test runs (48 prompts × 3 runs) and provided a firm statistical foundation for drawing conclusions.

The results provide clear evidence against the hypothesis of deliberate optimization. Across all tested dimensions, the data shows that pelicans are not drawn any better than other animals, bicycles are not drawn any better than other vehicles, and pelicans on bicycles are not drawn any better than the combined baseline performance of pelican and bicycle rendering would predict. The analysis notes that GLM-5.2 comes closest to showing an effect, displaying the largest relative boost on the exact pelican-bicycle combination and producing one sample that caught Willison's eye. However, even this effect is described as small and not statistically significant, insufficient to suggest deliberate optimization.

Castillo made the detailed results accessible through a filter view that allows readers to explore the breakdown by animal, vehicle, and combination, enabling independent verification of the findings. This transparency and systematic approach represent a substantial upgrade from prior anecdotal spot-checking and effectively answer the question: AI labs have not been pelicanmaxxing.

Context & Analysis

The question of whether AI labs might deliberately optimize for the pelican-bicycle prompt has been an informal, tongue-in-cheek benchmark among AI observers for some time. Simon Willison, the author of this blog, notes that he had attempted some spot-checking in the past by testing models on other animals and vehicles, but with far less systematic methodology. Dylan Castillo's approach represents a significant step forward in rigor: by testing a full factorial design (8 animals × 6 vehicles = 48 unique prompts), running each three times, and evaluating across seven different models, the study provides a much stronger empirical foundation than anecdotal observation.

The findings are clear and consistent across multiple dimensions. The analysis shows that pelicans are not drawn any better than other animals, bicycles are not drawn any better than other vehicles, and the combination of pelicans on bicycles does not show a performance boost beyond what you would expect from adding the individual performance levels. The one partial exception is GLM-5.2, which shows a slight boost on the exact pelican-bicycle combination, but this effect is described as small and not statistically significant. This thorough investigation effectively closes the door on the hypothesis that AI labs have engaged in deliberate, measurable optimization for this particular prompt.

FAQ

Which AI models were tested in this experiment?
The researcher tested GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen3.7-Max, GLM-5.2, and DeepSeek V4 Pro, and used GPT-5.6 Luna and Gemini 3.1 Flash-Lite to help evaluate the results.
What exactly did the test measure?
The test used 48 prompts created from 8 animals and 6 vehicles, running each prompt three times through the 7 models, then evaluating whether pelicans, bicycles, or pelicans on bicycles were drawn any better than the model's general performance would predict.

Get the latest AI Safety & Alignment news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime

1 minute a day. The AI essentials.

200+ sources · Email / LINE / Slack

Get it free →