Skip to content
← Home
Contemporary · Blog Post

The Pelican Test: Why Every AI Benchmark Eventually Dies

Simon Willison · 2026

"A benchmark that becomes famous enough to matter also becomes a target — and once labs and models start optimizing (deliberately or not) for whatever the benchmark measures, its correlation with real capability quietly decays, even while the benchmark itself keeps getting cited."

The idea

In 2024, Simon Willison started asking every new AI model to 'generate an SVG of a pelican riding a bicycle' as an informal way to compare them. For about a year, the quality of the pelican tracked real model quality surprisingly well. By mid-2026, that correlation had mostly broken: GPT-5.6 and Claude Fable 5's pelicans are outclassed by GLM-5.2, a model nobody seriously considers Fable-class.

Why it works
The takeaway — recall it first
Further reading

Read more about the topic

Up NextSuggested: Continues the theme of Worldview & Futurism

Situational Awareness: The Decade Ahead

"AGI by ~2027 and superintelligence shortly after will trigger a trillion-dollar compute buildout and a US-China race."

Leopold Aschenbrenner · EssayContinue→
Listen
0 / 3