The Pelican Test: Why Every AI Benchmark Eventually Dies
Simon Willison · 2026
"A benchmark that becomes famous enough to matter also becomes a target — and once labs and models start optimizing (deliberately or not) for whatever the benchmark measures, its correlation with real capability quietly decays, even while the benchmark itself keeps getting cited."
In 2024, Simon Willison started asking every new AI model to 'generate an SVG of a pelican riding a bicycle' as an informal way to compare them. For about a year, the quality of the pelican tracked real model quality surprisingly well. By mid-2026, that correlation had mostly broken: GPT-5.6 and Claude Fable 5's pelicans are outclassed by GLM-5.2, a model nobody seriously considers Fable-class.
This is a live, ongoing case of Goodhart's Law — 'when a measure becomes a target, it ceases to be a good measure' — playing out in public, one model release at a time. Willison is careful not to claim labs are deliberately training on his specific prompt; he thinks it's more likely that general-purpose capability gains and general-purpose overfitting to popular benchmark *styles* (SVG generation, in this case) have simply diverged from the underlying reasoning and agentic-tool-use ability that actually matters now. The tell is that the pelican test never measured what came to matter most: reliable multi-step tool use over long agentic conversations. A benchmark can decay in relevance even without anyone gaming it directly — the ground truth of what 'good' means can simply move faster than the benchmark can track it.
According to Willison, why did the pelican-riding-a-bicycle test stop correlating well with real model quality after its first year?
Read more about the topic
The explanation above is written with AI assistance. These are the originals — go to them to check it.
- Kimi K3, and what we can still learn from the pelican benchmarkSimon Willison's Weblog
- Welcome Inkling by Thinking MachinesHugging Face Blog
Situational Awareness: The Decade Ahead
"AGI by ~2027 and superintelligence shortly after will trigger a trillion-dollar compute buildout and a US-China race."