The Pelican Test: Why Every AI Benchmark Eventually Dies
Simon Willison · 2026
"A benchmark that becomes famous enough to matter also becomes a target — and once labs and models start optimizing (deliberately or not) for whatever the benchmark measures, its correlation with real capability quietly decays, even while the benchmark itself keeps getting cited."
In 2024, Simon Willison started asking every new AI model to 'generate an SVG of a pelican riding a bicycle' as an informal way to compare them. For about a year, the quality of the pelican tracked real model quality surprisingly well. By mid-2026, that correlation had mostly broken: GPT-5.6 and Claude Fable 5's pelicans are outclassed by GLM-5.2, a model nobody seriously considers Fable-class.
Read more about the topic
Situational Awareness: The Decade Ahead
"AGI by ~2027 and superintelligence shortly after will trigger a trillion-dollar compute buildout and a US-China race."