This is what happens when AI benchmark “Draw a pelican riding a bicycle” is tried on Gemini 3.1 Pro and Qwen3.6-35B-A3B

TL;DR AI
2 min readKey summary
At PyCon US 2026, Simon Willison revisited his long-running “pelican riding a bicycle” image benchmark.
He said newer models like Gemini 3.1 Pro and Qwen3.6-35B-A3B produced especially strong, more natural-looking results.
Other models, including Claude Sonnet 4.5, Claude Opus 4.7, Gemma 4, and GLM-5.1, also showed clear progress.
Willison stressed that the benchmark is useful for spotting model quirks and strengths, but it is far too narrow to judge overall AI quality.
