Even the best AI models lose about half their performance when charts get complicated, new benchmark finds

TL;DR AI
2 min readKey summary
Researchers unveiled RealChart2Code, a benchmark built from real Kaggle datasets that tests 14 AI models on chart replication, reproduction, and refinement.
Top proprietary models such as Claude 4.5 Opus, Gemini 3 Pro Preview, and GPT-5.1 performed far worse on complex visualizations than on simpler chart tasks.
Open-weight models dropped even more sharply, often producing invalid code or handling the underlying data incorrectly.
The benchmark highlights a major gap between current model ability on simple chart tasks and the demands of real-world visualization workflows.



