Switch language한국어
Back to the list

Even the best AI models lose about half their performance when charts get complicated, new benchmark finds

TL;DR AI

Key summary

2 min read
  1. Researchers unveiled RealChart2Code, a benchmark built from real Kaggle datasets that tests 14 AI models on chart replication, reproduction, and refinement.

  2. Top proprietary models such as Claude 4.5 Opus, Gemini 3 Pro Preview, and GPT-5.1 performed far worse on complex visualizations than on simpler chart tasks.

  3. Open-weight models dropped even more sharply, often producing invalid code or handling the underlying data incorrectly.

  4. The benchmark highlights a major gap between current model ability on simple chart tasks and the demands of real-world visualization workflows.

Read the original