Measuring the performance of our models on real-world tasks

TL;DR AI
2 min readKey summary
OpenAI launched GDPval, a new AI evaluation focused on realistic knowledge-work tasks across 44 occupations and 9 major U.S. industries.
The benchmark includes 1,320 tasks, with 220 gold tasks open-sourced, and tests deliverables like briefs, diagrams, spreadsheets, slides, and multimedia.
GDPval is meant to measure performance on economically valuable work, not just academic-style benchmarks, giving a clearer view of real-world usefulness.
It aims to make AI progress more evidence-based by showing how models handle everyday professional workflows in practice.



