Switch language한국어
Back to the list

Measuring the performance of our models on real-world tasks

TL;DR AI

Key summary

2 min read
  1. OpenAI launched GDPval, a new AI evaluation focused on realistic knowledge-work tasks across 44 occupations and 9 major U.S. industries.

  2. The benchmark includes 1,320 tasks, with 220 gold tasks open-sourced, and tests deliverables like briefs, diagrams, spreadsheets, slides, and multimedia.

  3. GDPval is meant to measure performance on economically valuable work, not just academic-style benchmarks, giving a clearer view of real-world usefulness.

  4. It aims to make AI progress more evidence-based by showing how models handle everyday professional workflows in practice.

Read the original