OpenAI claims that "AI capabilities may not be being measured correctly"

TL;DR AI
2 min readKey summary
OpenAI released a playbook for third-party evaluation of frontier models, arguing that tests must cover not just the model, but the harness, tools, budget, and attack conditions.
It warned that factors like reward hacking, refusal behavior, data contamination, broken tasks, and strategic underperformance can distort benchmark results.
The guidance pushes for better standards to measure capability and safety as AI systems become more tool-using and multi-step.
The move highlights the limits of traditional benchmarks for models such as GPT-5.5, GPT-5.4, and Codex, and calls for more reliable evaluation practices.



