Our First Proof submissions

TL;DR AI
2 min readKey summary
OpenAI says its internal model produced proof attempts for all 10 First Proof math problems.
Expert review judged at least five attempts likely correct, with several others still under review, while problem 2 is now believed wrong after further analysis.
The run used limited human supervision, and some answers were refined with ChatGPT for verification and formatting.
First Proof is designed to test whether AI can generate correct, checkable proofs in hard research domains, making it a stronger reasoning benchmark.



