Analysis finds Google AI Overviews is wrong 10 percent of the time

TL;DR AI
2 min readKey summary
The New York Times analyzed Google’s Gemini-powered AI Overviews and found it was accurate on about 90 percent of SimpleQA questions.
A startup called Oumi ran the SimpleQA benchmark of over 4,000 verifiable questions to test AI Overviews.
Oumi recorded 85% accuracy with Gemini 2.5 and 91% after the Gemini 3 update, but the analysis still documented notable factual errors.
Extrapolating the roughly 10% miss rate across all searches implies tens of millions of incorrect answers per day.
Examples include a wrong museum date for Bob Marley’s former home and a mistaken claim denying the Classical Music Hall of Fame.



