On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientists
TL;DR AI
2 min readKey summary
A study had 45 domain scientists assess 2,960 criticisms from human and AI reviews of 82 Nature-family papers.
AI reviewers often identified correct, significant, and well-supported criticisms better than humans in some cases.
But they still struggled with specialized subfield knowledge and long-context reasoning, and AI models tended to overlap with each other more than human reviewers.
The results offer rare large-scale evidence on where AI can strengthen peer review and where human expertise remains essential.
