Switch language한국어
Back to the list

Review Arcade: On the Human Alignment and Gameability of LLM Reviews

TL;DR AI

Key summary

2 min read
  1. A study of 1,000 ACL 2025 submissions found LLM-generated reviews only weakly matched human judgments.

  2. Review scores varied significantly across prompts and repeated runs, showing poor stability.

  3. Researchers also showed scores could be improved iteratively without substantive paper changes.

  4. The results raise concerns about using LLMs as reliable peer reviewers in conference systems.

Read the original