Switch language한국어
Back to the list

SoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones?

TL;DR AI

Key summary

2 min read
  1. Researchers introduced SoundnessBench, a benchmark built from 1,099 reconstructed ML proposals from ICLR submissions with reviewer soundness labels and source-paper audits.

  2. Across 12 frontier LLMs, the models showed a strong optimism bias: under standard prompts, they often marked weak ideas as sound.

  3. Stricter prompting changed the kinds of mistakes more than the overall failure rate, suggesting the issue is not solved by prompt tuning alone.

  4. The authors also examined contamination and other confounds, and concluded no single simple cause explains the behavior.

Read the original