Switch language한국어
Back to the list

Survey Finds 5 Cutting-Edge AIs Disagreed on Claims 67% of the Time

TL;DR AI

Key summary

2 min read
  1. Lenz tested 1,000 user-submitted claims across five frontier AI models and found disagreement in 672 cases.

  2. Only 328 claims received unanimous ratings, showing that the models often judged the same claim differently.

  3. Some claims split all five models, underscoring the inconsistency of AI fact-checking on real-world inputs.

  4. Lenz said human-labeled answers are being used to study where model judgments diverge from people.

Read the original