Switch language한국어
Back to the list

Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth

TL;DR AI

Key summary

2 min read
  1. Researchers introduced BonaFide, a benchmark with 3,066 labeled chain-of-thought examples across 13 tasks and 10 models.

  2. When major faithfulness metrics were tested against ground-truth labels, most performed near chance, showed bias, and struggled on longer reasoning traces.

  3. The study also found weak transferability across settings and high compute costs for many metrics.

  4. The results raise concerns that popular faithfulness scores may not reliably measure whether LLM reasoning traces are truly faithful.

Read the original