Switch language한국어
Back to the list

Convergence Without Understanding: When Language Models Agree on Representations but Disagree on Reasoning

TL;DR AI

Key summary

2 min read
  1. A study of 16 language models across eight families and 800 tasks found that hidden-state similarity often does not match shared reasoning.

  2. Models looked more alike on problems they answered incorrectly, suggesting representational similarity can rise when reasoning fails.

  3. Pre-decision states were similar, but post-decision states diverged, showing that internal trajectories split after choosing an answer.

  4. Although some shared information was decodable, it had little causal impact on outputs, challenging simple interpretability claims and ensemble assumptions.

Read the original