Switch language한국어
Back to the list

Where Does Authorship Signal Emerge in Encoder-Based Language Models?

TL;DR AI

Key summary

2 min read
  1. A new study finds that authorship attribution accuracy in encoder-based language models can vary by up to 4× just by changing the scoring method.

  2. Researchers compared models built on the same pretrained encoder and showed that the scorer, not the encoder itself, drove the performance gap.

  3. Interpretability and causal tests found stylistic signals across layers, but the scoring setup determined where those signals were consolidated during training.

  4. The result suggests some apparent model gains come from output scoring design rather than better language representations.

Read the original