Switch language한국어
Back to the list

LLMs believe false statements even after explicit warnings that they're false

TL;DR AI

Key summary

2 min read
  1. Researchers fine-tuned LLMs on documents with false claims, including explicit negations and corrections.

  2. Even when the training text said the claims were false, the models usually still treated them as true.

  3. The same pattern showed up in training aimed at reducing harmful or deceptive behavior.

  4. The results raise concerns that simple fact-checking or warning text may not reliably change model beliefs.

Read the original