LLMs believe false statements even after explicit warnings that they're false

TL;DR AI
2 min readKey summary
Researchers fine-tuned LLMs on documents with false claims, including explicit negations and corrections.
Even when the training text said the claims were false, the models usually still treated them as true.
The same pattern showed up in training aimed at reducing harmful or deceptive behavior.
The results raise concerns that simple fact-checking or warning text may not reliably change model beliefs.
