When Should Models Change Their Minds? Contextual Belief Management in Large Language Models

TL;DR AI
2 min readKey summary
Researchers introduced BeliefTrack, a benchmark for contextual belief management in multi-turn tasks with exact turn-level evaluation.
The benchmark covers Rule Discovery and Circuit Diagnosis, testing whether models can correctly store, update, or ignore beliefs over time.
Baseline large language models often fail to maintain accurate internal belief states during long interactions.
Training with belief-state rewards sharply reduces these failures, and representation-level steering also improves performance.
