Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning
TL;DR AI
2 min readKey summary
Researchers introduced SaliTrap, a benchmark for measuring salience bias in commonsense reasoning.
Across 12 leading LLMs, models often latched onto irrelevant explicit cues and missed the correct answer.
Context-free probing recovered most of those missed answers, pointing to knowledge suppression rather than missing knowledge.
Lightweight inference-time prompting reduced the problem without retraining.
The study shows that many commonsense failures in LLMs are driven by prompt framing and distractor sensitivity.
