Decoding the Critique Mechanism in Large Reasoning Models
TL;DR AI
2 min readKey summary
Researchers found that large reasoning models can often recover from injected arithmetic mistakes even when the error is not explicitly corrected in text.
The behavior appears to come from an internal critique mechanism that can flag reasoning problems beneath the surface of the chain of thought.
They also identified an interpretable critique vector in feature space that improves error detection and test-time performance across model families.
The finding suggests models may be steered to self-verify more reliably without retraining, improving robustness and accuracy.
