Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics
TL;DR AI
2 min readKey summary
Researchers found that tracking probe trajectories through a model’s internal reasoning states predicts future behavior better than single-point probes.
Trajectory-based analysis and signal-processing features separated future model states more effectively across multiple reasoning models and datasets.
Template-based training data performed about as well as dynamic labeling, while max-pooling beat average or last-token pooling.
The results point to a more reliable approach for safety monitoring and behavior prediction in advanced reasoning models.
