Confidence and Calibration of Activation Oracles for Reliable Interpretation of Language Model Internals
TL;DR AI
2 min readKey summary
Researchers benchmarked six confidence-estimation methods for activation oracles in two Qwen-family models.
Bootstrap mode frequency was the best-calibrated method, with lower expected calibration error than answer-word log-probability.
The findings suggest more trustworthy confidence scores for white-box interpretability and activation-based analysis.
The team also released code plus new oracle and target models for Qwen3.6-27B; log-probability remains a cheaper screening option.
