Confidence and Calibration of Activation Oracles for Reliable Interpretation of Language Model Internals

TL;DR AI
2 min readKey summary
A new arXiv paper studies how to calibrate confidence scores for activation-oracle interpretability methods.
Across 6,000 samples per setting, bootstrap mode frequency gave the best calibration among six methods tested.
Log-probability was a cheaper baseline, but it was consistently less reliable than bootstrap-based confidence.
The work aims to make natural-language explanations of model activations more trustworthy by improving uncertainty estimates.
