← 返回论文检索
ICLR 2026PosterAccept (Poster)

Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers

Thomas Heap, Tim Lawson, Lucy Farnik, Laurence Aitchison

University of Bristol

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Sparse autoencoders (SAEs) are widely used to extract sparse, interpretable latents from transformer activations. We test whether commonly used SAE quality metrics and automatic explanation pipelines can distinguish trained transformers from randomly initialized ones (e.g., where parameters are sampled i.i.d. from a Gaussian). Over a wide range of Pythia model sizes and multiple randomization schemes, we find that, in many settings, SAEs trained on randomly initialized transformers produce auto-interpretability scores and reconstruction metrics that are similar to those from trained models. These results show that high aggregate auto-interpretability scores do not, by themselves, guarantee that learned, computationally relevant features have been recovered. We therefore recommend treating common SAE metrics as useful but insufficient proxies for mechanistic interpretability and argue for routine randomized baselines and targeted measures of feature 'abstractness'.