← 返回论文检索
CVPR 2026

Global Information Thresholding for Sufficient and Necessary Circuits

Jegyeong Cho

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Mechanistic interpretability seeks circuits that are both sufficient on their own and necessary to a model's behavior under intervention, yet many circuit-discovery pipelines still rely on manually fixed circuit budgets. We propose a procedure that scores edges with sign-preserving integrated gradients and then chooses a single global cutoff by a performance-retention criterion for circuit selection. This makes circuit size a consequence of retained behavior rather than a manually selected budget. On the Mechanistic Interpretability Benchmark, this retention-calibrated thresholding yields circuits that are competitive on CPR/CMD across multiple tasks and models. On our GPT-2 IOI probability-space diagnostic, it improves sufficiency relative to the EAP-IG baseline while also improving the reported necessity diagnostics. Threshold ablations and stability checks further show that the selected operating point yields the best overall balance between self-sufficiency and necessity-oriented diagnostics.