← 返回论文检索
ICML 2026PosterAccept (regular)

Aggregate Models, Not Explanations: Improving Feature Importance Estimation

Joseph Paillard, Angel REYERO LOBO, Denis-Alexander Engemann, Thirion Bertrand

Inria; Roche Pharma Research & Early Development (pRED) · Institut de Mathémathiques de Toulouse & Inria Paris-Saclay · Roche Innovation Center Basel, Pharma Research & Early Development, F. Hoffmann-La Roche Ltd. · inria

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Feature-importance methods show promise for transforming machine learning (ML) models from predictive engines into tools for scientific discovery. However, expressive models can be unstable due to data sampling and algorithmic stochasticity, leading to inaccurate variable importance estimates, undermining their utility in critical biomedical applications. While ensembling offers a remedy, the choice between explaining a single ensemble model or aggregating individual model explanations is non-trivial due to the non-linearity of importance measures, and remains largely understudied. Our theoretical analysis, developed under assumptions accommodating complex state-of-the-art ML models, reveals that this choice is governed by a trade-off involving the model's excess risk. In contrast to prior literature, we show that ensembling at the model level provides more accurate variable-importance estimates, particularly for expressive models, by reducing this leading error term. We validate these findings on classical benchmarks and a large-scale proteomic study from the UK Biobank.