← 返回论文检索
ICML 2026PosterAccept (spotlight)

Position: Anthropomorphic Misalignment Research Needs Stronger Evidence

Vansh Gupta, Peter Nutter, Samuel Stante, Andreas Krause, Florian Tramer, Lukas Fluri, Xin Chen, Anna Hedström

ETHZ - ETH Zurich · ETH Zurich / MATS 10.0 · ETH Zurich · ETH Zürich · ETH AI Center (ETH Zürich)

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

We argue that many Anthropomorphized Misalignment Research (AMR) studies need stronger evidence to ensure that they can provide a robust foundation for critical safety decisions, such as model deployment and regulation. By evaluating failure modes across different misalignment concepts, such as deception, emergent misalignment, and sycophancy, we show how conceptual ambiguity, non-robust datasets and experimental design, and insufficient causal interventions can lead to overinterpretation of model behaviors. This position paper aims to offer guidance on evidentiary considerations that can help improve methodological rigor in AMR. To achieve this, we provide a clear call to action through a proposed framework of evidence levels and a diagnostic checklist. These shared standards will enable more productive scientific discourse and ensure that claims about AI risks rest on solid empirical foundations.