← 返回论文检索
ICLR 2026Blog Track PosterAccept (Poster)

Misalignments and RL Failure Modes in the Early Stage of Superintelligence

Shu Yang, Hanqi Yan, Di Wang

KAUST · King’s College London · King Abdullah University of Science and Technology

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

With the rapid ability grokking of frontier Large Models (LMs), there is growing attention and research focus on aligning them with human values and intent via large scale reinforcement learning and other techniques. However, as LMs are getting stronger and more agentic, their misalignment and deceptive behaviors are also emerging and becoming increasingly difficult for humans to pre-detect and keep track of. This blog post discusses current misalignment patterns, deceptive behaviors, RL failure modes, and emergent traits in modern large models to further AI safety discussions and advance the development of mitigation strategies for LM misbehaviors.