← 返回论文检索
ACL 2026aclfindings

Thinking Twice Makes Large Language Models Safer and More Helpful

Yutao Mou, Yuxiao Luo, Shikun Zhang, Wei Ye

Peking University

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2026.findings-acl.1812 ↗

摘要

Current safety alignment techniques for large language models (LLMs) struggle to balance harmlessness and helpfulness: improving safety often comes at the cost of degraded utility. Our preliminary study shows that guiding unaligned base models with safety-aware reasoning that includes explicit self-reflection can effectively defend jailbreak attacks while preserving response quality. This observation motivates internalizing and strengthening self-reflective reasoning capabilities within LLMs to achieve a better safety–utility trade-off. We propose Safety-aware Reflective Reasoning Optimization (SaRO), a two-stage framework: (1) Reasoning-style Warmup (RW) to internalize self-reflective reasoning, and (2) Self-reflective Reasoning Process Optimization (SRPO) to encourage reflection and correction. Experiments show that SaRO outperforms existing reasoning-based alignment methods, achieving a better balance of safety and helpfulness.