← 返回论文检索
ACL 2026longmain

DarwinTOD: LLM-Driven Lifelong Self-evolution for Task-oriented Dialog Systems

Shuyu Zhang, Yujie Liu, Xinru Wang, Cheng Zhang, Yanmin Zhu, Bin Li

Beijing Institute of Graphic Communication · University of Sydney · Tianjin University of Finance & Economics · Shanghai Jiaotong University · Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Chinese Academy of Sciences

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2026.acl-long.2050 ↗

摘要

Traditional task-oriented dialog systems are unable to evolve from ongoing interactions or adapt to new domains after deployment, that is a critical limitation in real-world dynamic environments. Continual learning approaches depend on episodic retraining with human-curated data, failing to achieve autonomy lifelong improvement. While evolutionary computation and LLM driven self-improvement offer promising mechanisms for dialog optimization, they lack a unified framework for holistic, iterative strategy refinement. To bridge this gap, we propose DarwinTOD, a lifelong self-evolving dialog framework that systematically integrates these two paradigms, enabling continuous strategy optimization from a zero-shot base without task-specific fine-tuning. DarwinTOD maintains an Evolvable Strategy Bank and operates through a dual-loop process: online multi-agent dialog execution with peer critique, and offline structured evolutionary operations that refine the strategy bank using accumulated feedback. This closed-loop design enables autonomous continuous improvement without human intervention. Extensive experiments show that DarwinTOD surpasses previous state-of-the-art methods and exhibits continuous performance gains throughout evolution. Our work provides a novel framework for building dialog systems with lifelong self-evolution capabilities.