← 返回论文检索
ACL 2026aclfindings

Learning to Translate by Translating: Stabilizing the Dual Loop via Semantic-Aware Self-Evolution

Kui Liu, Mingming Yin, Zuoli Tang, Zihao Li, Chilin Fu, Xiaolu Zhang, Jun Zhou, Lixin Zou, Chenliang Li

University of Science and Technology of China, Zhejiang University, nanjing university, Xiamen University, national university of singaore, National University of Singapore, Nanyang Technological University, Peking University, Jilin University, Tianjin University, Harbin Institute of Technology, Dalian University of Technology, Xi'an University of Electronic Science and Technology, Hong Kong University of Science and Technology, The Chinese University of Hong Kong, University of Hong Kong, Tsinghua University, Renmin University of China, Shanghai Jiaotong University, Fudan University, University of the Chinese Academy of Sciences, Huazhong University of Science and Technology, Central South University, South China University of Technology, SUN YAT-SEN UNIVERSITY, Northwest Polytechnical University Xi'an, Sichuan University, Soochow University, Shandong University, Northeastern University, Beijing Institute of Technology, Beijing University of Post and Telecommunications, National University of Defense Technology, East China Normal University, Beijing University of Aeronautics and Astronautics, Tongji University, Nanjing University of Science and Technology and Wuhan University · AntGroup · Ant Group · Wuhan University

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2026.findings-acl.723 ↗

摘要

Despite the remarkable success of Large Language Models (LLMs) in Machine Translation (MT), the scarcity of high-quality parallel corpora and the prohibitive cost of their acquisition constrain scalability. To this end, we propose Learning to Translate by Translating (LTT), an LLM-driven dual-learning framework that enables autonomous translation, achieving an 80.42% performance improvement over the base model. By adapting the cycle-consistency principle to the generative paradigm, LTT eliminates the need for parallel data. It employs a robust semantic-aware reward function that balances adequacy with reconstruction fidelity, effectively mitigating the reward hacking issues inherent in traditional unsupervised MT. Relying solely on monolingual data, our 8B model consistently outperforms significantly larger models (70B+) in low-resource settings and achieves parity with state-of-the-art supervised baselines on mainstream benchmarks. LTT thus offers a scalable, data-efficient paradigm for autonomous machine translation.