From Fragments to Facts: A Curriculum-Driven DPO Approach for Generating Hindi News Veracity Explanations
Tata Consultancy Services Limited, India · Indian Institute of Technology, Patna, Dhirubhai Ambani Institute Of Information and Communication Technology
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2026.findings-acl.2108 ↗
摘要
In an era of rampant misinformation, generating reliable news explanations is vital, especially for underrepresented languages like Hindi. Lacking robust automated tools, Hindi faces challenges in scaling misinformation detection. To bridge this gap, we propose DeFactoX, a novel framework integrating Direct Preference Optimization (DPO) with Curriculum learning to align machine-generated explanations with human reasoning. Fact-checked explanations from credible sources serve as preferred responses, while LLM outputs highlight system limitations and serve as non-preferred responses. At the core of this framework lies Hin-DPO, an enhanced variant of DPO that enriches the loss function with two novel parameters, Actuality and Finesse, enhancing explanation quality and consistency. Experiments with LLMs (Mistral, Llama, Gemma) and PLMs (mBART, mT5) confirm the framework’s effectiveness in generating coherent, contextually relevant explanations.