← 返回论文检索
ACL 2026longmain

Iterative Dual-Model Alignment for Story Evaluation

Bruce Qin, Dan Goldwasser

Purdue University

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2026.acl-long.648 ↗

摘要

Large language models (LLMs) can both evaluate and explain text quality; however, most existing evaluators operate as static classifiers and lack the ability to refine their reasoning through interaction. We propose an Iterative Alpha–Beta Learning framework that jointly trains two complementary 8B models: an Alpha (\alpha) classifier that assesses pairwise story engagement, and a Beta (\beta) generator that produces structured, rubric-guided comparative explanations. The two models co-evolve within a closed feedback loop: \alpha provides probabilistic preference signals to guide \beta’s Direct Preference Optimization (DPO), while \beta’s improved explanations are reintegrated to retrain \alpha via a KL-based contrastive objective. This dual optimization enables mutual learning: \alpha gains interpretability and robustness from \beta’s textual rationales, while \beta acquires stronger alignment and discriminative precision from \alpha’s confidence deltas. Experiments on human-annotated story-pair datasets HANNA show that the proposed system consistently outperforms strong single-model baselines in both accuracy and explanation quality across multiple iterative rounds.