← 返回论文检索
EMNLP 2025emnlpfindings

Training Medical QA Models Based on Mixed Rewards from Multiple-Choice and Open-Ended Questions

Yue Qiu, Yujan Ting, Pei Dong, Terrence Chen, Weijing Huang

United Imaging Intelligence · UII America

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2025.findings-emnlp.463 ↗

摘要

Reinforcement learning (RL) for large language models (LLMs) typically requires clear reward signals, which are often unavailable for open-ended (OE) questions where answer evaluation is ambiguous without scalable expert labeling. We investigate whether LLMs benefit from training on mixed data with varying reward clarity. Our approach combines Multiple-choice questions (MCQs), which offer clear binary rewards, with OE questions, for which we use simpler, potentially noisy rewards such as Jaccard similarity or LLM-based evaluators. We hypothesize that MCQs can stabilize training when mixed with OE questions. Our experiments show this mixed-data approach consistently improves medical question-answering performance across model scales.