← 返回论文检索
ICML 2026PosterAccept (regular)

From Knowledge to Inference: Formalizing Specialized Public Health Reasoning on GlobalHealthAtlas

Zhaokun Yan, Shan Xu, Wuzheng Dong, Zhaohan Liu, Lijie Feng, Chengxiao Dai, Chen Tianqi, Yingting Li, Yi Zhang, Yunpu Ma, Binfan Liu, Wenting Wei, Tongning Wu

China Academy of Information and Communications Technology · China Academy of Information and Communications Technolog · Nankai University · University of British Columbia · CRRC Industrial Academy Co., Ltd. · The University of Sydney · University of Hong Kong · Beijing University of Posts and Telecommunications · University of Science and Technology Beijing · LMU, Chair of Artificial Intelligence and Machine Learning · lut · Shanghai Artificial Intelligence Laboratory · china academy of information and communications technology

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Public health reasoning requires population-level inference grounded in scientific evidence, expert consensus, and safety constraints. However, it remains underexplored as a structured machine learning problem with limited supervised signals and benchmarks. We introduce GlobalHealthAtlas, a large-scale multilingual dataset of 280,210 instances spanning 15 public health domains and 17 languages, stratified into three difficulty levels from health literacy to epidemiological and policy reasoning. Instances are derived from openly available public health sources and labeled by language, domain, and difficulty to support supervised learning and slice-based evaluation. We further propose a large language model (LLM) assisted construction and quality-control pipeline with retrieval, duplication, evidence-grounding checks, and label validation to improve consistency at scale. Finally, we present a domain-aligned evaluator distilled from high confidence judgments of diverse LLMs to assess outputs along six dimensions: Accuracy, Reasoning, Completeness, Consensus Alignment, Terminology Norms, and Insightfulness. Together, these contributions enable reproducible training and evaluation of LLMs for public health reasoning beyond conventional QA benchmarks.