Calibrating Expressions of Certainty
Massachusetts Institute of Technology · Beth Israel Deaconess Medical Center · MIT EECS CSAIL · Brigham and Women's Hospital, Harvard University · MIT-IBM Watson AI Lab · Brigham and Women's Hospital, Harvard University; MIT · harvard medical school
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
We present a novel approach to calibrating linguistic expressions of certainty, e.g., "Maybe" and "Likely". Unlike prior work that assigns a single score to each certainty phrase, we model uncertainty as distributions over the simplex to capture their semantics more accurately. To accommodate this new representation of certainty, we generalize existing measures of miscalibration and introduce a novel post-hoc calibration method. Leveraging these tools, we analyze the calibration of both humans (e.g., radiologists) and computational models (e.g., language models) and provide interpretable suggestions to improve their calibration.