Enhancing In-Context Learning via Implicit Demonstration Augmentation
Peking University · Shanghai Jiaotong University, Wuhan University, Tsinghua University, Tsinghua University, Microsoft, University of the Chinese Academy of Sciences, Chinese Academy of Sciences, Beijing University of Aeronautics and Astronautics, South China University of Technology, SUN YAT-SEN UNIVERSITY, University of Electronic Science and Technology of China, Huazhong University of Science and Technology, Harbin Institute of Technology, Shandong University, nanjing university, Beijing University of Posts and Telecommunications, Shanghai Artificial Intelligence Laboratory, Shanghai University of Science and Technology, Tianjin University, Northeastern University, Southeast University, Xi’an Jiaotong University, Xiamen University, Fudan University, Renmin University of China, Nankai University, Meituan, Kuaishou- 快手科技, East China Normal University, Xi’an University of Electronic Science and Technology, Nanjing University of Aeronautics and Astronautics, Nanjing University of Science and Technology, Southern University of Science and Technology, Northwest Polytechnical University Xi’an, Chongqing University, Jilin University, Beijing Normal University, University of Science and Technology Beijing and Zhejiang University
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2024.acl-long.155 ↗
摘要
The emergence of in-context learning (ICL) enables large pre-trained language models (PLMs) to make predictions for unseen inputs without updating parameters. Despite its potential, ICL’s effectiveness heavily relies on the quality, quantity, and permutation of demonstrations, commonly leading to suboptimal and unstable performance. In this paper, we tackle this challenge for the first time from the perspective of demonstration augmentation. Specifically, we start with enriching representations of demonstrations by leveraging their deep feature distribution. We then theoretically reveal that when the number of augmented copies approaches infinity, the augmentation is approximately equal to a novel logit calibration mechanism integrated with specific statistical properties. This insight results in a simple yet highly efficient method that significantly improves the average and worst-case accuracy across diverse PLMs and tasks. Moreover, our method effectively reduces performance variance among varying demonstrations, permutations, and templates, and displays the capability to address imbalanced class distributions.