Is Human-Like Text Liked by Humans? Multilingual Human Detection and Preference Against AI
Institute for Computer Science, Artificial Intelligence and Technology · University of Melbourne · CNR · Mohamed bin Zayed University of Artificial Intelligence · Linköping University · Indian Institute of Technology, Delhi and Mohamed bin Zayed University of Artificial Intelligence · Institute of Science Tokyo · NII LLMC and Mohamed bin Zayed University of Artificial Intelligence · University of Oslo · Nebius · New York University Abu Dhabi · Institute for Computer Science, Artificial Intelligence and Technology, Mohamed bin Zayed University of Artificial Intelligence and Technische Universität Darmstadt
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2026.acl-long.639 ↗
摘要
Prior studies have shown that distinguishing text generated by Large Language Models (LLMs) from human-written one is highly challenging for humans, and often no better than random guessing. To verify the generalizability of this finding across languages and domains, we perform an extensive case study to identify the upper bound of human detection accuracy. Across 16 datasets covering 9 languages and 9 domains, 19 annotators achieved an average detection accuracy of 87.6%, thus challenging previous conclusions. We find that major gaps between human and machine text lie in concreteness, cultural nuances, and diversity. Prompting by explicitly explaining the distinctions in the prompts can partially bridge the gaps in over 50% of the cases. However, we also find that humans do not always prefer human-written text, particularly when they cannot clearly identify its source. We release our dataset, the human labels, and the annotator metadata at https://github.com/xnlp-lab/HumanEval-MGT.