← 返回论文检索
EMNLP 2024emnlpfindings

Look Who’s Talking Now: Covert Channels From Biased LLMs

Daniel Silva, Frederic Sala, Ryan Gabrys

Naval Information Warfare Center Pacific · University of Wisconsin, Madison · Naval Information Warfare Center

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2024.findings-emnlp.971 ↗

摘要

Large language model-based steganography encodes hidden messages into model-generated tokens. The key tradeoff is between how much hidden information can be introduced and how much the model can be perturbed. To address this tradeoff, we show how to adapt strategies previously used for LLM watermarking to encode large amounts of information. We tackle the practical (but difficult) setting where we do not have access to the full model when trying to recover the hidden information. Theoretically, we study the fundamental limits in how much steganographic information can be inserted into LLM-created outputs. We provide practical encoding schemes and present experimental results showing that our proposed strategies are nearly optimal.