← 返回论文检索
ICML 2026PosterAccept (regular)

Efficient LLM Moderation with Multi-Layer Latent Prototypes

Maciej Chrabaszcz, Filip Szatkowski, Bartosz Wójcik, Jan Dubiński, Tomasz Trzcinski, Sebastian Cygert

NASK National Research Institute / Warsaw University of Technology · Mistral.AI, Warsaw University of Technology · Jagiellonian University · Warsaw University of Technology · Warsaw University of Technology, Tooploox · NASK - National Research Institute. Gdańsk University of Technology

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Although modern LLMs are aligned with human values during post-training, robust moderation remains essential to prevent harmful outputs at deployment time. Existing approaches suffer from performance-efficiency trade-offs and are difficult to customize to user-specific requirements. Motivated by this gap, we introduce Multi-Layer Prototype Moderator (MLPM), a lightweight and highly customizable input moderation tool. We propose leveraging prototypes of intermediate representations across multiple layers to improve moderation quality while maintaining high efficiency. By design, our method adds negligible overhead to the generation pipeline and can be seamlessly applied to any model. MLPM achieves state-of-the-art performance on diverse moderation benchmarks and demonstrates strong scalability across model families of various sizes. Moreover, we show that it integrates smoothly into end-to-end moderation pipelines and further improves response safety when combined with output moderation techniques. Overall, our work provides a practical and adaptable solution for safe, robust, and efficient LLM deployment.