← 返回论文检索
ACL 2026aclfindings

ToxiTrace: Gradient-Aligned Training for Explainable Chinese Toxicity Detection

Boyang Li, Hongzhe Shou, Yuanyuan Liang, JingBin Zhang, Fang Zhou

East China Normal University

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2026.findings-acl.354 ↗

摘要

Existing Chinese toxic content detection methods mainly target sentence-level classification but often fail to provide readable and contiguous toxic evidence spans. We propose ToxiTrace, an explainability-oriented method for BERT-style encoders with three components: (1) CuSA, which refines encoder-derived saliency cues into fine-grained toxic spans with lightweight LLM guidance; (2) GCLoss, a gradient-constrained objective that concentrates token-level saliency on toxic evidence while suppressing irrelevant activations; and (3) ARCL, which constructs sample-specific contrastive reasoning pairs to sharpen the semantic boundary between toxic and non-toxic content. Experiments show that ToxiTrace improves classification accuracy and toxic span extraction while preserving efficient encoder-based inference and producing more coherent, human-readable explanations. The core training code is available at https://github.com/ZhouF-ECNU/ToxiTrace.