← 返回论文检索
EMNLP 2025emnlpfindings

Layer Duplication in LLMs

Neo Eyal, Nachum Dershowitz, Kfir Bar

Tel Aviv University, Technion · Reichman University

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2025.findings-emnlp.967 ↗

摘要

We investigate the effect of duplicating multihead self-attention layers in large language models (LLMs) across a range of language tasks, with and without fine-tuning. The results demonstrate that duplicating the initial layers once or twice often yields a significant performance boost. Attention analysis uncovered the underlying mechanisms driving the improvement when performing layer duplication. This method enhances LLM capabilities with or without additional training or labeled data.