← 返回论文检索
EMNLP 2025emnlpfindings

How a Bilingual LM Becomes Bilingual: Tracing Internal Representations with Sparse Autoencoders

Tatsuro Inaba, Go Kamoda, Kentaro Inui, Masaru Isonuma, Yusuke Miyao, Yohei Oseki, Yu Takagi, Benjamin Heinzerling

Mohamed bin Zayed University of Artificial Intelligence · Graduate University for Advanced Studies and National Institute for Japanese Language and Linguistics · MBZUAI, RIKEN and Tohoku University · National Institute of Informatics and Tohoku University · The University of Tokyo · University of Tokyo · Nagoya Institute of Technology · Tohoku University and RIKEN

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2025.findings-emnlp.725 ↗

摘要

This study explores how bilingual language models develop complex internal representations.We employ sparse autoencoders to analyze internal representations of bilingual language models with a focus on the effects of training steps, layers, and model sizes.Our analysis shows that language models first learn languages separately, and then gradually form bilingual alignments, particularly in the mid layers. We also found that this bilingual tendency is stronger in larger models.Building on these findings, we demonstrate the critical role of bilingual representations in model performance by employing a novel method that integrates decomposed representations from a fully trained model into a mid-training model.Our results provide insights into how language models acquire bilingual capabilities.