Can Small Vision–Language Models Perform Sign Language Translation?
Indian Association for the Cultivation of Science
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2026.findings-acl.1609 ↗
摘要
Vision-Language Models (VLMs) have shown strong generalization across multimodal tasks, but their capacity to handle sign language translation (SLT), which requires fine-grained spatiotemporal reasoning and linguistic understanding, remains unclear. In this study, we evaluate whether small VLMs (with \leq3B parameters) can perform SLT effectively. We perform supervised fine-tuning on four publicly available multilingual SLT datasets, including one German (DGS), two American (ASL), and one Indian (ISL), applying parameter-efficient LoRA to the language decoder while keeping the vision encoder frozen and training only the connector. To evaluate translation quality, we propose entity- and semantics-aware metrics tailored for SLT. We highlight the data imbalance issues present in the above widely used SLT datasets. Our analysis highlights the limitations in applying general-purpose VLMs to SLT, unlike their applicability in other tasks, and provides insights to inform future development of VLMs for SLP, which is essential for building inclusive AI applications.