← 返回论文检索
ACL 2026aclfindings

PictoEduca: Building a Dataset for Spanish Text-to-Pictogram Generation

Alfonso Manuel Paredes Umeres, Marco Antonio Sobrevilla Cabezudo

Pontificia Universidad Católica del Perú

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2026.findings-acl.1738 ↗

摘要

We present PictoEduca, the first large-scale Spanish text-to-pictogram dataset for augmentative and alternative communication (AAC), derived from primary educational materials and grounded in the ARASAAC pictogram repository. The dataset is released with a reproducible pipeline that combines automatic annotation with targeted expert correction, supporting scalable and high-quality corpus construction. We benchmark a rule-based system (ARAWORD) and neural models (T5, LLaMA) under direct text-to-pictogram and two-stage text-to-concept-to-pictogram settings. Results show that the rule-based system remains a strong baseline, while neural models benefit from explicit semantic abstraction, with the two-stage approach improving semantic coherence and reducing ambiguity. We further explore data selection strategies, demonstrating that combining domain similarity with a quality signal yields higher-quality silver data, reduces annotation effort, and improves model performance in low-resource regimes. PictoEduca enables reproducible evaluation and advances Spanish text-to-pictogram research.