UniMTR: Unified Recognition of Dual-style Traditional Mongolian Scripts via Contrastive Representation Alignment
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3754793 ↗
摘要
Traditional Mongolian script recognition poses unique challenges due to its vertical layout, complex morphology, and the coexistence of visually distinct writing styles, such as standard printed (White) and cursive calligraphic (Hawang) forms. Existing approaches typically rely on style-specific models, leading to limited generalization and increased computational cost. In this paper, we propose UniMTR, a unified and lightweight framework for dual-style Traditional Mongolian word recognition. UniMTR distills knowledge from two expert teacher networks into a compact student model through a novel contrastive distillation strategy. This strategy leverages cross-style positive pair construction, hard negative mining, and uncertainty-aware loss weighting to bridge the style gap. We also introduce a new glyph-code encoding scheme that captures context-dependent visual variants beyond Unicode representation. Experiments on the newly constructed benchmark MTR-Mix demonstrate that UniMTR outperforms state-of-the-art baselines, achieving 13.6% CER and 14.2% average style-wise CER, while reducing model size by over 75% and enabling real-time inference at 320 FPS. Our approach offers a practical and scalable solution for style-robust script recognition in resource-constrained scenarios.