← 返回论文检索
ACM Multimedia 2025Grand Challenges

HyMoENet: Mixture-of-Experts Enhanced CNN-Transformer Hybrid Framework for Classifying Anatomical Sites in Endoscopic ENT Images

Trong-Nhan Nguyen, Luan L. M. Nguyen, Phat-Dat To, Tran-Quoc Duy Nguyen, Anh-Huy Nguyen, Tuan Pham-Dang, Chu Lam Nguyen, Duy V. M. Nguyen

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3762092 ↗

摘要

Classifying anatomical regions in endoscopic ENT (ear, nose, and throat) images is challenging due to strong inter-class similarities, bilateral symmetry, and the scarcity of annotated datasets. To overcome these issues, we present HyMoENet. This novel hybrid deep learning architecture combines convolutional neural networks (CNNs) for localized feature extraction with Vision Transformers to represent global context. Furthermore, it leverages a sparse Mixture-of-Experts (MoE) technique to improve multi-perspective specialization. Our architecture makes use of parallel CNN-Transformer encoders, which are incorporated into a dynamic MoE layer that adaptively routes representations to the most appropriate experts. Concurrently, a semantic-preserving skip connection preserves global coherence. When tested on a clinically annotated ENT endoscopy dataset from Thong Nhat Hospital in Vietnam, HyMoENet outperformed both single-stream and conventional hybrid models with an accuracy of 97.50. These results demonstrate that integrating local-global representation learning and expert modularization enhances classification accuracy for anatomically similar structures. HyMoENet sets a new benchmark for automated ENT image processing, setting the groundwork for intelligent diagnostic systems in clinical endoscopy.