← 返回论文检索
ACM Multimedia 2025Grand Challenges

Higher-Order Vision-Language Fusion for Video Popularity Prediction

Kele Xu, Qisheng Xu, Binli Luo, Han Zhou 0003, Zengming Lin, Hui Geng, Xianhan Tan

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3763762 ↗

摘要

Predicting the popularity of social media videos involves estimating user engagement based on rich multimodal information embedded within the posts. Unlike static images, videos incorporate temporally evolving visual signals that, alongside associated metadata such as descriptions, hashtags, timestamps, and user attributes, offer valuable insights into their potential audience reach. Prior approaches typically extract features from different modalities independently and merge them via naïve concatenation, which overlooks the semantic discrepancy and interaction dynamics across modalities. To address these limitations, we propose a feature fusion framework that encodes and aligns video content and associated textual cues into a shared semantic space. By jointly modeling temporally structured visual features with context-aware textual embeddings, our method effectively captures cross-modal correlations that are crucial for discerning content virality patterns. In addition, we incorporate user-centric behavioral profiles and content creation dynamics, enriching the representation with personalized signals that reflect audience-specific preferences. Notably, our method achieves top-tier performance in the 2025 SMP challenge, ranking among the highest-performing entries. This strong empirical result underscores the value of deep semantic alignment across video, text, and user domains in accurately forecasting social media video popularity.