← 返回论文检索
ACM Multimedia 2024Tutorial Presentations

From Multimodal LLM to Human-level AI: Modality, Instruction, Reasoning and Beyond

Hao Fei 0001, Xiangtai Li, Haotian Liu, Fuxiao Liu, Zhuosheng Zhang 0001, Hanwang Zhang, Shuicheng Yan

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3664647.3689171 ↗

摘要

Artificial intelligence (AI) encompasses knowledge acquisition and real-world grounding across various modalities, including language, visual, auditory, and sensory data. Multimodal large language models (MLLMs) have thus recently garnered growing interest in both academia and industry, showing an unprecedented trend to achieve human-level AI. This tutorial aims to deliver a comprehensive review of cutting-edge research in MLLMs, focusing on three key areas: MLLM architecture design, instructional learning, and multimodal reasoning of MLLMs. We will explore technical advancements, synthesize key challenges, and discuss potential avenues for future research. All the resources and materials will be made available online. https://mllm2024.github.io/ACM-MM2024