Next Phase of Research on Multimodal Foundation Models: From Alignments to Content Generation and Quality Assessment
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3758125 ↗
摘要
AI as a concept has been around since the 1950s. With the recent advancements in machine learning technology, and the availability of big data and large computing resources, the scene is set for the explosive growth of AI. In particular, the emergence of Multimodal Foundation Models that offer significant capabilities in content comprehension, generation and reasoning has opened up opportunities for multimodal research and applications. The talk first reviews the trends and developments in Multimodal Foundation Models. It then outlines the advances in multilingual and multimodal alignments and discusses the emergence of language- and media-agnostic signals that appear to represent abstract concepts commonly used in human languages. These signals have been shown to have a positive impact on both the performance and safety of the resulting models, especially in enhancing those with low-training samples. To further improve the performance of the models, most current approaches employ reinforcement learning with various reward functions to achieve better controllable content generation and trust. To facilitate effective reinforcement learning, quality assessment is of pivotal importance, while it has largely been overlooked. This talk further presents recent approaches to quality assessment and its role in the generation of textual descriptions, videos, and 3D media. As research on Multimodal Foundation Models is still in an early stage, this talk concludes with directions for future research.