J* E* C* N* U* N* S* ›› 2026, Vol. 2026 ›› Issue (5): 109-118.doi: 10.3969/j.issn.1000-5641.2026.05.009

• Data Intelligent Technologies • Previous Articles    

Survey of audio-driven cross-modal interaction technologies

Mingshu TANG1, Minghe YU1,*(), Tiancheng ZHANG2, Ge YU2   

  1. 1. School of Software, Northeastern University, Shenyang 110169, China
    2. School of Computer Science and Engineering, Northeastern University, Shenyang 110169, China
  • Received:2026-07-08 Online:2026-09-25 Published:2026-09-12
  • Contact: Minghe YU E-mail:yuminghe@mail.neu.edu.cn

Abstract:

Multimodal interaction has become an important research direction in artificial intelligence. The audio modality carries rich temporal dynamic information and environmental semantics, playing a significant role in scenarios such as intelligent assistants, virtual digital humans, and autonomous driving. However, most existing surveys focus on the visual modality and lack a systematic review of audio-centric cross-modal interaction, making it difficult to comprehensively present the technical roadmap and developmental bottlenecks in this field. Taking audio as the dominant modality, this paper focuses on three major interaction lines—audio-text, audio-static image, and audio-dynamic video—and extends them to audio-driven three-dimensional motion generation. It systematically summarizes representative methods, technical routes, and application progress developed in recent years. On this basis, the paper analyzes the current development status and main limitations of different research directions from the perspectives of data foundation, time-series modeling, evaluation systems, and adaptability to real scenarios, with particular emphasis on the bottlenecks faced by multimodal audio datasets in terms of scale, annotation quality, and acquisition cost. Overall, audio-driven cross-modal interaction has achieved significant progress in unified representation learning, generative modeling, and real-scenario applications, and continues to advance toward stronger multi-granularity alignment capabilities, higher-quality data support, more robust evaluation systems, and more efficient deployment methods.

Key words: audio cross-modal interaction, cross-modal retrieval, multimodal learning, audio-visual alignment, data analysis

CLC Number: