华东师范大学学报(自然科学版) ›› 2026, Vol. 2026 ›› Issue (5): 109-118.doi: 10.3969/j.issn.1000-5641.2026.05.009

• 数据智能技术 • 上一篇    

音频驱动的跨模态交互技术综述

唐明曙1, 于明鹤1,*(), 张天成2, 于戈2   

  1. 1. 东北大学 软件学院, 沈阳 110169
    2. 东北大学 计算机科学与工程学院, 沈阳 110169
  • 收稿日期:2026-07-08 出版日期:2026-09-25 发布日期:2026-09-12
  • 通讯作者: 于明鹤 E-mail:yuminghe@mail.neu.edu.cn
  • 基金资助:
    国家自然科学基金 (62461146205, 62137001)

Survey of audio-driven cross-modal interaction technologies

Mingshu TANG1, Minghe YU1,*(), Tiancheng ZHANG2, Ge YU2   

  1. 1. School of Software, Northeastern University, Shenyang 110169, China
    2. School of Computer Science and Engineering, Northeastern University, Shenyang 110169, China
  • Received:2026-07-08 Online:2026-09-25 Published:2026-09-12
  • Contact: Minghe YU E-mail:yuminghe@mail.neu.edu.cn

摘要:

多模态交互是人工智能领域的重要研究方向, 其中音频模态承载了丰富的时序动态信息和环境语义, 在智能助手、虚拟数字人和自动驾驶等场景中具有重要作用. 然而, 现有综述多以视觉模态为中心, 对以音频为核心的跨模态交互缺乏系统性梳理, 难以全面呈现该领域的技术脉络与发展瓶颈. 本文以音频为主导模态, 围绕音频与文本、静态图像、动态视频3类交互主线, 并延伸至音频驱动3D运动生成方向, 系统总结近年来的代表性方法、技术路线和应用进展. 在此基础上, 本文从数据基础、时序建模、评测体系和真实场景适应性等方面分析不同研究方向的发展现状与主要局限, 重点分析多模态音频数据集在规模、标注质量和获取成本方面面临的瓶颈. 总体来看, 音频跨模态交互已在统一表征学习、生成式建模和真实场景应用中取得显著进展, 并正在向更强的多粒度对齐能力、更高质量的数据支撑、更鲁棒的评测体系和更高效的部署方法持续发展.

关键词: 音频跨模态交互, 跨模态检索, 多模态学习, 音视觉对齐, 数据分析

Abstract:

Multimodal interaction has become an important research direction in artificial intelligence. The audio modality carries rich temporal dynamic information and environmental semantics, playing a significant role in scenarios such as intelligent assistants, virtual digital humans, and autonomous driving. However, most existing surveys focus on the visual modality and lack a systematic review of audio-centric cross-modal interaction, making it difficult to comprehensively present the technical roadmap and developmental bottlenecks in this field. Taking audio as the dominant modality, this paper focuses on three major interaction lines—audio-text, audio-static image, and audio-dynamic video—and extends them to audio-driven three-dimensional motion generation. It systematically summarizes representative methods, technical routes, and application progress developed in recent years. On this basis, the paper analyzes the current development status and main limitations of different research directions from the perspectives of data foundation, time-series modeling, evaluation systems, and adaptability to real scenarios, with particular emphasis on the bottlenecks faced by multimodal audio datasets in terms of scale, annotation quality, and acquisition cost. Overall, audio-driven cross-modal interaction has achieved significant progress in unified representation learning, generative modeling, and real-scenario applications, and continues to advance toward stronger multi-granularity alignment capabilities, higher-quality data support, more robust evaluation systems, and more efficient deployment methods.

Key words: audio cross-modal interaction, cross-modal retrieval, multimodal learning, audio-visual alignment, data analysis

中图分类号: