| 1 |
Huang R J, Li M Z, Yang D C, et al. AudioGPT: understanding and generating speech, music, sound, and talking head [C]//Proceedings of the AAAI Conference on Artificial Intelligence. 2024, 38(21): 23802-23804.
|
| 2 |
Chu Y F, Xu J, Zhou X H, et al. Qwen-Audio: advancing universal audio understanding via unified large-scale audio-language models [PP/OL]. V1. arXiv (2023-11-13)[2026-07-11]. https://arxiv.org/abs/2311.07919.
|
| 3 |
Oncescu A M, Koepke A S, Henriques J F, et al. Audio retrieval with natural language queries [C]//Interspeech 2021. ISCA, 2021: 2411-2415.
|
| 4 |
Kim C D, Kim B, Lee H, et al. AudioCaps: generating captions for audios in the wild [C]//Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL-HLT). 2019: 119-132.
|
| 5 |
Drossos K, Lipping S, Virtanen T. Clotho: an audio captioning dataset [C]//2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2020: 736-740.
|
| 6 |
Xin Y F, Zou Y X. Improving audio-text retrieval via hierarchical cross-modal interaction and auxiliary captions [C]//Interspeech 2023. ISCA, 2023: 341-345.
|
| 7 |
Xie Y X, Zhu Z H, Zhuang X W, et al. GPA: global and prototype alignment for audio-text retrieval [C]//Interspeech 2024. ISCA, 2024: 5078-5082.
|
| 8 |
Elizalde B, Deshmukh S, Al Ismail M, et al. CLAP: learning audio concepts from natural language supervision [C]//2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023: 1-5.
|
| 9 |
Wu H H, Seetharaman P, Kumar K, et al. Wav2CLIP: learning robust audio representations from CLIP [C]//2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022: 4563-4567.
|
| 10 |
Iashin V, Xie W, Rahtu E, et al. Synchformer: efficient synchronization from sparse cues [C]//2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024: 5325-5329.
|
| 11 |
Chen S Y, Wang C Y, Chen Z Y, et al.. WavLM: large-scale self-supervised pre-training for full stack speech processing. IEEE Journal of Selected Topics in Signal Processing, 2022, 16 (6): 1505- 1518.
|
| 12 |
Baevski A, Zhou H, Mohamed A, et al. wav2vec 2.0: a framework for self-supervised learning of speech representations [C]//Advances in Neural Information Processing Systems (NeurIPS). 2020: 12449-12460.
|
| 13 |
Wang Y, Skerry-Ryan R J, Stanton D, et al. Tacotron: towards end-to-end speech synthesis [C]//Interspeech 2017. ISCA, 2017: 4006-4010.
|
| 14 |
Ren Y, Hu C, Tan X, et al. FastSpeech 2: fast and high-quality end-to-end text to speech [C]//International Conference on Learning Representations (ICLR). 2021.
|
| 15 |
Wang C, Chen S, Wu Y, et al. Neural codec language models are zero-shot text to speech synthesizers [PP/OL]. V1. arXiv (2023-01-05)[2026-07-11]. https://arxiv.org/abs/2301.02111.
|
| 16 |
Liu H, Chen Z, Yuan Y, et al. AudioLDM: text-to-audio generation with latent diffusion models [C]//International Conference on Machine Learning (ICML). PMLR, 2023: 21450-21474.
|
| 17 |
Kreuk F, Synnaeve G, Polyak A, et al. AudioGen: textually guided audio generation [PP/OL]. V1. arXiv (2022-09-30)[2026-07-11]. https://arxiv.org/abs/2209.15352.
|
| 18 |
Tang C, Yu W, Sun G, et al. SALMONN: towards generic hearing abilities for large language models [PP/OL]. V1. arXiv (2023-10-20)[2026-07-11]. https://arxiv.org/abs/2310.13289.
|
| 19 |
Deshmukh S, Elizalde B, Singh R, et al. Pengi: an audio language model for audio tasks [C]//Advances in Neural Information Processing Systems (NeurIPS). 2023: 18090-18108.
|
| 20 |
Gong Y, Liu A H, Luo H, et al. Listen, think, and understand [PP/OL]. V1. arXiv (2023-05-18)[2026-07-11]. https://arxiv.org/abs/2305.10790.
|
| 21 |
Aytar Y, Vondrick C, Torralba A. SoundNet: learning sound representations from unlabeled video [C]//Advances in Neural Information Processing Systems (NeurIPS). 2016: 892-900.
|
| 22 |
Girdhar R, El-Nouby A, Liu Z, et al. ImageBind: one embedding space to bind them all [C]//IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2023: 15180-15190.
|
| 23 |
Gong Y, Rouditchenko A, Liu A H, et al. Contrastive audio-visual masked autoencoder [C]//International Conference on Learning Representations (ICLR). 2023.
|
| 24 |
Deng J, Dong W, Socher R, et al. ImageNet: a large-scale hierarchical image database [C]//IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2009: 248-255.
|
| 25 |
Zhao H, Gan C, Rouditchenko A, et al. The sound of pixels [C]//European Conference on Computer Vision (ECCV). 2018: 570-586.
|
| 26 |
Tian Y, Li D, Xu C. Unified multisensory perception: weakly-supervised audio-visual video parsing [C]//European Conference on Computer Vision (ECCV). 2020: 436-454.
|
| 27 |
Shi B, Hsu W N, Lakhotia K, et al. Learning audio-visual speech representation by masked multimodal cluster prediction [C]//International Conference on Learning Representations (ICLR). 2022: 1-24.
|
| 28 |
Zhang W, Cun X, Wang X, et al. SadTalker: learning realistic 3D motion coefficients for stylized audio-driven single image talking face animation [C]//IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2023: 8652-8661.
|
| 29 |
Li J, Kang D, Pei W, et al. Audio2Gestures: generating diverse gestures from speech audio with conditional variational autoencoders [C]//IEEE International Conference on Computer Vision (ICCV). 2021: 11293-11302.
|
| 30 |
Liu H, Zhu Z, Becherini G, et al. EMAGE: towards unified holistic co-speech gesture generation via expressive masked audio gesture modeling [C]//IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2024: 1144-1154.
|
| 31 |
Aneja S, Sevastopolsky A, Kirschstein T, et al. GaussianSpeech: audio-driven personalized 3D gaussian avatars [C]//IEEE International Conference on Computer Vision (ICCV). 2025: 13065-13075.
|
| 32 |
郭星星, 肖雁南, 温佩芝, 等.. 基于注意力机制的音频驱动数字人脸视频生成方法. 计算机科学, 2026, 53 (2): 245- 252.
|
| 33 |
Chen X, Fang H, Lin T Y, et al. Microsoft COCO Captions: data collection and evaluation server [PP/OL]. V1. arXiv (2015-04-01)[2026-07-11]. https://arxiv.org/abs/1504.00325.
|
| 34 |
Piczak K J. ESC: dataset for environmental sound classification [C]//ACM International Conference on Multimedia (ACM MM). 2015: 1015-1018.
|
| 35 |
Fonseca E, Favory X, Pons J, et al.. FSD50K: an open dataset of human-labeled sound events. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2022, 30, 829- 852.
|
| 36 |
Mei X, Meng C, Liu H, et al.. WavCaps: a ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language multimodal research. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2024, 32, 3339- 3352.
|
| 37 |
Chen H, Xie W, Vedaldi A, et al. VGGSound: a large-scale audio-visual dataset [C]//2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2020: 721-725.
|
| 38 |
Gemmeke J F, Ellis D P W, Freedman D, et al. Audio set: an ontology and human-labeled dataset for audio events [C]//2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2017: 776-780.
|
| 39 |
Liu H, Zhu Z, Iwamoto N, et al. BEAT: a large-scale semantic and emotional multi-modal dataset for conversational gesture synthesis [C]//European Conference on Computer Vision (ECCV). 2022: 612-630.
|