华东师范大学学报(自然科学版) ›› 2026, Vol. 2026 ›› Issue (5): 153-166.doi: 10.3969/j.issn.1000-5641.2026.05.013

• 数据治理 • 上一篇    

扩散模型风险治理: 对齐、质量与真实性权衡视角

陈碟, 陈岑*(), 王延昊   

  1. 华东师范大学 数据科学与工程学院, 上海 200062
  • 收稿日期:2026-08-01 出版日期:2026-09-25 发布日期:2026-09-12
  • 通讯作者: 陈岑 E-mail:cenchen@dase.ecnu.edu.cn
  • 基金资助:
    中央引导地方资金项目 (黔科合中引地〔2025〕006号)

Risk governance in diffusion models: A perspective on trade-offs among alignment, quality, and truthfulness

Die CHEN, Cen CHEN*(), Yanhao WANG   

  1. School of Data Science and Engineering, East China Normal University, Shanghai 200062, China
  • Received:2026-08-01 Online:2026-09-25 Published:2026-09-12
  • Contact: Cen CHEN E-mail:cenchen@dase.ecnu.edu.cn

摘要:

扩散模型的快速发展正在重塑图像与视频生成范式, 并推动生成式人工智能从研究工具逐渐演化为数字内容基础设施. 然而, 生成能力提升的同时, 也伴随着内容安全、版权归属、公平偏见、真实性缺失以及隐私泄露等一系列治理风险. 现有研究较少关注不同安全目标之间的内在冲突: 过强的安全对齐可能削弱生成质量, 真实性约束又可能限制模型创造能力, 而更高的视觉逼真度则进一步放大深度伪造与隐私风险. 围绕上述问题, 本文从内容安全、版权与知识产权、公平性与偏见、幻觉与真实性、隐私泄露5个维度, 对扩散模型中的核心治理风险进行系统梳理, 并提出“对齐、生成质量、真实性” (Alignment–Quality–Truthfulness, AQT) 三元权衡框架, 用于统一分析扩散模型中的关键矛盾与底层机制. 进一步指出, 数据记忆、概念耦合、指导放大、世界模型缺失以及数据分布偏差, 共同构成当前扩散模型治理问题的重要来源. 本文希望通过 AQT 框架为扩散模型安全研究提供统一分析视角, 并推动生成模型从“高质量生成”走向“安全、真实且可信的生成”.

关键词: 扩散模型, 可信生成, 模型安全

Abstract:

Diffusion models are rapidly reshaping the paradigms of image and video generation, thus shifting generative artificial intelligence from a research-oriented tool toward a foundational infrastructure for digital-content creation. However, alongside the remarkable improvements in generative capability, a series of governance risks have emerged, including content safety, copyright ownership, fairness and bias, lack of truthfulness, and privacy leakage. Existing studies rarely focus on the intrinsic conflicts among different safety objectives: excessively strong safety alignments may degrade generation quality, constraints on truthfulness may suppress creative capability, and higher visual realism may increase the risks of deepfakes and privacy abuse. Hence, this paper systematically reviews the core governance risks of diffusion models from five perspectives: content safety, copyright and intellectual property, fairness and bias, hallucination and truthfulness, and privacy leakage. Furthermore, we propose an “alignment–quality–truthfulness” (AQT) trade-off framework to provide a unified perspective for analyzing the fundamental tensions and underlying mechanisms in diffusion models. We further argue that data memorization, concept entanglement, guidance amplification, the lack of world modeling, and data distribution bias jointly constitute the major sources of current governance challenges in diffusion models. The proposed AQT framework aims to provide a unified analytical perspective for diffusion-model safety research and facilitate the evolution of generative models from merely “high-quality generation” toward generation that is safe, truthful, and trustworthy.

Key words: diffusion models, trustworthy generation, model safety

中图分类号: