华东师范大学学报(自然科学版) ›› 2026, Vol. 2026 ›› Issue (4): 102-111.doi: 10.3969/j.issn.1000-5641.2026.04.011

• • 上一篇    

针对大语言模型的社会认知评测框架和问答基准

马毅鸣1, 林欣1,*(), 倪琴2, 余杨泽3, 邓赐平4, 贺樑1   

  1. 1. 华东师范大学 计算机科学与技术学院, 上海 200062
    2. 上海外国语大学 多语种智慧教育重点实验室, 上海 201620
    3. 上海师范大学 信息与机电工程学院, 上海 200233
    4. 华东师范大学 心理与认知科学学院, 上海 200062
  • 收稿日期:2024-12-03 出版日期:2026-07-25 发布日期:2026-07-18
  • 通讯作者: 林欣 E-mail:xlin@cs.ecnu.edu.cn
  • 基金资助:
    国家自然科学基金(6210020445); 国家重点研发计划(2021ZD0111000, 2021ZD0111004); 上海市自然科学基金(21511100101, 21511100102, 21ZR1446900)

Social cognition evaluation framework and question-answering benchmark for large language model

Yiming MA1, Xin LIN1,*(), Qin NI2, Yangze YU3, Ciping DENG4, Liang HE1   

  1. 1. School of Computer Science and Technology, East China Normal University, Shanghai 200062, China
    2. Key Laboratory of Multilingual Education with AI, Shanghai International Studies University, Shanghai 201620, China
    3. College of Information, Mechanical and Electrical Engineering, Shanghai Normal University, Shanghai 200233, China
    4. School of Psychology and Cognitive Science, East China Normal University, Shanghai 200062, China
  • Received:2024-12-03 Online:2026-07-25 Published:2026-07-18
  • Contact: Xin LIN E-mail:xlin@cs.ecnu.edu.cn

摘要:

大语言模型 (Large Language Model, LLM) 与人类的互动日益频繁, 评估它们在社会认知方面的能力显得尤为重要. 本文借鉴人类的社会认知理论, 设计了一个以心智理论 (Theory of Mind, ToM) 为核心的评估框架, 并利用该框架, 生成了一套针对LLM的社会认知问答基准. 该基准以选择题的方式呈现, 灵感来自心理学实验. 最后, 将该数据集应用于各种模型并以此研究 LLM 的 ToM 能力发展. 结果表明, 虽然 LLM 表现出新生的 ToM 能力, 但仍有很大的提升空间. 此外, 研究揭示了 LLM 和人类的 ToM 发展之间的相似之处.

关键词: 大语言模型, 心智理论, 社会认知, 问答基准

Abstract:

Considering the increasing interactions between large language model (LLM) and humans, the assessment of their proficiency in social cognition, a core component of understanding mutual perspectives, is crucial. Drawing from human social cognition theories, this study created an evaluation framework rooted in the theory of mind (ToM), which focuses on the ability to infer the mental states of others. By employing this framework, a social cognition assessment for LLM was developed and presented as a benchmark for multiple-choice questions, inspired by psychological experiments. By applying this benchmark to various commercial models and examining the development of LLM's ToM abilities, our findings indicate that while LLM exhibit nascent ToM skills, there is substantial room for improvement. Furthermore, our study revealed similarities between the development of ToM in LLM and humans.

Key words: large language model (LLM), theory of mind, social cognition, question-answering benchmark

中图分类号: