跳到正文
原文
arXiv · Human-Computer Interaction· Juan Manuel Contreras·· 3 小时前AI 评分57

LLM 原生心理测量工具揭示 25 个模型的自我报告与行为差距

An LLM-Native Psychometric Instrument Reveals a Self-Report--Behavior Gap Across 25 Models

AI 导读

论文构建由 LLM 专属行为组成的自陈量表,对 17 家开发者的 25 个 LLM 施测 300 题各 30 次,得到五个可复现因子。自陈结果与 2500 个开放式行为样本的人类评分几乎不相关,verbose 除外;Responsiveness 上自陈更贴近 LLM 评委而非人类。

来源:arXiv · Human-Computer Interaction · arxiv.org