Skip to content
Hot eventLive

LLM自评与行为差距研究覆盖25模型

1 reports1 sources4 hr ago updated

Get the story

AI overview

2026年10月8日,arXiv Human-Computer Interaction 发布一手研究,构建由 LLM 专属行为组成的自陈量表,对17家开发者的25个LLM施测,每模型300题、各30次,得到五个可复现因子。研究将自陈结果与2500个开放式行为样本的人类评分对比,发现两者几乎不相关,verbose 除外;在 Responsiveness 上,自陈更贴近 LLM 评委而非人类。论文据此指出模型自我报告与实际行为存在差距。

Generated from reports · updated 3 hr ago

Timeline

Follow the coverage from different angles.

Oct 8, 2026
  1. arXiv · Human-Computer Interaction
    LLM 原生心理测量工具揭示 25 个模型的自我报告与行为差距

    论文构建由 LLM 专属行为组成的自陈量表,对 17 家开发者的 25 个 LLM 施测 300 题各 30 次,得到五个可复现因子。自陈结果与 2500 个开放式行为样本的人类评分几乎不相关,verbose 除外;Responsiveness 上自陈更贴近 LLM 评委而非人类。

Heat trend

Current heat 9·Comparable peak 10(Oct 8)·Comparable change over 24 hours –

02.557.510Oct8Oct8Oct8Oct8

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.