Hot eventLive
研究:LLM说服能力评测方法间一致性弱
1 reports1 sources4 hr ago updated
Get the story
AI overview
2026年10月8日,arXiv Computers and Society 发表研究,将九种已发表的自动化评测方法统一到同一设置,在十五个 LLM 上运行并比较排名。结果显示各方法一致性很弱,平均 Spearman ρ=0.25;模型拒绝部分任务(多集中在操控类任务)使一致性下降约四分之一,通用能力主要影响理性说服评测。研究指出单一分数只反映特定设置,难以跨任务衡量模型说服力。
Generated from reports · updated 3 hr ago
Timeline
Follow the coverage from different angles.
Oct 8, 2026
- arXiv · Computers and SocietyLLM 的说服力取决于评测方式
研究将九种已发表的自动化评测方法统一到同一设置,在十五个 LLM 上运行并比较其排名。结果显示各方法一致性很弱,平均 Spearman ρ=0.25;模型拒绝部分任务(多集中在操控类任务)使一致性下降约四分之一,通用能力主要影响理性说服评测。单一分数只反映特定设置,难以跨任务衡量模型说服力。
Heat trend
Current heat 9·Comparable peak 10(Oct 8)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.