Skip to content
Hot eventLive

研究:LLM说服能力评测方法间一致性弱

1 reports1 sources4 hr ago updated

Get the story

AI overview

2026年10月8日,arXiv Computers and Society 发表研究,将九种已发表的自动化评测方法统一到同一设置,在十五个 LLM 上运行并比较排名。结果显示各方法一致性很弱,平均 Spearman ρ=0.25;模型拒绝部分任务(多集中在操控类任务)使一致性下降约四分之一,通用能力主要影响理性说服评测。研究指出单一分数只反映特定设置,难以跨任务衡量模型说服力。

Generated from reports · updated 3 hr ago

Timeline

Follow the coverage from different angles.

Oct 8, 2026
  1. arXiv · Computers and Society
    LLM 的说服力取决于评测方式

    研究将九种已发表的自动化评测方法统一到同一设置,在十五个 LLM 上运行并比较其排名。结果显示各方法一致性很弱,平均 Spearman ρ=0.25;模型拒绝部分任务(多集中在操控类任务)使一致性下降约四分之一,通用能力主要影响理性说服评测。单一分数只反映特定设置,难以跨任务衡量模型说服力。

Heat trend

Current heat 9·Comparable peak 10(Oct 8)·Comparable change over 24 hours –

02.557.510Oct8Oct8Oct8Oct8

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.