Skip to content
Hot eventLive

LLM-as-a-Judge 残差评判难度测量学研究

1 reports1 sources16 hr ago updated

Get the story

AI overview

2026年10月5日,arXiv Computation and Language 发布一项测量学研究(一手报道),题为《LLM-as-a-Judge 的评分对齐之外:残差评判难度的测量学分析》。研究指出,LLM-as-a-Judge 与人类评委的总分对齐并不能说明两者觉得同样的案例难判。作者在 SummEval 数据集上对 17 个开放权重 LLM 评委与人类评分分别拟合 Many-Facet Rasch Models,将分数分解为摘要质量、评分者严格度、维度严格度与评分阈值,并定义“残差难度”这一指标,以刻画总分对齐之外仍存在的评判分歧。该研究强调仅看总分对齐会掩盖评判难度结构上的差异。

Generated from reports · updated 16 hr ago

Timeline

Follow the coverage from different angles.

Oct 5, 2026
  1. arXiv · Computation and Language
    LLM-as-a-Judge 的评分对齐之外:残差评判难度的测量学分析

    一项测量学研究指出,LLM-as-a-Judge 与人类评委的总分对齐并不能说明两者觉得同样的案例难判。作者在 SummEval 上对 17 个开放权重 LLM 评委与人类评分分别拟合 Many-Facet Rasch Models,把分数分解为摘要质量、评分者严格度、维度严格度与评分阈值,并定义残差难度。

Heat trend

Current heat 6·Comparable peak 10(Oct 5)·Comparable change over 24 hours –

02.557.510Oct5Oct5Oct5Oct6

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.