跳到正文
原文
arXiv · Computation and Language· Longwei Cong, Sonja Hahn, Sebastian Gombert, Leon Camus, Fabian Zehner, Hendrik Drachsler, Ulf Kroehne·· 16 小时前AI 评分37

LLM-as-a-Judge 的评分对齐之外:残差评判难度的测量学分析

Evaluating LLM-as-a-Judge Beyond Score Alignment: A Psychometric Analysis of Residual Judging Difficulty

AI 导读

一项测量学研究指出,LLM-as-a-Judge 与人类评委的总分对齐并不能说明两者觉得同样的案例难判。作者在 SummEval 上对 17 个开放权重 LLM 评委与人类评分分别拟合 Many-Facet Rasch Models,把分数分解为摘要质量、评分者严格度、维度严格度与评分阈值,并定义残差难度。

来源:arXiv · Computation and Language · arxiv.org