LLM-as-a-Judge 残差评判难度测量学研究
Get the story
2026年10月5日,arXiv Computation and Language 发布一项测量学研究(一手报道),题为《LLM-as-a-Judge 的评分对齐之外:残差评判难度的测量学分析》。研究指出,LLM-as-a-Judge 与人类评委的总分对齐并不能说明两者觉得同样的案例难判。作者在 SummEval 数据集上对 17 个开放权重 LLM 评委与人类评分分别拟合 Many-Facet Rasch Models,将分数分解为摘要质量、评分者严格度、维度严格度与评分阈值,并定义“残差难度”这一指标,以刻画总分对齐之外仍存在的评判分歧。该研究强调仅看总分对齐会掩盖评判难度结构上的差异。
Generated from reports · updated 16 hr ago
Timeline
Follow the coverage from different angles.
- arXiv · Computation and LanguageLLM-as-a-Judge 的评分对齐之外:残差评判难度的测量学分析
一项测量学研究指出,LLM-as-a-Judge 与人类评委的总分对齐并不能说明两者觉得同样的案例难判。作者在 SummEval 上对 17 个开放权重 LLM 评委与人类评分分别拟合 Many-Facet Rasch Models,把分数分解为摘要质量、评分者严格度、维度严格度与评分阈值,并定义残差难度。
Heat trend
Current heat 6·Comparable peak 10(Oct 5)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.