O*NET-BENCH审计LLM裁判的职业测量有效性
Get the story
研究推出 O*NET-BENCH,基于 45,796 条工人评分构建,并在 4,501 条测试评分上评估 33 个评委配置,覆盖六个模型族。结果显示,25 个配置的 tie-aware pair accuracy 至少 0.60,但评委估计 3.0%-97.9% 的回复可接受,远偏离职业匹配工人的 61.1%。校准后分数最多解释 8.5% 的个体评分方差,表明排序一致不足以支撑职业测量。
Generated from reports · updated 16 hr ago
Timeline
Follow the coverage from different angles.
- arXiv · Artificial Intelligence排序正确,尺度错误:审计 LLM 评委的职业 AI 测量
研究推出 O*NET-BENCH,基于 45,796 条工人评分构建,在 4,501 条测试评分上评估 33 个评委配置(覆盖六个模型族)。25 个配置的 tie-aware pair accuracy 至少 0.60,但评委估计 3.0%-97.9% 的回复可接受,远偏离职业匹配工人的 61.1%。校准后分数最多解释 8.5% 的个体评分方差,排序一致不足以支撑职业测量。
Heat trend
Current heat 6·Comparable peak 10(Oct 5)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.