trajectory-judge:仅看结果的 LLM 裁判漏检 Agent 轨迹故障
Get the story
2026年10月5日,arXiv Software Engineering 频道发布论文,提出 trajectory-judge 测试台。该测试台使用确定性客服环境和一步故障注入器标注了400条轨迹,用于检验 LLM 裁判对 Agent 轨迹故障的识别能力。结果显示,只看请求和最终回复的14B裁判在四类不改变回复的故障上召回率为34%至76%,但配对区分度为零;按结果是否被破坏拆分也不能修复该偏差。论文据此表明,仅凭结果的评审会漏掉轨迹中的故障。
Generated from reports · updated 16 hr ago
Timeline
Follow the coverage from different angles.
- arXiv · Software Engineeringtrajectory-judge:仅看结果的 LLM 裁判会漏掉 Agent 轨迹中的什么
论文提出 trajectory-judge 测试台,用确定性客服环境和一步故障注入器标注 400 条轨迹,检验 LLM 裁判对 Agent 轨迹故障的识别。结果显示,只看请求和最终回复的 14B 裁判在四类不改变回复的故障上召回 34% 至 76%,但配对区分度为零;按结果是否被破坏拆分也不能修复该偏差。
Heat trend
Current heat 7·Comparable peak 10(Oct 5)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.