Skip to content
Hot eventLive

trajectory-judge:仅看结果的 LLM 裁判漏检 Agent 轨迹故障

1 reports1 sources16 hr ago updated

Get the story

AI overview

2026年10月5日,arXiv Software Engineering 频道发布论文,提出 trajectory-judge 测试台。该测试台使用确定性客服环境和一步故障注入器标注了400条轨迹,用于检验 LLM 裁判对 Agent 轨迹故障的识别能力。结果显示,只看请求和最终回复的14B裁判在四类不改变回复的故障上召回率为34%至76%,但配对区分度为零;按结果是否被破坏拆分也不能修复该偏差。论文据此表明,仅凭结果的评审会漏掉轨迹中的故障。

Generated from reports · updated 16 hr ago

Timeline

Follow the coverage from different angles.

Oct 5, 2026
  1. arXiv · Software Engineering
    trajectory-judge:仅看结果的 LLM 裁判会漏掉 Agent 轨迹中的什么

    论文提出 trajectory-judge 测试台,用确定性客服环境和一步故障注入器标注 400 条轨迹,检验 LLM 裁判对 Agent 轨迹故障的识别。结果显示,只看请求和最终回复的 14B 裁判在四类不改变回复的故障上召回 34% 至 76%,但配对区分度为零;按结果是否被破坏拆分也不能修复该偏差。

Heat trend

Current heat 7·Comparable peak 10(Oct 5)·Comparable change over 24 hours –

02.557.510Oct5Oct5Oct5Oct6

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.