arXiv · Artificial Intelligence· Caiqi Zhang, Rujun Han, Zifeng Wang, Zoey CuiZhu, Nigel Collier, Tomas Pfister, Chen-Yu Lee·· 3 hr agoSelectedAI score64
VeriHarness:扩展长周期任务的智能体验证
VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks
AI brief
论文提出 VeriHarness,将基础 LLM 转为带工作空间、证据工具和可复用验证技能的智能体验证器,通过分歧解析器与共识挑战者筛选并修正产出。在五个长周期基准和 Gemini 3.5 Flash、Claude Opus 4.8 上达到最高选择分,证据修订分别提升 6.2 与 6.4 分,并展示验证技能可从失败反馈中自我改进。
Why it matters
为无参考答案的长周期任务提供了可复用的验证技能与证据修订机制,读者可据此评估自身 agent 工作流的可靠性边界。
Source: arXiv · Artificial Intelligence · arxiv.org