SSCBench:评估工具调用LLM Agent故障注入测试的证据有效性
Get the story
SSCBench 提出一套测量协议,用于评估工具调用型 LLM 智能体故障注入测试的证据有效性。研究在两个 τ-bench 环境中展开,覆盖 4 种故障算子、5 种智能体配置,共进行 1,191 次故障执行。实验结果显示,在 44 条最终可见反证的已采纳运行中,仅 17 条在首次使用前收到,27 条在之后收到;自动轨迹分析可以恢复采纳行为,却无法可靠恢复首次错误依赖事件。论文据此主张,在把采纳解释为失败之前,应先报告证据条件与支撑总体。
Generated from reports · updated 3 hr ago
Timeline
Follow the coverage from different angles.
- arXiv · Software EngineeringSSCBench:评估工具调用型 LLM 智能体故障注入测试的证据有效性
SSCBench 提出测量协议,评估工具调用型 LLM 智能体故障注入测试的证据有效性,在两个 τ-bench 环境中覆盖 4 种故障算子、5 种智能体配置、1,191 次故障执行。实验显示,在 44 条最终可见反证的已采纳运行中,仅 17 条在首次使用前收到,27 条在之后收到;自动轨迹分析可恢复采纳,却无法可靠恢复首次错误依赖事件。论文主张应在把采纳解释为失败前报告证据条件与支撑总体。
Heat trend
Current heat 9·Comparable peak 10(Oct 9)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.