Skip to content
Hot eventLive

论文提出长程智能体GHOST风险及STAR-Guard防御

1 reports1 sources16 hr ago updated

Get the story

AI overview

2026年10月5日,arXiv(Artificial Intelligence,一手来源)发表论文,提出长程AI智能体中一种新型执行安全失效模式GHOST:在良性交互下,智能体仍可能违反多轮之前设定的安全约束,在GPT-5.5上的发生率为11.5%。研究从理论上证明,若每个安全前缀的残余违规风险下界不可求和,执行几乎必然进入风险区。论文进一步提出双层防御方法STAR-Guard,结合历史语义安全约束恢复与执行前确定性审计;在GPT-5.5实验设置下,使用该防御后未再观察到GHOST事件。

Generated from reports · updated 16 hr ago

Timeline

Follow the coverage from different angles.

Oct 5, 2026
  1. arXiv · Artificial Intelligence
    长程智能体中的 GHOST:跨轮次被忽视安全约束引发的治理风险

    研究提出长程 AI 智能体的新型执行安全失效模式 GHOST:在良性交互下,智能体仍可能违反多轮之前设定的安全约束,在 GPT-5.5 上发生率达 11.5%。理论上证明若每个安全前缀的残余违规风险下界不可求和,执行几乎必然进入风险区。研究进一步提出双层防御 STAR-Guard,结合历史语义安全约束恢复与执行前确定性审计,在 GPT-5.5 实验设置下未再观察到 GHOST 事件。

Heat trend

Current heat 7·Comparable peak 10(Oct 5)·Comparable change over 24 hours –

02.557.510Oct5Oct5Oct5Oct6

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.