Skip to content
arXiv · Artificial Intelligence· Chenglin Yang·· 4 hr agoAI score38

评测整体堆栈而非单层:智能体动作的确定性门与 LLM 门会独立失效吗?

Evaluate the Stack, Not the Layer: Do Deterministic and LLM Gates for Agent Actions Fail Independently?

AI brief

研究在三个语料库的 1,119 个已标注智能体动作上,检验运行时门控堆栈“错误相乘”的假设,堆栈含一层确定性规则与四个 LLM 评审。STRICT 定义下任意两个评审组合约为 1.2 至 1.4 层,规则层加一个评审为 1.86 至 2.09 层;评审耦合的难度占比在 31.8% 至 61.8% 间不可识别。

Source: arXiv · Artificial Intelligence · arxiv.org