arXiv · Artificial Intelligence· Chenglin Yang·· 3 小时前AI 评分38
评测整体堆栈而非单层:智能体动作的确定性门与 LLM 门会独立失效吗?
Evaluate the Stack, Not the Layer: Do Deterministic and LLM Gates for Agent Actions Fail Independently?
AI 导读
研究在三个语料库的 1,119 个已标注智能体动作上,检验运行时门控堆栈“错误相乘”的假设,堆栈含一层确定性规则与四个 LLM 评审。STRICT 定义下任意两个评审组合约为 1.2 至 1.4 层,规则层加一个评审为 1.86 至 2.09 层;评审耦合的难度占比在 31.8% 至 61.8% 间不可识别。
来源:arXiv · Artificial Intelligence · arxiv.org