研究评估编码Agent动态并发子代理的收益与失败模式
Get the story
2026年10月8日,arXiv 软件工程方向发布一项研究,评估编码智能体在长周期任务中启用动态并发子代理的端到端效果与调度行为。研究以 Codex、Claude Code 与 Kimi Code 为对象,通过对照开启动态并发策略与关闭状态下的执行表现进行分析。实验覆盖 354 个任务和 2,124 次执行,梳理了 13 种并发特有失败模式与 28 种可观察模式,并刻画动态并发带来优势的条件。研究指出,在较短任务中模型能力主导结果,而在长周期开发中,编排成为任务完成的关键。目前报道仅涉及该研究的方法与主要结论,未见后续验证或争议。
Generated from reports · updated 3 hr ago
Timeline
Follow the coverage from different angles.
- arXiv · Software Engineering子智能体并行工作:长周期编码任务中动态并发的优势与陷阱
研究通过对照 Codex、Claude Code 与 Kimi Code 在开启动态并发策略与关闭状态下的执行表现,评估子智能体并行工作在长周期编码任务中的端到端效果与调度行为。实验覆盖 354 个任务和 2,124 次执行,分析了 13 种并发特有失败模式与 28 种可观察模式,并刻画动态并发带来优势的条件。研究指出,模型能力在较短任务中主导结果,而长周期开发中编排成为任务完成的关键。
Heat trend
Current heat 9·Comparable peak 10(Oct 8)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.