Vals AI评测:多智能体团队成本高、质量提升有限
Get the story
Vals AI 在 Vibe Code Bench 上测试 GPT-6 Sol 和 Claude Opus 5.5,评估多智能体团队与单智能体的成本与质量差异。结果显示,智能体团队成本为单智能体的 1.8 倍至 5.1 倍,四组对比中仅一组出现显著质量提升。该研究还引用 Anthropic 自有测试,称增加智能体主要提升速度而非质量;OpenAI 研究者 Noam Brown 也持类似观点,认为多智能体主要买来速度而非质量。目前报道未提供各组具体任务细节或质量指标口径。
Generated from reports · updated 3 hr ago
Timeline
Follow the coverage from different angles.
- The Decoder研究称 AI 智能体团队耗费大量 token 但质量提升有限
Vals AI 在 Vibe Code Bench 上测试 GPT-6 Sol 和 Claude Opus 5.5,发现智能体团队成本为单智能体的 1.8 倍至 5.1 倍,四组对比中仅一组显著提升。Anthropic 自有测试也显示增加智能体主要提升速度,OpenAI 研究者 Noam Brown 称多智能体主要买来速度而非质量。
Heat trend
Current heat 9·Comparable peak 10(Oct 12)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.