Skip to content
Hot eventLive

Vals AI评测:多智能体团队成本高、质量提升有限

1 reports1 sources4 hr ago updated

Get the story

AI overview

Vals AI 在 Vibe Code Bench 上测试 GPT-6 Sol 和 Claude Opus 5.5,评估多智能体团队与单智能体的成本与质量差异。结果显示,智能体团队成本为单智能体的 1.8 倍至 5.1 倍,四组对比中仅一组出现显著质量提升。该研究还引用 Anthropic 自有测试,称增加智能体主要提升速度而非质量;OpenAI 研究者 Noam Brown 也持类似观点,认为多智能体主要买来速度而非质量。目前报道未提供各组具体任务细节或质量指标口径。

Generated from reports · updated 3 hr ago

Timeline

Follow the coverage from different angles.

Oct 11, 2026
  1. The Decoder
    研究称 AI 智能体团队耗费大量 token 但质量提升有限

    Vals AI 在 Vibe Code Bench 上测试 GPT-6 Sol 和 Claude Opus 5.5,发现智能体团队成本为单智能体的 1.8 倍至 5.1 倍,四组对比中仅一组显著提升。Anthropic 自有测试也显示增加智能体主要提升速度,OpenAI 研究者 Noam Brown 称多智能体主要买来速度而非质量。

Heat trend

Current heat 9·Comparable peak 10(Oct 12)·Comparable change over 24 hours –

02.557.510Oct12Oct12Oct12Oct12

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.