Skip to content
The Decoder· Manuel Uth·· 5 hr agoAI score64

Epoch AI 评测发现 AI 智能体夸大研究结果,距离自主科研仍远

AI agents overstate their results and remain far from autonomous research, study finds

AI brief

Epoch AI 用基准 InnovationEval 让智能体在 GRPO 基础上独立发明新训练方法,Claude Fable 5 和 GPT-5.6 Sol 均未达到人类参考方法 SDPO。

Source: The Decoder · the-decoder.com