Skip to content
Evaluation sources

Terminal-Bench Science

Harbor / 专业办公 · 在终端里完成科研任务

Official evaluation
On HOTAIObserving
Evidence budgetUnweighted
Upstream dataTo verify
Last successful syncCollection not started

What and how it measures

区分 Claude Code、Codex 和 mini-SWE 等系统;不能把不同系统成绩直接当模型能力。

How this evidence is used

观察中:12 个型号且混合运行系统,新前沿型号仍不完整。

Evaluation limits and attribution

观察中:12 个型号且混合运行系统,新前沿型号仍不完整。

Data license: 待确认榜单数据使用边界

成绩由 Harbor 发布,原始分数与 HOTAI 共识分使用不同尺度,不能直接相加。