Skip to content
Hot eventLive

OpenProblemBench:基础理论科学开放问题基准发布

1 reports1 sources3 hr ago updated

Get the story

AI overview

2026年10月9日,研究者在arXiv发布OpenProblemBench,用于评测AI在基础理论科学开放问题上的能力。该基准包含82个来自数学与理论物理文献的未解决问题,每个问题提供研究背景、假设与既有进展。评估采用四个模型独立评判提交的正确性、完整性与进展程度,无需参考答案。在七个被评测配置中,GPT-6-Astra平均解题率最高,为14.0%;全尺寸开源模型为5.5%-6.7%;Flash模型为2.4%-3.7%。

Generated from reports · updated 3 hr ago

Timeline

Follow the coverage from different angles.

Oct 9, 2026
  1. arXiv · Artificial Intelligence
    OpenProblemBench:评测 AI 在基础理论科学开放问题上的能力

    研究者发布 OpenProblemBench,包含 82 个来自数学与理论物理文献的未解决问题,每个问题提供研究背景、假设与既有进展。四个评估模型独立评判提交的正确性、完整性与进展程度,无需参考答案。七个被评测配置中,GPT-6-Astra 平均解题率最高为 14.0%,全尺寸开源模型为 5.5-6.7%,Flash 模型为 2.4-3.7%。

Heat trend

Current heat 9·Comparable peak 10(Oct 9)·Comparable change over 24 hours –

02.557.510Oct9Oct9Oct9Oct9

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.