OpenProblemBench:基础理论科学开放问题基准发布
Get the story
2026年10月9日,研究者在arXiv发布OpenProblemBench,用于评测AI在基础理论科学开放问题上的能力。该基准包含82个来自数学与理论物理文献的未解决问题,每个问题提供研究背景、假设与既有进展。评估采用四个模型独立评判提交的正确性、完整性与进展程度,无需参考答案。在七个被评测配置中,GPT-6-Astra平均解题率最高,为14.0%;全尺寸开源模型为5.5%-6.7%;Flash模型为2.4%-3.7%。
Generated from reports · updated 3 hr ago
Timeline
Follow the coverage from different angles.
- arXiv · Artificial IntelligenceOpenProblemBench:评测 AI 在基础理论科学开放问题上的能力
研究者发布 OpenProblemBench,包含 82 个来自数学与理论物理文献的未解决问题,每个问题提供研究背景、假设与既有进展。四个评估模型独立评判提交的正确性、完整性与进展程度,无需参考答案。七个被评测配置中,GPT-6-Astra 平均解题率最高为 14.0%,全尺寸开源模型为 5.5-6.7%,Flash 模型为 2.4-3.7%。
Heat trend
Current heat 9·Comparable peak 10(Oct 9)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.