Skip to content
Hot eventLive

ArcticQA/ArcticAbstain:北极科学LLM拒答数据集与基准

1 reports1 sources4 hr ago updated

Get the story

AI overview

2026年10月8日,研究者在arXiv(Computation and Language)发布 ArcticQA 数据集与 ArcticAbstain 基准,用于评测大语言模型在北极科学多选题中无有效选项时的拒答能力。数据集含194道题,基准对比答案存在与答案缺失两种条件,评测 Gemini、Claude、ChatGPT 系列共8个模型,记录9,312条回答。结果显示:答案存在时拒答率为0.0%至63.0%;将正确答案替换后,拒答率平均上升5.05个百分点。目前未见与此前报道相矛盾的信息。

Generated from reports · updated 3 hr ago

Timeline

Follow the coverage from different angles.

Oct 8, 2026
  1. arXiv · Computation and Language
    ArcticQA:面向北极科学 LLM 拒答能力的数据集与基准

    研究者发布 ArcticQA 数据集与 ArcticAbstain 基准,用于评测 LLM 在北极科学多选题中无有效选项时的拒答能力。数据集含 194 道题,基准对比答案存在与答案缺失两种条件,评测 Gemini、Claude、ChatGPT 系列共 8 个模型,记录 9,312 条回答。答案存在时拒答率为 0.0% 至 63.0%,替换正确答案后拒答率平均上升 5.05 个百分点。

Heat trend

Current heat 9·Comparable peak 10(Oct 8)·Comparable change over 24 hours –

02.557.510Oct8Oct8Oct8Oct8

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.