Skip to content
arXiv · Computation and Language· Benjamin Wilcox, Dawei Gao, Pradeeban Kathiravelu, Douglas Causey, Kewei Sha, Yunhe Feng·· 4 hr agoAI score32

ArcticQA:面向北极科学 LLM 拒答能力的数据集与基准

Arctic Questions, Missing Answers: A Dataset and Benchmark for LLM Abstention in Arctic Science

AI brief

研究者发布 ArcticQA 数据集与 ArcticAbstain 基准,用于评测 LLM 在北极科学多选题中无有效选项时的拒答能力。数据集含 194 道题,基准对比答案存在与答案缺失两种条件,评测 Gemini、Claude、ChatGPT 系列共 8 个模型,记录 9,312 条回答。答案存在时拒答率为 0.0% 至 63.0%,替换正确答案后拒答率平均上升 5.05 个百分点。

Source: arXiv · Computation and Language · arxiv.org