arXiv · Computation and Language· Benjamin Wilcox, Dawei Gao, Pradeeban Kathiravelu, Douglas Causey, Kewei Sha, Yunhe Feng·· 4 小时前AI 评分32
ArcticQA:面向北极科学 LLM 拒答能力的数据集与基准
Arctic Questions, Missing Answers: A Dataset and Benchmark for LLM Abstention in Arctic Science
AI 导读
研究者发布 ArcticQA 数据集与 ArcticAbstain 基准,用于评测 LLM 在北极科学多选题中无有效选项时的拒答能力。数据集含 194 道题,基准对比答案存在与答案缺失两种条件,评测 Gemini、Claude、ChatGPT 系列共 8 个模型,记录 9,312 条回答。答案存在时拒答率为 0.0% 至 63.0%,替换正确答案后拒答率平均上升 5.05 个百分点。
来源:arXiv · Computation and Language · arxiv.org