Skip to content
Hot eventLive

arXiv论文质疑健康AI分诊失败源于评估格式

1 reports1 sources16 hr ago updated

Get the story

AI overview

此前一项 Nature Medicine 研究报告,消费级健康 AI ChatGPT 在急诊分诊中出现 51.6% 的漏分诊,并将其归因于模型风险。2026年10月5日,一篇发表于 arXiv(Human-Computer Interaction)的论文对此提出质疑,认为该研究使用的考试式提示词与强制 A/B/C/D 输出更像测量工具,而非真实使用场景,因此失败可能反映评估格式问题而非模型能力本身。论文据此认为,原研究可能高估了消费级健康 AI 的分诊失误。目前该争议仍停留在方法学层面,报道未给出后续验证或原研究团队的回应。

Generated from reports · updated 16 hr ago

Timeline

Follow the coverage from different angles.

Oct 5, 2026
  1. arXiv · Human-Computer Interaction
    论文称评估格式而非模型能力导致消费级健康 AI 分诊失误被高估

    论文质疑一项 Nature Medicine 研究把 ChatGPT Health 51.6% 的急诊漏分诊归因于模型风险,认为其考试式提示词与强制 A/B/C/D 输出更像测量工具而非真实使用。

Heat trend

Current heat 7·Comparable peak 10(Oct 5)·Comparable change over 24 hours –

02.557.510Oct5Oct5Oct5Oct6

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.