Hot eventLive
arXiv论文质疑健康AI分诊失败源于评估格式
1 reports1 sources16 hr ago updated
Get the story
AI overview
此前一项 Nature Medicine 研究报告,消费级健康 AI ChatGPT 在急诊分诊中出现 51.6% 的漏分诊,并将其归因于模型风险。2026年10月5日,一篇发表于 arXiv(Human-Computer Interaction)的论文对此提出质疑,认为该研究使用的考试式提示词与强制 A/B/C/D 输出更像测量工具,而非真实使用场景,因此失败可能反映评估格式问题而非模型能力本身。论文据此认为,原研究可能高估了消费级健康 AI 的分诊失误。目前该争议仍停留在方法学层面,报道未给出后续验证或原研究团队的回应。
Generated from reports · updated 16 hr ago
LatestOct 5
arXiv论文称评估格式而非模型能力导致健康AI分诊失败被高估。Timeline
Follow the coverage from different angles.
Oct 5, 2026
- arXiv · Human-Computer Interaction论文称评估格式而非模型能力导致消费级健康 AI 分诊失误被高估
论文质疑一项 Nature Medicine 研究把 ChatGPT Health 51.6% 的急诊漏分诊归因于模型风险,认为其考试式提示词与强制 A/B/C/D 输出更像测量工具而非真实使用。
Heat trend
Current heat 7·Comparable peak 10(Oct 5)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.