arXiv · Human-Computer Interaction· David Fraile Navarro, Jialei Sheng, Farah Magrabi, Enrico Coiera·· 16 小时前AI 评分52
论文称评估格式而非模型能力导致消费级健康 AI 分诊失误被高估
Evaluation format, not model capability, drives measured triage failure in the assessment of consumer health AI
AI 导读
论文质疑一项 Nature Medicine 研究把 ChatGPT Health 51.6% 的急诊漏分诊归因于模型风险,认为其考试式提示词与强制 A/B/C/D 输出更像测量工具而非真实使用。
来源:arXiv · Human-Computer Interaction · arxiv.org