研究:临床医生语言模型使用方式与基准评测存在偏差
Get the story
2026年10月9日,arXiv(Computation and Language)发布一手研究,分析35个专科6,342名医护在八个月内向机构助手发送的127,833条查询。结果显示,文档与行政类查询占36.2%、知识检索占28.9%,诊断仅占3.7%,且逾三分之一查询按原样难以良好作答。研究者将RCQ-Map应用于58个公开基准,构建Clinical AI Benchmark Atlas,发现基准中位不含文档请求,任务构成与真实使用仅重合31%,据此提出评测需匹配真实临床使用。
Generated from reports · updated 3 hr ago
Timeline
Follow the coverage from different angles.
- arXiv · Computation and Language临床医生使用语言模型的方式与模型评测方式存在偏差
研究分析 35 个专科 6,342 名医护在八个月内向机构助手发送的 127,833 条查询,发现文档与行政占 36.2%、知识检索占 28.9%,诊断仅 3.7%,且逾三分之一查询按原样难以良好作答。将 RCQ-Map 应用于 58 个公开基准构建的 Clinical AI Benchmark Atlas 显示,基准中位不含文档请求,任务构成与真实使用仅重合 31%,评测需匹配真实临床使用。
Heat trend
Current heat 9·Comparable peak 10(Oct 9)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.