Skip to content
arXiv · Information Retrieval· Qingfang Liu, Qiao Jin, Joe D. Menke, Thorsten Kahnt, Zhiyong Lu·· 7 hr agoSelectedAI score60

AI 聊天机器人检索医学证据的效果:模型、用户角色与样本量的影响

Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions

AI brief

研究评估 Claude Sonnet 5、Gemini 3.1 Pro 和 ChatGPT GPT-5.5 在 20 个临床问题上的检索表现,ChatGPT 召回率 63.1%,显著高于 Claude 的 37.0% 和 Gemini 的 17.3%;研究者角色召回高于临床与患者角色;样本量每翻倍,检索几率提高 50%(OR 1.50,95% CI 1.24-1.81)。

Why it matters

三款主流大模型在医学问答中的检索召回差异显著,读者可借此评估自身工作流中模型选型的证据基础。

Source: arXiv · Information Retrieval · arxiv.org