arXiv · Artificial Intelligence· Yulin Fu (Beijing University of Posts and Telecommunications), Junren Wang (West China Hospital, Sichuan Provincial Engineering Research Center of Intelligent Diagnosis and Treatment of Breast Diseases), Guangjing Yang (Beijing University of Posts and Telecommunications), Zhangyuan Yu (Beijing University of Posts and Telecommunications), Wanran Sun (Beijing University of Posts and Telecommunications), Jiabao Zhou (Beijing University of Posts and Telecommunications), Jin Yin (West China Hospital, Sichuan Provincial Engineering Research Center of Intelligent Diagnosis and Treatment of Breast Diseases), Qicheng Lao (Beijing University of Posts and Telecommunications)·· 4 hr agoAI score37
MedBenchAgent:面向医疗 VLM 基准构建的系统化自动化
MedBenchAgent: Towards Systematic Automation of Medical VLM Benchmark Construction
AI brief
MedBenchAgent 将医疗视觉语言模型(VLM)基准构建形式化为约束编译,自动推导评测规范本身,包括评测内容、支撑每项任务的标注及证据到评测项的转化。
Source: arXiv · Artificial Intelligence · arxiv.org