arXiv · Computers and Society· David Gringras, Misha Salahshoor·· 3 小时前AI 评分58
前沿滞后:学术 AI 评测中的能力误报文献审计
Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation
AI 导读
研究系统审计了 OpenAlex 中 2022-2026 年间的 112,303 条 LLM 关键词匹配记录,发现论文评测的中位模型比同期前沿 LLM 低 10.45 ECI,且差距以每年 4.07 ECI 扩大。仅 18.4% 的全文明确标注评测日期,52.5% 的摘要以
来源:arXiv · Computers and Society · arxiv.org