Skip to content
Hot eventLive

研究揭示文档VLM检索-阅读差距

1 reports1 sources16 hr ago updated

Get the story

AI overview

2026年10月5日,arXiv计算机视觉方向发布一手研究,提出“检索—阅读鸿沟”概念:文档视觉语言模型(VLM)即使检索到正确页面,也常常不利用其中证据。研究团队构建了带可追溯证据的 FoveaDoc-Bench。实验显示,检索几乎覆盖全部证据页,但加入 CPU-OCR 文本后,严格准确率提升13至16个百分点;使用精确文本层时,增益约翻倍。该结果表明,检索成功并不等于阅读有效,文本提取质量对文档VLM表现有显著影响。

Generated from reports · updated 16 hr ago

Timeline

Follow the coverage from different angles.

Oct 5, 2026
  1. arXiv · Computer Vision
    检索到却读不懂:提取文本何时弥合文档视觉语言模型的检索—阅读鸿沟

    研究提出检索—阅读鸿沟:文档视觉语言模型(VLM)即使检索到正确页面,也常不利用其中证据。在带可追溯证据的 FoveaDoc-Bench 上,检索几乎覆盖全部证据页,但加入 CPU-OCR 文本使严格准确率提升 13 至 16 个百分点,精确文本层约使增益翻倍。

Heat trend

Current heat 7·Comparable peak 10(Oct 5)·Comparable change over 24 hours –

02.557.510Oct5Oct5Oct5Oct6

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.