Skip to content
Hot eventLive

研究:VLM小物体识别极限与局部分解规则

1 reports1 sources4 hr ago updated

Get the story

AI overview

该研究关注视觉语言模型(VLM)在大图中漏检小目标的问题。论文提出用目标侧向视觉 token 数 S 与单次调用覆盖内容 L 来解释这一现象:整图送入时,目标侧 token 至多为 m*sqrt(N/(WH));token 预算翻倍只让最大 S 提高约 41%,且达到所需 S 至少需要约 S^2 量级 token。研究据此说明局部放大在何种条件下是安全的。目前未见与早先报道矛盾之处。

Generated from reports · updated 3 hr ago

Timeline

Follow the coverage from different angles.

Oct 8, 2026
  1. arXiv · Computer Vision
    研究解释 VLM 为何漏检小目标及局部放大何时安全

    论文提出用目标侧向视觉 token 数 S 与单次调用覆盖内容 L 解释 VLM 在大图中漏检小目标:整图送入时目标侧 token 至多为 m*sqrt(N/(WH)),token 预算翻倍只让最大 S 提高约 41%,且达到所需 S 至少需要约 S^2 量级 token。

Heat trend

Current heat 9·Comparable peak 10(Oct 8)·Comparable change over 24 hours –

02.557.510Oct8Oct8Oct8Oct8

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.