arXiv · Computer Vision· L. D. M. S. Sai Teja, Ufaq Khan, N. Siva Gopala Krishna, Satyajit Tourani, Ashshak Sharifdeen, Fida Mohammad Thoker, Bernard Ghanem, Muhammad Haris Khan·· 3 小时前AI 评分42
Composition, Not Conversation: VLMs Lose the Scene, Not the Thread
Composition, Not Conversation: VLMs Lose the Scene, Not the Thread
AI 导读
研究提出 Layered-VQA 数据集,包含 93 个场景、300 个问题,每个图像分解为有序 RGBA 层,问题标注支持层、最小充分层和干扰层。评估 11 个 3B-32B 开源权重 VLM 与 2 个专有模型,共 187,200 次对话、1.74M 次交叉评判。发现三方面失败:场景碎片化显著降低准确率、Oracle Inversion、证据需求增加时 grounding 退化快于答案准确率。
来源:arXiv · Computer Vision · arxiv.org