Skip to content
Hot eventLive

MLLM零样本语言推理用于跨视角地理定位研究

1 reports1 sources3 hr ago updated

Get the story

AI overview

2026年10月7日,arXiv计算机视觉方向发布研究《What Words Keep of a Place》,提出一种零样本跨视角地理定位方法。该方法用多模态大语言模型(MLLM)把地面全景与卫星瓦片描述成结构化文本并比对,全程无需训练。研究在9,826对VIGOR数据上评估:全库按描述相似度排序,Recall@1仅0.39%;将候选缩到10个邻近瓦片后,MLLM评判器把随机排序效果翻倍,并追平强词法基线,还能指出字段一致或冲突。

Generated from reports · updated 3 hr ago

Timeline

Follow the coverage from different angles.

Oct 7, 2026
  1. arXiv · Computer Vision
    What Words Keep of a Place:跨视角地理定位的零样本语言推理研究

    研究提出零样本跨视角地理定位方法,用多模态大语言模型(MLLM)把地面全景与卫星瓦片描述成结构化文本并比对,全程无需训练。在 9,826 对 VIGOR 数据上,全库按描述相似度排序 Recall@1 仅 0.39%;缩到 10 个邻近瓦片后,MLLM 评判器将随机排序效果翻倍并追平强词法基线,还能指出字段一致或冲突。

Heat trend

Current heat 9·Comparable peak 10(Oct 7)·Comparable change over 24 hours –

02.557.510Oct7Oct7Oct7Oct7

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.