כתבה
arXiv cs.LG ·
What Words Keep of a Place: Zero-Shot Language Reasoning for Cross-View Geo-Localization
תקציר מקורי באנגליתarXiv:2610.07269v1 Announce Type: cross Abstract: Cross-view geo-localization is commonly solved as an image retrieval problem, matching a ground-level image against a database of satellite tiles through a jointly trained embedding. Such models are accurate, but they need large paired supervision and cannot show what evidence supports a match. In this paper, we study a different question: how much of this task can be solved through language alone? We prompt a multimodal large language model (MLLM) to describe each ground panorama and each satellite tile as structured text, and localize by comparing these descriptions. No component is trained. We evaluate on 9,826 VIGOR pairs from four U.S. cities, in three settings. First, the descriptions are faithful but not discriminative. They agree cl
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית