כתבה
arXiv cs.LG ·
More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe
תקציר מקורי באנגליתarXiv:2607.15942v2 Announce Type: replace-cross Abstract: Remote sensing vision-language models are increasingly expected to support open-ended reasoning over Earth Observation data and a variety of tasks. Most recent progress in this area has been driven by remote-sensing-specific architectural designs, often introducing new encoders, alignment modules, or task-specific fusion mechanisms. In this work, we challenge the necessity of such architectural specialization. We show that a generally capable vision-language model can achieve competitive or state-of-the-art performance at challenging remote sensing benchmarks, provided that it is trained at sufficient scale across diverse data and tasks. Our model uses a single language policy that can either answer directly in text or invoke a loca
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית