כתבה
arXiv cs.AI ·
GeoContext: סולם תפאורה אחד, שני תסביך כשל בגאולוגיה-לשון: תלות פשוטה במיקום המשתמש
GeoContext: One Context Ladder, Two Failure Modes in Vision-Language Geolocation: Flat Reliance on User-Provided Location Context and False Confirmation of Location Claims
GeoContext: סולם חדש לבדיקת גאולוגיה-לשון שבודק את יכולת המודלים למקום תמונות נתונות תפאורה של מיקום המשתמש. הסולם כולל שני תפקודים: GeoHint, מקום פתוח-מסתורי נתון תפאורה קצרה, ו-GeoVerify, סימון בינארי של האם התמונה נלקחה בתוך 150 מטרים ממקום טענה.
תקציר מקורי באנגליתarXiv:2609.05761v1 Announce Type: cross Abstract: Visual geolocation benchmarks typically ask a model where an image was captured without accounting for the location context that users often provide. We introduce GeoContext, a resource supporting two complementary tasks: GeoHint, open-ended localization given a true but coarse location hint, and GeoVerify, binary verification of whether an image was taken within 150 m of a claimed place. GeoContext constructs a context ladder by stratifying nearby reference points according to distance and referenceability, allowing the image to remain fixed while the supplied context varies. The benchmark covers 109 sites in 30 cities and evaluates five vision-language models using 21,933 GeoHint responses and 6,270 GeoVerify responses. Our evaluation rev
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית