כתבה
arXiv cs.AI ·
GaugeVLM: הקצאת פיקוח מרחבי
GaugeVLM: Structuring Spatial Supervision with Measured Geometric Interventions
GaugeVLM הוא מודל חדש שמשפר את היכולת להבין יחסים מרחביים. הוא משתמש בהתערבויות גאומטריות כדי ליצור נתונים מובנים. התוצאות מראות שיפור ב-10 מדדים שונים.
תקציר מקורי באנגליתarXiv:2609.38285v1 Announce Type: cross Abstract: Vision-language models (VLMs) can contradict themselves across views of the same spatial relation and fail to respond when that relation changes. Addressing these failures requires supervision that captures error magnitude and geometric dependencies across observations, both of which remain implicit in training on individual answers or ordinal preferences. Therefore, we introduce GaugeVLM, which makes this structure explicit through controlled object and camera interventions in explicit 3D scenes, producing linked observations with measured differences between spatial relations and shared truths across views. To translate this structure into learning signals, its core objective, GaugeDPO, converts measured errors into preference margins, di
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית