יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

SpatialCORE: תרגום-מובנה של סמכות-מודלי ספציפי לתחום החזותי-שפתי

SpatialCORE: Confidence-Aware Grounded Spatial Reasoning in Large Vision--Language Models
מודלי חזותי-שפתי גדולים סובלים מחולשה בתחום הספציפי. SpatialCORE פותרת זאת עם תשתית-מובנה של סמכות.
תקציר מקורי באנגליתarXiv:2609.38716v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have made remarkable progress across visual perception tasks, yet spatial reasoning remains a persistent weakness, especially for questions that require reasoning over visual space. Recent spatial-reasoning methods incorporate generated grounding, where models predict bounding boxes, masks, or other localization outputs for task-relevant objects as part of their reasoning trace. However, these approaches typically optimize final-answer correctness alone, allowing correct answers to be rewarded even when the model does not reason from confidently localized task-relevant objects. We introduce SpatialCORE (Spatially COnfident REasoning), a post-training framework that turns the model's own confidence in gen
קרא במקור המקורי