כתבה
arXiv cs.CL ·
SpaceCast-Bench: בדיקת יכולת התכנות המרחבית הטיפולית במודלי תצוגה-שפה
SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models
SpaceCast-Bench: בדיקת יכולת התכנות המרחבית הטיפולית במודלי תצוגה-שפה. המאמר עוסק בבניית בנק אימונים חדש, SpaceCast-Bench, ובבדיקת 21 מודלי תצוגה-שפה. התוצאות מציגות פער גדול בין המודלים החזקים לבין הביצועים האנושיים.
תקציר מקורי באנגליתarXiv:2610.12402v1 Announce Type: cross Abstract: Existing spatial reasoning benchmarks mainly test spatial perception: reading off relations already visible in the input. Yet real-world spatial intelligence demands predictive spatial reasoning: constructing a scene from observations, anticipating how an intervention changes it, and reasoning about the unseen outcome. We introduce SpaceCast-Bench, the first benchmark to directly and diagnostically evaluate this capability. Built around an observe-transform-infer framework, its 3,862 questions from 182 real-world scenes span 16 task types at three levels: static perception, local prediction, and global prediction, progressively requiring scene understanding, spatial state updating, and relational inference over unobserved outcomes. Evaluati
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית