כתבה
arXiv cs.AI ·
RoboSPA: האם דגמי VLA יכולים לעבור מאתגרים פשוטים ומשימות קצרות-תקופה?
RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?
דגמי VLA עדיין נתקלים בבעיות קשות בתחום הספציפיות המדויקות ותכנון תפעולי באורך זמן ארוך.
תקציר מקורי באנגליתarXiv:2609.05324v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation. However, existing datasets and benchmarks mainly evaluate task completion under predefined settings, offering limited insight into model reasoning under increasing spatial and procedural complexity. We introduce \textbf{RoboSPA} (\textbf{Robo}t \textbf{S}patial-\textbf{P}rocedural \textbf{A}ssessment), a large-scale robotic manipulation dataset and benchmark for diagnosing embodied reasoning in VLA models. \texttt{RoboSPA} focuses on two core dimensions, Fine-Grained Spatial Reasoning and Long-Horizon Procedural Planning, covering 10 task categories and 56 base tasks. Each task is instantiated across five difficulty levels, yieldi
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית