כתבה
arXiv cs.AI ·
Gondola: תכנון יישר-לשפה למניפולציה רובוטית
Gondola: Grounded Vision Language Planning for Robotic Manipulation
Gondola היא תכנות יישר-לשפה שמטפלת במניפולציה רובוטית. היא יוצרת תוכניות מבוקרות עם קישור פיקסל-לשפה. Gondola נבנתה על ידי צוות של CSHizhe.
תקציר מקורי באנגליתarXiv:2506.11261v2 Announce Type: replace-cross Abstract: Vision-language-action (VLA) models have shown promising progress in robotic manipulation. However, directly mapping visual observations and language instructions to low-level actions often results in limited interpretability and weak robustness in complex, long-horizon tasks. To address these challenges, we employ a modular manipulation framework that separates high-level planning from low-level control. At its core is Gondola, a grounded vision-language planning model that generates structured plans with explicit pixel-level object grounding before action execution. Given multi-view observations and planning history, Gondola predicts the next-step plan as interleaved textual instructions and multi-view segmentation masks correspon
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית