כתבה
arXiv cs.AI ·
רב-ערכיות ויזואלית-שפתית להשגה תוך-תפקודית דרך פריורים יצירתיים והשגה כללית
Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp Synthesis
מערכת רב-ערכיות ויזואלית-שפתית להשגה תוך-תפקודית, המשתמשת בפריורים יצירתיים והשגה כללית. המערכת, שנקראת AdaRoboVLG, תומכת בהשגה כללית של ידיים רב-ערכיות. המערכת משתמשת בפריורים יצירתיים כדי להשיג ידיים כלליות, ומשתמשת בהשגה כללית כדי להשיג ידיים רב-ערכיות. המערכת נבחנה באמצעות ניסויים בסימולציה ובעולם האמיתי.
תקציר מקורי באנגליתarXiv:2609.04096v2 Announce Type: replace-cross Abstract: This paper proposes AdaRoboVLG, a task-adaptive Vision-Language-Grasp (VLG) framework that supports generalizable grasp synthesis across different robotic hands. Unlike existing VLG methods that tightly couple foundation models with end-to-end grasp policies, AdaRoboVLG learns an efficient generalizable base policy that generates and evaluates physically feasible grasp candidates through explicit kinematic mapping and force-closure-based stability estimation, while offloading task-dependent understanding to specialized foundation-model modules. These modules provide composable priors that are integrated into the grasp synthesis process, enabling contextually adaptive grasp synthesis without retraining the underlying grasp policy. Th
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית