כתבה
arXiv cs.LG ·
אותו סצנה, משימה שונה
Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs
CRAFT משפר את הכללת המשימות במודלים של ראייה-שפה-פעולה. הוא עושה זאת על ידי העברת פיקוח מביצועים מדגומים לזוגות נגדיים. זה מאפשר למודלים ללמוד משימות חדשות ולהצליח בהן.
תקציר מקורי באנגליתarXiv:2610.00524v1 Announce Type: cross Abstract: Vision-language-action (VLA) models often struggle to generalize to skill combinations absent from their fine-tuning demonstrations, even when every constituent skill has been demonstrated. We focus on a vision shortcut as one failure mode: during fine-tuning, visual observations can serve as a proxy for the instruction, so a policy may execute a demonstrated combination associated with similar observations rather than the instructed combination. This motivates training with counterfactual pairs formed by holding a demonstration observation fixed while changing the instruction to specify an undemonstrated combination. These pairs, however, lack corresponding demonstrated action targets. Crucially, the currently required skill has already be
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית