כתבה
arXiv cs.CL ·
Video2Skill: מניפולציה למיומנויות גמישות
Video2Skill: From Streaming Experience to Reusable Embodied Skills
Video2Skill הוא אלגוריתם שמטרתו ללמוד מיומנויות גמישות מסרטונים. הוא בונה ספרייה של מיומנויות שימושיות ומאפשר לסוכנים לתכנ� את פעולותיהם בצורה יותר יעילה. המחקר בודק 19 מודלים שונים ומציע שיטה חדשה לשיפור היכולת ללמוד מיומנויות חדשות.
תקציר מקורי באנגליתarXiv:2609.36691v1 Announce Type: new Abstract: Manipulation behaviors vary widely across objects and scenes, but they share a small set of reusable skills, and planning with these skills helps embodied agents generalize to new tasks. Yet an agent can only plan with skills it knows. Recovering skills from observed experience, the inverse of planning, builds this knowledge over time and yields skill data for training future agents. Vision-Language Models (VLMs) describe individual manipulation events well, but can they organize a stream of events into reusable skills? We formulate this problem as Streaming Embodied Skill Discovery (SESD): a model watches videos in sequence and maintains a persistent skill library that shapes its later decisions. To systematically measure this ability, we in
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית