יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

CF-VLA: ייצוג פעולות יעיל ומפורט למדיניות VLA

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies
CF-VLA מציע ייצוג פעולות יעיל ומפורט למדיניות VLA, על ידי שימוש בשני שלבים: פיתוח נקודת התחלה מבוססת-פעולה ושיפור מקומי.
תקציר מקורי באנגליתarXiv:2604.24622v4 Announce Type: replace-cross Abstract: Flow-based vision-language-action (VLA) policies offer strong expressivity for action generation, but suffer from a fundamental inefficiency: multi-step inference is required to recover action structure from uninformative Gaussian noise, leading to a poor efficiency-quality trade-off under real-time constraints. We address this issue by rethinking the role of the starting point in generative action modeling. Instead of shortening the sampling trajectory, we propose CF-VLA, a coarse-to-fine two-stage formulation that restructures action generation into a coarse initialization step that constructs an action-aware starting point, followed by a single-step local refinement that corrects residual errors. Concretely, the coarse stage lear
קרא במקור המקורי