יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

DriftOPD: הפצת מדיניות VLA ברמת רצף

DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies
DriftOPD הוא כלי להפצת מדיניות Vision-Language-Action (VLA) ברמת רצף. הוא מאפשר אופטימיזציה של מדיניות VLA ללא צורך באינטראקציה מקוונת או מורה נפרד. DriftOPD משתמש בשיטת הפצה יחידה וב-Q-function critic שנלמד מדמואים אופליין.
תקציר מקורי באנגליתarXiv:2610.00317v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models increasingly rely on action experts that generate short action chunks under receding-horizon control. While chunk-level training is convenient across robot embodiments, it optimizes local action likelihood without explicitly accounting for long-horizon task success. Sequence-level reinforcement learning can address this limitation, but typically requires policy rollouts and closed-loop interaction, which are costly for real-robot manipulation. We introduce DriftOPD, a teacher-free, rollout-free framework for sequence-level on-policy distillation of continuous VLA action experts. We show that the sequence-level reverse Kullback-Leibler (KL) divergence decomposes into a chunk-level reverse-KL term and a fut
קרא במקור המקורי