כתבה
arXiv cs.AI ·
VLA-Precision: שיטה חדשה לשיפור דיוק וחוזרות של דגמי VLA
VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models
VLA-Precision היא שיטה חדשה לשיפור דיוק וחוזרות של דגמי VLA. השיטה משתמשת בשיטת קובוסטראפינג אסימטרי לשיפור דיוק וחוזרות של הדגמים. השיטה נבחנה בתשע תחומי כימיה והציגה תוצאות משמעותיות.
תקציר מקורי באנגליתarXiv:2609.04355v1 Announce Type: cross Abstract: Pretrained vision-language-action (VLA) models enable broad manipulation but remain unreliable in tasks demanding precision and repeatability. Applying real-world online reinforcement learning (RL) to VLA post-training enables autonomous trial-and-error improvement beyond demonstrations alone, but exposes two bottlenecks: 1) unreliable value signals can induce policy drift; 2) large-VLA overhead constrains throughput and sample efficiency. To address these challenges, we present VLA-Precision, an efficient real-world online RL framework featuring the Asymmetric Co-Bootstrapping (ACoB) algorithm and the ACoB-Stream architecture. Specifically, ACoB establishes asymmetric co-bootstrapping across timescales: early intervention-guided behavioral
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית