יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

VLA-Precision: שיטת חיבור אסימטרית לשיפור עדינות של דגמי VLA

VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models
שיטת שיפור עדינות לדגמי VLA, המאפשרת שיפור עדין של דגמי VLA באופן אונליין. השיטה משתמשת בשיטת חיבור אסימטרית כדי לשפר את העדינות של הדגמים. השיטה נבחנה בתשעה משימות כימיות גבוהות-דיוק והיא הציגה תוצאות משמעותיות.
תקציר מקורי באנגליתarXiv:2609.04355v1 Announce Type: cross Abstract: Pretrained vision-language-action (VLA) models enable broad manipulation but remain unreliable in tasks demanding precision and repeatability. Applying real-world online reinforcement learning (RL) to VLA post-training enables autonomous trial-and-error improvement beyond demonstrations alone, but exposes two bottlenecks: 1) unreliable value signals can induce policy drift; 2) large-VLA overhead constrains throughput and sample efficiency. To address these challenges, we present VLA-Precision, an efficient real-world online RL framework featuring the Asymmetric Co-Bootstrapping (ACoB) algorithm and the ACoB-Stream architecture. Specifically, ACoB establishes asymmetric co-bootstrapping across timescales: early intervention-guided behavioral
קרא במקור המקורי