יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

EvoCUA-1.5: רכיבת השתלמות חיה לאג'נטים לשימוש במחשב

EvoCUA-1.5: Online Reinforcement Learning for Multi-turn Computer-Use Agents
EvoCUA-1.5 ממשיך אג'נטים לשימוש במחשב ללמידת השתלמות חיה, שם הם משתכללים באמצעות רכיבת השתלמות חיה. האג'נטים נחשפים לסביבה פרקטית, והם משתכללים באמצעות תוצאות מעשיות. EvoCUA-1.5 פותח פרקטיקה יעילה לקידום למידת השתלמות חיה באג'נטים לשימוש במחשב.
תקציר מקורי באנגליתarXiv:2607.09773v2 Announce Type: replace-cross Abstract: Computer-use agents must solve long-horizon tasks through repeated interaction with partially observable, multimodal desktop environments. Although imitation learning and offline trajectory refinement provide strong priors, static traces cannot cover the causal feedback loop of real computer use: each action changes the screen state, future action space, and recovery options. EvoCUA-1.5 extends self-evolving computer-use agents from offline experience learning to online reinforcement learning, where policies interact with executable sandbox environments and improve from verifiable task outcomes. Online RL in this setting requires more than directly reusing single-turn language-RL recipes. Multi-turn interaction introduces context-ma
קרא במקור המקורי