כתבה
arXiv cs.AI ·
ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents
תקציר מקורי באנגליתarXiv:2609.40253v1 Announce Type: cross Abstract: Online training enables computer-use agents (CUAs) to improve through interaction with executable environments. However, existing methods primarily rely on sparse outcome rewards, which provide no supervision for intermediate actions. On-policy self-distillation (OPSD) offers token-level learning signals through privileged rescoring, but directly applying it to CUA online training presents two challenges: fixed guidance may become misaligned with the student's current state, and guidance-induced probability shifts may conflict with step-level correctness. We introduce ComputerSD, an online self-distillation method for CUAs that converts real-time feedback from executed GUI transitions into guidance for policy learning. A fine-tuned GUI anal
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית