יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

ComputerSD: אימון אונליין של סוכני מחשבים דרך פידבק בזמן אמת

ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents
אימון אונליין של סוכני מחשבים דרך פידבק בזמן אמת. המחקר מציג טכניקה חדשה לאימון סוכני מחשבים שמשתמשת בפידבק בזמן אמת. הטכניקה, שנקראת ComputerSD, מאפשרת לסוכני המחשבים ללמוד מהפידבק ולשפר את יכולותיהם. המחקר מציג תוצאות של ComputerSD בהשוואה לטכניקות אימון אחרות.
תקציר מקורי באנגליתarXiv:2609.40253v2 Announce Type: replace-cross Abstract: Online training enables computer-use agents (CUAs) to improve through interaction with executable environments. However, existing methods primarily rely on sparse outcome rewards, which provide no supervision for intermediate actions. On-policy self-distillation (OPSD) offers token-level learning signals through privileged rescoring, but directly applying it to CUA online training presents two challenges: fixed guidance may become misaligned with the student's current state, and guidance-induced probability shifts may conflict with step-level correctness. We introduce ComputerSD, an online self-distillation method for CUAs that converts real-time feedback from executed GUI transitions into guidance for policy learning. A fine-tuned
קרא במקור המקורי