יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

מעבר מחיקוי לגילוי תגמול

From Imitation to Reward Discovery: On-Policy Warmup for Agentic RL
חוקרים מציעים שיטה חדשה לאימון סוכנים בתחום הלמידה החיזוקית. השיטה, הנקראת On-Policy Warmup, מאפשרת לסוכנים ללמוד מניסיונם הוא ולשפר את ביצועיהם.
תקציר מקורי באנגליתarXiv:2609.39436v1 Announce Type: cross Abstract: Reinforcement learning with a verifiable reward (RLVR) offers a scalable approach to training language-model agents, yet sparse outcome rewards can leave early training with little signal for policy improvement. We identify an On-Policy Acceleration Phenomenon: in our main comparisons, RLVR initialized with on-policy distillation reaches high performance earlier in training and achieves both higher average performance during subsequent RLVR and higher final performance than the alternative baselines. Motivated by this observation, we study On-Policy Warmup (OPW), a teacher-guided stage in which the student trains with teacher supervision on its own interaction trajectories before transitioning to RLVR. Unlike imitation on fixed teacher-gene
קרא במקור המקורי