כתבה
arXiv cs.LG ·
מהטבעי למושג: חימום-על-מדיני לפעילות RL
From Imitation to Reward Discovery: On-Policy Warmup for Agentic RL
חימום-על-מדיני לפעילות RLVR: פיתוח יעיל של סוכני-שפה. חימום-על-מדיני (OPW) הוא שלב מודרך שבו התלמיד מתאמן תחת ניהול-מורה על-פי נתיבי-מעבר שלו עצמו, לפני עבריית לRLVR. OPW יעיל יותר מאימיטציה על-פי נתיבי-מעבר קבועים, ומספק תאוריה-תאורטית להסברה של הפעילות-על-מדיני.
תקציר מקורי באנגליתarXiv:2609.39436v1 Announce Type: new Abstract: Reinforcement learning with a verifiable reward (RLVR) offers a scalable approach to training language-model agents, yet sparse outcome rewards can leave early training with little signal for policy improvement. We identify an On-Policy Acceleration Phenomenon: in our main comparisons, RLVR initialized with on-policy distillation reaches high performance earlier in training and achieves both higher average performance during subsequent RLVR and higher final performance than the alternative baselines. Motivated by this observation, we study On-Policy Warmup (OPW), a teacher-guided stage in which the student trains with teacher supervision on its own interaction trajectories before transitioning to RLVR. Unlike imitation on fixed teacher-genera
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית