יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

קירור המדגם, לא הלומד: טמפרטורת הדגימה מזיזה את המצוק המיושן של GRPO

Cool the Sampler, Not the Learner: Sampling Temperature Moves the Staleness Cliff of Importance-Corrected GRPO
חוקרים גילו כי קירור המדגם, ולא הלומד, משפר את יציבות GRPO. הם בדקו את השפעת טמפרטורת הדגימה על היציבות, ומצאו כי היא יכולה למנוע הידרדרות בביצועים. המחקר בוצע עם מודל Qwen, והראה כי קירור המדגם יכול לשפר את התוצאות.
תקציר מקורי באנגליתarXiv:2609.36953v1 Announce Type: cross Abstract: Production RL for language models lets the sampler fall behind the learner and repairs the resulting mismatch with a truncated importance weight. We ask how long the sampler can go without a refresh under that correction, and find a cliff: on Qwen2.5-Math-1.5B and GSM8K, importance-corrected GRPO refreshed every 192 updates learns well for 180 steps and then degrades severely in all three data seeds before the refresh arrives. Published remedies for staleness act on the update; we act on the sampler instead. Decoupled cooling draws samples at temperature 0.8 while the learner, the reference model and the importance weights stay at temperature 1, with the behaviour probability recorded from the tempered distribution, so the learner's objecti
קרא במקור המקורי