כתבה
arXiv cs.AI ·
SERA: חלוקת תקציב רולאוט להפצת סיכוי סורבל-מאוזן ללמידת השלמה בעקביות
SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning
SERA מציגה חלוקת תקציב רולאוט שמפצה סיכוי סורבל-מאוזן ללמידת השלמה בעקביות. המאמר עוסק בשיפור ביצועי MaxRL על ידי שימוש בחלוקת תקציב רולאוט שמפצה את הסיכוי הסורבל-מאוזן. המחברים מציגים פתרון לבעיה זו ומדגימים את יעילותו במספר תרחישים.
תקציר מקורי באנגליתarXiv:2609.36552v1 Announce Type: cross Abstract: Maximum Likelihood Reinforcement Learning (MaxRL) targets prompt-wise log-success and has shown strong performance on reasoning tasks. Under finite rollout budgets, however, the estimator used by MaxRL attenuates each prompt's likelihood gradient by a factor that depends on its success probability and rollout count. Under uniform rollout allocation, the common rollout count fails to compensate for success-dependent attenuation, leaving low-success prompts more strongly attenuated and distorting their relative contributions to the expected aggregate gradient. We introduce SERA (Scale-Equalized Rollout Allocation), which redistributes a fixed rollout budget to approximately equalize these finite-rollout scaling factors. Building on our theore
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית