כתבה
arXiv cs.LG ·
CERO: הקצאת רולאוטים לאימון RL
CERO: Where and When to Allocate Rollouts for RL Post-Training
CERO הוא שיטה להקצאת רולאוטים לאימון RL. היא מאפשרת הקצאה יעילה של תקציב רולאוטים על פני כל האימון. CERO משתמשת בייצוג קומפקטי ובשיטות אופטימיזציה מקוונות.
תקציר מקורי באנגליתarXiv:2610.09679v1 Announce Type: new Abstract: Adaptive rollout methods for group-relative reinforcement learning typically allocate a fixed per-update budget across prompts. We instead study how to coordinate a finite rollout budget over the entire training horizon. We formulate this problem using a concave surrogate utility of cumulative prompt exposure and introduce CERO, an online primal dual scheduler for prompt admission and budget pacing. In our experiments, each admitted prompt receives a fixed-size response group. CERO instead adapts which prompts are selected, how often they are revisited across rounds, and how many groups are generated in each round. A compact Fenchel representation linearizes the dependence on cumulative exposure, while projected online gradient descent update
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית