כתבה
arXiv cs.CL ·
Headroom-Drift Replay
Headroom-Drift Replay: A Primitive for Principled Replay Control in GRPO
Headroom-Drift Replay הוא פרימיטיב חדש לבקרת ריפליי ב-GRPO. הוא מאפשר לבצע ריפליי מבוסס עקרונות, תוך שימוש בדרגון קבוצות וסינון על פי תאימות עם מדיניות נוכחית. השיטה הוכחה כיעילה בבחינות שונות, כולל ריפליי נאיבי ושיטות ריפליי רחבות יותר.
תקציר מקורי באנגליתarXiv:2609.03941v1 Announce Type: cross Abstract: RL-based post-training for reasoning models is increasingly bottlenecked by repeated fresh rollout generation, particularly in agentic settings where environment interaction dominates wall-clock cost. Replay can reduce this burden by reusing past trajectories, but existing methods typically embed it within larger training pipelines involving exploration, experience restructuring, or mixed-policy optimization. This makes replay's own contribution difficult to isolate. We ask a focused question: how far can principled replay selection alone go? We introduce Headroom-Drift Replay, a group-level replay control primitive for GRPO that separates reuse into two decisions. Headroom ranks stored groups by remaining learning value, while Drift gates
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית