כתבה
arXiv cs.LG ·
Headroom-Drift Replay: טכניקה חדשה לשימוש מחדש של נתונים באימון GRPO
Headroom-Drift Replay: A Primitive for Principled Replay Control in GRPO
אנו מציגים טכניקה חדשה לשימוש מחדש של נתונים באימון GRPO, המפחיתה את העומס על RL-based post-training למודלי סיבוכיות.
תקציר מקורי באנגליתarXiv:2609.03941v1 Announce Type: new Abstract: RL-based post-training for reasoning models is increasingly bottlenecked by repeated fresh rollout generation, particularly in agentic settings where environment interaction dominates wall-clock cost. Replay can reduce this burden by reusing past trajectories, but existing methods typically embed it within larger training pipelines involving exploration, experience restructuring, or mixed-policy optimization. This makes replay's own contribution difficult to isolate. We ask a focused question: how far can principled replay selection alone go? We introduce Headroom-Drift Replay, a group-level replay control primitive for GRPO that separates reuse into two decisions. Headroom ranks stored groups by remaining learning value, while Drift gates th
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית