יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

Headroom-Drift Replay: טכניקה חדשה לשימוש מחדש של נתונים באימון GRPO

Headroom-Drift Replay: A Primitive for Principled Replay Control in GRPO
אנו מציגים טכניקה חדשה לשימוש מחדש של נתונים באימון GRPO, המפחיתה את העומס על RL-based post-training למודלי סיבוכיות.
תקציר מקורי באנגליתarXiv:2609.03941v1 Announce Type: new Abstract: RL-based post-training for reasoning models is increasingly bottlenecked by repeated fresh rollout generation, particularly in agentic settings where environment interaction dominates wall-clock cost. Replay can reduce this burden by reusing past trajectories, but existing methods typically embed it within larger training pipelines involving exploration, experience restructuring, or mixed-policy optimization. This makes replay's own contribution difficult to isolate. We ask a focused question: how far can principled replay selection alone go? We introduce Headroom-Drift Replay, a group-level replay control primitive for GRPO that separates reuse into two decisions. Headroom ranks stored groups by remaining learning value, while Drift gates th
קרא במקור המקורי