כתבה
arXiv cs.LG ·
DataFlex-RL: פלטפורמת הערכה למדיניות נתונים
DataFlex-RL: An Evaluation Platform for RLVR Data Policies
DataFlex-RL היא פלטפורמת הערכה למדיניות נתונים ללמידת חיזוק עם תגמולים מאומתים. הפלטפורמה מאפשרת השוואה בין מדיניות נתונים שונות. הניסוי הראשי העריך 13 קונפיגורציות שונות עם Qwen2.5-7B-Base.
תקציר מקורי באנגליתarXiv:2609.06107v1 Announce Type: new Abstract: Data policies for reinforcement learning with verifiable rewards (RLVR) determine which rollouts are used, how strongly they are weighted, and which domains contribute to subsequent training batches. We introduce DataFlex-RL, an evaluation platform for comparing these choices under a common GRPO recipe. Our primary experiment evaluates 13 configurations across 12 matched seeds using Qwen2.5-7B-Base and 12 mathematics, logic, and science benchmarks. Uniform GRPO improves the domain-balanced average accuracy by 7.76 percentage points over the untrained checkpoint. None of the eight rollout-selection or reweighting methods achieves a paired 95% confidence interval that excludes zero relative to uniform sampling, and none of the three adaptive mi
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית