יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

למידה מעבר לדגימה: חילופי מסלול מודלים ללמידת חיזוק

Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR
GRAFT הוא כלי לחילופי מסלולים בין מודלים שונים, המאפשר למידה משותפת ושיפור ביצועים. הוא משתמש במשקלות תואמות וקיטוע חשיבות כדי להתגבר על פערים בין המודלים. GRAFT הראה שיפורים משמעותיים בביצועים בהשוואה לשיטות קודמות.
תקציר מקורי באנגליתarXiv:2609.37868v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) methods such as GRPO rely on successful self-generated trajectories, but finite rollout budgets can produce all-fail groups with no reward-based policy-gradient signal. While additional rollouts improve the chance of success at higher cost, successful trajectories missing from one model's rollouts may already have been discovered by another. Indeed, we observe that heterogeneous models often succeed on complementary prompts, creating opportunities for mutual learning without a designated stronger teacher. To exploit this complementarity, we propose GRAFT (Gated Replacement of Answer-Failed groups with peer Trajectories), an off-policy-aware framework that replaces all-fail groups with in
קרא במקור המקורי