כתבה
arXiv cs.AI ·
למידה מעבר למה שאתה סוגר: חלופה-פוליסי-מודע לשינוי נתיבי מודלים
Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR
מסגרת חלופה-פוליסי-מודע לשינוי נתיבי מודלים, המשפרת את ביצועי המודל. המאמר מציג פרוטוקול חדש לשינוי נתיבי מודלים, המאפשר למודלים שונים ללמוד זה מזה. הפרוטוקול, הנקרא GRAFT, משתמש בשינוי נתיבי מודלים כדי לשפר את ביצועי המודלים. המאמר מציג תוצאות של GRAFT על מודלים שונים, ומראה כי GRAFT משפר את ביצועי המודלים.
תקציר מקורי באנגליתarXiv:2609.37868v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) methods such as GRPO rely on successful self-generated trajectories, but finite rollout budgets can produce all-fail groups with no reward-based policy-gradient signal. While additional rollouts improve the chance of success at higher cost, successful trajectories missing from one model's rollouts may already have been discovered by another. Indeed, we observe that heterogeneous models often succeed on complementary prompts, creating opportunities for mutual learning without a designated stronger teacher. To exploit this complementarity, we propose GRAFT (Gated Replacement of Answer-Failed groups with peer Trajectories), an off-policy-aware framework that replaces all-fail groups with in
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית