כתבה
arXiv cs.LG ·
Train4Merge: חקירה מאורגנת של RL vs. SFT מורים למיזוג OPD-בסיסי
Train4Merge: A Controlled Single-Teacher Study of RL vs. SFT Teachers for OPD-Based Model Merging
במאמר זה, חוקרים חקרו את השפעת שיטות הלימוד של RL ו-SFT על יכולת המיזוג של דגמי למידת מכונה. התוצאות הראו ש-RL יוצרים מורים חזקים יותר, שמאפשרים לסטודנטים ללמוד טוב יותר.
תקציר מקורי באנגליתarXiv:2609.32303v2 Announce Type: replace-cross Abstract: Domain experts trained from a shared checkpoint can transfer their specialized capabilities to a single student through on-policy distillation (OPD). Existing research primarily focuses on improving this merging process, while the algorithms used to train the experts have received limited systematic comparison. We investigate which training algorithm produces teachers better suited to OPD through controlled single-teacher comparisons of supervised fine-tuning (SFT) and reinforcement learning (RL) across Agentic, Reasoning, and Perception. Teachers and students share the same Qwen3.5-9B initialization, and the two teacher types are compared at similar task performance. Our experiments show that RL teachers yield stronger students and
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית