כתבה
arXiv cs.LG ·
SMAT: טריינינג חדש: פשוט ומצטיין
SMAT: Simple and Efficient Merge-Aware Training
SMAT (Simple Merge-Aware Training) היא שיטת טריינינג שמשפרת את הביצועים של רכיבים מוזגים ללא טריינינג משותף. SMAT משתמשת בשלושה פעולות: סקייל, מאסק ופרטורב.
תקציר מקורי באנגליתarXiv:2609.33437v2 Announce Type: replace Abstract: Model merging integrates the capabilities of multiple experts without joint retraining, but standard expert training optimizes task loss alone and does not guarantee good performance after merging. Merge-aware training (MAT) aims to improve merged performance, but existing methods do not fully account for common merging operations and add training cost. We observe that, from an expert's perspective, common merging methods can be described by three operations: Scale reweights its own update, Mask removes selected coordinates, and Perturb adds updates from other experts. Based on this view, we introduce SMAT (Simple MAT), which jointly optimizes expert loss and expected loss at simulated merged parameters generated by sampling scaling coeff
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית