כתבה
arXiv cs.LG ·
על יכולות נוספות ואיחוד מודלים
On Emergent Capabilities and Model Merging
במאמר זה, החוקרים חקרו את השפעת איחוד מודלים על יכולות נוספות. התוצאות הראו שאיחוד מודלים לא יוצר יכולות חדשות, אלא משמר את היכולות הקיימות.
תקציר מקורי באנגליתarXiv:2609.24504v2 Announce Type: replace Abstract: Fine-tuned checkpoints and adapters now fill public repositories, and the most common operation applied to these artifacts is model merging: arithmetic on their weights that assembles capabilities cheaply. We ask what this operation does to emergent capabilities: behaviors an artifact carries that were never an explicit training target. Studying two independent testbeds (activation oracles and emergent-misaligned models) across three model families, we find that the answer is threefold. First, merging preserves an emergent capability that both parents carry: merging two misaligned checkpoints retains most of their broad misalignment across the whole mixing range. Second, merging cannot create an emergent capability that is superadditive i
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית