כתבה
arXiv cs.LG ·
האם שיטות תכנות יעילות לאימון חדש נבדלות באמת?
Are Parameter-Efficient Fine-tuning Methods Really Different?
חוקרים השוו שישה שיטות שונות לאימון חדש של מודלי שפה והתפשטות. הם חקרו את יעילותן, את הזיכרון שלהם ואת השינויים בגאומטריה של המודלים המוכנים מראש. התוצאות הראו ששיטות מסוימות יעילות יותר במקרים מסוימים, אך הן גם גורמות לאיבוד יותר של זיכרון. החוקרים גם גילו ששימור הגאומטריה של המודל המוכן מראש אינו חשוב כל כך, ושהשיטות השונות פועלות באופן שונה.
תקציר מקורי באנגליתarXiv:2610.09122v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) offers many parameterizations, yet their methodological and functional differences remain unclear. We compare six methods in language and diffusion models to examine how their parameterizations relate to task performance, forgetting, and changes in pretrained weight geometry. Motivated by the spectrum-preserving design of orthogonal fine-tuning (OFT), we first ask whether spectral preservation is itself important for adaptation and retention. We find that the selected LoRA-family methods also approximately preserve pretrained geometry, and that restoring their slightly drifted singular-value spectra largely preserves task performance, questioning the necessity of explicit geometric preservation. Beyond t
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית