כתבה
arXiv cs.AI ·
לאן להתאים חשיבות: טיפול נבחר בשכבות לשימור יכולת
Where to Adapt Matters: Layer-Selective Fine-Tuning for Capability Retention
טיפול נבחר בשכבות של מודלי Transformer מציע תוצאות שונות לגבי יכולת המודל לבצע משימות ספציפיות ולשמור יכולות כלליות. ניתן לשמור יכולות כלליות יותר של המודל על ידי טיפול נבחר בשכבות שלו. ניתן לשמור יכולות כלליות יותר של המודל על ידי טיפול נבחר בשכבות שלו.
תקציר מקורי באנגליתarXiv:2610.11620v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) enables large language models (LLMs) to adapt to specialized tasks, but often at the cost of degrading general capabilities acquired during pretraining. Existing approaches primarily mitigate this trade-off through data replay or regularization, relying on additional data or explicit optimization constraints. We instead focus on a different question: where should adaptation be applied? We find that fine-tuning different Transformer layers produces different target-task gains and degrees of capability degradation, suggesting that not all layers are equally suitable for adaptation. To characterize this difference, we use layer-wise empirical Fisher information to measure target-task sensitivity. However, c
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית