כתבה
arXiv cs.CL ·
Specializing Without Forgetting: Analyzing Knowledge Preservation in Multilingual Model Adaptation
תקציר מקורי באנגליתarXiv:2606.00284v2 Announce Type: replace Abstract: While continual pretraining (CPT) is a practical way to extend large language models to new languages, na\"ive finetuning often erodes existing capabilities through catastrophic forgetting. We investigate which model layers drive this trade-off, and whether interventions at these layers can guide knowledge preservation during adaptation. We interpolate gemma-3-4b model states before and after CPT on five language families to localize forgetting on reading comprehension and translation, finding that middle-layer reversion yields the largest comprehension recovery, while translation effects vary by language family and direction. Guided by these findings, we evaluate CPT strategies that leverage this layer information to mitigate forgetting:
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית