כתבה
arXiv cs.CL ·
אופטימיזציה מונחית גרדיאנט להתאמה של מודלי שפה גדולים
Beyond LoRA vs. Full Fine-Tuning: Gradient-Guided Optimizer Routing for LLM Adaptation
חוקרים הציעו שיטה חדשה להתאמה של מודלי שפה גדולים, MoLF, המשלבת בין Full Fine-Tuning ו-Low-Rank Adaptation. השיטה הוכחה כיעילה במגוון משימות ומודלים, כולל Qwen.
תקציר מקורי באנגליתarXiv:2605.07111v3 Announce Type: replace Abstract: Recent literature on fine-tuning Large Language Models highlights a fundamental debate. While Full Fine-Tuning (FFT) provides greater representational plasticity, Low-Rank Adaptation (LoRA) can match or surpass FFT performance while constraining updates to a low-rank space and potentially benefiting from additional regularization. Through empirical evaluation across diverse tasks (SQL, Medical QA, and Counterfactual Knowledge) and varying language models (Gemma-3-1B, Qwen2.5-1.5B, and Qwen2.5-3B), we observe both trends and find that the better static architecture depends on the task and model. Spectral and truncation analyses further show that endpoint compressibility alone does not explain these task differences, suggesting task-score s
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית