כתבה
arXiv cs.LG ·
TACO: אופטימייזר חדש לקיפול דגמי LLM
TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning
TACO הוא אופטימייזר חדש שמקטין את השימוש בזיכרון בעת קיפול דגמי LLM. הוא מאפשר קיפול מלא של מודלים בני 30-32B פרמטרים על גפי H100 בודד. TACO משתמש בגישה גאומטרית חדשה כדי לחסוך בזיכרון.
תקציר מקורי באנגליתarXiv:2610.02199v1 Announce Type: new Abstract: Full-parameter fine-tuning of large language models (LLMs) incurs substantial optimizer state memory overhead, limiting the model sizes that fit on modern GPUs. Existing approaches either compress optimizer state, abandon first-order gradients, or change the update geometry while retaining dense state. The recently introduced Muon optimizer reduces optimizer memory through matrix-valued updates. Still, its geometry differs from AdamW and can lead to performance degradation when fine-tuning AdamW-pretrained models. To reduce optimizer memory without sacrificing accuracy or computational efficiency in LLM fine-tuning, we propose Ternary Absolute-max Column-wise One-sparse optimizer, or TACO, which follows Muon's operator-norm steepest-descent v
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית