כתבה
arXiv cs.LG ·
האם למידת הייצוג יכולה להיות נפרדת מהקטיעה של התפוקה? עדכוני קוטבים ישנם תשובה
Can Representation Learning Decouple from Loss Minimization? Polar Updates Have an Answer
במאמר זה, החוקרים חוקרים את השאלה האם למידת הייצוג יכולה להיות נפרדת מהקטיעה של התפוקה. הם מציגים עדכוני קוטבים שיכולים להגביר את הקצב של הלמידה. התוצאות המוצגות במאמר זה יכולות להיות חשובות לפיתוח של מודלי LLM חדישים.
תקציר מקורי באנגליתarXiv:2609.36240v1 Announce Type: new Abstract: Does representation learning stop when the training loss stops improving? We study this question for matrix Muon, whose polar-normalised updates have a step length set by the gradient's rank rather than its norm. Near the edge of stability, full-batch Muon on teacher-student problems enters approximately period-2 loss oscillations that persist for thousands of steps: the cycle-mean loss stays flat or rises, yet the weights keep moving and the learned features continue to align with the teacher subspace. For linear teacher-student learning toys, we derive explicit cycle and alignment formulas and conditional plateau and decay bounds. For a population mean-field ReLU model, we prove that, under stated dimension, initialisation and small-head co
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית