כתבה
arXiv cs.LG ·
אופטימיזציה: תוכניות זיכרון להגדלת הגבול המעליו של האימון
Optimizer Memory Schedules for Outscaling the Overtraining Axis
במאמר זה, החוקרים חקרו את האופטימיזציה של תוכניות זיכרון כדי להגדיל את הגבול המעליו של האימון. הם השוו את רגישותם של ארבעה אופטימיזציות: AdamW, Muon, SOAP ו-ADANA. התוצאות הראו ש-ADANA עולה על AdamW, ושתוכניות זיכרון חשובות לאופטימיזציה.
תקציר מקורי באנגליתarXiv:2609.04577v1 Announce Type: new Abstract: We investigate how optimizers scale across the overtraining axis and show that relative optimizer performance and optimal hyperparameters change substantially with training horizon. In particular, we study how matrix-preconditioned methods (Muon and SOAP) and a momentum-scheduled method (ADANA) scale relative to AdamW. We compare these four optimizers across models from 51M to 253M parameters and overtraining (OT) factors from 1x to 256x, sweeping the base learning rate at every setting. The preferred learning rate schedule can reverse across the overtraining axis, the best weight decay coefficient scales approximately as sqrt(OT), and longer horizons generally favor longer fixed memory. ADANA's scaling advantage over AdamW persists after tun
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית