כתבה
arXiv cs.CL ·
PowerStep: אופטימיזציה אדפטיבית חסכונית בזיכרון
PowerStep: Memory-Efficient Adaptive Optimization via $\ell_p$-Norm Steepest Descent
PowerStep היא שיטה חדשה לאופטימיזציה אדפטיבית שחוסכת זיכרון. היא מיועדת לאימון מודלים גדולים כמו Transformers. PowerStep משיגה תוצאות תחרותיות עם זיכרון פחות מאשר AdamW.
תקציר מקורי באנגליתarXiv:2605.10335v2 Announce Type: replace-cross Abstract: Adaptive optimizers such as Adam are standard for training Transformers, but storing gradient first and second moments incurs substantial memory overhead. We introduce PowerStep, a memory-efficient optimizer that achieves coordinate-wise adaptivity without storing second-moment statistics. Motivated by $\ell_p$-norm steepest descent, PowerStep applies a signed-power transform directly to one momentum buffer. We establish a finite-horizon stationarity bound for exact, unregularized updates, with an $O(1/\sqrt{T})$ term and a noise-dependent residual. Experiments on Transformers from 124M to 235B parameters show competitive validation quality while halving $\texttt{fp32}$ optimizer-state memory relative to AdamW. Combined with uniform
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית