כתבה
arXiv cs.LG ·
התכנסות SGD בהסתברות גבוהה
High-Probability Convergence of SGD via Batched Updates
שיטת Batched SGD מאפשרת התכנסות בהסתברות גבוהה עבור אופטימיזציה בקנה מידה גדול. השיטה מחלקת דגימות לתקופות ומבצעת עדכון בודד לכל תקופה, תוך שימוש בהערכת גרדיאנט משופרת. הגישה מאפשרת ניתוח פשוט ויעיל עבור אובייקטיבים קמורים ולא קמורים.
תקציר מקורי באנגליתarXiv:2609.12765v1 Announce Type: cross Abstract: Stochastic gradient descent (SGD) is the primary workhorse for large-scale optimization. While the average behavior of its iterates, typically characterized by mean-squared error bounds, is well-understood, obtaining high-probability guarantees for the last iterate remains challenging. Prior approaches to this problem have either imposed restrictive assumptions (such as bounded domains or gradients) or relied on complex proofs involving auxiliary sequences. In this work, we propose Batched SGD, a simple variant that partitions online samples into epochs and performs a single update per epoch using a refined, low-variance gradient estimate. Our main contribution demonstrates that this batching mechanism enables a surprisingly simple high-pro
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית