יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

למידת חיזוק לאיטום גרדיאנט היברידי

Reinforcement Learning to Accelerate Primal-Dual Hybrid Gradient for Linear Programming
GALLOP משתמש בלמידת חיזוק לאיטום גרדיאנט היברידי עבור תוכנית ליניארית. השיטה מאפשרת למידה של פרמטרים רציפים וקבלת החלטות אוטומטיות. GALLOP הוכחה כיעילה במגוון תרחישים.
תקציר מקורי באנגליתarXiv:2610.01546v1 Announce Type: cross Abstract: Primal-dual hybrid gradient (PDHG) methods solve large-scale linear programs (LPs) using GPU-friendly matrix-vector products and projections, but their practical performance depends on coordinating algorithm parameters, acceleration, and restarts. We introduce GALLOP, which uses reinforcement learning to jointly learn continuous algorithm parameters and discrete restart decisions without differentiating through the solver. Its generalized accelerated PDHG update combines separate primal and dual extrapolation, history corrections, and restart anchoring with independently adjustable coefficients. We train a dimension-agnostic feedback policy using a groupwise proximal policy optimization objective that clips likelihood ratios separately for
קרא במקור המקורי