יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

למידת חיזוק עמידה באמצעות רגולריזציה

A Unified and Constrained View of Regularization-Based Robust Reinforcement Learning
חוקרים פיתחו שיטה חדשה ללמידת חיזוק עמידה באמצעות רגולריזציה. השיטה מאפשרת אימון מדיניות עמידה נגד הפרעות קלט. המחקר מציג תוצאות מבטיחות במשימות בקרה רציפות.
תקציר מקורי באנגליתarXiv:2609.13050v1 Announce Type: new Abstract: Regularization-based methods have become a standard approach for training Deep Reinforcement Learning policies against adversarial input perturbations. In this paper, we unify these methods by deriving new upper bounds on the performance gap between the nominal and worst-case policies. Each upper bound is expressed as an existing regularization objective plus a KL-divergence penalty between the nominal and worst-case policies, which further explains why adding a KL penalty improves robustness in practice. Building on these bounds, we formulate robust training as a constrained optimization problem, showing that existing methods correspond to the special case of a fixed Lagrange multiplier. We instead update the multiplier jointly with the poli
קרא במקור המקורי