יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

למידת חיזוק עם מדיניות דטרמיניסטית

Boundary-aware Reinforcement Learning for Hypercube State Spaces via Deterministic Policy Gradient
חוקרים פיתחו שיטת למידת חיזוק עם מדיניות דטרמיניסטית למרחבי מצב היפרקובייה. השיטה משתמשת במשוואת בלמן עם תנאי שפה נוימן. הניסויים הראו שיפור ביציבות הלמידה.
תקציר מקורי באנגליתarXiv:2610.09712v1 Announce Type: cross Abstract: We develop a continuous-time deterministic policy gradient framework for reinforcement learning with reflected state dynamics, where the state process is governed by a controlled reflected stochastic differential equation on a hypercube. Under suitable regularity assumptions, we establish the connection between the value function and the Neumann Bellman equation, introduce an advantage-rate function that yields a deterministic policy gradient formula, and prove the martingale characterization theorem. Motivated by these theoretical results, we propose a continuous-time deep deterministic policy gradient algorithm for reflected stochastic systems, in which the Neumann boundary condition is imposed via either soft penalization or hard archite
קרא במקור המקורי