יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

משפט גרדיאנט של פרמטרי האקראיות של הסביבה לצורך תכנון משותף בלמידת מכונה

Environment Parameter Gradient Theorem for Co-Design in Reinforcement Learning
במאמר זה, המחברים מפתחים משפט גרדיאנט של פרמטרי האקראיות של הסביבה, המאפשר תכנון משותף של פוליציה ופרמטרי הסביבה בלמידת מכונה.
תקציר מקורי באנגליתarXiv:2607.12590v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) is traditionally concerned with learning a control policy for a fixed environment. In many engineering systems, however, the environment itself is alterable, i.e., physical or operational parameters can be tuned to shape the system's transition dynamics and costs experienced by the RL agent. This motivates jointly optimizing both the policy and the environment design parameters. To this end, we establish an Environment Parameter Gradient Theorem --- a formal expression for the gradient of the RL's objective function with respect to environment parameters. The key theoretical device is a generalized action-value function $Q_{\pi,\xi}(s,a,\zeta)$, which comprises two copies of the environment parameters: $\
קרא במקור המקורי