כתבה
arXiv cs.LG ·
למידת תהליך חסכוני למשתנים סופר-אינפיניטיים
Learning Infinite-Horizon Average-Reward CMDPs via State Augmentation
במאמר זה, המחברים מציגים תהליך למידה חסכוני למשתנים סופר-אינפיניטיים. התהליך משתמש בשיטת תגבור ומציע רגרט ופגיעה בקונסטריינט עם סדר גודל √(T) עם הסברה גבוהה.
תקציר מקורי באנגליתarXiv:2609.39093v1 Announce Type: new Abstract: We study infinite-horizon average-reward constrained Markov decision processes (CMDPs) under the weakly communicating assumption. Existing high-probability guarantees for this setting either require computationally inefficient algorithms or have suboptimal dependence on the number of interactions $T$. We propose, to the best of our knowledge, the first computationally efficient algorithm that achieves $\widetilde{\mathcal{O}}(\sqrt{T})$ regret and cumulative constraint violation with high probability in the tabular setting. The $\sqrt{T}$ dependence is optimal up to logarithmic factors. Our approach incorporates cumulative constraint violation into the state and defines a reshaped reward through differences of a Huber potential. The added sta
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית