יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

למידת MDPs עם אילוצי הימור

Learning Chance-Constrained MDPs with Bellman Distributional Certificates
חוקרים פיתחו שיטה חדשה ללמידת MDPs עם אילוצי הימור, המאפשרת שליטה טובה יותר על הסיכון. השיטה משתמשת בתעודת Bellman היפרדותית, המאפשרת לבנות נסיגה של Bellman להסתברות הפרת אילוצים לפני בחירת מדיניות.
תקציר מקורי באנגליתarXiv:2609.30856v1 Announce Type: new Abstract: Safe reinforcement learning (RL) commonly enforces expected-cost constraints, but such expectation safety may fail to control the probability of rare high-cost trajectories. Chance-constrained MDPs (CCMDPs) impose a stronger probability-level requirement, but are widely viewed as harder because the chance constraint is nonconvex and depends on the full trajectory rather than a Bellman-linear expectation. In this paper, we reveal that this computational difficulty does not necessarily imply a higher statistical price. For tabular discounted CCMDPs with fixed bounded successor support and access to a certified planning oracle, we establish a model-based upper bound, with a matching lower bound up to logarithmic terms. Technically, our key idea
קרא במקור המקורי