יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

הסברה מאומתת של רונדינג למדיניות דקתית ב-MDPs

Bellman-Certified Rounding for Sparse Policy Deployment in MDPs
אנו מחקרים את השפעת סברה מאומתת של רונדינג על פריסת מדיניות דקתית ב-MDPs.
תקציר מקורי באנגליתarXiv:2610.00325v1 Announce Type: new Abstract: Continuous policy optimization may spread an update across many states, even when deployment permits only a few complete state-level changes. We study how much discounted return can be retained when continuous row mixtures are rounded to sparse binary policies in finite MDPs. Policy-dependent visitation couples the row edits, while long horizons make global curvature bounds conservative. From $2d+2$ Bellman solves, we derive reusable envelopes that support uniform and candidate-specific guarantees before rounding. A rank-two rational representation of each exchange further permits weighted curvature integration along the realized trajectory. We prove that linear dimension dependence is unavoidable when the budget scales, and that exact global
קרא במקור המקורי