יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

לשיפור אופטימלי של פוליציה

Towards Optimal Policy Improvement
אלגוריתמי RL יעילים לומדים לפתור מערכות תגובה-החלטה דרך שיפור פוליציה מחזורי. המאמר חוקר שיפור פוליציה מראשית, ומציג פעולן חדש שמשפר את הביצועים של חלוקי פוליציה.
תקציר מקורי באנגליתarXiv:2610.01566v1 Announce Type: new Abstract: Practical Reinforcement Learning (RL) algorithms learn to solve Markov Decision Processes (MDPs) through iterative policy improvement in the presence of approximate evaluation. We study policy improvement from first principles, defining optimal policy improvement as producing the best policy attainable in a single update under specified constraints. We show that optimal improvement restricted to a set of states is equivalent to solving an induced MDP, characterizing planning with an explicit or implicit model as a path towards optimal policy improvement. Because practical methods commonly solve such induced problems through iterative improvement in the form of greedification, we take steps towards optimal greedification under the central prac
קרא במקור המקורי