כתבה
arXiv cs.AI ·
מדיניות אחת או מדיניות רב-מוצרים ללמידת רפורמציה מבוססת פאזות?
Single or Multiple Policies for Phase-Structured Reinforcement Learning?
במאמר זה, נבחן את האפקטיביות של מדיניות אחת מול מדיניות רב-מוצרים בלמידת רפורמציה מבוססת פאזות. נמצא כי פעמים רבות, מדיניות רב-מוצרים יכולה להציג תוצאות טובות יותר, אך זה תלוי במספר גורמים.
תקציר מקורי באנגליתarXiv:2610.03475v1 Announce Type: cross Abstract: Many reinforcement-learning (RL) problems are non-stationary yet structured and can be decomposed into phases, each with its own transition probabilities and reward functions. When the phase sequence is known, the common solution augments the state with information to satisfy the Markovian property and applies standard RL techniques. However, prior work finds that the multi-policy approach for different phases can outperform a single state-augmented policy shared among the phases, for reasons that remain unclear. In this work, we first show that the shared policy can theoretically achieve performance of any multi-policy solution. However, whether a multi-policy solution can perform better than the corresponding single shared policy in pract
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית