כתבה
arXiv cs.LG ·
מדיניות אחת או מדיניות רב-מוצרים ללמידת החזרה של פאזות?
Single or Multiple Policies for Phase-Structured Reinforcement Learning?
במחקר זה נבחנה האפשרות של שימוש במדיניות אחת או מדיניות רב-מוצרים ללמידת החזרה של פאזות. התוצאות הראו שהשימוש במדיניות רב-מוצרים עשוי להציג יתרון, אך זה תלוי במספר גורמים.
תקציר מקורי באנגליתarXiv:2610.03475v1 Announce Type: new Abstract: Many reinforcement-learning (RL) problems are non-stationary yet structured and can be decomposed into phases, each with its own transition probabilities and reward functions. When the phase sequence is known, the common solution augments the state with information to satisfy the Markovian property and applies standard RL techniques. However, prior work finds that the multi-policy approach for different phases can outperform a single state-augmented policy shared among the phases, for reasons that remain unclear. In this work, we first show that the shared policy can theoretically achieve performance of any multi-policy solution. However, whether a multi-policy solution can perform better than the corresponding single shared policy in practic
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית