כתבה
arXiv cs.AI ·
תגמול נכון אינו מספיק
When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game
חוקרים בדקו את יכולתו של אלגוריתם PPO לנצל תוצאות מתמטיות במשחק ברוקר-סוחר. הם מצאו ש-PPO אינו מדויק עם זרם הזמנות לא מודע, אך הצליח לשפר את התוצאות כאשר הוא התבסס על מדיניות אנליטית.
תקציר מקורי באנגליתarXiv:2610.03598v1 Announce Type: cross Abstract: Reinforcement learning (RL) is increasingly used for financial optimal-control problems when complex dynamics make analytical strategies difficult to obtain. There are financial mathematics literactures which provides many solved models whose equations and controls could evaluate and guide learning; we ask whether RL can exploit these results. We place a proximal policy optimisation (PPO) agent in an analytically solved continuous-time broker--trader game. PPO replaces the broker and chooses its trading speed while interacting with an informed trader and stochastic uninformed order flow. We derive a finite-step reward from the broker's continuous-time payoff and verify its discrete implementation through grid refinement and an exact one-ste
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית