יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

למידה לפתור שליטה סטוכסטית עם דריפטים ותגמולים לא ידועים

Learning to Solve Stochastic Controls with Unknown Drifts and Running Rewards: Theory, Algorithms and Convergence
אנו מחקרים שליטה סטוכסטית עם דריפטים ותגמולים לא ידועים. אנו מפתחים אלגוריתמים RL יעילים ומדויקים ללמידה של פונקציות ערך אופטימליות ופוליציות שליטה.
תקציר מקורי באנגליתarXiv:2609.14972v1 Announce Type: new Abstract: We study continuous-time and possibly high-dimensional stochastic control problems where drift coefficients and running reward functions are unknown. Due to these missing model primitives, we take the exploratory, reinforcement learning (RL) framework of Wang, Zariphopoulou, and Zhou(2020) with relaxed controls and entropy regularization. The objective is to develop theoretically grounded, efficient and scalable RL algorithms to learn both the optimal value functions (which also solve the exploratory HJB equation) and optimal exploratory feedback control policies. When the diffusion coefficients do not contain control, we employ probabilistic representations of both the optimal value function and its gradient based on an auxiliary state proce
קרא במקור המקורי