כתבה
arXiv cs.LG ·
Learning High-Risk High-Precision Motion Control
תקציר מקורי באנגליתarXiv:2609.34851v2 Announce Type: replace Abstract: Deep reinforcement learning (DRL) algorithms for movement control are typically evaluated and benchmarked on sequential decision tasks where imprecise actions may be corrected with later actions, thus allowing high returns with noisy actions. In contrast, we focus on an under-researched class of high-risk, high-precision motion control problems where actions carry irreversible outcomes, driving sharp peaks and ridges to plague the state-action reward landscape. Using computational pool as a representative example of such problems, we propose and evaluate State-Conditioned Shooting (SCOOT), a novel DRL algorithm that builds on advantage-weighted regression (AWR) with three key modifications: 1) Performing policy optimization only using eli
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית