כתבה
arXiv cs.LG ·
עיצוב פרס רווח אופטימלי: מקרה מבחן חניה אוטונומית
Optimal Reward Shaping: Autonomous Car Parking Case Study
חוקרים פיתחו שיטה לעיצוב פרס רווח אופטימלי לחניה אוטונומית. השיטה משלבת מידע על הסביבה והתנאים, ומאפשרת לסוכנים ללמוד ולהתאים את התנהגותם בצורה יעילה יותר.
תקציר מקורי באנגליתarXiv:2607.23617v1 Announce Type: new Abstract: Designing effective reward functions for model-free reinforcement learning under non-holonomic constraints remains a persistent challenge, often resulting in severe local minima such as policy paralysis or over-conservative hazard avoidance. In this work, we present a parameterized reward shaping framework featuring coverage-gated alignment feedback, drive-direction switch regularization, and an aligned episode termination mechanism evaluated on an autonomous parallel parking task. Crucially, we show that environmental reward parameters and algorithmic hyperparameters are deeply co-dependent, requiring joint meta-optimization to achieve stable convergence. By employing surrogate-based Bayesian optimization, our co-optimized Deep Q-Network (DQ
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית