כתבה
arXiv cs.LG ·
בלמן אחד מאוחד ללמידת ריפוד בעלת תכונות בטיחותיות
A Unified Bellman Operator for Safety-Critical Reinforcement Learning
אחד המאמרים החדשים בתחום הלמידה והפיתוח של מודלי ריפוד, המציע פתרון ללמידה בעלת תכונות בטיחותיות. המאמר, שפורסם ב-arXiv, מציע פענוח חדש לבעיית הבטיחות בלמידה והפיתוח של מודלי ריפוד.
תקציר מקורי באנגליתarXiv:2610.12420v1 Announce Type: new Abstract: Reinforcement learning in safety-critical domains requires maximizing task performance while strictly adhering to safety constraints. Existing safe reinforcement learning paradigms typically force a trade-off: they either require a priori knowledge to provide strict safety guarantees (e.g., safety filters), or they enable joint learning but only satisfy safety constraints on average. In this work, we propose a novel Bellman operator that unifies performance and safety objectives into a joint value function. We show that temporal difference learning with the joint Bellman operator converges under a two-timescale stochastic approximation framework. On the fast timescale, the safety value of the learning joint policy is estimated, while the join
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית