יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

ניתוח סטטיסטי של למידת חיזוק הפוכה

Statistical analysis of Inverse Entropy-regularized Reinforcement Learning
חוקרים פיתחו גישה סטטיסטית ללמידת חיזוק הפוכה עם רגולריזציה של אנטרופיה. הגישה משלבת רגולריזציה של אנטרופיה עם שחזור ליניארי של פונקציית התגמול. התוצאות מראות קונברגנציה אופטימלית לפונקציית התגמול.
תקציר מקורי באנגליתarXiv:2512.06956v2 Announce Type: replace-cross Abstract: Inverse reinforcement learning aims to infer the reward function that explains expert behavior observed through trajectories of state--action pairs. A long-standing difficulty in classical IRL is the non-uniqueness of the recovered reward: many reward functions can induce the same optimal policy, rendering the inverse problem ill-posed. In this paper, we develop a statistical framework for Inverse Entropy-regularized Reinforcement Learning that resolves this ambiguity by combining entropy regularization with a least-squares reconstruction of the reward from the soft Bellman residual. This combination yields a unique and well-defined so-called least-squares reward consistent with the expert policy. We model the expert demonstrations
קרא במקור המקורי