יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

ביצוע IRL ללא תפנית: חידוש בשיקוף תגמול על ידי חיפוש ערך רצפי

Loop-Free Inverse Reinforcement Learning via Sequential Value Recovery with Q-Score Matching
חידוש בשיקוף תגמול: IRL ללא תפנית, על ידי חיפוש ערך רצפי. המאמר מציג חידוש בשיקוף תגמול, המבוסס על חיפוש ערך רצפי. החידוש נועד לשפר את יעילות הלמידה ולהפחית את התפנית בשיקוף תגמול.
תקציר מקורי באנגליתarXiv:2609.38955v1 Announce Type: new Abstract: Inverse Reinforcement Learning (IRL) aims to recover a reward function that explains expert demonstrations. Existing IRL methods typically rely on a bi-level optimization procedure that alternates between reward learning and policy optimization, leading to substantial computational burden and training instability. In this work, we introduce a different route that eliminates policy optimization entirely by leveraging diffusion policies. Our key insight is that a diffusion policy encodes the action-gradient structure of the optimal soft Q function, enabling reward learning to be cast as a sequence of value recovery problems, thereby allowing us to bypass reward-policy loops inherent in prior IRL methods. Specifically, our method proceeds in thr
קרא במקור המקורי