כתבה
arXiv cs.AI ·
ביצוע IRL ללא תלות בפוליציה: פיתוח רכיב Q-סקור
Loop-Free Inverse Reinforcement Learning via Sequential Value Recovery with Q-Score Matching
מחקר חדש מציע שיטה לביצוע IRL ללא תלות בפוליציה, על ידי שימוש ברכיב Q-סקור. השיטה, LFIRL, פועלת באופן פשוט, ללא תלות בפוליציה, ומשפרת את היעילות של הלמידה. המחקר מציע שיטה חדשה לביצוע IRL, על ידי שימוש ברכיב Q-סקור.
תקציר מקורי באנגליתarXiv:2609.38955v1 Announce Type: cross Abstract: Inverse Reinforcement Learning (IRL) aims to recover a reward function that explains expert demonstrations. Existing IRL methods typically rely on a bi-level optimization procedure that alternates between reward learning and policy optimization, leading to substantial computational burden and training instability. In this work, we introduce a different route that eliminates policy optimization entirely by leveraging diffusion policies. Our key insight is that a diffusion policy encodes the action-gradient structure of the optimal soft Q function, enabling reward learning to be cast as a sequence of value recovery problems, thereby allowing us to bypass reward-policy loops inherent in prior IRL methods. Specifically, our method proceeds in t
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית