כתבה
arXiv cs.LG ·
יתרונות מוכחים של רגולריזציה
Provable Benefits of Regularization: Fast Rates for Adversarial Imitation Learning
חוקרים רגולריזציה בלמידת חיקוי אדברסרית. התוצאות מראות יתרונות מוכחים בקצבי למידה מהירים. האלגוריתם Dually Regularized AIL משלב רגולריזציה של מדיניות ופרס.
תקציר מקורי באנגליתarXiv:2609.35698v3 Announce Type: replace Abstract: We study adversarial imitation learning (AIL), in which an agent learns to imitate expert demonstrations by optimizing a policy against an adversarial reward that distinguishes expert and learner behavior. Historically, reward regularization and entropy-based policy regularization are key components of empirically successful methods such as GAIL and LS-IQ, yet their finite-sample benefits remain underexplored. We establish fast rates for jointly regularized AIL in finite-horizon Markov decision processes with general function approximation. Our model-free algorithm, Dually Regularized AIL, combines KL policy regularization with a quadratic reward penalty weighted by expert and learner occupancies. With K online episodes and N expert traje
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית