כתבה
arXiv cs.LG ·
Provable Benefits of Regularization: Fast Rates for Adversarial Imitation Learning
תקציר מקורי באנגליתarXiv:2609.35698v2 Announce Type: replace Abstract: We study adversarial imitation learning (AIL), in which an agent learns to imitate expert demonstrations by optimizing a policy against an adversarial reward that distinguishes expert and learner behavior. Historically, reward regularization and entropy-based policy regularization are key components of empirically successful methods such as GAIL and LS-IQ, yet their finite-sample benefits remain underexplored. We establish fast rates for jointly regularized AIL in finite-horizon Markov decision processes with general function approximation. Our model-free algorithm, Dually Regularized AIL, combines KL policy regularization with a quadratic reward penalty weighted by expert and learner occupancies. With K online episodes and N expert traje
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית