כתבה
arXiv cs.AI ·
Temporal Self-Imitation Learning
תקציר מקורי באנגליתarXiv:2606.19752v3 Announce Type: replace-cross Abstract: Long-horizon policies trained with reinforcement learning can still achieve high return through inefficient interactions, while rare efficient behaviors discovered during training may be forgotten. We argue that temporal efficiency itself provides a source of self-supervision for reinforcement learning. We introduce Temporal Self-Imitation Learning (TSIL), a reinforcement learning framework that mines temporally efficient successful trajectories generated during learning and converts them into reusable supervision for future policy improvement. TSIL progressively refines learning using configuration-conditioned adaptive temporal targets derived from fast successful trajectories, while preserving and replaying efficient behaviors thr
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית