כתבה
arXiv cs.LG ·
From Trajectories to Instructions: Language-Conditioned Meta-Reinforcement Learning
תקציר מקורי באנגליתarXiv:2607.18830v1 Announce Type: new Abstract: Model-Agnostic Meta-Learning (MAML) is a widely used framework for reinforcement learning (RL) that enables efficient transfer by learning global policy parameters that can be rapidly adapted to new tasks. MAML training proceeds in two loops: an inner loop where the global parameters are adapted to task-specific parameters, and an outer loop where these task-specific parameters are evaluated and losses are back-propagated to improve the global parameters. Traditionally, the inner loop adaptation is performed by collecting trajectories from the task environment and applying gradient updates on the empirical expected return, which can be a costly operation. We note that it is the outer loop that drives the actual learning of global parameters,
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית