כתבה
arXiv cs.LG ·
Looped Actor: Depth-Recurrent Reasoning Models for Reinforcement Learning
תקציר מקורי באנגליתarXiv:2609.37432v1 Announce Type: new Abstract: Looped reasoning models repeatedly apply a shared set of parameters, enabling more computation without increasing the model size. These models also support input-dependent computation by dynamically deciding when to stop looping. Motivated by the recent success of looped transformers in language modeling and reasoning, we investigate whether dynamic looping can similarly benefit sequential decision-making. We provide a complexity-theoretic motivation for this approach by showing that there exist Markov decision processes in which a state-adaptive policy achieves the optimal return with asymptotically less expected computation than any optimal fixed-runtime policy. To learn compute-adaptive policies in practice, we introduce Looped Actor, a tr
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית