כתבה
arXiv cs.LG ·
Exploiting Exogenous Structure for Sample-Efficient Reinforcement Learning
תקציר מקורי באנגליתarXiv:2409.14557v4 Announce Type: replace-cross Abstract: We study a structured class of Markov Decision Processes, known as Exo-MDPs, in which the state space is partitioned into exogenous and endogenous components. Exogenous states evolve stochastically, independent of the agent's actions, while endogenous states evolve deterministically based on both state components and actions. Exo-MDPs capture many operations research settings, including inventory control, resource management, and ride-sharing. Our first contribution is structural: we establish a representational equivalence between discrete MDPs, Exo-MDPs, and discrete linear mixture MDPs. Our second contribution is statistical. We characterize the minimax regret of learning in Exo-MDPs when the effective dimension r is small relati
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית