כתבה
arXiv cs.LG ·
Lifted Bellman Linear Programming for Offline Reinforcement Learning
תקציר מקורי באנגליתarXiv:2609.24489v2 Announce Type: replace Abstract: Offline reinforcement learning (RL) typically trains a critic by minimizing a regression loss against bootstrapped value targets stabilized by target networks with exponential moving average (EMA) updates. Multi-step targets incorporate behavior-policy actions and therefore require off-policy correction. We instead impose in-sample Bellman optimality on the critic through inequality constraints. We formulate the Lifted Bellman Linear Program (LBLP), which lifts the linear programming characterization of Bellman optimality to the joint $(Q,V)$ space so that every constraint involves only state-action pairs in the dataset. Its unique minimizer is the in-sample optimal pair, and constraints along $K$-step segments of dataset trajectories lea
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית