כתבה
arXiv cs.LG ·
Decision-Focused Learning in MDPs: An Occupancy Measure Approach
תקציר מקורי באנגליתarXiv:2610.08384v1 Announce Type: new Abstract: In this work, we consider decision-focused learning (DFL) for a Markov decision process (MDP), where existing methods differentiate through the KKT conditions of the Bellman equation and require solving a linear system over all state-action pairs, limiting its scalability. We address this by reformulating the MDP as an occupancy measure-based linear program (LP), whose feasible region is induced by predicted dynamics, and we derive a closed-form gradient by identifying the active constraints in the feasible polyhedron via the pivoting algorithm. This occupancy measure-based LP layer raises two challenges: (1) LP's solution gradient is discontinuous when active constraints change, and (2) the LP backward cost still scales with the state size,
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית