כתבה
arXiv cs.LG ·
מדיניות חזקה וקצרה לתהליכי קבלת החלטים תת-מודולריים ב-MDPs דרך LP-Based Submodular Orienteering
Strong and Compact Policies for Submodular Markov Decision Processes via LP-Based Submodular Orienteering
המאמר מציג אלגוריתם LP-Based Submodular Orienteering לתהליכי קבלת החלטים תת-מודולריים ב-MDPs, המספק תוצאות טובות יותר מאלגוריתמים קיימים. האלגוריתם משתמש ברעיונות מה-Sherali-Adams hierarchy ו-Round-or-Cut. המאמר גם מציג תרחישים חדשים לאלגוריתם, כולל תרחישים עם גבולות זמן.
תקציר מקורי באנגליתarXiv:2609.15539v1 Announce Type: cross Abstract: Finding policies for Markov Decision Processes (MDPs) is a central problem in areas such as Reinforcement Learning and Operations Research. Here, we have to repeatedly choose an action that should be performed by an agent. Depending on the action and the current state of the agent, the agent collects a reward and randomly transitions into a new state. The goal is to maximize the reward in expectation over a finite time horizon of length $H$. We consider a recently introduced variant that generalizes the traditionally additive reward function in the model to a monotone submodular one, which allows for capturing a range of interesting applications. Without the stochastic component, this problem is equivalent to the Submodular Orienteering pro
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית