כתבה
arXiv cs.AI ·
ללמוד POMDPs מעבר לפעולות מלאות-דרגה ואובסרבביליות של מצב
Toward Learning POMDPs Beyond Full-Rank Actions and State Observability
במאמר זה, נראה כיצד ניתן ללמוד POMDPs (Partially Observable Markov Decision Processes) מעבר לפעולות מלאות-דרגה ואובסרבביליות של מצב. נראה כיצד ניתן ללמוד דגמי POMDPs על ידי שימוש במודלי תצפיות ומודלי תגובה. המאמר כולל תיאור של המודלים והכלים שהוצגו במאמר.
תקציר מקורי באנגליתarXiv:2601.18930v5 Announce Type: replace-cross Abstract: We are interested in enabling autonomous agents to learn and reason about systems with hidden states, such as locking mechanisms. We cast this problem as learning the parameters of a discrete Partially Observable Markov Decision Process (POMDP). The agent begins with knowledge of the POMDP's actions and observation spaces, but not its state space, transitions, or observation models. These properties must be constructed from a sequence of actions and observations. Spectral approaches to learning models of partially observable domains, such as Predictive State Representations (PSRs), learn representations of state that are sufficient to predict future outcomes. PSR models, however, do not have explicit transition and observation syste
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית