כתבה
arXiv cs.LG ·
אבסטרקציות מרכזיות של החלטות דרך הערכת תכונות-ק-פונקציות-פרקטליות
Decision-Centered Abstractions via Orthogonal Estimation of Difference-of-Q Functions
למידת החלטות רפורמנטית מחוץ לזמן-אמת, עם הערכת תכונות-ק-פונקציות-פרקטליות.
תקציר מקורי באנגליתarXiv:2406.08697v4 Announce Type: replace-cross Abstract: Offline reinforcement learning enables evaluation and optimization of sequential decisions from historical data, when it is not possible to deploy new policies online due to safety, cost, and other concerns. Big data advances enable rich state information, but may naively include reward- and action- irrelevant dynamics that are ultimately unnecessary for learning optimal actions. We introduce state abstractions that target preservation of the difference-of-Q functions, and we propose to learn these abstractions via causal machine learning of the difference-of-Q function and standard statistical sparse learning. Under a nonparametric additive-rewards model, we characterize when decision-centered abstractions are simpler than the full
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית