יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

גבולות PAC ללמידה ב-MDPs קונטקסטואליים

Minimax PAC Bounds for Learning in Exogenous Contextual MDPs
המאמר עוסק בגבולות PAC ללמידה ב-MDPs קונטקסטואליים. החוקרים הציגו פרקטיקה PAC שבה הלמד הופך לגישה לאורקלים סיימפלינג גם לפני וגם בזמן ההחלטה. הם חקרו את פוליסי אבלוג (PE), אבלוג של הערך הטוב ביותר (BVE) ואבלוג של הפוליסי הטובה ביותר (BPE).
תקציר מקורי באנגליתarXiv:2606.25170v2 Announce Type: replace-cross Abstract: We introduce a PAC framework in which the learner can access sampling oracles both before and at decision time. Sample complexity is measured by a pair $(n,m)$, where $n$ is the learning budget spent before a query is known and $m$ is the additional sampling budget per query. We demonstrate its relevance in discounted Markov decision processes with exogenous i.i.d.\ contexts revealed before acting. Contexts may affect both rewards and transitions but remain uncontrolled by the agent. The learner can sample the unknown context distribution and the transition kernel. We study policy evaluation (PE), best-value estimation (BVE), and best-policy extraction (BPE). When rewards and transitions are known, a variance-reduced algorithm solve
קרא במקור המקורי