יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

רכישת נתונים אופטימלית ללמידת חיזוק

Optimal Data Acquisition for Reinforcement Learning: A Large Deviations Perspective
פותחת מסגרת חדשה לרכישת נתונים בלמידת חיזוק, המשלבת תורת הסטיות הגדולות. הגישה מאפשרת רכישת נתונים יעילה יותר ומשפרת את ביצועי האלגוריתמים.
תקציר מקורי באנגליתarXiv:2605.28675v2 Announce Type: replace Abstract: Data acquisition efficiency is a central challenge in deploying reinforcement learning in business and healthcare operations, where interactions are costly, slow, and often involve humans in the loop. This paper develops a unified large deviations framework for data acquisition in infinite-horizon reinforcement learning. We introduce the exponential decay rate of the policy-selection error probability as a principled efficiency metric and derive a variational characterization of this rate via large deviations theory for Markov chains, yielding a nested optimization problem. Based on this characterization, we formalize two complementary notions of optimality in terms of the optimal solution of the nested problem. Because the resulting prog
קרא במקור המקורי