כתבה
arXiv cs.LG ·
למידת פוליצי עם תגובות חלשות
Policy Learning with Weak Signals
למידת פוליצי עם תגובות חלשות אינה נלמדת באופן כללי. עם זאת, כאשר תגובות הטיפול משתנות באופן רציף, ניתן לפתח מדיניות-אדפטיבית-מינימקס שמגיעה לאבדן-רווחת-סלבסטריק.
תקציר מקורי באנגליתarXiv:2610.10167v1 Announce Type: cross Abstract: Policy learning in digital experimentation faces three challenges: weak signal-to-noise ratios, rich covariate spaces, and massive data volumes. We formalize this regime by modeling treatment-effect estimates from increasingly fine covariate partitions as Gaussian observations with bounded signal-to-noise ratios. We establish that, in general, the optimal treatment policy is not learnable in this setting. Even learning the optimal policy value suffers from impractically slow rates. However, when treatment effects vary smoothly, we derive minimax-adaptive policies based on linear smoothers that achieve vanishing welfare regret. We demonstrate the practical value of our framework by applying it to large-scale real-world experiments at Netflix
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית