יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

למידה על-פוליצי לפתרון אופטימיזציה לינארית קונטקסטואלית עם פרסום חלקי

Decision-Focused On-Policy Learning for Contextual Linear Optimization with Partial Feedback
אופטימיזציה קונטקסטואלית לינארית עם פרסום חלקי: למידה על-פוליצי עם פיתוח חדש של אלגוריתם לפתרון אופטימיזציה.
תקציר מקורי באנגליתarXiv:2606.01081v2 Announce Type: replace Abstract: Decision-focused learning (DFL) trains predictive models by optimizing downstream decision quality rather than standalone prediction accuracy. For contextual linear optimization, most existing DFL methods assume offline data and full observations of the objective cost vector. We develop an on-policy learning method for sequential contextual linear optimization under partial feedback, generalizing the standard bandit feedback setting. Our method learns a stochastic predict-then-optimize policy that samples a cost-vector prediction from a conditional distribution and solves the resulting downstream linear optimization problem. To update this distributional model, we introduce a two-component hybrid gradient estimator. The first component is
קרא במקור המקורי