כתבה
arXiv cs.LG ·
Variance-Optimal Off-Policy Evaluation with Conjunct Effect Modeling
תקציר מקורי באנגליתarXiv:2610.08677v1 Announce Type: new Abstract: Off-policy evaluation (OPE) for contextual bandit policies becomes challenging when action-level importance weighting incurs excessive variance. Doubly robust (DR) estimation remains unbiased under common support but retains these high-variance action-level weights. A prior estimator, Off-policy evaluation with Conjunct Effect Model (OffCEM), replaces them with more stable cluster-level weights, at the cost of relying on local correctness of the reward model. In this paper, we show that, under the assumptions required by DR and OffCEM, there exists an unbiased family of estimators that interpolates between OffCEM and DR. Building on this result, we propose the Variance Optimal-CEM (VOCEM) estimator, which selects the interpolation coefficient
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית