כתבה
arXiv cs.LG ·
Budgeted Multi-Source Counterfactual Annotation for Off-Policy Evaluation
תקציר מקורי באנגליתarXiv:2610.10974v1 Announce Type: new Abstract: Off-policy evaluation (OPE) estimates the value of a target policy from logged data, but limited behavior-policy coverage can force high-variance reweighting or reward-model extrapolation. Counterfactual annotations can add evidence about unobserved actions, yet practical sources, including domain experts and large language models (LLMs), may be costly, biased, or noisy. We study budgeted acquisition of such annotations for contextual-bandit OPE. Given source-specific costs and error profiles, we formulate an integer allocation problem over context-action pairs and annotation sources to minimize the component of estimator variance that depends on the annotation plan. We characterize when annotations are valuable through a first-annotation thr
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית