כתבה
arXiv cs.LG ·
Amortized Off-Policy Evaluation for LLMs
תקציר מקורי באנגליתarXiv:2610.10848v1 Announce Type: new Abstract: Accurate evaluation is central to selecting which LLM to deploy, yet testing a candidate on live traffic exposes real users to an unvetted model. Teams therefore evaluate candidates offline, on data produced by already-deployed models. This is off-policy evaluation (OPE), and it faces two distribution shifts: as a model is updated in post-training, its responses diverge from the logged ones (policy shift), and the reward definition under which it is judged changes with business requirements (reward shift). Classical OPE methods are ill-suited to this continual-deployment setting because they are defined per task and require fitting from scratch on every new logged dataset or reward definition. To address this, we propose PFN-OPE, a prior-data
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית