כתבה
arXiv cs.CL ·
הערכה לא מוטית: Holdout Best-of-N
Holdout Best-of-N: Unbiased Evaluation and Its Cost
חוקרים שיטה להערכה לא מוטית של Best-of-N, המשמשת לבחירת הטוב ביותר מבין מועמדים. השיטה מבוססת על מטריצה קבועה של ציונים עצמאיים. המחקר מראה כי השיטה מספקת הערכה לא מוטית תחת תנאים מסוימים.
תקציר מקורי באנגליתarXiv:2610.08719v1 Announce Type: new Abstract: Reusing the scores that select a Best-of-$N$ winner can overstate its expected reward. We study evaluation from a fixed matrix of $K$ independent scores per candidate for a policy that selects using $J$ fresh scores. A single estimator based only on this matrix is exactly unbiased for expected judge reward under every independent, stable collection of candidate-specific score laws if and only if $J<K$, for every pool size $M\ge N\ge2$. At $J=K-1$, the selector deepens as $K$ grows. For independent Gaussian scores with common variance and fixed $M\ge N\ge2$, the unbiased minimax risk in this regime is of order $\sigma^2/\sqrt K$, attained by Holdout; allowing bias improves the rate to $\sigma^2/K$. For two candidates, we derive the minimum-var
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית