יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

אבחון יעיל של פוליציות Best-of-N להתאמה בזמן הריצה

Efficient Best-of-N policy evaluation for inference-time alignment
במאמר זה, המחברים מציגים פרוטוקול יעיל לאבחון ובחירת פוליציות Best-of-N להתאמה בזמן הריצה. הם מציעים אתונטרי של BoN-DR, שמשתמש באורדר-סטטיסטיקה של BoN כדי לבצע את האבחון. הם גם מציעים שני חוקי בחירה: (i) להגדיל את הערך המוערך של הפוליציה ו(2) להגדיל את הגבול הנמוך בביטחון על השיפור על פני הפוליציה המקורית. המחברים מדגימים את פרוטוקולם באמצעות ניסויים סינתטיים ו-GSM8K עם מודלים שונים.
תקציר מקורי באנגליתarXiv:2610.09250v1 Announce Type: new Abstract: Best-of-N (BoN) is a common inference-time alignment method that selects the highest-scoring response among N samples from a reference model. Evaluating BoN policies from logged data is challenging under sample-only access because standard off-policy estimators require density ratios that depend on unavailable response likelihoods. In this paper, we propose a sample-only framework for evaluating and selecting BoN policies without access to these likelihoods. We show that the order-statistic structure of BoN allows the required density ratios to be expressed through score-rank probabilities that are estimable from samples alone. We then develop a doubly robust estimator of the BoN policy value (BoN-DR) that efficiently reuses a shared auxiliar
קרא במקור המקורי