כתבה
arXiv cs.LG ·
מה שpass@k לא יכול למדוד: הערכה של תפוצה ושימור יכולת אחרי טירונים
What pass@k Cannot Measure: Evaluating Diversity and Capability Retention after Post-Training
במאמר זה, המחברים חוקרים את יעילות פרוטוקול pass@k, שמשמש להערכת יכולת של מודלי RL אחרי טירונים. הם מציגים תוצאות של טירונים של מודל Qwen2.5-1.5B-Instruct, שמצביעות על חולשות של pass@k. המחברים טוענים שפרוטוקול זה חסר בתכונה חשובה, וכי יש לבחון את התוצאות באופן שונה.
תקציר מקורי באנגליתarXiv:2610.07405v1 Announce Type: new Abstract: pass@$k$, the fraction of problems a model solves within $k$ sampled attempts, is the field's default protocol for deciding whether reinforcement-learning (RL) post-training on verifiable rewards improved a model. At the population level, pass@$k$ depends only on a problem's probability of a correct sample, with no term for how it is distributed across outputs. We show this gap is not academic. Training Qwen2.5-1.5B-Instruct on grade-school math with Group Relative Policy Optimization (GRPO) and with rejection-sampling fine-tuning (RFT, training on the model's own shortest verifier-passed rollout) moves three complementary diversity measures (token-level entropy, answer-level entropy, unique answers per prompt) in opposite directions, with ze
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית