יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

למידה מהפער בין Pass@K ל- Pass@1

Learning from the Gap Between Pass@K and Pass@1
במאמר זה, המחברים מציגים טכניקה חדשה לשיפור דגמי שפה גדולים, המבוססת על למידה מהפער בין Pass@K ל- Pass@1. הם מציעים טכניקה חדשה, GapFT, שמטרתה להעביר את היתרון של חיפוש לדגם. הם מדגימים את GapFT על שני דגמי שפה, Llama-3.1-8B ו-Mistral-7B, ומראים כי GapFT משיג תוצאות טובות יותר מאשר טכניקה קיימת, RFT.
תקציר מקורי באנגליתarXiv:2609.35793v2 Announce Type: replace Abstract: Sampling many responses and keeping one that passes a verifier lets large language models solve problems beyond their single-response ability, but this search must be paid again for every query, while many deployments answer with a single response. Post-training on verified responses can transfer the benefit of search into the model. With a fixed budget, selecting by correctness alone spends slots on problems the model already answers correctly, leaving fewer to correct its failures. To address this imbalance, we propose GapFT, which trains on the gap between Pass@K and Pass@1: problems that the source model fails with one response but solves within K samples. GapFT keeps the objective and training budget fixed and changes only which veri
קרא במקור המקורי