כתבה
arXiv cs.CL ·
אל תשבית: טכניקת תלוי-עצמית להטמעה עצמית לכיסוי
Don't Repeat Yourself: Self-Supervised Fine-Tuning for Coverage
טכניקת הטמעה עצמית שמגדילה את רמת הכיסוי והשונות בתוצאות המודל. המאמר עוסק בפיתוח טכניקה שמשפרת את יכולת המודל למצוא פתרונות שונים לבעיות, ומציע פתרון לבעיה של תפיסה חדשה.
תקציר מקורי באנגליתarXiv:2609.31688v2 Announce Type: replace Abstract: In verifiable domains such as math and coding, finding one correct solution among many attempts can matter more than the pass rate of each attempt. Post-training can concentrate large language model outputs around a few modes, while increasing sampling temperature has limited effectiveness. We introduce Don't Repeat Yourself Supervised Fine-Tuning (DRY-SFT), a post-training method that increases output diversity and coverage: the probability of at least one correct solution among many attempts. DRY-SFT has two stages. First, for each problem, sequentially generate K solutions, showing the model all prior attempts and asking for a different solution. Second, fine-tune on each attempt independently, removing prior attempts from the context.
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית