כתבה
arXiv cs.AI ·
טריינינג קדימה, דיסטילציה אחורה: הפעלת עצמית-דיסטילציה על-מדיני ללמידת מודלי שפה גדולים
Train Ahead, Distill Back: Bootstrapping On-Policy Self-Distillation for Large Language Models
במאמר זה, נראה כיצד ניתן לשפר מודלי שפה גדולים באמצעות דיסטילציה עצמית-על-מדיני. המאמר מציג טכניקה חדשה המכונה B-OPSD, המאפשרת למודל לטריינינג קדימה, להשיג טורים-פרטים יותר נאמנים, ולאחר מכן לשים דגש על יעדים יותר מידעיים. התוצאות המוצגות במאמר תומכות בטענה זו, ומראות שהטכניקה החדשה יעילה יותר מהטכניקה הקיימת.
תקציר מקורי באנגליתarXiv:2609.37132v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) improves large language models by letting a self-teacher with privileged information provide dense token-level supervision on the model's own trajectories. Yet existing methods typically construct the self-teacher from the current, initial, or slowly averaged policy state, leaving the quality of supervision constrained by the teacher's ability to exploit privileged information. We ask whether the model's own optimization progress can instead be recycled into a stronger self-teacher. In this paper, we introduce Bootstrapped On-Policy Self-Distillation (B-OPSD), which temporarily trains the policy ahead to obtain a future teacher, restores the student to the original policy state, and then uses the future teac
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית