כתבה
arXiv cs.CL ·
שיפור עצמי רקורסיבי באמצעות תרגום על-מדרגה להסתגלות למחשב
Recursive Self-Improvement via On-Policy Distillation for Reasoning
במאמר זה, המחברים מציגים שיטה חדשה לשיפור עצמי רקורסיבי, המבוססת על תרגום על-מדרגה. השיטה, הנקראת DCE+SRCL, מאפשרת למודל ללמוד מסגנונו הסגולי ולשפר את עצמו. המחברים מציגים תוצאות של ניסויים שהראו שהשיטה החדשה מבצעת טוב יותר מאשר שיטות קודמות.
תקציר מקורי באנגליתarXiv:2609.30652v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student model by having it generate trajectories, then matching its next-token predictions with an external teacher's next-token predictions. This provides dense, token-level supervision to the student. On-policy self-distillation (OPSD) eliminates the need for the external teacher. Specifically, a second frozen copy of the student model, now given the ground truth in its context, serves as the teacher. The student model only receives the problem and learns to mimic the privileged teacher model, while the teacher remains frozen throughout training. Previous work showed that freezing the teacher is useful for training stability, but we argue that this can prevent the teacher from incorporating the improvem
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית