יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

כאשר קטיעה הופכת את התיקון: דינמיקות הכשל של קטיעה נקודתית של קל-קרא קדימה-OPSD

When Clipping Reverses Correction: Failure Dynamics of Pointwise Forward-KL On-Policy Self-Distillation
במחקר זה נחקרה תופעה של קטיעה נקודתית של קל-קרא קדימה-OPSD, שהוצעה כדרך לשיפור תהליך הלמידה. נמצא כי קטיעה זו עשויה לגרום לכשל בתהליך הלמידה, ולהפכה של התיקון.
תקציר מקורי באנגליתarXiv:2609.38995v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) trains a student on its own generated responses using feedback from the same model conditioned on privileged information. On mathematical reasoning, the original OPSD study finds that stylistic tokens can dominate the training signal over math-related tokens, and that pointwise clipping of the forward KL objective stabilizes training. Pointwise clipping caps each vocabulary-wise forward KL term at a fixed threshold before summing over the vocabulary. Follow-up studies have adopted this clipping, but its effect on training has not been directly examined. In matched training runs differing only in whether clipping is applied, we observe that clipped runs produce substantially more repetitions that persist to t
קרא במקור המקורי