יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

E$^2$-OPSD: תיפופחת האנטרופיה המוגברת בהתפלגות האופטימלית של עצמ-התפלגות

E$^2$-OPSD: Taming Entropy Overshoot in On-Policy Self-Distillation
E$^2$-OPSD: פתרון לבעיה של אנטרופיה מוגברת בהתפלגות האופטימלית של עצמ-התפלגות. המאמר מציג פתרון חדש לבעיה של אנטרופיה מוגברת בהתפלגות האופטימלית של עצמ-התפלגות. הפתרון, E$^2$-OPSD, משתמש בשיטת התפלגות מודלים כדי לפחות את האנטרופיה המוגברת. המאמר מציג תוצאות של E$^2$-OPSD שהושוותה ל-OPSD.
תקציר מקורי באנגליתarXiv:2610.05048v3 Announce Type: replace Abstract: On-policy self-distillation (OPSD) provides dense token-level supervision without a second model: one network acts as teacher with the reference solution and as student with only the problem. We identify a specific failure mode of this recipe. During training, student token entropy rises past the teacher's and remains elevated, a pattern we call entropy overshoot. We trace it to both sides of distillation. The reference-conditioned teacher is confident along its answer-directed reasoning path, but this confidence transfers poorly to student-generated prefixes, making its supervision overly tied to answer-specific cues rather than reusable reasoning patterns; meanwhile, the forward KL used by OPSD continually diffuses the student's predict
קרא במקור המקורי