יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

דחיקות: לא כל הפשטות היא טובה: גבולות של תפיסה-עצמית עצמית ללמידה על-פוליסי

Denser $\neq$ Better: Limits of On-Policy Self-Distillation for Continual Post-Training
במאמר זה נחקרים גבולות של שיטת הלמידה העצמית-עצמית ללמידה על-פוליסי. נמצא כי שיטה זו יעילה רק כאשר הסימנים של המורה יציבים ומותאמים.
תקציר מקורי באנגליתarXiv:2607.01763v2 Announce Type: replace Abstract: Continual post-training enables foundation models to acquire new knowledge while preserving existing capabilities. Recent work suggests that on-policy learning can mitigate forgetting, with self-distillation as a particularly attractive approach. We revisit this optimistic claim through self-distillation policy optimization (SDPO). Our experiments show that SDPO accelerates in-domain specialization when teacher signals are stable and well aligned, but struggles to generalize out of distribution. In continual post-training, SDPO exhibits greater forgetting and can even collapse, whereas GRPO, the more established on-policy reinforcement learning method, adapts more conservatively and better preserves prior capabilities. Further analyses li
קרא במקור המקורי