יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

פרסום: תפאורה נחושה כשינוי באופן-משחק עצמי-התפשטות

Privileged Context as Drift in On-Policy Self-Distillation
במאמר זה נחקרה השפעת תפאורה נחושה על שינויי מדיניות באופן-משחק עצמי-התפשטות. המחקר נערך עם מודל Qwen2.5-7B.
תקציר מקורי באנגליתarXiv:2610.07842v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) trains a language model to match a copy of itself conditioned on privileged context. Existing work varies what privileged context contains and how it is produced while also changing models, data, and training setups, making the effects of privileged context design difficult to isolate. Motivated by efforts in continual learning to reduce catastrophic forgetting, we study how the choice of privileged context affects policy drift. Specifically, we vary two axes: content (a demonstration, feedback, or rephrase) and source (external, self-generated with a verifier, or self-generated without a verifier). We train Qwen2.5-7B with OPSD across these nine combinations and three datasets, measuring target-task accurac
קרא במקור המקורי