יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

מורה עצמי טוב פוגש את התלמיד בנקודתו

A Good Self-Teacher Meets the Student Where They Are: Joint On-Policy Learning and Teaching
שיטה חדשה ללמידת חיזוק משולבת עם הוראה עצמית, Joint On-Policy Learning and Teaching, משפרת את יעילות האימון והביצועים. השיטה מאפשרת למודל ללמוד מעצמו ולהורות את עצמו בו-זמנית.
תקציר מקורי באנגליתarXiv:2610.10447v1 Announce Type: new Abstract: Reinforcement Learning (RL) from outcome rewards suffers from sparse supervision, particularly on difficult, long-horizon tasks where successful trajectories are rare and costly to generate. On-Policy Distillation (OPD) offers an attractive alternative by providing dense token-level supervision from a stronger teacher along the student's own generations. Self-distillation methods further remove the need for a separate teacher model by conditioning the same policy on privileged information to serve as its own teacher. However, privileged conditioning alone does not guarantee that the resulting distillation update improves the student. Indeed, privileged information can lead the teacher to solve tasks through shortcuts unavailable to the studen
קרא במקור המקורי