כתבה
arXiv cs.AI ·
אילוף מדיניות ללא נתונים
Data-free On-policy Distillation
חוקרים גילו כי אילוף מדיניות (OPD) כמעט אדיש לנתוני האימון שלו. הם פיתחו שיטה חדשה, Data-free On-policy Distillation (DF-OPD), שמאפשרת אילוף ללא נתונים חיצוניים. DF-OPD משיגה תוצאות טובות יותר מאשר שיטות קודמות, ומציעה פוטנציאל לשיפור היעילות של אילוף מדיניות.
תקציר מקורי באנגליתarXiv:2609.14193v1 Announce Type: cross Abstract: On-policy distillation (OPD) has become a standard component of frontier post-training pipelines, yet how much its training data actually contributes has gone largely unexamined. On the two teacher-student pairings most common in practice, we find OPD almost indifferent to its data: eight prompts already match a 17k-problem dataset, and three independently built datasets whose difficulty and teacher-student KL differ several-fold produce nearly indistinguishable training curves. Two causes account for this. First, the unit of data in OPD is the state a prompt leads to, not the prompt itself: a single prompt keeps exposing new teacher correction as sampling continues, while the marginal value of additional prompts collapses after eight. Seco
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית