כתבה
arXiv cs.AI ·
מה חשוב בהתפשטות על-פוליצי: נקודת מבט על כושר נתונים ובחירת נתונים
What Matters in On-Policy Distillation? A Perspective on Data Efficiency and Data Selection
חוקרים חקרו כושר נתונים ובחירת נתונים בהתפשטות על-פוליצי. הם גילו שאפשר להשיג תוצאות טובות על ידי ניתוח נתונים ארוכים ובחירת דוגמאות קשות.
תקציר מקורי באנגליתarXiv:2609.05198v1 Announce Type: new Abstract: On-Policy Distillation (OPD) has emerged as a widely adopted post-training paradigm for enhancing large language models in reasoning domains. However, the data-centric mechanisms in OPD remain relatively underexplored. This paper presents a empirical study of data efficiency and data selection in OPD. We begin by investigating an extreme setting: training OPD on only one example, namely 1-shot OPD. Surprisingly, we find that 1-shot OPD is consistently effective across all sampled training examples and harder examples often yield superior performance gain. We next investigate what actually drives the student model's improvement in the training data. Our analysis reveals that the improvement is not driven by high token entropy, but the longer C
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית