כתבה
arXiv cs.AI ·
מעבר לניסוח: איך נתונים מעצבים העברה בתהליך אידוי מדיניות
Beyond Prompt Count: How Data Shapes Transfer in On-Policy Distillation
חוקרים בדקו כיצד בחירת ניסוחים משפיעה על העברה בין מורים לתלמידים בתהליך אידוי מדיניות. הם מצאו שמספר קטן של ניסוחים יכול להשיג ביצועים דומים לאלו של מאגר גדול. החוקרים גם בדקו את השפעת המורה על יעילות הניסוחים.
תקציר מקורי באנגליתarXiv:2609.37377v1 Announce Type: new Abstract: On-policy distillation (OPD) trains students using teacher feedback on their own sampled responses, yet how prompt choice shapes transfer across teacher-student pairs remains poorly understood. We systematically study prompt quantity, source, and selection across RL- and SFT-continuation pairs and cross-model settings. We find that OPD can be highly prompt-efficient: a few prompts can approach large-pool performance, with four DAPO prompts matching the observed mathematics score of 3,840 DeepMath prompts. However, prompt utility is relational rather than intrinsic: changing only the teacher can reverse the relative effectiveness of mathematics and code prompts. To characterize these transfer differences, we analyze parameter and functional ch
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית