כתבה
arXiv cs.LG ·
לדעת מתי לעצור ולאן להתחיל מחדש: הקצאת יעילות להטמעה מרובה-סובבים של תהליך הטמעה
Know When to Stop, Where to Restart: Accelerating Multi-Turn Agentic On-Policy Distillation
STRIDE (Stop-and-Restart on-policy Distillation acceleration) הוצג כפתרון להקצאת יעילות להטמעה מרובה-סובבים של תהליך הטמעה. השיטה משתמשת בטכניקות של עצירה מותאמת ואחסון של פרפקסים. STRIDE נבחנה על $ au^2$-bench והציגה תוצאות טובות יותר מאשר תהליך הטמעה המלא.
תקציר מקורי באנגליתarXiv:2609.14636v1 Announce Type: new Abstract: On-policy distillation (OPD) has become a standard approach for transferring capabilities from large teachers to compact students. Its cost, however, is dominated by autoregressive student rollouts and scales poorly in multi-turn agentic settings. Existing acceleration methods truncate or relocate the supervision signal according to fixed, offline budgets, despite substantial variation in teacher-signal reliability both within and across trajectories. Our empirical analysis on $\tau^2$-bench reveals a clear structure in this variation: informative supervision is concentrated in the prefix of each turn, and, most importantly for multi-turn agentic training, the cross-turn loss of teacher endorsement is temporally locked to the student's first
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית