יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

קבלו לדעת כשלהקטין, ואיפה להתחיל מחדש: הגברת המהירות של העברת יכולות באופן אוניברסלי

Know When to Stop, Where to Restart: Accelerating Multi-Turn Agentic On-Policy Distillation
במאמר זה, המחברים מציגים שיטה חדשה להגברת המהירות של העברת יכולות באופן אוניברסלי. השיטה, הנקראת STRIDE, משתמשת בטכניקה של עצירה מותאמת ובאחסון של פרפקסים. המחברים מציגים תוצאות של ניסויים שהראו שהשיטה מצליחה להגביר את המהירות של העברת יכולות באופן יעיל.
תקציר מקורי באנגליתarXiv:2609.14636v1 Announce Type: cross Abstract: On-policy distillation (OPD) has become a standard approach for transferring capabilities from large teachers to compact students. Its cost, however, is dominated by autoregressive student rollouts and scales poorly in multi-turn agentic settings. Existing acceleration methods truncate or relocate the supervision signal according to fixed, offline budgets, despite substantial variation in teacher-signal reliability both within and across trajectories. Our empirical analysis on $\tau^2$-bench reveals a clear structure in this variation: informative supervision is concentrated in the prefix of each turn, and, most importantly for multi-turn agentic training, the cross-turn loss of teacher endorsement is temporally locked to the student's firs
קרא במקור המקורי