יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

הטמעת תכונות של גוף תיקון בפלטפורמת OPD

Multi-Turn On-Policy Distillation with Prefix Replay
במאמר זה, המחברים פיתחו טכניקה חדשה להטמעת תכונות של גוף תיקון בפלטפורמת OPD. הם הציגו את Replayed-Prefix On-Policy Distillation (ReOPD), שמאפשרת לשלב תכונות של גוף תיקון בפלטפורמת OPD באופן יעיל ומהיר. המחברים הראו כי ReOPD יעיל יותר ומהיר יותר מ-OPD, ומאפשרת לשלב תכונות של גוף תיקון בפלטפורמת OPD באופן יעיל ומהיר.
תקציר מקורי באנגליתarXiv:2607.04763v3 Announce Type: replace Abstract: We study on-policy distillation (OPD) for agentic tasks, where an LLM agent interacts with an environment over multiple turns and a student imitates a teacher over these multi-turn interaction histories. Fully online OPD is costly because each update requires fresh student rollouts through the environment and teacher queries at visited histories. We propose Replayed-Prefix On-Policy Distillation (ReOPD), an off-environment alternative that reuses pre-collected teacher trajectories as replayed prefixes: the student acts at selected steps, while the teacher provides dense per-step supervision without executing new environment interactions. We show that multi-turn OPD introduces a prefix trap: making histories more student-on-policy improves
קרא במקור המקורי