יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

פעולה ראשונה, היגיון שני: האצת עיבוד לסוכנים רב-תורים

Act First, Reason Later: Accelerating On-Policy Distillation for Multi-Turn Agents via Reference-Conditioned Inverse Dynamics
חוקרים הציגו ActFirst-OPD, שיטה להאצת עיבוד לסוכנים רב-תורים. השיטה מאפשרת לסוכנים לפעול ראשונה ולהיגיון שני, ומשפרת את מהירות האימון. ניסויים הראו שהשיטה משיגה שיפורים משמעותיים במהירות האימון ובאיכות התוצאות.
תקציר מקורי באנגליתarXiv:2609.36608v1 Announce Type: cross Abstract: On-policy distillation (OPD) trains multi-turn language agents with dense teacher supervision on student-generated responses. However, standard think-then-act rollouts require lengthy reasoning before each short action, delaying environment transitions and experience collection. Generating actions directly reduces this delay but can degrade rollout quality. To address this, we propose ActFirst-OPD, an act-first, reason-later training framework that decouples environment interaction from full-response generation. The student infers and executes actions through reference-conditioned inverse dynamics using its current interaction context and a reference next observation, and switches to autonomous next-action prediction when the resulting tran
קרא במקור המקורי