יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

PG-SFT: שיפור יכולות ושימור יכולות באימון חדש של סוכנים ללא גישה לאינטרנט

PG-SFT: Balancing Capability Acquisition and Retention in Offline Agent Fine-Tuning
במאמר זה, המחברים חוקרים את האופן שבו ניתן לשפר יכולות ולשמור יכולות באימון חדש של סוכנים ללא גישה לאינטרנט. הם מציגים את PG-SFT, תוכנה שמשתמשת במידע על תוצאות הסוכן כדי לאימון אותו. התוכנה נבדקה במספר תרחישים והתוצאות היו טובות. המחברים מציעים את PG-SFT כאלטרנטיבה לאימון סוכנים ללא גישה לאינטרנט.
תקציר מקורי באנגליתarXiv:2610.00949v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) on offline agent trajectories is the standard approach for training specialized tool-using agents, but forcing models to imitate reasoning and actions token by token may harm other capabilities (e.g., general reasoning, tool calling, code generation) of the base model. In this work, we focus on studying \emph{how to better balance the trade-off between acquiring new capabilities and preserving existing ones during agent trace SFT}. By comparing several baselines in our setup, standard SFT improves the target benchmark while lowering several non-target benchmark scores; meanwhile, simply constraining distributional drift using KL penalty or limiting the update magnitude did not avoid this regression trend. Motivate
קרא במקור המקורי