יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

מערכות כמו תשתית: העשרת תגובות להפעלת סוכנים עצמאיים במשימות ארוכות-טווח

Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks
במאמר זה, נציגים פרדיגמה חדשה להעשרת תגובות לסוכנים עצמאיים במשימות ארוכות-טווח. הפרדיגמה, הקרויה FEEs, מציעה להעשיר את הסביבה שבה פועל הסוכן, כדי לסייע לו ללמוד ולהתפתח. המחברים מציגים תוצאות מחקר שמדגימות את יעילות FEEs במשימות ארוכות-טוור.
תקציר מקורי באנגליתarXiv:2609.08404v1 Announce Type: new Abstract: Large Language Models demonstrate remarkable proficiency in static reasoning, yet training them as autonomous agents through Reinforcement Learning (RL) for long-horizon tasks is often hindered by severe reward sparsity. While conventional \textit{agent-side warming} up via supervised fine-tuning (SFT) can alleviate this, it is frequently limited by data scarcity and constrained exploration. To address this, we propose a paradigm shift to \textit{environment-side adaptation} by constructing \textbf{F}eedback-\textbf{E}nriched \textbf{E}nvironments (\textbf{FEEs}). Through a pilot study, we establish a feedback design strategy that reformulates environments by transitioning from action guidance to observation enrichment during the later stages
קרא במקור המקורי