כתבה
arXiv cs.AI ·
AREX-2: שיפור סוכנים עצמיים
AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks
AREX-2 הוא פרויקט לשיפור יכולת השיפור העצמי של סוכנים. הוא משתמש בנתונים ממשימות תכנות ולמידת מכונה. הסוכן מבוסס על Qwen3.8-27B ומראה תוצאות חזקות במשימות שונות.
תקציר מקורי באנגליתarXiv:2609.38288v1 Announce Type: new Abstract: We present AREX-2, an effort to advance the self-improving capability of LLM agents, which we define as the ability to iteratively refine a solution at test time. This ability rests on two complementary capabilities: reflection, which produces a solution better than the current one, and long-horizon execution, which keeps the iteration effective over many rounds. We hypothesize that both capabilities are domain-agnostic, and can therefore be learned in scenarios that are well suited for supervision. Accordingly, we synthesize long-horizon improvement trajectories from machine learning and algorithmic programming tasks, two domains that offer verifiable feedback and reward sustained iteration. Trained on this data, our agent, built on Qwen3.8-
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית