כתבה
arXiv cs.AI ·
T1: תכנות סופי ללמידת רפלקסיה למשימות ארוכות-טווח
T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks
T1, מודל חדש, הצליח לבצע משימות ארוכות-טווח כמו תכנות וגילוי מדעי. הוא עובד על פלטפורמה ענקת ענן ומסוגל לבצע יותר מ-300 פעולות כל אחת. T1 גם עובד עם מודל GPT-5 והצליח להשיג תוצאות טובות יותר.
תקציר מקורי באנגליתarXiv:2609.11042v1 Announce Type: cross Abstract: Agent usage is shifting toward long-horizon tasks such as coding and scientific discovery, among which terminal tasks are especially important. We introduce T1, a Mixture-of-Experts model of 122B total trained with reinforcement learning, operating a real shell in a cloud sandbox for up to 300+ tool-call turns per task, rewarded by executing each task's own verifier. We provide a comprehensive recipe: First, an aggressively warm-started to stabilize actor-critic training, with a dense process reward scoring trajectories by the absolute number of passing verifiers. Second, stable optimization through TITO construction, training on the exact sampled token identifiers with drift repair at turn boundaries, and rollout routing replay, recording
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית