כתבה
arXiv cs.LG ·
A-Evolve-Training: טריינינג אוטונומית של מודל 30B
A-Evolve-Training: Autonomous Post-Training of a 30B Model
מערכת אוטונומית של טריינינג פוסט-טריינינג של מודל 30B. המערכת טריינה את המודל ללא התערבות אנושית, והגיעה לתוצאות טובות יותר מאלה של חוקרים אנושיים. המערכת גם זיהתה שהמדד שלה הפך למוטה, ושינתה את הדגש שלה.
תקציר מקורי באנגליתarXiv:2606.20657v3 Announce Type: replace-cross Abstract: Post-training a frontier model is normally weeks of human work: proposing data and recipe changes, launching runs, reading evals, deciding what to keep. We report an autonomous system that runs this loop with no human in the loop, post-training a 30B Nemotron across four rounds over multiple weeks. The autonomously produced model reaches a held-out score of 0.86 against the top human submission's 0.87 on the public NVIDIA Nemotron-Reasoning Challenge leaderboard, placing 8th of ~4000 at the time of writing. More striking than the number: the loop detected that its own dev metric had stopped tracking external performance on the weakest domain -- candidates drove dev to record highs without moving the external target -- and revised it
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית