כתבה
Dwarkesh Podcast ·
קדימה באימון מודלים מגיעה בעיקר מנתונים
Pretraining progress is mostly coming from data
חקר התקדמות האימון של מודלים מראה כי רוב הקידום מגיע משיפור נתונים. המחקר בדק את השיפורים במודלים ונתונים בין 2019 ל-2025, ומצא כי 3.24 פעמים יותר יעילות חישובית הגיעה משיפור נתונים מאשר משיפור מודלים.
תקציר מקורי באנגליתHow much of the rapid progress in AI that we’ve seen over the last few years 1 has come from data versus model improvements? The answer has big implications for the economics of frontier labs and the pace of future progress. We investigate this question at a relatively small scale, and for pretraining specifically, from 2019 to 2025. During each of those years, a new open model recipe was published which codified that year’s publicly known algorithmic tweaks (for example, improvements in architecture, optimizer, initializations, learning rate schedule, hyperparams, etc). And during each of those years, there was also a new public data corpus (produced by broader scrapes and new curation/extraction/filtering techniques). We train combinations of these year-representative model r
קרא במקור המקורי
dwarkesh.com
פתח כתבה מקורית