וידאו
YT AI Engineer ·
מציאות מסובכת של קנה מידה: נתונים סינתטיים ופרה-אימון
The Messy Reality of Scale: Synthetic Data and Pre-Training — Marah Abdin & Robert McHardy, poolside
▶ צפה כאן — בלי לצאת מהאתר
חברת poolside מפתחת נתונים סינתטיים ופרה-אימון לקודים אגנטיים. החברה משתמשת בצינור מרובשבבות ליצירת נתונים סינתטיים, ומאמנת מודלים עם 118 מיליארד פרמטרים. התוצאות הראשוניות מראות עליונות על מודל GLM 4.5 Air.
תקציר מקורי באנגליתGood code data runs out, so poolside manufactures more of it, and the hard part is making it teach. Their synthetic pipeline pairs templates with supplementary context and spreads generations across an axis of phrasing, with difficulty tuned so a task is neither trivial nor so hard the model learns nothing from it. Multistage pipelines port existing data into new shapes, swapping character styles or plots and turning single prompts into multi turn chats, while an orchestrator polices every generation and drops the ones that miss. On the training side the team trusts nothing: run two replicas of the same model on the same data and they must return the same number, or the run gets killed. That is how the messy failures surface. Broken GPUs show up as a spiky loss curve, a numerical precision
קרא במקור המקורי
youtube.com
פתח כתבה מקורית