יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

וידאו YT AI Engineer ·

Lessons from Generating 12 Trillion Synthetic Tokens — Bogdan Gaza, DatologyAI

▶ צפה כאן — בלי לצאת מהאתר
תקציר מקורי באנגליתThe web has roughly 30 trillion tokens of usable text. Frontier pre-training needs more. Here's how DatologyAI generates the rest. Bogdan Gaza, co-founder and CTO of DatologyAI, shares the engineering lessons from running synthetic data jobs at trillion-token scale, including a recent run of about 12 trillion tokens across web, math and code. He covers BeyondWeb, DatologyAI's rephrasing-based synthetic data recipe, and the move from a split Slurm/Kubernetes setup to a single Ray, KubeRay and vLLM pipeline on EKS on HyperPod. He then walks through four bottlenecks: S3 metadata, GPU failures, cross-cluster scheduling and inference tuning. In this talk: • Why seeded rephrasing beats asking a model for synthetic data from scratch • Batching S3 metadata fetches to cut 9–11 days down to about 2
קרא במקור המקורי