כתבה
arXiv cs.LG ·
Repetition Mismatch: Why Data Mixture Experiments Don't Scale and How to Fix Them
תקציר מקורי באנגליתarXiv:2606.07597v2 Announce Type: replace Abstract: Pre-training data mixtures are commonly tuned by running small-scale experiments and extrapolating to the target training budget. When high-quality data is scarce and must be repeated, this extrapolation frequently fails, but the source of the failure has not been isolated. We show that a primary culprit is a repetition mismatch: because high-quality datasets are small, their repetition rate changes as the training budget grows, shifting the optimal mixture in ways that small-scale proxy experiments do not anticipate. A subsampling procedure that matches the target repetition rate controls for this effect. In a two-source setting combining limited high-quality data with web crawl, a single repetition-controlled experiment using only 1/16
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית