כתבה
arXiv cs.LG ·
Learned, Then Lost: A Measured Single-Example Counterfactual in Pre-training
תקציר מקורי באנגליתarXiv:2608.19168v2 Announce Type: replace Abstract: A single training example's contribution to a finished model is normally estimated rather than measured, because measuring it takes two expensive full pre-training runs that differ in one row of one batch. We ran that counterfactual 24 times at a small scale. We trained 32 GPT-2 models at 124M parameters from scratch on OpenWebText, over four conditions and eight seeds. At step 200 of 9,536, at peak learning rate, we replaced one row of a 256-row batch with a fixed context injection carrying a 194-token passage. The three injected conditions are: 1. fluent prose with a corpus-attested subject, 2. fluent prose with a fabricated subject matched to it within 0.14% on full-batch gradient delta, and 3. random keyboard characters. The fourth co
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית