כתבה
arXiv cs.LG ·
Conditional Transfer from Controlled Pretraining Mixtures to Code
תקציר מקורי באנגליתarXiv:2610.11548v1 Announce Type: new Abstract: Synthetic tasks are increasingly used both as probes of language-model capability and as pretraining data. Both uses are often justified by loss reduction: falling loss is treated as informative, and faster loss reduction with more sampling as evidence that a task is worth sampling. We separate three signals. A task is diagnostic when its loss tracks global pretraining progress; it is teachable when its loss responds to its own token budget; and a data source transfers when including it improves a downstream target. We study controlled pretraining in which 70% of the corpus is fixed general Python and the remaining 30% is a simplex over three source families: OpenCodeInstruct, a curated suite of 12 code-adjacent synthetic tasks, and 15 litera
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית