כתבה
arXiv cs.AI ·
Beyond Rephrasing: Book-Level Organization Improves Synthetic Textbook Data for Mid-Training
תקציר מקורי באנגליתarXiv:2607.28109v2 Announce Type: replace Abstract: Synthetic textbook data has improved language model pre-training, but prior work largely treats the benefit as a property of generated content or local rewriting style. We study a different factor: whether related content is organized into coherent book-level documents. We contribute both a scalable synthesis pipeline and controlled evidence that this organization matters. The pipeline retrieves source material from a pre-training corpus, clusters it into topical units, plans hierarchical tables of contents, and assembles source-grounded sections into complete books (our Full setting), yielding 686K textbooks (32B tokens) across 15,000+ disciplines. Replacing natural books in a mid-training mix with this corpus improves downstream perform
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית