כתבה
arXiv cs.CL ·
DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data
תקציר מקורי באנגליתarXiv:2607.24717v1 Announce Type: new Abstract: Pretraining data processing is critical to the downstream performance of Large Language Models (LLMs). However, many existing approaches define a fixed processing strategy at the corpus or domain level and apply it uniformly to many examples, without adapting to the needs of each example. We propose DataOrchestra, a framework that unifies different processing operations and orchestrates an example-specific pipeline for each example. Given a chunk of pretraining data, an orchestrator decides whether to drop, untouch, or clean it. For a chunk to be cleaned, it selects one or more downstream operations, ranging from programmatic editing to different forms of LLM-based rewriting. For each rewriting step, it further generates a concrete instructio
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית