כתבה
arXiv cs.CL ·
Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training
תקציר מקורי באנגליתarXiv:2506.01732v4 Announce Type: replace Abstract: Large Language Models (LLMs) are pre-trained on large amounts of data from different sources and domains. Such datasets often contain trillions of tokens, including large portions of copyrighted or proprietary content, which raises questions about the legal use of such models. This underscores the need for truly open pre-training data that complies with data security regulations. In this paper, we introduce Common Corpus, the largest open dataset for LLM pre-training. The data assembled in Common Corpus are either uncopyrighted or under open licenses, totaling about two trillion tokens. The dataset contains a wide variety of languages, ranging from the high-resource European languages to some low-resource languages rarely represented in p
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית