כתבה
arXiv cs.CL ·
אוסף נתונים של 41 מיליארד טוקנים מהווב הפורטוגזי
Fine PT-PT Web: A High-Quality 41 Billion Tokens Data Collection of the European Portuguese Web
חוקרים פיתחו פייפליין יעיל לקידום נתונים מהווב הפורטוגזי. הם השתמשו במודלים עם זיהוי שפה ודדופליקציה משוקללת. התוצאה היא אוסף נתונים נקי ומייצג של 41 מיליארד טוקנים.
תקציר מקורי באנגליתarXiv:2609.07699v2 Announce Type: replace Abstract: Curating Web corpora for regional language variants like European Portuguese (PT-PT) is heavily bottlenecked by dialectal overlap (mainly with PT-BR) and data processing scale. This paper presents an efficient pipeline to curate a production-ready PT-PT corpus from the Portuguese Web, spanning 411 TB of raw data from Arquivo.pt. We introduce a novel post-scraping block that removes boilerplate and line duplicates prior to filtering. This early-stage intervention increases final document yield by 19.04% by rescuing valid text that standard heuristic filters prematurely discard. Integrated with rigorous language identification, weighted fuzzy deduplication, and neural quality classification, our pipeline offers a scalable framework and a cl
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית