כתבה
arXiv cs.CL ·
NorBERTo: A ModernBERT Model Trained for Portuguese with 331 Billion Tokens Corpus
תקציר מקורי באנגליתarXiv:2605.00086v2 Announce Type: replace Abstract: High-quality corpora are essential for advancing Natural Language Processing (NLP) in Portuguese. Building on previous encoder-only models such as BERTimbau and Albertina PT-BR, we introduce NorBERTo, a modern encoder based on the ModernBERT architecture, featuring long-context support and efficient attention mechanisms. NorBERTo is trained on Aurora-PT, a newly curated Brazilian Portuguese corpus comprising 331 billion GPT-2 tokens collected from diverse web sources and existing multilingual datasets. We systematically benchmark NorBERTo against Strong baselines on semantic similarity, textual entailment and classification tasks using standardized datasets such as ASSIN 2 and PLUE. On PLUE, NorBERTo-large achieves the best results among
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית