כתבה
arXiv cs.CL ·
TabiBERT: מודל ModernBERT גדול-קנה מידה לטורקית
TabiBERT: A Large-Scale ModernBERT Foundation Model and A Unified Benchmark for Turkish
TabiBERT הוא מודל ModernBERT גדול-קנה מידה לטורקית, מאומן מאפס עם 1 טריליון טוקנים. המודל תומך באורך הקשב של 8,192 טוקנים ומוביל בחמישה מתוך שמונה קטגוריות.
תקציר מקורי באנגליתarXiv:2512.23065v4 Announce Type: replace Abstract: The introduction of BERT established encoder-only transformer models as a foundational paradigm in natural language processing. Encoder-only models remain the standard tool for classification, tagging and retrieval, where contextual representations and low inference cost matter more than text generation, yet Turkish lacks a monolingual encoder trained from scratch with the advances consolidated in ModernBERT (rotary positional embeddings, FlashAttention, refined normalization). We introduce TabiBERT, a monolingual Turkish encoder based on the ModernBERT architecture, pretrained from scratch for one trillion tokens sampled from an 86.58B-token multi-domain corpus of web text (72%), scientific publications (19%), source code (6%) and mathem
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית