כתבה
arXiv cs.CL ·
טיני-סקייל BERT סיני: השוואה מומחה של MLM, WWM ו-MacBERT
Tiny-Scale Chinese BERT Pretraining: A Controlled Comparison of MLM, WWM, and MacBERT Strategies
במאמר זה, נערכה השוואה של שלושה רגישויות טיני-סקייל BERT סיני: MLM, WWM ו-MacBERT. התוצאות הראו ש-MLM היה המודל המצליח ביותר, בעוד WWM היה המודל הטוב ביותר בקטגוריות פרפלקסיטי ו-MLM hit rate.
תקציר מקורי באנגליתarXiv:2610.08879v1 Announce Type: new Abstract: Pretraining strategies significantly impact the quality of language models, yet existing comparisons of Masked Language Modeling (MLM), Whole Word Masking (WWM), and MacBERT-style replacement have focused primarily on base-scale models (>=110M parameters). This paper presents a controlled comparison of these three strategies on a tiny-scale Chinese BERT model (4 layers, 256 hidden dimensions, 8.7M parameters). Under identical architecture, corpus (1.29M sentences from Chinese Wikipedia), and hyperparameters, we train three models from scratch and evaluate them across five intrinsic dimensions: perplexity, MLM hit rate, semantic discrimination, grammatical judgment, and contextual sensitivity. At tiny scale, MLM achieves the best overall intri
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית