יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

תופעת Grokking במודלים שפה

A Pre-Training Analogue of Grokking in Language Models: Tracing Delayed Grammatical Generalization
חוקרים תופעת Grokking במודלים שפה, שבהם המודלים מכללים הבנה דקה אחרי תקופה ארוכה. המחקר מתמקד ב-BLiMP, ומראה שווקטורים גרמתיים הופכים ליותר מנבאים.
תקציר מקורי באנגליתarXiv:2606.00230v2 Announce Type: replace Abstract: Grokking, the phenomenon in which neural networks generalize long after fitting their training data, has been studied in supervised settings on many epochs. LLM pre-training instead involves next-token prediction over an unlabeled corpus, with limited data repetition and no explicit train/validation split. To address this, we propose an exposure-based framework that enables the study of grokking-like dynamics during LLM pre-training. We ground our evaluation in BLiMP minimal pairs, which provide controlled grammatical contrasts. For every BLiMP minimal pair, we identify a critical phrase, the smallest continuous span that captures the grammatical contrast and the phenomenon-relevant context. Examples whose critical phrase appears in the p
קרא במקור המקורי