יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

השגת רגישות טקסטואלית בדגמי שפה על ידי התאמה הירוריסטית ולמידה של סופרטוקן

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning
דגמי שפה נתקעים בתקן טקסטואלי קבוע. חוקרים הציגו פתרון לבעיה זו. הם פיתחו פרקטיקה להתאמה הירוריסטית של דגמי שפה, וכן פיתחו סופרטוקן לשפה. פרויקט זה יכול לשפר את יעילות דגמי שפה.
תקציר מקורי באנגליתarXiv:2505.09738v2 Announce Type: replace Abstract: Pretrained language models (LLMs) are often constrained by their fixed tokenization schemes, leading to inefficiencies and performance limitations, particularly for multilingual or specialized applications. This tokenizer lock-in presents significant challenges. standard methods to overcome this often require prohibitive computational resources. Although tokenizer replacement with heuristic initialization aims to reduce this burden, existing methods often require exhaustive residual fine-tuning and still may not fully preserve semantic nuances or adequately address the underlying compression inefficiencies. Our framework introduces two innovations: first, Tokenadapt, a model-agnostic tokenizer transplantation method, and second, novel pre
קרא במקור המקורי