כתבה
arXiv cs.AI ·
LLM-Microscope: גילוי תפקיד הנקודות בזיכרון ההקשרי של טרנספורמרים
LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers
LLM-Microscope הוא כלי פתוח לבדיקת זיכרון הקשרי של מודלים גדולים. הוא מראה כי נקודות ומילים קטנות משחקות תפקיד חשוב בשמירה על הקשר. הכלי מאפשר לבדוק את ההשפעה של מחיקת נקודות ומילים על ביצועי המודל.
תקציר מקורי באנגליתarXiv:2502.15007v2 Announce Type: replace-cross Abstract: We introduce methods to quantify how Large Language Models (LLMs) encode and store contextual information, revealing that tokens often seen as minor (e.g., determiners, punctuation) carry surprisingly high context. Notably, removing these tokens -- especially stopwords, articles, and commas -- consistently degrades performance on MMLU and BABILong-4k, even if removing only irrelevant tokens. Our analysis also shows a strong correlation between contextualization and linearity, where linearity measures how closely the transformation from one layer's embeddings to the next can be approximated by a single linear mapping. These findings underscore the hidden importance of filler tokens in maintaining context. For further exploration, we
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית