כתבה
arXiv cs.LG ·
שכבות LLM מתקנות זו את זו מיד
LLM Layers Immediately Correct Each Other
חוקרים גילו מנגנון תיקון בשכבות ה-Transformer של מודלים LLM. המנגנון, TLCM, מאפשר לשכבות סמוכות לתקן זו את זו. TLCM נמצא ב-5 מתוך 7 משפחות מודלים פתוחות ופועל על רוב הטוקנים.
תקציר מקורי באנגליתarXiv:2609.07876v1 Announce Type: cross Abstract: Recent methods in language model interpretability employ techniques such as sparse autoencoders to decompose residual stream contributions into linear, semantically meaningful features. Such methods are commonly interpreted as identifying features that persist in the residual stream and that subsequent layers build upon. We challenge this view by identifying the Transformer Layer Correction Mechanism (TLCM), wherein adjacent transformer layers systematically counteract portions of each other's contributions. TLCM appears in 5 out of 7 major open-source model families and activates across nearly all tokens in diverse texts. We show that TLCM emerges during pretraining, operates most strongly on contextually dependent tokens, and adaptively c
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית