כתבה
arXiv cs.CL ·
חללי מושגים: חישוב מעבר לעדשת הלוגיט
Concept Subspaces Compute Beyond the Logit Lens: A Weights-Only Test for Locating Representations Upstream of Readout
חוקרים פיתחו שיטה חדשה לאיתור ייצוגים במודלים, מעבר לעדשת הלוגיט. השיטה מודדת את החפיפה בין תת-מרחבים מושגיים לבין הכיוונים הדומיננטיים של מטריצת האילוף. המחקר בודק את השיטה על מודלים שונים ומראה תוצאות מבטיחות.
תקציר מקורי באנגליתarXiv:2609.39263v1 Announce Type: new Abstract: A concept subspace's effect on model behavior does not establish how it relates to the output readout. We introduce a two-sided geometric diagnostic that measures an extracted subspace's overlap with the dominant right-singular directions of the unembedding matrix, evaluated against output-oriented positive controls. Given an extracted basis, the raw diagnostic requires only model weights. Our testbed is the Format-Agnostic Reasoning Subspace (FARS), a ten-dimensional basis extracted from eighteen reasoning concepts expressed in six surface forms. Across nine rank-matched estimators and twenty-six models, four activation-derived concept estimators carry only 0.38--0.80% mean energy in the top-ten readout span. Final-layer PCA carries 3.56%, e
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית