כתבה
arXiv cs.LG ·
Beyond Linear Concepts: Discovering and Aligning Non-Linear Concept Manifolds in Large Language Models
תקציר מקורי באנגליתarXiv:2610.01821v1 Announce Type: new Abstract: Understanding information processing in large language models (LLMs) requires dissecting the geometric organization of their internal token representations. While existing mechanistic interpretability (MI) methods seek to extract concepts, they are constrained by a strong linearity assumption challenged by evidence of non-linear feature manifolds. We move beyond linear concepts by adapting Non-Linear Multi-Dimensional Concept Discovery (NLMCD) from computer vision to token-level LLM activations, modeling concepts as low-dimensional manifolds. To compare concept manifolds across layers and models, we introduce a concept-based alignment (CBA) score, a generalized Rand index that measures geometric proximity without explicit feature matching. Ou
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית