כתבה
arXiv cs.LG ·
Manifold Projection and Iterative Autoencoder Refinement for Masked Language Modeling
תקציר מקורי באנגליתarXiv:2609.30288v1 Announce Type: cross Abstract: In Transformer-based masked language models, attention is the primary mechanism for context mixing, but there are other ways to mix data across tokens. Recent attention-free mixers replace attention with fixed or hypernetwork-generated MLPs, alternating their dynamic, content-dependent weighting for computational simplicity. We build an alternative that gets the same property from a low-rank bottleneck autoencoder. We replace attention with a stack of autoencoder-based mixing modules, one operating over local neighborhoods, one over the full sequence, and one across attention heads, each compressing and reconstructing its input through a bottleneck, and its width is a hyperparameter rather than a training effect. In masked positions, we int
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית