כתבה
arXiv cs.CL ·
פרויקציה למישור ושיפור רציפי של אוטואנקודר
Manifold Projection and Iterative Autoencoder Refinement for Masked Language Modeling
במאמר זה, המחברים מציגים פרויקציה למישור ושיפור רציפי של אוטואנקודר למודלי שפה חסויים. הם מציעים שיטה חדשה שבמקום שימוש בתפקידי התייחסות, הם משתמשים בפרויקציה למישור ובשיפור רציפי של אוטואנקודר. השיטה החדשה מציעה תוצאות טובות יותר מאשר השיטה המקורית, והיא יעילה יותר.
תקציר מקורי באנגליתarXiv:2609.30288v1 Announce Type: new Abstract: In Transformer-based masked language models, attention is the primary mechanism for context mixing, but there are other ways to mix data across tokens. Recent attention-free mixers replace attention with fixed or hypernetwork-generated MLPs, alternating their dynamic, content-dependent weighting for computational simplicity. We build an alternative that gets the same property from a low-rank bottleneck autoencoder. We replace attention with a stack of autoencoder-based mixing modules, one operating over local neighborhoods, one over the full sequence, and one across attention heads, each compressing and reconstructing its input through a bottleneck, and its width is a hyperparameter rather than a training effect. In masked positions, we intro
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית