כתבה
arXiv cs.LG ·
TANGO: Treating Tokens as Operators
תקציר מקורי באנגליתarXiv:2608.22117v2 Announce Type: replace Abstract: Transformers separate cross-token mixing in self-attention from token-wise transformation in feed-forward networks. We ask whether combining these operations can lower predictive loss under fixed data and parameter budgets. To do so, we introduce the Token-Aggregated Nonlinear Gating Operator (TANGO) model. TANGO computes a nonlinear feature-wise gate at each source token. Attention averages these gates for each destination. The average modulates a linear projection of the destination and forms the diagonal core of a source-conditioned linear operator. We test this proposal by comparing full-prefix and windowed TANGO with looped and untied Transformers, the Gated Attention Unit (GAU), and Fast Linear Attention with a Single Head (FLASH) o
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית