יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

חזרה על תקן צורת המודלים של Transformer

Revisiting the Shape Convention of Transformer Language Models
במאמר זה, חוקרים חוקרים את תקן צורת המודלים של Transformer ומציעים תצורה חדשה של MLPs.
תקציר מקורי באנגליתarXiv:2602.06471v2 Announce Type: replace-cross Abstract: The architectural shape of dense Transformers has remained remarkably stable: narrow-wide-narrow feed-forward networks (FFNs) consume most non-embedding parameters. Motivated by theoretical and empirical evidences that residual wide-narrow-wide (hourglass) MLPs remain expressive despite bottlenecks, we revisit whether this architectural convention is necessary for dense language models. We study Hourglass Transformers, which replace the conventional FFN with residual stacks of hourglass sub-MLPs and use hourglass attention to decouple residual-stream width from attention width. This exposes a practical depth-width trade-off: compressing the FFN intermediate dimension allows wider hidden states and fewer layers at matched parameter b
קרא במקור המקורי