כתבה
arXiv cs.LG ·
Do Vision Transformers Need All-to-All Attention? Global Communication Through Elastic Learned Cores
תקציר מקורי באנגליתarXiv:2605.12491v2 Announce Type: replace-cross Abstract: Vision Transformers (ViTs) learn rich visual-semantic representations through all-to-all self-attention among patch tokens. However, this design implicitly assumes that direct pairwise patch interactions are necessary for effective representation learning. In this work, we challenge this assumption and show that representations supporting both global recognition and dense prediction can be learned without direct patch-to-patch interaction. We propose VECA (Visual Elastic-Core Attention), a vision transformer with core-periphery structured attention mediated by a small set of learned cores. Patch tokens exchange global information exclusively through these cores, while the full set of dense patches are preserved and iteratively updat
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית