יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

קידוד עמדה יחסית בתשומת לב NoPE גלובלית

How Local Mixing Encodes Relative Position in Global NoPE Attention
חוקרים גילו כיצד שילוב של שכבות מיקסינג מקומיות ותשומת לב גלובלית NoPE יכול לקודד עמדה יחסית. המחקר מציע הסבר תאורטי וניסויי לתופעה, ויכול לשפר את הבנתנו את האופן בו מודלים היברידיים מקודדים עמדה.
תקציר מקורי באנגליתarXiv:2609.38109v1 Announce Type: new Abstract: The attention operation is naively position invariant. However, positional information is fundamental to natural language, and therefore a variety of explicit position encodings have been developed in transformer-based models, such as rotary position encoding (RoPE). Although explicit position encodings have long been assumed to be required, recent methods that interleave local mixing layers, such as sliding window attention (SWA) and gated linear attention, while not encoding position (NoPE) in global attention layers has recently been shown to be successful at scale. How and why this approach works is not well-understood. In this paper, we develop an explanation of how hybrid models of this sort can implicitly encode position at global NoPE
קרא במקור המקורי