יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

PHBA: Prefix-State Hybrid Block Attention

PHBA: Prefix-State Hybrid Block Attention - חידוש בארכיטקטורה למודלי זיכרון ארוך-תקופה. המאמר מציג חידוש בשם PHBA, שמשלב ריבוי-מצבים ובלוקים עם תשומת-לב-קצרה. המודל נבחן באמצעות טכנולוגיית Triton, שמאפשרת זרימה של קבוצות-בלוקים ומצבי-ראש-קצרים.
תקציר מקורי באנגליתarXiv:2610.08527v1 Announce Type: new Abstract: Hybrid architectures combining linear sequence models with softmax attention provide an effective balance between efficient long-context modeling and precise token retrieval. Existing designs such as Native Hybrid Attention (NHA) combine compressed long-term states with sliding-window attention, but their exact attention is restricted to a fixed local window. In this work, we introduce Prefix-State Hybrid Block Attention (PHBA), which replaces local sliding-window attention with top-k block-sparse retrieval and couples each retrieved block with a compact prefix state summarizing its preceding context. The prefix states are constructed by a gated linear recurrence at block boundaries and retrieved together with the corresponding token blocks,
קרא במקור המקורי