כתבה
arXiv cs.LG ·
LampAttention: פיצוח תפקידים מעודכן בעל ראייה מוקדמת לאקסלרטורים ייעודיים
LampAttention: Look-Ahead Mixed-Precision FlashAttention for Dedicated Accelerators
LampAttention היא פיצוח תפקידים חדשה שמשתמשת בראייה מוקדמת ומעודכן בעל ראייה מוקדמת לאקסלרטורים ייעודיים. הפיצוח זה יכול לשפר את יעילות החישוביים של מערכות עם גמישות גבוהה.
תקציר מקורי באנגליתarXiv:2609.39361v1 Announce Type: new Abstract: While most attention logits can be computed in low precision without degrading numerical stability, current attention kernels fail to exploit this phenomenon. We introduce a novel hardware-algorithm co-design in the form of mixed-precision FlashAttention. Our method accumulates key-query products and evaluates their exponentials in 8-bit formats, then adaptively identifies sensitive sub-blocks and recomputes them in 16-bit formats. We propose the specifications for a dedicated accelerator capable of executing this pipeline efficiently. Simulated experiments with Qwen3 and Gemma 3 show that rerouting a selective minority of sub-blocks to high precision is sufficient to recover the baseline model performance.
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית