כתבה
arXiv cs.AI ·
Hybrid Latent Attention למודלי שפה מחזורי
Hybrid Latent Attention for Looped Language Models
מודלי שפה מחזוריים נשפרו עם Hybrid Latent Attention, שמציעה שיפור בביצועי גיבוי וביצועי גיבוי
תקציר מקורי באנגליתarXiv:2610.07940v1 Announce Type: cross Abstract: Looped language models apply the same stack of layers T times to each token, which deepens the model without adding parameters but multiplies its key-value (KV) cache by T. The larger cache limits how many sequences a GPU can decode at once and slows each decoding step, which reads the whole cache. We propose Hybrid Latent Attention (HLA), which keeps exact keys and values within a sliding window of W recent tokens and stores each older token as a compact latent that the query of each loop reads directly, without reconstructing keys and values. We uptrain HLA on Ouro looped models (T=4) with 1.4B and 2.6B parameters, keeping the pretrained weights frozen and training only the added parameters to reproduce the original attention. The cache s
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית