כתבה
arXiv cs.CL ·
Hybrid Latent Attention for Looped Language Models
תקציר מקורי באנגליתarXiv:2610.07940v1 Announce Type: new Abstract: Looped language models apply the same stack of layers T times to each token, which deepens the model without adding parameters but multiplies its key-value (KV) cache by T. The larger cache limits how many sequences a GPU can decode at once and slows each decoding step, which reads the whole cache. We propose Hybrid Latent Attention (HLA), which keeps exact keys and values within a sliding window of W recent tokens and stores each older token as a compact latent that the query of each loop reads directly, without reconstructing keys and values. We uptrain HLA on Ouro looped models (T=4) with 1.4B and 2.6B parameters, keeping the pretrained weights frozen and training only the added parameters to reproduce the original attention. The cache shr
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית