כתבה
arXiv cs.LG ·
התפשטות נוסטלגית: זיכרון משותף במודלי Looped Transformers
The Surprising Effectiveness of Shared Memory in Looped Transformers
מחקר חדש מצא ששיתוף זיכרון במודלי Looped Transformers יכול לשפר את האיכות ללא גידול במספר הפרמטרים. המחקר חקר את השפעת שיתוף זיכרון על האיכות של המודלים.
תקציר מקורי באנגליתarXiv:2610.02383v1 Announce Type: new Abstract: Looped Transformers apply the same layers several times per token, adding compute to improve quality without more parameters. Each recursion, however, writes its own key-value cache, so memory still grows with compute. Inference-time techniques can shrink this cache at a cost in quality. We pretrain looped language models to share memory: only the first recursion writes a cache, and later recursions read it while keeping a short window of their own. Surprisingly, we find that sharing memory does not cost quality and instead improves it. At 150M-1B parameters, our Looped Prediction Transformer (LPT) and its hybrid variant set a new quality-memory frontier for looped models: with five recursions, the hybrid lowers validation perplexity on FineW
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית