יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

מערכת תרגום עצמית: צמצום זיכרון KV-מספרי קצבי עם תנועה תקשורת כללית

A Self-Pruning Transformer: Extreme KV-Cache Compression with Universal Attention
מערכת תרגום עצמית: צמצום זיכרון KV-מספרי קצבי עם תנועה תקשורת כללית. פיתוח חדש של פונקציות דעיכה לצמצום זיכרון KV-מספרי קצבי במערכות LLM. המערכת החדשה, Universal Attention, מציעה צמצום זיכרון KV-מספרי קצבי של 25 פעמים באורך 16,000.
תקציר מקורי באנגליתarXiv:2610.09051v1 Announce Type: new Abstract: The large KV-cache size of modern LLMs creates a barrier to efficient deployment. Recent work has explored replacing attention layers' RoPE positional embeddings with alternative decay-based mechanisms, which can then be used to prune KV-cache during inference. However, these decay functions have limited expressivity, and in practice devolve into sliding-window-like eviction patterns. In this work, we propose a unifying framework for complementary and novel decay mechanisms, capturing complex key statistics and interactions while preserving expressive RoPE embeddings and Softmax attention. The resulting Universal Attention is a highly expressive and end-to-end trainable architecture, whose composite decay mechanism acts as a natural, $\textit
קרא במקור המקורי