כתבה
arXiv cs.CL ·
QuantMLA: פונקציות-מאולתרות דו-נתיביות להצפנה נמוכת-ביט ל-MLA KV Caching
QuantMLA: Function-Aligned Dual-Path Quantization for Low-Bit MLA KV Caching
QuantMLA היא פלטפורמה של דו-נתיביות להצפנה נמוכת-ביט של MLA KV Caching. היא מאפשרת הצפנה נמוכת-ביט של נתיבי התכנות וה-RoPE ב-MLA KV Caching.
תקציר מקורי באנגליתarXiv:2609.36760v1 Announce Type: cross Abstract: Multi-Head Latent Attention (MLA) enables expressive multi-head attention with compact caches for its content and decoupled RoPE paths, yet cache memory still scales linearly with context length and batch size. In this work, we establish a systematic model of MLA's dual-path quantization errors, characterizing their distinct effects on attention-output distortion and explaining the pronounced amplification of RoPE-path errors. Guided by this analysis, we introduce QuantMLA, a function-aligned framework for low-bit dual-path quantization. We derive path-specific transformation spaces that preserve full-precision computation while remaining fully fusible into model parameters offline, eliminating online transformation overhead. Within these s
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית