כתבה
arXiv cs.CL ·
QuantMLA: קוונטיזציה כפולת מסלול
QuantMLA: Function-Aligned Dual-Path Quantization for Low-Bit MLA KV Caching
QuantMLA היא שיטה לקוונטיזציה כפולת מסלול ל-MLA. היא מאפשרת אחסון INT4 משותף של קשי תוכן ו-RoPE עם פגיעה מינימלית בדיוק. QuantMLA משפרת ביצועים על גבי מודלים MLA שונים.
תקציר מקורי באנגליתarXiv:2609.36760v2 Announce Type: replace-cross Abstract: Multi-Head Latent Attention (MLA) enables expressive multi-head attention with compact caches for its content and decoupled RoPE paths, yet cache memory still scales linearly with context length and batch size. In this work, we establish a systematic model of MLA's dual-path quantization errors, characterizing their distinct effects on attention-output distortion and explaining the pronounced amplification of RoPE-path errors. Guided by this analysis, we introduce QuantMLA, a function-aligned framework for low-bit dual-path quantization. We derive path-specific transformation spaces that preserve full-precision computation while remaining fully fusible into model parameters offline, eliminating online transformation overhead. Within
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית