יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

DAMP: קיצור-מודע-מערכת-של-פסיק-מעורבת-לקיחת-מידע-מסוגמן-בעקב-פסיק-קבוע

DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization
DAMP: קיצור-מודע-מערכת-של-פסיק-מעורבת-לקיחת-מידע-מסוגמן-בעקב-פסיק-קבוע. פיתוח של מערכת של פסיק-מעורבת-לקיחת-מידע-מסוגמן-בעקב-פסיק-קבוע, שמשמשת לקיצור-מודע-מערכת-של-פסיק-מעורבת-לקיחת-מידע-מסוגמן-בעקב-פסיק-קבוע.
תקציר מקורי באנגליתarXiv:2608.27513v2 Announce Type: replace-cross Abstract: Complex reasoning and agentic applications increasingly rely on long-context inference, where growing KV caches increase both memory usage and decoding overhead. Hybrid models reduce these costs by combining Softmax Attention with Gated DeltaNet (GDN) or Kimi Delta Attention (KDA), which maintain fixed-size recurrent states. These states are commonly stored in FP32 and consume substantial GPU memory, while their updates are limited by memory bandwidth. Quantization can reduce both storage footprint and memory traffic, but we find that uniform INT8 and FP8 degrade complex reasoning accuracy, while INT4 and NVFP4 collapse it to near zero. To our knowledge, this is the first study of post-training recurrent-state quantization for GDN a
קרא במקור המקורי