כתבה
arXiv cs.CL ·
LinearARD: Linear-Memory Attention Distillation for RoPE Restoration
תקציר מקורי באנגליתarXiv:2604.00004v2 Announce Type: replace Abstract: The extension of context windows in Large Language Models is typically facilitated by scaling positional encodings followed by lightweight Continual Pre-Training (CPT). While effective for processing long sequences, this paradigm often disrupts original model capabilities, leading to performance degradation on standard short-text benchmarks. We propose LinearARD, a self-distillation method that restores Rotary Position Embeddings (RoPE)-scaled students through attention-structure consistency with a frozen native-RoPE teacher. Rather than matching opaque hidden states, LinearARD aligns the row-wise distributions of dense $Q/Q$, $K/K$, and $V/V$ self-relation matrices to directly supervise attention dynamics. To overcome the quadratic memor
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית