כתבה
arXiv cs.LG ·
RoPE at the End of Its Rope? Theory, Diagnosis, and Mitigation of Long-Context Failures
תקציר מקורי באנגליתarXiv:2609.39929v1 Announce Type: new Abstract: Long-context failures of RoPE-based language models can arise from RoPE's intrinsic tradeoff between maintaining stable token preferences and distinguishing nearby positions. Determining which weakness to address, and how, requires a more precise characterization of RoPE's behavior in trained models across context lengths. We address a key limitation of prior theory by allowing unequal query-key scales across RoPE frequencies, which aligns well with practical empirical observations. Our theory makes both vulnerabilities measurable for individual heads and inputs, and quantifies how high-frequency components support positional sensitivity while potentially disrupting semantic stability. We also derive a theoretical context-length bound beyond
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית