כתבה
arXiv cs.LG ·
RoPE attention is an exact forward-pass gradient step with softmax intact
תקציר מקורי באנגליתarXiv:2609.06685v1 Announce Type: cross Abstract: We derive an exact gradient-step representation of the RoPE-softmax forward pass. For every deterministic RoPE-softmax attention head with arbitrary affine projection weights, we construct a query-dependent effective matrix $\Delta M_i$ satisfying $y_i = \mu_i + u_i^\top \Delta M_i$, where $\mu_i$ is the uniform mean of the attended values and $u_i$ is the augmented query input. The construction applies the classical exponential divided difference $\rho = \phi_1$ to retain the softmax exactly. Its positive coefficients give a unit gradient-step representation on a query-conditioned quadratic objective. The same function connects the RoPE generator to exact positional finite differences. We derive a tokenwise formula for the error of reusing
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית