כתבה
arXiv cs.LG ·
Attention Kernels for Learning Maps Between Heavy-Tailed Measures
תקציר מקורי באנגליתarXiv:2610.00564v1 Announce Type: new Abstract: Operator learning on probability measures can be accomplished with transformers. For measures with polynomial tails, the exponential weighting in softmax can make the corresponding measure-level attention integrals diverge. This motivates replacing the exponential with slower-growing functions. We construct two benchmarks for operator learning on measures with closed-form targets. We use these benchmarks to study attention kernel growth and data transformation in post-norm transformers. Without data transformation, the softmax models exhibit ensemble collapse on both heavy-tailed benchmarks, while the three slower-growing kernels avoid collapse. Symlog preprocessing allows softmax to avoid collapse on the matrix inverse task but not on the sh
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית