כתבה
arXiv cs.LG ·
Learned Queries and Keys Are All You Need: Replacing the Value Projection with Structured Transforms
תקציר מקורי באנגליתarXiv:2609.36698v1 Announce Type: new Abstract: To reduce the number of parameters and cache memory requirements of transformers we introduce dual-headed transformers instead of three heads. We studied Walsh-Hadamard Transform (WHT), Discrete Cosine Transform (DCT), Discrete Fourier Transform, filterbank based Shearlet Transform, and Multiplication-Avoiding (MA) operators to construct dual heads. We combine spatial patches and their orthogonal transforms (or Shearlet and MA operators) in a structure similar to the attention block. We obtained better results than triple headed transformers in ImageNet. Extensive simulation examples are presented.
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית