כתבה
arXiv cs.LG ·
Why shared attention vectors fail: a case for outcome-indexed tuning
תקציר מקורי באנגליתarXiv:2609.08615v1 Announce Type: new Abstract: Dimensional attention in learning is often implemented as a globally shared attention vector, where each stimulus dimension corresponds to a single scalar. These scalars are learned by models through gradient-descent on error, where predictive features acquire more salience. We show that under multi-outcome learning, where models predict more than one outcome, this shared vector becomes unstable; it collapses to its bounds and prevents the models from learning meaningful attentional tunings for learning and generalization. We address this by introducing an outcome-indexed attentional matrix that converts globally shared attentional tuning into an outcome-indexed representation. We present an analysis of the unstable shared vectors and derive
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית