כתבה
arXiv cs.LG ·
Mutual Equilibrium: Multimodal Representation Learning through Reciprocal Feedback
תקציר מקורי באנגליתarXiv:2609.39456v1 Announce Type: new Abstract: This work proposes a mutual feedback architecture, MEQ, that refines the two inputs, of possibly different modalities, into a pair of coupled embeddings such that each embedding reflects the information of the other. The core idea is to incorporate continuous interchange of information between the two inputs. This idea leads to a mutual feedback architecture consisting of two components whose outputs are fed back into the other. The final output of this model is defined as the fixed point of this interaction. We provide theoretical analysis that offers interpretation of this model as well as design choices to prevent failure cases. We show the benefits of MEQ through classification and visual grounding tasks spanning various datasets. Quantit
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית