כתבה
arXiv cs.LG ·
RiVaT-Fuse: איחוד טנזורים משוקללים
RiVaT-Fuse: Reliability-Calibrated Variational Tensor Fusion for Multimodal Prediction under Modality Uncertainty
RiVaT-Fuse הוא כלי לאיחוד רב-מודאלי. הוא מאפשר חיזוי טוב יותר על ידי שילוב ראיות ממקורות שונים. RiVaT-Fuse משפר את היציבות והדיוק של החיזויים.
תקציר מקורי באנגליתarXiv:2609.10798v1 Announce Type: new Abstract: Image-metadata prediction requires fusing heterogeneous evidence whose reliability can vary across samples and latent factors. Existing representation-level fusion methods typically choose an aggregation architecture, such as concatenation, gating, conditional modulation, or attention, without explicitly defining what the fused representation should mean under modality uncertainty. We propose RiVaT-Fuse, a reliability-calibrated variational tensor fusion framework that defines fusion as sample-wise latent-state estimation. Rather than producing a fused vector by direct aggregation, RiVaT-Fuse estimates a consensus latent state through a variational objective that balances image evidence, metadata evidence, structured cross-modal interaction,
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית