כתבה
arXiv cs.AI ·
MoEMB: Scaling Universal Multimodal Embeddings with Efficient Mixture-of-Experts Models
תקציר מקורי באנגליתarXiv:2609.08663v1 Announce Type: cross Abstract: Universal multimodal embedding (UME) increasingly demands encoder's capacity for handling a broad range of tasks and modalities with increased complexity. Prior scaling methods either increase the representation size, retrieval effort, or scales the encoder into a heavy multimodal LLM. Recent works, such as Think-Then-Embed (TTE), explore scaling via reasoning tokens. However, embedding models are hard to scale up: increasing parameters directly tradeoffs for the large training batch size that contrastive learning needs, and retrieval has to be served under tight latency. Moreover, UME tasks are diverse in complexity, where scaling up embedders can bring significant redundant computation. In this work, we propose MOEMB, which instead scales
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית