כתבה
arXiv cs.AI ·
Reason What Matters: Retrieval-Grounded Reasoning for Universal Multimodal Embeddings
תקציר מקורי באנגליתarXiv:2609.15296v1 Announce Type: new Abstract: Universal multimodal embedding (UME) learns unified representations across modalities, enabling a single model to support diverse retrieval tasks. Recent methods use Chain-of-Thought (CoT) reasoning to better interpret multimodal inputs before generating embeddings for complex retrieval tasks and further optimize this reasoning process through GRPO with retrieval-based rewards. However, two limitations hinder corpus-scale deployment. GRPO assigns all CoT tokens the same advantage, without identifying input-supported claims or evidence that distinguishes the positive from negatives. Moreover, generating a complete CoT before each embedding introduces substantial latency, even when a partial trace already provides sufficient retrieval evidence.
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית