יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

גשר בין הקרבה המשמעותית להקרבה היישומית ב-RAG המודלי המרובד

Bridging the Semantic-Utility Gap in Multimodal RAG via Generator-in-the-Loop Alignment
במאמר זה, נציג פרקטיקה חדשה לצמצום הפער בין הקרבה משמעותית להקרבה יישומית ב-RAG המודלי המרובד. הפרקטיקה, המכונה 'גשר בין הקרבה המשמעותית להקרבה היישומית', משתמשת באלגוריתם של הסתברות כדי לצמצם את הפער. הפרקטיקה נבחנה על ידי ניסויים שהראו כי היא יעילה יותר מאשר פרקטיקות אחרות. המאמר כולל גם דיון בהיבטים התאורטיים של הפרקטיקה.
תקציר מקורי באנגליתarXiv:2609.08188v1 Announce Type: new Abstract: Vision-language models (VLMs) augmented with retrieval-augmented generation (RAG) benefit from access to external evidence. However, standard retrievers and rerankers optimize for semantic similarity rather than answer utility, creating a preference gap: documents that appear relevant may not help the generator produce a correct answer. Motivated by this, we propose a two-stage generator-in-the-loop alignment framework that closes this gap without human document-level relevance annotations. Our framework consists of two stages: in Stage 1, a VLM generates a hypothetical text passage from the image-query pair, which is used as the retrieval query for dense text search, bridging the image-to-text modality gap. In Stage 2, a cross-encoder rerank
קרא במקור המקורי