כתבה
arXiv cs.LG ·
Multimodal CoLRAG-TF: שיטה חדשה לאיחזור מידע מ-PDFs מורכבים
Multimodal CoLRAG-TF: Triple-Filtered Retrieval for Complex PDFs
חוקרים פיתחו את Multimodal CoLRAG-TF, שיטה לאיחזור מידע מ-PDFs מורכבים. השיטה משלבת טכניקות שונות, כולל עיבוד תמונות וטקסט, כדי לשפר את האיחזור. השיטה נבדקה על אוסף של PDFs והראתה תוצאות משופרות.
תקציר מקורי באנגליתarXiv:2607.20517v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) over heterogeneous PDF collections remains challenging due to multimodal content, domain-specific terminology, and the need for multi-hop reasoning across dispersed evidence. We present Multimodal CoLRAG-TF, a four-axis fusion architecture that integrates dense text embeddings, BM25 keyword matching, knowledge-graph triple filtering, and image-based similarity for robust retrieval over complex documents. Our system constructs a multimodal index of 2,403 blocks extracted from 43 Japanese disaster lesson PDFs, supported by a hybrid OCR pipeline and LLM-based caption generation. To enhance compositional reasoning, we extract 11,414 OpenIE triples and index them with FAISS, enabling sub-second triple lookup an
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית