יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

Rethinking RAG in Long Videos: What to Retrieve and How to Use It?

תקציר מקורי באנגליתarXiv:2606.13141v2 Announce Type: replace Abstract: Retrieval-augmented generation is extending beyond text to long videos, where query-relevant chunks can be represented across multiple modalities and temporal granularities. Progress in this setting, VideoRAG, is limited by two gaps: existing benchmarks allow queries to be answered without the video, obscuring retrieval errors, and prior methods apply a single modality-granularity configuration per query, ignoring chunk-level variability. We address both by introducing V-RAGBench, a benchmark of $\langle$query, evidence chunk, answer$\rangle$ triplets over hour-scale videos, in which each answer depends on a unique evidence chunk, enabling decoupled evaluation of retrieval and generation, and CARVE, a training-free method that runs parall
קרא במקור המקורי