כתבה
arXiv cs.AI ·
אבד בשיחה או אבד בתרגום? זיהוי ירידה ביצועים ב-RAG
Lost in Conversation or Lost in Translation? Diagnosing Multi-Turn Degradation in RAG
במחקר זה נחקרה ירידה ביצועים של מודלי RAG בשיחות רב-טור. נמצא כי רב-טור גורם לירידה בדיוק וביישור של המודלים.
תקציר מקורי באנגליתarXiv:2609.36700v1 Announce Type: cross Abstract: When conversing with large language models (LLMs), users often begin with a simple question and build towards a multi-hop question through follow-up turns. Retrieval-augmented generation (RAG) and its graph-based variant (GraphRAG) have become the dominant approaches for grounding LLM responses in external evidence, yet both are evaluated almost exclusively on single-turn, fully specified queries. We systematically investigate this evaluation mismatch through a large-scale simulation study. Building on prior work on multi-turn LLM evaluation, we transform questions from multi-hop question answering (QA) benchmarks into underspecified conversations and evaluate ten LLM assistants with eight retrieval systems across 1.5 million simulated conv
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית