יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

בדיקת עמידות אחזור של מודלי שפה גדולים

Evaluating the Retrieval Robustness of Large Language Models
חוקרים בדיקת עמידות אחזור של מודלי שפה גדולים. הם בודקים האם שיטות אחזור משפרות באופן עקבי את הביצועים. הניסויים כוללים 11 מודלים, ביניהם GPT, Qwen ו-Claude.
תקציר מקורי באנגליתarXiv:2505.21870v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) generally enhances large language models' (LLMs) ability to solve knowledge-intensive tasks. But RAG could also lead to performance degradation due to imperfect retrieval and the model's limited ability to leverage retrieved content. In this work, we evaluate the robustness of LLMs in practical RAG setups (henceforth retrieval robustness). We focus on three research questions: (1) whether RAG is always better than non-RAG; (2) whether more retrieved documents always lead to better performance; and (3) whether document order impacts results. To facilitate this study, we establish a benchmark of 1,891 samples spanning five datasets across three task categories, each with documents retrieved using b
קרא במקור המקורי