כתבה
arXiv cs.CL ·
RAG: מדידת עמידות
In RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Document Poisoning
חוקרים בדקו כיצד מודל Llama 3.1 8B מגיב להרעלת טקסטים מושגים. הדיוק ירד מ-77.9% ל-43.5% כאשר כל הטקסטים המושגים הורעלו.
תקציר מקורי באנגליתarXiv:2609.09243v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) grounds a language model in retrieved documents, which reduces hallucination but creates a new attack surface: if retrieved text is tampered with, the model may repeat the falsehood. We study how much a small quantized model, Llama 3.1 8B, degrades when a fraction of its retrieved context is poisoned. Three corruption strategies are tested, entity swap, number swap, and negation, each applied to zero, one, two, or three of the three retrieved passages, over a factorial sweep of 588 runs on a fact-checking task built from FEVER. Accuracy falls from 77.9% on clean context to 43.5% when all three passages are corrupted. Entity swap flips the largest share of answers that were correct on clean context. Numbe
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית