כתבה
arXiv cs.CL ·
BadRAG: זיהוי פגיעויות בדגמי שפה גדולים
BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models
BadRAG הוא איום חדש שמאפשר לתוקף לכוון את התגובה של מודלי שפה גדולים על ידי החדרת מעברים מזיקים לבסיס הידע. התוקף יכול ליצור מעברים שיורטו רק כאשר מילים מסוימות מופיעות בשאילתות המשתמש. המחקר הראה כי החדרת 10 מעברים מזיקים בלבד יכולה להגיע לשיעור הצלחה של 98.2%.
תקציר מקורי באנגליתarXiv:2406.00083v3 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by retrieving relevant information from external knowledge bases to provide more accurate, contextually informed, and up-to-date responses. However, this reliance on external knowledge introduces significant security vulnerabilities, as many RAG systems (e.g., Google Search) rely on large and unsanitized data repositories (e.g., Reddit). In this paper, we unveil a novel threat in which attackers steer the RAG system's response by injecting malicious passages into its knowledge base. When a user's query contains attacker-specified trigger words, the RAG retrieves and refers to these malicious passages, enabling the attacker to steer the response without alteri
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית