כתבה
arXiv cs.AI ·
RAG-PIBench: מבחן ספירה לזיהוי פצצות-הזרע במערכות RAG
RAG-PIBench: A Leakage-Aware Benchmark for Prompt-Injection Detection in Trustworthy RAG Systems
נחשפה מערכת מבחן חדשה לזיהוי פצצות-הזרע במערכות RAG. המבחן, RAG-PIBench, כולל 4,876 דוגמאות תחומיות והציג תוצאות טובות למערכות שונות.
תקציר מקורי באנגליתarXiv:2610.08571v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems are vulnerable to prompt-injection attacks embedded in retrieved content. We introduce RAG-PIBench, a benchmark for RAG-style prompt-injection detection containing 4,876 contextual examples across frozen train, validation, and protected-test splits. Using a leakage-aware construction pipeline and strict evaluation protocol, we compare keyword-based, semantic-reference, TF-IDF, and transformer-based detectors. DistilBERT achieves the best protected-test performance (F1 = 0.896, PR-AUC = 0.968), while TF-IDF SVM and logistic regression remain competitive. Our results demonstrate the value of leakage-aware benchmark design and strong sparse baselines for reliable prompt-injection detection in RAG sy
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית