יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

התקפות פרומפט נגד LLMs

From ASR to ASP: Evaluating Prompt Attack Vulnerabilities Against Open-Source LLMs
חוקרים בדקו התקפות פרומפט נגד 14 LLMs פתוחים ו-3 סגורים, ומצאו פגיעות במודלים כמו StableLM2 ו-Mistral. הם הציעו מדד ASP להערכת הצלחת ההתקפות.
תקציר מקורי באנגליתarXiv:2505.14368v3 Announce Type: replace-cross Abstract: Recent studies demonstrate that Large Language Models (LLMs) are vulnerable to attacks that generate harmful or sensitive outputs. As open-source LLMs are increasingly adopted in high-impact applications such as finance, law, and healthcare, systematically investigating their security risks is becoming increasingly important towards a trustworthy LLM era. This paper comprehensively studies effective prompt injection attacks against 14 widely used open-source and three closed-source LLMs on five attack benchmarks. Moreover, existing evaluation metrics mostly only consider the attack success rate, overlooking uncertainty in model responses. Our proposed Attack Success Probability (ASP) additionally captures uncertain behaviors for eva
קרא במקור המקורי