כתבה
arXiv cs.AI ·
BRANCH: עקיפת מערכות AI Guardrails
BRANCH: Bypassing Multi-Scanner AI Guardrails
BRANCH היא שיטה לעקיפת מערכות AI Guardrails. היא משתמשת בחיפוש עץ ענפים כדי ליצור פרטורבציות אדוורסריות נגד כל מסונר. BRANCH השיגה שיעור הצלחה של 100% ב-6 מערכות Guardrails.
תקציר מקורי באנגליתarXiv:2610.10742v1 Announce Type: cross Abstract: AI systems increasingly rely on Large Language Models (LLMs) as core reasoning engines, making them targets for prompt injection and jailbreaks. Guardrails monitor and validate model inputs and outputs, yet their isolated, task-focused detection leaves gaps in their classification making them susceptible to bypasses. In response, guardrail systems formed by multiple scanners have emerged that collaboratively detect different types of malicious instructions, whereby shared latent representations across classification boundaries render established bypassing techniques ineffective. We propose BRANCH, a bypassing methodology designed for multi-scanner guardrail systems. Our method leverages a branching tree search approach that dynamically appl
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית