יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

ניתוח הטעיה הגנתית נגד תקיפות אוטומטיות מונחות מודל

Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems
מחקר זה בוחן את היעילות של הטעיה הגנתית נגד תקיפות אוטומטיות מונחות מודל על מערכות AI אגנטיות. המחקר מציג אלגוריתם להטעיה הגנתית הנקרא Contextual Misdirection via Progressive Engagement (CMPE), שמטרתו להפחית את הצלחת התוקף. הניסויים הראו ירידה משמעותית בשיעור ההצלחה של התוקף.
תקציר מקורי באנגליתarXiv:2606.20470v4 Announce Type: replace-cross Abstract: Agentic AI systems increasingly rely on language-model components to interpret instructions, process external data, invoke tools, and coordinate with other agents. These capabilities make prompt-injection and jailbreak attacks more consequential, especially as attackers adopt model-guided automation to scale probing, prompt refinement, and response evaluation. This work analyzes the resulting attack-defense setting through a probabilistic model of a target system, its defense mechanism, and the attacker's automated judge. Our analysis shows that conventional detect-and-block defenses can allow attacker success rate (ASR) to approach one as the query budget grows, since predictable refusals provide useful feedback to automated search
קרא במקור המקורי