יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

פרדוקס היכולת: איך אודיטורים חכמים יותר הופכים מערכות רב-סוכנים לפחות בטוחות

The Capability Paradox: How Smarter Auditors Make Multi-Agent Systems Less Secure
מחקר מציג פרדוקס בו אודיטורים חכמים יותר מעלים את רמת ההתקפות המוצלחות במערכות רב-סוכנים. התוצאות מראות כי הגדלת יכולת הסוכנים יכולה לפגוע בביטחון המערכת. המחקר מציע פתרון באמצעות אימות הסוכנים באופן א-סימטרי.
תקציר מקורי באנגליתarXiv:2605.17480v3 Announce Type: replace Abstract: Multi-agent systems extend large language models (LLMs) by decomposing tasks among specialized agents, but their distributed decision process creates new attack surfaces. We identify semantic hijacking, an attack in which harmful requests are concealed within domain-specific narratives and propagated to a Manager through Worker reports, without any syntactic injection primitives. Across 42,000 adversarial trials over 12 Manager models and 7 Worker configurations, we uncover a capability paradox: as Worker capability increases, the mean system-level Attack Success Rate (ASR) increases from 18.4% to 63.9%, peaking at 94.4%. To explain this effect, we conduct multi-level mediation analysis on two independent datasets (47,807 interactions). T
קרא במקור המקורי