יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

במסתור: בחנות ניידת: בדיקת בטיחות סוכני LLM נגד התקפות פירוק עם DECOMPBENCH

Hidden in Plain Sight: Benchmarking Agent Safety Against Decomposition Attacks with DECOMPBENCH
בדיקת בטיחות סוכני LLM נגד התקפות פירוק, שבהן תפקוד רע נפרק למשימות פשוטות שאינן נתפסות כמסוכנות.
תקציר מקורי באנגליתarXiv:2606.13994v2 Announce Type: replace-cross Abstract: LLM-based Agents are becoming increasingly capable and widely deployed, creating growing incentives for adversarial misuse in the real-world. A key emerging threat is Decomposition Attacks \cite{glukhov2024breach, jones2024adversaries} in which a harmful task is broken into simpler, benign subtasks that evade safety mechanisms when executed separately but cumulatively fulfill the malicious intent. Although recent benchmarks assess agent safety in multi-turn and multi-tool-use settings, they do not explicitly capture this form of decompositional misuse and may not represent realistic adversarial execution flows. To this end, we introduce DeCompBench, a benchmark designed specifically to evaluate agentic safety under decomposition att
קרא במקור המקורי