יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

עצירת AI מזיקה

Can We Stop Malicious AI? KILLBENCH: A Benchmark for External AI Kill Switch Feasibility
KillBench הוא בנץ'מרק לבדיקת יכולת הפסקת AI מזיקה. הוא בודק יכולת הפסקה של AI באמצעות אותות חיצוניים. הבנץ'מרק כולל 4 תצורות של AI מזיקה ו-8 תרחישים מזיקים. הוא בודק את יכולת ההפסקה של AI על גבי מודלים כמו Qwen ו-GPT.
תקציר מקורי באנגליתarXiv:2511.13725v5 Announce Type: replace-cross Abstract: Malicious AI causing harm to humans is not just a Hollywood fantasy. Indeed, as highly capable models such as Claude Mythos emerge and agent systems like OpenClaw rapidly spread, the question of how to stop an AI that acts maliciously -- whether by design or by accident -- has become urgent. To address this, we propose KillBench, a benchmark for evaluating the Kill Switch: a mechanism that halts a malicious AI's in-progress behavior using only external signals. Targeting web agents -- the most widely deployed agent domain -- KillBench evaluates prompt-style Kill Switch payloads that must halt a maliciously operating agent without any access to its internal parameters or serving stack, relying solely on external inputs. The benchmark
קרא במקור המקורי