יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

התקפת Option-Channel על מודלים מסוגננים

One Word Opens the Gate: The Option-Channel Attack on Typed Decision Models as Agent Guardrails
חוקרים גילו פרצת אבטחה במודלים מסוגננים המשמשים כשומרי גדר במערכות אוטונומיות. ההתקפה, הנקראת Option-Channel, מאפשרת לתוקף לעקוף את המודל ולבצע פעולות לא מורשות. המחקר בדק שבעה מודלים פתוחים ומצא כי הם רגישים להתקפה.
תקציר מקורי באנגליתarXiv:2610.12292v1 Announce Type: new Abstract: A typed decision model reads a piece of text and returns a probability over caller-defined options, each with a short written definition, generating no text. Recent work places these models in agent systems as guardrails: the component that reads a proposed tool call or incoming message and decides whether to allow it. We evaluate seven open-weight models in that role and report the two error directions separately: a fail-open error allows a prohibited action and is a vulnerability; a fail-closed error blocks a permitted one and is only a cost. On prompt-injection, jailbreak and toxic-content screening, accuracy at the allow-or-block decision ranges from 36% to 72% against a chance level of 50%. A low error rate in one direction only reflects
קרא במקור המקורי