יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

Kill-Chain Canaries: עקבות תקיפה בשלבי הפצה של פלאשבאקים וחמש מודלי LLM

Kill-Chain Canaries: Stage-Level Tracking of Prompt Injection Across Attack Surfaces and Five Production LLMs
במאמר זה, המחברים מציגים שיטה לעקוב אחר פלאשבאקים בשלבי הפצה, ובודקים את רגישותם של חמש מודלי LLM. הם מצאו שהמודלים השונים רגישים במידה שונה לפלאשבאקים, ושיטתם יכולה לשמש ככלי לפיקוח על רגישותם של המודלים.
תקציר מקורי באנגליתarXiv:2603.28013v4 Announce Type: replace-cross Abstract: Multi-agent LLM systems now read documents, web pages and tool results on behalf of users, yet their resistance to prompt injection is usually reported as one number: did the attack succeed? We introduce a kill-chain canary method that plants a unique token in every injected payload and records the furthest of four stages it reaches (Exposed -> Persisted -> Relayed -> Executed), across 950 runs, five production LLMs, six attack surfaces, and five defense conditions. Exposure was 100% among runs that called the tool; the outcomes differ downstream. Claude Haiku 4.5 and Claude Sonnet 4.5 executed none of their 164 text-surface attacks, and in the text relay the canary token never appeared in a memory write (0/40); GPT-4o-mini executed
קרא במקור המקורי