יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

הגנה מפני סכנות: רגישות הצגה תקיפה בבניית-אב-בטיחות

Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks
במאמר זה, חוקרים חוקרים את השפעת ההצגה על תוצאות בניית-אב-בטיחות ל-LLM. הם מציגים תרחיש-שמירה-הצגה-רגישה (TPRS) כדי לבדוק את השפעת ההצגה על תוצאות הבניית-אב. התוצאות מציגות כי תוצאות הבניית-אב תלויות בהצגה ולא רק במודל.
תקציר מקורי באנגליתarXiv:2610.03585v1 Announce Type: cross Abstract: Security benchmarks for LLM-based agents often report the attack success rate (ASR) as a measure of model robustness and use these scores to compare different models and defense mechanisms, assuming that they describe the security of the agent. In this paper, we explore whether it also influences the benchmark's measurement. To measure the effect of the benchmark representation, we introduce threat-preserving representation sensitivity (TPRS), which measures how much the ASR changes when we change the agent-visible representation while holding the underlying task, harmful action, security policy, ground truth, environment, and the evaluation criteria fixed. On Agent Security Bench (ASB), replacing threat-related tool names with threat-neutr
קרא במקור המקורי