כתבה
arXiv cs.AI ·
ביטחון סביר: חקירה במודל סטקלברג של ביטחון AI
Defensive Sufficiency in a Stackelberg Model of AI Security
במאמר זה, החוקרים חקרו את הביטחון של מערכות AI בעזרת מודל סטקלברג. הם חקרו את השימוש במידע שנאסף מבדיקות אוטומטיות, טיפול צבאי אנושי ותגובה לאסונות. הם גם חקרו את השימוש בתקיפה ובתיקון.
תקציר מקורי באנגליתarXiv:2610.09892v2 Announce Type: replace-cross Abstract: Feedback from automated testing, human red teaming, and incident response can strengthen an AI system's defenses when discovered failures lead to effective repairs. We study when this feedback process provides sufficient protection and when investing in it is economically worthwhile. We begin by showing that an attack surface composed of finite number of inputs is defended with probability 1 if every unresolved attack has a persistent chance of discovery, repairs are effective, and subsequent updates preserve earlier protection. We derive completion-time bounds and extend the analysis to growing attack surfaces, repairs that generalize across related attacks, and multiple discovery mechanisms. These results distinguish eventual prot
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית