כתבה
arXiv cs.LG ·
הגבלת המודל, חסרת המערכת: מדידה ואחריות בניהול AI התקפי
Restricting the Model, Missing the System: Measurement and Accountability in Offensive AI Governance
מדידת יכולת ה-AI ההתקפית נוטה להגזים בנזק ולהפליא למודל, ולא למערכת כולה.
תקציר מקורי באנגליתarXiv:2605.09504v3 Announce Type: replace-cross Abstract: We show that the instruments used to measure AI offensive capability fail in two ways: (i) they overstate harm, and (ii) they credit the model with capability that belongs to the surrounding system. We argue that restricting access to a model is therefore necessary but not sufficient and that policy and procurement also need system-level, harm-grounded capability assessment. In June 2026, two frontier models were suspended under US export controls, reportedly prompted by a jailbreak that asked a model to read a codebase and fix its flaws. This finding measured an elicitation \emph{system} of model, prompt, and task. We support our argument with a study of an open-source framework in which lightweight large language model (LLM) agent
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית