כתבה
arXiv cs.CL ·
PatchBench: מדידת נזק משני בתיקון פעילות
PatchBench: Measuring Collateral Damage in Activation Patching
PatchBench הוא בנק אמתי של 400 כשלונות פריצה מודל-ספציפיים. הוא בודק תיקונים בהתנהגות המודל ומודד נזק משני. המחקר מראה כי תיקונים יכולים לגרום לנזק משני משמעותי, גם אם הם עוברים בדיקות.
תקציר מקורי באנגליתarXiv:2610.10276v1 Announce Type: cross Abstract: An LLM safety patch can pass a benchmark while still being a poor repair. This risk is especially acute for jailbreak repairs, where the goal is to correct a specific unsafe behaviour without changing unrelated behaviours. A patch may block exact evaluation prompts yet fail on close harmful variants, or suppress harmful behaviour by over-refusing benign prompts that share its wording or structure. Existing protocols primarily test whether models can be broken, while aggregate metrics (attack success, refusal rates, global capability) cannot distinguish selective repairs from broader local suppression. To address this gap, we introduce PatchBench, a benchmark of empirically observed model-specific jailbreak failures inducing actionable harmf
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית