כתבה
arXiv cs.LG ·
PatchBench: Measuring Collateral Damage in Activation Patching
תקציר מקורי באנגליתarXiv:2610.10276v1 Announce Type: new Abstract: An LLM safety patch can pass a benchmark while still being a poor repair. This risk is especially acute for jailbreak repairs, where the goal is to correct a specific unsafe behaviour without changing unrelated behaviours. A patch may block exact evaluation prompts yet fail on close harmful variants, or suppress harmful behaviour by over-refusing benign prompts that share its wording or structure. Existing protocols primarily test whether models can be broken, while aggregate metrics (attack success, refusal rates, global capability) cannot distinguish selective repairs from broader local suppression. To address this gap, we introduce PatchBench, a benchmark of empirically observed model-specific jailbreak failures inducing actionable harmful
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית