כתבה
arXiv cs.AI ·
Behind Harmful Compliance: Behavioral and Mechanistic Divergence Across LLM Jailbreaks
תקציר מקורי באנגליתarXiv:2604.18510v2 Announce Type: replace-cross Abstract: Open-weight language models can be rendered unsafe through several parameter-level interventions, yet models with matched harmful compliance can exhibit fundamentally different failure modes. We compare harmful supervised fine-tuning (SFT), harmful reinforcement learning with verifiable rewards (RLVR), and refusal-feature abliteration in Qwen2.5-7B and Llama-3.1-8B using harmfulness, capability, self-audit, safety reflection, representation similarity, and refusal-direction repair. All three routes reach near-ceiling harmfulness, but SFT causes the broadest capability loss and representational drift; abliteration yields localized, family-dependent refusal-feature suppression; and RLVR largely preserves base-model capability, explici
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית