כתבה
arXiv cs.CL ·
עמידות בטיחות בקנה מידה גדול
How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE
חוקרים בדקו עמידות בטיחות במודל GLM-5.3-Flash בן 320B פרמטרים. התקפת כיוון יחיד הצליחה לשחזר את היכולת לסרב, אך רק כאשר נערכו עריכות במספר רכיבים.
תקציר מקורי באנגליתarXiv:2609.09793v1 Announce Type: cross Abstract: Directional ablation removes an aligned language model's ability to refuse by projecting a single "refusal direction" out of the weights that write the residual stream. It needs no gradient-based training and no optimization, only a few hundred contrastive prompts, which makes it the canonical white-box attack on open-weight alignment. However, it has been established only on dense models up to roughly 70B parameters. We study whether it survives the shift to frontier mixture-of-experts (MoE) models whose residual streams are no longer a single tensor and whose weights ship quantized. We apply it to GLM-5.3-Flash (320B parameters, 288 routed experts, a four-wide hyper-connection residual, block-FP8). The attack survives the architecture, bu
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית