כתבה
arXiv cs.LG ·
גאומטריה של דחייה: למה סיכון אחרי-כתיבה הוא רך והסיכון בזמן הכשרה נשאר
The Geometry of Refusal: Why Post-Hoc Safety Is Fragile and Pretraining-Time Safety Persists
הסיכון האחרי-כתיבה (RLHF) הוא רך וניתן להפריע על ידי התאמות תפריט. חוקרים חשפו זאת במאמרם, שבו הם חקרו את יציבות הסיכון במודלי LLM. הם גילו שהסיכון האחרי-כתיבה נותן תגובה דלה ומפוצלת, ולא יכול למחוק יכולות קיימות. בעקבות זאת, הם הציעו שיטה חדשה לסיכון, הכוללת סיכון בזמן הכשרה.
תקציר מקורי באנגליתarXiv:2609.06934v2 Announce Type: replace Abstract: Post-hoc safety training (RLHF, DPO) is the dominant way to align large language models, yet jailbreaks (Zou et al., 2023), fine-tuning attacks (Qi et al., 2024), and activation-space edits (Arditi et al., 2024) keep recovering the behaviors it was meant to remove. We give this fragility one geometric explanation and follow it into pretraining. We measure the safety update $\Delta = W_{safe} - W_{base}$ against the curvature of the model's capabilities (the empirical Fisher of a capability loss). Across five model families, post-hoc safety lands in a suppression regime: $\Delta$ is nearly orthogonal to the capability directions, and its small in-subspace part concentrates on a few high-curvature ones. The update is thin but sharp, a refus
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית