כתבה
arXiv cs.AI ·
סירוב קורא רק חלק ממה שהמודל יודע
Refusal Reads Only a Slice of What the Model Knows: Harm-Keyed Routing and Its Exceptions Across Model Families
חוקרים בדקו כיצד מודלים שונים, כולל Llama ו-Qwen, קובעים החלטות סירוב. התוצאות מראות שהמודלים קוראים רק חלק מהידע שלהם, ושההחלטות מבוססות על גורמים שונים.
תקציר מקורי באנגליתarXiv:2609.14759v1 Announce Type: cross Abstract: Alignment applied after pretraining is shallow in a measurable way: a single direction in a model's residual stream can be edited out, and the model stops refusing harmful requests. That fact says how easily refusal can be removed, not what the refusal decision was reading in the first place. We ask what it reads, and we separate that from what the model comprehends. Across four open-weight models spanning three families, moral comprehension is native to pretraining: a low-rank moral subspace crystallizes during pretraining, and alignment rotates it once without rebuilding it. The refusal gate, in contrast, is a fresh post-training construction with only a weak pretraining precursor, written into a narrow control-token channel where the ref
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית