כתבה
arXiv cs.AI ·
Skin-Deep: ניתוח גאומטרי לאבחון רגישות להתפרצות בייצוגי מודלי שפה גדולים
Skin-Deep: A Geometric Diagnostic for Alignment Fragility in Large Language Model Representations
ניתוח חדש לאבחון רגישות להתפרצות בייצוגי מודלי שפה גדולים. הניתוח, שנקרא SKIN-DEEP, חוקר את הפעילות הנשארת של המודל באמצעות ניתוח גאומטרי. התוצאות מצביעות על כך שהמודלים הגדולים עשויים להיות רגישים לשינויים באופן שלא ניתן לזיהוי.
תקציר מקורי באנגליתarXiv:2606.22676v2 Announce Type: replace Abstract: Refusal on a safety benchmark does not reveal how stable that behavior will remain after model updates. Benign downstream fine-tuning can weaken refusal, yet behavioral evaluations typically expose this fragility only after an intervention. We introduce SKIN-DEEP, a geometric diagnostic that examines the unmodified model's residual-stream activations. It compares aligned and base checkpoints to identify safety-separating directions, tests their behavioral relevance through ablation, and summarizes the layer-wise pattern in the Geometric Fragility Score (GFS). Across twenty-one instruction-tuned models, harmful requests and benign instructions exhibit a recurring low-rank separation pattern. Selected direction ablations weaken refusal, wit
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית