יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

סירוב מקומי, נזק מועבר: שכבות בטיחות תחת עדינות דגימות

Refusal Localizes, the Damage Relocates: Safety Layers Under Few-Sample Fine-Tuning
חוקרים בדקו את השפעת דגימות מוגבלות על מודלי LLM. הם מצאו כי שכבות בטיחות מסוימות יכולות לשמור על סירוב לבקשות מזיקות, אפילו כאשר המודל מאולץ ללמוד מדגימות מוגבלות. החוקרים השתמשו במודל Llama-3.1-8B כדי לבדוק את השיטה.
תקציר מקורי באנגליתarXiv:2610.00320v1 Announce Type: cross Abstract: Fine-tuning adapts aligned large language models (LLMs) to downstream tasks, but a few dozen harmful examples can remove their refusal of harmful requests. Prior work localizes safety-related behavior to specific layers, directions, and tokens, suggesting targets for protection. We test whether successful localization and recovery support defenses that survive changes in the attack. Across six checkpoints from four model families, harmful and benign prompts remain linearly separable after attack, and patching full clean hidden states into the compromised model restores refusal at a reproducible transition depth. Building on a prior layer-freezing defense, we freeze every layer up to this depth and repeat the attack. At a hundred harmful exa
קרא במקור המקורי