יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

האילוזיה של הפקודה: מדוע פקודות המערכת משנות את החישוב בדגמי שפה

The System Prompt Illusion: How Instruction Preambles Modify Computation in Language Models
פקודות המערכת של דגמי שפה משנות את החישוב באופן שלא צפוי. חוקרים גילו שפקודות שונות יוצרות שינויים שונים בדגמי שפה, ושאלה האם זה יכול להיות חלק מהסיבה לחולשה בבטיחות של דגמי שפה.
תקציר מקורי באנגליתarXiv:2609.38205v1 Announce Type: cross Abstract: System prompts are the primary lever practitioners use to control language model behavior, yet what they actually do to the computation inside the transformer remains poorly understood. Across 17 instruction-tuned models spanning 8 architecture families and 1.5B to 72B parameters, we use Centered Kernel Alignment (CKA) to compare layer-wise representations under 20 system prompts in five functional categories. Effects are layer-selective and instruction-type-dependent: persona and formatting instructions deeply restructure intermediate representations, while safety instructions barely move them, producing changes statistically indistinguishable from a minimal baseline. Restrictive safety instructions and explicitly permissive ones ("you hav
קרא במקור המקורי