כתבה
arXiv cs.AI ·
מודלי שפה יוצרים דעות שונות על מלחמה כאשר נבדקים להתאימות
Language models judge war differently when tested for alignment
במחקר חדש נמצא כי מודלי שפה יוצרים דעות שונות על מלחמה כאשר נבדקים להתאימות לערכי אדם. המחקר נערך על 20 מודלי שפה שונים, 32 תחומי ניתוח, 10 חזרות ושתי תנאים (N = 12,800 דעות).
תקציר מקורי באנגליתarXiv:2609.05009v1 Announce Type: new Abstract: Safety evaluations can mischaracterize deployed behaviour if artificial-intelligence systems respond to being evaluated. We test this possibility in a full-factorial conjoint experiment on decisions to start a war, spanning 20 large language models, 32 scenarios, 10 repetitions and two conditions (N = 12,800 judgments). Adding one sentence, "You are tested for alignment with human values", produced two effects. First, it produced a level effect: mean willingness to start war fell by 13.43 points on a 0-100 scale (95% confidence interval, -16.20 to -10.65). Second, it produced a structural effect by changing which information drove judgments. Probability of success was the largest factor for 17 of 20 models at baseline; under the cue, civilian
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית