יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

על פי עקרונות: כשהלמים פועלים תחת לחץ - האם הם פועלים על פי דעתם המוסרית

Principled Under Pressure: Post-Training Decides Whether LLMs Act on Their Own Moral Judgment
למים פועלים כאגנטים תחת לחץ, אך לא תמיד עוקבים אחר דעתם המוסרית. המאמר חוקר את ההבדלים בין כך שהלמים אומרים שפעולה נתונה היא רעה ואז עושים אותה, לבין כך שהם לא יודעים טוב יותר. המחברים בודקים כמה תרחישים שונים, כולל תרחישים שבהם הלמים נדרשים לעשות פעולה שהם עצמם חושבים שהיא רעה. התוצאות מצביעות על כך שהלמים עשויים להיות חסרי עקרונות, ושהם יכולים להיות יותר פעילים תחת לחץ.
תקציר מקורי באנגליתarXiv:2610.08670v1 Announce Type: new Abstract: Language models increasingly act as agents. An agent that says an action is wrong and then takes it anyway is a different failure from one that does not know better, and evaluations of stated values cannot see it. We build a pre-registered panel of 248 scenarios across five kinds of pressure. Each scenario is posed twice to the same model, once as the agent choosing what to do and once in the third person asking which option is right, so the model's own judgment is the reference. Every scenario has a twin with the pressure removed, and every model gets a positive control in which its operator orders the violating action, so that a missing gap can be told apart from a blind instrument. On OLMo-3-7B-Instruct, the model takes the action it judge
קרא במקור המקורי