כתבה
arXiv cs.CL ·
LLM: החלטות מוסריות תחת לחץ
Principled Under Pressure: Post-Training Decides Whether LLMs Act on Their Own Moral Judgment
חוקרים בדקו כיצד מודלים שפה גדולים (LLM) מתנהגים תחת לחץ. הם השוו את התנהגותם של OLMo-3-7B-Instruct, Llama-3.1-8B-Instruct, Tulu 3 ו-Qwen2.5-7B-Instruct ב-248 תרחישים שונים. התוצאות הראו שהמודלים מתנהגים באופן שונה תחת לחץ, ושהדרך בה הם מאומנים משפיעה על התנהגותם.
תקציר מקורי באנגליתarXiv:2610.08670v1 Announce Type: cross Abstract: Language models increasingly act as agents. An agent that says an action is wrong and then takes it anyway is a different failure from one that does not know better, and evaluations of stated values cannot see it. We build a pre-registered panel of 248 scenarios across five kinds of pressure. Each scenario is posed twice to the same model, once as the agent choosing what to do and once in the third person asking which option is right, so the model's own judgment is the reference. Every scenario has a twin with the pressure removed, and every model gets a positive control in which its operator orders the violating action, so that a missing gap can be told apart from a blind instrument. On OLMo-3-7B-Instruct, the model takes the action it jud
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית