יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

אם LLMs מאמינים בהאשמה או בהאשמה עצמה? מדידת שינויי האמונה ב-Werewolf

Do LLMs Trust the Accuser or the Accusation? Measuring Belief Shifts in Werewolf
במחקר חדש, נבדקו LLMs במשחק Werewolf כדי לבדוק את יכולתם לשנות את דעותיהם בעקבות האשמות. התוצאות הראו שה-LLMs עדיין נוטים להאמין בהאשמות, גם אם האשמה היא זו של כלב-עריק.
תקציר מקורי באנגליתarXiv:2609.12446v1 Announce Type: cross Abstract: Social-deduction games such as Werewolf are increasingly used to evaluate LLM agents, but existing evaluations often rely on final game outcomes. We propose a belief-shift evaluation benchmark in Werewolf for analyzing communication skills through belief updating. Using LLM-played games, we annotate suspicion and accusation messages and measure how an observing village-side model's beliefs change after each message. We evaluate 40 open-weight LLM configurations on 1,224 annotated messages. Our results show that larger models better distinguish true wolves from villagers based on game history, but accusations still strongly influence their beliefs. Models become more suspicious of the accused target and less suspicious of the accuser, especi
קרא במקור המקורי