כתבה
arXiv cs.CL ·
הערכת התאמה של נטיות התנהגות במודלים LLM
Evaluating Alignment of Behavioral Dispositions in LLMs
חוקרים פיתחו את STAR, כלי לבדיקת התאמה של נטיות התנהגות במודלים LLM. המחקר בדק 25 מודלים, כולל LLMs, ומצא פערים בהתאמה לנטיות התנהגות אנושיות. הנתונים והקוד זמינים לציבור.
תקציר מקורי באנגליתarXiv:2602.11328v2 Announce Type: replace Abstract: As people turn to LLMs for social advice, understanding their behavior in such contexts becomes essential. In this work, we focus on behavioral dispositions: the underlying tendencies that shape responses in social contexts. We introduce STAR, a framework for studying how closely the dispositions expressed by LLMs align with those of humans. STAR builds on established psychological questionnaires, adapting their items into realistic advice-seeking scenarios, as self-report may not transfer to actual advisory behavior. Using STAR, we construct a dataset of 23k scenarios, each validated by 3 raters and annotated with preferences from 10 participants. Across 25 LLMs, we find that (1) when human consensus is high, frontier models can fail to
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית