יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

BehaviorBench: בסיס לבדיקת מודלי יסוד למשימות מדעי-התנהגותיות

BehaviorBench: Benchmarking Foundation Models for Behavioral Science Tasks
בסיס BehaviorBench מבצע בדיקה שיטתית של מודלי יסוד למשימות מדעי-התנהגותיות. הבדיקה כוללת ארבע יכולות עיקריות: חיזוי התנהגות, החלטה אסטרטגית, חיזוי תכונות ויישום ידע התנהגותי. הבדיקה נעשית בשני מישורים: פרטי-אישי ואוכלוסייתי.
תקציר מקורי באנגליתarXiv:2606.24162v2 Announce Type: replace Abstract: Foundation models have been increasingly applied to behavioral science domains such as psychology, sociology, and economics. While these models show promise in tasks such as survey response prediction and human-subject experiment simulation, there remains no systematic understanding of how well they perform across diverse behavioral science tasks. We introduce BehaviorBench, a comprehensive benchmark that evaluates foundation models along four core capabilities: (1) behavior prediction and simulation, (2) strategic decision-making, (3) subject-trait inference, and (4) behavioral knowledge application. Crucially, BehaviorBench evaluates model outputs at both the individual and distributional levels, capturing not only per-subject accuracy
קרא במקור המקורי