כתבה
arXiv cs.LG ·
BehaviorBench: Benchmarking Foundation Models for Behavioral Science Tasks
תקציר מקורי באנגליתarXiv:2606.24162v2 Announce Type: replace-cross Abstract: Foundation models have been increasingly applied to behavioral science domains such as psychology, sociology, and economics. While these models show promise in tasks such as survey response prediction and human-subject experiment simulation, there remains no systematic understanding of how well they perform across diverse behavioral science tasks. We introduce BehaviorBench, a comprehensive benchmark that evaluates foundation models along four core capabilities: (1) behavior prediction and simulation, (2) strategic decision-making, (3) subject-trait inference, and (4) behavioral knowledge application. Crucially, BehaviorBench evaluates model outputs at both the individual and distributional levels, capturing not only per-subject acc
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית