יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

כשמשתמשים סינתטיים נכשלים

When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses
מחקר חדש בודק את יכולתם של מודלים גדולים לשפה (LLM) לחקות תגובות אנושיות בסקרים. התוצאות מראות כי המודלים נכשלים בשני תחומים: הם אינם מצליחים לחקות תגובות אנושיות ברמה האישית, והם מייחסים חשיבות רבה מדי לזהות האישית בעת חיזוי דעות.
תקציר מקורי באנגליתarXiv:2607.26348v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product, policy, and market decisions. We ask when this substitution is valid and when it fails, and package the answer as an evaluation framework for intelligent synthetic-user systems. A single protocol, run across four models spanning two families and an 8B-to-frontier capability range, is applied to two independent domains of real human-response data: U.S. general social attitudes (General Social Survey) and cross-cultural values (World Values Survey). Every model is benchmarked against a suite of non-LLM baselines fit on held-out human data. Under demographic prompting and the survey-simulation protocols we
קרא במקור המקורי