יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

הערכת שיחות מדומות ב-GPT

Evaluating LLM-Simulated Conversations in Modeling Inconsistent and Uncollaborative Behaviors in Human Social Interaction
חוקרים פיתחו את CoCoEval, שיטה להערכת שיחות מדומות ב-GPT-4.1, GPT-5.1 ו-Claude Opus 4. התוצאות מראות פערים בין שיחות אנושיות למדומות.
תקציר מקורי באנגליתarXiv:2603.17094v2 Announce Type: replace Abstract: Simulating human conversations using large language models (LLMs) has emerged as a scalable methodology for modeling human social interaction. This paper reconsiders the evaluation of simulated conversations by explicitly recognizing that human conversations inherently involve inconsistent and uncollaborative behaviors, such as misunderstandings and interruptions. Since these behaviors contribute to the complexity of human social interaction, we argue that LLM-simulated conversations should reproduce them at frequencies comparable to those observed in human conversations. To support a detailed and interpretable evaluation of these behaviors, we introduce CoCoEval, a framework consisting of an evaluation scheme based on turn-level detectio
קרא במקור המקורי