כתבה
arXiv cs.CL ·
Large Language Models Exhibit Human-Like Bayesian Hypocrisy
תקציר מקורי באנגליתarXiv:2609.35779v1 Announce Type: cross Abstract: Given recent achievements of large language models (LLMs), frontier models are expected to perform well on Bayesian reasoning tasks, at least as well as humans. Furthermore, there is no reason to expect that LLMs will condemn others who offer those very same Bayesian judgments, a fallibility observed in human decision-making (Cao, et al., 2019). In 5 experiments with 48 experimental conditions employing over 5,000 trials, GPT-4o and Claude 3.7 Sonnet were tested on two variations of a Bayesian reasoning task. We also assessed LLM evaluation of the competence and morality of a hypothetical person who had offered the same reasoning task as them. LLMs hovered near human performance on the Bayesian task, though their reasoning was more rule-bas
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית