כתבה
arXiv cs.CL ·
האם היגיון משפר עומק פסיכולוגי במודלי שפה גדולים?
Does Reasoning Improve Psychological Depth in Large Language Models? It Depends on Who's Judging
חוקרים בדקו את השפעת ההיגיון על עומק פסיכולוגי בסיפורים קצרים. הם השוו את GPT-5 ו-GPT-4o, ו-DeepSeek-R1 ו-DeepSeek-V3. התוצאות הראו שהעדפות אנושיות לא הראו יתרון אוניברסלי להיגיון.
תקציר מקורי באנגליתarXiv:2609.13773v1 Announce Type: cross Abstract: LLM-as-a-Judge evaluators are increasingly used to score open-ended generation, yet a judge's correlation with human ratings on its development set may not guarantee valid measurement when outputs are closely matched and human preferences are subjective. We study this failure mode through psychological depth in short stories. Seven human readers and an LLM-judge ensemble selected on the original scalar Psychological Depth Scale dataset ($\rho = 0.646$) evaluated 60 blinded, prompt-matched story pairs from GPT-5 vs.\ GPT-4o and DeepSeek-R1 vs.\ DeepSeek-V3. Human preferences showed no universal reasoning advantage: GPT-5 was modestly preferred over GPT-4o (60.0--62.9\%), whereas DeepSeek-R1 trailed V3 (42.9\%), and inter-reader agreement was
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית