כתבה
arXiv cs.LG ·
האם רישום טוב יותר את המחשבה במודלי שפה גדולים? זה תלוי במי שמסקן
Does Reasoning Improve Psychological Depth in Large Language Models? It Depends on Who's Judging
מחקר חדש מצא שמודלי שפה גדולים לא תמיד משפרים את העומק הנפשי של התוכן שהם יוצרים. המחקר, שפורסם בארכיווה, ניתח 60 סיפורים שנוצרו על ידי מודלי שפה שונים ומצא שהבחירה של האנשים בין הסיפורים הייתה לא קבועה. המודל של GPT-5, לדוגמה, נבחר על ידי 60-62.9% מהקוראים, בעוד שהמודל DeepSeek-R1 נבחר רק ב-42.9%. המחקר מצא גם שהמודלים השונים נבחרו על ידי האנשים בצורה שונה, ושהבחירה הייתה תלויה במי שקרא.
תקציר מקורי באנגליתarXiv:2609.13773v1 Announce Type: new Abstract: LLM-as-a-Judge evaluators are increasingly used to score open-ended generation, yet a judge's correlation with human ratings on its development set may not guarantee valid measurement when outputs are closely matched and human preferences are subjective. We study this failure mode through psychological depth in short stories. Seven human readers and an LLM-judge ensemble selected on the original scalar Psychological Depth Scale dataset ($\rho = 0.646$) evaluated 60 blinded, prompt-matched story pairs from GPT-5 vs.\ GPT-4o and DeepSeek-R1 vs.\ DeepSeek-V3. Human preferences showed no universal reasoning advantage: GPT-5 was modestly preferred over GPT-4o (60.0--62.9\%), whereas DeepSeek-R1 trailed V3 (42.9\%), and inter-reader agreement was n
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית