יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

הערכת מודלי שפה גדולים לאבחון דיכאון

"Mirror" Large Language Model Evaluations of Depression are Criterion Contaminated
חוקרים גילו שהערכות מודלי שפה גדולים לאבחון דיכאון עלולות להיות מוטות. המחקר הראה כי השימוש בשאילתות זהות לאבחון דיכאון יכול ליצור תוצאות מוטות. מודלים כמו LLaMA יכולים לחזות תוצאות ברמה גבוהה, אך יש לבחון את התוצאות בזהירות.
תקציר מקורי באנגליתarXiv:2508.05830v3 Announce Type: replace Abstract: Large Language Model (LLM) studies that use language responses elicited from depression assessments to predict scores on those same assessments often report near-perfect prediction of depression. We refer to these as "Mirror" evaluations and demonstrate an applied case of criterion contamination. N = 110 participants completed both structured diagnostic depression interviews (Mirror condition) and life history interviews ("Non-Mirror" condition). LLMs were prompted to predict depression scores in each condition. As expected, Mirror evaluations were near-perfect. However, Non-Mirror evaluations also displayed prediction sizes considered outstanding in psychology. Further, both Mirror and Non-Mirror predictions correlated with Patient Healt
קרא במקור המקורי