כתבה
arXiv cs.AI ·
דירוג מודלי שפה של דיכאון משקף יותר את הרועה מאשר המטופל
Language-model ratings of depression reflect the rater more than the patient
מחקר חדש בדק את היכולת של מודלי שפה לדרג רמות דיכאון. התוצאות הראו שהדירוגים משקפים יותר את הרועה מאשר המטופל. המחקר השתמש ב-11 מודלים פתוחים וב-189 ראיונות כדי לבדוק את התוצאות.
תקציר מקורי באנגליתarXiv:2610.08501v1 Announce Type: cross Abstract: Depression has no diagnostic blood test. Language models promise tireless, consistent assessment, but can accurate raters disagree about individuals? We pre-registered 880 language-model raters, crossing 11 open models with prompting and scoring choices, and applied them to 189 interviews against the eight-item Patient Health Questionnaire. Model choice explained 30.0% of summed-symptom score variance, stable participant differences 10.5%. Two randomly drawn raters with area under the receiver operating characteristic curve (AUC) >= 0.70 disagreed on screening decisions for 40% of participants, on average. Average over-rating governed how many were flagged, yet equal-capacity raters chose differently for about one participant in five. A loc
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית