יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

דירוג מודלי שפה של דיכאון משקף יותר את המדרג ופחות את המטופל

Language-model ratings of depression reflect the rater more than the patient
חוקרים בדקו 880 מודלי שפה שונים עם פרומפטים ואופציות ציון שונות, והגיעו למסקנה שהבחירה של המודל היא זו שמשפיעה על הדירוג. הניסוי הראה שאפילו מדרגים עם יכולת זיהוי גבוהה יכולים לחלוק על החלטות הסקרינג.
תקציר מקורי באנגליתarXiv:2610.08501v1 Announce Type: new Abstract: Depression has no diagnostic blood test. Language models promise tireless, consistent assessment, but can accurate raters disagree about individuals? We pre-registered 880 language-model raters, crossing 11 open models with prompting and scoring choices, and applied them to 189 interviews against the eight-item Patient Health Questionnaire. Model choice explained 30.0% of summed-symptom score variance, stable participant differences 10.5%. Two randomly drawn raters with area under the receiver operating characteristic curve (AUC) >= 0.70 disagreed on screening decisions for 40% of participants, on average. Average over-rating governed how many were flagged, yet equal-capacity raters chose differently for about one participant in five. A locke
קרא במקור המקורי