יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

בין סטרקטור לשרשרת-מחשבה: השוואת קריטריות פלט-LLM להערכת חומרה של דיכאון

Structure vs. Chain-of-Thought: Evaluating LLM Criteria Extraction for Depression Severity
במאמר זה, נבדוק את יעילות קריטריות פלט-LLM להערכת חומרה של דיכאון. נשווה בין שני גישות: קריטריות פלט ושרשרת-מחשבה. נבדוק גם את השפעת קליברציה על תוצאות.
תקציר מקורי באנגליתarXiv:2609.39049v1 Announce Type: cross Abstract: A large language model (LLM) can rate depression severity directly from a social media post or mark which clinical criteria the post shows and let code turn the count into a label. The latter is easier to audit because a clinician can check each marked criterion. We compare these approaches on two Reddit corpora using three LLMs (from 9B to frontier scale) and two questionnaires (PHQ-9, BDI-II), and measure agreement with quadratic weighted kappa. For the two frontier models, criteria extraction scores above chain-of-thought on one corpus only when its decision thresholds are fitted on labeled data. Neither model's gain is significant, with or without recalibrating chain-of-thought on the same labels. With thresholds fixed a priori from PHQ
קרא במקור המקורי