כתבה
arXiv cs.CL ·
KCSAT-ML: חקירת מודלי תקיפה עם קוהורט-ארצי-אנושי-קשיות
KCSAT-ML: Probing Reasoning Models with Nationwide-Cohort Human Difficulty
במאמר זה, חברת Naver הציגה את KCSAT-ML, בסיס נתונים חדש לבדיקת מודלי תקיפה. הבסיס כולל 664 תרגילים במתמטיקה, כולל 339 תרגילים עיקריים עם תקלות רשמיות. המאמר חשף שלושה תבניות: (1) דיוק נמוך נקלע לקצה הגבוה של תקלות אנושיות; (2) גידול בקיטוב (TTS) גורם לשימוש בסימנים רובאיים לינארי עם קצב תקלות; (3) TTS גורם להפלה בין תקיפה נגד-סקאלה על תרגילים קשים ותקיפה יתר-סקאלה על תרגילים קלים.
תקציר מקורי באנגליתarXiv:2606.10403v3 Announce Type: replace Abstract: Math reasoning benchmarks have proliferated, yet most lack a per-item difficulty signal grounded in actual human performance. We introduce KCSAT-ML, a decade (2014-2025) of Korean College Scholastic Ability Test (KCSAT; Suneung) mathematics: 664 problems with a 339-item core set carrying official per-item error rates from nationwide cohorts of hundreds of thousands of examinees. We pair the benchmark with Difficulty-aligned Reasoning Gain (DRG): a score-orthogonal metric that asks whether a model's mistakes concentrate on the items humans found hard, or on items humans found easy. Together they expose, across a wide range of VLMs (and LLMs with OCR), three patterns: (i) low-budget accuracy collapses on the high-human-error tail at every m
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית