יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

STRUCTURALCOST: מאגר נתונים לדירוג קושי עיבוד משפטים

STRUCTURALCOST: A controlled reading time dataset for modeling human sentence processing difficulty
STRUCTURALCOST הוא מאגר נתונים לדירוג קושי עיבוד משפטים. הוא כולל 475 משתתפים ו-40,800 תצפיות. המחקר בודק את הקושי בעיבוד משפטים עם תלות נושא-פועל. התוצאות מראות שמודלים שונים מחקים את הקושי, אך לא לגמרי.
תקציר מקורי באנגליתarXiv:2610.08208v1 Announce Type: new Abstract: We introduce STRUCTURALCOST, a self-paced reading dataset of 475 participants and 40,800 observations isolating the processing cost of long-distance subject-verb dependency resolution. We replicate a low-powered psycholinguistic finding at NLP scale, namely that human reading times at the main verb increase with dependency length, driven by syntactic embedding beyond linear distance. Different language models -- spanning n-gram models, SSMs, and transformers -- partially mirror this graded difficulty profile, yet underestimate the integration cost humans incur, with a gap that persists across architectures and model sizes. This suggests these models capture the predictive component of human processing but not the full integration cost that wo
קרא במקור המקורי