כתבה
arXiv cs.CL ·
STRUCTURALCOST: מאגר נתונים לדירוג קושי עיבוד משפטים
STRUCTURALCOST: A controlled reading time dataset for modeling human sentence processing difficulty
STRUCTURALCOST הוא מאגר נתונים לדירוג קושי עיבוד משפטים. הוא כולל 475 משתתפים ו-40,800 תצפיות. המחקר בודק את הקושי בעיבוד משפטים עם תלות נושא-פועל. התוצאות מראות שמודלים שונים מחקים את הקושי, אך לא לגמרי.
תקציר מקורי באנגליתarXiv:2610.08208v1 Announce Type: new Abstract: We introduce STRUCTURALCOST, a self-paced reading dataset of 475 participants and 40,800 observations isolating the processing cost of long-distance subject-verb dependency resolution. We replicate a low-powered psycholinguistic finding at NLP scale, namely that human reading times at the main verb increase with dependency length, driven by syntactic embedding beyond linear distance. Different language models -- spanning n-gram models, SSMs, and transformers -- partially mirror this graded difficulty profile, yet underestimate the integration cost humans incur, with a gap that persists across architectures and model sizes. This suggests these models capture the predictive component of human processing but not the full integration cost that wo
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית