כתבה
arXiv cs.CL ·
A Systematic Analysis of the Predictive Power of LM Surprisal in Reading Chinese
תקציר מקורי באנגליתarXiv:2610.04898v2 Announce Type: replace Abstract: This study analyzes the predictive power of LM-derived, token-level surprisal on Mandarin Chinese reading times. We first propose the Shortest Matching Sequence (SMS), an alignment scheme that maps between the word segmentation assumed by eye-tracking corpora and the LMs' subword tokenization, as the two tokenizations often disagree in the context of Mandarin Chinese. Then, using a suite of Chinese-Pythia models (14M-1.4B) trained on scratch with 30B tokens, we examine how well surprisal predicts first fixation duration, gaze duration, and total reading time in three paragraph-level eye-tracking corpora of Mandarin Chinese (GECO-CN, HKP, and MECO). Contrary to previous null findings, our results show that surprisal is predictive of Chines
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית