כתבה
arXiv cs.CL ·
Internal Knowledge Without External Expression: Probing the Generalization Boundary of a Classical Chinese Language Model
תקציר מקורי באנגליתarXiv:2604.14180v3 Announce Type: replace Abstract: We train a 318M-parameter Transformer language model from scratch on a curated corpus of 1.56 billion tokens of pure Classical Chinese, with zero English characters or Arabic numerals. Through systematic out-of-distribution (OOD) testing, we ask whether the model distinguishes known from unknown inputs, and whether it expresses that distinction in its generated text. We find a clear dissociation between internal and external uncertainty. Internally, the model exhibits a perplexity jump ratio of 2.39x between real and fabricated historical events (p = 8.9e-11, n = 92 per group), with semi-fabricated events (real figures + fictional actions) showing the highest perplexity (4.24x, p = 1.1e-16), demonstrating genuine factual encoding beyond s
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית