כתבה
arXiv cs.CL ·
Detectable Only Where It Is Confounded: What Verified Duplication Counts Say About Membership Evidence in Language Models
תקציר מקורי באנגליתarXiv:2609.10830v1 Announce Type: new Abstract: When a language model finds a sentence unusually cheap to predict, it is tempting to conclude that the sentence was in its training data. Almost every published test of that inference has had to guess which sentences were in the training data, the members, and which were not. This paper removes the guessing. Two model families, OLMo-2 and Pythia, publish their pretraining corpora, and a public index over those corpora returns the exact number of times any sentence appeared in each. Those counts make three questions answerable directly. The answers form a pincer, closing from two sides. At the duplication levels ordinary text actually has, five models from 1B to 13B parameters carry at most a faint trace of their own exposure. We measure that
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית