כתבה
arXiv cs.CL ·
דפסים של חזרה: מה סטטיסטיקות חזרה יכולות לספר על הוכחות חברות במודלי שפה
Detectable Only Where It Is Confounded: What Verified Duplication Counts Say About Membership Evidence in Language Models
מחקר מצא כי מודלי שפה נזכרים בצורה רפויה בנתוני הלמידה שלהם
תקציר מקורי באנגליתarXiv:2609.10830v2 Announce Type: replace Abstract: When a language model finds a sentence unusually cheap to predict, it is tempting to conclude that the sentence was in its training data. Almost every published test of that inference has had to guess which sentences were in the training data, the members, and which were not. This paper removes the guessing. Two model families, OLMo-2 and Pythia, publish their pretraining corpora, and a public index over those corpora returns the exact number of times any sentence appeared in each. Those counts make three questions answerable directly. The answers form a pincer, closing from two sides. At the duplication levels ordinary text actually has, five models from 1B to 13B parameters carry at most a faint trace of their own exposure. We measure t
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית