כתבה
arXiv cs.CL ·
SynthSentry: איתור נתונים סינתטיים בנתוני אימון מודלי שפה
SynthSentry: Detecting Synthetic Data Contamination in Language Model Training Data
SynthSentry הוא כלי לאיתור נתונים סינתטיים בנתוני אימון של מודלי שפה. הוא משתמש במדדים סטטיסטיים כדי לזהות נתונים מלאכותיים. הכלי נבדק על נתונים מלאכותיים שנוצרו על ידי מודלים שונים.
תקציר מקורי באנגליתarXiv:2609.12353v1 Announce Type: new Abstract: Large language models trained recursively on their own or other models' outputs undergo model collapse, in which distributional tails and factual accuracy deteriorate while fluency survives. Prior work diagnoses collapse after training; the actionable problem is screening a corpus of unknown provenance before training. We introduce SynthSentry, a corpus-level, model-agnostic contamination signal requiring no access to the generating model, no generation history, and no synthetic labels. The score is a distributional divergence over three statistics: lexical diversity collapse, n-gram tail truncation, and perplexity variance across reference models. We evaluate on corpora contaminated by small open-weight generators and an instruction-tuned op
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית