כתבה
arXiv cs.LG ·
ספקטרום רגישות לאופטימיזציה רציפה של דגימות טקסט
A Fragility Spectrum for Recursive Language-Model Training
נמצא כי דגימות טקסט שנוצרו על ידי דגימות טקסט קודמות גורמות לאובדן של תכונות הדגימה. נמצא כי תכונות הדגימה נקבעות על ידי תכונות המודל עצמו, ולא על ידי גודל הפרמטרים.
תקציר מקורי באנגליתarXiv:2609.11149v1 Announce Type: cross Abstract: Model-generated text is finding its way back into training corpora, and there is plenty of evidence that training on such data over and over collapses output diversity. Prior work has studied the phenomenon itself: which protocols and which data mixtures cause collapse. But different models behave very differently under the same process. We fix one recursive contamination protocol and let 13 publicly released checkpoints form an ecosystem that shares a common corpus for five generations. The unique 4-gram outcome after five generations ranges from 0.187 to 0.940 across checkpoints, a roughly five-fold spread: some models are barely touched, others degenerate into repetitive fragments. Changing the composition of the shared pool or mixing in
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית