כתבה
arXiv cs.LG ·
ביקורתיות בפירוק חד-פעמי ובפחתת-סימנים של נתוני רנדומל עם תכנות-אנומלי
Criticality in Dissimilar Decomposition and Undersampling of Random Datasets with Anomalies
חוקרים ניסו להבין איך נתוני AI שנוצרו על ידי LLM קיימים משפיעים על פירוק חבצ'י של נתוני LLM חדשים. הם גילו תגובה פאזה, וכי נתוני AI עצמם יוצרים פירוק חד-פעמי חזק.
תקציר מקורי באנגליתarXiv:2609.13201v1 Announce Type: new Abstract: Training datasets for upcoming LLMs would include a significant amount of AI text/image data generated from current LLMs. In such a scenario, it is important to understand how this affects batch decompositions and thereby, the performance of the resultant new LLM. In this paper, we consider AI generated data as anomalies ``linked" to main data points and study decomposition and undersampling properties of the overall random dataset. We use redundancy graphs and iteration techniques to obtain bounds for the minimum size of a strongly dissimilar (SD) decomposition and demonstrate a phase transition phenomena, wherein the minimum size is essentially determined by the \emph{main} data points when the number of anomalies is small and is ``taken" o
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית