כתבה
arXiv cs.AI ·
קיום קובע דיוק: זיהום שקט במאגרי נתונים של דטקטור
Prevalence Determines Precision:Silent Contamination in Detector-Defined Datasets
מאגרי נתונים שנבנו על ידי דטקטורים יכולים להיות מזוהמים באירועים פנטומים, המשפיעים על הדיוק. זה נכון גם למודלים כמו Claude ו-GPT-5.
תקציר מקורי באנגליתarXiv:2609.11449v1 Announce Type: cross Abstract: Many ML datasets are constructed by running a detector, heuristic, or model over candidate pools; accepted items become labels. Dataset precision is then governed by true-positive prevalence in each pool via Bayes, not solely by detector quality. Using one instrument and period, we hold a detector-defined event dataset plus an independent official index labeling every detected item as real or phantom. One detector, three pools yield phantom rates 81.7%, 9.0%, and 0.0%. Transferring precision from the two high-rate pools to the low-rate pool predicts 0.955 versus measured 0.183, a +422% error; the Bayes expression predicts all three within 3.3%. The detected response curve is an exact convex combination of a true-event and a phantom componen
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית