כתבה
arXiv cs.LG ·
קיום קובע דיוק: זיהום סמוי בנתוני דטקטור
Prevalence Determines Precision:Silent Contamination in Detector-Defined Datasets
דיוק נתוני דטקטור מושפע מקיום אמיתי של אירועים, ולא רק מאיכות הדטקטור. זיהום סמוי בנתוני דטקטור יכול להשפיע על תוצאות הדטקטור.
תקציר מקורי באנגליתarXiv:2609.11449v1 Announce Type: new Abstract: Many ML datasets are constructed by running a detector, heuristic, or model over candidate pools; accepted items become labels. Dataset precision is then governed by true-positive prevalence in each pool via Bayes, not solely by detector quality. Using one instrument and period, we hold a detector-defined event dataset plus an independent official index labeling every detected item as real or phantom. One detector, three pools yield phantom rates 81.7%, 9.0%, and 0.0%. Transferring precision from the two high-rate pools to the low-rate pool predicts 0.955 versus measured 0.183, a +422% error; the Bayes expression predicts all three within 3.3%. The detected response curve is an exact convex combination of a true-event and a phantom component
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית