יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

נתונים שמשימים, תקן שמפלט: הערכת זיהום בקורפוסים ציבוריים של MRI לסיווג גידולי מוח

Scores That Hold, Benchmarks That Leak: Measuring Dataset Contamination in Public Brain-Tumor MRI Classification
במאמר זה, נבחן את איכות הקורפוסים הציבוריים של MRI לסיווג גידולי מוח. נמצא כי קורפוסים אלה נזיהו, וזיהום זה יכול להשפיע על תוצאות המודלים.
תקציר מקורי באנגליתarXiv:2610.00421v1 Announce Type: cross Abstract: Automated classification of brain tumors from MRI is a heavily published application of deep learning in medical imaging, with reported accuracies on public benchmarks routinely exceeding 98%. However, accuracy does not capture a critical dimension of benchmark quality: dataset integrity, defined as the independence of test from training data at the image, patient, and acquisition-source levels. We introduce a three-layer contamination framework comprising duplicate, patient, and source-label leakage to assess the public corpora on which this literature rests. We audit the three most widely used corpora against a chest-radiograph negative control and quantify each layer's effect on measured performance across nine architectures and three ev
קרא במקור המקורי