יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

דיסקרימינטיבי ספאן כגורם ליוזמות נתונים סינתטיות על ידי תיקון מחלק

Discriminative Span as a Predictor of Synthetic Data Utility via Classifier Reconstruction
במאמר זה, המחברים מציגים מדד חדש שמצפה לאיכות נתונים סינתטיים ללא צורך באימון מודל. המדד פועל במרחב האפימינג של מודל יסודי שהוכשר מראש ומייצג את הקבצים דרך וקטורי ההבדלים בין הדגמים. המחברים מבודקים האם וקטור המשקל של מחלק המופעל יכול להיות מוצג תוך כדי הספייס של השונויות שהוזכרו. המדד נבדק על ידי המחברים במספר קבצי נתונים וארכיטקטורות.
תקציר מקורי באנגליתarXiv:2605.09697v4 Announce Type: replace-cross Abstract: In many real-world computer vision applications, including medical imaging and industrial inspection, binary classification tasks are characterized by a severe scarcity of positive samples. A widely adopted solution is to generate synthetic positive data using image-to-image transformations applied to negative samples. However, a fundamental challenge remains: how can we reliably assess whether such synthetic data will improve downstream model performance? In this work, we propose a geometry-driven metric that predicts the utility of synthetic data without requiring model training. Our approach operates in the embedding space of a pre-trained foundation model and represents the dataset through difference vectors between samples. We
קרא במקור המקורי