יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

דיסקרימינטיבי ספן כפוטר סינתטי לנתוני יעילות דרך תיקון קלאסיפייר

Discriminative Span as a Predictor of Synthetic Data Utility via Classifier Reconstruction
במאמר זה, המחברים מציגים חידוש חדשני שמאפשר לגרוע את יעילות הנתונים הסינתטיים ללא צורך באימון דגם. החידוש, המבוסס על גאומטריה, פועל במרחב האימבדינג של דגם יסודי ומייצג את הקבוצת הנתונים דרך וקטורי ההבדלים בין דגימות. המחברים מדגימים את החידוש במספר קבוצות נתונים וארכיטקטורות שונות.
תקציר מקורי באנגליתarXiv:2605.09697v5 Announce Type: replace-cross Abstract: In many real-world computer vision applications, including medical imaging and industrial inspection, binary classification tasks are characterized by a severe scarcity of positive samples. A widely adopted solution is to generate synthetic positive data using image-to-image transformations applied to negative samples. However, a fundamental challenge remains: how can we reliably assess whether such synthetic data will improve downstream model performance? In this work, we propose a geometry-driven metric that predicts the utility of synthetic data without requiring model training. Our approach operates in the embedding space of a pre-trained foundation model and represents the dataset through difference vectors between samples. We
קרא במקור המקורי