כתבה
arXiv cs.LG ·
גודל האימון המינימלי לדוגמנים טבלאיים עמוקים
Below what training size do deep tabular generators stop beating trivial baselines? A preregistered benchmark on a size ladder of clinical and standard datasets
מחקר חדש בודק את ביצועיהם של דוגמנים טבלאיים עמוקים על מנת לקבוע האם הם מצליחים לעקוף בסיסי טריביאליים בגודלי אימון שונים. המחקר השתמש ב-8 סטים של נתונים ציבוריים ו-4 סטים קליניים קטנים.
תקציר מקורי באנגליתarXiv:2610.03500v1 Announce Type: new Abstract: Deep tabular generative models are benchmarked on datasets with tens of thousands of rows; clinical datasets have hundreds. We preregistered and ran a size-ladder benchmark to find where the two regimes diverge: 8 public datasets subsampled from 200 to 20,000 training rows, seven generators (independent marginals, Gaussian copula, SMOTE, unconditional SMOTE, CTGAN, TVAE, TabDDPM) with a fixed 20-trial tuning budget and 5 evaluation seeds, plus 4 natively small clinical datasets at true size, for 2,220 committed runs in total. The primary metric is the AUROC of fixed classifiers trained on synthetic and tested on real data. In 23 of 24 (dataset, deep model) pairs no deep model ever beats the best trivial baseline by more than seed noise, at an
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית