יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

ReLaG: פלטפורמה סקאלאבילית שמגדירה חלוקות סטטיסטיות עם יחסי גומלין

ReLaG: A Scalable Framework Generalizing Random Splits to Data with Latent Relations
ReLaG היא פלטפורמה שמגדירה חלוקות סטטיסטיות עם יחסי גומלין בין דגימות. היא משמשת לטיפול במידע עם יחסי גומלין, כגון נתוני ביוכימיה.
תקציר מקורי באנגליתarXiv:2609.38538v1 Announce Type: new Abstract: Random splitting can yield non-independent train--test subsets when a dataset contains related samples, as is common in certain applications such as biochemical studies. This leads to overly optimistic generalization estimates. Here, we introduce ReLaG, a modality-agnostic framework that models sample relatedness through a hierarchical latent-variable process and infers groups of related samples using proximity graphs and community detection to produce independent train--test subsets. Across molecular and protein datasets, ReLaG matches existing relation-aware methods while scaling substantially better, enabling splits at previously impractical dataset sizes. We further introduce a label-free procedure that adapts the splitting resolution to
קרא במקור המקורי