כתבה
arXiv cs.LG ·
ממשוואות התנהגות למוצא: תיוח דגמי יסוד טבולריים לנתוני אימון סינתטיים
From Behavior to Provenance: Attributing Tabular Foundation Models to Synthetic Pretraining Data
חוקרים פיתחו שיטה לתיוח דגמי יסוד טבולריים לנתוני אימון סינתטיים. השיטה מאפשרת לבדוק את השפעת הנתונים על התנהגות המודל. הניסויים הראו שהסרת 5% מהנתונים הסינתטיים המשפיעים ביותר גרמה לירידה בביצועים.
תקציר מקורי באנגליתarXiv:2610.02347v1 Announce Type: new Abstract: Training-data attribution aims to identify which training examples shape model behavior, yet validating such claims is difficult because causal training influence is rarely observable. We argue that controlled synthetic pretraining makes attribution experimentally testable. Using O'PRIOR, a provenance-rich synthetic task generator for tabular foundation models, we construct a testbed in which every pretraining task carries explicit lineage over structural mechanisms, missingness, confounding, shortcuts, and distribution shift. We combine behavior-conditioned attribution with counterfactual retraining and provenance-aware interventions to test both task-level faithfulness and mechanism-level consistency. On held-out real tasks, removing the to
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית