כתבה
arXiv cs.LG ·
GRASP: תיאור גאומטרי-מודל חיבורי להסתגלות סקאלאבילית של נתוני תגובה
GRASP: Geometry-aware Residual Alignment for Scalable Pretraining Data Attribution
GRASP הוא תיאור גאומטרי-מודל חיבורי שמטרתו לשפר את תגובת הנתונים בסקאלה. הוא משתמש בפניית חיבורי קוורטית כדי לתאר את התגובה של קבוצות של נתונים. GRASP נבחן באמצעות תרגילי תגובה והוכיח יעילות גבוהה.
תקציר מקורי באנגליתarXiv:2606.06892v3 Announce Type: replace Abstract: Scalable data attribution methods typically assign isolated utility scores to individual training examples. This prevalent additive assumption fundamentally fails to capture critical subset dynamics, including data redundancy and complementary coverage. In this work, we reframe attribution as subset-level counterfactual utility prediction and introduce GRASP, an interaction-aware surrogate. Grounded in a theoretical smoothness lower bound, GRASP explicitly models subset interactions through a quadratic geometric penalty. To achieve pretraining-scale efficiency without relying on hidden oracle tuning, we couple low-dimensional feature sketches with a strictly finite lower-confidence bound selection protocol. Extensive subset-retraining eva
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית