כתבה
arXiv cs.CL ·
האם פירוש יכול לחזות אחריות על נתונים לא ראויים?
Can Interpretation Predict Behavior on Unseen Data?
במאמר זה, המחברים חוקרים את האפשרות לשימוש בפירוש כדי לחזות אחריות של דגימות נתונים לא ראויים. הם מציעים ומדגימים חדשנות זו באמצעות שימוש בפרטי המודל כדי לחזות אחריות על נתונים לא ראויים. המחברים מציעים חדשנות זו כדי להבין טוב יותר את המודלים ולשפר את יכולתם לפעול באופן יעיל.
תקציר מקורי באנגליתarXiv:2507.06445v4 Announce Type: replace-cross Abstract: Interpretability research often predicts model responses to targeted mechanistic interventions. But can we predict responses to unseen input data? We propose and demonstrate this alternate objective by using model internals to predict their out-of-distribution (OOD) behavior. We train hundreds of Transformers on simple synthetic tasks, where perfect in-distribution accuracy is compatible with multiple OOD generalization rules. We successfully use attention patterns -- observed only on in-distribution data -- to predict which rule each model follows on OOD data. Our experiments decouple the mechanistic faithfulness of our interpretation from its predictive value; ablations reveal such internal patterns can suppress rather than suppor
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית