יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

PhysSAE: פיזיקה-לא-מוסברת עם אוטואנקודרים דקים

PhysSAE: Mechanistic Interpretability with Sparse Autoencoders
PhysSAE היא תשתית פיזיקה-לא-מוסברת שמשתמשת באוטואנקודרים דקים כדי לחשב את המייצגים הספריים של PINNs. התשתית נבחנה על שש משפחות של PDEs, והתוצאות הראו שהאוטואנקודרים יכולים לגלות תכונות פיזיקליות ולזהות את הגורמים הגורמים להן.
תקציר מקורי באנגליתarXiv:2609.07061v1 Announce Type: new Abstract: Physics-Informed Neural Networks (PINNs) embed PDE residuals into neural network training, but their internal representations remain opaque: it is unknown what physical features their hidden layers encode or whether those features have a localized causal role. We present PhysSAE, a mechanistic interpretability framework that trains overcomplete sparse autoencoders (SAEs) on PINN penultimate-layer activations and evaluates dictionary atoms through direct causal intervention in the original frozen hidden state: $h_{\mathrm{cf}} = h - \alpha z_k d_k$, bypassing the SAE decoder entirely. Across six PDE families, with 3 PINN seeds and 3 SAE seeds each---we show that (i) Our discovered SAE atoms align with independently-defined physical observables
קרא במקור המקורי