כתבה
arXiv cs.LG ·
מודלי דיפוזיה: פירוק ספורדי של טרקטוריות פעולה
Residualized Temporal Sparse Autoencoders for Interpreting Diffusion Models
מודלי דיפוזיה: פירוק ספורדי של טרקטוריות פעולה. ניתוח של טרקטוריות פעולה של מודלי דיפוזיה על ידי פירוק ספורדי.
תקציר מקורי באנגליתarXiv:2605.27813v2 Announce Type: replace-cross Abstract: Text-to-image diffusion models generate images by iterative denoising, so their internal layers produce trajectories of activations rather than single static representations. Sparse autoencoders (SAEs) have recently been used to decompose diffusion activations into interpretable features, but most approaches analyze individual timesteps or condition on time rather than learning from full trajectories. Training one SAE on whole trajectories would make each feature a single trajectory across timesteps, but adjacent activations are largely linearly predictable from one another, so such an SAE spends its latents on content carried forward from step to step. We introduce residualized temporal SAEs (ReSAE), which fit linear predictors bet
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית