יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

כמה קול נשאר באמבדינג? סקירה של הפוך של קודרים לקול

How Much Audio Is Left In An Embedding? An Inversion Audit Of Audio Encoders
במאמר זה, נחקרה הסיבתיות של קודרים לקול על ידי שיקוף מקורי. התוצאות חשובות להבנת תכונות הקודרים.
תקציר מקורי באנגליתarXiv:2610.12250v1 Announce Type: cross Abstract: Pretrained audio encoders are reused for downstream tasks that are often unknown when the encoder is trained, so their usefulness depends partly on which signal properties survive the pretext objective. We study this retained information through paired source reconstruction. Using a shared Stable Audio Open latent diffusion decoder, we reconstruct five-second, 44.1-kHz stereo music from frozen representations produced by supervised classifiers (VGGish, ConvNeXt), an audio-text contrastive model (CLAP), and a waveform-reconstruction model (EnCodec). These objectives impose different pressures to preserve source detail, while their exposed interfaces vary substantially in temporal and spectral resolution. Evaluating on the Million Song Datase
קרא במקור המקורי