כתבה
arXiv cs.LG ·
Feature Encoding in VAE-based Audio Decoders: Effects of Input, Depth and Distribution
תקציר מקורי באנגליתarXiv:2610.07966v1 Announce Type: cross Abstract: Neural audio synthesis models like the Realtime Audio Variational autoEncoder (RAVE) achieve impressive genera tion quality, yet how their internal representations encode musical features remains poorly understood. We present a systematic layer-wise and cross-layer cluster analysis of RAVE decoder activations across three models trained on different musical domains, tested with four stimulus types. We then evaluate architectural generalization with a general purpose EnCodec model. For RAVE, we find that synthetic stimuli are encoded well across models and audio features (pitch |\r{ho}|=0.45, 5.1x the null, BPM |\r{ho}| = 0.76, 8.6x the null). These results are reduced but still substantively apparent when using natural audio (mean across fe
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית