כתבה
arXiv cs.AI ·
Revisiting Input Time-frequency Representations in Multi-pitch Estimation for Vocal Ensembles
תקציר מקורי באנגליתarXiv:2610.03656v1 Announce Type: cross Abstract: Multi-pitch estimation in vocal ensembles is challenging because singers occupy overlapping pitch ranges and often sing at closely spaced fundamental frequencies, causing their harmonics to overlap in time-frequency representations. Existing models commonly use harmonic constant-Q transform (HCQT)-based representations to provide frequency-adaptive resolution, at the cost of expensive feature extraction when training mixtures are generated on the fly. We revisit this design and compare HCQT with a linear short-time Fourier transform (STFT), whose frequency bins are directly provided as model inputs. Despite its fixed frequency resolution and the absence of a pitch-aligned input grid, the linear STFT outperforms HCQT while substantially redu
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית