[论文] Revisiting Input Time-frequency Representations in Multi-pitch Estimat...
研究领域: ML 作者: Junyoung Koh, Hao-Wen Dong 发布时间: 2026-10-02 arXiv: 2610.03656
论文概要
研究领域: ML 作者: Junyoung Koh, Hao-Wen Dong 发布时间: 2026-10-02 arXiv: 2610.03656
中文摘要
声乐合奏的多音高估计具有挑战性,因为歌手的音域重叠且基频间隔很近,导致谐波在时频表示中重叠。现有模型通常使用谐波恒定 Q 变换(HCQT)表示来提供频率自适应分辨率,但当训练混合在线生成时,特征提取代价高昂。我们重新审视这一设计,将 HCQT 与线性短时傅里叶变换(STFT)进行比较——后者的频率 bin 直接作为模型输入。尽管频率分辨率固定且缺乏音高对齐的输入网格,线性 STFT 的性能仍优于 HCQT,同时大幅降低了特征提取成本。进一步分析表明,更长的分析窗或更广的频谱覆盖没有额外改进,而将输入限制在预测音高范围内会降低线性 STFT 的优势。这些结果表明,更细的频率分辨率不一定能改善声乐合奏的多音高估计,而更短的分析窗对时变声乐音高可能更有效。
原文摘要
Multi-pitch estimation in vocal ensembles is challenging because singers occupy overlapping pitch ranges and often sing at closely spaced fundamental frequencies, causing their harmonics to overlap in time-frequency representations. Existing models commonly use harmonic constant-Q transform (HCQT)-based representations to provide frequency-adaptive resolution, at the cost of expensive feature extraction when training mixtures are generated on the fly. We revisit this design and compare HCQT with a linear short-time Fourier transform (STFT), whose frequency bins are directly provided as model inputs. Despite its fixed frequency resolution and the absence of a pitch-aligned input grid, the linear STFT outperforms HCQT while substantially reducing feature-extraction cost. Further analysis sho...
*自动采集于 2026-10-06*
#论文 #arXiv #ML #小凯