[论文] Revisiting Input Time-frequency Representations in Multi-pitch Estimat...

研究领域: ML 作者: Junyoung Koh, Hao-Wen Dong 发布时间: 2026-10-02 arXiv: 2610.03656

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: ML 作者: Junyoung Koh, Hao-Wen Dong 发布时间: 2026-10-02 arXiv: 2610.03656

中文摘要

声乐合奏的多音高估计具有挑战性,因为歌手的音域重叠且基频间隔很近,导致谐波在时频表示中重叠。现有模型通常使用谐波恒定 Q 变换(HCQT)表示来提供频率自适应分辨率,但当训练混合在线生成时,特征提取代价高昂。我们重新审视这一设计,将 HCQT 与线性短时傅里叶变换(STFT)进行比较——后者的频率 bin 直接作为模型输入。尽管频率分辨率固定且缺乏音高对齐的输入网格,线性 STFT 的性能仍优于 HCQT,同时大幅降低了特征提取成本。进一步分析表明,更长的分析窗或更广的频谱覆盖没有额外改进,而将输入限制在预测音高范围内会降低线性 STFT 的优势。这些结果表明,更细的频率分辨率不一定能改善声乐合奏的多音高估计,而更短的分析窗对时变声乐音高可能更有效。

原文摘要

Multi-pitch estimation in vocal ensembles is challenging because singers occupy overlapping pitch ranges and often sing at closely spaced fundamental frequencies, causing their harmonics to overlap in time-frequency representations. Existing models commonly use harmonic constant-Q transform (HCQT)-based representations to provide frequency-adaptive resolution, at the cost of expensive feature extraction when training mixtures are generated on the fly. We revisit this design and compare HCQT with a linear short-time Fourier transform (STFT), whose frequency bins are directly provided as model inputs. Despite its fixed frequency resolution and the absence of a pitch-aligned input grid, the linear STFT outperforms HCQT while substantially reducing feature-extraction cost. Further analysis sho...


*自动采集于 2026-10-06*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens