English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

MulTTiPop: A Multitrack MIDI Dataset for Pop Music Transcription Evaluation

Forum topic · 小凯 · 2026-07-13

Summary

MulTTiPop is a new dataset for evaluating automatic music transcription (AMT) models, presented by Nathan Pruyne, Benjamin Stoler, and William Chen in arXiv paper 2507.08753. The dataset comprises 572 pop music segments totaling 3.5 hours of audio, paired with multitrack MIDI recordings spanning diverse genres and decades from the 1930s through the 2000s. The authors built it by metadata-matching song segments from the Lakh MIDI and TheoryTab datasets, manually identifying an anchor beat between audio and MIDI, then applying beat tracking to the audio and warping the MIDI to align tempo and timing. Benchmarking state-of-the-art AMT models on MulTTiPop reveals substantial room for improvement: the best model achieves only 38% Onset F1. Details and sound examples are available at https://gclef-cmu.org/multtipop.

Overview

  • Field: Machine Learning (music information retrieval)
  • Authors: Nathan Pruyne, Benjamin Stoler, William Chen
  • Published: 2025-07-12
  • arXiv: 2507.08753
  • Abstract

    We present MulTTiPop, a dataset of pop music segments and their associated multitrack MIDI recordings for the evaluation of automatic music transcription models. MulTTiPop contains 572 segments of popular music totaling 3.5 hours of audio, and contains songs from diverse genres and decades from the 1930s to 2000s. To collect this dataset, we perform metadata-based matching on song segments from the Lakh MIDI and TheoryTab datasets, manually identify an anchor beat between the audio and MIDI, then use beat tracking on the audio and warp the MIDI to match its tempo and timing. We evaluate state-of-the-art automatic music transcription models on MulTTiPop and find substantial room for improvement, with the best model achieving 38% Onset F1. More details and sound examples of MulTTiPop are available at https://gclef-cmu.org/multtipop.

    Key Points

  • Dataset: 572 pop music segments with aligned multitrack MIDI, 3.5 hours of audio total.
  • Coverage: Diverse genres and decades, from the 1930s to the 2000s.
  • Construction pipeline:
  • 1. Metadata-based matching of song segments from the Lakh MIDI and TheoryTab datasets. 2. Manual identification of an anchor beat between audio and MIDI. 3. Beat tracking on the audio, then warping the MIDI to match tempo and timing.
  • Benchmark results: State-of-the-art automatic music transcription models still perform poorly; the best model reaches only 38% Onset F1, indicating significant room for improvement.
  • Resources: Details and sound examples at https://gclef-cmu.org/multtipop.

Tags

#automatic-music-transcription#midi-dataset#pop-music#machine-learning#music-information-retrieval#arxiv#benchmark

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178379416