English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Winning Approach for Querying Sounds by Vocal Imitation: AES AIMLA 2025 Challenge Technical Report

Forum topic · 小凯 · 2026-08-21

Summary

A technical report by Aditya Bhattacharjee, Christos Plachouras, Sungkyun Chang, and Emmanouil Benetos describes the winning submission to the AES AIMLA 2025 Challenge on querying sound effects by vocal imitation. The authors investigate two complementary fine-tuning strategies: contrastive learning with a frozen, pretrained CED (Contrastive Audio Encoder) encoder, and joint contrastive-triplet learning with semi-hard negatives using a MobileNetV3 encoder. Vocal imitation query systems allow users to find sound effects by imitating them vocally, requiring robust embedding models that align vocal imitations with target sound recordings. The report has been updated after the challenge to include additional details released post-competition. The work is available on arXiv as paper 2608.19174 and was originally published on zhichai.net.

Overview

This technical report describes the winning submission to the AES AIMLA 2025 Challenge on querying sound effects by vocal imitation — the task of retrieving target sounds from a database using a user's vocal imitation as the query.

  • Research area: Machine Learning (ML)
  • Authors: Aditya Bhattacharjee, Christos Plachouras, Sungkyun Chang, Emmanouil Benetos
  • arXiv: 2608.19174
  • Abstract

    > This technical report describes our winning submission to the AES AIMLA 2025 Challenge on querying sound effects by vocal imitation. We investigate two complementary fine-tuning strategies: contrastive learning with a frozen, pretrained CED encoder, and joint contrastive-triplet learning with semi-hard negatives using a MobileNetV3 encoder. This report has been updated for posterity to include details released after the challenge.

    Key Points

  • Task: Retrieving sound effects by matching them against vocal imitations.
  • Strategy 1: Contrastive learning using a frozen, pretrained CED encoder.
  • Strategy 2: Joint contrastive-triplet learning with semi-hard negatives, based on a MobileNetV3 encoder.
  • Result: The approach won the AES AIMLA 2025 Challenge.
---

*Auto-collected on 2026-08-21.*

Tags

#machine-learning#audio-retrieval#vocal-imitation#contrastive-learning#triplet-learning#sound-effects#aes-aimla-2025#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633736