English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Listen, Think, and Understand: LTU and the OpenAQA Dataset (May 2023, arXiv)

Forum topic · 小凯 · 2026-07-05

Summary

"Listen, Think, and Understand" (arXiv:2305.10790, May 2023) by Yuan Gong, Hongyin Luo, Alexander H. Liu, Leonid Karlinsky, and James Glass introduces LTU, a model that connects audio perception with higher-level reasoning, and the OpenAQA dataset used to train it. The work addresses a gap between traditional audio models that focus on low-level recognition (speech transcription, sound event classification) and the ability to answer open-ended, reasoning-based questions about audio content. The forum entry catalogs the paper alongside related audio question answering and multimodal retrieval resources, including Clotho-AQA, and situates it within multi-modal question answering research. The OpenAQA dataset contains millions of audio question-answering pairs designed to teach models both listening (perception) and thinking (reasoning) over audio inputs. This post summarizes the paper's positioning, its dataset contribution, evaluation considerations, and practical implications for building audio-language systems, with cross-references to related cross-modal retrieval and audio QA work.

Listen, Think, and Understand: LTU and the OpenAQA Dataset (May 2023, arXiv)

Overview

This forum entry catalogs the paper Listen, Think, and Understand (arXiv:2305.10790, May 2023).

  • Authors / Affiliations: Yuan Gong, Hongyin Luo, Alexander H. Liu, Leonid Karlinsky, James Glass (MIT CSAIL / MIT-IBM Watson AI Lab)
  • Source: https://arxiv.org/abs/2305.10790
  • Category: Multi-modal, Question Answering, Audio-Language Models
  • Key Points

  • The paper introduces LTU (Listen, Think, Understand), a model designed to bridge low-level audio perception (listening) with higher-level reasoning and open-ended question answering (thinking/understanding) over audio.
  • It contributes the OpenAQA dataset, a large-scale collection of open-ended audio question-answering pairs built to train models on both perception and reasoning tasks.
  • Motivation: conventional audio models handle recognition tasks (speech transcription, sound event classification) but lack the ability to reason about or discuss audio content in natural language the way vision-language models do for images.
  • The work aligns with the broader trend of connecting pre-trained audio encoders to large language models, similar to how vision encoders are coupled with LLMs for multimodal chat.
  • Related Entries

  • Clotho-AQA: A Crowdsourced Dataset for Audio Question Answering
  • Cross-Modal Retrieval: A Systematic Review of Methods and Future Directions
  • Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models
  • Notes

    Detailed quantitative results should be verified against the original PDF at arXiv:2305.10790. Readers following the retrieval-to-generation pipeline may find this paper relevant to audio-grounded RAG and multimodal QA systems.

    References

  • Original paper: Listen, Think, and Understand. arXiv:2305.10790.

Tags

#audio-language-models#question-answering#multimodal#datasets#audio-understanding#ltu#openaqa#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208767