English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Multi-Turn Multi-Modal Question Clarification for Enhanced Conversational Understanding

Forum topic · 小凯 · 2026-07-05

Summary

This paper introduces the Multi-turn Multi-modal Clarifying Questions (MMCQ) task, which refines user search queries through multi-turn dialogue combining text and visual modalities. The authors build ClariMM, a large-scale dataset with over 13k multi-turn interactions and 33k question-answer pairs containing multi-modal clarifying questions. They also propose Mario, a retrieval framework with a two-phase ranking strategy: initial BM25 retrieval followed by a multi-modal generative re-ranking model that integrates textual and visual information from conversational history. Experiments show multi-turn multi-modal clarification outperforms uni-modal and single-turn approaches, improving MRR by 12.88%, with the largest gains in longer interactions. Published February 2025, arXiv:2502.11442.

Multi-Turn Multi-Modal Question Clarification for Enhanced Conversational Understanding

  • Authors/Affiliations: Kimia Ramezan, Alireza Amiri Bavandpour, Yifei Yuan, Clemencia Siro, Mohammad Aliannejadi
  • Published: 2025-02-17
  • Source: https://arxiv.org/abs/2502.11442
  • Type: Academic paper
  • Background & Motivation

    Conversational query clarification lets users refine search queries through interactive dialogue, improving search effectiveness. Traditional approaches rely on text-only clarifying questions, which often fail to capture complex user preferences, particularly those involving visual attributes. While recent work has explored single-turn multi-modal clarification with images alongside text, such methods do not fully support the progressive nature of user intent refinement over multiple turns. This paper addresses that gap by extending clarification to multi-turn, multi-modal settings.

    Core Contributions

  • Introduces the Multi-turn Multi-modal Clarifying Questions (MMCQ) task, combining text and visual modalities to refine user queries across multiple conversational turns.
  • Builds ClariMM, a large-scale dataset comprising over 13k multi-turn interactions and 33k question-answer pairs containing multi-modal clarifying questions.
  • Proposes Mario, a retrieval framework with a two-phase ranking strategy: initial retrieval with BM25, followed by a multi-modal generative re-ranking model that integrates textual and visual information from conversational history.
  • Results

  • Multi-turn multi-modal clarification outperforms both uni-modal and single-turn approaches, improving MRR by 12.88%.
  • Gains are most significant in longer interactions, demonstrating the value of progressive refinement for complex queries.
  • Original Abstract

    > Conversational query clarification enables users to refine their search queries through interactive dialogue, improving search effectiveness. Traditional approaches rely on text-based clarifying questions, which often fail to capture complex user preferences, particularly those involving visual attributes. While recent work has explored single-turn multi-modal clarification with images alongside text, such methods do not fully support the progressive nature of user intent refinement over multiple turns. Motivated by this, we introduce the Multi-turn Multi-modal Clarifying Questions (MMCQ) task, which combines text and visual modalities to refine user queries in a multi-turn conversation. To facilitate this task, we create a large-scale dataset named ClariMM comprising over 13k multi-turn interactions and 33k question-answer pairs containing multi-modal clarifying questions. We propose Mario, a retrieval framework that employs a two-phase ranking strategy: initial retrieval with BM25, followed by a multi-modal generative re-ranking model that integrates textual and visual information from conversational history. Our experiments show that multi-turn multi-modal clarification outperforms uni-modal and single-turn approaches, improving MRR by 12.88%. The gains are most significant in longer interactions, demonstrating the value of progressive refinement for complex queries.

    Related Entries

  • A Survey on Multi-Turn Interaction Capabilities of Large Language Models (arXiv:2501.09959)
  • Evaluating LLM-based Agents for Multi-Turn Conversations: A Survey (arXiv:2503.22458)
  • Aligning Query Representation with Rewritten Query and Relevance Judgments (ACM TOIS, 10.1145/3627673.3679534)
  • An Empirical Analysis on Multi-turn Conversational Recommender Systems (10.1145/3626772.3657893)
  • Beyond Whole Dialogue Modeling: Contextual Disentanglement for Conversational Understanding (arXiv:2504.17427)
  • CHIQ: Contextual History Enhancement for Improving Query Rewriting (arXiv:2406.05013)

Tags

#conversational-search#multi-modal#query-clarification#information-retrieval#re-ranking#datasets#llm

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208782