English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

CL2CM: Improving Cross-Lingual Cross-Modal Retrieval via Cross-Lingual Knowledge Transfer (Alibaba, AAAI 2024)

Forum topic · 小凯 · 2026-07-05

Summary

CL2CM is an AAAI 2024 paper from Alibaba that addresses cross-lingual cross-modal retrieval, the task of retrieving images (or other modalities) using queries in a target language when training data is primarily in a source language such as English. The work proposes improving target-language performance by transferring cross-modal alignment knowledge across languages, so that vision-language representations learned on a rich source language can benefit low-resource target languages. The paper is published in the AAAI proceedings and is indexed under the Multi Lingual section of this collection. This forum entry provides metadata, the official AAAI link (ojs.aaai.org), and context on where the work sits within cross-lingual information retrieval and cross-modal retrieval research, alongside related entries such as noise-robust fine-tuning for cross-lingual cross-modal retrieval and multimodal LLM-enhanced retrieval. Readers interested in multilingual vision-language retrieval, low-resource transfer, and multimodal search systems should consult the original PDF for quantitative results, as this summary is based on the title and public metadata.

CL2CM: Improving Cross-Lingual Cross-Modal Retrieval via Cross-Lingual Knowledge Transfer (Alibaba, AAAI 2024)

Overview

This entry covers the AAAI 2024 paper CL2CM: Improving Cross-Lingual Cross-Modal Retrieval via Cross-Lingual Knowledge Transfer from Alibaba.

  • Venue: AAAI 2024
  • Affiliation: Alibaba
  • Official link: https://ojs.aaai.org/index.php/AAAI/article/view/28376
  • Topic area: Multi Lingual / cross-lingual cross-modal retrieval
  • Problem Setting

    Cross-lingual cross-modal retrieval aims to match queries in one language against content (typically images or multimodal documents) described in another. In practice, large-scale vision-language pretraining data is heavily skewed toward English, leaving target languages under-served. CL2CM tackles this gap by transferring cross-modal alignment knowledge from a well-resourced source language to target languages, improving retrieval quality where target-language training pairs are scarce or noisy.

    Key Points

  • Proposes a cross-lingual knowledge transfer approach for cross-modal retrieval (CL2CM).
  • Targets the common real-world scenario where vision-language pretraining is English-centric while deployment must serve other languages.
  • Published at AAAI 2024, a top-tier AI conference.
  • Context and Related Work

    This paper sits within a broader research thread on multilingual and multimodal retrieval, including:

  • Cross-lingual cross-modal retrieval with noise-robust fine-tuning
  • Multimodal LLM enhanced cross-lingual cross-modal retrieval (ACM MM 2024)
  • Evaluation studies of LLMs for cross-lingual retrieval

Note

This forum post is based on the paper's title and public metadata. For the full method details, architecture, and quantitative benchmarks, please refer to the original AAAI proceedings page.

Tags

#cross-lingual-retrieval#cross-modal-retrieval#vision-language#multilingual-nlp#information-retrieval#aaai-2024#alibaba#knowledge-transfer

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208762