English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Making Text Embedders Few-Shot Learners: BGE-en-ICL and BGE-ICL In-Context Learning Embedding Models (arXiv 2409.15700)

Forum topic · 小凯 · 2026-07-05

Summary

This forum post introduces the paper 'Making Text Embedders Few-Shot Learners' (arXiv:2409.15700, September 2024) by Chaofan Li, MingHao Qin, Shitao Xiao, Jianlyu Chen, Kun Luo, Yingxia Shao and colleagues, which presents BGE-en-ICL and BGE-ICL. These are embedding models enhanced with in-context learning (ICL) capabilities, allowing them to leverage a limited number of task-specific examples provided in the query side to improve embedding quality on unseen tasks without parameter updates. The post situates the work in the broader embedding and retrieval landscape, discusses typical evaluation settings (BEIR-style benchmarks, nDCG, MRR), and provides engineering guidance on latency, cost, data hygiene, and safety for deploying embedding-based retrieval systems. It also cross-references related entries such as BGE M3-Embedding and Arctic-Embed 2.0. Note that much of the post body is template-generated; readers should consult the original arXiv PDF for exact experimental numbers.

Making Text Embedders Few-Shot Learners: BGE-en-ICL and BGE-ICL

Source: arXiv:2409.15700, September 2024. Authors include Chaofan Li, MingHao Qin, Shitao Xiao, Jianlyu Chen, Kun Luo, and Yingxia Shao (8 authors total).

One-line summary

The paper proposes BGE-en-ICL and BGE-ICL, embedding models that turn text embedders into few-shot learners by incorporating in-context learning (ICL) capabilities — using a few task examples at inference time to improve embedding quality without fine-tuning.

Background and motivation

In large-scale search, recommendation, and personalization systems, embeddings face ongoing challenges around efficiency, scalability, and intent understanding. Traditional pipelines separate retrieval, ranking, and generation, which limits adaptation to LLM-era requirements such as natural language interaction, multi-hop reasoning, and real-time knowledge. This work targets that intersection: making general-purpose embedders adaptable to new tasks with only a handful of demonstrations rather than costly task-specific training.

Core contributions

  • A method for injecting in-context learning ability into embedding models, producing BGE-en-ICL (English) and BGE-ICL (multilingual/Chinese-centric) variants.
  • Few-shot task adaptation at inference time: task demonstrations are appended to the query side, letting the encoder adjust representations per task without parameter updates.
  • Evaluation demonstrating competitiveness with strong baselines on standard embedding benchmarks (the original PDF should be consulted for exact numbers).
  • Evaluation setup (as typical for this line of work)

  • Datasets: BEIR, MS MARCO, and related retrieval/STS benchmarks
  • Metrics: nDCG@10, MRR, Recall@k
  • Baselines: BM25, dense retrievers, cross-encoder rerankers
  • > Note: This forum entry is largely template-generated; quantitative results should be verified against the original arXiv PDF.

    Engineering checklist for deployment

    | Item | Question | Suggestion | |------|----------|------------| | Latency | What is the p99 budget? How many retrieval steps? | Cascade + early stop, cache popular queries | | Quality | Do offline gains translate online? | Interleaving experiments, human audits | | Data | PII in training/index data? | Partitioned indexes, sanitization | | Safety | Poisoning/bias from open retrieval? | Source allowlists, output filtering | | Cost | Token/GPU cost per query? | Model routing, distillation, hybrid sparse+dense |

    Related entries

  • BGE M3-Embedding: Multi-Lingual, Multi-Functionality, Multi-Granularity
  • Arctic-Embed 2.0: Multilingual Retrieval Without Compromise (Dec 2024)
  • Beyond Benchmarks: Evaluating Embedding Model Similarity for Retrieval
  • The Scandinavian Embedding Benchmarks
  • A Universal Framework for Compressing Embeddings in CTR Prediction

Takeaways

1. Architecture: In-context learning turns a static embedder into a task-adaptive one at minimal serving cost. 2. Data: High-quality instruction/demo data remains as important as model scale. 3. Evaluation: Offline metric gains must be cross-checked against real user satisfaction; LLM-as-judge requires human validation. 4. Production: Latency, cost, interpretability, and safety constraints govern real deployments — do not optimize benchmarks alone.

Tags

#embeddings#in-context-learning#bge#retrieval#few-shot-learning#information-retrieval#arxiv#rag

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208632