MedExpQA: Multilingual Benchmarking of Large Language Models for Medical Question Answering
Overview
MedExpQA: Multilingual benchmarking of Large Language Models for Medical Question Answering is a research paper published in September 2024 in the journal *Artificial Intelligence in Medicine* (Elsevier).
- Source: ScienceDirect article page
- Venue: Artificial Intelligence in Medicine, September 2024
- Topic: Vertical-domain AI — medical question answering with LLMs
- Provides a standardized, reproducible protocol for cross-lingual medical QA evaluation.
- Enables comparison of LLM performance in medical languages beyond English, highlighting gaps in low-resource settings.
- Complements related cross-lingual evaluation efforts (e.g., "Better to Ask in English?" studies) indexed in this collection.
- An interpretable ensemble of graph and language models
- Better to Ask in English: Cross-Lingual Evaluation of Large Language Models
What the paper addresses
The paper introduces a multilingual benchmark for medical question answering, motivated by the observation that most evaluations of large language models on medical knowledge are English-centric. Since clinical practice and medical education take place in many languages worldwide, multilingual evaluation is essential to understand whether LLMs generalize their medical competence across languages.
Significance for the community
Notes
The listing above is based on the paper's public metadata. For quantitative results, benchmark statistics, and model rankings, please refer to the original article via the ScienceDirect link.